--- title: NVIDIA Dynamo colorFrom: blue colorTo: green sdk: static pinned: false --- # NVIDIA Dynamo NVIDIA Dynamo is an open-source, low-latency, modular inference framework for serving generative AI models in distributed environments. It scales inference workloads across GPU fleets with intelligent resource scheduling and request routing, optimized memory management, and data transfer. ## Get Started - [NVIDIA Dynamo repository](https://github.com/ai-dynamo/dynamo) - [AI Dynamo on GitHub](https://github.com/ai-dynamo) - [NVIDIA Dynamo documentation](https://docs.nvidia.com/dynamo/latest/) ## What It Provides - Distributed and disaggregated inference serving - Support for SGLang, TensorRT-LLM, and vLLM backends - Tools, examples, evaluation datasets, and reproducible parser fixtures ## Datasets and Evaluation This organization publishes Dynamo-related datasets and evaluation fixtures. Each dataset includes a dataset card that describes its purpose, schema, provenance, license, versioning, and usage. For questions or corrections, open an issue in the relevant [AI Dynamo GitHub repository](https://github.com/ai-dynamo).