nvidia
/

Cosmos3-Super-Image2Video

@@ -10,6 +10,8 @@ tags:
   - cosmos3
   - vllm-omni
   - diffusers
   - image-to-video
   - video-generation
 countDownloads:
@@ -211,6 +213,7 @@ Our AI models are designed and/or optimized to run on NVIDIA GPU-accelerated sys
 - [PyTorch](https://github.com/nvidia/cosmos3)
 - [vLLM-Omni](https://github.com/vllm-project/vllm-omni)
 - [Hugging Face Diffusers](https://huggingface.co/docs/diffusers/en/index)
 **Supported Hardware Microarchitecture Compatibility:**
@@ -527,6 +530,12 @@ Example output generated by Diffusers:
 <video controls width="832" height="480" src="https://huggingface.co/nvidia/Cosmos3-Super-Image2Video/resolve/main/assets/example_output_diffusers.mp4"></video>
 ## Limitations
 Cosmos3 may produce imperfect outputs in challenging scenarios. Generation artifacts include temporal inconsistency, unstable camera or object motion, imprecise physical interactions, inaccurate audio-video synchronization, and action-state drift — especially in long-horizon or high-resolution outputs. Reasoning may also be incorrect: object states, causal relationships, spatial geometry, temporal ordering, agent intent, and future outcomes can be misinferred, and complex or long-context inputs may yield hallucinated entities, inconsistent interpretations, or implausible predictions. Because the model lacks an explicit physics simulator, 3D geometry, 4D space-time evolution, object permanence, contact dynamics, and physical laws are only approximated — producing artifacts such as disappearing or morphing objects, unrealistic collisions, and physically implausible motions. Quality further degrades in out-of-distribution environments, safety-critical edge cases, and domains underrepresented in training.
@@ -535,7 +544,7 @@ Cosmos3 outputs should not be treated as physically accurate simulation, reliabl
 ## Inference
-**Acceleration Engine:** [PyTorch](https://pytorch.org/), [vLLM](https://github.com/vllm-project/vllm), [vLLM-Omni](https://github.com/vllm-project/vllm-omni), [Hugging Face Diffusers](https://github.com/huggingface/diffusers)
 **Test Hardware:** GB200 and H100

   - cosmos3
   - vllm-omni
   - diffusers
+  - sglang
+  - sglang-diffusion
   - image-to-video
   - video-generation
 countDownloads:
 - [PyTorch](https://github.com/nvidia/cosmos3)
 - [vLLM-Omni](https://github.com/vllm-project/vllm-omni)
 - [Hugging Face Diffusers](https://huggingface.co/docs/diffusers/en/index)
+- [SGLang](https://sgl-project.github.io/)
 **Supported Hardware Microarchitecture Compatibility:**
 <video controls width="832" height="480" src="https://huggingface.co/nvidia/Cosmos3-Super-Image2Video/resolve/main/assets/example_output_diffusers.mp4"></video>
+### SGLang
+[SGLang Diffusion](https://sgl-project.github.io/diffusion) can serve `nvidia/Cosmos3-Super-Image2Video` through OpenAI-compatible video generation endpoints.
+For complete serving instructions and request examples, see the [Cosmos3 SGLang cookbook](https://lmsysorg.mintlify.app/cookbook/diffusion/Cosmos/Cosmos3).
 ## Limitations
 Cosmos3 may produce imperfect outputs in challenging scenarios. Generation artifacts include temporal inconsistency, unstable camera or object motion, imprecise physical interactions, inaccurate audio-video synchronization, and action-state drift — especially in long-horizon or high-resolution outputs. Reasoning may also be incorrect: object states, causal relationships, spatial geometry, temporal ordering, agent intent, and future outcomes can be misinferred, and complex or long-context inputs may yield hallucinated entities, inconsistent interpretations, or implausible predictions. Because the model lacks an explicit physics simulator, 3D geometry, 4D space-time evolution, object permanence, contact dynamics, and physical laws are only approximated — producing artifacts such as disappearing or morphing objects, unrealistic collisions, and physically implausible motions. Quality further degrades in out-of-distribution environments, safety-critical edge cases, and domains underrepresented in training.
 ## Inference
+**Acceleration Engine:** [PyTorch](https://pytorch.org/), [vLLM](https://github.com/vllm-project/vllm), [vLLM-Omni](https://github.com/vllm-project/vllm-omni), [Hugging Face Diffusers](https://github.com/huggingface/diffusers), [SGLang](https://sgl-project.github.io/), [SGLang Diffusion](https://sgl-project.github.io/diffusion)
 **Test Hardware:** GB200 and H100