[HF team] add pipeline tag to readme or better discoverability

#15
by linoyts HF Staff - opened
README.md CHANGED
@@ -1,20 +1,18 @@
1
  ---
2
  license: other
3
  license_name: openmdw1.1-license
4
- license_link: >-
5
- https://openmdw.ai/license/1-1/
6
  library_name: cosmos
7
  tags:
8
- - nvidia
9
- - cosmos
10
- - cosmos3
11
- - vllm
12
- - vllm-omni
13
- - sglang
14
- - sglang-diffusion
15
- - diffusers
16
- - text, image, video, audio, and action generation
17
- - omnimodel
18
  ---
19
 
20
  # **Cosmos 3: Omnimodal World Models for Physical AI**
@@ -169,7 +167,6 @@ Our AI models are designed and/or optimized to run on NVIDIA GPU-accelerated sys
169
  - [PyTorch](https://github.com/nvidia/cosmos3)
170
  - [vLLM-Omni](https://github.com/vllm-project/vllm-omni)
171
  - [Hugging Face Diffusers](https://huggingface.co/docs/diffusers/en/index)
172
- - [SGLang](https://github.com/sgl-project/sglang)
173
 
174
  **Supported Hardware Microarchitecture Compatibility:**
175
 
@@ -200,8 +197,6 @@ Raw data from internal and external sources is transformed into training-ready d
200
 
201
  Training datasets passed through multiple layers of automated and manual safeguards designed to reduce the presence of harmful or policy-violating content across categories including weapons and weapons-related instructional content, criminal planning, child sexual abuse material (CSAM), non-consensual intimate imagery (NCII), sexual content involving minors, harassment, hate speech, profanity, threats and incitement to violence, self-harm or suicide-related content, and graphic violence. Data sources are reviewed for licensing compatibility, provenance, and alignment with internal data governance and safety policies before admission into training corpora. Automated filtering pipelines combine multiple detection strategies: hash-matching against known CSAM and NCII reference databases; classifier-based moderation models trained for explicit sexual content, hate speech, violence, weapons imagery, and other restricted categories; keyword and regex-based screening for criminal-planning, threats, and self-harm phrases in text data; metadata and provenance heuristics for source-level risk signals; and embedding-based anomaly detection to surface samples that fall outside expected distributions. Human review and targeted audits supplement automated filtering for selected datasets, benchmark construction, and safety-sensitive evaluation. For multimodal Physical AI data (robotics, autonomous driving, industrial scenes), additional filtering targets invalid action trajectories, physically implausible interactions, and unsafe control sequences. Synthetic and simulation-generated data are evaluated through internal validation before inclusion. Benchmark evaluations and red-team testing are applied post-training to surface remaining safety gaps across world generation, reasoning, audio, and action tasks. No large-scale data-filtering process can guarantee complete removal of all harmful content; residual risks may remain, particularly in rare edge cases or open-world deployment settings. Ongoing monitoring and dataset review continue post-release.
202
 
203
- - For more information about the datasets used to train this model, please see the [Public Summary of Training Content](https://docs.nvidia.com/cosmos/latest/_downloads/e482b7114ce8dbfbb07d2d4b42cafe4e/training-content-Cosmos-3.pdf).
204
-
205
  **Data Modality and Training Data Size**
206
 
207
  | Modality | Reasoning Data Sample Count | Generation Data Sample Count |
@@ -921,69 +916,6 @@ Example output:
921
 
922
  <video controls width="1280" height="720" src="https://huggingface.co/nvidia/Cosmos3-Nano/resolve/main/assets/example_t2v_diffusers_output.mp4"></video>
923
 
924
- ### SGLang
925
-
926
- [SGLang Diffusion](https://docs.sglang.io/docs/sglang-diffusion/index) can serve `nvidia/Cosmos3-Nano` through OpenAI-compatible image and video generation endpoints. Install SGLang from the main branch with diffusion dependencies, then start a server:
927
-
928
- ```shell
929
- git clone --branch main https://github.com/sgl-project/sglang.git
930
- cd sglang
931
- pip install -e "python[diffusion]"
932
- pip install "cosmos-guardrail==0.3.1"
933
-
934
- sglang serve --model-path nvidia/Cosmos3-Nano
935
- ```
936
-
937
- Cosmos 3 support in SGLang Diffusion currently requires the SGLang main branch. Switch to a stable SGLang release once Cosmos 3 support is included there.
938
-
939
- For a video-specialized checkpoint, use `Cosmos3-Super-Image2Video` with multiple GPUs:
940
-
941
- ```shell
942
- sglang serve \
943
- --model-path nvidia/Cosmos3-Super-Image2Video \
944
- --num-gpus 4
945
- ```
946
-
947
- Supported SGLang endpoints:
948
-
949
- | Mode | Endpoint | Notes |
950
- | --- | --- | --- |
951
- | Text to image | `POST /v1/images/generations` | Returns base64 image data by default |
952
- | Text to video | `POST /v1/videos` | Creates an async job; poll `GET /v1/videos/{id}` and download `/content` |
953
- | Image to video | `POST /v1/videos` | Upload the conditioning image with `input_reference` |
954
-
955
- Example text-to-video request:
956
-
957
- ```shell
958
- job_id=$(curl -sS -X POST http://localhost:30000/v1/videos \
959
- --form-string "prompt=A small warehouse robot moves a blue box across a clean floor." \
960
- --form-string "negative_prompt=blurry, distorted, low quality" \
961
- --form-string "size=1280x720" \
962
- --form-string "num_frames=81" \
963
- --form-string "fps=24" \
964
- --form-string "num_inference_steps=35" \
965
- --form-string "guidance_scale=4.0" \
966
- --form-string "flow_shift=10.0" \
967
- --form-string "seed=42" \
968
- --form-string 'extra_params={"guardrails":true,"use_resolution_template":false,"use_duration_template":false}' \
969
- | python -c 'import json, sys; print(json.load(sys.stdin)["id"])')
970
-
971
- while true; do
972
- status=$(curl -sS "http://localhost:30000/v1/videos/${job_id}" \
973
- | python -c 'import json, sys; print(json.load(sys.stdin)["status"])')
974
- [ "$status" = "completed" ] && break
975
- [ "$status" = "failed" ] && exit 1
976
- sleep 1
977
- done
978
-
979
- curl -sS -L "http://localhost:30000/v1/videos/${job_id}/content" \
980
- -o cosmos3_t2v_output.mp4
981
- ```
982
-
983
- SGLang accepts Cosmos 3 request options including `max_sequence_length`, `flow_shift`, `extra_params.guardrails`, `extra_params.use_resolution_template`, and `extra_params.use_duration_template`. Video-to-video, video-with-sound, and action generation are not supported by SGLang yet.
984
-
985
- For complete serving instructions and request examples, see the [Cosmos3 SGLang cookbook](https://docs.sglang.io/cookbook/diffusion/Cosmos/Cosmos3).
986
-
987
  ## Limitations
988
 
989
  Cosmos3 may produce imperfect outputs in challenging scenarios. Generation artifacts include temporal inconsistency, unstable camera or object motion, imprecise physical interactions, inaccurate audio-video synchronization, and action-state drift — especially in long-horizon or high-resolution outputs. Reasoning may also be incorrect: object states, causal relationships, spatial geometry, temporal ordering, agent intent, and future outcomes can be misinferred, and complex or long-context inputs may yield hallucinated entities, inconsistent interpretations, or implausible predictions. Because the model lacks an explicit physics simulator, 3D geometry, 4D space-time evolution, object permanence, contact dynamics, and physical laws are only approximated — producing artifacts such as disappearing or morphing objects, unrealistic collisions, and physically implausible motions. Quality further degrades in out-of-distribution environments, safety-critical edge cases, and domains underrepresented in training.
@@ -992,7 +924,7 @@ Cosmos3 outputs should not be treated as physically accurate simulation, reliabl
992
 
993
  ## Inference
994
 
995
- **Acceleration Engine:** [PyTorch](https://pytorch.org/), [vLLM](https://github.com/vllm-project/vllm), [vLLM-Omni](https://github.com/vllm-project/vllm-omni), [Hugging Face Diffusers](https://github.com/huggingface/diffusers), [SGLang](https://github.com/sgl-project/sglang), [SGLang Diffusion](https://docs.sglang.io/docs/sglang-diffusion/index)
996
 
997
  **Test Hardware:** GB200 and H100
998
 
@@ -1004,4 +936,4 @@ Please make sure you have proper rights and permissions for all input image and
1004
 
1005
  Users are responsible for model inputs and outputs. Users are responsible for ensuring safe integration of this model, including implementing guardrails as well as other safety mechanisms, prior to deployment.
1006
 
1007
- For more detailed information on ethical considerations for this model, please see the Model Card++ [Explainability](EXPLAINABILITY.md), [Bias](BIAS.md), [Safety & Security](SAFETY.md), and [Privacy](PRIVACY.md) subcards. Please report model quality, risk, security vulnerabilities or NVIDIA AI Concerns [here](https://www.nvidia.com/en-us/support/submit-security-vulnerability/).
 
1
  ---
2
  license: other
3
  license_name: openmdw1.1-license
4
+ license_link: https://openmdw.ai/license/1-1/
 
5
  library_name: cosmos
6
  tags:
7
+ - nvidia
8
+ - cosmos
9
+ - cosmos3
10
+ - vllm
11
+ - vllm-omni
12
+ - diffusers
13
+ - text, image, video, audio, and action generation
14
+ - omnimodel
15
+ pipeline_tag: any-to-any
 
16
  ---
17
 
18
  # **Cosmos 3: Omnimodal World Models for Physical AI**
 
167
  - [PyTorch](https://github.com/nvidia/cosmos3)
168
  - [vLLM-Omni](https://github.com/vllm-project/vllm-omni)
169
  - [Hugging Face Diffusers](https://huggingface.co/docs/diffusers/en/index)
 
170
 
171
  **Supported Hardware Microarchitecture Compatibility:**
172
 
 
197
 
198
  Training datasets passed through multiple layers of automated and manual safeguards designed to reduce the presence of harmful or policy-violating content across categories including weapons and weapons-related instructional content, criminal planning, child sexual abuse material (CSAM), non-consensual intimate imagery (NCII), sexual content involving minors, harassment, hate speech, profanity, threats and incitement to violence, self-harm or suicide-related content, and graphic violence. Data sources are reviewed for licensing compatibility, provenance, and alignment with internal data governance and safety policies before admission into training corpora. Automated filtering pipelines combine multiple detection strategies: hash-matching against known CSAM and NCII reference databases; classifier-based moderation models trained for explicit sexual content, hate speech, violence, weapons imagery, and other restricted categories; keyword and regex-based screening for criminal-planning, threats, and self-harm phrases in text data; metadata and provenance heuristics for source-level risk signals; and embedding-based anomaly detection to surface samples that fall outside expected distributions. Human review and targeted audits supplement automated filtering for selected datasets, benchmark construction, and safety-sensitive evaluation. For multimodal Physical AI data (robotics, autonomous driving, industrial scenes), additional filtering targets invalid action trajectories, physically implausible interactions, and unsafe control sequences. Synthetic and simulation-generated data are evaluated through internal validation before inclusion. Benchmark evaluations and red-team testing are applied post-training to surface remaining safety gaps across world generation, reasoning, audio, and action tasks. No large-scale data-filtering process can guarantee complete removal of all harmful content; residual risks may remain, particularly in rare edge cases or open-world deployment settings. Ongoing monitoring and dataset review continue post-release.
199
 
 
 
200
  **Data Modality and Training Data Size**
201
 
202
  | Modality | Reasoning Data Sample Count | Generation Data Sample Count |
 
916
 
917
  <video controls width="1280" height="720" src="https://huggingface.co/nvidia/Cosmos3-Nano/resolve/main/assets/example_t2v_diffusers_output.mp4"></video>
918
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
919
  ## Limitations
920
 
921
  Cosmos3 may produce imperfect outputs in challenging scenarios. Generation artifacts include temporal inconsistency, unstable camera or object motion, imprecise physical interactions, inaccurate audio-video synchronization, and action-state drift — especially in long-horizon or high-resolution outputs. Reasoning may also be incorrect: object states, causal relationships, spatial geometry, temporal ordering, agent intent, and future outcomes can be misinferred, and complex or long-context inputs may yield hallucinated entities, inconsistent interpretations, or implausible predictions. Because the model lacks an explicit physics simulator, 3D geometry, 4D space-time evolution, object permanence, contact dynamics, and physical laws are only approximated — producing artifacts such as disappearing or morphing objects, unrealistic collisions, and physically implausible motions. Quality further degrades in out-of-distribution environments, safety-critical edge cases, and domains underrepresented in training.
 
924
 
925
  ## Inference
926
 
927
+ **Acceleration Engine:** [PyTorch](https://pytorch.org/), [vLLM](https://github.com/vllm-project/vllm), [vLLM-Omni](https://github.com/vllm-project/vllm-omni), [Hugging Face Diffusers](https://github.com/huggingface/diffusers)
928
 
929
  **Test Hardware:** GB200 and H100
930
 
 
936
 
937
  Users are responsible for model inputs and outputs. Users are responsible for ensuring safe integration of this model, including implementing guardrails as well as other safety mechanisms, prior to deployment.
938
 
939
+ For more detailed information on ethical considerations for this model, please see the Model Card++ [Explainability](EXPLAINABILITY.md), [Bias](BIAS.md), [Safety & Security](SAFETY.md), and [Privacy](PRIVACY.md) subcards. Please report model quality, risk, security vulnerabilities or NVIDIA AI Concerns [here](https://www.nvidia.com/en-us/support/submit-security-vulnerability/).
modular_model_index.json DELETED
@@ -1,75 +0,0 @@
1
- {
2
- "_blocks_class_name": "Cosmos3OmniBlocks",
3
- "_class_name": "Cosmos3OmniModularPipeline",
4
- "_diffusers_version": "0.39.0.dev0",
5
- "text_tokenizer": [
6
- "transformers",
7
- "Qwen2TokenizerFast",
8
- {
9
- "pretrained_model_name_or_path": "nvidia/Cosmos3-Nano",
10
- "revision": null,
11
- "subfolder": "text_tokenizer",
12
- "type_hint": [
13
- "transformers",
14
- "Qwen2TokenizerFast"
15
- ],
16
- "variant": null
17
- }
18
- ],
19
- "vae": [
20
- "diffusers",
21
- "AutoencoderKLWan",
22
- {
23
- "pretrained_model_name_or_path": "nvidia/Cosmos3-Nano",
24
- "revision": null,
25
- "subfolder": "vae",
26
- "type_hint": [
27
- "diffusers",
28
- "AutoencoderKLWan"
29
- ],
30
- "variant": null
31
- }
32
- ],
33
- "transformer": [
34
- "diffusers",
35
- "Cosmos3OmniTransformer",
36
- {
37
- "pretrained_model_name_or_path": "nvidia/Cosmos3-Nano",
38
- "revision": null,
39
- "subfolder": "transformer",
40
- "type_hint": [
41
- "diffusers",
42
- "Cosmos3OmniTransformer"
43
- ],
44
- "variant": null
45
- }
46
- ],
47
- "scheduler": [
48
- "diffusers",
49
- "UniPCMultistepScheduler",
50
- {
51
- "pretrained_model_name_or_path": "nvidia/Cosmos3-Nano",
52
- "revision": null,
53
- "subfolder": "scheduler",
54
- "type_hint": [
55
- "diffusers",
56
- "UniPCMultistepScheduler"
57
- ],
58
- "variant": null
59
- }
60
- ],
61
- "sound_tokenizer": [
62
- "diffusers",
63
- "Cosmos3AVAEAudioTokenizer",
64
- {
65
- "pretrained_model_name_or_path": "nvidia/Cosmos3-Nano",
66
- "revision": null,
67
- "subfolder": "sound_tokenizer",
68
- "type_hint": [
69
- "diffusers",
70
- "Cosmos3AVAEAudioTokenizer"
71
- ],
72
- "variant": null
73
- }
74
- ]
75
- }
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
sound_tokenizer/diffusion_pytorch_model.safetensors CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:a4b6da5975bf89f6853fe589ad4752d281ac79fbdfad52ea90537fa080b4b9c2
3
- size 1985176840
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:9d4c61cde38acfb0cad9048a140c3533750277a8462b19dc08450d9fe1ad9879
3
+ size 1892409600