Setup Guide
System Requirements
- NVIDIA GPUs with Ampere architecture (RTX 30 Series, A100) or newer
- NVIDIA driver compatible with CUDA 12.6
- Linux x86-64
- glibc>=2.31 (e.g Ubuntu >=22.04)
- Python 3.10
Installation
Clone the repository
git clone git@github.com:nvidia-cosmos/cosmos-predict2.git
cd cosmos-predict2
Option 1: Virtual environment
Install system dependencies:
curl -LsSf https://astral.sh/uv/install.sh | sh
source $HOME/.local/bin/env
Install the package into a new environment:
uv sync --extra cu126
source .venv/bin/activate
Or, install the package into the active environment (e.g. conda):
uv sync --extra cu126 --active --inexact
Option 2: Docker container
Please make sure you have access to Docker on your machine and the NVIDIA Container Toolkit is installed.
Build and run the container:
docker run --gpus all --rm -v .:/workspace -v /workspace/.venv -it $(docker build -q .)
Downloading Checkpoints
- Get a Hugging Face Access Token with
Readpermission - Install Hugging Face CLI:
uv tool install -U "huggingface_hub[cli]" - Login:
hf auth login - Accept the Llama-Guard-3-8B terms.
To download a specific model:
./scripts/download_checkpoints.py --model_types <model_type> --model_sizes <model_size>
| Models | Link | Download Arguments | Notes |
|---|---|---|---|
| Cosmos-Predict2-0.6B-Text2Image | 🤗 Huggingface | --model_types text2image --model_sizes 0.6B |
N/A |
| Cosmos-Predict2-2B-Text2Image | 🤗 Huggingface | --model_types text2image --model_sizes 2B |
N/A |
| Cosmos-Predict2-14B-Text2Image | 🤗 Huggingface | --model_types text2image --model_sizes 14B |
N/A |
| Cosmos-Predict2-2B-Video2World | 🤗 Huggingface | --model_types video2world --model_sizes 2B |
Download 720P, 16FPS by default. Supports 480P and 720P resolution. Supports 10FPS and 16FPS |
| Cosmos-Predict2-14B-Video2World | 🤗 Huggingface | --model_types video2world --model_sizes 14B |
Download 720P, 16FPS by default. Supports 480P and 720P resolution. Supports 10FPS and 16FPS |
| Cosmos-Predict2-2B-Sample-Action-Conditioned | 🤗 Huggingface | --model_types sample_action_conditioned |
Supports 480P and 4FPS. |
| Cosmos-Predict2-14B-Sample-GR00T-Dreams-GR1 | 🤗 Huggingface | --model_types sample_gr00t_dreams_gr1 |
Supports 480P and 16FPS. |
| Cosmos-Predict2-14B-Sample-GR00T-Dreams-DROID | 🤗 Huggingface | --model_types sample_gr00t_dreams_droid |
Supports 480P and 16FPS. |
To download all the checkpoints (requires ~250GB of disk space), run:
./scripts/download_checkpoints.py
To see the full list of options, run:
./scripts/download_checkpoints.py --help
Troubleshooting
CUDA/GPU Issues
CUDA driver version insufficient: Update NVIDIA drivers to latest version compatible with CUDA 12.6+
Out of Memory (OOM) errors: Use 2B models instead of 14B, or reduce batch size/resolution
Missing CUDA libraries: Set paths with
export CUDA_HOME=$CONDA_PREFIX
Installation Issues
Conda environment conflicts: Create fresh environment with
conda create -n cosmos-predict2-clean python=3.10 -yFlash-attention build failures: Install build tools with
apt-get install build-essentialTransformer engine linking errors: Reinstall with
pip install --force-reinstall transformer-engine==1.12.0
For other issues, check GitHub Issues.