StreamRender-H3 / README.md
yitongl's picture
Clarify licenses for the included upstream runtime models
ee01d6c verified
|
Raw History Blame
5.13 kB
metadata
license: other
license_name: minimax-h3-community
license_link: LICENSE
tags:
  - video-generation
  - streaming
  - interactive
  - world-model
  - minimax-h3

StreamRender-H3: Interactive Coding Worlds with Real-time Streaming Diffusion Render

Project page: https://nvlabs.github.io/Sana/Sol-Engine/StreamRender-H3/

StreamRender-H3 splits an interactive world into a Code Game Engine and a Streaming Diffusion Render. The engine keeps the world state in code (maps, physics, rules) and streams a semantic map for every frame. A two-step streaming video model built on MiniMax-H3 Ref2VA renders those maps into photorealistic video as the player drives, with the look set by a reference image and a prompt.

  • Two denoising steps per chunk in the distilled student.
  • Diffusion render of 217.96 ms per streaming step on 8 GB200 GPUs, a 3.67× speedup over the PyTorch baseline.
  • Fast Causal H3 VAE decodes 2 latent frames at 720p in 24 ms on one GB200, built by a self-improving agent loop.

Released runtime assets

This repository contains the complete model assets for the streaming runtime.

Directory Component Weight provenance
backbone/ Full 50-layer streaming H3 renderer, 1008b step224 EMA, LoRA merged Trained renderer
causal_decoder/ Fast Causal H3 VAE, distilled ViT24 decoder step6000 Trained decoder
qwen_model/, qwen_dcp/ H3-compatible Qwen conditioning model in safetensors and DCP formats Unchanged upstream model; DCP is a storage conversion
tae/ Official H3 TAE weights for streaming reference encoding Unchanged upstream weights
audio_vae/ H3 audio codec required by the model contract Unchanged upstream weights
tokenizer/, qwen_processor/ Conditioning tokenizer and processor Upstream runtime assets
configs/, assets.json Runtime configuration, environment versions, relative asset manifest Deployment configuration
reference_image/ Default B10_e-track-3_bluehour reference picture Example initialization asset

There are five distinct model components. Qwen's two storage formats are the same model. The causal decoder is constructed directly from its architecture configuration and student checkpoint; no native H3 video VAE checkpoint is needed. Streaming reference encoding uses TAE, and video decoding uses the trained causal decoder.

Download and use

Code and full deployment instructions: NVlabs/Sana — StreamRender-H3.

from huggingface_hub import snapshot_download
snapshot_download("Efficient-Large-Model/StreamRender-H3",
                  local_dir="./StreamRender-H3")

The complete assets are approximately 206.87 GB. Reconstruct the exact merged backbone from its nine safetensors shards (another 66.28 GB of disk):

python StreamRender-H3/download_assets.py
cd Sana/models/streamrender_h3
python scripts/preflight.py --assets /absolute/path/to/StreamRender-H3/assets.json
GPUS=8 PYTHON_BIN=python bash scripts/launch_local.sh \
  --config /absolute/path/to/StreamRender-H3/configs/runtime.json \
  --assets /absolute/path/to/StreamRender-H3/assets.json

Use the documented Python/CUDA environment and install frontend dependencies first. The bundle includes compatible Qwen DCP weights, so no separate DCP conversion or personal model repository is required. The default prompt and reference image are B10_e-track-3_bluehour. Alternative semantic-video examples and their media download instructions are in the runtime's examples/ folder.

The runtime uses eight GPUs in one NVLink domain, two full denoising evaluations, CFG=1 and S=2/W=2/C=2. Qwen conditions on the prompt and first picture once per session. SHA256SUMS records file integrity. Each component retains its applicable upstream license; see LICENSE, LICENSE-Qwen.txt, LICENSE-TAE.txt, and NOTICE.

License

The H3 renderer, distilled causal decoder and H3 audio codec follow the MiniMax H3 Community License, including its territory, use and redistribution terms; see LICENSE and NOTICE. Qwen components retain their Apache-2.0 license (LICENSE-Qwen.txt). TAE retains its upstream MIT license and attribution (LICENSE-TAE.txt). This bundle is not covered by a single blanket license. The runtime code's license does not replace model licenses.

Citation

@misc{streamrenderh3_2026,
  title        = {StreamRender-H3: Interactive Coding Worlds with Real-time Streaming Diffusion Render},
  author       = {Yanzuo Lu and Tian Ye and Yitong Li and Shuchen Xue and Haozhe Liu and Song Han and Enze Xie},
  year         = {2026},
  howpublished = {\url{https://nvlabs.github.io/Sana/Sol-Engine/StreamRender-H3/}},
  note         = {Project page}
}

Acknowledgements

We thank the MiniMax-H3 team for the base video generation model and native video codec. Our streaming renderer and distilled causal decoder build on these components. Their upstream license and attribution notices remain applicable.