Text-to-Video
Diffusers
Safetensors
MiniMax H3
video
audio
text-to-audio-video
distillation
dmd2
few-step
fastvideo
fasth3
Instructions to use FastVideo/FastVideo-Minimax-FastH3-Preview-v0.1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use FastVideo/FastVideo-Minimax-FastH3-Preview-v0.1 with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("FastVideo/FastVideo-Minimax-FastH3-Preview-v0.1", dtype=torch.bfloat16, device_map="cuda") prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k" image = pipe(prompt).images[0] - Notebooks
- Google Colab
- Kaggle
| license: other | |
| license_name: minimax-h3-community | |
| license_link: LICENSE | |
| base_model: MiniMaxAI/MiniMax-H3 | |
| library_name: diffusers | |
| pipeline_tag: text-to-video | |
| tags: | |
| - text-to-video | |
| - video | |
| - audio | |
| - text-to-audio-video | |
| - distillation | |
| - dmd2 | |
| - few-step | |
| - minimax-h3 | |
| - fastvideo | |
| - fasth3 | |
| # FastVideo-Minimax-FastH3-Preview-v0.1 | |
| **A few-step (4-step) distillation preview of [MiniMax-H3](https://huggingface.co/MiniMaxAI/MiniMax-H3)**, the | |
| 33B dual-modality (video + audio) diffusion transformer β distilled with data-free | |
| DMD2 by the [FastVideo](https://github.com/hao-ai-lab/FastVideo) team. | |
| The base model samples with 50 denoising steps; this student walks a 4-step grid | |
| on the release's shift-12 rectified-flow schedule (12.5Γ fewer transformer | |
| evaluations), generating synchronized video and audio in one pipeline call. | |
| > **Preview status (v0.1):** this is an early training checkpoint (step 1400 of a | |
| > 4000-step run) published for evaluation and integration work. Sample quality is | |
| > still maturing; expect a stronger release checkpoint from the same run. | |
| ## What's in the repo | |
| Diffusers-format (modular pipeline) layout. Only the `transformer/` weights differ | |
| from the base release β the distilled student, in bf16. All other components | |
| (Qwen3-VL text encoder, video/audio VAEs, tokenizer, processor, schedulers) are | |
| unmodified copies of the base release, included so the repo is self-contained. | |
| The student was trained with block-sparse video attention (VSA, 64-token tiles, | |
| 90% sparsity) and carries its trained sparse-gate parameters | |
| (`attn.to_gate_compress`); it can be run dense (default) or with VSA for | |
| additional inference speedup. | |
| ## Usage (FastVideo) | |
| ```python | |
| from fastvideo import VideoGenerator | |
| gen = VideoGenerator.from_pretrained( | |
| "FastVideo/FastVideo-Minimax-FastVideo-Minimax-FastH3-Preview-v0.1", | |
| num_gpus=1, | |
| ) | |
| video = gen.generate_video( | |
| prompt="<your H3-format multimodal prompt>", | |
| num_inference_steps=4, # the distilled grid | |
| guidance_scale=1.0, # the base model is guidance-distilled | |
| ) | |
| ``` | |
| Prompts follow the MiniMax-H3 multimodal prompt format | |
| (`integrated_multimodal_description: ... overall_soundscape: ...`); see the base | |
| model card for the prompting guide. | |
| ## Training summary | |
| - **Method:** data-free DMD2 (distribution matching distillation) β student / | |
| frozen teacher / trained fake-score critic, backward-simulation rollout | |
| (the student walks its own 4-step sampling grid during training), x0-space | |
| critic regression, shifted score-time sampling matched to the dual video/audio | |
| noise clocks (shifts 12 / 3). | |
| - **Student grid:** 4 steps on the release sampler's shift-12 schedule. | |
| - **Attention:** student trained with VSA block-sparse attention (64-token tiles, | |
| 90% video-tile sparsity); teacher and critic dense. | |
| - **Data:** text prompts only (data-free) β ~258k prompts (VidProM-H3 + | |
| synthetic t2va prompt set); no video data used. | |
| - **Precision:** fp32 master weights, bf16 compute. | |
| - **Hardware:** 32Γ NVIDIA GB200. | |
| ## Limitations | |
| - Preview checkpoint β quality below the base model's 50-step sampling, | |
| especially on fine motion and audio detail; improves with training. | |
| - Inherits all content limitations and usage restrictions of the base model. | |
| - The 4-step grid is what the student was trained for; other step counts are | |
| off-distribution. | |
| ## License | |
| Distributed under the MiniMax H3 Community License (see [LICENSE](LICENSE)), | |
| inherited from the base model. Review the license (including its territory and | |
| acceptable-use terms) before use or redistribution. | |
| ## Notes | |
| - The `transformer_ref` component (reference-conditioning variant) is **not** | |
| packaged here; its entry in `modular_model_index.json` points at the base | |
| `MiniMaxAI/MiniMax-H3` repo and is fetched from there if used. This preview | |
| distills the text-to-video+audio path only. | |