Text-to-Video
Diffusers
Safetensors
MiniMax H3
video
audio
text-to-audio-video
distillation
dmd2
few-step
fastvideo
fasth3
Instructions to use FastVideo/FastVideo-Minimax-FastH3-Preview-v0.1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use FastVideo/FastVideo-Minimax-FastH3-Preview-v0.1 with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("FastVideo/FastVideo-Minimax-FastH3-Preview-v0.1", dtype=torch.bfloat16, device_map="cuda") prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k" image = pipe(prompt).images[0] - Notebooks
- Google Colab
- Kaggle
File size: 3,868 Bytes
41421a3 46fca7f 41421a3 31f8e8d 41421a3 31f8e8d 41421a3 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 | ---
license: other
license_name: minimax-h3-community
license_link: LICENSE
base_model: MiniMaxAI/MiniMax-H3
library_name: diffusers
pipeline_tag: text-to-video
tags:
- text-to-video
- video
- audio
- text-to-audio-video
- distillation
- dmd2
- few-step
- minimax-h3
- fastvideo
- fasth3
---
# FastVideo-Minimax-FastH3-Preview-v0.1
**A few-step (4-step) distillation preview of [MiniMax-H3](https://huggingface.co/MiniMaxAI/MiniMax-H3)**, the
33B dual-modality (video + audio) diffusion transformer — distilled with data-free
DMD2 by the [FastVideo](https://github.com/hao-ai-lab/FastVideo) team.
The base model samples with 50 denoising steps; this student walks a 4-step grid
on the release's shift-12 rectified-flow schedule (12.5× fewer transformer
evaluations), generating synchronized video and audio in one pipeline call.
> **Preview status (v0.1):** this is an early training checkpoint (step 1400 of a
> 4000-step run) published for evaluation and integration work. Sample quality is
> still maturing; expect a stronger release checkpoint from the same run.
## What's in the repo
Diffusers-format (modular pipeline) layout. Only the `transformer/` weights differ
from the base release — the distilled student, in bf16. All other components
(Qwen3-VL text encoder, video/audio VAEs, tokenizer, processor, schedulers) are
unmodified copies of the base release, included so the repo is self-contained.
The student was trained with block-sparse video attention (VSA, 64-token tiles,
90% sparsity) and carries its trained sparse-gate parameters
(`attn.to_gate_compress`); it can be run dense (default) or with VSA for
additional inference speedup.
## Usage (FastVideo)
```python
from fastvideo import VideoGenerator
gen = VideoGenerator.from_pretrained(
"FastVideo/FastVideo-Minimax-FastVideo-Minimax-FastH3-Preview-v0.1",
num_gpus=1,
)
video = gen.generate_video(
prompt="<your H3-format multimodal prompt>",
num_inference_steps=4, # the distilled grid
guidance_scale=1.0, # the base model is guidance-distilled
)
```
Prompts follow the MiniMax-H3 multimodal prompt format
(`integrated_multimodal_description: ... overall_soundscape: ...`); see the base
model card for the prompting guide.
## Training summary
- **Method:** data-free DMD2 (distribution matching distillation) — student /
frozen teacher / trained fake-score critic, backward-simulation rollout
(the student walks its own 4-step sampling grid during training), x0-space
critic regression, shifted score-time sampling matched to the dual video/audio
noise clocks (shifts 12 / 3).
- **Student grid:** 4 steps on the release sampler's shift-12 schedule.
- **Attention:** student trained with VSA block-sparse attention (64-token tiles,
90% video-tile sparsity); teacher and critic dense.
- **Data:** text prompts only (data-free) — ~258k prompts (VidProM-H3 +
synthetic t2va prompt set); no video data used.
- **Precision:** fp32 master weights, bf16 compute.
- **Hardware:** 32× NVIDIA GB200.
## Limitations
- Preview checkpoint — quality below the base model's 50-step sampling,
especially on fine motion and audio detail; improves with training.
- Inherits all content limitations and usage restrictions of the base model.
- The 4-step grid is what the student was trained for; other step counts are
off-distribution.
## License
Distributed under the MiniMax H3 Community License (see [LICENSE](LICENSE)),
inherited from the base model. Review the license (including its territory and
acceptable-use terms) before use or redistribution.
## Notes
- The `transformer_ref` component (reference-conditioning variant) is **not**
packaged here; its entry in `modular_model_index.json` points at the base
`MiniMaxAI/MiniMax-H3` repo and is fetched from there if used. This preview
distills the text-to-video+audio path only.
|