File size: 3,868 Bytes
41421a3
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
46fca7f
41421a3
 
31f8e8d
41421a3
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
31f8e8d
41421a3
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
---
license: other
license_name: minimax-h3-community
license_link: LICENSE
base_model: MiniMaxAI/MiniMax-H3
library_name: diffusers
pipeline_tag: text-to-video
tags:
- text-to-video
- video
- audio
- text-to-audio-video
- distillation
- dmd2
- few-step
- minimax-h3
- fastvideo
- fasth3
---

# FastVideo-Minimax-FastH3-Preview-v0.1

**A few-step (4-step) distillation preview of [MiniMax-H3](https://huggingface.co/MiniMaxAI/MiniMax-H3)**, the
33B dual-modality (video + audio) diffusion transformer — distilled with data-free
DMD2 by the [FastVideo](https://github.com/hao-ai-lab/FastVideo) team.

The base model samples with 50 denoising steps; this student walks a 4-step grid
on the release's shift-12 rectified-flow schedule (12.5× fewer transformer
evaluations), generating synchronized video and audio in one pipeline call.

> **Preview status (v0.1):** this is an early training checkpoint (step 1400 of a
> 4000-step run) published for evaluation and integration work. Sample quality is
> still maturing; expect a stronger release checkpoint from the same run.

## What's in the repo

Diffusers-format (modular pipeline) layout. Only the `transformer/` weights differ
from the base release — the distilled student, in bf16. All other components
(Qwen3-VL text encoder, video/audio VAEs, tokenizer, processor, schedulers) are
unmodified copies of the base release, included so the repo is self-contained.

The student was trained with block-sparse video attention (VSA, 64-token tiles,
90% sparsity) and carries its trained sparse-gate parameters
(`attn.to_gate_compress`); it can be run dense (default) or with VSA for
additional inference speedup.

## Usage (FastVideo)

```python
from fastvideo import VideoGenerator

gen = VideoGenerator.from_pretrained(
    "FastVideo/FastVideo-Minimax-FastVideo-Minimax-FastH3-Preview-v0.1",
    num_gpus=1,
)
video = gen.generate_video(
    prompt="<your H3-format multimodal prompt>",
    num_inference_steps=4,   # the distilled grid
    guidance_scale=1.0,      # the base model is guidance-distilled
)
```

Prompts follow the MiniMax-H3 multimodal prompt format
(`integrated_multimodal_description: ... overall_soundscape: ...`); see the base
model card for the prompting guide.

## Training summary

- **Method:** data-free DMD2 (distribution matching distillation) — student /
  frozen teacher / trained fake-score critic, backward-simulation rollout
  (the student walks its own 4-step sampling grid during training), x0-space
  critic regression, shifted score-time sampling matched to the dual video/audio
  noise clocks (shifts 12 / 3).
- **Student grid:** 4 steps on the release sampler's shift-12 schedule.
- **Attention:** student trained with VSA block-sparse attention (64-token tiles,
  90% video-tile sparsity); teacher and critic dense.
- **Data:** text prompts only (data-free) — ~258k prompts (VidProM-H3 +
  synthetic t2va prompt set); no video data used.
- **Precision:** fp32 master weights, bf16 compute.
- **Hardware:** 32× NVIDIA GB200.

## Limitations

- Preview checkpoint — quality below the base model's 50-step sampling,
  especially on fine motion and audio detail; improves with training.
- Inherits all content limitations and usage restrictions of the base model.
- The 4-step grid is what the student was trained for; other step counts are
  off-distribution.

## License

Distributed under the MiniMax H3 Community License (see [LICENSE](LICENSE)),
inherited from the base model. Review the license (including its territory and
acceptable-use terms) before use or redistribution.

## Notes

- The `transformer_ref` component (reference-conditioning variant) is **not**
  packaged here; its entry in `modular_model_index.json` points at the base
  `MiniMaxAI/MiniMax-H3` repo and is fetched from there if used. This preview
  distills the text-to-video+audio path only.