--- license: other license_name: minimax-h3-community license_link: LICENSE base_model: MiniMaxAI/MiniMax-H3 library_name: diffusers pipeline_tag: text-to-video tags: - text-to-video - video - audio - text-to-audio-video - distillation - dmd2 - few-step - minimax-h3 - fastvideo - fasth3 --- # FastVideo-Minimax-FastH3-Preview-v0.1 **A few-step (4-step) distillation preview of [MiniMax-H3](https://huggingface.co/MiniMaxAI/MiniMax-H3)**, the 33B dual-modality (video + audio) diffusion transformer — distilled with data-free DMD2 by the [FastVideo](https://github.com/hao-ai-lab/FastVideo) team. The base model samples with 50 denoising steps; this student walks a 4-step grid on the release's shift-12 rectified-flow schedule (12.5× fewer transformer evaluations), generating synchronized video and audio in one pipeline call. > **Preview status (v0.1):** this is an early training checkpoint (step 1400 of a > 4000-step run) published for evaluation and integration work. Sample quality is > still maturing; expect a stronger release checkpoint from the same run. ## What's in the repo Diffusers-format (modular pipeline) layout. Only the `transformer/` weights differ from the base release — the distilled student, in bf16. All other components (Qwen3-VL text encoder, video/audio VAEs, tokenizer, processor, schedulers) are unmodified copies of the base release, included so the repo is self-contained. The student was trained with block-sparse video attention (VSA, 64-token tiles, 90% sparsity) and carries its trained sparse-gate parameters (`attn.to_gate_compress`); it can be run dense (default) or with VSA for additional inference speedup. ## Usage (FastVideo) ```python from fastvideo import VideoGenerator gen = VideoGenerator.from_pretrained( "FastVideo/FastVideo-Minimax-FastVideo-Minimax-FastH3-Preview-v0.1", num_gpus=1, ) video = gen.generate_video( prompt="", num_inference_steps=4, # the distilled grid guidance_scale=1.0, # the base model is guidance-distilled ) ``` Prompts follow the MiniMax-H3 multimodal prompt format (`integrated_multimodal_description: ... overall_soundscape: ...`); see the base model card for the prompting guide. ## Training summary - **Method:** data-free DMD2 (distribution matching distillation) — student / frozen teacher / trained fake-score critic, backward-simulation rollout (the student walks its own 4-step sampling grid during training), x0-space critic regression, shifted score-time sampling matched to the dual video/audio noise clocks (shifts 12 / 3). - **Student grid:** 4 steps on the release sampler's shift-12 schedule. - **Attention:** student trained with VSA block-sparse attention (64-token tiles, 90% video-tile sparsity); teacher and critic dense. - **Data:** text prompts only (data-free) — ~258k prompts (VidProM-H3 + synthetic t2va prompt set); no video data used. - **Precision:** fp32 master weights, bf16 compute. - **Hardware:** 32× NVIDIA GB200. ## Limitations - Preview checkpoint — quality below the base model's 50-step sampling, especially on fine motion and audio detail; improves with training. - Inherits all content limitations and usage restrictions of the base model. - The 4-step grid is what the student was trained for; other step counts are off-distribution. ## License Distributed under the MiniMax H3 Community License (see [LICENSE](LICENSE)), inherited from the base model. Review the license (including its territory and acceptable-use terms) before use or redistribution. ## Notes - The `transformer_ref` component (reference-conditioning variant) is **not** packaged here; its entry in `modular_model_index.json` points at the base `MiniMaxAI/MiniMax-H3` repo and is fetched from there if used. This preview distills the text-to-video+audio path only.