File size: 11,269 Bytes
4d76c8b
df252a6
 
 
 
4d76c8b
 
 
df252a6
 
 
 
4d76c8b
 
df252a6
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
cb54f74
 
 
 
 
 
 
 
445cdf8
cb54f74
 
445cdf8
 
 
 
cb54f74
 
 
 
 
 
 
 
 
 
 
df252a6
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
---
title: FastH3 4-step Preview
emoji: 🎬
colorFrom: red
colorTo: gray
sdk: gradio
sdk_version: 6.26.0
app_file: app.py
short_description: 4-step MiniMax-H3 with sparse attention β€” video + audio
python_version: "3.12"
startup_duration_timeout: 1h
suggested_hardware: zero-a10g
---

# FastH3 4-step Preview (VSA, data-free) β€” MiniMax-H3 in four forward passes

[`FastVideo/FastVideo-FastH3-4-step-Preview-v1-VSA-DataFree`](https://huggingface.co/FastVideo/FastVideo-FastH3-4-step-Preview-v1-VSA-DataFree)
is a **data-free DMD2 distillation** of [`MiniMaxAI/MiniMax-H3`](https://huggingface.co/MiniMaxAI/MiniMax-H3), the
33B dual-modality DiT that emits a video *and* its synchronized soundtrack from one denoising pass. The student
keeps the teacher's architecture exactly and only replaces `transformer/`, so it is a drop-in for the same
`diffusers` modular pipeline β€” but it needs **four** DiT forwards instead of thirty, and it was distilled **with
Video Sparse Attention on**.

Everything runs unquantized at **bfloat16**.

## The sampling contract: five grid points, four forwards

`MiniMaxH3Scheduler.set_timesteps(n)` builds `linspace(1, 0, n)`, applies the shift
`Οƒ' = sΒ·Οƒ / (1 + (sβˆ’1)Β·Οƒ)`, and then drops the trailing zero when it forms the timesteps β€” so **`n` sigma grid
points drive `n βˆ’ 1` model evaluations**. The distilled ladder is `t = 999, 749, 500, 250 β†’ 0`: five points, four
forwards.

This Space therefore fixes `num_inference_steps = 5`, matching the checkpoint's own
`fastvideo_inference.json` (`num_inference_steps: 5`, `transformer_forwards: 4`,
`dmd_denoising_steps: [999, 749, 500, 250]`, `guidance_scale: 1.0`). It is not a knob. There is no CFG and no
negative prompt β€” the teacher is guidance-distilled and the student inherits that.

## Video Sparse Attention is not optional here

This checkpoint is the **VSA** variant. Its `fastvideo_inference.json` pins
`attention_backend: VIDEO_SPARSE_ATTN_H3`, `vsa_sparsity: 0.9`, `vsa_tile_size: 64`, and the transformer ships 50
trained `attn.to_gate_compress` tensors (~3.6 GiB) that only the sparse path consumes. FastVideo publishes a
separate `…-Dense-DataFree` checkpoint for people who want dense β€” running *this* one dense is running it
off-distribution.

The published kernel is `vsa_kernel: sm100a`, which is GB200-only, and the ZeroGPU pool is **sm120** (RTX PRO 6000
Blackwell). FastVideo's other officially supported route is `--vsa-kernel triton`, which is pure Triton and
architecture-agnostic β€” so this Space vendors those two files verbatim from
[FastVideo](https://github.com/hao-ai-lab/FastVideo) (Apache-2.0) into `vsa_kernel/` and ports the H3 backend on top
of them:

| File | What it is |
|---|---|
| `vsa_kernel/block_sparse_attn_triton.py` | FastVideo's Triton block-sparse attention, verbatim except that the autotune sweep is collapsed to the single config it lands on for Blackwell (re-enable it with `H3_VSA_AUTOTUNE=1`). |
| `vsa_kernel/index.py` | FastVideo's `map_to_index` / `topk_index_to_map`, verbatim. |
| `vsa_h3.py` | The port: `MiniMaxH3VSAAttnProcessor`, a `diffusers` attention processor reproducing `MiniMaxH3VSABackend`. |

`vsa_h3.py` follows FastVideo's `video_sparse_attn_h3.py` step for step: 64-token `(4, 4, 4)` tiles over the
post-patchify video grid, segment-pure prefix tiles, per-head fp32 pooled tile scores,
`topk = max(1, min(⌈(1 βˆ’ sparsity)Β·n_video_tilesβŒ‰, n_video_tiles))`, prefix keys exempt (always selected) and prefix
queries always dense, plus the compression branch `softmax(scores) @ pool(v)` broadcast back over each tile row and
scaled by the trained gate with **no** activation. The tile geometry is derived per-forward from the pipeline's own
`token_tags` / `position_ids`, so the `[text | cond | audio | video]` packing stays authoritative.

`diffusers` 0.40.0 has no `to_gate_compress`, so `vsa_h3.add_gate_compress_modules()` patches
`MiniMaxH3TransformerBlock.__init__` **before** the pipeline loads; otherwise the 50 gate tensors load as
"unexpected keys" and are silently dropped. Blocks whose gate is all-zero have it removed again after load, exactly
as FastVideo does.

A hidden `/selftest` API endpoint runs the ported kernel at sparsity 0 against `F.scaled_dot_product_attention` on
a real packed layout, so a wrong tile order or transpose is caught without spending a generation.

## Split across two Spaces

MiniMax-H3 is ~196 GiB in bfloat16 and a ZeroGPU Space is evicted at **150 GB of storage**, so no single
unquantized Space can hold it. Cutting `MiniMaxH3Blocks` at its `text_encoder` step splits it in two, and both
halves fit:

| Space | Subfolders | Download |
|---|---|---|
| [`multimodalart/qwen3vl-conditioner`](https://huggingface.co/spaces/multimodalart/qwen3vl-conditioner) | `text_encoder/` + `tokenizer/` + `processor/` | 66.7 GB β†’ 62.15 GiB bf16 |
| this one | `transformer/` + `vae/` + `audio_vae/` | ~85 GB β†’ 75.7 GiB resident |

The conditioner is a **public** Space and is unchanged by the distillation β€” the student's
`modular_model_index.json` points its `text_encoder` at the same Qwen3-VL weights β€” so this Space reuses it as-is
and calls it over the gradio API for every request.

`h3_split_blocks.py` subclasses the pipeline's blocks with the `text_encoder` step removed. Dropping the step drops
the components it declares, so `load_components` resolves only `transformer` / `vae` / `audio_vae` / the two
schedulers, and `prompt_embeds` + `text_token_tags` become ordinary required pipeline inputs. The wire format is
those two tensors β€” `(1, num_text_tokens, 5120)` bfloat16 and `(num_text_tokens,)` int64 β€” in one safetensors file
with the resolved `height` / `width` / `num_frames` in its metadata header.

## Text-to-video+audio only

The preview distills **the T2VA path only**, so this Space exposes no keyframe / reference inputs β€” the student's
`transformer_ref` tower is not packaged with the checkpoint. Image conditioning is what
[`multimodalart/minimax-h3`](https://huggingface.co/spaces/multimodalart/minimax-h3) (the undistilled 30-step
teacher) is for.

## Prompt format

MiniMax-H3 is trained on a structured multimodal caption, not a bare sentence:

```
integrated_multimodal_description: <shots, camera, subjects, action, lighting>
overall_soundscape: <diegetic sound>
non_diegetic_music: <score>
```

**Expand prompt** (on by default) sends a short prompt through the conditioner's Qwen3-VL prompt rewriter, which
writes that structure for you and returns it. Turn it off when you have already written a full-format prompt β€” the
examples that carry MiniMax's own official prompts do exactly that.

## Examples

The two long examples are MiniMax's own published prompts, taken verbatim from the base model's repo (Apache-2.0
code / docs):

- the starship-bridge two-shot from `scripts/readme/reproducible-768p-t2va-request.sh`
- the bakery two-shot, Case 1 of `docs/VIDEO_PROMPT_WRITING_GUIDE_base_en.md`

## Generation constraints

Fixed by the checkpoint: 24 fps, a 768 pixel short edge at the training canvas, `num_frames` snapped up to the next
`17Β·n + 5`. The distillation's operating point is **1344Γ—768 Γ— 124 frames (β‰ˆ5 s)** β€” that is the default, and it is
exactly the layout VSA was tuned on (`grid = 37Γ—24Γ—42`, 37 296 video rows, 672 tiles, 66 selected). The duration
slider reaches 8 s and the canvas dropdown offers smaller/faster grids, both outside the distilled operating point,
so quality degrades gracefully rather than being guaranteed. The 8 s ceiling is a VRAM limit, not a model limit:
75.7 GiB of weights sit resident on a 95.0 GiB card and the sparse working set grows with sequence length.

The checkpoint itself is a **preview**.

## Measured

On this Space, over `gradio_client`. Startup β€” the ~85 GB download plus the load, with no placement and therefore no
pack β€” is **78 s**.

| Canvas Γ— frames | Packed rows | Video tiles kept | Denoise + decode | Per forward | Peak allocated |
|---|---|---|---|---|---|
| 544Γ—544 Γ— 56 (2.3 s) | 4 913 | 10/100 | 6 s | 1.4 s | 81.05 GiB |
| 1344Γ—768 Γ— 124 (5.2 s) | 37 296 | 66/660 | 59–72 s | 14.7–18.0 s | 83.70 GiB |
| 1344Γ—768 Γ— 192 (8.0 s) | 57 456 | 102/1020 | 87 s | 21.9 s | 88.00 GiB |

Sparse attention makes the cost **linear** in the packed rows rather than quadratic, so `get_duration` drops the
quadratic term a dense path needs. The spread on the 5 s point is how fast a slice of the pool the request lands on;
the fit is taken at the slow end (`1.95e-3 Β· rows`) so a slow slice is not aborted mid-video. That books 98 s at the
5 s default and 143 s at the 8 s maximum.

A cold worker pays a one-time **68 s**: 75.7 GiB across PCIe (~11 s) plus the Triton JIT of the vendored kernels
(~57 s). The same 544Γ—544 request measures 74 s cold against 6 s warm. That is booked only on the first request a
process serves rather than padded onto every request.

At 88.00 GiB on a 95.0 GiB card the 8 s ceiling leaves ~7 GiB β€” which is why the slider stops at 8 s and not 10 s.

The vendored kernels are checked against `F.scaled_dot_product_attention` on the live GPU by the hidden `/selftest`
endpoint: at `sparsity = 0` the ported path reproduces dense attention to **4.1e-03 relative / 0.999996 cosine**,
which is bf16 rounding.

## Placement

`H3_PLACEMENT=lazy`: the weights move onto the card on the first GPU call and stay there. `spaces`' startup
`torch.pack()` would write a second on-disk copy of every resident CUDA tensor, and 85 + 75 GB exceeds the 150 GB
quota, so packing is not an option here. The one-time `.to("cuda")` plus the Triton JIT of the block-sparse kernels
lands inside the first request of a cold worker; after that there is **no offloading in the request path at all**.

## Space variables

| Variable | Default | Meaning |
|---|---|---|
| `H3_MODEL_REPO` | `FastVideo/FastVideo-FastH3-4-step-Preview-v1-VSA-DataFree` | The distilled diffusers-layout checkpoint. Public. |
| `H3_CONDITIONER` | `multimodalart/qwen3vl-conditioner` | The public Space this one asks for embeddings. |
| `H3_ATTENTION` | `vsa` | The ported VSA-H3 backend. Any other value is passed to `set_attention_backend` as a dense escape hatch (e.g. `_native_cudnn`). |
| `H3_VSA_SPARSITY` | `0.9` | The checkpoint's trained sparsity. |
| `H3_VSA_AUTOTUNE` | unset | Set to `1` to restore FastVideo's full Triton autotune sweep instead of the pinned Blackwell config. |
| `H3_PLACEMENT` | `lazy` | `lazy` moves all weights onto the card on the first GPU call; `offload` hands placement to `ComponentsManager.enable_auto_cpu_offload`. |
| `H3_GPU_SIZE` | `xlarge` | ZeroGPU allocation size. `large` (48 GB) does not fit 75 GiB of weights. |

## Secrets

None. Every weight this Space downloads is public, and the conditioner is a public Space called on the requesting
user's own ZeroGPU token.

## License

The weights are under the **MiniMax H3 Community License**, inherited from the base model β€” it carries territory and
acceptable-use restrictions. Read
[the base model's license](https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/LICENSE) before using outputs. The
vendored kernels in `vsa_kernel/` are Apache-2.0, from
[hao-ai-lab/FastVideo](https://github.com/hao-ai-lab/FastVideo).
</content>
</invoke>