Mike0021's picture
phase 0 baseline
d4ceaf5 verified
|
Raw
History Blame Contribute Delete
9.87 kB
---
title: FastH3 4-step Preview (VSA)
emoji: 🎬
colorFrom: red
colorTo: gray
sdk: gradio
sdk_version: 6.26.0
app_file: app.py
short_description: 4-step MiniMax-H3 sparse attention β€” video + audio
python_version: "3.12"
startup_duration_timeout: 1h
suggested_hardware: zero-a10g
---
# FastH3 4-step Preview (VSA, data-free) β€” MiniMax-H3 in four forward passes
[`FastVideo/FastVideo-FastH3-4-step-Preview-v1-VSA-DataFree`](https://huggingface.co/FastVideo/FastVideo-FastH3-4-step-Preview-v1-VSA-DataFree)
is a **data-free DMD2 distillation** of [`MiniMaxAI/MiniMax-H3`](https://huggingface.co/MiniMaxAI/MiniMax-H3), the
33B dual-modality DiT that emits a video *and* its synchronized soundtrack from one denoising pass. The student
keeps the teacher's architecture exactly and only replaces `transformer/`, so it is a drop-in for the same
`diffusers` modular pipeline β€” but it needs **four** DiT forwards instead of thirty, and it was distilled **with
Video Sparse Attention on**.
Everything runs unquantized at **bfloat16**.
## The sampling contract: five grid points, four forwards
`MiniMaxH3Scheduler.set_timesteps(n)` builds `linspace(1, 0, n)`, applies the shift
`Οƒ' = sΒ·Οƒ / (1 + (sβˆ’1)Β·Οƒ)`, and then drops the trailing zero when it forms the timesteps β€” so **`n` sigma grid
points drive `n βˆ’ 1` model evaluations**. The distilled ladder is `t = 999, 749, 500, 250 β†’ 0`: five points, four
forwards.
This Space therefore fixes `num_inference_steps = 5`, matching the checkpoint's own
`fastvideo_inference.json` (`num_inference_steps: 5`, `transformer_forwards: 4`,
`dmd_denoising_steps: [999, 749, 500, 250]`, `guidance_scale: 1.0`). It is not a knob. There is no CFG and no
negative prompt β€” the teacher is guidance-distilled and the student inherits that.
## Video Sparse Attention is not optional here
This checkpoint is the **VSA** variant. Its `fastvideo_inference.json` pins
`attention_backend: VIDEO_SPARSE_ATTN_H3`, `vsa_sparsity: 0.9`, `vsa_tile_size: 64`, and the transformer ships 50
trained `attn.to_gate_compress` tensors (~3.6 GiB) that only the sparse path consumes. FastVideo publishes a
separate `…-Dense-DataFree` checkpoint for people who want dense β€” running *this* one dense is running it
off-distribution.
The published kernel is `vsa_kernel: sm100a`, which is GB200-only, and the ZeroGPU pool is **sm120** (RTX PRO 6000
Blackwell). FastVideo's other officially supported route is `--vsa-kernel triton`, which is pure Triton and
architecture-agnostic β€” so this Space vendors those two files verbatim from
[FastVideo](https://github.com/hao-ai-lab/FastVideo) (Apache-2.0) into `vsa_kernel/` and ports the H3 backend on top
of them:
| File | What it is |
|---|---|
| `vsa_kernel/block_sparse_attn_triton.py` | FastVideo's Triton block-sparse attention, verbatim except that the autotune sweep is collapsed to the single config it lands on for Blackwell (re-enable it with `H3_VSA_AUTOTUNE=1`). |
| `vsa_kernel/index.py` | FastVideo's `map_to_index` / `topk_index_to_map`, verbatim. |
| `vsa_h3.py` | The port: `MiniMaxH3VSAAttnProcessor`, a `diffusers` attention processor reproducing `MiniMaxH3VSABackend`. |
`vsa_h3.py` follows FastVideo's `video_sparse_attn_h3.py` step for step: 64-token `(4, 4, 4)` tiles over the
post-patchify video grid, segment-pure prefix tiles, per-head fp32 pooled tile scores,
`topk = max(1, min(⌈(1 βˆ’ sparsity)Β·n_video_tilesβŒ‰, n_video_tiles))`, prefix keys exempt (always selected) and prefix
queries always dense, plus the compression branch `softmax(scores) @ pool(v)` broadcast back over each tile row and
scaled by the trained gate with **no** activation. The tile geometry is derived per-forward from the pipeline's own
`token_tags` / `position_ids`, so the `[text | cond | audio | video]` packing stays authoritative.
`diffusers` 0.40.0 has no `to_gate_compress`, so `vsa_h3.add_gate_compress_modules()` patches
`MiniMaxH3TransformerBlock.__init__` **before** the pipeline loads; otherwise the 50 gate tensors load as
"unexpected keys" and are silently dropped. Blocks whose gate is all-zero have it removed again after load, exactly
as FastVideo does.
A hidden `/selftest` API endpoint runs the ported kernel at sparsity 0 against `F.scaled_dot_product_attention` on
a real packed layout, so a wrong tile order or transpose is caught without spending a generation.
## Split across two Spaces
MiniMax-H3 is ~196 GiB in bfloat16 and a ZeroGPU Space is evicted at **150 GB of storage**, so no single
unquantized Space can hold it. Cutting `MiniMaxH3Blocks` at its `text_encoder` step splits it in two, and both
halves fit:
| Space | Subfolders | Download |
|---|---|---|
| [`multimodalart/qwen3vl-conditioner`](https://huggingface.co/spaces/multimodalart/qwen3vl-conditioner) | `text_encoder/` + `tokenizer/` + `processor/` | 66.7 GB β†’ 62.15 GiB bf16 |
| this one | `transformer/` + `vae/` + `audio_vae/` | ~85 GB β†’ 75.7 GiB resident |
The conditioner is a **public** Space and is unchanged by the distillation β€” the student's
`modular_model_index.json` points its `text_encoder` at the same Qwen3-VL weights β€” so this Space reuses it as-is
and calls it over the gradio API for every request.
`h3_split_blocks.py` subclasses the pipeline's blocks with the `text_encoder` step removed. Dropping the step drops
the components it declares, so `load_components` resolves only `transformer` / `vae` / `audio_vae` / the two
schedulers, and `prompt_embeds` + `text_token_tags` become ordinary required pipeline inputs. The wire format is
those two tensors β€” `(1, num_text_tokens, 5120)` bfloat16 and `(num_text_tokens,)` int64 β€” in one safetensors file
with the resolved `height` / `width` / `num_frames` in its metadata header.
## Text-to-video+audio only
The preview distills **the T2VA path only**, so this Space exposes no keyframe / reference inputs β€” the student's
`transformer_ref` tower is not packaged with the checkpoint. Image conditioning is what
[`multimodalart/minimax-h3`](https://huggingface.co/spaces/multimodalart/minimax-h3) (the undistilled 30-step
teacher) is for.
## Prompt format
MiniMax-H3 is trained on a structured multimodal caption, not a bare sentence:
```
integrated_multimodal_description: <shots, camera, subjects, action, lighting>
overall_soundscape: <diegetic sound>
non_diegetic_music: <score>
```
**Expand prompt** (on by default) sends a short prompt through the conditioner's Qwen3-VL prompt rewriter, which
writes that structure for you and returns it. Turn it off when you have already written a full-format prompt β€” the
examples that carry MiniMax's own official prompts do exactly that.
## Generation constraints
Fixed by the checkpoint: 24 fps, a 768 pixel short edge at the training canvas, `num_frames` snapped up to the next
`17Β·n + 5`. The distillation's operating point is **1344Γ—768 Γ— 124 frames (β‰ˆ5 s)** β€” that is the default. The 8 s
slider ceiling and the smaller canvases sit outside the distilled operating point, so quality degrades gracefully
rather than being guaranteed. The 8 s ceiling is a VRAM limit, not a model limit: 75.7 GiB of weights sit resident
on a 95.0 GiB card and the sparse working set grows with sequence length.
The checkpoint itself is a **preview**.
## Latency instrumentation
Every response carries `gen_s` β€” the server-side wall time of the pipeline call (four transformer forwards plus the
two decoders), measured with `time.perf_counter()` inside the `@spaces.GPU` function β€” along with the per-forward
times and the peak CUDA allocation in the report line, and a per-request id proving each artifact is freshly
generated. Historical numbers below were measured on the reference deployment of this exact configuration, over
`gradio_client`; this Space's own numbers are in its `.auto/log.jsonl`.
| Canvas Γ— frames | Packed rows | Video tiles kept | Denoise + decode | Peak allocated |
|---|---|---|---|---|
| 544Γ—544 Γ— 56 (2.3 s) | 4 913 | 10/100 | 6 s | 81.05 GiB |
| 1344Γ—768 Γ— 124 (5.2 s) | 37 296 | 66/660 | 59–72 s | 83.70 GiB |
| 1344Γ—768 Γ— 192 (8.0 s) | 57 456 | 102/1020 | 87 s | 88.00 GiB |
A cold worker pays a one-time **~68 s** (75.7 GiB across PCIe ~11 s + Triton JIT of the vendored kernels ~57 s),
booked only on the first request a process serves.
## Space variables
| Variable | Default | Meaning |
|---|---|---|
| `H3_MODEL_REPO` | `FastVideo/FastVideo-FastH3-4-step-Preview-v1-VSA-DataFree` | The distilled diffusers-layout checkpoint. Public. |
| `H3_CONDITIONER` | `multimodalart/qwen3vl-conditioner` | The public Space this one asks for embeddings. |
| `H3_ATTENTION` | `vsa` | The ported VSA-H3 backend. Any other value is passed to `set_attention_backend` as a dense escape hatch (e.g. `_native_cudnn`). |
| `H3_VSA_SPARSITY` | `0.9` | The checkpoint's trained sparsity. |
| `H3_VSA_AUTOTUNE` | unset | Set to `1` to restore FastVideo's full Triton autotune sweep instead of the pinned Blackwell config. |
| `H3_VSA_WARPS` | `4` | `num_warps` of the pinned Triton block-sparse config. |
| `H3_PLACEMENT` | `lazy` | `lazy` moves all weights onto the card on the first GPU call; `offload` hands placement to `ComponentsManager.enable_auto_cpu_offload`. |
| `H3_GPU_SIZE` | `xlarge` | ZeroGPU allocation size. `large` (48 GB) does not fit 75 GiB of weights. |
## Secrets
None. Every weight this Space downloads is public, and the conditioner is a public Space called on the requesting
user's own ZeroGPU token.
## License
The weights are under the **MiniMax H3 Community License**, inherited from the base model β€” it carries territory and
acceptable-use restrictions. Read
[the base model's license](https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/LICENSE) before using outputs. The
vendored kernels in `vsa_kernel/` are Apache-2.0, from [hao-ai-lab/FastVideo](https://github.com/hao-ai-lab/FastVideo).