File size: 4,343 Bytes
39083bd daef8bc 39083bd a8b97d3 39083bd a8b97d3 daef8bc a8b97d3 39083bd | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 | ---
license: other
license_name: minimax-h3-community-license-agreement
license_link: https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/LICENSE
base_model: MiniMaxAI/MiniMax-H3
tags:
- comfyui
- nvfp4
- video
- quantized
---
# MiniMax H3 ref2va β NVFP4 (ComfyUI-native)
> π **Use this model hosted (API + playground):**
> [modelslab.com/models/minimax/minimax-hailuo03-reference-to-video](https://modelslab.com/models/minimax/minimax-hailuo03-reference-to-video)
NVFP4 quantization of the [MiniMax H3](https://huggingface.co/MiniMaxAI/MiniMax-H3)
**reference-to-video** DiT (`ref2va`), in ComfyUI's native quant layout β loads with
the stock `UNETLoader` on any Blackwell GPU (sm_120: RTX 5090 / RTX PRO 6000, and
newer). Derived from
[Comfy-Org/MiniMax-H3](https://huggingface.co/Comfy-Org/MiniMax-H3)'s
`minimax_h3_ref2va_bf16.safetensors`.
## Files
| File | Size | GEMM path |
|---|---|---|
| `diffusion_models/minimax_h3_ref2va_nvfp4.safetensors` | 38.6 GB | native FP4 tensor-core |
| `diffusion_models/minimax_h3_ref2va_nvfp4_fpmm.safetensors` | 38.6 GB | dequant β bf16 (quality-safe, same path the official NVFP4 text encoder uses) |
Quantization policy mirrors the official `int8_convrot` release: the 50 main
blocks' `attn.qkv_proj`, `attn.out_proj`, `mlp.fc1`, `mlp.fc2` (200 layers) go
to NVFP4 (E2M1, per-16 FP8-E4M3 block scales + global FP32 scale); everything
quality-critical β adaln/modulation, norms, patch/condition projections, token
refiner, final layers β stays bf16. Weight-only PTQ (no activation calibration),
produced with the included `convert_nvfp4.py` via ComfyUI's `comfy_kitchen`
`quantize_nvfp4` kernels.
## Usage (ComfyUI)
Drop into `ComfyUI/models/diffusion_models/` and select in `UNETLoader`
(weight_dtype `default`) inside the official
[MiniMax H3 r2v template](https://github.com/Comfy-Org/workflow_templates/blob/main/templates/video_minimax_h3_r2v.json).
Pair with the official `qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors` text
encoder and both H3 VAEs from
[Comfy-Org/MiniMax-H3](https://huggingface.co/Comfy-Org/MiniMax-H3).
## Benchmarks (RTX PRO 6000 Blackwell Max-Q, 96 GB, SageAttention, 1344Γ768Γ124f, 20 steps)
All on the same card, same seed, SageAttention. **torch cu130 build strongly
recommended** β it enables comfy_kitchen's optimized CUDA kernels (on cu128
the FP4 path runs emulated and is ~2.3Γ slower).
| DiT variant | Size | s/step (cu130) | s/step (cu128) | Notes |
|---|---|---|---|---|
| **nvfp4 (this repo)** | 38.6 GB | **8.47** | 19.8 (emulated) | full model, **1.5Γ faster than bf16** |
| convrot W4A4 (measured, not shipped) | 37.4 GB | 9.69 | β | visible text-rendering artifacts |
| bf16 (reference) | 66.3 GB | 12.76 | 14.65 | baseline quality |
| **nvfp4_fpmm (this repo)** | 38.6 GB | β | 14.99 | ties bf16 on cu128, 42 % less VRAM |
| pruned int8_convrot (official) | 21.0 GB | β | 22.7 | pruned arch, W8A16 dequant path |
Stacking ComfyUI's `EasyCache` (threshold 0.2, quality-gated) cuts wall time a
further ~1.4Γ.
## End-to-end speed (full optimized stack)
600W RTX PRO 6000 Blackwell, `nvfp4` + cu130 + SageAttention +
fp16_accumulation + fused adaln/gate kernels + EasyCache 0.2, 20 steps,
1344Γ768 @ 24 fps with native stereo audio:
| Video length | Wall time | Notes |
|---|---|---|
| **15 s** (362 frames β the model's one-shot ceiling) | **8 min 07 s** | vs ~25 min on the stock template path (~3Γ) |
| **5 s** (124 frames) | **~2 min** (hot server) | measured through a production API |
| 15 s draft (16 steps @ 960Γ544) | ~3.5β4 min | preview tier |
Important: H3's trained range tops out at 362 frames (~15 s) β longer one-shot
generations collapse regardless of settings (verified empirically at 736
frames with multiple schedules/shifts).
Same-seed visual quality of `nvfp4` vs bf16: no quantization artifacts observed
(trajectory divergence only β diffusion is chaotic under any weight
perturbation; see sample). The W4A4 experiment degraded fine text rendering,
which is why it is not shipped.
A same-seed sample generated with this checkpoint is in
[`assets/sample_r2v_5s.mp4`](assets/sample_r2v_5s.mp4).
## License
MiniMax H3 Community License. This is a derivative work of MiniMaxAI/MiniMax-H3;
see the [license](https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/LICENSE)
for terms.
|