comfyui
nvfp4
video
quantized
File size: 4,343 Bytes
39083bd
 
 
 
 
 
 
 
 
 
 
 
 
 
daef8bc
 
 
39083bd
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
a8b97d3
 
 
39083bd
a8b97d3
 
 
 
 
 
 
 
daef8bc
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
a8b97d3
 
 
 
 
39083bd
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
---
license: other
license_name: minimax-h3-community-license-agreement
license_link: https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/LICENSE
base_model: MiniMaxAI/MiniMax-H3
tags:
- comfyui
- nvfp4
- video
- quantized
---

# MiniMax H3 ref2va β€” NVFP4 (ComfyUI-native)

> πŸš€ **Use this model hosted (API + playground):**
> [modelslab.com/models/minimax/minimax-hailuo03-reference-to-video](https://modelslab.com/models/minimax/minimax-hailuo03-reference-to-video)

NVFP4 quantization of the [MiniMax H3](https://huggingface.co/MiniMaxAI/MiniMax-H3)
**reference-to-video** DiT (`ref2va`), in ComfyUI's native quant layout β€” loads with
the stock `UNETLoader` on any Blackwell GPU (sm_120: RTX 5090 / RTX PRO 6000, and
newer). Derived from
[Comfy-Org/MiniMax-H3](https://huggingface.co/Comfy-Org/MiniMax-H3)'s
`minimax_h3_ref2va_bf16.safetensors`.

## Files

| File | Size | GEMM path |
|---|---|---|
| `diffusion_models/minimax_h3_ref2va_nvfp4.safetensors` | 38.6 GB | native FP4 tensor-core |
| `diffusion_models/minimax_h3_ref2va_nvfp4_fpmm.safetensors` | 38.6 GB | dequant β†’ bf16 (quality-safe, same path the official NVFP4 text encoder uses) |

Quantization policy mirrors the official `int8_convrot` release: the 50 main
blocks' `attn.qkv_proj`, `attn.out_proj`, `mlp.fc1`, `mlp.fc2` (200 layers) go
to NVFP4 (E2M1, per-16 FP8-E4M3 block scales + global FP32 scale); everything
quality-critical β€” adaln/modulation, norms, patch/condition projections, token
refiner, final layers β€” stays bf16. Weight-only PTQ (no activation calibration),
produced with the included `convert_nvfp4.py` via ComfyUI's `comfy_kitchen`
`quantize_nvfp4` kernels.

## Usage (ComfyUI)

Drop into `ComfyUI/models/diffusion_models/` and select in `UNETLoader`
(weight_dtype `default`) inside the official
[MiniMax H3 r2v template](https://github.com/Comfy-Org/workflow_templates/blob/main/templates/video_minimax_h3_r2v.json).
Pair with the official `qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors` text
encoder and both H3 VAEs from
[Comfy-Org/MiniMax-H3](https://huggingface.co/Comfy-Org/MiniMax-H3).

## Benchmarks (RTX PRO 6000 Blackwell Max-Q, 96 GB, SageAttention, 1344Γ—768Γ—124f, 20 steps)

All on the same card, same seed, SageAttention. **torch cu130 build strongly
recommended** β€” it enables comfy_kitchen's optimized CUDA kernels (on cu128
the FP4 path runs emulated and is ~2.3Γ— slower).

| DiT variant | Size | s/step (cu130) | s/step (cu128) | Notes |
|---|---|---|---|---|
| **nvfp4 (this repo)** | 38.6 GB | **8.47** | 19.8 (emulated) | full model, **1.5Γ— faster than bf16** |
| convrot W4A4 (measured, not shipped) | 37.4 GB | 9.69 | β€” | visible text-rendering artifacts |
| bf16 (reference) | 66.3 GB | 12.76 | 14.65 | baseline quality |
| **nvfp4_fpmm (this repo)** | 38.6 GB | β€” | 14.99 | ties bf16 on cu128, 42 % less VRAM |
| pruned int8_convrot (official) | 21.0 GB | β€” | 22.7 | pruned arch, W8A16 dequant path |

Stacking ComfyUI's `EasyCache` (threshold 0.2, quality-gated) cuts wall time a
further ~1.4Γ—.

## End-to-end speed (full optimized stack)

600W RTX PRO 6000 Blackwell, `nvfp4` + cu130 + SageAttention +
fp16_accumulation + fused adaln/gate kernels + EasyCache 0.2, 20 steps,
1344Γ—768 @ 24 fps with native stereo audio:

| Video length | Wall time | Notes |
|---|---|---|
| **15 s** (362 frames β€” the model's one-shot ceiling) | **8 min 07 s** | vs ~25 min on the stock template path (~3Γ—) |
| **5 s** (124 frames) | **~2 min** (hot server) | measured through a production API |
| 15 s draft (16 steps @ 960Γ—544) | ~3.5–4 min | preview tier |

Important: H3's trained range tops out at 362 frames (~15 s) β€” longer one-shot
generations collapse regardless of settings (verified empirically at 736
frames with multiple schedules/shifts).

Same-seed visual quality of `nvfp4` vs bf16: no quantization artifacts observed
(trajectory divergence only β€” diffusion is chaotic under any weight
perturbation; see sample). The W4A4 experiment degraded fine text rendering,
which is why it is not shipped.

A same-seed sample generated with this checkpoint is in
[`assets/sample_r2v_5s.mp4`](assets/sample_r2v_5s.mp4).

## License

MiniMax H3 Community License. This is a derivative work of MiniMaxAI/MiniMax-H3;
see the [license](https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/LICENSE)
for terms.