MiniMax-H3 · NVFP4-rotated (single-file, 16 GB)
A rotated-NVFP4 quantization of the base MiniMax-H3 video + synchronized-audio DiT, as one self-contained ~12.5 GB file that runs fully resident on a single 16 GB Blackwell card, zero CPU offload, on the native nunchaku W4A4 kernel — at bf16-class sharpness. This is the quality / stability line (full 8-step). For a ~5×-faster 4-step version see FastH3-NVFP4-rotated.
Loads with our
H3RotNVFP4Loadernode — get it at H3-RotNVFP4-ComfyUI-Loader (why a custom node, full method, dependency versions there).
Why this line (vs FastH3)
Both lines are the same rotated-NVFP4 tech and both fit 16 GB resident. The difference is denoising steps:
- This (base, 8-step): more steps → temporally stable and crisp even on the hard scenes — neon night, rain, water spray render clean, no flicker. Slower.
- FastH3 (4-step): ~5× faster, but glow/motion scenes (neon/water) show temporal flicker inherent to the 4-step distillation.
If you care about clean neon/water/rain or maximum sharpness, use this line. If you want speed on static/portrait content, use FastH3.
Measured (1× RTX 5070 Ti 16 GB, warm, DiT-only)
- Genuine 16 GB residency, zero offload — ~0.9 GB CPU offload, peak VRAM 13.6 GB, 5.7 s/step. System RAM ~31 GB to load (single file, no 80 GB shell floor).
- 480p (832×480) full 124-frame clip ~66 s warm. Audio preserved (clean 32 kHz stereo).
- Max clip length on a single 16 GB card (dynamic mode): 480p ~12.9 s (310 frames), 768p ~5.17 s (124 frames) — resolution-driven (same ceiling as the FastH3 line).
- Only 8 steps for this quality line — vs the ~10 a stock-nvfp4 recipe typically needs. Per step it's on par with community quants on the same card (same nunchaku W4A4 kernel; rotation overhead offset). (The 4-step FastH3 line is faster still.)
- File 12.5 GB (T1, all-NVFP4 blocks + curve-basis AdaLN + bf16 islands).
Why "rotated"?
The rotated in the name is a runtime block-256 Hadamard applied before the NVFP4 W4A4 GEMM (QuaRot-style — it spreads outliers so the 4-bit grid isn't dominated by a few large values). A clean same-weights check — identical weights and settings, only rotation toggled — confirms it earns its place: rotation keeps crisp edges where un-rotated NVFP4 hazes over (measurably sharper on every no-reference IQA metric, ~+70% on a Laplacian-variance sharpness proxy), at a cost of about 15% per step and a different per-seed draw. Un-rotated NVFP4 isn't broken — fine on easy scenes at enough steps — but where detail counts rotation is the sharper half. Full method + measurements: H3-RotNVFP4-ComfyUI-Loader.
Samples
![]() neon rain street — clean reflections |
![]() sci-fi rooftop — flying cars |
![]() sailboat / sea |
Full MP4 in samples/. These are 8-step renders — temporally clean and crisp even on neon/water, where the 4-step FastH3 line needs care. (GIFs are video-only previews.)
Setup
- Install the loader node + nunchaku — see H3-RotNVFP4-ComfyUI-Loader (verified
torch 2.12.1+cu130+ a nunchaku cu13/torch2.12 wheel; Linux + Blackwell RTX 50-series / GB10). - Download
h3_base_T1.safetensorsintoComfyUI/models/diffusion_models/. - Load
workflow.json: H3RotNVFP4Loader (ckpt_name= the file), text encoder on CPU (CLIPLoader device=cpu), 8 steps. Log marker:[rot] single-file wrapped 200 linears (200 rotation-NVFP4 + 0 bf16-protected).
Credits & license
MiniMax-H3 (MiniMaxAI) · nunchaku / SVDQuant (NVFP4 W4A4 kernel) · ComfyUI. Derivative under the MiniMax-H3 Community License — commercial use with attribution; US/UK/EU/South Korea users apply to MiniMax; >$20M/yr needs separate authorization. Verify for your use.
Model tree for pottokao/MiniMax-H3-NVFP4-rotated
Base model
MiniMaxAI/MiniMax-H3

