MiniMax-H3 · NVFP4-rotated (single-file, 16 GB)

A rotated-NVFP4 quantization of the base MiniMax-H3 video + synchronized-audio DiT, as one self-contained ~12.5 GB file that runs fully resident on a single 16 GB Blackwell card, zero CPU offload, on the native nunchaku W4A4 kernel — at bf16-class sharpness. This is the quality / stability line (full 8-step). For a ~5×-faster 4-step version see FastH3-NVFP4-rotated.

Loads with our H3RotNVFP4Loader node — get it at H3-RotNVFP4-ComfyUI-Loader (why a custom node, full method, dependency versions there).

Why this line (vs FastH3)

Both lines are the same rotated-NVFP4 tech and both fit 16 GB resident. The difference is denoising steps:

  • This (base, 8-step): more steps → temporally stable and crisp even on the hard scenes — neon night, rain, water spray render clean, no flicker. Slower.
  • FastH3 (4-step): ~5× faster, but glow/motion scenes (neon/water) show temporal flicker inherent to the 4-step distillation.

If you care about clean neon/water/rain or maximum sharpness, use this line. If you want speed on static/portrait content, use FastH3.

Measured (1× RTX 5070 Ti 16 GB, warm, DiT-only)

  • Genuine 16 GB residency, zero offload — ~0.9 GB CPU offload, peak VRAM 13.6 GB, 5.7 s/step. System RAM ~31 GB to load (single file, no 80 GB shell floor).
  • 480p (832×480) full 124-frame clip ~66 s warm. Audio preserved (clean 32 kHz stereo).
  • Max clip length on a single 16 GB card (dynamic mode): 480p ~12.9 s (310 frames), 768p ~5.17 s (124 frames) — resolution-driven (same ceiling as the FastH3 line).
  • Only 8 steps for this quality line — vs the ~10 a stock-nvfp4 recipe typically needs. Per step it's on par with community quants on the same card (same nunchaku W4A4 kernel; rotation overhead offset). (The 4-step FastH3 line is faster still.)
  • File 12.5 GB (T1, all-NVFP4 blocks + curve-basis AdaLN + bf16 islands).

Why "rotated"?

The rotated in the name is a runtime block-256 Hadamard applied before the NVFP4 W4A4 GEMM (QuaRot-style — it spreads outliers so the 4-bit grid isn't dominated by a few large values). A clean same-weights check — identical weights and settings, only rotation toggled — confirms it earns its place: rotation keeps crisp edges where un-rotated NVFP4 hazes over (measurably sharper on every no-reference IQA metric, ~+70% on a Laplacian-variance sharpness proxy), at a cost of about 15% per step and a different per-seed draw. Un-rotated NVFP4 isn't broken — fine on easy scenes at enough steps — but where detail counts rotation is the sharper half. Full method + measurements: H3-RotNVFP4-ComfyUI-Loader.

Samples


neon rain street — clean reflections

sci-fi rooftop — flying cars

sailboat / sea

Full MP4 in samples/. These are 8-step renders — temporally clean and crisp even on neon/water, where the 4-step FastH3 line needs care. (GIFs are video-only previews.)

Setup

  1. Install the loader node + nunchaku — see H3-RotNVFP4-ComfyUI-Loader (verified torch 2.12.1+cu130 + a nunchaku cu13/torch2.12 wheel; Linux + Blackwell RTX 50-series / GB10).
  2. Download h3_base_T1.safetensors into ComfyUI/models/diffusion_models/.
  3. Load workflow.json: H3RotNVFP4Loader (ckpt_name = the file), text encoder on CPU (CLIPLoader device=cpu), 8 steps. Log marker: [rot] single-file wrapped 200 linears (200 rotation-NVFP4 + 0 bf16-protected).

Credits & license

MiniMax-H3 (MiniMaxAI) · nunchaku / SVDQuant (NVFP4 W4A4 kernel) · ComfyUI. Derivative under the MiniMax-H3 Community License — commercial use with attribution; US/UK/EU/South Korea users apply to MiniMax; >$20M/yr needs separate authorization. Verify for your use.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for pottokao/MiniMax-H3-NVFP4-rotated

Finetuned
(115)
this model