MiniMax-H3 Β· FastH3-NVFP4-rotated (single-file, 16 GB)
A rotated-NVFP4 quantization of FastH3 β FastVideo's 4-step, VSA-distilled MiniMax-H3 (video + synchronized audio) β as one self-contained ~12.8 GB file that runs fully resident on a single 16 GB Blackwell card, zero CPU offload, on the native nunchaku W4A4 kernel. This is the fast line (4-step, ~5Γ fewer steps). For the higher-quality 8-step line see MiniMax-H3-NVFP4-rotated.
Loads with our
H3RotNVFP4Loadernode β get it at H3-RotNVFP4-ComfyUI-Loader (why a custom node, full method, and dependency versions are there).
Measured (1Γ RTX 5070 Ti 16 GB, warm, DiT-only)
- Genuine 16 GB residency, zero offload β ~1.2 GB CPU offload (base-class, NOT the 26 GB an un-shrunk AdaLN needs), peak VRAM 13.7 GB.
- 480p (832Γ480): ~5.7 s/step Γ 4 β full 124-frame clip in ~25β30 s of denoise.
- Max clip length on a single 16 GB card (dynamic mode): 480p ~12.9 s (310 frames), 768p ~5.17 s (124 frames) β the ceiling is resolution-driven (same for the base line). VSA runs much faster per step but is shorter (480p ~2.8 s / 68 frames; 768p won't fit).
- Only 4 steps. FastH3 is 4-step distilled β vs the
10 a stock-nvfp4 recipe typically needs. Per step it's on par with community quants on the same card (rotation overhead offset by nunchaku's efficient W4A4 kernel), so end-to-end it's **1.9Γ faster**. - Audio preserved (clean 32 kHz stereo). File 12.8 GB (T1, all-NVFP4 blocks + rank-32 curve-basis AdaLN + bf16 islands).
Sampler β use Euler (important)
Run with the euler sampler + simple schedule (sigma shift 12), CFG 1.0, 4 steps β the settings FastH3 was distilled for. This is not our quantization: it's a standard sampler choice, and it matters at 4 steps.
eulerβ coherent, temporally stable video (recommended default; the includedworkflow.jsonuses it).res_multistep(a 2nd-order solver) looks a touch sharper but strobes on high-contrast neon / rain / water β the composition churns frame-to-frame. Avoid it here.- Want more sharpness/motion on the stable Euler base? Raise steps (6β8) β trades some speed. For maximum sharpness, the 8-step base line is crisper still.
Pick sampler/steps per scene and taste β the model is the same either way.
Why "rotated"?
The rotated in the name is a runtime block-256 Hadamard applied before the NVFP4 W4A4 GEMM (QuaRot-style β it spreads outliers so the 4-bit grid isn't dominated by a few large values). A clean same-weights check β identical weights and settings, only rotation toggled β confirms it earns its place: the rotated render keeps crisp edges and sharp wet-pavement reflections, while the un-rotated version hazes over, softening mid-ground detail and neon glyphs (measurably sharper on every no-reference IQA metric, ~+70% on a Laplacian-variance sharpness proxy). It isn't free β the runtime Hadamard costs about 15% per step and gives a different draw for a given seed. Un-rotated NVFP4 isn't broken (fine on easy scenes at enough steps), but where detail counts rotation is the sharper half β so we ship rotated. Full method + measurements: H3-RotNVFP4-ComfyUI-Loader.
Samples
![]() close-up face β eyelashes, iris, skin |
![]() forest waterfall β moss, spray |
![]() sailboat / sea β rigging, sun-glitter |
![]() sci-fi rooftop β flying cars, neon |
Full MP4 + synchronized audio in samples/. Rotated-NVFP4 vs fp8 vs bf16 quality comparison (frames + IQA table) in quality/ β same perceptual sharpness as fp8/bf16, but NVFP4 runs the native fp4 kernel (fp8 falls back to bf16 GEMM). (GIFs are video-only previews; use the euler sampler per above for coherent motion.)
Setup
- Install the loader node + nunchaku β see H3-RotNVFP4-ComfyUI-Loader (verified
torch 2.12.1+cu130+ a nunchaku cu13/torch2.12 wheel; Linux + Blackwell RTX 50-series / GB10). - Download
h3_fasth3_T1.safetensorsintoComfyUI/models/diffusion_models/. - Load
workflow.json: H3RotNVFP4Loader (ckpt_name= the file), text encoder on CPU (CLIPLoader device=cpu), 4 steps. Log marker:[rot] single-file wrapped 200 linears (200 rotation-NVFP4 + 0 bf16-protected).
Credits & license
MiniMax-H3 (MiniMaxAI) Β· FastVideo (FastH3 4-step + VSA distillation) Β· nunchaku / SVDQuant (NVFP4 W4A4 kernel) Β· ComfyUI. Derivative under the MiniMax-H3 Community License β commercial use with attribution; US/UK/EU/South Korea users apply to MiniMax; >$20M/yr needs separate authorization. Verify for your use.
Model tree for pottokao/MiniMax-H3-FastH3-NVFP4-rotated
Base model
MiniMaxAI/MiniMax-H3


