MiniMax-H3 Β· FastH3-NVFP4-rotated (single-file, 16 GB)

A rotated-NVFP4 quantization of FastH3 β€” FastVideo's 4-step, VSA-distilled MiniMax-H3 (video + synchronized audio) β€” as one self-contained ~12.8 GB file that runs fully resident on a single 16 GB Blackwell card, zero CPU offload, on the native nunchaku W4A4 kernel. This is the fast line (4-step, ~5Γ— fewer steps). For the higher-quality 8-step line see MiniMax-H3-NVFP4-rotated.

Loads with our H3RotNVFP4Loader node β€” get it at H3-RotNVFP4-ComfyUI-Loader (why a custom node, full method, and dependency versions are there).

Measured (1Γ— RTX 5070 Ti 16 GB, warm, DiT-only)

  • Genuine 16 GB residency, zero offload β€” ~1.2 GB CPU offload (base-class, NOT the 26 GB an un-shrunk AdaLN needs), peak VRAM 13.7 GB.
  • 480p (832Γ—480): ~5.7 s/step Γ— 4 β†’ full 124-frame clip in ~25–30 s of denoise.
  • Max clip length on a single 16 GB card (dynamic mode): 480p ~12.9 s (310 frames), 768p ~5.17 s (124 frames) β€” the ceiling is resolution-driven (same for the base line). VSA runs much faster per step but is shorter (480p ~2.8 s / 68 frames; 768p won't fit).
  • Only 4 steps. FastH3 is 4-step distilled β€” vs the 10 a stock-nvfp4 recipe typically needs. Per step it's on par with community quants on the same card (rotation overhead offset by nunchaku's efficient W4A4 kernel), so end-to-end it's **1.9Γ— faster**.
  • Audio preserved (clean 32 kHz stereo). File 12.8 GB (T1, all-NVFP4 blocks + rank-32 curve-basis AdaLN + bf16 islands).

Sampler β€” use Euler (important)

Run with the euler sampler + simple schedule (sigma shift 12), CFG 1.0, 4 steps β€” the settings FastH3 was distilled for. This is not our quantization: it's a standard sampler choice, and it matters at 4 steps.

  • euler β†’ coherent, temporally stable video (recommended default; the included workflow.json uses it).
  • res_multistep (a 2nd-order solver) looks a touch sharper but strobes on high-contrast neon / rain / water β€” the composition churns frame-to-frame. Avoid it here.
  • Want more sharpness/motion on the stable Euler base? Raise steps (6–8) β€” trades some speed. For maximum sharpness, the 8-step base line is crisper still.

Pick sampler/steps per scene and taste β€” the model is the same either way.

Why "rotated"?

The rotated in the name is a runtime block-256 Hadamard applied before the NVFP4 W4A4 GEMM (QuaRot-style β€” it spreads outliers so the 4-bit grid isn't dominated by a few large values). A clean same-weights check β€” identical weights and settings, only rotation toggled β€” confirms it earns its place: the rotated render keeps crisp edges and sharp wet-pavement reflections, while the un-rotated version hazes over, softening mid-ground detail and neon glyphs (measurably sharper on every no-reference IQA metric, ~+70% on a Laplacian-variance sharpness proxy). It isn't free β€” the runtime Hadamard costs about 15% per step and gives a different draw for a given seed. Un-rotated NVFP4 isn't broken (fine on easy scenes at enough steps), but where detail counts rotation is the sharper half β€” so we ship rotated. Full method + measurements: H3-RotNVFP4-ComfyUI-Loader.

Samples


close-up face β€” eyelashes, iris, skin

forest waterfall β€” moss, spray

sailboat / sea β€” rigging, sun-glitter

sci-fi rooftop β€” flying cars, neon

Full MP4 + synchronized audio in samples/. Rotated-NVFP4 vs fp8 vs bf16 quality comparison (frames + IQA table) in quality/ β€” same perceptual sharpness as fp8/bf16, but NVFP4 runs the native fp4 kernel (fp8 falls back to bf16 GEMM). (GIFs are video-only previews; use the euler sampler per above for coherent motion.)

Setup

  1. Install the loader node + nunchaku β€” see H3-RotNVFP4-ComfyUI-Loader (verified torch 2.12.1+cu130 + a nunchaku cu13/torch2.12 wheel; Linux + Blackwell RTX 50-series / GB10).
  2. Download h3_fasth3_T1.safetensors into ComfyUI/models/diffusion_models/.
  3. Load workflow.json: H3RotNVFP4Loader (ckpt_name = the file), text encoder on CPU (CLIPLoader device=cpu), 4 steps. Log marker: [rot] single-file wrapped 200 linears (200 rotation-NVFP4 + 0 bf16-protected).

Credits & license

MiniMax-H3 (MiniMaxAI) Β· FastVideo (FastH3 4-step + VSA distillation) Β· nunchaku / SVDQuant (NVFP4 W4A4 kernel) Β· ComfyUI. Derivative under the MiniMax-H3 Community License β€” commercial use with attribution; US/UK/EU/South Korea users apply to MiniMax; >$20M/yr needs separate authorization. Verify for your use.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for pottokao/MiniMax-H3-FastH3-NVFP4-rotated

Finetuned
(115)
this model