MiniMax-H3-NVFP4 / README.md
lilcheaty's picture
Upload README.md with huggingface_hub
d0b6e4f verified
|
Raw
History Blame Contribute Delete
9.17 kB
metadata
license: other
license_name: minimax-h3-community-license-agreement
license_link: https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/LICENSE
tags:
  - comfyui
  - nvfp4
  - quantized
  - video
  - text-to-video
base_model: MiniMaxAI/MiniMax-H3
base_model_relation: quantized

MiniMax H3 β€” NVFP4

NVFP4 quantizations of the MiniMax-H3 ref2va diffusion transformer for ComfyUI.

NVFP4 requires an NVIDIA Blackwell GPU (RTX 50-series, RTX PRO 6000, B200). On Ada, Hopper or older the NVFP4 path is emulated β€” use Comfy-Org's int8_convrot files instead.

Which file do I want?

file size s/it VRAM (DiT) notes
minimax_h3_ref2va_pruned_nvfp4.safetensors 12.5 GB 1.90 11.9 GB recommended
minimax_h3_ref2va_nvfp4_mixed.safetensors 24.4 GB 1.92 ~20 GB from unpruned bf16
minimax_h3_ref2va_nvfp4_full.safetensors 18.7 GB 1.91 ~16 GB experimental

Take pruned_nvfp4 unless you have a specific reason not to: half the size of the alternatives at identical speed, and it leaves the modulation path at full precision.

Why the pruned base is the right one to quantize

Comfy-Org's pruned checkpoint is not lossily pruned β€” it is a structural refactor of AdaLN, and understanding it explains the whole table above.

In the bf16 model, AdaLN modulation dominates the parameter count:

group bf16 pruned
adaln_proj 13.04B (39.4%) 0.04B (0.2%)
mlp 12.02B 11.56B
attn 8.02B 7.71B
token_refiner / norms / embedders 0.05B 0.80B
total 33.12B 20.11B

The bf16 model projects a 5376-dim conditioning vector into modulation parameters per block. The pruned model replaces this with an 8-dim timestep table (adaln_t_table, shape [1025, 8]) feeding adaln_proj.linear of shape [96768, 8]. Because modulation depends only on the timestep, that 5376-wide projection was almost entirely redundant β€” 13.04B parameters collapse to 0.04B, a ~326x reduction.

This matters for quantization because AdaLN is the part you least want to quantize: it emits the scale and shift applied to every residual stream, so error there is multiplicative and compounds across all 50 blocks and every sampling step. In the bf16 model you face a bad choice β€” protect AdaLN and produce a ~36 GB file (larger than the 34 GB int8 it should beat), or quantize 39% of the model and hope. In the pruned model the problem disappears: AdaLN is already tiny, so you keep it at full precision for free and quantize only attn+mlp, which are error-tolerant.

Comfy-Org's pruned_int8_convrot quantizes exactly those 200 attn/mlp layers to int8_convrot and leaves everything else alone. pruned_nvfp4 takes that same set to NVFP4.

Measured

RTX PRO 6000 Blackwell (96 GB), ComfyUI 0.30.0, ref2va, 864x480, 39 frames, 20 steps, res_multistep / beta, three matched seeds:

model size staged VRAM s/it
pruned_int8_convrot (Comfy-Org) 21.0 GB 19,995 MB 2.17
pruned_nvfp4 (this repo) 12.5 GB 11,944 MB 1.90

-40% file size, -8.0 GB VRAM, -12.4% sampling time.

At ~12 GB for the DiT, a 32 GB card (RTX 5090) becomes viable if the text encoder is offloaded to CPU after encoding β€” it runs once per prompt, not once per sampling step.

How pruned_nvfp4 was built

The obvious tool does not work. The StarNodes model converter passes non-floating-point tensors through untouched:

if not tensor.dtype.is_floating_point:
    return tensor          # already-int8 weights are copied verbatim

so running it on an int8 checkpoint silently produces a byte-identical file. A real dequantize -> requantize is required. comfy-kitchen exposes both halves:

params = ckt.TensorWiseINT8Layout.Params(
    scale=weight_scale, orig_dtype=torch.bfloat16, orig_shape=qdata.shape,
    is_weight=True, convrot=True, convrot_groupsize=256)
deq = ckt.TensorWiseINT8Layout.dequantize(qdata, params)              # -> bf16
nq, nparams = ckt.TensorCoreNVFP4Layout.quantize(deq.contiguous())    # -> NVFP4
tensors = ckt.TensorCoreNVFP4Layout.state_dict_tensors(nq, nparams)

Per-layer config lives in a comfy_quant uint8 tensor holding JSON, e.g. {"format": "int8_tensorwise", "convrot": true, "convrot_groupsize": 256}, rewritten to {"format": "nvfp4"} on output. Note quantize() requires bf16/fp16 β€” float32 raises Unsupported dtype code.

Full script: pruned_to_nvfp4.py in this repo. 200 layers, ~6 seconds on one GPU.

Prompting: H3 wants a structured IR, not prose

Read this before blaming the weights for bad output. H3 was trained on the structured output of H3-Context-IR, a preprocessing model that rewrites a plain request into labelled sections; MiniMax's model card calls it "critical to the quality of the final output". ComfyUI passes your raw string straight to the DiT, so you must write that structure yourself.

Official guides: base Β· ref

For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.

integrated_multimodal_description: [Shot 1] Live-action, cinematic, the young woman shown in
<Picture 1> remains beside the rain-covered train window, preserving her appearance and the
carriage layout. The camera trucks right with small amplitude at slow speed as she lifts her
gaze toward the passing city lights. The quiet, breathy young woman (S1) says:
<d>[English] I get off at the next station.</d> She folds the letter along its existing crease.

overall_soundscape: The train wheels produce a steady metallic rhythm beneath a low
ventilation hum. Rain ticks against the window while paper rustles softly in her hands.

non_diegetic_music: Sustained cello notes at a slow tempo with widely spaced piano tones.

Dialogue must be explicit or you get gibberish. Speech is generated jointly with video, so saying that someone speaks without giving the words yields correct prosody and mouth shapes with no lexical content. Speaker identity, action and delivery go outside <d>; only the language tag and verbatim words go inside. Use stable IDs (S1), (S2), and (S1,S2) for simultaneous speech.

Other essentials: [Shot 1] carries no timestamp, later shots use [Shot N] At MM:SS.mmm; aim for 350-500 words of description; write camera motion as type + amplitude + speed; reference tags must appear in the order the inputs were connected. ref2va accepts up to 9 reference images, and 3-4 varied shots hold identity far better than one.

Companion files (mirrored, not ours)

file precision origin
text_encoders/qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors NVFP4-AWQ Comfy-Org, unmodified
vae/minimax_h3_video_vae_fp16.safetensors FP16 Comfy-Org, unmodified
vae/minimax_h3_audio_vae_fp32.safetensors FP32 Comfy-Org, unmodified

The NVFP4 text encoder is Comfy-Org's work, not ours. Only the minimax_h3_ref2va_*nvfp4*.safetensors diffusion models here are new.

VAEs are deliberately not quantized: they are small, run once per generation rather than per step, and decode straight to pixels and audio samples where error is immediately visible. The text encoder being NVFP4 buys VRAM, not speed β€” it also runs once per prompt.

Usage

πŸ“‚ ComfyUI/models/
β”œβ”€β”€ πŸ“‚ diffusion_models/
β”‚   └── minimax_h3_ref2va_pruned_nvfp4.safetensors
β”œβ”€β”€ πŸ“‚ text_encoders/
β”‚   └── qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors
└── πŸ“‚ vae/
    β”œβ”€β”€ minimax_h3_video_vae_fp16.safetensors
    └── minimax_h3_audio_vae_fp32.safetensors

Recommended stack total: 33.7 GB. Use the official R2V template, swapping the diffusion model. Requires ComfyUI >= 0.30.0 (native H3 support in comfy/ldm/minimax/). CLIPLoader type must be minimax; sampler res_multistep; frame length must satisfy 17n+5.

Honest limitations

  • Quality was compared against pruned_int8_convrot at three matched seeds with no visible degradation, but this is not a rigorous evaluation β€” no FVD, no human study, no long-duration or 2K testing.
  • pruned_nvfp4 is doubly quantized (bf16 -> int8_convrot by Comfy-Org -> NVFP4 here). Error from both passes compounds. It held up in testing, but that is a real caveat.
  • Only ref2va is converted. fl2va is not included.
  • Benchmarks are single-GPU, one card, one resolution.

Failure cases are welcome in the discussions tab β€” concrete artifacts beat aggregate scores.

License

Inherits the MiniMax-H3 Community License Agreement from the original model.