--- license: other license_name: minimax-h3-community-license-agreement license_link: LICENSE base_model: MiniMaxAI/MiniMax-H3 library_name: comfyui pipeline_tag: image-text-to-video tags: - minimax-h3 - comfyui - quantization - int8 - w4 - nvfp4 - video - audio - fl2va - ref2va --- # MiniMax-H3 Stock ComfyUI Quants Community FL2VA and Ref2VA diffusion-transformer checkpoints for [`MiniMaxAI/MiniMax-H3`](https://huggingface.co/MiniMaxAI/MiniMax-H3). All files in this repository retain all 50 transformer blocks and use the stock ComfyUI fused-QKV and time-table layout. No custom node or ComfyUI core patch is required. These are community conversions, not official MiniMax or ComfyOrg releases. ## Naming - No runtime marker in the filename means **stock ComfyUI compatible**. - `FL2VA` is text/first-frame/last-frame-to-audio-video generation. - `Ref2VA` is reference-image/video/audio-to-audio-video generation. - Quantized tensor counts, retained BF16 islands, GPU class, and expected memory class are documented here instead of being encoded in filenames. - The patch-required dynamic-time, separate-QKV editions use the explicit `DT-sQKV` marker and live in the separate [MiniMax-H3-DynTime-sQKV](https://huggingface.co/DmitryDB/MiniMax-H3-DynTime-sQKV) repository. ## Choose a checkpoint Download one FL2VA or Ref2VA checkpoint from the same profile row. | Profile | Direct downloads | File size | GPU class and quant layout | |---|---|---:|---| | **INT8 ConvRot HQ** | [FL2VA](https://huggingface.co/DmitryDB/MiniMax-H3-ComfyUI-Quants/resolve/main/FL2VA/MiniMax-H3_FL2VA-INT8-ConvRot-HQ.safetensors?download=true) · [Ref2VA](https://huggingface.co/DmitryDB/MiniMax-H3-ComfyUI-Quants/resolve/main/Ref2VA/MiniMax-H3_Ref2VA-INT8-ConvRot-HQ.safetensors?download=true) | 21.908 GiB | **32 GB+ · RTX 30/40.** 145 INT8 ConvRot + 63 BF16 semantic matrices. Largest BF16 island. A 24 GB RTX 4090 loader test offloaded about 0.955 GiB. | | **INT8 ConvRot** | [FL2VA](https://huggingface.co/DmitryDB/MiniMax-H3-ComfyUI-Quants/resolve/main/FL2VA/MiniMax-H3_FL2VA-INT8-ConvRot.safetensors?download=true) · [Ref2VA](https://huggingface.co/DmitryDB/MiniMax-H3-ComfyUI-Quants/resolve/main/Ref2VA/MiniMax-H3_Ref2VA-INT8-ConvRot.safetensors?download=true) | 20.940 GiB | **24 GB · RTX 30/40.** 170 INT8 ConvRot + 38 BF16 semantic matrices. Fully resident in the RTX 4090 loader test. | | **INT8 ConvRot Lite** | [FL2VA](https://huggingface.co/DmitryDB/MiniMax-H3-ComfyUI-Quants/resolve/main/FL2VA/MiniMax-H3_FL2VA-INT8-ConvRot-Lite.safetensors?download=true) · [Ref2VA](https://huggingface.co/DmitryDB/MiniMax-H3-ComfyUI-Quants/resolve/main/Ref2VA/MiniMax-H3_Ref2VA-INT8-ConvRot-Lite.safetensors?download=true) | 20.330 GiB | **24 GB · RTX 30/40.** 185 INT8 ConvRot + 23 BF16 semantic matrices. Leaves more memory for the rest of the workflow. | | **W8/W4 ConvRot** | [FL2VA](https://huggingface.co/DmitryDB/MiniMax-H3-ComfyUI-Quants/resolve/main/FL2VA/MiniMax-H3_FL2VA-W8W4-ConvRot.safetensors?download=true) · [Ref2VA](https://huggingface.co/DmitryDB/MiniMax-H3-ComfyUI-Quants/resolve/main/Ref2VA/MiniMax-H3_Ref2VA-W8W4-ConvRot.safetensors?download=true) | 13.565 GiB | **16 GB · RTX 30/40.** 86 W8 + 114 W4 main matrices; the eight token-refiner matrices remain BF16. | | **W4 ConvRot** | [FL2VA](https://huggingface.co/DmitryDB/MiniMax-H3-ComfyUI-Quants/resolve/main/FL2VA/MiniMax-H3_FL2VA-W4-ConvRot.safetensors?download=true) · [Ref2VA](https://huggingface.co/DmitryDB/MiniMax-H3-ComfyUI-Quants/resolve/main/Ref2VA/MiniMax-H3_Ref2VA-W4-ConvRot.safetensors?download=true) | 10.067 GiB | **12 GB · RTX 30/40.** 200 W4 main matrices + 8 INT8 token-refiner matrices. | | **W4 ConvRot Offload** | [FL2VA](https://huggingface.co/DmitryDB/MiniMax-H3-ComfyUI-Quants/resolve/main/FL2VA/MiniMax-H3_FL2VA-W4-ConvRot-Offload.safetensors?download=true) · [Ref2VA](https://huggingface.co/DmitryDB/MiniMax-H3-ComfyUI-Quants/resolve/main/Ref2VA/MiniMax-H3_Ref2VA-W4-ConvRot-Offload.safetensors?download=true) | 9.708 GiB | **8 GB + CPU offload · RTX 30/40.** All 208 main and token-refiner matrices use W4. | | **NVFP4 HQ** | [FL2VA](https://huggingface.co/DmitryDB/MiniMax-H3-ComfyUI-Quants/resolve/main/FL2VA/MiniMax-H3_FL2VA-NVFP4-HQ.safetensors?download=true) · [Ref2VA](https://huggingface.co/DmitryDB/MiniMax-H3-ComfyUI-Quants/resolve/main/Ref2VA/MiniMax-H3_Ref2VA-NVFP4-HQ.safetensors?download=true) | 13.597 GiB | **16–24 GB · RTX 50/Blackwell.** 170 NVFP4 + 30 BF16 main matrices; the eight token-refiner matrices remain BF16. | | **NVFP4** | [FL2VA](https://huggingface.co/DmitryDB/MiniMax-H3-ComfyUI-Quants/resolve/main/FL2VA/MiniMax-H3_FL2VA-NVFP4.safetensors?download=true) · [Ref2VA](https://huggingface.co/DmitryDB/MiniMax-H3-ComfyUI-Quants/resolve/main/Ref2VA/MiniMax-H3_Ref2VA-NVFP4.safetensors?download=true) | 10.862 GiB | **8–12 GB · RTX 50/Blackwell.** All 208 main and token-refiner matrices use block-scaled NVFP4. | Checkpoint size is not full-workflow peak VRAM. Resolution, frame count, attention backend, text encoder, VAEs, and ComfyUI offload settings also affect memory use. RTX 50 classifications are architecture-based; no full generation run was performed on an RTX 5090. NVFP4 here is block-scaled NVFP4, not AWQ. ## Measured RTX 4090 loader results FL2VA and Ref2VA were tested independently through stock ComfyUI. Each test also executed a real quantized INT8 projection. | Checkpoint | Loaded weights | Peak reserved | Free after load | Result | |---|---:|---:|---:|---| | `MiniMax-H3_*VA-INT8-ConvRot-Lite.safetensors` | 100% | 20.424 GiB | 2.072 GiB | PASS | | `MiniMax-H3_*VA-INT8-ConvRot.safetensors` | 100% | 21.025 GiB | 1.471 GiB | PASS | | `MiniMax-H3_*VA-INT8-ConvRot-HQ.safetensors` | 95.6% | 21.002 GiB | 1.494 GiB | PASS; about 0.955 GiB offloaded | These are loader/kernel measurements, not complete prompt-to-decoded-video VRAM peaks. ## Quantization and preserved components All 16 diffusion checkpoints: - retain all 50 transformer blocks; - use the fused `qkv_proj = cat(Q,K,V)` layout expected by stock ComfyUI; - use a rank-16 FP32, 4,097-point time table; - retain 51 independent FP32 AdaLN projections; - keep norms, conditioning projections, patch projections, output heads, and other small or sensitive tensors in source precision; - load without a custom loader or core patch in tested ComfyUI commit `14b05228`. INT8, W8, and W4 weights use ConvRot/Hadamard rotation with group size 256, per-row FP32 scales, and deterministic scale search. | INT8 profile | BF16 attention-output blocks | BF16 MLP `fc2` blocks | Token refiner | |---|---|---|---| | `INT8-ConvRot-Lite` | 0, 1, 2, 3, 4, 5, 6, 7, 9, 15, 19, 38, 45, 49 | 49 | Eight BF16 matrices | | `INT8-ConvRot` | 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 17, 19, 20, 27, 38, 43, 44, 45, 46, 47, 49 | 39, 45, 49 | Eight BF16 matrices | | `INT8-ConvRot-HQ` | all blocks 0–49 | 29, 39, 44, 45, 49 | Eight BF16 matrices | ## Stock layout versus DT-sQKV | Feature | This repository | DT-sQKV repository | |---|---|---| | Attention storage | Fused `qkv_proj` | Separate `q_proj`, `k_proj`, `v_proj` | | Attention execution | One fused call | Three projection calls | | Original FP32 `time_embedder` | Replaced by measured time table | Retained and evaluated at runtime | | `adaln_t_table` | FP32 `[4097,16]` | Absent | | Per-block AdaLN | 51 independent FP32 rank-16 projections | 51 independent FP32 rank-16 projections | | ComfyUI | Stock | Core patch required | The time table does not remove timestep conditioning. It interpolates a compact representation of the original measured time curve. Maximum measured table interpolation error is below `0.001%`; sampled end-to-end AdaLN relative error is approximately `3e-7` to `4e-7` across 19 timesteps. ## Validation Every released checkpoint passed: 1. exact key, shape, dtype, and quantization-inventory checks; 2. sampled reconstruction against its original FL2VA or Ref2VA HF shards; 3. a 19-timestep FP32 AdaLN numerical comparison; 4. complete CPU load as `MiniMaxH3Model` in clean ComfyUI commit `14b05228`; 5. remote byte-size and LFS SHA-256 verification. Reports under `reports/` retain their historical internal profile names so the published validation provenance remains intact. BF16 samples were checked bit-for-bit. A representative INT8 QKV sample has relative L2 error `0.008814`. A prompt-to-decoded-video perceptual A/B score has not been measured. ## Installation and required components Place one selected FL2VA or Ref2VA checkpoint in: ```text ComfyUI/models/diffusion_models/ ``` A complete workflow also needs the separately maintained Qwen3-VL MiniMax-H3 text encoder and these shared VAEs: | File | Role | |---|---| | `vae/MiniMax-H3_VideoVAE-FP16.safetensors` | Video latent encoder and decoder | | `vae/MiniMax-H3_AudioVAE-FP32.safetensors` | Audio latent encoder and decoder | No text encoder is included in this repository. ## License and attribution Use is subject to the included MiniMax-H3 community license. The base model is by MiniMax. This community conversion is not endorsed by MiniMax or ComfyOrg.