DmitryDB's picture
Update model card, filenames, and checksums
adc7972 verified
|
Raw
History Blame Contribute Delete
9.18 kB
---
license: other
license_name: minimax-h3-community-license-agreement
license_link: LICENSE
base_model: MiniMaxAI/MiniMax-H3
library_name: comfyui
pipeline_tag: image-text-to-video
tags:
- minimax-h3
- comfyui
- quantization
- int8
- w4
- nvfp4
- video
- audio
- fl2va
- ref2va
---
# MiniMax-H3 Stock ComfyUI Quants
Community FL2VA and Ref2VA diffusion-transformer checkpoints for
[`MiniMaxAI/MiniMax-H3`](https://huggingface.co/MiniMaxAI/MiniMax-H3).
All files in this repository retain all 50 transformer blocks and use the
stock ComfyUI fused-QKV and time-table layout. No custom node or ComfyUI core
patch is required.
These are community conversions, not official MiniMax or ComfyOrg releases.
## Naming
- No runtime marker in the filename means **stock ComfyUI compatible**.
- `FL2VA` is text/first-frame/last-frame-to-audio-video generation.
- `Ref2VA` is reference-image/video/audio-to-audio-video generation.
- Quantized tensor counts, retained BF16 islands, GPU class, and expected
memory class are documented here instead of being encoded in filenames.
- The patch-required dynamic-time, separate-QKV editions use the explicit
`DT-sQKV` marker and live in the separate
[MiniMax-H3-DynTime-sQKV](https://huggingface.co/DmitryDB/MiniMax-H3-DynTime-sQKV)
repository.
## Choose a checkpoint
Download one FL2VA or Ref2VA checkpoint from the same profile row.
| Profile | Direct downloads | File size | GPU class and quant layout |
|---|---|---:|---|
| **INT8 ConvRot HQ** | [FL2VA](https://huggingface.co/DmitryDB/MiniMax-H3-ComfyUI-Quants/resolve/main/FL2VA/MiniMax-H3_FL2VA-INT8-ConvRot-HQ.safetensors?download=true) · [Ref2VA](https://huggingface.co/DmitryDB/MiniMax-H3-ComfyUI-Quants/resolve/main/Ref2VA/MiniMax-H3_Ref2VA-INT8-ConvRot-HQ.safetensors?download=true) | 21.908 GiB | **32 GB+ · RTX 30/40.** 145 INT8 ConvRot + 63 BF16 semantic matrices. Largest BF16 island. A 24 GB RTX 4090 loader test offloaded about 0.955 GiB. |
| **INT8 ConvRot** | [FL2VA](https://huggingface.co/DmitryDB/MiniMax-H3-ComfyUI-Quants/resolve/main/FL2VA/MiniMax-H3_FL2VA-INT8-ConvRot.safetensors?download=true) · [Ref2VA](https://huggingface.co/DmitryDB/MiniMax-H3-ComfyUI-Quants/resolve/main/Ref2VA/MiniMax-H3_Ref2VA-INT8-ConvRot.safetensors?download=true) | 20.940 GiB | **24 GB · RTX 30/40.** 170 INT8 ConvRot + 38 BF16 semantic matrices. Fully resident in the RTX 4090 loader test. |
| **INT8 ConvRot Lite** | [FL2VA](https://huggingface.co/DmitryDB/MiniMax-H3-ComfyUI-Quants/resolve/main/FL2VA/MiniMax-H3_FL2VA-INT8-ConvRot-Lite.safetensors?download=true) · [Ref2VA](https://huggingface.co/DmitryDB/MiniMax-H3-ComfyUI-Quants/resolve/main/Ref2VA/MiniMax-H3_Ref2VA-INT8-ConvRot-Lite.safetensors?download=true) | 20.330 GiB | **24 GB · RTX 30/40.** 185 INT8 ConvRot + 23 BF16 semantic matrices. Leaves more memory for the rest of the workflow. |
| **W8/W4 ConvRot** | [FL2VA](https://huggingface.co/DmitryDB/MiniMax-H3-ComfyUI-Quants/resolve/main/FL2VA/MiniMax-H3_FL2VA-W8W4-ConvRot.safetensors?download=true) · [Ref2VA](https://huggingface.co/DmitryDB/MiniMax-H3-ComfyUI-Quants/resolve/main/Ref2VA/MiniMax-H3_Ref2VA-W8W4-ConvRot.safetensors?download=true) | 13.565 GiB | **16 GB · RTX 30/40.** 86 W8 + 114 W4 main matrices; the eight token-refiner matrices remain BF16. |
| **W4 ConvRot** | [FL2VA](https://huggingface.co/DmitryDB/MiniMax-H3-ComfyUI-Quants/resolve/main/FL2VA/MiniMax-H3_FL2VA-W4-ConvRot.safetensors?download=true) · [Ref2VA](https://huggingface.co/DmitryDB/MiniMax-H3-ComfyUI-Quants/resolve/main/Ref2VA/MiniMax-H3_Ref2VA-W4-ConvRot.safetensors?download=true) | 10.067 GiB | **12 GB · RTX 30/40.** 200 W4 main matrices + 8 INT8 token-refiner matrices. |
| **W4 ConvRot Offload** | [FL2VA](https://huggingface.co/DmitryDB/MiniMax-H3-ComfyUI-Quants/resolve/main/FL2VA/MiniMax-H3_FL2VA-W4-ConvRot-Offload.safetensors?download=true) · [Ref2VA](https://huggingface.co/DmitryDB/MiniMax-H3-ComfyUI-Quants/resolve/main/Ref2VA/MiniMax-H3_Ref2VA-W4-ConvRot-Offload.safetensors?download=true) | 9.708 GiB | **8 GB + CPU offload · RTX 30/40.** All 208 main and token-refiner matrices use W4. |
| **NVFP4 HQ** | [FL2VA](https://huggingface.co/DmitryDB/MiniMax-H3-ComfyUI-Quants/resolve/main/FL2VA/MiniMax-H3_FL2VA-NVFP4-HQ.safetensors?download=true) · [Ref2VA](https://huggingface.co/DmitryDB/MiniMax-H3-ComfyUI-Quants/resolve/main/Ref2VA/MiniMax-H3_Ref2VA-NVFP4-HQ.safetensors?download=true) | 13.597 GiB | **16–24 GB · RTX 50/Blackwell.** 170 NVFP4 + 30 BF16 main matrices; the eight token-refiner matrices remain BF16. |
| **NVFP4** | [FL2VA](https://huggingface.co/DmitryDB/MiniMax-H3-ComfyUI-Quants/resolve/main/FL2VA/MiniMax-H3_FL2VA-NVFP4.safetensors?download=true) · [Ref2VA](https://huggingface.co/DmitryDB/MiniMax-H3-ComfyUI-Quants/resolve/main/Ref2VA/MiniMax-H3_Ref2VA-NVFP4.safetensors?download=true) | 10.862 GiB | **8–12 GB · RTX 50/Blackwell.** All 208 main and token-refiner matrices use block-scaled NVFP4. |
Checkpoint size is not full-workflow peak VRAM. Resolution, frame count,
attention backend, text encoder, VAEs, and ComfyUI offload settings also affect
memory use. RTX 50 classifications are architecture-based; no full generation
run was performed on an RTX 5090. NVFP4 here is block-scaled NVFP4, not AWQ.
## Measured RTX 4090 loader results
FL2VA and Ref2VA were tested independently through stock ComfyUI. Each test
also executed a real quantized INT8 projection.
| Checkpoint | Loaded weights | Peak reserved | Free after load | Result |
|---|---:|---:|---:|---|
| `MiniMax-H3_*VA-INT8-ConvRot-Lite.safetensors` | 100% | 20.424 GiB | 2.072 GiB | PASS |
| `MiniMax-H3_*VA-INT8-ConvRot.safetensors` | 100% | 21.025 GiB | 1.471 GiB | PASS |
| `MiniMax-H3_*VA-INT8-ConvRot-HQ.safetensors` | 95.6% | 21.002 GiB | 1.494 GiB | PASS; about 0.955 GiB offloaded |
These are loader/kernel measurements, not complete prompt-to-decoded-video
VRAM peaks.
## Quantization and preserved components
All 16 diffusion checkpoints:
- retain all 50 transformer blocks;
- use the fused `qkv_proj = cat(Q,K,V)` layout expected by stock ComfyUI;
- use a rank-16 FP32, 4,097-point time table;
- retain 51 independent FP32 AdaLN projections;
- keep norms, conditioning projections, patch projections, output heads, and
other small or sensitive tensors in source precision;
- load without a custom loader or core patch in tested ComfyUI commit
`14b05228`.
INT8, W8, and W4 weights use ConvRot/Hadamard rotation with group size 256,
per-row FP32 scales, and deterministic scale search.
| INT8 profile | BF16 attention-output blocks | BF16 MLP `fc2` blocks | Token refiner |
|---|---|---|---|
| `INT8-ConvRot-Lite` | 0, 1, 2, 3, 4, 5, 6, 7, 9, 15, 19, 38, 45, 49 | 49 | Eight BF16 matrices |
| `INT8-ConvRot` | 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 17, 19, 20, 27, 38, 43, 44, 45, 46, 47, 49 | 39, 45, 49 | Eight BF16 matrices |
| `INT8-ConvRot-HQ` | all blocks 0–49 | 29, 39, 44, 45, 49 | Eight BF16 matrices |
## Stock layout versus DT-sQKV
| Feature | This repository | DT-sQKV repository |
|---|---|---|
| Attention storage | Fused `qkv_proj` | Separate `q_proj`, `k_proj`, `v_proj` |
| Attention execution | One fused call | Three projection calls |
| Original FP32 `time_embedder` | Replaced by measured time table | Retained and evaluated at runtime |
| `adaln_t_table` | FP32 `[4097,16]` | Absent |
| Per-block AdaLN | 51 independent FP32 rank-16 projections | 51 independent FP32 rank-16 projections |
| ComfyUI | Stock | Core patch required |
The time table does not remove timestep conditioning. It interpolates a compact
representation of the original measured time curve. Maximum measured table
interpolation error is below `0.001%`; sampled end-to-end AdaLN relative error
is approximately `3e-7` to `4e-7` across 19 timesteps.
## Validation
Every released checkpoint passed:
1. exact key, shape, dtype, and quantization-inventory checks;
2. sampled reconstruction against its original FL2VA or Ref2VA HF shards;
3. a 19-timestep FP32 AdaLN numerical comparison;
4. complete CPU load as `MiniMaxH3Model` in clean ComfyUI commit `14b05228`;
5. remote byte-size and LFS SHA-256 verification.
Reports under `reports/` retain their historical internal profile names so the
published validation provenance remains intact. BF16 samples were checked
bit-for-bit. A representative INT8 QKV sample has relative L2 error
`0.008814`. A prompt-to-decoded-video perceptual A/B score has not been
measured.
## Installation and required components
Place one selected FL2VA or Ref2VA checkpoint in:
```text
ComfyUI/models/diffusion_models/
```
A complete workflow also needs the separately maintained Qwen3-VL MiniMax-H3
text encoder and these shared VAEs:
| File | Role |
|---|---|
| `vae/MiniMax-H3_VideoVAE-FP16.safetensors` | Video latent encoder and decoder |
| `vae/MiniMax-H3_AudioVAE-FP32.safetensors` | Audio latent encoder and decoder |
No text encoder is included in this repository.
## License and attribution
Use is subject to the included MiniMax-H3 community license. The base model is
by MiniMax. This community conversion is not endorsed by MiniMax or ComfyOrg.