DmitryDB's picture
Update model card, filenames, and checksums
b660b69 verified
|
Raw
History Blame Contribute Delete
7.52 kB
---
license: other
license_name: minimax-h3-community-license-agreement
license_link: LICENSE
base_model: MiniMaxAI/MiniMax-H3
library_name: comfyui
pipeline_tag: image-text-to-video
tags:
- minimax-h3
- comfyui
- quantization
- int8
- video
- audio
- fl2va
- ref2va
- dynamic-time
- separate-qkv
- experimental
---
# MiniMax-H3 DynTime sQKV Quants
> **Experimental: a ComfyUI core patch is required.** These FL2VA and Ref2VA
> checkpoints retain the original FP32 runtime time MLP and physically separate
> Q, K, and V projections. They do not execute correctly in stock ComfyUI.
Community mixed-precision INT8 conversions of
[`MiniMaxAI/MiniMax-H3`](https://huggingface.co/MiniMaxAI/MiniMax-H3).
All 50 transformer blocks are retained. The repository is separate from the
[stock-compatible quants](https://huggingface.co/DmitryDB/MiniMax-H3-ComfyUI-Quants)
so patch-required files cannot be confused with ordinary ComfyUI checkpoints.
These are community derivatives, not official MiniMax or ComfyOrg releases.
## Naming
- `FL2VA` is text/first-frame/last-frame-to-audio-video generation.
- `Ref2VA` is reference-image/video/audio-to-audio-video generation.
- `DT-sQKV` means **dynamic-time conditioning with physically separate Q, K,
and V projections**.
- A filename without `DT-sQKV` belongs to the stock-compatible repository.
- Exact INT8/BF16 inventories and GPU classes are documented here instead of
being encoded in the filenames.
## Choose a checkpoint
| Profile | Direct downloads | File size | GPU class and quant layout |
|---|---|---:|---|
| **DT-sQKV INT8 ConvRot** | [FL2VA](https://huggingface.co/DmitryDB/MiniMax-H3-DynTime-sQKV/resolve/main/FL2VA/MiniMax-H3_FL2VA-DT-sQKV-INT8-ConvRot.safetensors?download=true) · [Ref2VA](https://huggingface.co/DmitryDB/MiniMax-H3-DynTime-sQKV/resolve/main/Ref2VA/MiniMax-H3_Ref2VA-DT-sQKV-INT8-ConvRot.safetensors?download=true) | 20.999 GiB | **24 GB · RTX 30/40/50.** 170 INT8 + 30 BF16 main semantic matrices; 270 physical INT8 modules; BF16 token refiner. Patch required. |
| **DT-sQKV INT8 ConvRot HQ** | [FL2VA](https://huggingface.co/DmitryDB/MiniMax-H3-DynTime-sQKV/resolve/main/FL2VA/MiniMax-H3_FL2VA-DT-sQKV-INT8-ConvRot-HQ.safetensors?download=true) · [Ref2VA](https://huggingface.co/DmitryDB/MiniMax-H3-DynTime-sQKV/resolve/main/Ref2VA/MiniMax-H3_Ref2VA-DT-sQKV-INT8-ConvRot-HQ.safetensors?download=true) | 27.994 GiB | **32 GB+ · RTX 30/40/50.** 74 INT8 + 126 BF16 main semantic matrices; 122 physical INT8 modules; BF16 token refiner. Patch required. |
The 24/32 GB classes are capacity guidance, not full-workflow peak guarantees.
Resolution, frame count, the text encoder, VAEs, attention backend, and offload
settings also affect memory use. Moving a 21–28 GiB model across PCIe at every
denoising step can make these editions impractical on 8–16 GB GPUs; use a
stock-compatible W4/W8W4/NVFP4 checkpoint for those memory classes.
## Required ComfyUI patch
Apply:
```text
patches/ComfyUI-MiniMax-H3-DT-sQKV.patch
```
The patch changes MiniMax-H3 model detection, construction, loading, and
forward execution. A loader merely accepting the checkpoint is not sufficient:
the patched forward path must evaluate the original time MLP and the separate
Q/K/V modules.
## What DT-sQKV preserves
| Feature | Stock-compatible quants | These DT-sQKV quants |
|---|---|---|
| Transformer blocks | All 50 retained | All 50 retained |
| Attention storage | Fused `qkv_proj = cat(Q,K,V)` | Physical `q_proj`, `k_proj`, `v_proj` tensors |
| Attention execution | One fused projection call | Three projection calls |
| Original `time_embedder` | Replaced by measured table | Retained in FP32 |
| `time_embedder.proj_in` | Absent | FP32 weight `[5376,256]`, bias `[5376]` |
| `time_embedder.proj_out` | Absent | FP32 weight `[2688,5376]`, bias `[2688]` |
| `adaln_t_table` | FP32 `[4097,16]` | Absent |
| `adaln_curve_basis` | Absent | FP32 `[2688,16]` |
| `adaln_curve_mean` | Absent | FP32 `[2688]` |
| Per-block AdaLN | 51 independent FP32 rank-16 projections | 51 independent FP32 rank-16 projections |
| ComfyUI | Stock | Included core patch required |
The original FP32 time path runs for every requested timestep:
```text
256 -> 5,376 -> 2,688
full_t = SiLU(original_time_embedder(t))
coords = (full_t - mean) @ basis[2,688 x 16]
AdaLN_i(t) = independent_projection_i(coords)
```
The shared rank-16 basis removes redundant input width from the 51 large AdaLN
projections. It does not replace the original time MLP and does not merge the
per-block AdaLN layers. Full-time relative reconstruction error is about
`3e-7`; measured basis orthogonality residual is below `6e-7`.
## Separate Q/K/V layout
The original Diffusers checkpoints contain separate `to_q`, `to_k`, and
`to_v` tensors. These files retain that layout through loading and execution:
- 50 main transformer attention blocks;
- 2 token-refiner attention blocks;
- 156 physical Q/K/V weights;
- three projection calls per attention block;
- no fused `qkv_proj` modules.
Released Q/K/V tensors were checked bit-for-bit against their corresponding
contiguous slices in the stock-compatible fused checkpoint.
## Quantization profiles
Both profiles use ConvRot/Hadamard group size 256, deterministic scale search,
and per-row FP32 scales for INT8 weights. Norms, patch projections, output
heads, the time MLP, rank-16 basis, and all AdaLN projections retain source
precision.
| Profile | Main semantic matrices | Physical INT8 modules | Token refiner | Time/AdaLN path |
|---|---|---:|---|---|
| `DT-sQKV-INT8-ConvRot` | 170 INT8 + 30 BF16 | 270 | BF16 | Original FP32 time MLP, basis, mean, and 51 FP32 AdaLN projections |
| `DT-sQKV-INT8-ConvRot-HQ` | 74 INT8 + 126 BF16 | 122 | BF16 | Original FP32 time MLP, basis, mean, and 51 FP32 AdaLN projections |
The standard profile keeps 30 high-risk main matrices in BF16. The HQ profile
keeps every attention-output projection and every MLP `fc2` projection in
BF16, together with the 26 highest-error QKV groups. All 50 HQ `fc1`
projections remain INT8.
## Validation
Every checkpoint passed:
1. exact key, shape, dtype, and quantization-inventory validation;
2. bitwise Q/K/V split verification;
3. dynamic-time reconstruction comparison;
4. complete CPU load through patched ComfyUI as `MiniMaxH3Model`;
5. remote LFS byte-size and SHA-256 verification.
Reports under `reports/` retain their historical internal profile names so the
validation provenance remains intact. Tests used clean ComfyUI commit
`14b05228` plus the included patch. Future ComfyUI revisions may require the
same small core changes to be forward-ported.
No full prompt-to-decoded-video perceptual A/B score is claimed. The 32 GB
profile was structurally validated but is not claimed to remain fully resident
on a 24 GB GPU.
## Installation
1. Use a ComfyUI revision compatible with the included patch.
2. Apply `patches/ComfyUI-MiniMax-H3-DT-sQKV.patch` and restart ComfyUI.
3. Place one selected checkpoint in `ComfyUI/models/diffusion_models/`.
4. Use the matching FL2VA or Ref2VA workflow.
A complete workflow also requires the MiniMax-H3 Qwen3-VL text encoder and the
video/audio VAEs from the stock-compatible repository. They are not duplicated
here.
## License and attribution
Use is subject to the included MiniMax-H3 community license. The base model is
by MiniMax. ComfyUI and its quantization runtimes are separate upstream
projects. This community conversion is not endorsed by MiniMax or ComfyOrg.