File size: 7,524 Bytes
2d2fe97 b660b69 2d2fe97 b660b69 2d2fe97 b660b69 2d2fe97 b660b69 2d2fe97 b660b69 2d2fe97 b660b69 2d2fe97 b660b69 2d2fe97 b660b69 2d2fe97 b660b69 2d2fe97 b660b69 2d2fe97 b660b69 2d2fe97 b660b69 2d2fe97 c2fcf27 b660b69 c2fcf27 2d2fe97 b660b69 2d2fe97 b660b69 2d2fe97 b660b69 c2fcf27 b660b69 c2fcf27 b660b69 c2fcf27 b660b69 2d2fe97 b660b69 2d2fe97 c2fcf27 2d2fe97 c2fcf27 b660b69 2d2fe97 b660b69 2d2fe97 c2fcf27 b660b69 2d2fe97 c2fcf27 b660b69 c2fcf27 2d2fe97 b660b69 2d2fe97 b660b69 2d2fe97 c2fcf27 b660b69 c2fcf27 b660b69 c2fcf27 b660b69 c2fcf27 b660b69 c2fcf27 b660b69 c2fcf27 b660b69 c2fcf27 b660b69 2d2fe97 c2fcf27 2d2fe97 c2fcf27 b660b69 c2fcf27 2d2fe97 b660b69 2d2fe97 c2fcf27 2d2fe97 c2fcf27 b660b69 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 | ---
license: other
license_name: minimax-h3-community-license-agreement
license_link: LICENSE
base_model: MiniMaxAI/MiniMax-H3
library_name: comfyui
pipeline_tag: image-text-to-video
tags:
- minimax-h3
- comfyui
- quantization
- int8
- video
- audio
- fl2va
- ref2va
- dynamic-time
- separate-qkv
- experimental
---
# MiniMax-H3 DynTime sQKV Quants
> **Experimental: a ComfyUI core patch is required.** These FL2VA and Ref2VA
> checkpoints retain the original FP32 runtime time MLP and physically separate
> Q, K, and V projections. They do not execute correctly in stock ComfyUI.
Community mixed-precision INT8 conversions of
[`MiniMaxAI/MiniMax-H3`](https://huggingface.co/MiniMaxAI/MiniMax-H3).
All 50 transformer blocks are retained. The repository is separate from the
[stock-compatible quants](https://huggingface.co/DmitryDB/MiniMax-H3-ComfyUI-Quants)
so patch-required files cannot be confused with ordinary ComfyUI checkpoints.
These are community derivatives, not official MiniMax or ComfyOrg releases.
## Naming
- `FL2VA` is text/first-frame/last-frame-to-audio-video generation.
- `Ref2VA` is reference-image/video/audio-to-audio-video generation.
- `DT-sQKV` means **dynamic-time conditioning with physically separate Q, K,
and V projections**.
- A filename without `DT-sQKV` belongs to the stock-compatible repository.
- Exact INT8/BF16 inventories and GPU classes are documented here instead of
being encoded in the filenames.
## Choose a checkpoint
| Profile | Direct downloads | File size | GPU class and quant layout |
|---|---|---:|---|
| **DT-sQKV INT8 ConvRot** | [FL2VA](https://huggingface.co/DmitryDB/MiniMax-H3-DynTime-sQKV/resolve/main/FL2VA/MiniMax-H3_FL2VA-DT-sQKV-INT8-ConvRot.safetensors?download=true) · [Ref2VA](https://huggingface.co/DmitryDB/MiniMax-H3-DynTime-sQKV/resolve/main/Ref2VA/MiniMax-H3_Ref2VA-DT-sQKV-INT8-ConvRot.safetensors?download=true) | 20.999 GiB | **24 GB · RTX 30/40/50.** 170 INT8 + 30 BF16 main semantic matrices; 270 physical INT8 modules; BF16 token refiner. Patch required. |
| **DT-sQKV INT8 ConvRot HQ** | [FL2VA](https://huggingface.co/DmitryDB/MiniMax-H3-DynTime-sQKV/resolve/main/FL2VA/MiniMax-H3_FL2VA-DT-sQKV-INT8-ConvRot-HQ.safetensors?download=true) · [Ref2VA](https://huggingface.co/DmitryDB/MiniMax-H3-DynTime-sQKV/resolve/main/Ref2VA/MiniMax-H3_Ref2VA-DT-sQKV-INT8-ConvRot-HQ.safetensors?download=true) | 27.994 GiB | **32 GB+ · RTX 30/40/50.** 74 INT8 + 126 BF16 main semantic matrices; 122 physical INT8 modules; BF16 token refiner. Patch required. |
The 24/32 GB classes are capacity guidance, not full-workflow peak guarantees.
Resolution, frame count, the text encoder, VAEs, attention backend, and offload
settings also affect memory use. Moving a 21–28 GiB model across PCIe at every
denoising step can make these editions impractical on 8–16 GB GPUs; use a
stock-compatible W4/W8W4/NVFP4 checkpoint for those memory classes.
## Required ComfyUI patch
Apply:
```text
patches/ComfyUI-MiniMax-H3-DT-sQKV.patch
```
The patch changes MiniMax-H3 model detection, construction, loading, and
forward execution. A loader merely accepting the checkpoint is not sufficient:
the patched forward path must evaluate the original time MLP and the separate
Q/K/V modules.
## What DT-sQKV preserves
| Feature | Stock-compatible quants | These DT-sQKV quants |
|---|---|---|
| Transformer blocks | All 50 retained | All 50 retained |
| Attention storage | Fused `qkv_proj = cat(Q,K,V)` | Physical `q_proj`, `k_proj`, `v_proj` tensors |
| Attention execution | One fused projection call | Three projection calls |
| Original `time_embedder` | Replaced by measured table | Retained in FP32 |
| `time_embedder.proj_in` | Absent | FP32 weight `[5376,256]`, bias `[5376]` |
| `time_embedder.proj_out` | Absent | FP32 weight `[2688,5376]`, bias `[2688]` |
| `adaln_t_table` | FP32 `[4097,16]` | Absent |
| `adaln_curve_basis` | Absent | FP32 `[2688,16]` |
| `adaln_curve_mean` | Absent | FP32 `[2688]` |
| Per-block AdaLN | 51 independent FP32 rank-16 projections | 51 independent FP32 rank-16 projections |
| ComfyUI | Stock | Included core patch required |
The original FP32 time path runs for every requested timestep:
```text
256 -> 5,376 -> 2,688
full_t = SiLU(original_time_embedder(t))
coords = (full_t - mean) @ basis[2,688 x 16]
AdaLN_i(t) = independent_projection_i(coords)
```
The shared rank-16 basis removes redundant input width from the 51 large AdaLN
projections. It does not replace the original time MLP and does not merge the
per-block AdaLN layers. Full-time relative reconstruction error is about
`3e-7`; measured basis orthogonality residual is below `6e-7`.
## Separate Q/K/V layout
The original Diffusers checkpoints contain separate `to_q`, `to_k`, and
`to_v` tensors. These files retain that layout through loading and execution:
- 50 main transformer attention blocks;
- 2 token-refiner attention blocks;
- 156 physical Q/K/V weights;
- three projection calls per attention block;
- no fused `qkv_proj` modules.
Released Q/K/V tensors were checked bit-for-bit against their corresponding
contiguous slices in the stock-compatible fused checkpoint.
## Quantization profiles
Both profiles use ConvRot/Hadamard group size 256, deterministic scale search,
and per-row FP32 scales for INT8 weights. Norms, patch projections, output
heads, the time MLP, rank-16 basis, and all AdaLN projections retain source
precision.
| Profile | Main semantic matrices | Physical INT8 modules | Token refiner | Time/AdaLN path |
|---|---|---:|---|---|
| `DT-sQKV-INT8-ConvRot` | 170 INT8 + 30 BF16 | 270 | BF16 | Original FP32 time MLP, basis, mean, and 51 FP32 AdaLN projections |
| `DT-sQKV-INT8-ConvRot-HQ` | 74 INT8 + 126 BF16 | 122 | BF16 | Original FP32 time MLP, basis, mean, and 51 FP32 AdaLN projections |
The standard profile keeps 30 high-risk main matrices in BF16. The HQ profile
keeps every attention-output projection and every MLP `fc2` projection in
BF16, together with the 26 highest-error QKV groups. All 50 HQ `fc1`
projections remain INT8.
## Validation
Every checkpoint passed:
1. exact key, shape, dtype, and quantization-inventory validation;
2. bitwise Q/K/V split verification;
3. dynamic-time reconstruction comparison;
4. complete CPU load through patched ComfyUI as `MiniMaxH3Model`;
5. remote LFS byte-size and SHA-256 verification.
Reports under `reports/` retain their historical internal profile names so the
validation provenance remains intact. Tests used clean ComfyUI commit
`14b05228` plus the included patch. Future ComfyUI revisions may require the
same small core changes to be forward-ported.
No full prompt-to-decoded-video perceptual A/B score is claimed. The 32 GB
profile was structurally validated but is not claimed to remain fully resident
on a 24 GB GPU.
## Installation
1. Use a ComfyUI revision compatible with the included patch.
2. Apply `patches/ComfyUI-MiniMax-H3-DT-sQKV.patch` and restart ComfyUI.
3. Place one selected checkpoint in `ComfyUI/models/diffusion_models/`.
4. Use the matching FL2VA or Ref2VA workflow.
A complete workflow also requires the MiniMax-H3 Qwen3-VL text encoder and the
video/audio VAEs from the stock-compatible repository. They are not duplicated
here.
## License and attribution
Use is subject to the included MiniMax-H3 community license. The base model is
by MiniMax. ComfyUI and its quantization runtimes are separate upstream
projects. This community conversion is not endorsed by MiniMax or ComfyOrg.
|