File size: 7,524 Bytes
2d2fe97
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
b660b69
 
2d2fe97
 
 
b660b69
2d2fe97
b660b69
 
 
2d2fe97
b660b69
 
 
 
 
2d2fe97
b660b69
2d2fe97
b660b69
2d2fe97
b660b69
 
 
 
 
 
 
2d2fe97
b660b69
2d2fe97
b660b69
 
 
 
2d2fe97
b660b69
 
 
 
 
2d2fe97
b660b69
2d2fe97
b660b69
2d2fe97
c2fcf27
b660b69
c2fcf27
2d2fe97
b660b69
 
 
 
2d2fe97
b660b69
2d2fe97
b660b69
c2fcf27
b660b69
 
c2fcf27
b660b69
 
 
c2fcf27
 
 
 
b660b69
2d2fe97
b660b69
2d2fe97
c2fcf27
 
 
 
 
 
2d2fe97
c2fcf27
b660b69
 
 
2d2fe97
b660b69
2d2fe97
c2fcf27
b660b69
2d2fe97
c2fcf27
 
b660b69
c2fcf27
 
2d2fe97
b660b69
 
 
 
2d2fe97
b660b69
 
 
 
2d2fe97
c2fcf27
 
b660b69
 
c2fcf27
b660b69
 
 
 
c2fcf27
 
 
b660b69
c2fcf27
b660b69
c2fcf27
 
b660b69
 
c2fcf27
b660b69
 
 
 
c2fcf27
b660b69
 
 
2d2fe97
c2fcf27
2d2fe97
c2fcf27
b660b69
c2fcf27
 
2d2fe97
b660b69
 
 
2d2fe97
c2fcf27
2d2fe97
c2fcf27
 
b660b69
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
---
license: other
license_name: minimax-h3-community-license-agreement
license_link: LICENSE
base_model: MiniMaxAI/MiniMax-H3
library_name: comfyui
pipeline_tag: image-text-to-video
tags:
  - minimax-h3
  - comfyui
  - quantization
  - int8
  - video
  - audio
  - fl2va
  - ref2va
  - dynamic-time
  - separate-qkv
  - experimental
---

# MiniMax-H3 DynTime sQKV Quants

> **Experimental: a ComfyUI core patch is required.** These FL2VA and Ref2VA
> checkpoints retain the original FP32 runtime time MLP and physically separate
> Q, K, and V projections. They do not execute correctly in stock ComfyUI.

Community mixed-precision INT8 conversions of
[`MiniMaxAI/MiniMax-H3`](https://huggingface.co/MiniMaxAI/MiniMax-H3).
All 50 transformer blocks are retained. The repository is separate from the
[stock-compatible quants](https://huggingface.co/DmitryDB/MiniMax-H3-ComfyUI-Quants)
so patch-required files cannot be confused with ordinary ComfyUI checkpoints.

These are community derivatives, not official MiniMax or ComfyOrg releases.

## Naming

- `FL2VA` is text/first-frame/last-frame-to-audio-video generation.
- `Ref2VA` is reference-image/video/audio-to-audio-video generation.
- `DT-sQKV` means **dynamic-time conditioning with physically separate Q, K,
  and V projections**.
- A filename without `DT-sQKV` belongs to the stock-compatible repository.
- Exact INT8/BF16 inventories and GPU classes are documented here instead of
  being encoded in the filenames.

## Choose a checkpoint

| Profile | Direct downloads | File size | GPU class and quant layout |
|---|---|---:|---|
| **DT-sQKV INT8 ConvRot** | [FL2VA](https://huggingface.co/DmitryDB/MiniMax-H3-DynTime-sQKV/resolve/main/FL2VA/MiniMax-H3_FL2VA-DT-sQKV-INT8-ConvRot.safetensors?download=true) · [Ref2VA](https://huggingface.co/DmitryDB/MiniMax-H3-DynTime-sQKV/resolve/main/Ref2VA/MiniMax-H3_Ref2VA-DT-sQKV-INT8-ConvRot.safetensors?download=true) | 20.999 GiB | **24 GB · RTX 30/40/50.** 170 INT8 + 30 BF16 main semantic matrices; 270 physical INT8 modules; BF16 token refiner. Patch required. |
| **DT-sQKV INT8 ConvRot HQ** | [FL2VA](https://huggingface.co/DmitryDB/MiniMax-H3-DynTime-sQKV/resolve/main/FL2VA/MiniMax-H3_FL2VA-DT-sQKV-INT8-ConvRot-HQ.safetensors?download=true) · [Ref2VA](https://huggingface.co/DmitryDB/MiniMax-H3-DynTime-sQKV/resolve/main/Ref2VA/MiniMax-H3_Ref2VA-DT-sQKV-INT8-ConvRot-HQ.safetensors?download=true) | 27.994 GiB | **32 GB+ · RTX 30/40/50.** 74 INT8 + 126 BF16 main semantic matrices; 122 physical INT8 modules; BF16 token refiner. Patch required. |

The 24/32 GB classes are capacity guidance, not full-workflow peak guarantees.
Resolution, frame count, the text encoder, VAEs, attention backend, and offload
settings also affect memory use. Moving a 21–28 GiB model across PCIe at every
denoising step can make these editions impractical on 8–16 GB GPUs; use a
stock-compatible W4/W8W4/NVFP4 checkpoint for those memory classes.

## Required ComfyUI patch

Apply:

```text
patches/ComfyUI-MiniMax-H3-DT-sQKV.patch
```

The patch changes MiniMax-H3 model detection, construction, loading, and
forward execution. A loader merely accepting the checkpoint is not sufficient:
the patched forward path must evaluate the original time MLP and the separate
Q/K/V modules.

## What DT-sQKV preserves

| Feature | Stock-compatible quants | These DT-sQKV quants |
|---|---|---|
| Transformer blocks | All 50 retained | All 50 retained |
| Attention storage | Fused `qkv_proj = cat(Q,K,V)` | Physical `q_proj`, `k_proj`, `v_proj` tensors |
| Attention execution | One fused projection call | Three projection calls |
| Original `time_embedder` | Replaced by measured table | Retained in FP32 |
| `time_embedder.proj_in` | Absent | FP32 weight `[5376,256]`, bias `[5376]` |
| `time_embedder.proj_out` | Absent | FP32 weight `[2688,5376]`, bias `[2688]` |
| `adaln_t_table` | FP32 `[4097,16]` | Absent |
| `adaln_curve_basis` | Absent | FP32 `[2688,16]` |
| `adaln_curve_mean` | Absent | FP32 `[2688]` |
| Per-block AdaLN | 51 independent FP32 rank-16 projections | 51 independent FP32 rank-16 projections |
| ComfyUI | Stock | Included core patch required |

The original FP32 time path runs for every requested timestep:

```text
256 -> 5,376 -> 2,688
full_t = SiLU(original_time_embedder(t))
coords = (full_t - mean) @ basis[2,688 x 16]
AdaLN_i(t) = independent_projection_i(coords)
```

The shared rank-16 basis removes redundant input width from the 51 large AdaLN
projections. It does not replace the original time MLP and does not merge the
per-block AdaLN layers. Full-time relative reconstruction error is about
`3e-7`; measured basis orthogonality residual is below `6e-7`.

## Separate Q/K/V layout

The original Diffusers checkpoints contain separate `to_q`, `to_k`, and
`to_v` tensors. These files retain that layout through loading and execution:

- 50 main transformer attention blocks;
- 2 token-refiner attention blocks;
- 156 physical Q/K/V weights;
- three projection calls per attention block;
- no fused `qkv_proj` modules.

Released Q/K/V tensors were checked bit-for-bit against their corresponding
contiguous slices in the stock-compatible fused checkpoint.

## Quantization profiles

Both profiles use ConvRot/Hadamard group size 256, deterministic scale search,
and per-row FP32 scales for INT8 weights. Norms, patch projections, output
heads, the time MLP, rank-16 basis, and all AdaLN projections retain source
precision.

| Profile | Main semantic matrices | Physical INT8 modules | Token refiner | Time/AdaLN path |
|---|---|---:|---|---|
| `DT-sQKV-INT8-ConvRot` | 170 INT8 + 30 BF16 | 270 | BF16 | Original FP32 time MLP, basis, mean, and 51 FP32 AdaLN projections |
| `DT-sQKV-INT8-ConvRot-HQ` | 74 INT8 + 126 BF16 | 122 | BF16 | Original FP32 time MLP, basis, mean, and 51 FP32 AdaLN projections |

The standard profile keeps 30 high-risk main matrices in BF16. The HQ profile
keeps every attention-output projection and every MLP `fc2` projection in
BF16, together with the 26 highest-error QKV groups. All 50 HQ `fc1`
projections remain INT8.

## Validation

Every checkpoint passed:

1. exact key, shape, dtype, and quantization-inventory validation;
2. bitwise Q/K/V split verification;
3. dynamic-time reconstruction comparison;
4. complete CPU load through patched ComfyUI as `MiniMaxH3Model`;
5. remote LFS byte-size and SHA-256 verification.

Reports under `reports/` retain their historical internal profile names so the
validation provenance remains intact. Tests used clean ComfyUI commit
`14b05228` plus the included patch. Future ComfyUI revisions may require the
same small core changes to be forward-ported.

No full prompt-to-decoded-video perceptual A/B score is claimed. The 32 GB
profile was structurally validated but is not claimed to remain fully resident
on a 24 GB GPU.

## Installation

1. Use a ComfyUI revision compatible with the included patch.
2. Apply `patches/ComfyUI-MiniMax-H3-DT-sQKV.patch` and restart ComfyUI.
3. Place one selected checkpoint in `ComfyUI/models/diffusion_models/`.
4. Use the matching FL2VA or Ref2VA workflow.

A complete workflow also requires the MiniMax-H3 Qwen3-VL text encoder and the
video/audio VAEs from the stock-compatible repository. They are not duplicated
here.

## License and attribution

Use is subject to the included MiniMax-H3 community license. The base model is
by MiniMax. ComfyUI and its quantization runtimes are separate upstream
projects. This community conversion is not endorsed by MiniMax or ComfyOrg.