MiniMax-H3 DynTime sQKV Quants

Experimental: a ComfyUI core patch is required. These FL2VA and Ref2VA checkpoints retain the original FP32 runtime time MLP and physically separate Q, K, and V projections. They do not execute correctly in stock ComfyUI.

Community mixed-precision INT8 conversions of MiniMaxAI/MiniMax-H3. All 50 transformer blocks are retained. The repository is separate from the stock-compatible quants so patch-required files cannot be confused with ordinary ComfyUI checkpoints.

These are community derivatives, not official MiniMax or ComfyOrg releases.

Naming

  • FL2VA is text/first-frame/last-frame-to-audio-video generation.
  • Ref2VA is reference-image/video/audio-to-audio-video generation.
  • DT-sQKV means dynamic-time conditioning with physically separate Q, K, and V projections.
  • A filename without DT-sQKV belongs to the stock-compatible repository.
  • Exact INT8/BF16 inventories and GPU classes are documented here instead of being encoded in the filenames.

Choose a checkpoint

Profile Direct downloads File size GPU class and quant layout
DT-sQKV INT8 ConvRot FL2VA · Ref2VA 20.999 GiB 24 GB · RTX 30/40/50. 170 INT8 + 30 BF16 main semantic matrices; 270 physical INT8 modules; BF16 token refiner. Patch required.
DT-sQKV INT8 ConvRot HQ FL2VA · Ref2VA 27.994 GiB 32 GB+ · RTX 30/40/50. 74 INT8 + 126 BF16 main semantic matrices; 122 physical INT8 modules; BF16 token refiner. Patch required.

The 24/32 GB classes are capacity guidance, not full-workflow peak guarantees. Resolution, frame count, the text encoder, VAEs, attention backend, and offload settings also affect memory use. Moving a 21–28 GiB model across PCIe at every denoising step can make these editions impractical on 8–16 GB GPUs; use a stock-compatible W4/W8W4/NVFP4 checkpoint for those memory classes.

Required ComfyUI patch

Apply:

patches/ComfyUI-MiniMax-H3-DT-sQKV.patch

The patch changes MiniMax-H3 model detection, construction, loading, and forward execution. A loader merely accepting the checkpoint is not sufficient: the patched forward path must evaluate the original time MLP and the separate Q/K/V modules.

What DT-sQKV preserves

Feature Stock-compatible quants These DT-sQKV quants
Transformer blocks All 50 retained All 50 retained
Attention storage Fused qkv_proj = cat(Q,K,V) Physical q_proj, k_proj, v_proj tensors
Attention execution One fused projection call Three projection calls
Original time_embedder Replaced by measured table Retained in FP32
time_embedder.proj_in Absent FP32 weight [5376,256], bias [5376]
time_embedder.proj_out Absent FP32 weight [2688,5376], bias [2688]
adaln_t_table FP32 [4097,16] Absent
adaln_curve_basis Absent FP32 [2688,16]
adaln_curve_mean Absent FP32 [2688]
Per-block AdaLN 51 independent FP32 rank-16 projections 51 independent FP32 rank-16 projections
ComfyUI Stock Included core patch required

The original FP32 time path runs for every requested timestep:

256 -> 5,376 -> 2,688
full_t = SiLU(original_time_embedder(t))
coords = (full_t - mean) @ basis[2,688 x 16]
AdaLN_i(t) = independent_projection_i(coords)

The shared rank-16 basis removes redundant input width from the 51 large AdaLN projections. It does not replace the original time MLP and does not merge the per-block AdaLN layers. Full-time relative reconstruction error is about 3e-7; measured basis orthogonality residual is below 6e-7.

Separate Q/K/V layout

The original Diffusers checkpoints contain separate to_q, to_k, and to_v tensors. These files retain that layout through loading and execution:

  • 50 main transformer attention blocks;
  • 2 token-refiner attention blocks;
  • 156 physical Q/K/V weights;
  • three projection calls per attention block;
  • no fused qkv_proj modules.

Released Q/K/V tensors were checked bit-for-bit against their corresponding contiguous slices in the stock-compatible fused checkpoint.

Quantization profiles

Both profiles use ConvRot/Hadamard group size 256, deterministic scale search, and per-row FP32 scales for INT8 weights. Norms, patch projections, output heads, the time MLP, rank-16 basis, and all AdaLN projections retain source precision.

Profile Main semantic matrices Physical INT8 modules Token refiner Time/AdaLN path
DT-sQKV-INT8-ConvRot 170 INT8 + 30 BF16 270 BF16 Original FP32 time MLP, basis, mean, and 51 FP32 AdaLN projections
DT-sQKV-INT8-ConvRot-HQ 74 INT8 + 126 BF16 122 BF16 Original FP32 time MLP, basis, mean, and 51 FP32 AdaLN projections

The standard profile keeps 30 high-risk main matrices in BF16. The HQ profile keeps every attention-output projection and every MLP fc2 projection in BF16, together with the 26 highest-error QKV groups. All 50 HQ fc1 projections remain INT8.

Validation

Every checkpoint passed:

  1. exact key, shape, dtype, and quantization-inventory validation;
  2. bitwise Q/K/V split verification;
  3. dynamic-time reconstruction comparison;
  4. complete CPU load through patched ComfyUI as MiniMaxH3Model;
  5. remote LFS byte-size and SHA-256 verification.

Reports under reports/ retain their historical internal profile names so the validation provenance remains intact. Tests used clean ComfyUI commit 14b05228 plus the included patch. Future ComfyUI revisions may require the same small core changes to be forward-ported.

No full prompt-to-decoded-video perceptual A/B score is claimed. The 32 GB profile was structurally validated but is not claimed to remain fully resident on a 24 GB GPU.

Installation

  1. Use a ComfyUI revision compatible with the included patch.
  2. Apply patches/ComfyUI-MiniMax-H3-DT-sQKV.patch and restart ComfyUI.
  3. Place one selected checkpoint in ComfyUI/models/diffusion_models/.
  4. Use the matching FL2VA or Ref2VA workflow.

A complete workflow also requires the MiniMax-H3 Qwen3-VL text encoder and the video/audio VAEs from the stock-compatible repository. They are not duplicated here.

License and attribution

Use is subject to the included MiniMax-H3 community license. The base model is by MiniMax. ComfyUI and its quantization runtimes are separate upstream projects. This community conversion is not endorsed by MiniMax or ComfyOrg.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for DmitryDB/MiniMax-H3-DynTime-sQKV

Finetuned
(16)
this model