File size: 6,049 Bytes
831022f | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 | ---
language: en
tags:
- video-generation
- minimax-h3
- ref2va
- int8
- convrot
- turbo
- h3ddle
- pulpcut
license: other
license_name: minimax-h3-community-license
license_link: https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/939557dc319dd91227e30195a763f272ba7f8765/LICENSE
base_model: MiniMaxAI/MiniMax-H3
pretty_name: PulpCut MiniMax H3 Ref2VA Turbo INT8 ConvRot
---
# MiniMax H3 Ref2VA Turbo · pruned INT8 ConvRot (single file)
## What this repository is
A single-file MiniMax H3 **Ref2VA** diffusion transformer — the omni-reference
variant, which conditions generation on ordered reference images — with the
lightx2v **turbo step-distillation merged into the weights**, quantized in the
same pruned **INT8 ConvRot** layout as the Comfy-Org release. It is a drop-in
replacement for `minimax_h3_ref2va_pruned_int8_convrot.safetensors` in any
runtime that reads the optimized INT8 single-file layout — including
[H3ddle](https://github.com/AlexanderIstomin/h3ddle), the open-source native
macOS app it was built for.
This file is **not a standalone model**. It needs the rest of the optimized
package (Qwen3-VL-32B INT8 text encoder, video/audio VAEs, tokenizer) from
[Comfy-Org/MiniMax-H3](https://huggingface.co/Comfy-Org/MiniMax-H3), and the
FL2VA transformer alongside it if you also want prompt-only and keyframe
generation.
## Why this merge was made and republished
Every published turbo LoRA for MiniMax H3 targets the **FL2VA** transformer.
Reference generation therefore had no fast path at all: on Apple silicon a
short reference clip costs roughly fifty minutes at the base model's twenty
denoising passes, with no distilled alternative to fall back on. Since the two
transformers are structurally identical — 932 tensors, identical names and
shapes — the FL2VA distillation can be merged into the Ref2VA weights, which
is what this file is.
## Does the distillation survive the transfer?
The adapter was distilled against FL2VA, and reference-identity transfer is
exactly the capability step distillation is most likely to damage, so this was
measured rather than assumed. Matched control on an M1 Pro — same prompt, same
reference image, same seed, 448², adapter as the only variable:
| | passes | wall time | identity vs reference (SSIM) |
|---|---:|---:|---:|
| base Ref2VA | 20 | 49.0 min | 0.426 |
| this file | 8 | 20.2 min | 0.476 |
**2.4x faster with identity transfer intact.** The reference subject's
markings, build, and setting carry through; detail in fur and whiskers is
slightly sharper than the base at its full pass count. The SSIM difference is
modest and comes from a single prompt, so treat it as evidence that identity
transfer is *not degraded* rather than proof that it improves.
Recommended settings: 8 denoising passes, euler sampling, all 50 blocks.
## Known behavior and limitations
The pruned ComfyUI conversion of the turbo LoRA **drops all 51 AdaLN adapter
pairs** (the source targets AdaLN input dimension 2688, while pruned
"compact-curve" models use dimension 8), and its own metadata warns that
four-step distillation behaviour may therefore be degraded. We use 8 passes
rather than 4 for that reason. The same caveat applies to our FL2VA turbo
file; see
[ModelTC/Minimax-H3-Turbo#7](https://github.com/ModelTC/Minimax-H3-Turbo/issues/7).
## How the merge/quantization is done (high level)
For each of the 208 quantized projections, the BF16 pruned Ref2VA base weight
is merged with `strength × B·A` (rank-64, strength 1.0, `ema_pruned` variant),
rotated by the grouped 256-wide Hadamard transform used by the ConvRot runtime
kernels, and requantized with symmetric per-row absmax INT8 scales.
Token-refiner adapters merge losslessly in BF16. All other tensors are copied
byte-identical from the official INT8 file. The pipeline reproduces the
official quantizer exactly: run at strength 0 against the Ref2VA base it
regenerates the official file with all 3,046,400 scales identical and 1,711 of
19.27 billion int8 values differing (rounding ties).
## Source and attribution
- Original model: [MiniMaxAI/MiniMax-H3](https://huggingface.co/MiniMaxAI/MiniMax-H3)
- Pruned INT8 ConvRot base + shared package files: [Comfy-Org/MiniMax-H3](https://huggingface.co/Comfy-Org/MiniMax-H3)
- Turbo distillation LoRA: [ModelTC/Minimax-H3-Turbo](https://github.com/ModelTC/Minimax-H3-Turbo) (lightx2v team)
- Pruned ComfyUI LoRA conversion: [drbaph/MiniMax-H3-Turbo-Lora-ComfyUI](https://huggingface.co/drbaph/MiniMax-H3-Turbo-Lora-ComfyUI)
## Licensing
Derivative of MiniMax H3 weights; the
[MiniMax H3 Community License Agreement](https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/939557dc319dd91227e30195a763f272ba7f8765/LICENSE)
applies. By downloading you agree to its terms.
## What this file is used for in H3ddle
H3ddle installs it as a managed model alongside the reference-capable package:
the app verifies the SHA-256 below, reuses the shared package files it already
has via hardlinks, and defaults the model to 8 passes. Published by
[PulpCut](https://huggingface.co/PulpCut), whose editor family shares the
local-first media generation stack that H3ddle implements in the open.
## Safety and intended use
Intended for local, personal video generation. The merge changes speed
characteristics, not the base model's content behavior; all usage restrictions
of the MiniMax H3 Community License apply unchanged.
## File inventory
| File | Bytes | SHA-256 |
|---|---|---|
| `minimax_h3_ref2va_pruned_turbo_int8_convrot.safetensors` | 20,970,379,854 | `e64cef63bc2785bcd72e6103c52aa78c6cd2c4f9870a7ce79675083fd65cf2e7` |
## Reproducibility references
The conversion is a single dependency-free Python script,
[`Scripts/convert-turbo-package.py`](https://github.com/AlexanderIstomin/h3ddle/blob/main/Scripts/convert-turbo-package.py)
in the H3ddle repository, including the strength-0 self-check used to validate
the pipeline against the official file.
## Contact
Open an issue in the [H3ddle repository](https://github.com/AlexanderIstomin/h3ddle/issues).
|