File size: 12,450 Bytes
dd7a16d
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
166a35f
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
977526a
 
 
 
 
 
 
 
 
 
 
 
 
 
 
dd7a16d
 
 
 
 
166a35f
 
 
dd7a16d
 
 
 
 
 
 
 
 
 
166a35f
 
dd7a16d
 
166a35f
 
dd7a16d
 
 
 
 
 
 
 
 
 
 
 
 
166a35f
 
dd7a16d
 
 
 
 
 
 
 
 
 
 
 
 
 
166a35f
dd7a16d
 
 
 
 
 
 
 
 
 
 
 
166a35f
dd7a16d
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
---
license: other
license_name: minimax-h3-community-license
license_link: https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/LICENSE
base_model: MiniMaxAI/MiniMax-H3
library_name: mlx
pipeline_tag: text-to-video
tags:
- mlx
- apple-silicon
- text-to-video
- image-to-video
- audio-video-generation
- diffusion
---

# MiniMax-H3 MLX — 8-bit full pipeline (Apple Silicon)

MLX (Apple Silicon) build of the [**MiniMax-H3**](https://huggingface.co/MiniMaxAI/MiniMax-H3) 33B
joint video+audio diffusion transformer, packaged as a **self-contained pipeline** (DiT + text
encoder + video VAE + audio VAE + tokenizer + configs) for the
[PipeNetwork MLX port](https://github.com/PipeNetwork/minimax-h3-mlx).

> Powered by MiniMax H3.
>
> **These files are modified.** Quantized to 8-bit, text encoder rebuilt from the published
> `MiniMax-H3-FL2VA-MLX-Serve-8bit` checkpoint, tested and documented on an M4 Max. See
> `PATCHES.md` and the *Validation report* below.

## Runtime pin

This repository was built and validated against:

- Base: `PipeNetwork/minimax-h3-mlx` commit `b2f7e4d2` (2026-08-10)
- Patches (branch `video-lab-8bit-fixes`): `7210b93e6df86bf9c7206091c9542b6983c10c30` — see `PATCHES.md`
- MLX `0.32.0` (Python 3.12), mlx-vlm `0.6.13`

## Quick start

```bash
git clone https://github.com/PipeNetwork/minimax-h3-mlx
cd minimax-h3-mlx
git checkout b2f7e4d2   # base commit this repo was validated with
python3 -m venv .venv && .venv/bin/pip install -r requirements.txt

# Text-to-video+audio, validated settings (768x448, 16 sigma points -> 15 forwards)
.venv/bin/python scripts/generate.py "a red fox leaps over a mossy log in a misty forest" \
  -c <this repo> -t <this repo>/transformer -s 16 --seed 42 \
  --height 448 --width 768 -o fox.mp4

# Turbo LoRA (4-step EMA, included as turbo_lora_4step_ema.safetensors):
# merge it into a copy of the transformer with scripts/merge_h3_turbo.py,
# then run with -s 8 (see the Turbo LoRA section).
```

The pipeline needs only this repository: `-c` supplies the VAEs, text encoder and tokenizer;
`-t` points at the quantized DiT. Native canvas is a 768px short edge (e.g. 1344x768 16:9);
smaller canvases are off-distribution and degrade quickly.

## Image-to-Video (FL2VA)

This repository is **I2V-capable** (partition `fl2va`, tasks `t2va` + `fl2va` in `model_index.json`). The image is **optional** — present as a keyframe conditioning row via `--image` / `--anchor`, or omitted for pure T2V.

**Processor + vision tower required for I2V.** The `processor/` directory (7 files) and the 3 vision shards (`text_encoder/model-00005..00007-of-00008-vision.safetensors` + `vision_quant_config.json`, 529 `model.visual.*` tensors, 89 quantized) must be present. Without them the pipeline cannot encode images (falls back to T2V only).

**Canonical I2V command (validated):**

```bash
# 704×544 portrait (height 704 width 544) or 512×384 lightweight variant — both portrait 9:16 / 4:3
# Bridge image: inputs/bridge-20260813-181546-krea2_turbo-selfie-style-vertical-shot-a-young-woman-in-h.png
PROMPT=$(cat <<'PROMPT_EOF'
Vertical 9:16 TikTok-style UGC selfie video, handheld smartphone feel, natural indoor daylight near a window. A friendly creator speaks directly to camera with natural blinking, subtle head nods, and a warm smile. Add small human imperfections: a tiny hesitation, a soft breath, a quick smile mid-sentence, and a micro-pause before the last line. Realistic skin texture, stable identity, no face warping, minimal flicker, clean audio with natural room tone.

No subtitles. No on-screen text. No logos. No watermarks.

The creator says (exactly, with the same pacing and hesitations):
"Okay, entonces… eh… un datazo. Si estás trabado con tu código, solo da el primer pasito… como, abre tu terminal y escribe fran. (sonríe) Así de simple. Te vas a sorprender de lo rápido que todo se vuelve más fácil."
PROMPT_EOF
)

.venv/bin/python scripts/generate.py "$PROMPT" \
  -c <this repo> -t <this repo>/transformer -s 8 --seed 1996783985 \
  --height 704 --width 544 -d 5 \
  --image inputs/bridge-20260813-181546-krea2_turbo-selfie-style-vertical-shot-a-young-woman-in-h.png \
  --anchor first --release-encoder \
  -o i2v-704x544.mp4

# Lightweight alternative (faster, same identity):
.venv/bin/python scripts/generate.py "$PROMPT" \
  -c <this repo> -t <this repo>/transformer -s 8 --seed 1996783985 \
  --height 512 --width 384 -d 5 \
  --image inputs/bridge-20260813-181546-krea2_turbo-selfie-style-vertical-shot-a-young-woman-in-h.png \
  --anchor first --release-encoder \
  -o i2v-384x512.mp4
```

Valid `anchors` are `first` and `last` (FL2VA). Omit `--image`/`--anchor` for text-to-video. The image is consumed by `text_encoder` via `processor/` (`Qwen2VLImageProcessorPil` fallback when torch absent) and the vision tower, then patched as video-conditioning rows (see `minimal I2V encode` in `PATCHES.md`).

## Example outputs

![MiniMax-H3 MLX comparison — 8-bit vs turbo variants](images/comparative.png)

Red fox on a mossy log in a misty forest (prompt above, seed 42). Columns are time (0.4s → 3.0s);
rows: 8-bit 16-step, 8-bit + Turbo LoRA, 4-bit + Turbo LoRA at 1344x768.

Example clip (played inline by the Hub):

| File | Build | Settings |
|---|---|---|
| `examples/v3_768x448_16steps_8bit.mp4` | 8-bit (this repo) | 768x448, 16 steps, 33.9 min |

Turbo variants: see rows 2-3 above (a 1344x768 4-bit video re-render is planned).

## Model layout

| Path | Role | Representation |
|---|---|---|
| `transformer/transformer.safetensors` | 33B DiT (joint video+audio) | 8-bit affine, group 64 (AdaLN 8-bit), ~35.3 GB |
| `text_encoder/` (5+3 shards) | Qwen3-VL-32B conditioner (64 layers, truncated to layer 50) | 8-bit quantized, ~27.5 GB language (5 shards) + ~0.76 GB vision (3 shards `model-00005..00007-of-00008-vision.safetensors`, 529 `model.visual.*` tensors) |
| `text_encoder/vision_quant_config.json` | Vision tower quantization recipe (affine 8-bit g64) | json, 89 quantized tensors (`blocks.*`, `merger.*`, `deepstack_merger_list.*`) |
| `processor/` | Qwen3-VL processor (image preprocessing for I2V) | 7 files: `preprocessor_config.json`, `video_preprocessor_config.json`, `chat_template.json`, `tokenizer.json`, `tokenizer_config.json`, `vocab.json`, `merges.txt` (~11.6 MB) |
| `video_vae/` | Tiled causal video VAE (17-frame chunks, latents_mean/std) | fp16/bf16, 5.2 GB |
| `audio_vae/` | DAC encoder + BigVGAN vocoder, stereo 32 kHz | fp32, 0.6 GB |
| `tokenizer/` | Qwen3-VL tokenizer | json/vocab/merges |
| `model_index.json` | Pipeline config (fl2va partition, sigma shift video 12.0 / audio 3.0) | json |
| `turbo_lora_4step_ema.safetensors` | 4-step EMA Turbo LoRA (larryvrh, Apache-2.0) | 758 MB |
| `scripts/merge_h3_turbo.py` | Merge the Turbo LoRA into a transformer copy | python |

## Memory and speed (measured, M4 Max 68.7 GB)

- Resident during generation: DiT ~21.5 GB + text encoder ~22 GB + VAEs ~6 GB (AdaLN projections
  are precomputed and dropped, freeing ~13.8 GB). With `--release-encoder`, the text encoder is **freed after conditioning** (`prompt_embeds` already built, before the denoise loop), freeing ~27.5 GB (log: `released text encoder after conditioning (freeing ~27.5 GB)` + `mx.clear_cache()` + `tree_flatten` sizing). Peak RSS drops by ~22 GB resident; overhead <1 s.
- 768x448, 16 steps: **33.9 min** (119 s/step), peak RSS 24.7 GB, no swap. With turbo merge: ~24 min (8 steps). With **turbo8** (pure 8-bit, no merge) at 768x448, 8 steps (7 forwards): **~18.5 min (139 s/step)** vs turbo merge **251 s/step****~1.6–1.8× per-pixel** speedup (same canvas, fewer denoise forwards).
- 1344x768 (native 16:9), 8 steps: **~2 h** (996 s/step), peak ~57 GB + compressed memory; requires
  an otherwise idle machine (jetsam kills it under heavy ambient load).
- 704×544 I2V (portrait 9:16) turbo merge, 8 steps (7 forwards): **~33 min (251 s/step)**; turbo8: **~139 s/step** (measured 768x448 baseline, scales linearly per pixel). Lightweight 384×512 I2V is proportionally faster (off-distribution, useful for wiring checks).
- TeaCache: measured **0 skips / 7 forwards** with default last-block hook (`--teacache`, `thresh 0.2`, `start 3`, `compute_last_step`) at 768x448 turbo8 — probe feature pre-final-norm of last block (block 49/50) barely moves, so gate `rel_l1 <= 0.2` never fires; with `thresh 0.35` only 1/7 skips and still **no wall saving** (partial forward already traverses 49/50 blocks, only saves final norm + heads). Overhead +8–10% without skip. **Recommendation: keep TeaCache OFF by default**; if wall saving is needed, use an earlier hook `--teacache-layer 40` with `--teacache-thresh 0.25–0.35` (trades quality for ~20% block saving per skip). See `PATCHES.md` and `tmp/S9_teacache_spike_RESULT.md`.
- Attention is dense (MiniMax has not released sparse attention); quantization does not reduce the
  attention cost — it exists to fit memory.

## Validation report (2026-08-16, M4 Max 68.7 GB)

Prompt used everywhere: *"a red fox leaps over a mossy log in a misty forest"*, seed 42. Every clip
was judged by an independent vision-capable model on extracted frames (not statistics alone):

| Build | Settings | Time | Visual verdict |
|---|---|---|---|
| 8-bit (this repo) | 768x448, 16 steps | 33.9 min | **9/10** — coherent fox walking a mossy log; stable background; no melting |
| 8-bit + Turbo LoRA | 768x448, 8 steps | ~24 min | High confidence — "impressive for a 4-step turbo LoRA" |
| 4-bit + Turbo LoRA | 1344x768, 8 steps | 2 h 3 min | **95%** — full leap arc; no melting, no banding (frames; mp4 mux was a local script bug, not the model) |
| I2V E2E canonical (FL2VA, turbo merge) | 544×704 (704h×544w) 9:16, 8 steps (7 forwards), 5 s, seed 1996783985, `--image bridge-...png --anchor first --release-encoder` | ~33 min (251 s/step) | Keyframe anchored, identity preserved; requires `processor/` + vision tower (529 tensors); image optional (FL2VA otherwise uses key) |
| I2V Turbo8 (pure 8-bit, FL2VA) | 768×448, 8 steps (7 forwards), 5 s | ~18.5 min (139 s/step, ~1.6–1.8× per-pixel vs turbo merge) | Same I2V prompt/image, pure 8-bit DiT without LoRA merge; faster, same wiring |

Key findings:
- The VAE decoder is **not** the source of artifacts: parity with the reference diffusers
  implementation is exact (video 1.2e-6, audio 2.3e-8). The 17-frame chunked decode with 5-frame
  crossfade is the reference design; a periodic ~1.2x seam rhythm is architectural and masked at
  native resolution.
- **Canvas is the dominant quality factor**: the same VAE roundtrip scores 21.4 dB at 384x256 vs
  27.9 dB at 768x448. Keep the short edge at 768.
- First-generation failures (flat field, collapsed latents) were traced to an incompatible
  third-party serve-pack conversion (inert `condition_proj`, exploding text-encoder activations),
  not to the pipeline; see `PATCHES.md` and the report in `BIG_GOAL.md` of the source workspace.

## Turbo LoRA

`turbo_lora_4step_ema.safetensors` (also present as `turbo_lora.safetensors` — hardlink alias, same 779,849,816 bytes) is the 4-step EMA distilled LoRA from
[larryvrh/MiniMax-H3-Turbo-Lora](https://huggingface.co/larryvrh/MiniMax-H3-Turbo-Lora)
(Apache-2.0; sha256 `5a6eeba1…`). Merge it into a transformer copy:

```bash
.venv/bin/python scripts/merge_h3_turbo.py --base transformer --lora turbo_lora_4step_ema.safetensors -o transformer-turbo
# then generate with the merged dir as -t and -s 8 (7 forwards)
```

8 steps denoise cleanly with the merged weights; 4 steps can collapse to noise (see port issue #5).

## Patches vs upstream

`PATCHES.md` documents **five** patch families on branch `video-lab` over base commit `b2f7e4d2` (2026-08-10): (a) DiT QKV blocked layout, (b) text-encoder 8-bit + processor fallback + positional scatter fix, (c) `release_text_encoder` headroom, (d) two-phase TeaCache, (e) I2V processor + vision shards. See `PATCHES.md` for per-file details and line references. The remaining pipeline is otherwise byte-identical to the base commit.

## License

- **Weights**: MiniMax H3 Community License (included as `LICENSE`) — not open source; mark
  modified files, keep the attribution, territorial exclusions apply (EU/UK/KR/US excluded per the
  mlx-serve pack note). By downloading you accept those terms.
- **Port code**: Apache-2.0 (PipeNetwork/minimax-h3-mlx).
- **Turbo LoRA**: Apache-2.0 (larryvrh/MiniMax-H3-Turbo-Lora).