--- license: other license_name: minimax-h3-community-license license_link: https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/LICENSE base_model: MiniMaxAI/MiniMax-H3 library_name: mlx pipeline_tag: text-to-video tags: - mlx - apple-silicon - text-to-video - image-to-video - audio-video-generation - diffusion --- # MiniMax-H3 MLX — 8-bit full pipeline (Apple Silicon) MLX (Apple Silicon) build of the [**MiniMax-H3**](https://huggingface.co/MiniMaxAI/MiniMax-H3) 33B joint video+audio diffusion transformer, packaged as a **self-contained pipeline** (DiT + text encoder + video VAE + audio VAE + tokenizer + configs) for the [PipeNetwork MLX port](https://github.com/PipeNetwork/minimax-h3-mlx). > Powered by MiniMax H3. > > **These files are modified.** Quantized to 8-bit, text encoder rebuilt from the published > `MiniMax-H3-FL2VA-MLX-Serve-8bit` checkpoint, tested and documented on an M4 Max. See > `PATCHES.md` and the *Validation report* below. ## Runtime pin This repository was built and validated against: - Base: `PipeNetwork/minimax-h3-mlx` commit `b2f7e4d2` (2026-08-10) - Patches (branch `video-lab-8bit-fixes`): `7210b93e6df86bf9c7206091c9542b6983c10c30` — see `PATCHES.md` - MLX `0.32.0` (Python 3.12), mlx-vlm `0.6.13` ## Quick start ```bash git clone https://github.com/PipeNetwork/minimax-h3-mlx cd minimax-h3-mlx git checkout b2f7e4d2 # base commit this repo was validated with python3 -m venv .venv && .venv/bin/pip install -r requirements.txt # Text-to-video+audio, validated settings (768x448, 16 sigma points -> 15 forwards) .venv/bin/python scripts/generate.py "a red fox leaps over a mossy log in a misty forest" \ -c -t /transformer -s 16 --seed 42 \ --height 448 --width 768 -o fox.mp4 # Turbo LoRA (4-step EMA, included as turbo_lora_4step_ema.safetensors): # merge it into a copy of the transformer with scripts/merge_h3_turbo.py, # then run with -s 8 (see the Turbo LoRA section). ``` The pipeline needs only this repository: `-c` supplies the VAEs, text encoder and tokenizer; `-t` points at the quantized DiT. Native canvas is a 768px short edge (e.g. 1344x768 16:9); smaller canvases are off-distribution and degrade quickly. ## Image-to-Video (FL2VA) This repository is **I2V-capable** (partition `fl2va`, tasks `t2va` + `fl2va` in `model_index.json`). The image is **optional** — present as a keyframe conditioning row via `--image` / `--anchor`, or omitted for pure T2V. **Processor + vision tower required for I2V.** The `processor/` directory (7 files) and the 3 vision shards (`text_encoder/model-00005..00007-of-00008-vision.safetensors` + `vision_quant_config.json`, 529 `model.visual.*` tensors, 89 quantized) must be present. Without them the pipeline cannot encode images (falls back to T2V only). **Canonical I2V command (validated):** ```bash # 704×544 portrait (height 704 width 544) or 512×384 lightweight variant — both portrait 9:16 / 4:3 # Bridge image: inputs/bridge-20260813-181546-krea2_turbo-selfie-style-vertical-shot-a-young-woman-in-h.png PROMPT=$(cat <<'PROMPT_EOF' Vertical 9:16 TikTok-style UGC selfie video, handheld smartphone feel, natural indoor daylight near a window. A friendly creator speaks directly to camera with natural blinking, subtle head nods, and a warm smile. Add small human imperfections: a tiny hesitation, a soft breath, a quick smile mid-sentence, and a micro-pause before the last line. Realistic skin texture, stable identity, no face warping, minimal flicker, clean audio with natural room tone. No subtitles. No on-screen text. No logos. No watermarks. The creator says (exactly, with the same pacing and hesitations): "Okay, entonces… eh… un datazo. Si estás trabado con tu código, solo da el primer pasito… como, abre tu terminal y escribe fran. (sonríe) Así de simple. Te vas a sorprender de lo rápido que todo se vuelve más fácil." PROMPT_EOF ) .venv/bin/python scripts/generate.py "$PROMPT" \ -c -t /transformer -s 8 --seed 1996783985 \ --height 704 --width 544 -d 5 \ --image inputs/bridge-20260813-181546-krea2_turbo-selfie-style-vertical-shot-a-young-woman-in-h.png \ --anchor first --release-encoder \ -o i2v-704x544.mp4 # Lightweight alternative (faster, same identity): .venv/bin/python scripts/generate.py "$PROMPT" \ -c -t /transformer -s 8 --seed 1996783985 \ --height 512 --width 384 -d 5 \ --image inputs/bridge-20260813-181546-krea2_turbo-selfie-style-vertical-shot-a-young-woman-in-h.png \ --anchor first --release-encoder \ -o i2v-384x512.mp4 ``` Valid `anchors` are `first` and `last` (FL2VA). Omit `--image`/`--anchor` for text-to-video. The image is consumed by `text_encoder` via `processor/` (`Qwen2VLImageProcessorPil` fallback when torch absent) and the vision tower, then patched as video-conditioning rows (see `minimal I2V encode` in `PATCHES.md`). ## Example outputs ![MiniMax-H3 MLX comparison — 8-bit vs turbo variants](images/comparative.png) Red fox on a mossy log in a misty forest (prompt above, seed 42). Columns are time (0.4s → 3.0s); rows: 8-bit 16-step, 8-bit + Turbo LoRA, 4-bit + Turbo LoRA at 1344x768. Example clip (played inline by the Hub): | File | Build | Settings | |---|---|---| | `examples/v3_768x448_16steps_8bit.mp4` | 8-bit (this repo) | 768x448, 16 steps, 33.9 min | Turbo variants: see rows 2-3 above (a 1344x768 4-bit video re-render is planned). ## Model layout | Path | Role | Representation | |---|---|---| | `transformer/transformer.safetensors` | 33B DiT (joint video+audio) | 8-bit affine, group 64 (AdaLN 8-bit), ~35.3 GB | | `text_encoder/` (5+3 shards) | Qwen3-VL-32B conditioner (64 layers, truncated to layer 50) | 8-bit quantized, ~27.5 GB language (5 shards) + ~0.76 GB vision (3 shards `model-00005..00007-of-00008-vision.safetensors`, 529 `model.visual.*` tensors) | | `text_encoder/vision_quant_config.json` | Vision tower quantization recipe (affine 8-bit g64) | json, 89 quantized tensors (`blocks.*`, `merger.*`, `deepstack_merger_list.*`) | | `processor/` | Qwen3-VL processor (image preprocessing for I2V) | 7 files: `preprocessor_config.json`, `video_preprocessor_config.json`, `chat_template.json`, `tokenizer.json`, `tokenizer_config.json`, `vocab.json`, `merges.txt` (~11.6 MB) | | `video_vae/` | Tiled causal video VAE (17-frame chunks, latents_mean/std) | fp16/bf16, 5.2 GB | | `audio_vae/` | DAC encoder + BigVGAN vocoder, stereo 32 kHz | fp32, 0.6 GB | | `tokenizer/` | Qwen3-VL tokenizer | json/vocab/merges | | `model_index.json` | Pipeline config (fl2va partition, sigma shift video 12.0 / audio 3.0) | json | | `turbo_lora_4step_ema.safetensors` | 4-step EMA Turbo LoRA (larryvrh, Apache-2.0) | 758 MB | | `scripts/merge_h3_turbo.py` | Merge the Turbo LoRA into a transformer copy | python | ## Memory and speed (measured, M4 Max 68.7 GB) - Resident during generation: DiT ~21.5 GB + text encoder ~22 GB + VAEs ~6 GB (AdaLN projections are precomputed and dropped, freeing ~13.8 GB). With `--release-encoder`, the text encoder is **freed after conditioning** (`prompt_embeds` already built, before the denoise loop), freeing ~27.5 GB (log: `released text encoder after conditioning (freeing ~27.5 GB)` + `mx.clear_cache()` + `tree_flatten` sizing). Peak RSS drops by ~22 GB resident; overhead <1 s. - 768x448, 16 steps: **33.9 min** (119 s/step), peak RSS 24.7 GB, no swap. With turbo merge: ~24 min (8 steps). With **turbo8** (pure 8-bit, no merge) at 768x448, 8 steps (7 forwards): **~18.5 min (139 s/step)** vs turbo merge **251 s/step** — **~1.6–1.8× per-pixel** speedup (same canvas, fewer denoise forwards). - 1344x768 (native 16:9), 8 steps: **~2 h** (996 s/step), peak ~57 GB + compressed memory; requires an otherwise idle machine (jetsam kills it under heavy ambient load). - 704×544 I2V (portrait 9:16) turbo merge, 8 steps (7 forwards): **~33 min (251 s/step)**; turbo8: **~139 s/step** (measured 768x448 baseline, scales linearly per pixel). Lightweight 384×512 I2V is proportionally faster (off-distribution, useful for wiring checks). - TeaCache: measured **0 skips / 7 forwards** with default last-block hook (`--teacache`, `thresh 0.2`, `start 3`, `compute_last_step`) at 768x448 turbo8 — probe feature pre-final-norm of last block (block 49/50) barely moves, so gate `rel_l1 <= 0.2` never fires; with `thresh 0.35` only 1/7 skips and still **no wall saving** (partial forward already traverses 49/50 blocks, only saves final norm + heads). Overhead +8–10% without skip. **Recommendation: keep TeaCache OFF by default**; if wall saving is needed, use an earlier hook `--teacache-layer 40` with `--teacache-thresh 0.25–0.35` (trades quality for ~20% block saving per skip). See `PATCHES.md` and `tmp/S9_teacache_spike_RESULT.md`. - Attention is dense (MiniMax has not released sparse attention); quantization does not reduce the attention cost — it exists to fit memory. ## Validation report (2026-08-16, M4 Max 68.7 GB) Prompt used everywhere: *"a red fox leaps over a mossy log in a misty forest"*, seed 42. Every clip was judged by an independent vision-capable model on extracted frames (not statistics alone): | Build | Settings | Time | Visual verdict | |---|---|---|---| | 8-bit (this repo) | 768x448, 16 steps | 33.9 min | **9/10** — coherent fox walking a mossy log; stable background; no melting | | 8-bit + Turbo LoRA | 768x448, 8 steps | ~24 min | High confidence — "impressive for a 4-step turbo LoRA" | | 4-bit + Turbo LoRA | 1344x768, 8 steps | 2 h 3 min | **95%** — full leap arc; no melting, no banding (frames; mp4 mux was a local script bug, not the model) | | I2V E2E canonical (FL2VA, turbo merge) | 544×704 (704h×544w) 9:16, 8 steps (7 forwards), 5 s, seed 1996783985, `--image bridge-...png --anchor first --release-encoder` | ~33 min (251 s/step) | Keyframe anchored, identity preserved; requires `processor/` + vision tower (529 tensors); image optional (FL2VA otherwise uses key) | | I2V Turbo8 (pure 8-bit, FL2VA) | 768×448, 8 steps (7 forwards), 5 s | ~18.5 min (139 s/step, ~1.6–1.8× per-pixel vs turbo merge) | Same I2V prompt/image, pure 8-bit DiT without LoRA merge; faster, same wiring | Key findings: - The VAE decoder is **not** the source of artifacts: parity with the reference diffusers implementation is exact (video 1.2e-6, audio 2.3e-8). The 17-frame chunked decode with 5-frame crossfade is the reference design; a periodic ~1.2x seam rhythm is architectural and masked at native resolution. - **Canvas is the dominant quality factor**: the same VAE roundtrip scores 21.4 dB at 384x256 vs 27.9 dB at 768x448. Keep the short edge at 768. - First-generation failures (flat field, collapsed latents) were traced to an incompatible third-party serve-pack conversion (inert `condition_proj`, exploding text-encoder activations), not to the pipeline; see `PATCHES.md` and the report in `BIG_GOAL.md` of the source workspace. ## Turbo LoRA `turbo_lora_4step_ema.safetensors` (also present as `turbo_lora.safetensors` — hardlink alias, same 779,849,816 bytes) is the 4-step EMA distilled LoRA from [larryvrh/MiniMax-H3-Turbo-Lora](https://huggingface.co/larryvrh/MiniMax-H3-Turbo-Lora) (Apache-2.0; sha256 `5a6eeba1…`). Merge it into a transformer copy: ```bash .venv/bin/python scripts/merge_h3_turbo.py --base transformer --lora turbo_lora_4step_ema.safetensors -o transformer-turbo # then generate with the merged dir as -t and -s 8 (7 forwards) ``` 8 steps denoise cleanly with the merged weights; 4 steps can collapse to noise (see port issue #5). ## Patches vs upstream `PATCHES.md` documents **five** patch families on branch `video-lab` over base commit `b2f7e4d2` (2026-08-10): (a) DiT QKV blocked layout, (b) text-encoder 8-bit + processor fallback + positional scatter fix, (c) `release_text_encoder` headroom, (d) two-phase TeaCache, (e) I2V processor + vision shards. See `PATCHES.md` for per-file details and line references. The remaining pipeline is otherwise byte-identical to the base commit. ## License - **Weights**: MiniMax H3 Community License (included as `LICENSE`) — not open source; mark modified files, keep the attribution, territorial exclusions apply (EU/UK/KR/US excluded per the mlx-serve pack note). By downloading you accept those terms. - **Port code**: Apache-2.0 (PipeNetwork/minimax-h3-mlx). - **Turbo LoRA**: Apache-2.0 (larryvrh/MiniMax-H3-Turbo-Lora).