File size: 5,758 Bytes
3ac5208
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
---
license: other
license_name: minimax-h3-community-license
license_link: https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/LICENSE
base_model: MiniMaxAI/MiniMax-H3
tags:
- mlx
- apple-silicon
- text-to-video
- image-to-video
- audio-video-generation
- diffusion
pipeline_tag: image-text-to-video
library_name: mlx
---

# MiniMax-H3-MLX-8bit

MLX (Apple Silicon) build of the [**MiniMax-H3**](https://huggingface.co/MiniMaxAI/MiniMax-H3) diffusion
transformer, quantized to **8-bit** (group size 64).

> Powered by MiniMax H3.

**These files are modified.** The transformer weights have been converted to MLX and quantized;
they are not MiniMax's originals. Everything else about the model is unchanged.

## What this is

MiniMax-H3 generates **synchronized video and audio** together. It is not a language model: a 33B
diffusion transformer denoises video and audio latents jointly over one packed sequence, conditioned
by a frozen Qwen3-VL-32B encoder, with separate video and audio VAEs. Running it needs the pipeline
code, not just these weights:

```bash
git clone https://github.com/PipeNetwork/minimax-h3-mlx
cd minimax-h3-mlx && pip install -r requirements.txt
python scripts/generate.py "a red fox leaps over a mossy log" -o fox.mp4
```

This repository holds the **transformer only**. The VAEs and the text encoder come from the
[upstream release](https://huggingface.co/MiniMaxAI/MiniMax-H3); the pipeline loads them directly.

## Size

| | |
|---|---|
| on disk | 35.3 GB |
| resident during generation | **21.47 GB** |

The gap is deliberate. ~13B of H3's 33B parameters are the per-block AdaLN projections, whose only
input is the timestep embedding. For a fixed sampler schedule every modulation tensor a run needs is
precomputed once into a small table, and the projections are then dropped — so they are on disk but
never resident. The table scales with step count, not model size: measured at **145 MB for a 9-step
schedule** and 745 MB for 40 steps, against the 26 GB it replaces.

Those projections are quantized to **8-bit** here. That was measured, not assumed: quantizing them
shifts the modulation table by **0.25%**, an order of magnitude less than the 8-bit core's own
velocity error, and takes 12.2 GB off this download. (4-bit AdaLN is measurably worse — 0.77% on the
table, 2.8% on its worst tensor — and is not used at any core width.)

## How the widths compare

Measured with teacher forcing — one bfloat16 trajectory recorded, each variant re-predicting the
velocity at those same latents, so the difference is quantization error alone rather than trajectory
divergence. 20 paired observations per variant, aggregated with a paired bootstrap.

| bits | video rel-L2 [95% CI] | audio rel-L2 | video cosine |
|---:|---|---:|---:|
| 8 | 0.0329 [0.0277, 0.0381] | 0.0130 | 0.99941 |
| 6 | 0.0611 [0.0501, 0.0728] | 0.0274 | 0.99791 |
| 4 | 0.1649 [0.1324, 0.1971] | 0.1016 | 0.98456 |
| 3 | 0.2842 [0.2362, 0.3358] | 0.2341 | 0.95635 |

Every interval is disjoint from its neighbours, so the ranking is solid. Two things worth noting:
the steepest step is **6 to 4 bits** (2.7x), not at the low end; and audio degrades faster in
relative terms than video (its share of the error climbs from 0.40x at 8-bit to 0.82x at 3-bit),
plausibly because audio is a small fraction of the packed rows and has less redundancy to absorb it.

## Why 8, 6 and 4 bits only

Velocity error ranks the widths but does not say where output stops being usable — the scheduler
integrates velocity, so per-step error compounds along the trajectory. That has to be generated to
be seen. The same prompt, seed and settings were rendered through each checkpoint and compared to
bfloat16:

| build | PSNR vs bf16 | correlation | outcome |
|---|---:|---:|---|
| 8-bit | **27.6 dB** | 0.959 | near-identical |
| 4-bit | 22.0 dB | 0.854 | cooler colour, background artifacting, subject intact |
| 3-bit | 16.3 dB | 0.740 | **subject destroyed** |

At 3 bits the scene is gone — no animal, no log, just a textured field. It is built but **not
published**. Notably it does not degrade by blurring: its per-frame variance *rises* (54.7 against
bfloat16's 37.1) as structure is replaced by high-frequency noise, so a sharpness metric would have
scored it as healthy. 2-bit is not published either; extrapolation puts it near 50% velocity error.

6-bit was not rendered separately — it is bracketed by 8-bit and 4-bit, which both pass.

## Read this before choosing a quant

MiniMax has not released its sparse-attention implementation, so inference runs **dense** attention
over tens of thousands of rows. On an M3 Ultra a single denoising step costs about **8.8 minutes**
for a 5-second clip (37,966 packed rows) and **1.04 hours** for 15 seconds (109,318 rows).

Quantization does not change that. The bottleneck is attention FLOPs, which quantization does not
reduce; the linear layers are ~42% of the work at 5 s and ~20% at 15 s, so a 4-bit build is worth
roughly **1.2-1.4x end to end**. Choose a quant to *fit* H3 on your machine, not to make it quick.

## Licence

Governed by the [MiniMax H3 Community License](https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/LICENSE),
a copy of which is included in this repository. It is **not** an open-source licence. Notably:
redistribution must carry the agreement and mark modified files; commercial products above $20M
yearly revenue need separate authorization from MiniMax; and **the grant is territorially limited**
(worldwide, excluding the Excluded Territories defined in the agreement). By downloading these
weights you accept those terms.

The MLX port code is Apache-2.0 and lives at [https://github.com/PipeNetwork/minimax-h3-mlx](https://github.com/PipeNetwork/minimax-h3-mlx).