Instructions to use Giniiki/Z-Image-Turbo-mlx-4bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use Giniiki/Z-Image-Turbo-mlx-4bit with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("Giniiki/Z-Image-Turbo-mlx-4bit", dtype=torch.bfloat16, device_map="cuda") prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k" image = pipe(prompt).images[0] - MLX
How to use Giniiki/Z-Image-Turbo-mlx-4bit with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] huggingface-cli download --local-dir Z-Image-Turbo-mlx-4bit Giniiki/Z-Image-Turbo-mlx-4bit
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Draw Things
- DiffusionBee
- Atomic Chat
Z-Image-Turbo, 4-bit MLX
A 4-bit MLX conversion of Tongyi-MAI/Z-Image-Turbo.
5.91 GB instead of 32.84 GB. The model, the architecture and the outputs are
Tongyi-MAI's; the only thing done here is a change of numeric format.
The original publishes its 6.15 B-parameter diffusion transformer in float32. That is 24.62 GB of the 32.84 GB download for weights that no Apple-silicon runtime evaluates at float32. At 50 Mbit/s the original is about 1 h 27 m of downloading; this is about 16 minutes. On a phone that is the difference between a model you can offer and one you cannot.
What was changed
Stated plainly, as Apache-2.0 Β§4(b) requires.
- The transformer's 275 weight matrices were quantized to 4 bits, group
size 64, affine mode β
mlx.core.quantize(w, group_size=64, bits=4). That is 6,152,880,128 of its 6,154,908,736 parameters. 24.62 GB becomes 3.47 GB. - The text encoder's 253 weight matrices were quantized identically, including the tied vocabulary embedding. That is 4,022,272,000 of its 4,022,468,096 parameters. 8.04 GB becomes 2.26 GB.
- Everything not quantized was written as bfloat16. In the transformer that is a float32 β bfloat16 conversion; in the text encoder and the VAE the originals are already bfloat16 and the bytes are unchanged.
- The VAE was copied byte for byte. No quantization, no cast.
- Nothing was dropped, renamed, retrained, fine-tuned, merged or pruned. Every tensor in the original has a counterpart here, under its original name. The configs, the scheduler and the tokenizer are the original files, with one addition noted below.
Nothing else. This is not a different model and it is not tuned; it is the same weights at a smaller numeric width.
Why bfloat16 and never float16
Measured on these weights on an Apple M2, not assumed:
- The transformer's first forward returns 65536 of 65536 non-finite values at
float16 and a clean
[-6.75, 6.97]at bfloat16. The caption state it is conditioned on reaches 1.31e4, and the squares an RMSNorm takes of that are 1.7e8 against float16's 65504 ceiling. The float16 run decodes to a black frame. - The text encoder is about 4x more accurate in bfloat16 β 0.0043 against 0.0176 relative L2 versus transformers on the same 73-token prompt.
- The VAE returns NaN for every pixel at float16 and finite output at
bfloat16. Its
vae/config.jsonsaysforce_upcast: true, and it is right to.
So: bfloat16 for every unquantized tensor, and the scales and biases beside each quantized one are bfloat16 too.
What was deliberately left unquantized
| Left alone | Size | Reason |
|---|---|---|
| The whole VAE (244 tensors) | 167.7 MB | Almost entirely Conv2d, which is not Quantizable in mlx-swift at all. 0.17 GB is not worth a custom path. |
all_final_layer.2-1.adaLN_modulation.1.weight |
983,040 params, 1.4 MB at bf16 | It is index .1 of a two-element Sequential(SiLU, Linear). mlx-swift's Module.update(modules:) is handed [.none, QuantizedLinear] for that array, finds neither a value nor a dictionary at values.first, and throws unexpectedStructure from inside a try! β a crash, not a wrong image. A transformer block's adaLN_modulation is a one-element array and quantizes normally; only the final layer has the hole. |
RMSNorm gains, all biases, cap_pad_token |
1,045,568 params | Rank-1 or tiny; quantize does not touch them and there is nothing to gain. |
The text encoder's last decoder layer (layers.35) and its final norm are
kept, though a Z-Image pipeline stops at layer 34 and never evaluates them.
They are split across the published shards rather than sitting in a tail, so
dropping them saves about 100 MB out of 5.91 GB and costs the ability to load
this repository with every key accounted for. Not a good trade.
Naming: this repository keeps the original tensor names
There are two ways to publish an MLX quantization and they are not compatible.
- Emit the framework's post-quantization module tree. That is what
deepsweet/Z-Image-Turbo-6B-MLX-Q4does. It is convenient for exactly one port β the one whose module names the file was written from β and unreadable to every other, including ours. - Keep the checkpoint's own tensor names. Any loader that already knows how
to read
Tongyi-MAI/Z-Image-Turbokeeps its key mapping and only has to learn the packing. That is what this repository does.
The cost of the second choice is that diffusers has no convention for a quantized checkpoint under those names, so the layout below is an invention. It is written out in full so it can be implemented from this page alone.
The packing convention
For every quantized module at prefix P, the original P.weight of shape
[out, in] is replaced by three tensors under the same prefix:
| Tensor | dtype | Shape | Meaning |
|---|---|---|---|
P.weight |
uint32 |
[out, in / 8] |
The 4-bit codes, eight per word |
P.scales |
bfloat16 |
[out, in / 64] |
One scale per group of 64 |
P.biases |
bfloat16 |
[out, in / 64] |
One zero point per group |
P.bias, where the module has one, is untouched and stays a bfloat16 vector.
Element j of a word occupies bits [4j, 4j+4), least significant nibble
first, and the code is an unsigned integer in 0β¦15. Groups run along the input
axis. To reconstruct:
code = (weight[o, i / 8] >> (4 * (i % 8))) & 0xF
w[o, i] = code * scales[o, i / 64] + biases[o, i / 64]
which is mlx.core.dequantize(weight, scales, biases, group_size=64, bits=4, mode="affine"), and in mlx-swift
dequantized(weight, scales: scales, biases: biases, groupSize: 64, bits: 4).
That equivalence was checked by hand-decoding a tensor with the expression above
and diffing it against dequantized β they agree to one bfloat16 ULP, which is
the arithmetic and not the layout.
A tensor whose name does not have .scales beside it in the same file is
not quantized: read it as-is.
transformer/config.json and text_encoder/config.json each carry an added
"quantization": { "group_size": 64, "bits": 4 }
so a loader can discover the format from the config rather than sniffing key
names. It is the only edit made to any config file: the block is appended as
text and every other byte of both files is upstream's, so no value was
re-serialized and no precision was rewritten. model_index.json,
vae/config.json, scheduler/scheduler_config.json and all four tokenizer
files are byte-for-byte upstream's.
One trap for anyone writing the loader
The packed weights are uint32. A loader that casts every incoming tensor to
the model dtype β the usual value.astype(dtype) in a load loop β will
reinterpret eight packed weights as one float and produce silent garbage. Cast
floating-point tensors only.
Files
model_index.json
LICENSE
scheduler/scheduler_config.json
transformer/config.json (+ quantization block)
transformer/diffusion_pytorch_model-0000{1,2,3}-of-00003.safetensors 3.465 GB
transformer/diffusion_pytorch_model.safetensors.index.json
text_encoder/config.json (+ quantization block)
text_encoder/generation_config.json
text_encoder/model-0000{1,2,3}-of-00003.safetensors 2.263 GB
text_encoder/model.safetensors.index.json
vae/config.json
vae/diffusion_pytorch_model.safetensors 0.168 GB
tokenizer/{tokenizer.json,tokenizer_config.json,vocab.json,merges.txt} 0.016 GB
| Component | Original | Here |
|---|---|---|
| Transformer (6.15 B) | 24.620 GB, float32 | 3.465 GB |
| Text encoder, Qwen3-4B (4.02 B) | 8.045 GB, bfloat16 | 2.263 GB |
| VAE | 0.168 GB, bfloat16 | 0.168 GB (copied) |
| Tokenizer + configs | 0.012 GB | 0.016 GB |
| Total | 32.844 GB | 5.912 GB |
Each shard keeps the membership of the shard it came from, so the shard a tensor
lives in is the same one it lived in upstream. Both *.safetensors.index.json
files were regenerated: weight_map covers all 1071 transformer and 904 text
encoder tensors, and metadata.total_size is the exact summed payload
(3,465,052,288 and 2,262,920,192 bytes).
Verification
Everything below was run on an Apple M2 (16 GB) against this directory.
It loads with every key checked. The converted weights were loaded through
the reference implementation's own loader with all three checks on β no unused
keys, every model key set, no shape mismatch. The check is not vacuous: removing
the quantization block from transformer/config.json, so the loader builds a
dense module tree, makes it fail immediately with
mismatchedSize(path: ["x_embedder", "weight"], expected [3840, 64], actual [3840, 8]).
It produces a bit-identical image to the original. Same prompt, seed 42, 512 px, 8 steps, flow-match Euler, guidance 0:
| PNG MD5 | pixels | |
|---|---|---|
Tongyi-MAI/Z-Image-Turbo (32.84 GB), quantized to 4-bit at load |
8366dc53f2958d0ef946efa3f07a2f69 |
mean 132.09, sd 56.14, 0.08% clipped |
| This repository | 8366dc53f2958d0ef946efa3f07a2f69 |
mean 132.09, sd 56.14, 0.08% clipped |
Max absolute per-channel difference over all 786,432 float32 values: 0.0.
This is expected rather than lucky β the runtime path is float32 β bfloat16 β
quantize, and this conversion is the same three steps done ahead of time β but
it is the claim worth checking, and it was checked rather than assumed.
Per-tensor quantization error, dequantized and measured against the float32 original:
| Tensor | Relative L2 | SNR |
|---|---|---|
layers.10.feed_forward.w1.weight [10240, 3840] |
0.0947 | 20.5 dB |
cap_embedder.1.weight [3840, 2560] |
0.0911 | 20.8 dB |
all_x_embedder.2-1.weight [3840, 64] |
0.0762 | 22.4 dB |
model.embed_tokens.weight [151936, 2560] |
0.0911 | 20.8 dB |
model.layers.20.self_attn.q_proj.weight [4096, 2560] |
0.0928 | 20.7 dB |
Consistent across components and shapes, which is what argues against a packing error β a wrong nibble order or a mismatched group size shows up as one tensor far off the others, not as a uniform 9%.
Runtime, 512 px, 8 steps, staged so the encoder is released before the transformer loads: 20 s to load and forward once, about 9.0 s per denoising step, 2.4 GB peak resident.
What 4 bits costs, honestly
This conversion adds nothing to what the runtime already did, so the real question is what 4-bit costs against a wider format. The same prompt and seed at 8-bit (group 32) produces a different but equally sharp and photorealistic sample β the fox is framed differently, the fur and whisker detail is comparable, neither is degraded. Per-pixel that reads as PSNR 20.6 dB and correlation 0.91 against the 4-bit image, but those numbers are measuring two different samples of the same prompt, not fidelity loss. Sampling is chaotic: any change to the velocity field moves the trajectory.
Not measured: an unquantized bfloat16 reference. The transformer alone is 12.3 GB at bfloat16 and the attempt drove this 16 GB machine to 10.2 GB of swap without completing a single step. So there is no measurement here of 4-bit against full precision β only against 8-bit, and only on one prompt and one seed. If you need that comparison, run it on a machine with the memory for it.
Attribution
Derivative of Tongyi-MAI/Z-Image-Turbo,
Apache License 2.0. The full licence text ships as LICENSE in this
repository. The model card, the model itself and every architectural decision in
it are Tongyi-MAI's; see the original repository, the
project site and the
technical report.
The text encoder is Qwen3-4B, also Apache-2.0. The tokenizer files are the originals, unmodified.
Everything Tongyi-MAI states about the model applies unchanged: it is an 8-step distilled model, guidance should be 0, and the recommended resolution is 1024.
Quantized with mlx-swift's MLX.quantized. No content, capability or safety
property of the original was evaluated here.
- Downloads last month
- -
4-bit
Model tree for Giniiki/Z-Image-Turbo-mlx-4bit
Base model
Tongyi-MAI/Z-Image-Turbo