kruatech commited on
Commit
529e7ec
·
verified ·
1 Parent(s): b3217a5

Upload folder using huggingface_hub

Browse files
.gitattributes CHANGED
@@ -1,35 +1,19 @@
1
- *.7z filter=lfs diff=lfs merge=lfs -text
2
- *.arrow filter=lfs diff=lfs merge=lfs -text
3
  *.bin filter=lfs diff=lfs merge=lfs -text
4
- *.bz2 filter=lfs diff=lfs merge=lfs -text
5
  *.ckpt filter=lfs diff=lfs merge=lfs -text
6
- *.ftz filter=lfs diff=lfs merge=lfs -text
7
  *.gz filter=lfs diff=lfs merge=lfs -text
8
  *.h5 filter=lfs diff=lfs merge=lfs -text
9
- *.joblib filter=lfs diff=lfs merge=lfs -text
10
- *.lfs.* filter=lfs diff=lfs merge=lfs -text
11
- *.mlmodel filter=lfs diff=lfs merge=lfs -text
12
- *.model filter=lfs diff=lfs merge=lfs -text
13
  *.msgpack filter=lfs diff=lfs merge=lfs -text
14
  *.npy filter=lfs diff=lfs merge=lfs -text
15
  *.npz filter=lfs diff=lfs merge=lfs -text
16
  *.onnx filter=lfs diff=lfs merge=lfs -text
17
- *.ot filter=lfs diff=lfs merge=lfs -text
18
- *.parquet filter=lfs diff=lfs merge=lfs -text
19
- *.pb filter=lfs diff=lfs merge=lfs -text
20
- *.pickle filter=lfs diff=lfs merge=lfs -text
21
- *.pkl filter=lfs diff=lfs merge=lfs -text
22
  *.pt filter=lfs diff=lfs merge=lfs -text
23
  *.pth filter=lfs diff=lfs merge=lfs -text
24
- *.rar filter=lfs diff=lfs merge=lfs -text
25
- *.safetensors filter=lfs diff=lfs merge=lfs -text
26
- saved_model/**/* filter=lfs diff=lfs merge=lfs -text
27
- *.tar.* filter=lfs diff=lfs merge=lfs -text
28
  *.tar filter=lfs diff=lfs merge=lfs -text
29
- *.tflite filter=lfs diff=lfs merge=lfs -text
30
- *.tgz filter=lfs diff=lfs merge=lfs -text
31
- *.wasm filter=lfs diff=lfs merge=lfs -text
32
- *.xz filter=lfs diff=lfs merge=lfs -text
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
- *.zst filter=lfs diff=lfs merge=lfs -text
35
- *tfevents* filter=lfs diff=lfs merge=lfs -text
 
 
 
 
 
1
+ *.safetensors filter=lfs diff=lfs merge=lfs -text
 
2
  *.bin filter=lfs diff=lfs merge=lfs -text
 
3
  *.ckpt filter=lfs diff=lfs merge=lfs -text
 
4
  *.gz filter=lfs diff=lfs merge=lfs -text
5
  *.h5 filter=lfs diff=lfs merge=lfs -text
 
 
 
 
6
  *.msgpack filter=lfs diff=lfs merge=lfs -text
7
  *.npy filter=lfs diff=lfs merge=lfs -text
8
  *.npz filter=lfs diff=lfs merge=lfs -text
9
  *.onnx filter=lfs diff=lfs merge=lfs -text
 
 
 
 
 
10
  *.pt filter=lfs diff=lfs merge=lfs -text
11
  *.pth filter=lfs diff=lfs merge=lfs -text
 
 
 
 
12
  *.tar filter=lfs diff=lfs merge=lfs -text
 
 
 
 
13
  *.zip filter=lfs diff=lfs merge=lfs -text
14
+ **/tokenizer/tokenizer.json filter=lfs diff=lfs merge=lfs -text
15
+ **/tokenizer/vocab.json filter=lfs diff=lfs merge=lfs -text
16
+ **/tokenizer/merges.txt filter=lfs diff=lfs merge=lfs -text
17
+ **/lm_tokenizer/tokenizer.json filter=lfs diff=lfs merge=lfs -text
18
+ **/lm_tokenizer/vocab.json filter=lfs diff=lfs merge=lfs -text
19
+ **/lm_tokenizer/merges.txt filter=lfs diff=lfs merge=lfs -text
LICENSES.md ADDED
@@ -0,0 +1,22 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Licences in this repository
2
+
3
+ Each folder is a derivative of an upstream model and carries that model's licence.
4
+
5
+ | folder | upstream | licence |
6
+ | --------------------------------- | --------------------------------- | ---------- |
7
+ | `shared/qwen3-4b-text-encoder` | inside the two image repos below | Apache-2.0 |
8
+ | `shared/qwen3-4b-text-encoder-q4` | inside the two image repos below | Apache-2.0 |
9
+ | `z-image-turbo` | Tongyi-MAI/Z-Image-Turbo | Apache-2.0 |
10
+ | `z-image-turbo-q4` | Tongyi-MAI/Z-Image-Turbo | Apache-2.0 |
11
+ | `flux2-klein-4b` | black-forest-labs/FLUX.2-klein-4B | Apache-2.0 |
12
+ | `flux2-klein-4b-q4` | black-forest-labs/FLUX.2-klein-4B | Apache-2.0 |
13
+ | `ace-step-1.5-turbo` | ACE-Step/Ace-Step1.5 | MIT |
14
+
15
+ Apache-2.0: https://www.apache.org/licenses/LICENSE-2.0
16
+ MIT: https://opensource.org/license/mit
17
+
18
+ Conversion tooling and the model cards in this repository: MIT.
19
+
20
+ The upstream cards state usage restrictions and responsible-use commitments that
21
+ redistribution does not repeal. See in particular the out-of-scope use section of
22
+ the FLUX.2 [klein] card: https://huggingface.co/black-forest-labs/FLUX.2-klein-4B
README.md CHANGED
@@ -1,5 +1,185 @@
1
  ---
2
  license: other
3
  license_name: mixed-apache-2.0-and-mit
4
- license_link: LICENSE
 
 
 
 
 
 
 
 
 
5
  ---
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
  ---
2
  license: other
3
  license_name: mixed-apache-2.0-and-mit
4
+ license_link: https://huggingface.co/kruatech/studio-mlx/blob/main/LICENSES.md
5
+ library_name: mlx
6
+ tags:
7
+ - mlx
8
+ - mlx-swift
9
+ - apple-silicon
10
+ - text-to-image
11
+ - image-editing
12
+ - text-to-music
13
+ - quantized
14
  ---
15
+
16
+ <div align="center">
17
+
18
+ # studio-mlx
19
+
20
+ **Image and music models converted for native MLX inference on Apple silicon**
21
+
22
+ ![MLX](https://img.shields.io/badge/MLX-0.32-1f6feb) ![platform](https://img.shields.io/badge/platform-Apple%20silicon-black) ![precision](https://img.shields.io/badge/precision-bf16%20and%204--bit-6f42c1) [![licence](https://img.shields.io/badge/licence-Apache--2.0%20and%20MIT-brightgreen)](https://huggingface.co/kruatech/studio-mlx/blob/main/LICENSES.md)
23
+
24
+ </div>
25
+
26
+ Three models and the text encoder the two image models share. Every folder is a self-contained bundle with its own card, so you download only what you need.
27
+
28
+ Weights are not retrained. Tensor layouts were adapted for MLX and the 4-bit folders are quantized, so parameters are mathematically equivalent to upstream rather than byte-identical.
29
+
30
+ ## Quick start
31
+
32
+ Pick one line. The image models need the shared encoder, the music model does not.
33
+
34
+ ```bash
35
+ pip install "huggingface_hub[hf_xet]"
36
+
37
+ # Z-Image-Turbo - 6.14 GiB
38
+ hf download kruatech/studio-mlx --local-dir bundles \
39
+ --include "shared/qwen3-4b-text-encoder-q4/*" "z-image-turbo-q4/*"
40
+
41
+ # FLUX.2 [klein] 4B - 4.79 GiB
42
+ hf download kruatech/studio-mlx --local-dir bundles \
43
+ --include "shared/qwen3-4b-text-encoder-q4/*" "flux2-klein-4b-q4/*"
44
+
45
+ # ACE-Step 1.5 Turbo - 9.38 GiB
46
+ hf download kruatech/studio-mlx --local-dir bundles --include "ace-step-1.5-turbo/*"
47
+
48
+ ```
49
+
50
+ The `-q4` folders sit at the top level rather than inside the bf16 ones, so an include pattern fetches one precision without dragging in the other.
51
+
52
+ ## What is inside
53
+
54
+ | folder | what it does | precision | size | licence |
55
+ |----------------------------------------------------------------------|--------------------------------------------------------|-----------|-----------|------------|
56
+ | [`shared/qwen3-4b-text-encoder`](shared/qwen3-4b-text-encoder) | Shared text encoder for both image models | `bf16` | 7.51 GiB | Apache-2.0 |
57
+ | [`shared/qwen3-4b-text-encoder-q4`](shared/qwen3-4b-text-encoder-q4) | Shared text encoder for both image models | `4-bit` | 2.59 GiB | Apache-2.0 |
58
+ | [`z-image-turbo`](z-image-turbo) | Text to image, 6B, 8 steps | `bf16` | 11.62 GiB | Apache-2.0 |
59
+ | [`z-image-turbo-q4`](z-image-turbo-q4) | Text to image, 6B, 8 steps | `4-bit` | 3.55 GiB | Apache-2.0 |
60
+ | [`flux2-klein-4b`](flux2-klein-4b) | Text to image and multi-reference editing, 4B, 4 steps | `bf16` | 7.39 GiB | Apache-2.0 |
61
+ | [`flux2-klein-4b-q4`](flux2-klein-4b-q4) | Text to image and multi-reference editing, 4B, 4 steps | `4-bit` | 2.20 GiB | Apache-2.0 |
62
+ | [`ace-step-1.5-turbo`](ace-step-1.5-turbo) | Text to music, 48 kHz stereo | `bf16` | 9.38 GiB | MIT |
63
+
64
+ ## bf16 or 4-bit
65
+
66
+ **Use the 4-bit folders.** Quantization is MLX affine: 4 bits per weight with a bf16 scale and bias for every group of 64, about 4.5 bits per weight and 28-30% of the bf16 size. Normalizations, modulations and the encoder's embedding table stay in bf16, and matrix multiplication still runs in bf16 - the gain is memory and load time, not integer arithmetic.
67
+
68
+ | | bf16 | 4-bit |
69
+ |-----------------------------------|-----------|----------|
70
+ | Z-Image DiT | 11.46 GiB | 3.40 GiB |
71
+ | klein DiT | 7.22 GiB | 2.03 GiB |
72
+ | shared encoder | 7.51 GiB | 2.58 GiB |
73
+ | klein 512 px, 4 steps, wall clock | 170 s | 73 s |
74
+
75
+ Output quality was indistinguishable: the same prompt and seed gave two clean images that differ the way two seeds differ. The bf16 folders exist as the accuracy reference a port can be checked against.
76
+
77
+ ## Bundle layout
78
+
79
+ ```
80
+ <folder>/
81
+ ├── manifest.json SHA-256, byte sizes and tensor counts for every file
82
+ ├── config/ component configs copied from upstream
83
+ ├── tokenizer/ BPE plus chat_format.json, where applicable
84
+ └── weights/ safetensors in MLX tensor layout
85
+ ```
86
+
87
+ `manifest.json` records the quantization settings, so a loader rebuilds the exact layout without being told. Verify integrity before first use: the manifest carries SHA-256 per file, which catches corruption that size checks miss.
88
+
89
+ `tokenizer/chat_format.json` holds the Qwen chat template already rendered into a prefix and a suffix for each pipeline, so a runtime needs no Jinja at all.
90
+
91
+ <details>
92
+ <summary><b>Tensor layout</b></summary>
93
+
94
+ MLX convolutions expect channels last, PyTorch expects channels first, so convolution weights are permuted during conversion. Everything else keeps its upstream shape.
95
+
96
+ | kind | PyTorch | MLX |
97
+ |-----------------------------|---------------------|---------------------|
98
+ | `Conv2d.weight` | `(out, in, kH, kW)` | `(out, kH, kW, in)` |
99
+ | `Linear`, norms, embeddings | `(out, in)` | unchanged |
100
+
101
+ No `weight_norm` and no 3-D convolutions appear in these models, so no folding was needed. Parameter names match the upstream checkpoints, so weights load without remapping.
102
+
103
+ </details>
104
+
105
+ <details>
106
+ <summary><b>Verification numbers</b></summary>
107
+
108
+ Every module was compared against the upstream reference on fixed inputs in float32. The metric is `rel_max = max|a-b| / max|a|`.
109
+
110
+ | module | rel_max | reference |
111
+ |-------------------------------------------------|-----------|---------------------------------|
112
+ | Qwen3 encoder, hidden_states[-2] | 2.7e-07 | transformers |
113
+ | Qwen3 encoder, layers 9/18/27 with padding mask | 6.0e-07 | transformers |
114
+ | Z-Image DiT | 4.2e-06 | diffusers |
115
+ | Flux2 DiT | 3.7e-07 | diffusers |
116
+ | VAE decoder | 1.3e-05 | diffusers |
117
+ | VAE encoder | 5.3e-06 | diffusers |
118
+ | VAE tiled decode | 5.1e-06 | diffusers tiled_decode |
119
+ | sigma schedule, static shift | 3.2e-08 | FlowMatchEulerDiscreteScheduler |
120
+ | sigma schedule, exponential dynamic shift | 7.7e-08 | FlowMatchEulerDiscreteScheduler |
121
+ | position ids, latent packing, patchify | bit-exact | pipeline helpers |
122
+
123
+ A Swift MLX implementation was then checked against the Python one:
124
+
125
+ | module | rel_max | note |
126
+ |------------------------------|-----------------|------------------------------------------------|
127
+ | tokenization, both pipelines | bit-exact | same ids |
128
+ | Qwen3 encoder, bf16 | 2.2e-04 | Z-Image branch |
129
+ | Qwen3 encoder, bf16 | 4.6e-03 | klein branch, 512 tokens with mask |
130
+ | Qwen3 encoder, 4-bit | 3.0e-04 | proves the quantized layout is rebuilt exactly |
131
+ | Z-Image DiT, float32 | 3.7e-07 | 8 layers |
132
+ | Flux2 DiT, float32 | 1.0e-06 | 2 double + 2 single blocks |
133
+ | VAE decode / tiled / encode | 1e-05 or better | both VAE classes |
134
+ | VAE BatchNorm statistics | 0 | exact |
135
+
136
+ The bf16 figures for a full-depth DiT are larger - 2e-02 for both models - and that is rounding order, not a defect: at float32 the same code agrees to 1e-06, and the deviation grows with depth from a bf16-level 1e-04 per block. Comparing two bf16 implementations below 1e-02 is not meaningful for a 30-block network.
137
+
138
+ </details>
139
+
140
+ <details>
141
+ <summary><b>Performance on an M1 Max</b></summary>
142
+
143
+ Mac Studio, Apple M1 Max, 10-core CPU, 24-core GPU, 32 GB, macOS 26.5. Release build, 4-bit weights, batch 1, warm page cache.
144
+
145
+ | run | per step | VAE decode |
146
+ |--------------------------------------|----------|------------|
147
+ | Z-Image, 512 px, 4 steps | 2.6 s | 1.5 s |
148
+ | klein, 512 px, 4 steps | 2.3 s | 0.9 s |
149
+ | klein, 1024 px, 4 steps | 8.0 s | 0.2 s |
150
+ | klein, 512 px + one 512 px reference | 4.1 s | 0.9 s |
151
+
152
+ Reading the weights dominates a cold run and depends entirely on the storage: on an external volume delivering about 0.2 GiB/s, the 4-bit encoder took 22-32 s and a 4-bit DiT 18-35 s, against 60-107 s for the same weights in bf16. On internal storage expect these to be several times shorter. Compute is unaffected: the numbers above are steady state.
153
+
154
+ Stages run one at a time and release their weights, so peak memory is set by the largest single component rather than their sum. The 4-bit DiT is 2.03 GiB for klein and 3.40 GiB for Z-Image; activations at 1024 px add to that, and no separate peak measurement was made.
155
+
156
+ </details>
157
+
158
+ <details>
159
+ <summary><b>Limitations</b></summary>
160
+
161
+ - Seeds are not compatible with the upstream pipelines, which use `torch.Generator`. The same prompt gives comparable images, never the same file.
162
+ - Both models are distilled: CFG is not applied, and step counts are low by design (8 for Z-Image, 4 for klein).
163
+ - Z-Image's Omni mode is not covered: it needs a SigLIP encoder that is not part of this bundle.
164
+ - Tiled VAE decode is available for high resolutions and gives a result that differs slightly from a single pass, exactly as it does in diffusers.
165
+ - Batches larger than one are not implemented.
166
+
167
+ </details>
168
+
169
+ ## Licences
170
+
171
+ Derivatives of three upstream models under two licences. Each folder carries the licence of its upstream model.
172
+
173
+ | folder | upstream | licence |
174
+ |-----------------------------------|-----------------------------------------------------------------------------------------------|-----------------------------------------------------------|
175
+ | `shared/qwen3-4b-text-encoder` | [black-forest-labs/FLUX.2-klein-4B](https://huggingface.co/black-forest-labs/FLUX.2-klein-4B) | [Apache-2.0](https://www.apache.org/licenses/LICENSE-2.0) |
176
+ | `shared/qwen3-4b-text-encoder-q4` | [black-forest-labs/FLUX.2-klein-4B](https://huggingface.co/black-forest-labs/FLUX.2-klein-4B) | [Apache-2.0](https://www.apache.org/licenses/LICENSE-2.0) |
177
+ | `z-image-turbo` | [Tongyi-MAI/Z-Image-Turbo](https://huggingface.co/Tongyi-MAI/Z-Image-Turbo) | [Apache-2.0](https://www.apache.org/licenses/LICENSE-2.0) |
178
+ | `z-image-turbo-q4` | [Tongyi-MAI/Z-Image-Turbo](https://huggingface.co/Tongyi-MAI/Z-Image-Turbo) | [Apache-2.0](https://www.apache.org/licenses/LICENSE-2.0) |
179
+ | `flux2-klein-4b` | [black-forest-labs/FLUX.2-klein-4B](https://huggingface.co/black-forest-labs/FLUX.2-klein-4B) | [Apache-2.0](https://www.apache.org/licenses/LICENSE-2.0) |
180
+ | `flux2-klein-4b-q4` | [black-forest-labs/FLUX.2-klein-4B](https://huggingface.co/black-forest-labs/FLUX.2-klein-4B) | [Apache-2.0](https://www.apache.org/licenses/LICENSE-2.0) |
181
+ | `ace-step-1.5-turbo` | [ACE-Step/Ace-Step1.5](https://huggingface.co/ACE-Step/Ace-Step1.5) | [MIT](https://opensource.org/license/mit) |
182
+
183
+ The shared text encoder is redistributed from inside the two image repositories, both Apache-2.0. Conversion tooling and these cards: MIT.
184
+
185
+ The upstream cards state usage restrictions and responsible-use commitments that redistribution does not repeal. See in particular the out-of-scope use section of the FLUX.2 [klein] card.
ace-step-1.5-turbo/README.md ADDED
@@ -0,0 +1,213 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: mit
3
+ base_model: ACE-Step/Ace-Step1.5
4
+ tags:
5
+ - music-generation
6
+ - text-to-music
7
+ - mlx
8
+ - mlx-swift
9
+ - apple-silicon
10
+ - macos
11
+ library_name: mlx
12
+ pipeline_tag: text-to-audio
13
+ ---
14
+
15
+ # ACE-Step 1.5 Turbo for native MLX inference on Apple Silicon
16
+
17
+ Converted upstream weights for a native macOS implementation of ACE-Step 1.5.
18
+ The inference path runs entirely through Swift and MLX — no Python, PyTorch,
19
+ Transformers, Diffusers, subprocess or local server.
20
+
21
+ Weights are not retrained and not quantized. Tensor layouts were adapted for
22
+ MLX and VAE `weight_norm` was folded during conversion, so parameters are
23
+ mathematically equivalent to upstream rather than byte-identical.
24
+
25
+ > This folder is part of the combined `kruatech/studio-mlx` repository.
26
+ > The status note below applies to ACE-Step only: the Swift runtime for the
27
+ > image models in this repository is implemented, the one for ACE-Step is not.
28
+
29
+ ## Status
30
+
31
+ The Swift runtime is not published yet and will live in a separate repository,
32
+ so this repository currently provides the model bundle only. To try ACE-Step 1.5
33
+ today, use the upstream repository.
34
+
35
+ ## Supported tasks
36
+
37
+ All four turbo task types work:
38
+
39
+ | task | what it does |
40
+ |---|---|
41
+ | `text2music` | generate from a text description and lyrics |
42
+ | `repaint` | regenerate a time range, keep the rest bit-exact ¹ |
43
+ | `cover` | follow an existing track's structure via 5 Hz FSQ codes |
44
+ | `cover-nofsq` | same, through the full 25 Hz latent — stays closer to the source |
45
+
46
+ ¹ Outside the range the original PCM is spliced back, so it matches bit for bit
47
+ except in the crossfade windows at the boundaries (0.025 s each by default).
48
+
49
+ Also implemented: CFG with APG guidance, DCW wavelet correction (on by default,
50
+ as upstream), tiled VAE encode and decode, retake variations, velocity
51
+ stabilisation, second-order Heun sampler, LM forward pass.
52
+
53
+ Not implemented: generating audio codes from text end to end (the LM's
54
+ 217204-entry BPE tokenizer is not ported), constrained metadata decoding,
55
+ batches larger than one. Metadata is supplied explicitly.
56
+
57
+ Not applicable to turbo: `use_adg` and flow-edit morphing need a base model;
58
+ `extract`, `lego` and `complete` are base-only tasks.
59
+
60
+ ## Pipeline
61
+
62
+ ```
63
+ prompt ──▶ Qwen3-Embedding-0.6B ──┐
64
+ lyrics ──▶ Qwen3-Embedding-0.6B ──┼──▶ CondEncoder ──▶ encoder_hidden_states
65
+ reference latent ─────────────────┘
66
+
67
+ noise [1, T, 64] ──▶ DiT, 24 layers, 8 steps ──▶ latent ──▶ VAE ──▶ 48 kHz stereo
68
+ ```
69
+
70
+ `T` is latent frames at 25 Hz: 30 s of audio is 750 frames. The VAE upsamples by
71
+ 1920 per frame.
72
+
73
+ ## Files
74
+
75
+ Required for `text2music`, `repaint`, `cover` — 6.34 GB:
76
+
77
+ | path | size | contents |
78
+ |---|---|---|
79
+ | `weights/dit.safetensors` | 4.79 GB | DiT, FSQ quantizer, detokenizer, attention pooler — 677 tensors, BF16 |
80
+ | `weights/text_encoder.safetensors` | 1.19 GB | Qwen3-Embedding-0.6B, 310 tensors, BF16 |
81
+ | `weights/vae.safetensors` | 337 MB | AutoencoderOobleck, 291 tensors: 146 encoder + 145 decoder |
82
+ | `weights/silence.bin` | 3.8 MB | silence latent, F32 `[15000, 64]` in NLC |
83
+ | `config/dit.json`, `config/text.json` | 4 KB | model configs |
84
+ | `tokenizer/` | 14 MB | Qwen3-Embedding BPE |
85
+ | `manifest.json` | 3 KB | SHA-256, sizes and tensor counts |
86
+
87
+ Optional, 3.71 GB: `optional/lm.safetensors` and `optional/lm_tokenizer/` hold
88
+ the 1.7B LM. `text2music` never loads them, and cover works without them —
89
+ audio is encoded and tokenised directly. The LM is only needed to produce codes
90
+ from a text description instead of a reference track.
91
+
92
+ Verify integrity before first use: `manifest.json` carries SHA-256 for every
93
+ file, which catches corruption that size and header checks miss.
94
+
95
+ ## Tensor layout
96
+
97
+ MLX convolutions expect channels last, PyTorch expects channels first, so
98
+ convolution weights are permuted during conversion:
99
+
100
+ | kind | PyTorch | MLX |
101
+ |---|---|---|
102
+ | `Conv1d` | `(out, in, k)` | `(out, k, in)` |
103
+ | `ConvTranspose1d` | `(in, out, k)` | `(out, k, in)` |
104
+ | Snake `alpha`, `beta` | `(1, C, 1)` | `(1, 1, C)` |
105
+ | `Linear` | `(out, in)` | unchanged |
106
+
107
+ `weight_norm` pairs are folded into single tensors: `w = v · (g / ‖v‖)`,
108
+ computed in F32 with an F64 accumulator and rounded once. This removes 74
109
+ tensors from the VAE.
110
+
111
+ ## Memory
112
+
113
+ Peak MLX active memory per stage, F32, 60 s of audio with CFG 3.0:
114
+
115
+ | stage | peak active |
116
+ |---|---|
117
+ | text encoder | 2.1 GB |
118
+ | CondEncoder | 2.8 GB |
119
+ | DiT | 7.2 GB |
120
+ | VAE decoder, tiled | 4.1 GB |
121
+
122
+ Stages run sequentially and release their weights, so the figures do not add
123
+ up. DiT sets the ceiling; with the buffer cache the working set is about 9.2 GB.
124
+
125
+ **16 GB is the tested minimum.** An M1 Pro with 16 GB ran all four tasks
126
+ without swap growth. 8 GB is not expected to fit.
127
+
128
+ The VAE must be tiled — a single-pass 30 s decode drove the MLX buffer cache to
129
+ 25 GB. With tiling the cache stayed under 400 MB across all tested lengths, 10 s
130
+ to 180 s. Overlaps of 16, 32 and 64 latent frames each produced output identical
131
+ to a single pass.
132
+
133
+ DiT must run in F32. BF16 halves its memory but fails accuracy at
134
+ `rel_max 1.45e-01`.
135
+
136
+ ## Performance
137
+
138
+ Warm cache, release build, 8 diffusion steps, batch 1, F32. RTF is generation
139
+ time over output duration.
140
+
141
+ **M1 Pro, 16 GB** — 10-core CPU, 16-core GPU, macOS 26.5, Swift 6.3.2:
142
+
143
+ ```
144
+ mode wall dit s/step vae peak MB RTF
145
+ 30 s, guidance 1.0 8.6s 3.65s 0.46s 3.78s 6602 0.287
146
+ 60 s, guidance 1.0 15.6s 6.81s 0.85s 7.46s 6806 0.260
147
+ 30 s, guidance 3.0 11.6s 6.91s 0.86s 3.51s 6761 0.387
148
+ 60 s, guidance 3.0 22.1s 13.38s 1.67s 7.40s 7180 0.368
149
+ 30 s, Heun sampler 16.7s 11.93s 1.49s 3.55s 6761 0.557
150
+ ```
151
+
152
+ **M1 Max, 32 GB** — 10-core CPU, 24-core GPU, macOS 26.5, Swift 6.3.3, CFG 3.0:
153
+
154
+ ```
155
+ 30 s: text 0.15s cond 0.19s dit 4.47s vae 2.28s · 7.69s wall
156
+ 60 s: dit 8.77s (1.10 s/step) · RTF 0.146
157
+ ```
158
+
159
+ The M1 Max is about 1.5× faster per diffusion step, tracking GPU core count
160
+ rather than memory. Peak memory is the same on both.
161
+
162
+ CFG doubles the batch and roughly doubles diffusion time. Heun doubles model
163
+ evaluations per step. A first run after process start is several times slower
164
+ while 6.3 GB of weights come off disk — measure the second run.
165
+
166
+ ## Verification
167
+
168
+ Every module was compared against the upstream reference on fixed inputs in
169
+ F32. The metric is `rel_max = max_abs / std(reference)`, threshold `1e-03`:
170
+
171
+ | module | rel_max | reference |
172
+ |---|---|---|
173
+ | tokenization | bit-exact | `transformers` |
174
+ | text encoder, 28 layers | 3.39e-05 | `transformers` |
175
+ | CondEncoder | 5.98e-04 | `transformers` |
176
+ | DiT, 24 layers | 4.35e-05 | `diffusers` |
177
+ | VAE encoder | 5.38e-04 | `diffusers` |
178
+ | VAE decoder | cosine 1.000000 | `diffusers` |
179
+ | FSQ | 3.40e-07 | `vector_quantize_pytorch` ¹ |
180
+ | detokenizer | 1.15e-05 | `transformers` |
181
+ | audio tokenizer | 1.31e-05 | `transformers` ² |
182
+ | LM, 28 layers | 2.48e-04 | `transformers` ³ |
183
+ | CFG / APG | 4.95e-05 | independent MLX implementation ⁴ |
184
+
185
+ ¹ Codebook matched exactly on all 64000 entries.
186
+ ² All 50 indices matched exactly.
187
+ ³ argmax matched at all 50 positions.
188
+ ⁴ APG has no PyTorch reference upstream; validated on identical noise.
189
+
190
+ Reference versions: torch 2.13.0, transformers 5.14.1, diffusers 0.39.0,
191
+ mlx 0.32.0.
192
+
193
+ Verify the VAE against an F32 weight file, not the BF16 one shipped here.
194
+ `weight_norm` folding produces new values, and storing them in BF16 quantises
195
+ them: `rel_max 5.11e-02` against BF16 versus `4.78e-05` against F32. For
196
+ inference BF16 is fine — median signal-to-error through encode and decode is
197
+ 49.7 dB.
198
+
199
+ ## Limitations
200
+
201
+ Seeds are not compatible with the upstream pipeline, which uses
202
+ `torch.Generator`. The same prompt gives comparable music, never the same file.
203
+
204
+ Minimum length is 5.12 s. Step count is fixed at 8; this is the turbo model.
205
+
206
+ ## Licenses
207
+
208
+ MIT, same as upstream.
209
+
210
+ - DiT, VAE, LM: converted from
211
+ [ACE-Step 1.5](https://github.com/ace-step/ACE-Step-1.5), MIT.
212
+ - Text encoder and tokenizer: Qwen3-Embedding-0.6B, redistributed unchanged.
213
+ - Conversion tooling and this card: MIT.
flux2-klein-4b-q4/README.md ADDED
@@ -0,0 +1,130 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ library_name: mlx
4
+ tags:
5
+ - mlx
6
+ - mlx-swift
7
+ - apple-silicon
8
+ base_model:
9
+ - black-forest-labs/FLUX.2-klein-4B
10
+ ---
11
+
12
+ <div align="center">
13
+
14
+ # FLUX.2 [klein] 4B for MLX, 4-bit
15
+
16
+ **Text to image and multi-reference editing, 4B, 4 steps**
17
+
18
+ ![precision](https://img.shields.io/badge/precision-4--bit-6f42c1) ![size](https://img.shields.io/badge/size-2.20%20GiB-1f6feb) [![licence](https://img.shields.io/badge/licence-Apache--2.0-brightgreen)](https://www.apache.org/licenses/LICENSE-2.0) [![part of](https://img.shields.io/badge/part%20of-studio--mlx-black)](https://huggingface.co/kruatech/studio-mlx)
19
+
20
+ </div>
21
+
22
+ 4-bit version of FLUX.2 [klein] 4B. Everything in the bf16 card applies.
23
+
24
+ Converted from [black-forest-labs/FLUX.2-klein-4B](https://huggingface.co/black-forest-labs/FLUX.2-klein-4B) for native MLX inference on Apple silicon.
25
+
26
+ ## Download
27
+
28
+ Also download `shared/qwen3-4b-text-encoder-q4`.
29
+
30
+ ```bash
31
+ pip install "huggingface_hub[hf_xet]"
32
+
33
+ hf download kruatech/studio-mlx --local-dir bundles \
34
+ --include "shared/qwen3-4b-text-encoder-q4/*" "flux2-klein-4b-q4/*"
35
+ ```
36
+
37
+ ## Files
38
+
39
+ | component | class | precision | tensors | size |
40
+ |-------------|-------------------------|-----------------|---------|----------|
41
+ | transformer | Flux2Transformer2DModel | 4-bit, group 64 | 381 | 2.03 GiB |
42
+ | vae | AutoencoderKLFlux2 | keep | 251 | 160 MiB |
43
+
44
+ Folder total: **2.20 GiB**. `manifest.json` carries SHA-256, byte sizes and tensor counts for every file.
45
+
46
+ ## How this model works in MLX
47
+
48
+ - 106 linear layers quantized. The VAE is not quantized.
49
+
50
+ <details>
51
+ <summary><b>Tensor layout</b></summary>
52
+
53
+ MLX convolutions expect channels last, PyTorch expects channels first, so convolution weights are permuted during conversion. Everything else keeps its upstream shape.
54
+
55
+ | kind | PyTorch | MLX |
56
+ |-----------------------------|---------------------|---------------------|
57
+ | `Conv2d.weight` | `(out, in, kH, kW)` | `(out, kH, kW, in)` |
58
+ | `Linear`, norms, embeddings | `(out, in)` | unchanged |
59
+
60
+ No `weight_norm` and no 3-D convolutions appear in these models, so no folding was needed. Parameter names match the upstream checkpoints, so weights load without remapping.
61
+
62
+ </details>
63
+
64
+ <details>
65
+ <summary><b>Verification numbers</b></summary>
66
+
67
+ Every module was compared against the upstream reference on fixed inputs in float32. The metric is `rel_max = max|a-b| / max|a|`.
68
+
69
+ | module | rel_max | reference |
70
+ |-------------------------------------------------|-----------|---------------------------------|
71
+ | Qwen3 encoder, hidden_states[-2] | 2.7e-07 | transformers |
72
+ | Qwen3 encoder, layers 9/18/27 with padding mask | 6.0e-07 | transformers |
73
+ | Z-Image DiT | 4.2e-06 | diffusers |
74
+ | Flux2 DiT | 3.7e-07 | diffusers |
75
+ | VAE decoder | 1.3e-05 | diffusers |
76
+ | VAE encoder | 5.3e-06 | diffusers |
77
+ | VAE tiled decode | 5.1e-06 | diffusers tiled_decode |
78
+ | sigma schedule, static shift | 3.2e-08 | FlowMatchEulerDiscreteScheduler |
79
+ | sigma schedule, exponential dynamic shift | 7.7e-08 | FlowMatchEulerDiscreteScheduler |
80
+ | position ids, latent packing, patchify | bit-exact | pipeline helpers |
81
+
82
+ A Swift MLX implementation was then checked against the Python one:
83
+
84
+ | module | rel_max | note |
85
+ |------------------------------|-----------------|------------------------------------------------|
86
+ | tokenization, both pipelines | bit-exact | same ids |
87
+ | Qwen3 encoder, bf16 | 2.2e-04 | Z-Image branch |
88
+ | Qwen3 encoder, bf16 | 4.6e-03 | klein branch, 512 tokens with mask |
89
+ | Qwen3 encoder, 4-bit | 3.0e-04 | proves the quantized layout is rebuilt exactly |
90
+ | Z-Image DiT, float32 | 3.7e-07 | 8 layers |
91
+ | Flux2 DiT, float32 | 1.0e-06 | 2 double + 2 single blocks |
92
+ | VAE decode / tiled / encode | 1e-05 or better | both VAE classes |
93
+ | VAE BatchNorm statistics | 0 | exact |
94
+
95
+ The bf16 figures for a full-depth DiT are larger - 2e-02 for both models - and that is rounding order, not a defect: at float32 the same code agrees to 1e-06, and the deviation grows with depth from a bf16-level 1e-04 per block. Comparing two bf16 implementations below 1e-02 is not meaningful for a 30-block network.
96
+
97
+ </details>
98
+
99
+ <details>
100
+ <summary><b>Performance on an M1 Max</b></summary>
101
+
102
+ Mac Studio, Apple M1 Max, 10-core CPU, 24-core GPU, 32 GB, macOS 26.5. Release build, 4-bit weights, batch 1, warm page cache.
103
+
104
+ | run | per step | VAE decode |
105
+ |--------------------------------------|----------|------------|
106
+ | Z-Image, 512 px, 4 steps | 2.6 s | 1.5 s |
107
+ | klein, 512 px, 4 steps | 2.3 s | 0.9 s |
108
+ | klein, 1024 px, 4 steps | 8.0 s | 0.2 s |
109
+ | klein, 512 px + one 512 px reference | 4.1 s | 0.9 s |
110
+
111
+ Reading the weights dominates a cold run and depends entirely on the storage: on an external volume delivering about 0.2 GiB/s, the 4-bit encoder took 22-32 s and a 4-bit DiT 18-35 s, against 60-107 s for the same weights in bf16. On internal storage expect these to be several times shorter. Compute is unaffected: the numbers above are steady state.
112
+
113
+ Stages run one at a time and release their weights, so peak memory is set by the largest single component rather than their sum. The 4-bit DiT is 2.03 GiB for klein and 3.40 GiB for Z-Image; activations at 1024 px add to that, and no separate peak measurement was made.
114
+
115
+ </details>
116
+
117
+ <details>
118
+ <summary><b>Limitations</b></summary>
119
+
120
+ - Seeds are not compatible with the upstream pipelines, which use `torch.Generator`. The same prompt gives comparable images, never the same file.
121
+ - Both models are distilled: CFG is not applied, and step counts are low by design (8 for Z-Image, 4 for klein).
122
+ - Z-Image's Omni mode is not covered: it needs a SigLIP encoder that is not part of this bundle.
123
+ - Tiled VAE decode is available for high resolutions and gives a result that differs slightly from a single pass, exactly as it does in diffusers.
124
+ - Batches larger than one are not implemented.
125
+
126
+ </details>
127
+
128
+ ## Licence
129
+
130
+ [Apache-2.0](https://www.apache.org/licenses/LICENSE-2.0), the licence of the upstream model. Conversion tooling and this card: MIT. The upstream card states usage restrictions that redistribution does not repeal.
flux2-klein-4b/README.md ADDED
@@ -0,0 +1,134 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ library_name: mlx
4
+ tags:
5
+ - mlx
6
+ - mlx-swift
7
+ - apple-silicon
8
+ base_model:
9
+ - black-forest-labs/FLUX.2-klein-4B
10
+ ---
11
+
12
+ <div align="center">
13
+
14
+ # FLUX.2 [klein] 4B for MLX, bf16
15
+
16
+ **Text to image and multi-reference editing, 4B, 4 steps**
17
+
18
+ ![precision](https://img.shields.io/badge/precision-bf16-6f42c1) ![size](https://img.shields.io/badge/size-7.39%20GiB-1f6feb) [![licence](https://img.shields.io/badge/licence-Apache--2.0-brightgreen)](https://www.apache.org/licenses/LICENSE-2.0) [![part of](https://img.shields.io/badge/part%20of-studio--mlx-black)](https://huggingface.co/kruatech/studio-mlx)
19
+
20
+ </div>
21
+
22
+ 4B rectified flow transformer, text to image and multi-reference editing, 4 steps, distilled so guidance is not applied.
23
+
24
+ Converted from [black-forest-labs/FLUX.2-klein-4B](https://huggingface.co/black-forest-labs/FLUX.2-klein-4B) for native MLX inference on Apple silicon.
25
+
26
+ ## Download
27
+
28
+ Also download `shared/qwen3-4b-text-encoder`.
29
+
30
+ ```bash
31
+ pip install "huggingface_hub[hf_xet]"
32
+
33
+ hf download kruatech/studio-mlx --local-dir bundles \
34
+ --include "shared/qwen3-4b-text-encoder/*" "flux2-klein-4b/*"
35
+ ```
36
+
37
+ ## Files
38
+
39
+ | component | class | precision | tensors | size |
40
+ |-------------|-------------------------|-----------|---------|----------|
41
+ | transformer | Flux2Transformer2DModel | keep | 169 | 7.22 GiB |
42
+ | vae | AutoencoderKLFlux2 | keep | 251 | 160 MiB |
43
+
44
+ Folder total: **7.39 GiB**. `manifest.json` carries SHA-256, byte sizes and tensor counts for every file.
45
+
46
+ ## How this model works in MLX
47
+
48
+ - 5 double-stream blocks with separate text and image modulation, then 20 single-stream blocks where QKV and the MLP share one fused projection.
49
+ - Rotary embedding over four axes, 32 each, theta 2000.
50
+ - No `scaling_factor`: latents are patchified 2x2 into 128 channels and normalized with the VAE's BatchNorm running statistics, which are in `weights/vae.safetensors`.
51
+ - Sigmas use the exponential dynamic shift with mu from the upstream empirical formula.
52
+ - Reference-image editing needs no KV cache: reference tokens and their ids are appended to the sequence and the prediction is sliced back. Each 1024 px reference adds 4096 tokens, which roughly doubles the cost of a step.
53
+
54
+ <details>
55
+ <summary><b>Tensor layout</b></summary>
56
+
57
+ MLX convolutions expect channels last, PyTorch expects channels first, so convolution weights are permuted during conversion. Everything else keeps its upstream shape.
58
+
59
+ | kind | PyTorch | MLX |
60
+ |-----------------------------|---------------------|---------------------|
61
+ | `Conv2d.weight` | `(out, in, kH, kW)` | `(out, kH, kW, in)` |
62
+ | `Linear`, norms, embeddings | `(out, in)` | unchanged |
63
+
64
+ No `weight_norm` and no 3-D convolutions appear in these models, so no folding was needed. Parameter names match the upstream checkpoints, so weights load without remapping.
65
+
66
+ </details>
67
+
68
+ <details>
69
+ <summary><b>Verification numbers</b></summary>
70
+
71
+ Every module was compared against the upstream reference on fixed inputs in float32. The metric is `rel_max = max|a-b| / max|a|`.
72
+
73
+ | module | rel_max | reference |
74
+ |-------------------------------------------------|-----------|---------------------------------|
75
+ | Qwen3 encoder, hidden_states[-2] | 2.7e-07 | transformers |
76
+ | Qwen3 encoder, layers 9/18/27 with padding mask | 6.0e-07 | transformers |
77
+ | Z-Image DiT | 4.2e-06 | diffusers |
78
+ | Flux2 DiT | 3.7e-07 | diffusers |
79
+ | VAE decoder | 1.3e-05 | diffusers |
80
+ | VAE encoder | 5.3e-06 | diffusers |
81
+ | VAE tiled decode | 5.1e-06 | diffusers tiled_decode |
82
+ | sigma schedule, static shift | 3.2e-08 | FlowMatchEulerDiscreteScheduler |
83
+ | sigma schedule, exponential dynamic shift | 7.7e-08 | FlowMatchEulerDiscreteScheduler |
84
+ | position ids, latent packing, patchify | bit-exact | pipeline helpers |
85
+
86
+ A Swift MLX implementation was then checked against the Python one:
87
+
88
+ | module | rel_max | note |
89
+ |------------------------------|-----------------|------------------------------------------------|
90
+ | tokenization, both pipelines | bit-exact | same ids |
91
+ | Qwen3 encoder, bf16 | 2.2e-04 | Z-Image branch |
92
+ | Qwen3 encoder, bf16 | 4.6e-03 | klein branch, 512 tokens with mask |
93
+ | Qwen3 encoder, 4-bit | 3.0e-04 | proves the quantized layout is rebuilt exactly |
94
+ | Z-Image DiT, float32 | 3.7e-07 | 8 layers |
95
+ | Flux2 DiT, float32 | 1.0e-06 | 2 double + 2 single blocks |
96
+ | VAE decode / tiled / encode | 1e-05 or better | both VAE classes |
97
+ | VAE BatchNorm statistics | 0 | exact |
98
+
99
+ The bf16 figures for a full-depth DiT are larger - 2e-02 for both models - and that is rounding order, not a defect: at float32 the same code agrees to 1e-06, and the deviation grows with depth from a bf16-level 1e-04 per block. Comparing two bf16 implementations below 1e-02 is not meaningful for a 30-block network.
100
+
101
+ </details>
102
+
103
+ <details>
104
+ <summary><b>Performance on an M1 Max</b></summary>
105
+
106
+ Mac Studio, Apple M1 Max, 10-core CPU, 24-core GPU, 32 GB, macOS 26.5. Release build, 4-bit weights, batch 1, warm page cache.
107
+
108
+ | run | per step | VAE decode |
109
+ |--------------------------------------|----------|------------|
110
+ | Z-Image, 512 px, 4 steps | 2.6 s | 1.5 s |
111
+ | klein, 512 px, 4 steps | 2.3 s | 0.9 s |
112
+ | klein, 1024 px, 4 steps | 8.0 s | 0.2 s |
113
+ | klein, 512 px + one 512 px reference | 4.1 s | 0.9 s |
114
+
115
+ Reading the weights dominates a cold run and depends entirely on the storage: on an external volume delivering about 0.2 GiB/s, the 4-bit encoder took 22-32 s and a 4-bit DiT 18-35 s, against 60-107 s for the same weights in bf16. On internal storage expect these to be several times shorter. Compute is unaffected: the numbers above are steady state.
116
+
117
+ Stages run one at a time and release their weights, so peak memory is set by the largest single component rather than their sum. The 4-bit DiT is 2.03 GiB for klein and 3.40 GiB for Z-Image; activations at 1024 px add to that, and no separate peak measurement was made.
118
+
119
+ </details>
120
+
121
+ <details>
122
+ <summary><b>Limitations</b></summary>
123
+
124
+ - Seeds are not compatible with the upstream pipelines, which use `torch.Generator`. The same prompt gives comparable images, never the same file.
125
+ - Both models are distilled: CFG is not applied, and step counts are low by design (8 for Z-Image, 4 for klein).
126
+ - Z-Image's Omni mode is not covered: it needs a SigLIP encoder that is not part of this bundle.
127
+ - Tiled VAE decode is available for high resolutions and gives a result that differs slightly from a single pass, exactly as it does in diffusers.
128
+ - Batches larger than one are not implemented.
129
+
130
+ </details>
131
+
132
+ ## Licence
133
+
134
+ [Apache-2.0](https://www.apache.org/licenses/LICENSE-2.0), the licence of the upstream model. Conversion tooling and this card: MIT. The upstream card states usage restrictions that redistribution does not repeal.
shared/qwen3-4b-text-encoder-q4/README.md ADDED
@@ -0,0 +1,130 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ library_name: mlx
4
+ tags:
5
+ - mlx
6
+ - mlx-swift
7
+ - apple-silicon
8
+ base_model:
9
+ - black-forest-labs/FLUX.2-klein-4B
10
+ ---
11
+
12
+ <div align="center">
13
+
14
+ # Qwen3-4B text encoder for MLX, 4-bit
15
+
16
+ **Shared text encoder for both image models**
17
+
18
+ ![precision](https://img.shields.io/badge/precision-4--bit-6f42c1) ![size](https://img.shields.io/badge/size-2.59%20GiB-1f6feb) [![licence](https://img.shields.io/badge/licence-Apache--2.0-brightgreen)](https://www.apache.org/licenses/LICENSE-2.0) [![part of](https://img.shields.io/badge/part%20of-studio--mlx-black)](https://huggingface.co/kruatech/studio-mlx)
19
+
20
+ </div>
21
+
22
+ 4-bit version of the shared text encoder. Everything in the bf16 card applies.
23
+
24
+ Converted from [black-forest-labs/FLUX.2-klein-4B](https://huggingface.co/black-forest-labs/FLUX.2-klein-4B) for native MLX inference on Apple silicon.
25
+
26
+ ## Download
27
+
28
+ Pair it with `z-image-turbo-q4` or `flux2-klein-4b-q4`.
29
+
30
+ ```bash
31
+ pip install "huggingface_hub[hf_xet]"
32
+
33
+ hf download kruatech/studio-mlx --local-dir bundles \
34
+ --include "shared/qwen3-4b-text-encoder-q4/*"
35
+ ```
36
+
37
+ ## Files
38
+
39
+ | component | class | precision | tensors | size |
40
+ |--------------|------------------|-----------------|---------|----------|
41
+ | text_encoder | Qwen3ForCausalLM | 4-bit, group 64 | 877 | 2.58 GiB |
42
+
43
+ Folder total: **2.59 GiB**. `manifest.json` carries SHA-256, byte sizes and tensor counts for every file.
44
+
45
+ ## How this model works in MLX
46
+
47
+ - The embedding table is left in bf16 on purpose: quantizing it moved the encoder output by 3e-01 in testing, which is not worth 0.5 GiB.
48
+ - 245 linear layers are quantized, 4 bits with a bf16 scale and bias per group of 64.
49
+
50
+ <details>
51
+ <summary><b>Tensor layout</b></summary>
52
+
53
+ MLX convolutions expect channels last, PyTorch expects channels first, so convolution weights are permuted during conversion. Everything else keeps its upstream shape.
54
+
55
+ | kind | PyTorch | MLX |
56
+ |-----------------------------|---------------------|---------------------|
57
+ | `Conv2d.weight` | `(out, in, kH, kW)` | `(out, kH, kW, in)` |
58
+ | `Linear`, norms, embeddings | `(out, in)` | unchanged |
59
+
60
+ No `weight_norm` and no 3-D convolutions appear in these models, so no folding was needed. Parameter names match the upstream checkpoints, so weights load without remapping.
61
+
62
+ </details>
63
+
64
+ <details>
65
+ <summary><b>Verification numbers</b></summary>
66
+
67
+ Every module was compared against the upstream reference on fixed inputs in float32. The metric is `rel_max = max|a-b| / max|a|`.
68
+
69
+ | module | rel_max | reference |
70
+ |-------------------------------------------------|-----------|---------------------------------|
71
+ | Qwen3 encoder, hidden_states[-2] | 2.7e-07 | transformers |
72
+ | Qwen3 encoder, layers 9/18/27 with padding mask | 6.0e-07 | transformers |
73
+ | Z-Image DiT | 4.2e-06 | diffusers |
74
+ | Flux2 DiT | 3.7e-07 | diffusers |
75
+ | VAE decoder | 1.3e-05 | diffusers |
76
+ | VAE encoder | 5.3e-06 | diffusers |
77
+ | VAE tiled decode | 5.1e-06 | diffusers tiled_decode |
78
+ | sigma schedule, static shift | 3.2e-08 | FlowMatchEulerDiscreteScheduler |
79
+ | sigma schedule, exponential dynamic shift | 7.7e-08 | FlowMatchEulerDiscreteScheduler |
80
+ | position ids, latent packing, patchify | bit-exact | pipeline helpers |
81
+
82
+ A Swift MLX implementation was then checked against the Python one:
83
+
84
+ | module | rel_max | note |
85
+ |------------------------------|-----------------|------------------------------------------------|
86
+ | tokenization, both pipelines | bit-exact | same ids |
87
+ | Qwen3 encoder, bf16 | 2.2e-04 | Z-Image branch |
88
+ | Qwen3 encoder, bf16 | 4.6e-03 | klein branch, 512 tokens with mask |
89
+ | Qwen3 encoder, 4-bit | 3.0e-04 | proves the quantized layout is rebuilt exactly |
90
+ | Z-Image DiT, float32 | 3.7e-07 | 8 layers |
91
+ | Flux2 DiT, float32 | 1.0e-06 | 2 double + 2 single blocks |
92
+ | VAE decode / tiled / encode | 1e-05 or better | both VAE classes |
93
+ | VAE BatchNorm statistics | 0 | exact |
94
+
95
+ The bf16 figures for a full-depth DiT are larger - 2e-02 for both models - and that is rounding order, not a defect: at float32 the same code agrees to 1e-06, and the deviation grows with depth from a bf16-level 1e-04 per block. Comparing two bf16 implementations below 1e-02 is not meaningful for a 30-block network.
96
+
97
+ </details>
98
+
99
+ <details>
100
+ <summary><b>Performance on an M1 Max</b></summary>
101
+
102
+ Mac Studio, Apple M1 Max, 10-core CPU, 24-core GPU, 32 GB, macOS 26.5. Release build, 4-bit weights, batch 1, warm page cache.
103
+
104
+ | run | per step | VAE decode |
105
+ |--------------------------------------|----------|------------|
106
+ | Z-Image, 512 px, 4 steps | 2.6 s | 1.5 s |
107
+ | klein, 512 px, 4 steps | 2.3 s | 0.9 s |
108
+ | klein, 1024 px, 4 steps | 8.0 s | 0.2 s |
109
+ | klein, 512 px + one 512 px reference | 4.1 s | 0.9 s |
110
+
111
+ Reading the weights dominates a cold run and depends entirely on the storage: on an external volume delivering about 0.2 GiB/s, the 4-bit encoder took 22-32 s and a 4-bit DiT 18-35 s, against 60-107 s for the same weights in bf16. On internal storage expect these to be several times shorter. Compute is unaffected: the numbers above are steady state.
112
+
113
+ Stages run one at a time and release their weights, so peak memory is set by the largest single component rather than their sum. The 4-bit DiT is 2.03 GiB for klein and 3.40 GiB for Z-Image; activations at 1024 px add to that, and no separate peak measurement was made.
114
+
115
+ </details>
116
+
117
+ <details>
118
+ <summary><b>Limitations</b></summary>
119
+
120
+ - Seeds are not compatible with the upstream pipelines, which use `torch.Generator`. The same prompt gives comparable images, never the same file.
121
+ - Both models are distilled: CFG is not applied, and step counts are low by design (8 for Z-Image, 4 for klein).
122
+ - Z-Image's Omni mode is not covered: it needs a SigLIP encoder that is not part of this bundle.
123
+ - Tiled VAE decode is available for high resolutions and gives a result that differs slightly from a single pass, exactly as it does in diffusers.
124
+ - Batches larger than one are not implemented.
125
+
126
+ </details>
127
+
128
+ ## Licence
129
+
130
+ [Apache-2.0](https://www.apache.org/licenses/LICENSE-2.0), the licence of the upstream model. Conversion tooling and this card: MIT. The upstream card states usage restrictions that redistribution does not repeal.
shared/qwen3-4b-text-encoder/README.md ADDED
@@ -0,0 +1,133 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ library_name: mlx
4
+ tags:
5
+ - mlx
6
+ - mlx-swift
7
+ - apple-silicon
8
+ base_model:
9
+ - black-forest-labs/FLUX.2-klein-4B
10
+ ---
11
+
12
+ <div align="center">
13
+
14
+ # Qwen3-4B text encoder for MLX, bf16
15
+
16
+ **Shared text encoder for both image models**
17
+
18
+ ![precision](https://img.shields.io/badge/precision-bf16-6f42c1) ![size](https://img.shields.io/badge/size-7.51%20GiB-1f6feb) [![licence](https://img.shields.io/badge/licence-Apache--2.0-brightgreen)](https://www.apache.org/licenses/LICENSE-2.0) [![part of](https://img.shields.io/badge/part%20of-studio--mlx-black)](https://huggingface.co/kruatech/studio-mlx)
19
+
20
+ </div>
21
+
22
+ The text encoder both image models in this repository use. It is the Qwen3-4B encoder shipped inside the FLUX.2 [klein] and Z-Image-Turbo repositories; the two copies are identical except for the MLP of layer 35, which no pipeline reads.
23
+
24
+ Converted from [black-forest-labs/FLUX.2-klein-4B](https://huggingface.co/black-forest-labs/FLUX.2-klein-4B) for native MLX inference on Apple silicon.
25
+
26
+ ## Download
27
+
28
+ Pair it with `z-image-turbo` or `flux2-klein-4b`. On its own it generates nothing.
29
+
30
+ ```bash
31
+ pip install "huggingface_hub[hf_xet]"
32
+
33
+ hf download kruatech/studio-mlx --local-dir bundles \
34
+ --include "shared/qwen3-4b-text-encoder/*"
35
+ ```
36
+
37
+ ## Files
38
+
39
+ | component | class | precision | tensors | size |
40
+ |--------------|------------------|-----------|---------|----------|
41
+ | text_encoder | Qwen3ForCausalLM | keep | 398 | 7.51 GiB |
42
+
43
+ Folder total: **7.51 GiB**. `manifest.json` carries SHA-256, byte sizes and tensor counts for every file.
44
+
45
+ ## How this model works in MLX
46
+
47
+ - Z-Image consumes `hidden_states[-2]`, the output of layer n-1, so 35 of 36 layers run.
48
+ - klein concatenates `hidden_states[9]`, `[18]` and `[27]` into 7680 features, so 27 of 36 layers run.
49
+ - Neither pipeline touches the last layer, so a loader can skip it entirely.
50
+ - Padding matters for klein and not for Z-Image: Z-Image drops padded positions with a mask afterwards, while klein feeds all 512 tokens to the transformer, so the attention mask has to be applied.
51
+ - `tokenizer/chat_format.json` holds the Qwen chat template already rendered into a prefix and a suffix for each pipeline, so a runtime needs no Jinja. klein renders with thinking disabled, which still injects an empty `<think></think>` block.
52
+
53
+ <details>
54
+ <summary><b>Tensor layout</b></summary>
55
+
56
+ MLX convolutions expect channels last, PyTorch expects channels first, so convolution weights are permuted during conversion. Everything else keeps its upstream shape.
57
+
58
+ | kind | PyTorch | MLX |
59
+ |-----------------------------|---------------------|---------------------|
60
+ | `Conv2d.weight` | `(out, in, kH, kW)` | `(out, kH, kW, in)` |
61
+ | `Linear`, norms, embeddings | `(out, in)` | unchanged |
62
+
63
+ No `weight_norm` and no 3-D convolutions appear in these models, so no folding was needed. Parameter names match the upstream checkpoints, so weights load without remapping.
64
+
65
+ </details>
66
+
67
+ <details>
68
+ <summary><b>Verification numbers</b></summary>
69
+
70
+ Every module was compared against the upstream reference on fixed inputs in float32. The metric is `rel_max = max|a-b| / max|a|`.
71
+
72
+ | module | rel_max | reference |
73
+ |-------------------------------------------------|-----------|---------------------------------|
74
+ | Qwen3 encoder, hidden_states[-2] | 2.7e-07 | transformers |
75
+ | Qwen3 encoder, layers 9/18/27 with padding mask | 6.0e-07 | transformers |
76
+ | Z-Image DiT | 4.2e-06 | diffusers |
77
+ | Flux2 DiT | 3.7e-07 | diffusers |
78
+ | VAE decoder | 1.3e-05 | diffusers |
79
+ | VAE encoder | 5.3e-06 | diffusers |
80
+ | VAE tiled decode | 5.1e-06 | diffusers tiled_decode |
81
+ | sigma schedule, static shift | 3.2e-08 | FlowMatchEulerDiscreteScheduler |
82
+ | sigma schedule, exponential dynamic shift | 7.7e-08 | FlowMatchEulerDiscreteScheduler |
83
+ | position ids, latent packing, patchify | bit-exact | pipeline helpers |
84
+
85
+ A Swift MLX implementation was then checked against the Python one:
86
+
87
+ | module | rel_max | note |
88
+ |------------------------------|-----------------|------------------------------------------------|
89
+ | tokenization, both pipelines | bit-exact | same ids |
90
+ | Qwen3 encoder, bf16 | 2.2e-04 | Z-Image branch |
91
+ | Qwen3 encoder, bf16 | 4.6e-03 | klein branch, 512 tokens with mask |
92
+ | Qwen3 encoder, 4-bit | 3.0e-04 | proves the quantized layout is rebuilt exactly |
93
+ | Z-Image DiT, float32 | 3.7e-07 | 8 layers |
94
+ | Flux2 DiT, float32 | 1.0e-06 | 2 double + 2 single blocks |
95
+ | VAE decode / tiled / encode | 1e-05 or better | both VAE classes |
96
+ | VAE BatchNorm statistics | 0 | exact |
97
+
98
+ The bf16 figures for a full-depth DiT are larger - 2e-02 for both models - and that is rounding order, not a defect: at float32 the same code agrees to 1e-06, and the deviation grows with depth from a bf16-level 1e-04 per block. Comparing two bf16 implementations below 1e-02 is not meaningful for a 30-block network.
99
+
100
+ </details>
101
+
102
+ <details>
103
+ <summary><b>Performance on an M1 Max</b></summary>
104
+
105
+ Mac Studio, Apple M1 Max, 10-core CPU, 24-core GPU, 32 GB, macOS 26.5. Release build, 4-bit weights, batch 1, warm page cache.
106
+
107
+ | run | per step | VAE decode |
108
+ |--------------------------------------|----------|------------|
109
+ | Z-Image, 512 px, 4 steps | 2.6 s | 1.5 s |
110
+ | klein, 512 px, 4 steps | 2.3 s | 0.9 s |
111
+ | klein, 1024 px, 4 steps | 8.0 s | 0.2 s |
112
+ | klein, 512 px + one 512 px reference | 4.1 s | 0.9 s |
113
+
114
+ Reading the weights dominates a cold run and depends entirely on the storage: on an external volume delivering about 0.2 GiB/s, the 4-bit encoder took 22-32 s and a 4-bit DiT 18-35 s, against 60-107 s for the same weights in bf16. On internal storage expect these to be several times shorter. Compute is unaffected: the numbers above are steady state.
115
+
116
+ Stages run one at a time and release their weights, so peak memory is set by the largest single component rather than their sum. The 4-bit DiT is 2.03 GiB for klein and 3.40 GiB for Z-Image; activations at 1024 px add to that, and no separate peak measurement was made.
117
+
118
+ </details>
119
+
120
+ <details>
121
+ <summary><b>Limitations</b></summary>
122
+
123
+ - Seeds are not compatible with the upstream pipelines, which use `torch.Generator`. The same prompt gives comparable images, never the same file.
124
+ - Both models are distilled: CFG is not applied, and step counts are low by design (8 for Z-Image, 4 for klein).
125
+ - Z-Image's Omni mode is not covered: it needs a SigLIP encoder that is not part of this bundle.
126
+ - Tiled VAE decode is available for high resolutions and gives a result that differs slightly from a single pass, exactly as it does in diffusers.
127
+ - Batches larger than one are not implemented.
128
+
129
+ </details>
130
+
131
+ ## Licence
132
+
133
+ [Apache-2.0](https://www.apache.org/licenses/LICENSE-2.0), the licence of the upstream model. Conversion tooling and this card: MIT. The upstream card states usage restrictions that redistribution does not repeal.
z-image-turbo-q4/README.md ADDED
@@ -0,0 +1,130 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ library_name: mlx
4
+ tags:
5
+ - mlx
6
+ - mlx-swift
7
+ - apple-silicon
8
+ base_model:
9
+ - Tongyi-MAI/Z-Image-Turbo
10
+ ---
11
+
12
+ <div align="center">
13
+
14
+ # Z-Image-Turbo for MLX, 4-bit
15
+
16
+ **Text to image, 6B, 8 steps**
17
+
18
+ ![precision](https://img.shields.io/badge/precision-4--bit-6f42c1) ![size](https://img.shields.io/badge/size-3.55%20GiB-1f6feb) [![licence](https://img.shields.io/badge/licence-Apache--2.0-brightgreen)](https://www.apache.org/licenses/LICENSE-2.0) [![part of](https://img.shields.io/badge/part%20of-studio--mlx-black)](https://huggingface.co/kruatech/studio-mlx)
19
+
20
+ </div>
21
+
22
+ 4-bit version of Z-Image-Turbo. Everything in the bf16 card applies.
23
+
24
+ Converted from [Tongyi-MAI/Z-Image-Turbo](https://huggingface.co/Tongyi-MAI/Z-Image-Turbo) for native MLX inference on Apple silicon.
25
+
26
+ ## Download
27
+
28
+ Also download `shared/qwen3-4b-text-encoder-q4`.
29
+
30
+ ```bash
31
+ pip install "huggingface_hub[hf_xet]"
32
+
33
+ hf download kruatech/studio-mlx --local-dir bundles \
34
+ --include "shared/qwen3-4b-text-encoder-q4/*" "z-image-turbo-q4/*"
35
+ ```
36
+
37
+ ## Files
38
+
39
+ | component | class | precision | tensors | size |
40
+ |-------------|--------------------------|-----------------|---------|----------|
41
+ | transformer | ZImageTransformer2DModel | 4-bit, group 64 | 999 | 3.40 GiB |
42
+ | vae | AutoencoderKL | keep | 244 | 160 MiB |
43
+
44
+ Folder total: **3.55 GiB**. `manifest.json` carries SHA-256, byte sizes and tensor counts for every file.
45
+
46
+ ## How this model works in MLX
47
+
48
+ - 239 linear layers quantized; normalizations, modulations and timestep MLPs stay in bf16. The VAE is not quantized - at 160 MiB there is nothing to gain.
49
+
50
+ <details>
51
+ <summary><b>Tensor layout</b></summary>
52
+
53
+ MLX convolutions expect channels last, PyTorch expects channels first, so convolution weights are permuted during conversion. Everything else keeps its upstream shape.
54
+
55
+ | kind | PyTorch | MLX |
56
+ |-----------------------------|---------------------|---------------------|
57
+ | `Conv2d.weight` | `(out, in, kH, kW)` | `(out, kH, kW, in)` |
58
+ | `Linear`, norms, embeddings | `(out, in)` | unchanged |
59
+
60
+ No `weight_norm` and no 3-D convolutions appear in these models, so no folding was needed. Parameter names match the upstream checkpoints, so weights load without remapping.
61
+
62
+ </details>
63
+
64
+ <details>
65
+ <summary><b>Verification numbers</b></summary>
66
+
67
+ Every module was compared against the upstream reference on fixed inputs in float32. The metric is `rel_max = max|a-b| / max|a|`.
68
+
69
+ | module | rel_max | reference |
70
+ |-------------------------------------------------|-----------|---------------------------------|
71
+ | Qwen3 encoder, hidden_states[-2] | 2.7e-07 | transformers |
72
+ | Qwen3 encoder, layers 9/18/27 with padding mask | 6.0e-07 | transformers |
73
+ | Z-Image DiT | 4.2e-06 | diffusers |
74
+ | Flux2 DiT | 3.7e-07 | diffusers |
75
+ | VAE decoder | 1.3e-05 | diffusers |
76
+ | VAE encoder | 5.3e-06 | diffusers |
77
+ | VAE tiled decode | 5.1e-06 | diffusers tiled_decode |
78
+ | sigma schedule, static shift | 3.2e-08 | FlowMatchEulerDiscreteScheduler |
79
+ | sigma schedule, exponential dynamic shift | 7.7e-08 | FlowMatchEulerDiscreteScheduler |
80
+ | position ids, latent packing, patchify | bit-exact | pipeline helpers |
81
+
82
+ A Swift MLX implementation was then checked against the Python one:
83
+
84
+ | module | rel_max | note |
85
+ |------------------------------|-----------------|------------------------------------------------|
86
+ | tokenization, both pipelines | bit-exact | same ids |
87
+ | Qwen3 encoder, bf16 | 2.2e-04 | Z-Image branch |
88
+ | Qwen3 encoder, bf16 | 4.6e-03 | klein branch, 512 tokens with mask |
89
+ | Qwen3 encoder, 4-bit | 3.0e-04 | proves the quantized layout is rebuilt exactly |
90
+ | Z-Image DiT, float32 | 3.7e-07 | 8 layers |
91
+ | Flux2 DiT, float32 | 1.0e-06 | 2 double + 2 single blocks |
92
+ | VAE decode / tiled / encode | 1e-05 or better | both VAE classes |
93
+ | VAE BatchNorm statistics | 0 | exact |
94
+
95
+ The bf16 figures for a full-depth DiT are larger - 2e-02 for both models - and that is rounding order, not a defect: at float32 the same code agrees to 1e-06, and the deviation grows with depth from a bf16-level 1e-04 per block. Comparing two bf16 implementations below 1e-02 is not meaningful for a 30-block network.
96
+
97
+ </details>
98
+
99
+ <details>
100
+ <summary><b>Performance on an M1 Max</b></summary>
101
+
102
+ Mac Studio, Apple M1 Max, 10-core CPU, 24-core GPU, 32 GB, macOS 26.5. Release build, 4-bit weights, batch 1, warm page cache.
103
+
104
+ | run | per step | VAE decode |
105
+ |--------------------------------------|----------|------------|
106
+ | Z-Image, 512 px, 4 steps | 2.6 s | 1.5 s |
107
+ | klein, 512 px, 4 steps | 2.3 s | 0.9 s |
108
+ | klein, 1024 px, 4 steps | 8.0 s | 0.2 s |
109
+ | klein, 512 px + one 512 px reference | 4.1 s | 0.9 s |
110
+
111
+ Reading the weights dominates a cold run and depends entirely on the storage: on an external volume delivering about 0.2 GiB/s, the 4-bit encoder took 22-32 s and a 4-bit DiT 18-35 s, against 60-107 s for the same weights in bf16. On internal storage expect these to be several times shorter. Compute is unaffected: the numbers above are steady state.
112
+
113
+ Stages run one at a time and release their weights, so peak memory is set by the largest single component rather than their sum. The 4-bit DiT is 2.03 GiB for klein and 3.40 GiB for Z-Image; activations at 1024 px add to that, and no separate peak measurement was made.
114
+
115
+ </details>
116
+
117
+ <details>
118
+ <summary><b>Limitations</b></summary>
119
+
120
+ - Seeds are not compatible with the upstream pipelines, which use `torch.Generator`. The same prompt gives comparable images, never the same file.
121
+ - Both models are distilled: CFG is not applied, and step counts are low by design (8 for Z-Image, 4 for klein).
122
+ - Z-Image's Omni mode is not covered: it needs a SigLIP encoder that is not part of this bundle.
123
+ - Tiled VAE decode is available for high resolutions and gives a result that differs slightly from a single pass, exactly as it does in diffusers.
124
+ - Batches larger than one are not implemented.
125
+
126
+ </details>
127
+
128
+ ## Licence
129
+
130
+ [Apache-2.0](https://www.apache.org/licenses/LICENSE-2.0), the licence of the upstream model. Conversion tooling and this card: MIT. The upstream card states usage restrictions that redistribution does not repeal.
z-image-turbo/README.md ADDED
@@ -0,0 +1,134 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ library_name: mlx
4
+ tags:
5
+ - mlx
6
+ - mlx-swift
7
+ - apple-silicon
8
+ base_model:
9
+ - Tongyi-MAI/Z-Image-Turbo
10
+ ---
11
+
12
+ <div align="center">
13
+
14
+ # Z-Image-Turbo for MLX, bf16
15
+
16
+ **Text to image, 6B, 8 steps**
17
+
18
+ ![precision](https://img.shields.io/badge/precision-bf16-6f42c1) ![size](https://img.shields.io/badge/size-11.62%20GiB-1f6feb) [![licence](https://img.shields.io/badge/licence-Apache--2.0-brightgreen)](https://www.apache.org/licenses/LICENSE-2.0) [![part of](https://img.shields.io/badge/part%20of-studio--mlx-black)](https://huggingface.co/kruatech/studio-mlx)
19
+
20
+ </div>
21
+
22
+ 6B single-stream DiT (S3-DiT) plus the flux-dev VAE, text to image, 8 steps, no CFG.
23
+
24
+ Converted from [Tongyi-MAI/Z-Image-Turbo](https://huggingface.co/Tongyi-MAI/Z-Image-Turbo) for native MLX inference on Apple silicon.
25
+
26
+ ## Download
27
+
28
+ Also download `shared/qwen3-4b-text-encoder`. This folder has no `tokenizer/`: the tokenizer lives with the encoder, and without it a prompt cannot be built.
29
+
30
+ ```bash
31
+ pip install "huggingface_hub[hf_xet]"
32
+
33
+ hf download kruatech/studio-mlx --local-dir bundles \
34
+ --include "shared/qwen3-4b-text-encoder/*" "z-image-turbo/*"
35
+ ```
36
+
37
+ ## Files
38
+
39
+ | component | class | precision | tensors | size |
40
+ |-------------|--------------------------|-----------|---------|-----------|
41
+ | transformer | ZImageTransformer2DModel | bf16 | 521 | 11.46 GiB |
42
+ | vae | AutoencoderKL | keep | 244 | 160 MiB |
43
+
44
+ Folder total: **11.62 GiB**. `manifest.json` carries SHA-256, byte sizes and tensor counts for every file.
45
+
46
+ ## How this model works in MLX
47
+
48
+ - 30 main layers, 2 noise-refiner and 2 context-refiner layers, dim 3840, 30 heads.
49
+ - Rotary embedding over three axes with dims 32/48/48 and theta 256, applied to adjacent pairs rather than split halves.
50
+ - Image and caption tokens are padded to a multiple of 32 with learned pad tokens, and the sequence is `[image, caption]`.
51
+ - The scheduler uses the static shift of 3.0 from `scheduler_config.json`, and the prediction is negated before the Euler step.
52
+ - Latents scale by `1/0.3611` and shift by `0.1159` before decoding.
53
+
54
+ <details>
55
+ <summary><b>Tensor layout</b></summary>
56
+
57
+ MLX convolutions expect channels last, PyTorch expects channels first, so convolution weights are permuted during conversion. Everything else keeps its upstream shape.
58
+
59
+ | kind | PyTorch | MLX |
60
+ |-----------------------------|---------------------|---------------------|
61
+ | `Conv2d.weight` | `(out, in, kH, kW)` | `(out, kH, kW, in)` |
62
+ | `Linear`, norms, embeddings | `(out, in)` | unchanged |
63
+
64
+ No `weight_norm` and no 3-D convolutions appear in these models, so no folding was needed. Parameter names match the upstream checkpoints, so weights load without remapping.
65
+
66
+ </details>
67
+
68
+ <details>
69
+ <summary><b>Verification numbers</b></summary>
70
+
71
+ Every module was compared against the upstream reference on fixed inputs in float32. The metric is `rel_max = max|a-b| / max|a|`.
72
+
73
+ | module | rel_max | reference |
74
+ |-------------------------------------------------|-----------|---------------------------------|
75
+ | Qwen3 encoder, hidden_states[-2] | 2.7e-07 | transformers |
76
+ | Qwen3 encoder, layers 9/18/27 with padding mask | 6.0e-07 | transformers |
77
+ | Z-Image DiT | 4.2e-06 | diffusers |
78
+ | Flux2 DiT | 3.7e-07 | diffusers |
79
+ | VAE decoder | 1.3e-05 | diffusers |
80
+ | VAE encoder | 5.3e-06 | diffusers |
81
+ | VAE tiled decode | 5.1e-06 | diffusers tiled_decode |
82
+ | sigma schedule, static shift | 3.2e-08 | FlowMatchEulerDiscreteScheduler |
83
+ | sigma schedule, exponential dynamic shift | 7.7e-08 | FlowMatchEulerDiscreteScheduler |
84
+ | position ids, latent packing, patchify | bit-exact | pipeline helpers |
85
+
86
+ A Swift MLX implementation was then checked against the Python one:
87
+
88
+ | module | rel_max | note |
89
+ |------------------------------|-----------------|------------------------------------------------|
90
+ | tokenization, both pipelines | bit-exact | same ids |
91
+ | Qwen3 encoder, bf16 | 2.2e-04 | Z-Image branch |
92
+ | Qwen3 encoder, bf16 | 4.6e-03 | klein branch, 512 tokens with mask |
93
+ | Qwen3 encoder, 4-bit | 3.0e-04 | proves the quantized layout is rebuilt exactly |
94
+ | Z-Image DiT, float32 | 3.7e-07 | 8 layers |
95
+ | Flux2 DiT, float32 | 1.0e-06 | 2 double + 2 single blocks |
96
+ | VAE decode / tiled / encode | 1e-05 or better | both VAE classes |
97
+ | VAE BatchNorm statistics | 0 | exact |
98
+
99
+ The bf16 figures for a full-depth DiT are larger - 2e-02 for both models - and that is rounding order, not a defect: at float32 the same code agrees to 1e-06, and the deviation grows with depth from a bf16-level 1e-04 per block. Comparing two bf16 implementations below 1e-02 is not meaningful for a 30-block network.
100
+
101
+ </details>
102
+
103
+ <details>
104
+ <summary><b>Performance on an M1 Max</b></summary>
105
+
106
+ Mac Studio, Apple M1 Max, 10-core CPU, 24-core GPU, 32 GB, macOS 26.5. Release build, 4-bit weights, batch 1, warm page cache.
107
+
108
+ | run | per step | VAE decode |
109
+ |--------------------------------------|----------|------------|
110
+ | Z-Image, 512 px, 4 steps | 2.6 s | 1.5 s |
111
+ | klein, 512 px, 4 steps | 2.3 s | 0.9 s |
112
+ | klein, 1024 px, 4 steps | 8.0 s | 0.2 s |
113
+ | klein, 512 px + one 512 px reference | 4.1 s | 0.9 s |
114
+
115
+ Reading the weights dominates a cold run and depends entirely on the storage: on an external volume delivering about 0.2 GiB/s, the 4-bit encoder took 22-32 s and a 4-bit DiT 18-35 s, against 60-107 s for the same weights in bf16. On internal storage expect these to be several times shorter. Compute is unaffected: the numbers above are steady state.
116
+
117
+ Stages run one at a time and release their weights, so peak memory is set by the largest single component rather than their sum. The 4-bit DiT is 2.03 GiB for klein and 3.40 GiB for Z-Image; activations at 1024 px add to that, and no separate peak measurement was made.
118
+
119
+ </details>
120
+
121
+ <details>
122
+ <summary><b>Limitations</b></summary>
123
+
124
+ - Seeds are not compatible with the upstream pipelines, which use `torch.Generator`. The same prompt gives comparable images, never the same file.
125
+ - Both models are distilled: CFG is not applied, and step counts are low by design (8 for Z-Image, 4 for klein).
126
+ - Z-Image's Omni mode is not covered: it needs a SigLIP encoder that is not part of this bundle.
127
+ - Tiled VAE decode is available for high resolutions and gives a result that differs slightly from a single pass, exactly as it does in diffusers.
128
+ - Batches larger than one are not implemented.
129
+
130
+ </details>
131
+
132
+ ## Licence
133
+
134
+ [Apache-2.0](https://www.apache.org/licenses/LICENSE-2.0), the licence of the upstream model. Conversion tooling and this card: MIT. The upstream card states usage restrictions that redistribution does not repeal.