Instructions to use kruatech/studio-mlx with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use kruatech/studio-mlx with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] hf download kruatech/studio-mlx --local-dir studio-mlx
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
Upload folder using huggingface_hub
Browse files- .gitattributes +7 -23
- LICENSES.md +22 -0
- README.md +181 -1
- ace-step-1.5-turbo/README.md +213 -0
- flux2-klein-4b-q4/README.md +130 -0
- flux2-klein-4b/README.md +134 -0
- shared/qwen3-4b-text-encoder-q4/README.md +130 -0
- shared/qwen3-4b-text-encoder/README.md +133 -0
- z-image-turbo-q4/README.md +130 -0
- z-image-turbo/README.md +134 -0
.gitattributes
CHANGED
|
@@ -1,35 +1,19 @@
|
|
| 1 |
-
*.
|
| 2 |
-
*.arrow filter=lfs diff=lfs merge=lfs -text
|
| 3 |
*.bin filter=lfs diff=lfs merge=lfs -text
|
| 4 |
-
*.bz2 filter=lfs diff=lfs merge=lfs -text
|
| 5 |
*.ckpt filter=lfs diff=lfs merge=lfs -text
|
| 6 |
-
*.ftz filter=lfs diff=lfs merge=lfs -text
|
| 7 |
*.gz filter=lfs diff=lfs merge=lfs -text
|
| 8 |
*.h5 filter=lfs diff=lfs merge=lfs -text
|
| 9 |
-
*.joblib filter=lfs diff=lfs merge=lfs -text
|
| 10 |
-
*.lfs.* filter=lfs diff=lfs merge=lfs -text
|
| 11 |
-
*.mlmodel filter=lfs diff=lfs merge=lfs -text
|
| 12 |
-
*.model filter=lfs diff=lfs merge=lfs -text
|
| 13 |
*.msgpack filter=lfs diff=lfs merge=lfs -text
|
| 14 |
*.npy filter=lfs diff=lfs merge=lfs -text
|
| 15 |
*.npz filter=lfs diff=lfs merge=lfs -text
|
| 16 |
*.onnx filter=lfs diff=lfs merge=lfs -text
|
| 17 |
-
*.ot filter=lfs diff=lfs merge=lfs -text
|
| 18 |
-
*.parquet filter=lfs diff=lfs merge=lfs -text
|
| 19 |
-
*.pb filter=lfs diff=lfs merge=lfs -text
|
| 20 |
-
*.pickle filter=lfs diff=lfs merge=lfs -text
|
| 21 |
-
*.pkl filter=lfs diff=lfs merge=lfs -text
|
| 22 |
*.pt filter=lfs diff=lfs merge=lfs -text
|
| 23 |
*.pth filter=lfs diff=lfs merge=lfs -text
|
| 24 |
-
*.rar filter=lfs diff=lfs merge=lfs -text
|
| 25 |
-
*.safetensors filter=lfs diff=lfs merge=lfs -text
|
| 26 |
-
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
|
| 27 |
-
*.tar.* filter=lfs diff=lfs merge=lfs -text
|
| 28 |
*.tar filter=lfs diff=lfs merge=lfs -text
|
| 29 |
-
*.tflite filter=lfs diff=lfs merge=lfs -text
|
| 30 |
-
*.tgz filter=lfs diff=lfs merge=lfs -text
|
| 31 |
-
*.wasm filter=lfs diff=lfs merge=lfs -text
|
| 32 |
-
*.xz filter=lfs diff=lfs merge=lfs -text
|
| 33 |
*.zip filter=lfs diff=lfs merge=lfs -text
|
| 34 |
-
*.
|
| 35 |
-
*
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
*.safetensors filter=lfs diff=lfs merge=lfs -text
|
|
|
|
| 2 |
*.bin filter=lfs diff=lfs merge=lfs -text
|
|
|
|
| 3 |
*.ckpt filter=lfs diff=lfs merge=lfs -text
|
|
|
|
| 4 |
*.gz filter=lfs diff=lfs merge=lfs -text
|
| 5 |
*.h5 filter=lfs diff=lfs merge=lfs -text
|
|
|
|
|
|
|
|
|
|
|
|
|
| 6 |
*.msgpack filter=lfs diff=lfs merge=lfs -text
|
| 7 |
*.npy filter=lfs diff=lfs merge=lfs -text
|
| 8 |
*.npz filter=lfs diff=lfs merge=lfs -text
|
| 9 |
*.onnx filter=lfs diff=lfs merge=lfs -text
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 10 |
*.pt filter=lfs diff=lfs merge=lfs -text
|
| 11 |
*.pth filter=lfs diff=lfs merge=lfs -text
|
|
|
|
|
|
|
|
|
|
|
|
|
| 12 |
*.tar filter=lfs diff=lfs merge=lfs -text
|
|
|
|
|
|
|
|
|
|
|
|
|
| 13 |
*.zip filter=lfs diff=lfs merge=lfs -text
|
| 14 |
+
**/tokenizer/tokenizer.json filter=lfs diff=lfs merge=lfs -text
|
| 15 |
+
**/tokenizer/vocab.json filter=lfs diff=lfs merge=lfs -text
|
| 16 |
+
**/tokenizer/merges.txt filter=lfs diff=lfs merge=lfs -text
|
| 17 |
+
**/lm_tokenizer/tokenizer.json filter=lfs diff=lfs merge=lfs -text
|
| 18 |
+
**/lm_tokenizer/vocab.json filter=lfs diff=lfs merge=lfs -text
|
| 19 |
+
**/lm_tokenizer/merges.txt filter=lfs diff=lfs merge=lfs -text
|
LICENSES.md
ADDED
|
@@ -0,0 +1,22 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Licences in this repository
|
| 2 |
+
|
| 3 |
+
Each folder is a derivative of an upstream model and carries that model's licence.
|
| 4 |
+
|
| 5 |
+
| folder | upstream | licence |
|
| 6 |
+
| --------------------------------- | --------------------------------- | ---------- |
|
| 7 |
+
| `shared/qwen3-4b-text-encoder` | inside the two image repos below | Apache-2.0 |
|
| 8 |
+
| `shared/qwen3-4b-text-encoder-q4` | inside the two image repos below | Apache-2.0 |
|
| 9 |
+
| `z-image-turbo` | Tongyi-MAI/Z-Image-Turbo | Apache-2.0 |
|
| 10 |
+
| `z-image-turbo-q4` | Tongyi-MAI/Z-Image-Turbo | Apache-2.0 |
|
| 11 |
+
| `flux2-klein-4b` | black-forest-labs/FLUX.2-klein-4B | Apache-2.0 |
|
| 12 |
+
| `flux2-klein-4b-q4` | black-forest-labs/FLUX.2-klein-4B | Apache-2.0 |
|
| 13 |
+
| `ace-step-1.5-turbo` | ACE-Step/Ace-Step1.5 | MIT |
|
| 14 |
+
|
| 15 |
+
Apache-2.0: https://www.apache.org/licenses/LICENSE-2.0
|
| 16 |
+
MIT: https://opensource.org/license/mit
|
| 17 |
+
|
| 18 |
+
Conversion tooling and the model cards in this repository: MIT.
|
| 19 |
+
|
| 20 |
+
The upstream cards state usage restrictions and responsible-use commitments that
|
| 21 |
+
redistribution does not repeal. See in particular the out-of-scope use section of
|
| 22 |
+
the FLUX.2 [klein] card: https://huggingface.co/black-forest-labs/FLUX.2-klein-4B
|
README.md
CHANGED
|
@@ -1,5 +1,185 @@
|
|
| 1 |
---
|
| 2 |
license: other
|
| 3 |
license_name: mixed-apache-2.0-and-mit
|
| 4 |
-
license_link:
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 5 |
---
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
---
|
| 2 |
license: other
|
| 3 |
license_name: mixed-apache-2.0-and-mit
|
| 4 |
+
license_link: https://huggingface.co/kruatech/studio-mlx/blob/main/LICENSES.md
|
| 5 |
+
library_name: mlx
|
| 6 |
+
tags:
|
| 7 |
+
- mlx
|
| 8 |
+
- mlx-swift
|
| 9 |
+
- apple-silicon
|
| 10 |
+
- text-to-image
|
| 11 |
+
- image-editing
|
| 12 |
+
- text-to-music
|
| 13 |
+
- quantized
|
| 14 |
---
|
| 15 |
+
|
| 16 |
+
<div align="center">
|
| 17 |
+
|
| 18 |
+
# studio-mlx
|
| 19 |
+
|
| 20 |
+
**Image and music models converted for native MLX inference on Apple silicon**
|
| 21 |
+
|
| 22 |
+
   [](https://huggingface.co/kruatech/studio-mlx/blob/main/LICENSES.md)
|
| 23 |
+
|
| 24 |
+
</div>
|
| 25 |
+
|
| 26 |
+
Three models and the text encoder the two image models share. Every folder is a self-contained bundle with its own card, so you download only what you need.
|
| 27 |
+
|
| 28 |
+
Weights are not retrained. Tensor layouts were adapted for MLX and the 4-bit folders are quantized, so parameters are mathematically equivalent to upstream rather than byte-identical.
|
| 29 |
+
|
| 30 |
+
## Quick start
|
| 31 |
+
|
| 32 |
+
Pick one line. The image models need the shared encoder, the music model does not.
|
| 33 |
+
|
| 34 |
+
```bash
|
| 35 |
+
pip install "huggingface_hub[hf_xet]"
|
| 36 |
+
|
| 37 |
+
# Z-Image-Turbo - 6.14 GiB
|
| 38 |
+
hf download kruatech/studio-mlx --local-dir bundles \
|
| 39 |
+
--include "shared/qwen3-4b-text-encoder-q4/*" "z-image-turbo-q4/*"
|
| 40 |
+
|
| 41 |
+
# FLUX.2 [klein] 4B - 4.79 GiB
|
| 42 |
+
hf download kruatech/studio-mlx --local-dir bundles \
|
| 43 |
+
--include "shared/qwen3-4b-text-encoder-q4/*" "flux2-klein-4b-q4/*"
|
| 44 |
+
|
| 45 |
+
# ACE-Step 1.5 Turbo - 9.38 GiB
|
| 46 |
+
hf download kruatech/studio-mlx --local-dir bundles --include "ace-step-1.5-turbo/*"
|
| 47 |
+
|
| 48 |
+
```
|
| 49 |
+
|
| 50 |
+
The `-q4` folders sit at the top level rather than inside the bf16 ones, so an include pattern fetches one precision without dragging in the other.
|
| 51 |
+
|
| 52 |
+
## What is inside
|
| 53 |
+
|
| 54 |
+
| folder | what it does | precision | size | licence |
|
| 55 |
+
|----------------------------------------------------------------------|--------------------------------------------------------|-----------|-----------|------------|
|
| 56 |
+
| [`shared/qwen3-4b-text-encoder`](shared/qwen3-4b-text-encoder) | Shared text encoder for both image models | `bf16` | 7.51 GiB | Apache-2.0 |
|
| 57 |
+
| [`shared/qwen3-4b-text-encoder-q4`](shared/qwen3-4b-text-encoder-q4) | Shared text encoder for both image models | `4-bit` | 2.59 GiB | Apache-2.0 |
|
| 58 |
+
| [`z-image-turbo`](z-image-turbo) | Text to image, 6B, 8 steps | `bf16` | 11.62 GiB | Apache-2.0 |
|
| 59 |
+
| [`z-image-turbo-q4`](z-image-turbo-q4) | Text to image, 6B, 8 steps | `4-bit` | 3.55 GiB | Apache-2.0 |
|
| 60 |
+
| [`flux2-klein-4b`](flux2-klein-4b) | Text to image and multi-reference editing, 4B, 4 steps | `bf16` | 7.39 GiB | Apache-2.0 |
|
| 61 |
+
| [`flux2-klein-4b-q4`](flux2-klein-4b-q4) | Text to image and multi-reference editing, 4B, 4 steps | `4-bit` | 2.20 GiB | Apache-2.0 |
|
| 62 |
+
| [`ace-step-1.5-turbo`](ace-step-1.5-turbo) | Text to music, 48 kHz stereo | `bf16` | 9.38 GiB | MIT |
|
| 63 |
+
|
| 64 |
+
## bf16 or 4-bit
|
| 65 |
+
|
| 66 |
+
**Use the 4-bit folders.** Quantization is MLX affine: 4 bits per weight with a bf16 scale and bias for every group of 64, about 4.5 bits per weight and 28-30% of the bf16 size. Normalizations, modulations and the encoder's embedding table stay in bf16, and matrix multiplication still runs in bf16 - the gain is memory and load time, not integer arithmetic.
|
| 67 |
+
|
| 68 |
+
| | bf16 | 4-bit |
|
| 69 |
+
|-----------------------------------|-----------|----------|
|
| 70 |
+
| Z-Image DiT | 11.46 GiB | 3.40 GiB |
|
| 71 |
+
| klein DiT | 7.22 GiB | 2.03 GiB |
|
| 72 |
+
| shared encoder | 7.51 GiB | 2.58 GiB |
|
| 73 |
+
| klein 512 px, 4 steps, wall clock | 170 s | 73 s |
|
| 74 |
+
|
| 75 |
+
Output quality was indistinguishable: the same prompt and seed gave two clean images that differ the way two seeds differ. The bf16 folders exist as the accuracy reference a port can be checked against.
|
| 76 |
+
|
| 77 |
+
## Bundle layout
|
| 78 |
+
|
| 79 |
+
```
|
| 80 |
+
<folder>/
|
| 81 |
+
├── manifest.json SHA-256, byte sizes and tensor counts for every file
|
| 82 |
+
├── config/ component configs copied from upstream
|
| 83 |
+
├── tokenizer/ BPE plus chat_format.json, where applicable
|
| 84 |
+
└── weights/ safetensors in MLX tensor layout
|
| 85 |
+
```
|
| 86 |
+
|
| 87 |
+
`manifest.json` records the quantization settings, so a loader rebuilds the exact layout without being told. Verify integrity before first use: the manifest carries SHA-256 per file, which catches corruption that size checks miss.
|
| 88 |
+
|
| 89 |
+
`tokenizer/chat_format.json` holds the Qwen chat template already rendered into a prefix and a suffix for each pipeline, so a runtime needs no Jinja at all.
|
| 90 |
+
|
| 91 |
+
<details>
|
| 92 |
+
<summary><b>Tensor layout</b></summary>
|
| 93 |
+
|
| 94 |
+
MLX convolutions expect channels last, PyTorch expects channels first, so convolution weights are permuted during conversion. Everything else keeps its upstream shape.
|
| 95 |
+
|
| 96 |
+
| kind | PyTorch | MLX |
|
| 97 |
+
|-----------------------------|---------------------|---------------------|
|
| 98 |
+
| `Conv2d.weight` | `(out, in, kH, kW)` | `(out, kH, kW, in)` |
|
| 99 |
+
| `Linear`, norms, embeddings | `(out, in)` | unchanged |
|
| 100 |
+
|
| 101 |
+
No `weight_norm` and no 3-D convolutions appear in these models, so no folding was needed. Parameter names match the upstream checkpoints, so weights load without remapping.
|
| 102 |
+
|
| 103 |
+
</details>
|
| 104 |
+
|
| 105 |
+
<details>
|
| 106 |
+
<summary><b>Verification numbers</b></summary>
|
| 107 |
+
|
| 108 |
+
Every module was compared against the upstream reference on fixed inputs in float32. The metric is `rel_max = max|a-b| / max|a|`.
|
| 109 |
+
|
| 110 |
+
| module | rel_max | reference |
|
| 111 |
+
|-------------------------------------------------|-----------|---------------------------------|
|
| 112 |
+
| Qwen3 encoder, hidden_states[-2] | 2.7e-07 | transformers |
|
| 113 |
+
| Qwen3 encoder, layers 9/18/27 with padding mask | 6.0e-07 | transformers |
|
| 114 |
+
| Z-Image DiT | 4.2e-06 | diffusers |
|
| 115 |
+
| Flux2 DiT | 3.7e-07 | diffusers |
|
| 116 |
+
| VAE decoder | 1.3e-05 | diffusers |
|
| 117 |
+
| VAE encoder | 5.3e-06 | diffusers |
|
| 118 |
+
| VAE tiled decode | 5.1e-06 | diffusers tiled_decode |
|
| 119 |
+
| sigma schedule, static shift | 3.2e-08 | FlowMatchEulerDiscreteScheduler |
|
| 120 |
+
| sigma schedule, exponential dynamic shift | 7.7e-08 | FlowMatchEulerDiscreteScheduler |
|
| 121 |
+
| position ids, latent packing, patchify | bit-exact | pipeline helpers |
|
| 122 |
+
|
| 123 |
+
A Swift MLX implementation was then checked against the Python one:
|
| 124 |
+
|
| 125 |
+
| module | rel_max | note |
|
| 126 |
+
|------------------------------|-----------------|------------------------------------------------|
|
| 127 |
+
| tokenization, both pipelines | bit-exact | same ids |
|
| 128 |
+
| Qwen3 encoder, bf16 | 2.2e-04 | Z-Image branch |
|
| 129 |
+
| Qwen3 encoder, bf16 | 4.6e-03 | klein branch, 512 tokens with mask |
|
| 130 |
+
| Qwen3 encoder, 4-bit | 3.0e-04 | proves the quantized layout is rebuilt exactly |
|
| 131 |
+
| Z-Image DiT, float32 | 3.7e-07 | 8 layers |
|
| 132 |
+
| Flux2 DiT, float32 | 1.0e-06 | 2 double + 2 single blocks |
|
| 133 |
+
| VAE decode / tiled / encode | 1e-05 or better | both VAE classes |
|
| 134 |
+
| VAE BatchNorm statistics | 0 | exact |
|
| 135 |
+
|
| 136 |
+
The bf16 figures for a full-depth DiT are larger - 2e-02 for both models - and that is rounding order, not a defect: at float32 the same code agrees to 1e-06, and the deviation grows with depth from a bf16-level 1e-04 per block. Comparing two bf16 implementations below 1e-02 is not meaningful for a 30-block network.
|
| 137 |
+
|
| 138 |
+
</details>
|
| 139 |
+
|
| 140 |
+
<details>
|
| 141 |
+
<summary><b>Performance on an M1 Max</b></summary>
|
| 142 |
+
|
| 143 |
+
Mac Studio, Apple M1 Max, 10-core CPU, 24-core GPU, 32 GB, macOS 26.5. Release build, 4-bit weights, batch 1, warm page cache.
|
| 144 |
+
|
| 145 |
+
| run | per step | VAE decode |
|
| 146 |
+
|--------------------------------------|----------|------------|
|
| 147 |
+
| Z-Image, 512 px, 4 steps | 2.6 s | 1.5 s |
|
| 148 |
+
| klein, 512 px, 4 steps | 2.3 s | 0.9 s |
|
| 149 |
+
| klein, 1024 px, 4 steps | 8.0 s | 0.2 s |
|
| 150 |
+
| klein, 512 px + one 512 px reference | 4.1 s | 0.9 s |
|
| 151 |
+
|
| 152 |
+
Reading the weights dominates a cold run and depends entirely on the storage: on an external volume delivering about 0.2 GiB/s, the 4-bit encoder took 22-32 s and a 4-bit DiT 18-35 s, against 60-107 s for the same weights in bf16. On internal storage expect these to be several times shorter. Compute is unaffected: the numbers above are steady state.
|
| 153 |
+
|
| 154 |
+
Stages run one at a time and release their weights, so peak memory is set by the largest single component rather than their sum. The 4-bit DiT is 2.03 GiB for klein and 3.40 GiB for Z-Image; activations at 1024 px add to that, and no separate peak measurement was made.
|
| 155 |
+
|
| 156 |
+
</details>
|
| 157 |
+
|
| 158 |
+
<details>
|
| 159 |
+
<summary><b>Limitations</b></summary>
|
| 160 |
+
|
| 161 |
+
- Seeds are not compatible with the upstream pipelines, which use `torch.Generator`. The same prompt gives comparable images, never the same file.
|
| 162 |
+
- Both models are distilled: CFG is not applied, and step counts are low by design (8 for Z-Image, 4 for klein).
|
| 163 |
+
- Z-Image's Omni mode is not covered: it needs a SigLIP encoder that is not part of this bundle.
|
| 164 |
+
- Tiled VAE decode is available for high resolutions and gives a result that differs slightly from a single pass, exactly as it does in diffusers.
|
| 165 |
+
- Batches larger than one are not implemented.
|
| 166 |
+
|
| 167 |
+
</details>
|
| 168 |
+
|
| 169 |
+
## Licences
|
| 170 |
+
|
| 171 |
+
Derivatives of three upstream models under two licences. Each folder carries the licence of its upstream model.
|
| 172 |
+
|
| 173 |
+
| folder | upstream | licence |
|
| 174 |
+
|-----------------------------------|-----------------------------------------------------------------------------------------------|-----------------------------------------------------------|
|
| 175 |
+
| `shared/qwen3-4b-text-encoder` | [black-forest-labs/FLUX.2-klein-4B](https://huggingface.co/black-forest-labs/FLUX.2-klein-4B) | [Apache-2.0](https://www.apache.org/licenses/LICENSE-2.0) |
|
| 176 |
+
| `shared/qwen3-4b-text-encoder-q4` | [black-forest-labs/FLUX.2-klein-4B](https://huggingface.co/black-forest-labs/FLUX.2-klein-4B) | [Apache-2.0](https://www.apache.org/licenses/LICENSE-2.0) |
|
| 177 |
+
| `z-image-turbo` | [Tongyi-MAI/Z-Image-Turbo](https://huggingface.co/Tongyi-MAI/Z-Image-Turbo) | [Apache-2.0](https://www.apache.org/licenses/LICENSE-2.0) |
|
| 178 |
+
| `z-image-turbo-q4` | [Tongyi-MAI/Z-Image-Turbo](https://huggingface.co/Tongyi-MAI/Z-Image-Turbo) | [Apache-2.0](https://www.apache.org/licenses/LICENSE-2.0) |
|
| 179 |
+
| `flux2-klein-4b` | [black-forest-labs/FLUX.2-klein-4B](https://huggingface.co/black-forest-labs/FLUX.2-klein-4B) | [Apache-2.0](https://www.apache.org/licenses/LICENSE-2.0) |
|
| 180 |
+
| `flux2-klein-4b-q4` | [black-forest-labs/FLUX.2-klein-4B](https://huggingface.co/black-forest-labs/FLUX.2-klein-4B) | [Apache-2.0](https://www.apache.org/licenses/LICENSE-2.0) |
|
| 181 |
+
| `ace-step-1.5-turbo` | [ACE-Step/Ace-Step1.5](https://huggingface.co/ACE-Step/Ace-Step1.5) | [MIT](https://opensource.org/license/mit) |
|
| 182 |
+
|
| 183 |
+
The shared text encoder is redistributed from inside the two image repositories, both Apache-2.0. Conversion tooling and these cards: MIT.
|
| 184 |
+
|
| 185 |
+
The upstream cards state usage restrictions and responsible-use commitments that redistribution does not repeal. See in particular the out-of-scope use section of the FLUX.2 [klein] card.
|
ace-step-1.5-turbo/README.md
ADDED
|
@@ -0,0 +1,213 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: mit
|
| 3 |
+
base_model: ACE-Step/Ace-Step1.5
|
| 4 |
+
tags:
|
| 5 |
+
- music-generation
|
| 6 |
+
- text-to-music
|
| 7 |
+
- mlx
|
| 8 |
+
- mlx-swift
|
| 9 |
+
- apple-silicon
|
| 10 |
+
- macos
|
| 11 |
+
library_name: mlx
|
| 12 |
+
pipeline_tag: text-to-audio
|
| 13 |
+
---
|
| 14 |
+
|
| 15 |
+
# ACE-Step 1.5 Turbo for native MLX inference on Apple Silicon
|
| 16 |
+
|
| 17 |
+
Converted upstream weights for a native macOS implementation of ACE-Step 1.5.
|
| 18 |
+
The inference path runs entirely through Swift and MLX — no Python, PyTorch,
|
| 19 |
+
Transformers, Diffusers, subprocess or local server.
|
| 20 |
+
|
| 21 |
+
Weights are not retrained and not quantized. Tensor layouts were adapted for
|
| 22 |
+
MLX and VAE `weight_norm` was folded during conversion, so parameters are
|
| 23 |
+
mathematically equivalent to upstream rather than byte-identical.
|
| 24 |
+
|
| 25 |
+
> This folder is part of the combined `kruatech/studio-mlx` repository.
|
| 26 |
+
> The status note below applies to ACE-Step only: the Swift runtime for the
|
| 27 |
+
> image models in this repository is implemented, the one for ACE-Step is not.
|
| 28 |
+
|
| 29 |
+
## Status
|
| 30 |
+
|
| 31 |
+
The Swift runtime is not published yet and will live in a separate repository,
|
| 32 |
+
so this repository currently provides the model bundle only. To try ACE-Step 1.5
|
| 33 |
+
today, use the upstream repository.
|
| 34 |
+
|
| 35 |
+
## Supported tasks
|
| 36 |
+
|
| 37 |
+
All four turbo task types work:
|
| 38 |
+
|
| 39 |
+
| task | what it does |
|
| 40 |
+
|---|---|
|
| 41 |
+
| `text2music` | generate from a text description and lyrics |
|
| 42 |
+
| `repaint` | regenerate a time range, keep the rest bit-exact ¹ |
|
| 43 |
+
| `cover` | follow an existing track's structure via 5 Hz FSQ codes |
|
| 44 |
+
| `cover-nofsq` | same, through the full 25 Hz latent — stays closer to the source |
|
| 45 |
+
|
| 46 |
+
¹ Outside the range the original PCM is spliced back, so it matches bit for bit
|
| 47 |
+
except in the crossfade windows at the boundaries (0.025 s each by default).
|
| 48 |
+
|
| 49 |
+
Also implemented: CFG with APG guidance, DCW wavelet correction (on by default,
|
| 50 |
+
as upstream), tiled VAE encode and decode, retake variations, velocity
|
| 51 |
+
stabilisation, second-order Heun sampler, LM forward pass.
|
| 52 |
+
|
| 53 |
+
Not implemented: generating audio codes from text end to end (the LM's
|
| 54 |
+
217204-entry BPE tokenizer is not ported), constrained metadata decoding,
|
| 55 |
+
batches larger than one. Metadata is supplied explicitly.
|
| 56 |
+
|
| 57 |
+
Not applicable to turbo: `use_adg` and flow-edit morphing need a base model;
|
| 58 |
+
`extract`, `lego` and `complete` are base-only tasks.
|
| 59 |
+
|
| 60 |
+
## Pipeline
|
| 61 |
+
|
| 62 |
+
```
|
| 63 |
+
prompt ──▶ Qwen3-Embedding-0.6B ──┐
|
| 64 |
+
lyrics ──▶ Qwen3-Embedding-0.6B ──┼──▶ CondEncoder ──▶ encoder_hidden_states
|
| 65 |
+
reference latent ─────────────────┘
|
| 66 |
+
|
| 67 |
+
noise [1, T, 64] ──▶ DiT, 24 layers, 8 steps ──▶ latent ──▶ VAE ──▶ 48 kHz stereo
|
| 68 |
+
```
|
| 69 |
+
|
| 70 |
+
`T` is latent frames at 25 Hz: 30 s of audio is 750 frames. The VAE upsamples by
|
| 71 |
+
1920 per frame.
|
| 72 |
+
|
| 73 |
+
## Files
|
| 74 |
+
|
| 75 |
+
Required for `text2music`, `repaint`, `cover` — 6.34 GB:
|
| 76 |
+
|
| 77 |
+
| path | size | contents |
|
| 78 |
+
|---|---|---|
|
| 79 |
+
| `weights/dit.safetensors` | 4.79 GB | DiT, FSQ quantizer, detokenizer, attention pooler — 677 tensors, BF16 |
|
| 80 |
+
| `weights/text_encoder.safetensors` | 1.19 GB | Qwen3-Embedding-0.6B, 310 tensors, BF16 |
|
| 81 |
+
| `weights/vae.safetensors` | 337 MB | AutoencoderOobleck, 291 tensors: 146 encoder + 145 decoder |
|
| 82 |
+
| `weights/silence.bin` | 3.8 MB | silence latent, F32 `[15000, 64]` in NLC |
|
| 83 |
+
| `config/dit.json`, `config/text.json` | 4 KB | model configs |
|
| 84 |
+
| `tokenizer/` | 14 MB | Qwen3-Embedding BPE |
|
| 85 |
+
| `manifest.json` | 3 KB | SHA-256, sizes and tensor counts |
|
| 86 |
+
|
| 87 |
+
Optional, 3.71 GB: `optional/lm.safetensors` and `optional/lm_tokenizer/` hold
|
| 88 |
+
the 1.7B LM. `text2music` never loads them, and cover works without them —
|
| 89 |
+
audio is encoded and tokenised directly. The LM is only needed to produce codes
|
| 90 |
+
from a text description instead of a reference track.
|
| 91 |
+
|
| 92 |
+
Verify integrity before first use: `manifest.json` carries SHA-256 for every
|
| 93 |
+
file, which catches corruption that size and header checks miss.
|
| 94 |
+
|
| 95 |
+
## Tensor layout
|
| 96 |
+
|
| 97 |
+
MLX convolutions expect channels last, PyTorch expects channels first, so
|
| 98 |
+
convolution weights are permuted during conversion:
|
| 99 |
+
|
| 100 |
+
| kind | PyTorch | MLX |
|
| 101 |
+
|---|---|---|
|
| 102 |
+
| `Conv1d` | `(out, in, k)` | `(out, k, in)` |
|
| 103 |
+
| `ConvTranspose1d` | `(in, out, k)` | `(out, k, in)` |
|
| 104 |
+
| Snake `alpha`, `beta` | `(1, C, 1)` | `(1, 1, C)` |
|
| 105 |
+
| `Linear` | `(out, in)` | unchanged |
|
| 106 |
+
|
| 107 |
+
`weight_norm` pairs are folded into single tensors: `w = v · (g / ‖v‖)`,
|
| 108 |
+
computed in F32 with an F64 accumulator and rounded once. This removes 74
|
| 109 |
+
tensors from the VAE.
|
| 110 |
+
|
| 111 |
+
## Memory
|
| 112 |
+
|
| 113 |
+
Peak MLX active memory per stage, F32, 60 s of audio with CFG 3.0:
|
| 114 |
+
|
| 115 |
+
| stage | peak active |
|
| 116 |
+
|---|---|
|
| 117 |
+
| text encoder | 2.1 GB |
|
| 118 |
+
| CondEncoder | 2.8 GB |
|
| 119 |
+
| DiT | 7.2 GB |
|
| 120 |
+
| VAE decoder, tiled | 4.1 GB |
|
| 121 |
+
|
| 122 |
+
Stages run sequentially and release their weights, so the figures do not add
|
| 123 |
+
up. DiT sets the ceiling; with the buffer cache the working set is about 9.2 GB.
|
| 124 |
+
|
| 125 |
+
**16 GB is the tested minimum.** An M1 Pro with 16 GB ran all four tasks
|
| 126 |
+
without swap growth. 8 GB is not expected to fit.
|
| 127 |
+
|
| 128 |
+
The VAE must be tiled — a single-pass 30 s decode drove the MLX buffer cache to
|
| 129 |
+
25 GB. With tiling the cache stayed under 400 MB across all tested lengths, 10 s
|
| 130 |
+
to 180 s. Overlaps of 16, 32 and 64 latent frames each produced output identical
|
| 131 |
+
to a single pass.
|
| 132 |
+
|
| 133 |
+
DiT must run in F32. BF16 halves its memory but fails accuracy at
|
| 134 |
+
`rel_max 1.45e-01`.
|
| 135 |
+
|
| 136 |
+
## Performance
|
| 137 |
+
|
| 138 |
+
Warm cache, release build, 8 diffusion steps, batch 1, F32. RTF is generation
|
| 139 |
+
time over output duration.
|
| 140 |
+
|
| 141 |
+
**M1 Pro, 16 GB** — 10-core CPU, 16-core GPU, macOS 26.5, Swift 6.3.2:
|
| 142 |
+
|
| 143 |
+
```
|
| 144 |
+
mode wall dit s/step vae peak MB RTF
|
| 145 |
+
30 s, guidance 1.0 8.6s 3.65s 0.46s 3.78s 6602 0.287
|
| 146 |
+
60 s, guidance 1.0 15.6s 6.81s 0.85s 7.46s 6806 0.260
|
| 147 |
+
30 s, guidance 3.0 11.6s 6.91s 0.86s 3.51s 6761 0.387
|
| 148 |
+
60 s, guidance 3.0 22.1s 13.38s 1.67s 7.40s 7180 0.368
|
| 149 |
+
30 s, Heun sampler 16.7s 11.93s 1.49s 3.55s 6761 0.557
|
| 150 |
+
```
|
| 151 |
+
|
| 152 |
+
**M1 Max, 32 GB** — 10-core CPU, 24-core GPU, macOS 26.5, Swift 6.3.3, CFG 3.0:
|
| 153 |
+
|
| 154 |
+
```
|
| 155 |
+
30 s: text 0.15s cond 0.19s dit 4.47s vae 2.28s · 7.69s wall
|
| 156 |
+
60 s: dit 8.77s (1.10 s/step) · RTF 0.146
|
| 157 |
+
```
|
| 158 |
+
|
| 159 |
+
The M1 Max is about 1.5× faster per diffusion step, tracking GPU core count
|
| 160 |
+
rather than memory. Peak memory is the same on both.
|
| 161 |
+
|
| 162 |
+
CFG doubles the batch and roughly doubles diffusion time. Heun doubles model
|
| 163 |
+
evaluations per step. A first run after process start is several times slower
|
| 164 |
+
while 6.3 GB of weights come off disk — measure the second run.
|
| 165 |
+
|
| 166 |
+
## Verification
|
| 167 |
+
|
| 168 |
+
Every module was compared against the upstream reference on fixed inputs in
|
| 169 |
+
F32. The metric is `rel_max = max_abs / std(reference)`, threshold `1e-03`:
|
| 170 |
+
|
| 171 |
+
| module | rel_max | reference |
|
| 172 |
+
|---|---|---|
|
| 173 |
+
| tokenization | bit-exact | `transformers` |
|
| 174 |
+
| text encoder, 28 layers | 3.39e-05 | `transformers` |
|
| 175 |
+
| CondEncoder | 5.98e-04 | `transformers` |
|
| 176 |
+
| DiT, 24 layers | 4.35e-05 | `diffusers` |
|
| 177 |
+
| VAE encoder | 5.38e-04 | `diffusers` |
|
| 178 |
+
| VAE decoder | cosine 1.000000 | `diffusers` |
|
| 179 |
+
| FSQ | 3.40e-07 | `vector_quantize_pytorch` ¹ |
|
| 180 |
+
| detokenizer | 1.15e-05 | `transformers` |
|
| 181 |
+
| audio tokenizer | 1.31e-05 | `transformers` ² |
|
| 182 |
+
| LM, 28 layers | 2.48e-04 | `transformers` ³ |
|
| 183 |
+
| CFG / APG | 4.95e-05 | independent MLX implementation ⁴ |
|
| 184 |
+
|
| 185 |
+
¹ Codebook matched exactly on all 64000 entries.
|
| 186 |
+
² All 50 indices matched exactly.
|
| 187 |
+
³ argmax matched at all 50 positions.
|
| 188 |
+
⁴ APG has no PyTorch reference upstream; validated on identical noise.
|
| 189 |
+
|
| 190 |
+
Reference versions: torch 2.13.0, transformers 5.14.1, diffusers 0.39.0,
|
| 191 |
+
mlx 0.32.0.
|
| 192 |
+
|
| 193 |
+
Verify the VAE against an F32 weight file, not the BF16 one shipped here.
|
| 194 |
+
`weight_norm` folding produces new values, and storing them in BF16 quantises
|
| 195 |
+
them: `rel_max 5.11e-02` against BF16 versus `4.78e-05` against F32. For
|
| 196 |
+
inference BF16 is fine — median signal-to-error through encode and decode is
|
| 197 |
+
49.7 dB.
|
| 198 |
+
|
| 199 |
+
## Limitations
|
| 200 |
+
|
| 201 |
+
Seeds are not compatible with the upstream pipeline, which uses
|
| 202 |
+
`torch.Generator`. The same prompt gives comparable music, never the same file.
|
| 203 |
+
|
| 204 |
+
Minimum length is 5.12 s. Step count is fixed at 8; this is the turbo model.
|
| 205 |
+
|
| 206 |
+
## Licenses
|
| 207 |
+
|
| 208 |
+
MIT, same as upstream.
|
| 209 |
+
|
| 210 |
+
- DiT, VAE, LM: converted from
|
| 211 |
+
[ACE-Step 1.5](https://github.com/ace-step/ACE-Step-1.5), MIT.
|
| 212 |
+
- Text encoder and tokenizer: Qwen3-Embedding-0.6B, redistributed unchanged.
|
| 213 |
+
- Conversion tooling and this card: MIT.
|
flux2-klein-4b-q4/README.md
ADDED
|
@@ -0,0 +1,130 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: apache-2.0
|
| 3 |
+
library_name: mlx
|
| 4 |
+
tags:
|
| 5 |
+
- mlx
|
| 6 |
+
- mlx-swift
|
| 7 |
+
- apple-silicon
|
| 8 |
+
base_model:
|
| 9 |
+
- black-forest-labs/FLUX.2-klein-4B
|
| 10 |
+
---
|
| 11 |
+
|
| 12 |
+
<div align="center">
|
| 13 |
+
|
| 14 |
+
# FLUX.2 [klein] 4B for MLX, 4-bit
|
| 15 |
+
|
| 16 |
+
**Text to image and multi-reference editing, 4B, 4 steps**
|
| 17 |
+
|
| 18 |
+
  [](https://www.apache.org/licenses/LICENSE-2.0) [](https://huggingface.co/kruatech/studio-mlx)
|
| 19 |
+
|
| 20 |
+
</div>
|
| 21 |
+
|
| 22 |
+
4-bit version of FLUX.2 [klein] 4B. Everything in the bf16 card applies.
|
| 23 |
+
|
| 24 |
+
Converted from [black-forest-labs/FLUX.2-klein-4B](https://huggingface.co/black-forest-labs/FLUX.2-klein-4B) for native MLX inference on Apple silicon.
|
| 25 |
+
|
| 26 |
+
## Download
|
| 27 |
+
|
| 28 |
+
Also download `shared/qwen3-4b-text-encoder-q4`.
|
| 29 |
+
|
| 30 |
+
```bash
|
| 31 |
+
pip install "huggingface_hub[hf_xet]"
|
| 32 |
+
|
| 33 |
+
hf download kruatech/studio-mlx --local-dir bundles \
|
| 34 |
+
--include "shared/qwen3-4b-text-encoder-q4/*" "flux2-klein-4b-q4/*"
|
| 35 |
+
```
|
| 36 |
+
|
| 37 |
+
## Files
|
| 38 |
+
|
| 39 |
+
| component | class | precision | tensors | size |
|
| 40 |
+
|-------------|-------------------------|-----------------|---------|----------|
|
| 41 |
+
| transformer | Flux2Transformer2DModel | 4-bit, group 64 | 381 | 2.03 GiB |
|
| 42 |
+
| vae | AutoencoderKLFlux2 | keep | 251 | 160 MiB |
|
| 43 |
+
|
| 44 |
+
Folder total: **2.20 GiB**. `manifest.json` carries SHA-256, byte sizes and tensor counts for every file.
|
| 45 |
+
|
| 46 |
+
## How this model works in MLX
|
| 47 |
+
|
| 48 |
+
- 106 linear layers quantized. The VAE is not quantized.
|
| 49 |
+
|
| 50 |
+
<details>
|
| 51 |
+
<summary><b>Tensor layout</b></summary>
|
| 52 |
+
|
| 53 |
+
MLX convolutions expect channels last, PyTorch expects channels first, so convolution weights are permuted during conversion. Everything else keeps its upstream shape.
|
| 54 |
+
|
| 55 |
+
| kind | PyTorch | MLX |
|
| 56 |
+
|-----------------------------|---------------------|---------------------|
|
| 57 |
+
| `Conv2d.weight` | `(out, in, kH, kW)` | `(out, kH, kW, in)` |
|
| 58 |
+
| `Linear`, norms, embeddings | `(out, in)` | unchanged |
|
| 59 |
+
|
| 60 |
+
No `weight_norm` and no 3-D convolutions appear in these models, so no folding was needed. Parameter names match the upstream checkpoints, so weights load without remapping.
|
| 61 |
+
|
| 62 |
+
</details>
|
| 63 |
+
|
| 64 |
+
<details>
|
| 65 |
+
<summary><b>Verification numbers</b></summary>
|
| 66 |
+
|
| 67 |
+
Every module was compared against the upstream reference on fixed inputs in float32. The metric is `rel_max = max|a-b| / max|a|`.
|
| 68 |
+
|
| 69 |
+
| module | rel_max | reference |
|
| 70 |
+
|-------------------------------------------------|-----------|---------------------------------|
|
| 71 |
+
| Qwen3 encoder, hidden_states[-2] | 2.7e-07 | transformers |
|
| 72 |
+
| Qwen3 encoder, layers 9/18/27 with padding mask | 6.0e-07 | transformers |
|
| 73 |
+
| Z-Image DiT | 4.2e-06 | diffusers |
|
| 74 |
+
| Flux2 DiT | 3.7e-07 | diffusers |
|
| 75 |
+
| VAE decoder | 1.3e-05 | diffusers |
|
| 76 |
+
| VAE encoder | 5.3e-06 | diffusers |
|
| 77 |
+
| VAE tiled decode | 5.1e-06 | diffusers tiled_decode |
|
| 78 |
+
| sigma schedule, static shift | 3.2e-08 | FlowMatchEulerDiscreteScheduler |
|
| 79 |
+
| sigma schedule, exponential dynamic shift | 7.7e-08 | FlowMatchEulerDiscreteScheduler |
|
| 80 |
+
| position ids, latent packing, patchify | bit-exact | pipeline helpers |
|
| 81 |
+
|
| 82 |
+
A Swift MLX implementation was then checked against the Python one:
|
| 83 |
+
|
| 84 |
+
| module | rel_max | note |
|
| 85 |
+
|------------------------------|-----------------|------------------------------------------------|
|
| 86 |
+
| tokenization, both pipelines | bit-exact | same ids |
|
| 87 |
+
| Qwen3 encoder, bf16 | 2.2e-04 | Z-Image branch |
|
| 88 |
+
| Qwen3 encoder, bf16 | 4.6e-03 | klein branch, 512 tokens with mask |
|
| 89 |
+
| Qwen3 encoder, 4-bit | 3.0e-04 | proves the quantized layout is rebuilt exactly |
|
| 90 |
+
| Z-Image DiT, float32 | 3.7e-07 | 8 layers |
|
| 91 |
+
| Flux2 DiT, float32 | 1.0e-06 | 2 double + 2 single blocks |
|
| 92 |
+
| VAE decode / tiled / encode | 1e-05 or better | both VAE classes |
|
| 93 |
+
| VAE BatchNorm statistics | 0 | exact |
|
| 94 |
+
|
| 95 |
+
The bf16 figures for a full-depth DiT are larger - 2e-02 for both models - and that is rounding order, not a defect: at float32 the same code agrees to 1e-06, and the deviation grows with depth from a bf16-level 1e-04 per block. Comparing two bf16 implementations below 1e-02 is not meaningful for a 30-block network.
|
| 96 |
+
|
| 97 |
+
</details>
|
| 98 |
+
|
| 99 |
+
<details>
|
| 100 |
+
<summary><b>Performance on an M1 Max</b></summary>
|
| 101 |
+
|
| 102 |
+
Mac Studio, Apple M1 Max, 10-core CPU, 24-core GPU, 32 GB, macOS 26.5. Release build, 4-bit weights, batch 1, warm page cache.
|
| 103 |
+
|
| 104 |
+
| run | per step | VAE decode |
|
| 105 |
+
|--------------------------------------|----------|------------|
|
| 106 |
+
| Z-Image, 512 px, 4 steps | 2.6 s | 1.5 s |
|
| 107 |
+
| klein, 512 px, 4 steps | 2.3 s | 0.9 s |
|
| 108 |
+
| klein, 1024 px, 4 steps | 8.0 s | 0.2 s |
|
| 109 |
+
| klein, 512 px + one 512 px reference | 4.1 s | 0.9 s |
|
| 110 |
+
|
| 111 |
+
Reading the weights dominates a cold run and depends entirely on the storage: on an external volume delivering about 0.2 GiB/s, the 4-bit encoder took 22-32 s and a 4-bit DiT 18-35 s, against 60-107 s for the same weights in bf16. On internal storage expect these to be several times shorter. Compute is unaffected: the numbers above are steady state.
|
| 112 |
+
|
| 113 |
+
Stages run one at a time and release their weights, so peak memory is set by the largest single component rather than their sum. The 4-bit DiT is 2.03 GiB for klein and 3.40 GiB for Z-Image; activations at 1024 px add to that, and no separate peak measurement was made.
|
| 114 |
+
|
| 115 |
+
</details>
|
| 116 |
+
|
| 117 |
+
<details>
|
| 118 |
+
<summary><b>Limitations</b></summary>
|
| 119 |
+
|
| 120 |
+
- Seeds are not compatible with the upstream pipelines, which use `torch.Generator`. The same prompt gives comparable images, never the same file.
|
| 121 |
+
- Both models are distilled: CFG is not applied, and step counts are low by design (8 for Z-Image, 4 for klein).
|
| 122 |
+
- Z-Image's Omni mode is not covered: it needs a SigLIP encoder that is not part of this bundle.
|
| 123 |
+
- Tiled VAE decode is available for high resolutions and gives a result that differs slightly from a single pass, exactly as it does in diffusers.
|
| 124 |
+
- Batches larger than one are not implemented.
|
| 125 |
+
|
| 126 |
+
</details>
|
| 127 |
+
|
| 128 |
+
## Licence
|
| 129 |
+
|
| 130 |
+
[Apache-2.0](https://www.apache.org/licenses/LICENSE-2.0), the licence of the upstream model. Conversion tooling and this card: MIT. The upstream card states usage restrictions that redistribution does not repeal.
|
flux2-klein-4b/README.md
ADDED
|
@@ -0,0 +1,134 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: apache-2.0
|
| 3 |
+
library_name: mlx
|
| 4 |
+
tags:
|
| 5 |
+
- mlx
|
| 6 |
+
- mlx-swift
|
| 7 |
+
- apple-silicon
|
| 8 |
+
base_model:
|
| 9 |
+
- black-forest-labs/FLUX.2-klein-4B
|
| 10 |
+
---
|
| 11 |
+
|
| 12 |
+
<div align="center">
|
| 13 |
+
|
| 14 |
+
# FLUX.2 [klein] 4B for MLX, bf16
|
| 15 |
+
|
| 16 |
+
**Text to image and multi-reference editing, 4B, 4 steps**
|
| 17 |
+
|
| 18 |
+
  [](https://www.apache.org/licenses/LICENSE-2.0) [](https://huggingface.co/kruatech/studio-mlx)
|
| 19 |
+
|
| 20 |
+
</div>
|
| 21 |
+
|
| 22 |
+
4B rectified flow transformer, text to image and multi-reference editing, 4 steps, distilled so guidance is not applied.
|
| 23 |
+
|
| 24 |
+
Converted from [black-forest-labs/FLUX.2-klein-4B](https://huggingface.co/black-forest-labs/FLUX.2-klein-4B) for native MLX inference on Apple silicon.
|
| 25 |
+
|
| 26 |
+
## Download
|
| 27 |
+
|
| 28 |
+
Also download `shared/qwen3-4b-text-encoder`.
|
| 29 |
+
|
| 30 |
+
```bash
|
| 31 |
+
pip install "huggingface_hub[hf_xet]"
|
| 32 |
+
|
| 33 |
+
hf download kruatech/studio-mlx --local-dir bundles \
|
| 34 |
+
--include "shared/qwen3-4b-text-encoder/*" "flux2-klein-4b/*"
|
| 35 |
+
```
|
| 36 |
+
|
| 37 |
+
## Files
|
| 38 |
+
|
| 39 |
+
| component | class | precision | tensors | size |
|
| 40 |
+
|-------------|-------------------------|-----------|---------|----------|
|
| 41 |
+
| transformer | Flux2Transformer2DModel | keep | 169 | 7.22 GiB |
|
| 42 |
+
| vae | AutoencoderKLFlux2 | keep | 251 | 160 MiB |
|
| 43 |
+
|
| 44 |
+
Folder total: **7.39 GiB**. `manifest.json` carries SHA-256, byte sizes and tensor counts for every file.
|
| 45 |
+
|
| 46 |
+
## How this model works in MLX
|
| 47 |
+
|
| 48 |
+
- 5 double-stream blocks with separate text and image modulation, then 20 single-stream blocks where QKV and the MLP share one fused projection.
|
| 49 |
+
- Rotary embedding over four axes, 32 each, theta 2000.
|
| 50 |
+
- No `scaling_factor`: latents are patchified 2x2 into 128 channels and normalized with the VAE's BatchNorm running statistics, which are in `weights/vae.safetensors`.
|
| 51 |
+
- Sigmas use the exponential dynamic shift with mu from the upstream empirical formula.
|
| 52 |
+
- Reference-image editing needs no KV cache: reference tokens and their ids are appended to the sequence and the prediction is sliced back. Each 1024 px reference adds 4096 tokens, which roughly doubles the cost of a step.
|
| 53 |
+
|
| 54 |
+
<details>
|
| 55 |
+
<summary><b>Tensor layout</b></summary>
|
| 56 |
+
|
| 57 |
+
MLX convolutions expect channels last, PyTorch expects channels first, so convolution weights are permuted during conversion. Everything else keeps its upstream shape.
|
| 58 |
+
|
| 59 |
+
| kind | PyTorch | MLX |
|
| 60 |
+
|-----------------------------|---------------------|---------------------|
|
| 61 |
+
| `Conv2d.weight` | `(out, in, kH, kW)` | `(out, kH, kW, in)` |
|
| 62 |
+
| `Linear`, norms, embeddings | `(out, in)` | unchanged |
|
| 63 |
+
|
| 64 |
+
No `weight_norm` and no 3-D convolutions appear in these models, so no folding was needed. Parameter names match the upstream checkpoints, so weights load without remapping.
|
| 65 |
+
|
| 66 |
+
</details>
|
| 67 |
+
|
| 68 |
+
<details>
|
| 69 |
+
<summary><b>Verification numbers</b></summary>
|
| 70 |
+
|
| 71 |
+
Every module was compared against the upstream reference on fixed inputs in float32. The metric is `rel_max = max|a-b| / max|a|`.
|
| 72 |
+
|
| 73 |
+
| module | rel_max | reference |
|
| 74 |
+
|-------------------------------------------------|-----------|---------------------------------|
|
| 75 |
+
| Qwen3 encoder, hidden_states[-2] | 2.7e-07 | transformers |
|
| 76 |
+
| Qwen3 encoder, layers 9/18/27 with padding mask | 6.0e-07 | transformers |
|
| 77 |
+
| Z-Image DiT | 4.2e-06 | diffusers |
|
| 78 |
+
| Flux2 DiT | 3.7e-07 | diffusers |
|
| 79 |
+
| VAE decoder | 1.3e-05 | diffusers |
|
| 80 |
+
| VAE encoder | 5.3e-06 | diffusers |
|
| 81 |
+
| VAE tiled decode | 5.1e-06 | diffusers tiled_decode |
|
| 82 |
+
| sigma schedule, static shift | 3.2e-08 | FlowMatchEulerDiscreteScheduler |
|
| 83 |
+
| sigma schedule, exponential dynamic shift | 7.7e-08 | FlowMatchEulerDiscreteScheduler |
|
| 84 |
+
| position ids, latent packing, patchify | bit-exact | pipeline helpers |
|
| 85 |
+
|
| 86 |
+
A Swift MLX implementation was then checked against the Python one:
|
| 87 |
+
|
| 88 |
+
| module | rel_max | note |
|
| 89 |
+
|------------------------------|-----------------|------------------------------------------------|
|
| 90 |
+
| tokenization, both pipelines | bit-exact | same ids |
|
| 91 |
+
| Qwen3 encoder, bf16 | 2.2e-04 | Z-Image branch |
|
| 92 |
+
| Qwen3 encoder, bf16 | 4.6e-03 | klein branch, 512 tokens with mask |
|
| 93 |
+
| Qwen3 encoder, 4-bit | 3.0e-04 | proves the quantized layout is rebuilt exactly |
|
| 94 |
+
| Z-Image DiT, float32 | 3.7e-07 | 8 layers |
|
| 95 |
+
| Flux2 DiT, float32 | 1.0e-06 | 2 double + 2 single blocks |
|
| 96 |
+
| VAE decode / tiled / encode | 1e-05 or better | both VAE classes |
|
| 97 |
+
| VAE BatchNorm statistics | 0 | exact |
|
| 98 |
+
|
| 99 |
+
The bf16 figures for a full-depth DiT are larger - 2e-02 for both models - and that is rounding order, not a defect: at float32 the same code agrees to 1e-06, and the deviation grows with depth from a bf16-level 1e-04 per block. Comparing two bf16 implementations below 1e-02 is not meaningful for a 30-block network.
|
| 100 |
+
|
| 101 |
+
</details>
|
| 102 |
+
|
| 103 |
+
<details>
|
| 104 |
+
<summary><b>Performance on an M1 Max</b></summary>
|
| 105 |
+
|
| 106 |
+
Mac Studio, Apple M1 Max, 10-core CPU, 24-core GPU, 32 GB, macOS 26.5. Release build, 4-bit weights, batch 1, warm page cache.
|
| 107 |
+
|
| 108 |
+
| run | per step | VAE decode |
|
| 109 |
+
|--------------------------------------|----------|------------|
|
| 110 |
+
| Z-Image, 512 px, 4 steps | 2.6 s | 1.5 s |
|
| 111 |
+
| klein, 512 px, 4 steps | 2.3 s | 0.9 s |
|
| 112 |
+
| klein, 1024 px, 4 steps | 8.0 s | 0.2 s |
|
| 113 |
+
| klein, 512 px + one 512 px reference | 4.1 s | 0.9 s |
|
| 114 |
+
|
| 115 |
+
Reading the weights dominates a cold run and depends entirely on the storage: on an external volume delivering about 0.2 GiB/s, the 4-bit encoder took 22-32 s and a 4-bit DiT 18-35 s, against 60-107 s for the same weights in bf16. On internal storage expect these to be several times shorter. Compute is unaffected: the numbers above are steady state.
|
| 116 |
+
|
| 117 |
+
Stages run one at a time and release their weights, so peak memory is set by the largest single component rather than their sum. The 4-bit DiT is 2.03 GiB for klein and 3.40 GiB for Z-Image; activations at 1024 px add to that, and no separate peak measurement was made.
|
| 118 |
+
|
| 119 |
+
</details>
|
| 120 |
+
|
| 121 |
+
<details>
|
| 122 |
+
<summary><b>Limitations</b></summary>
|
| 123 |
+
|
| 124 |
+
- Seeds are not compatible with the upstream pipelines, which use `torch.Generator`. The same prompt gives comparable images, never the same file.
|
| 125 |
+
- Both models are distilled: CFG is not applied, and step counts are low by design (8 for Z-Image, 4 for klein).
|
| 126 |
+
- Z-Image's Omni mode is not covered: it needs a SigLIP encoder that is not part of this bundle.
|
| 127 |
+
- Tiled VAE decode is available for high resolutions and gives a result that differs slightly from a single pass, exactly as it does in diffusers.
|
| 128 |
+
- Batches larger than one are not implemented.
|
| 129 |
+
|
| 130 |
+
</details>
|
| 131 |
+
|
| 132 |
+
## Licence
|
| 133 |
+
|
| 134 |
+
[Apache-2.0](https://www.apache.org/licenses/LICENSE-2.0), the licence of the upstream model. Conversion tooling and this card: MIT. The upstream card states usage restrictions that redistribution does not repeal.
|
shared/qwen3-4b-text-encoder-q4/README.md
ADDED
|
@@ -0,0 +1,130 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: apache-2.0
|
| 3 |
+
library_name: mlx
|
| 4 |
+
tags:
|
| 5 |
+
- mlx
|
| 6 |
+
- mlx-swift
|
| 7 |
+
- apple-silicon
|
| 8 |
+
base_model:
|
| 9 |
+
- black-forest-labs/FLUX.2-klein-4B
|
| 10 |
+
---
|
| 11 |
+
|
| 12 |
+
<div align="center">
|
| 13 |
+
|
| 14 |
+
# Qwen3-4B text encoder for MLX, 4-bit
|
| 15 |
+
|
| 16 |
+
**Shared text encoder for both image models**
|
| 17 |
+
|
| 18 |
+
  [](https://www.apache.org/licenses/LICENSE-2.0) [](https://huggingface.co/kruatech/studio-mlx)
|
| 19 |
+
|
| 20 |
+
</div>
|
| 21 |
+
|
| 22 |
+
4-bit version of the shared text encoder. Everything in the bf16 card applies.
|
| 23 |
+
|
| 24 |
+
Converted from [black-forest-labs/FLUX.2-klein-4B](https://huggingface.co/black-forest-labs/FLUX.2-klein-4B) for native MLX inference on Apple silicon.
|
| 25 |
+
|
| 26 |
+
## Download
|
| 27 |
+
|
| 28 |
+
Pair it with `z-image-turbo-q4` or `flux2-klein-4b-q4`.
|
| 29 |
+
|
| 30 |
+
```bash
|
| 31 |
+
pip install "huggingface_hub[hf_xet]"
|
| 32 |
+
|
| 33 |
+
hf download kruatech/studio-mlx --local-dir bundles \
|
| 34 |
+
--include "shared/qwen3-4b-text-encoder-q4/*"
|
| 35 |
+
```
|
| 36 |
+
|
| 37 |
+
## Files
|
| 38 |
+
|
| 39 |
+
| component | class | precision | tensors | size |
|
| 40 |
+
|--------------|------------------|-----------------|---------|----------|
|
| 41 |
+
| text_encoder | Qwen3ForCausalLM | 4-bit, group 64 | 877 | 2.58 GiB |
|
| 42 |
+
|
| 43 |
+
Folder total: **2.59 GiB**. `manifest.json` carries SHA-256, byte sizes and tensor counts for every file.
|
| 44 |
+
|
| 45 |
+
## How this model works in MLX
|
| 46 |
+
|
| 47 |
+
- The embedding table is left in bf16 on purpose: quantizing it moved the encoder output by 3e-01 in testing, which is not worth 0.5 GiB.
|
| 48 |
+
- 245 linear layers are quantized, 4 bits with a bf16 scale and bias per group of 64.
|
| 49 |
+
|
| 50 |
+
<details>
|
| 51 |
+
<summary><b>Tensor layout</b></summary>
|
| 52 |
+
|
| 53 |
+
MLX convolutions expect channels last, PyTorch expects channels first, so convolution weights are permuted during conversion. Everything else keeps its upstream shape.
|
| 54 |
+
|
| 55 |
+
| kind | PyTorch | MLX |
|
| 56 |
+
|-----------------------------|---------------------|---------------------|
|
| 57 |
+
| `Conv2d.weight` | `(out, in, kH, kW)` | `(out, kH, kW, in)` |
|
| 58 |
+
| `Linear`, norms, embeddings | `(out, in)` | unchanged |
|
| 59 |
+
|
| 60 |
+
No `weight_norm` and no 3-D convolutions appear in these models, so no folding was needed. Parameter names match the upstream checkpoints, so weights load without remapping.
|
| 61 |
+
|
| 62 |
+
</details>
|
| 63 |
+
|
| 64 |
+
<details>
|
| 65 |
+
<summary><b>Verification numbers</b></summary>
|
| 66 |
+
|
| 67 |
+
Every module was compared against the upstream reference on fixed inputs in float32. The metric is `rel_max = max|a-b| / max|a|`.
|
| 68 |
+
|
| 69 |
+
| module | rel_max | reference |
|
| 70 |
+
|-------------------------------------------------|-----------|---------------------------------|
|
| 71 |
+
| Qwen3 encoder, hidden_states[-2] | 2.7e-07 | transformers |
|
| 72 |
+
| Qwen3 encoder, layers 9/18/27 with padding mask | 6.0e-07 | transformers |
|
| 73 |
+
| Z-Image DiT | 4.2e-06 | diffusers |
|
| 74 |
+
| Flux2 DiT | 3.7e-07 | diffusers |
|
| 75 |
+
| VAE decoder | 1.3e-05 | diffusers |
|
| 76 |
+
| VAE encoder | 5.3e-06 | diffusers |
|
| 77 |
+
| VAE tiled decode | 5.1e-06 | diffusers tiled_decode |
|
| 78 |
+
| sigma schedule, static shift | 3.2e-08 | FlowMatchEulerDiscreteScheduler |
|
| 79 |
+
| sigma schedule, exponential dynamic shift | 7.7e-08 | FlowMatchEulerDiscreteScheduler |
|
| 80 |
+
| position ids, latent packing, patchify | bit-exact | pipeline helpers |
|
| 81 |
+
|
| 82 |
+
A Swift MLX implementation was then checked against the Python one:
|
| 83 |
+
|
| 84 |
+
| module | rel_max | note |
|
| 85 |
+
|------------------------------|-----------------|------------------------------------------------|
|
| 86 |
+
| tokenization, both pipelines | bit-exact | same ids |
|
| 87 |
+
| Qwen3 encoder, bf16 | 2.2e-04 | Z-Image branch |
|
| 88 |
+
| Qwen3 encoder, bf16 | 4.6e-03 | klein branch, 512 tokens with mask |
|
| 89 |
+
| Qwen3 encoder, 4-bit | 3.0e-04 | proves the quantized layout is rebuilt exactly |
|
| 90 |
+
| Z-Image DiT, float32 | 3.7e-07 | 8 layers |
|
| 91 |
+
| Flux2 DiT, float32 | 1.0e-06 | 2 double + 2 single blocks |
|
| 92 |
+
| VAE decode / tiled / encode | 1e-05 or better | both VAE classes |
|
| 93 |
+
| VAE BatchNorm statistics | 0 | exact |
|
| 94 |
+
|
| 95 |
+
The bf16 figures for a full-depth DiT are larger - 2e-02 for both models - and that is rounding order, not a defect: at float32 the same code agrees to 1e-06, and the deviation grows with depth from a bf16-level 1e-04 per block. Comparing two bf16 implementations below 1e-02 is not meaningful for a 30-block network.
|
| 96 |
+
|
| 97 |
+
</details>
|
| 98 |
+
|
| 99 |
+
<details>
|
| 100 |
+
<summary><b>Performance on an M1 Max</b></summary>
|
| 101 |
+
|
| 102 |
+
Mac Studio, Apple M1 Max, 10-core CPU, 24-core GPU, 32 GB, macOS 26.5. Release build, 4-bit weights, batch 1, warm page cache.
|
| 103 |
+
|
| 104 |
+
| run | per step | VAE decode |
|
| 105 |
+
|--------------------------------------|----------|------------|
|
| 106 |
+
| Z-Image, 512 px, 4 steps | 2.6 s | 1.5 s |
|
| 107 |
+
| klein, 512 px, 4 steps | 2.3 s | 0.9 s |
|
| 108 |
+
| klein, 1024 px, 4 steps | 8.0 s | 0.2 s |
|
| 109 |
+
| klein, 512 px + one 512 px reference | 4.1 s | 0.9 s |
|
| 110 |
+
|
| 111 |
+
Reading the weights dominates a cold run and depends entirely on the storage: on an external volume delivering about 0.2 GiB/s, the 4-bit encoder took 22-32 s and a 4-bit DiT 18-35 s, against 60-107 s for the same weights in bf16. On internal storage expect these to be several times shorter. Compute is unaffected: the numbers above are steady state.
|
| 112 |
+
|
| 113 |
+
Stages run one at a time and release their weights, so peak memory is set by the largest single component rather than their sum. The 4-bit DiT is 2.03 GiB for klein and 3.40 GiB for Z-Image; activations at 1024 px add to that, and no separate peak measurement was made.
|
| 114 |
+
|
| 115 |
+
</details>
|
| 116 |
+
|
| 117 |
+
<details>
|
| 118 |
+
<summary><b>Limitations</b></summary>
|
| 119 |
+
|
| 120 |
+
- Seeds are not compatible with the upstream pipelines, which use `torch.Generator`. The same prompt gives comparable images, never the same file.
|
| 121 |
+
- Both models are distilled: CFG is not applied, and step counts are low by design (8 for Z-Image, 4 for klein).
|
| 122 |
+
- Z-Image's Omni mode is not covered: it needs a SigLIP encoder that is not part of this bundle.
|
| 123 |
+
- Tiled VAE decode is available for high resolutions and gives a result that differs slightly from a single pass, exactly as it does in diffusers.
|
| 124 |
+
- Batches larger than one are not implemented.
|
| 125 |
+
|
| 126 |
+
</details>
|
| 127 |
+
|
| 128 |
+
## Licence
|
| 129 |
+
|
| 130 |
+
[Apache-2.0](https://www.apache.org/licenses/LICENSE-2.0), the licence of the upstream model. Conversion tooling and this card: MIT. The upstream card states usage restrictions that redistribution does not repeal.
|
shared/qwen3-4b-text-encoder/README.md
ADDED
|
@@ -0,0 +1,133 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: apache-2.0
|
| 3 |
+
library_name: mlx
|
| 4 |
+
tags:
|
| 5 |
+
- mlx
|
| 6 |
+
- mlx-swift
|
| 7 |
+
- apple-silicon
|
| 8 |
+
base_model:
|
| 9 |
+
- black-forest-labs/FLUX.2-klein-4B
|
| 10 |
+
---
|
| 11 |
+
|
| 12 |
+
<div align="center">
|
| 13 |
+
|
| 14 |
+
# Qwen3-4B text encoder for MLX, bf16
|
| 15 |
+
|
| 16 |
+
**Shared text encoder for both image models**
|
| 17 |
+
|
| 18 |
+
  [](https://www.apache.org/licenses/LICENSE-2.0) [](https://huggingface.co/kruatech/studio-mlx)
|
| 19 |
+
|
| 20 |
+
</div>
|
| 21 |
+
|
| 22 |
+
The text encoder both image models in this repository use. It is the Qwen3-4B encoder shipped inside the FLUX.2 [klein] and Z-Image-Turbo repositories; the two copies are identical except for the MLP of layer 35, which no pipeline reads.
|
| 23 |
+
|
| 24 |
+
Converted from [black-forest-labs/FLUX.2-klein-4B](https://huggingface.co/black-forest-labs/FLUX.2-klein-4B) for native MLX inference on Apple silicon.
|
| 25 |
+
|
| 26 |
+
## Download
|
| 27 |
+
|
| 28 |
+
Pair it with `z-image-turbo` or `flux2-klein-4b`. On its own it generates nothing.
|
| 29 |
+
|
| 30 |
+
```bash
|
| 31 |
+
pip install "huggingface_hub[hf_xet]"
|
| 32 |
+
|
| 33 |
+
hf download kruatech/studio-mlx --local-dir bundles \
|
| 34 |
+
--include "shared/qwen3-4b-text-encoder/*"
|
| 35 |
+
```
|
| 36 |
+
|
| 37 |
+
## Files
|
| 38 |
+
|
| 39 |
+
| component | class | precision | tensors | size |
|
| 40 |
+
|--------------|------------------|-----------|---------|----------|
|
| 41 |
+
| text_encoder | Qwen3ForCausalLM | keep | 398 | 7.51 GiB |
|
| 42 |
+
|
| 43 |
+
Folder total: **7.51 GiB**. `manifest.json` carries SHA-256, byte sizes and tensor counts for every file.
|
| 44 |
+
|
| 45 |
+
## How this model works in MLX
|
| 46 |
+
|
| 47 |
+
- Z-Image consumes `hidden_states[-2]`, the output of layer n-1, so 35 of 36 layers run.
|
| 48 |
+
- klein concatenates `hidden_states[9]`, `[18]` and `[27]` into 7680 features, so 27 of 36 layers run.
|
| 49 |
+
- Neither pipeline touches the last layer, so a loader can skip it entirely.
|
| 50 |
+
- Padding matters for klein and not for Z-Image: Z-Image drops padded positions with a mask afterwards, while klein feeds all 512 tokens to the transformer, so the attention mask has to be applied.
|
| 51 |
+
- `tokenizer/chat_format.json` holds the Qwen chat template already rendered into a prefix and a suffix for each pipeline, so a runtime needs no Jinja. klein renders with thinking disabled, which still injects an empty `<think></think>` block.
|
| 52 |
+
|
| 53 |
+
<details>
|
| 54 |
+
<summary><b>Tensor layout</b></summary>
|
| 55 |
+
|
| 56 |
+
MLX convolutions expect channels last, PyTorch expects channels first, so convolution weights are permuted during conversion. Everything else keeps its upstream shape.
|
| 57 |
+
|
| 58 |
+
| kind | PyTorch | MLX |
|
| 59 |
+
|-----------------------------|---------------------|---------------------|
|
| 60 |
+
| `Conv2d.weight` | `(out, in, kH, kW)` | `(out, kH, kW, in)` |
|
| 61 |
+
| `Linear`, norms, embeddings | `(out, in)` | unchanged |
|
| 62 |
+
|
| 63 |
+
No `weight_norm` and no 3-D convolutions appear in these models, so no folding was needed. Parameter names match the upstream checkpoints, so weights load without remapping.
|
| 64 |
+
|
| 65 |
+
</details>
|
| 66 |
+
|
| 67 |
+
<details>
|
| 68 |
+
<summary><b>Verification numbers</b></summary>
|
| 69 |
+
|
| 70 |
+
Every module was compared against the upstream reference on fixed inputs in float32. The metric is `rel_max = max|a-b| / max|a|`.
|
| 71 |
+
|
| 72 |
+
| module | rel_max | reference |
|
| 73 |
+
|-------------------------------------------------|-----------|---------------------------------|
|
| 74 |
+
| Qwen3 encoder, hidden_states[-2] | 2.7e-07 | transformers |
|
| 75 |
+
| Qwen3 encoder, layers 9/18/27 with padding mask | 6.0e-07 | transformers |
|
| 76 |
+
| Z-Image DiT | 4.2e-06 | diffusers |
|
| 77 |
+
| Flux2 DiT | 3.7e-07 | diffusers |
|
| 78 |
+
| VAE decoder | 1.3e-05 | diffusers |
|
| 79 |
+
| VAE encoder | 5.3e-06 | diffusers |
|
| 80 |
+
| VAE tiled decode | 5.1e-06 | diffusers tiled_decode |
|
| 81 |
+
| sigma schedule, static shift | 3.2e-08 | FlowMatchEulerDiscreteScheduler |
|
| 82 |
+
| sigma schedule, exponential dynamic shift | 7.7e-08 | FlowMatchEulerDiscreteScheduler |
|
| 83 |
+
| position ids, latent packing, patchify | bit-exact | pipeline helpers |
|
| 84 |
+
|
| 85 |
+
A Swift MLX implementation was then checked against the Python one:
|
| 86 |
+
|
| 87 |
+
| module | rel_max | note |
|
| 88 |
+
|------------------------------|-----------------|------------------------------------------------|
|
| 89 |
+
| tokenization, both pipelines | bit-exact | same ids |
|
| 90 |
+
| Qwen3 encoder, bf16 | 2.2e-04 | Z-Image branch |
|
| 91 |
+
| Qwen3 encoder, bf16 | 4.6e-03 | klein branch, 512 tokens with mask |
|
| 92 |
+
| Qwen3 encoder, 4-bit | 3.0e-04 | proves the quantized layout is rebuilt exactly |
|
| 93 |
+
| Z-Image DiT, float32 | 3.7e-07 | 8 layers |
|
| 94 |
+
| Flux2 DiT, float32 | 1.0e-06 | 2 double + 2 single blocks |
|
| 95 |
+
| VAE decode / tiled / encode | 1e-05 or better | both VAE classes |
|
| 96 |
+
| VAE BatchNorm statistics | 0 | exact |
|
| 97 |
+
|
| 98 |
+
The bf16 figures for a full-depth DiT are larger - 2e-02 for both models - and that is rounding order, not a defect: at float32 the same code agrees to 1e-06, and the deviation grows with depth from a bf16-level 1e-04 per block. Comparing two bf16 implementations below 1e-02 is not meaningful for a 30-block network.
|
| 99 |
+
|
| 100 |
+
</details>
|
| 101 |
+
|
| 102 |
+
<details>
|
| 103 |
+
<summary><b>Performance on an M1 Max</b></summary>
|
| 104 |
+
|
| 105 |
+
Mac Studio, Apple M1 Max, 10-core CPU, 24-core GPU, 32 GB, macOS 26.5. Release build, 4-bit weights, batch 1, warm page cache.
|
| 106 |
+
|
| 107 |
+
| run | per step | VAE decode |
|
| 108 |
+
|--------------------------------------|----------|------------|
|
| 109 |
+
| Z-Image, 512 px, 4 steps | 2.6 s | 1.5 s |
|
| 110 |
+
| klein, 512 px, 4 steps | 2.3 s | 0.9 s |
|
| 111 |
+
| klein, 1024 px, 4 steps | 8.0 s | 0.2 s |
|
| 112 |
+
| klein, 512 px + one 512 px reference | 4.1 s | 0.9 s |
|
| 113 |
+
|
| 114 |
+
Reading the weights dominates a cold run and depends entirely on the storage: on an external volume delivering about 0.2 GiB/s, the 4-bit encoder took 22-32 s and a 4-bit DiT 18-35 s, against 60-107 s for the same weights in bf16. On internal storage expect these to be several times shorter. Compute is unaffected: the numbers above are steady state.
|
| 115 |
+
|
| 116 |
+
Stages run one at a time and release their weights, so peak memory is set by the largest single component rather than their sum. The 4-bit DiT is 2.03 GiB for klein and 3.40 GiB for Z-Image; activations at 1024 px add to that, and no separate peak measurement was made.
|
| 117 |
+
|
| 118 |
+
</details>
|
| 119 |
+
|
| 120 |
+
<details>
|
| 121 |
+
<summary><b>Limitations</b></summary>
|
| 122 |
+
|
| 123 |
+
- Seeds are not compatible with the upstream pipelines, which use `torch.Generator`. The same prompt gives comparable images, never the same file.
|
| 124 |
+
- Both models are distilled: CFG is not applied, and step counts are low by design (8 for Z-Image, 4 for klein).
|
| 125 |
+
- Z-Image's Omni mode is not covered: it needs a SigLIP encoder that is not part of this bundle.
|
| 126 |
+
- Tiled VAE decode is available for high resolutions and gives a result that differs slightly from a single pass, exactly as it does in diffusers.
|
| 127 |
+
- Batches larger than one are not implemented.
|
| 128 |
+
|
| 129 |
+
</details>
|
| 130 |
+
|
| 131 |
+
## Licence
|
| 132 |
+
|
| 133 |
+
[Apache-2.0](https://www.apache.org/licenses/LICENSE-2.0), the licence of the upstream model. Conversion tooling and this card: MIT. The upstream card states usage restrictions that redistribution does not repeal.
|
z-image-turbo-q4/README.md
ADDED
|
@@ -0,0 +1,130 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: apache-2.0
|
| 3 |
+
library_name: mlx
|
| 4 |
+
tags:
|
| 5 |
+
- mlx
|
| 6 |
+
- mlx-swift
|
| 7 |
+
- apple-silicon
|
| 8 |
+
base_model:
|
| 9 |
+
- Tongyi-MAI/Z-Image-Turbo
|
| 10 |
+
---
|
| 11 |
+
|
| 12 |
+
<div align="center">
|
| 13 |
+
|
| 14 |
+
# Z-Image-Turbo for MLX, 4-bit
|
| 15 |
+
|
| 16 |
+
**Text to image, 6B, 8 steps**
|
| 17 |
+
|
| 18 |
+
  [](https://www.apache.org/licenses/LICENSE-2.0) [](https://huggingface.co/kruatech/studio-mlx)
|
| 19 |
+
|
| 20 |
+
</div>
|
| 21 |
+
|
| 22 |
+
4-bit version of Z-Image-Turbo. Everything in the bf16 card applies.
|
| 23 |
+
|
| 24 |
+
Converted from [Tongyi-MAI/Z-Image-Turbo](https://huggingface.co/Tongyi-MAI/Z-Image-Turbo) for native MLX inference on Apple silicon.
|
| 25 |
+
|
| 26 |
+
## Download
|
| 27 |
+
|
| 28 |
+
Also download `shared/qwen3-4b-text-encoder-q4`.
|
| 29 |
+
|
| 30 |
+
```bash
|
| 31 |
+
pip install "huggingface_hub[hf_xet]"
|
| 32 |
+
|
| 33 |
+
hf download kruatech/studio-mlx --local-dir bundles \
|
| 34 |
+
--include "shared/qwen3-4b-text-encoder-q4/*" "z-image-turbo-q4/*"
|
| 35 |
+
```
|
| 36 |
+
|
| 37 |
+
## Files
|
| 38 |
+
|
| 39 |
+
| component | class | precision | tensors | size |
|
| 40 |
+
|-------------|--------------------------|-----------------|---------|----------|
|
| 41 |
+
| transformer | ZImageTransformer2DModel | 4-bit, group 64 | 999 | 3.40 GiB |
|
| 42 |
+
| vae | AutoencoderKL | keep | 244 | 160 MiB |
|
| 43 |
+
|
| 44 |
+
Folder total: **3.55 GiB**. `manifest.json` carries SHA-256, byte sizes and tensor counts for every file.
|
| 45 |
+
|
| 46 |
+
## How this model works in MLX
|
| 47 |
+
|
| 48 |
+
- 239 linear layers quantized; normalizations, modulations and timestep MLPs stay in bf16. The VAE is not quantized - at 160 MiB there is nothing to gain.
|
| 49 |
+
|
| 50 |
+
<details>
|
| 51 |
+
<summary><b>Tensor layout</b></summary>
|
| 52 |
+
|
| 53 |
+
MLX convolutions expect channels last, PyTorch expects channels first, so convolution weights are permuted during conversion. Everything else keeps its upstream shape.
|
| 54 |
+
|
| 55 |
+
| kind | PyTorch | MLX |
|
| 56 |
+
|-----------------------------|---------------------|---------------------|
|
| 57 |
+
| `Conv2d.weight` | `(out, in, kH, kW)` | `(out, kH, kW, in)` |
|
| 58 |
+
| `Linear`, norms, embeddings | `(out, in)` | unchanged |
|
| 59 |
+
|
| 60 |
+
No `weight_norm` and no 3-D convolutions appear in these models, so no folding was needed. Parameter names match the upstream checkpoints, so weights load without remapping.
|
| 61 |
+
|
| 62 |
+
</details>
|
| 63 |
+
|
| 64 |
+
<details>
|
| 65 |
+
<summary><b>Verification numbers</b></summary>
|
| 66 |
+
|
| 67 |
+
Every module was compared against the upstream reference on fixed inputs in float32. The metric is `rel_max = max|a-b| / max|a|`.
|
| 68 |
+
|
| 69 |
+
| module | rel_max | reference |
|
| 70 |
+
|-------------------------------------------------|-----------|---------------------------------|
|
| 71 |
+
| Qwen3 encoder, hidden_states[-2] | 2.7e-07 | transformers |
|
| 72 |
+
| Qwen3 encoder, layers 9/18/27 with padding mask | 6.0e-07 | transformers |
|
| 73 |
+
| Z-Image DiT | 4.2e-06 | diffusers |
|
| 74 |
+
| Flux2 DiT | 3.7e-07 | diffusers |
|
| 75 |
+
| VAE decoder | 1.3e-05 | diffusers |
|
| 76 |
+
| VAE encoder | 5.3e-06 | diffusers |
|
| 77 |
+
| VAE tiled decode | 5.1e-06 | diffusers tiled_decode |
|
| 78 |
+
| sigma schedule, static shift | 3.2e-08 | FlowMatchEulerDiscreteScheduler |
|
| 79 |
+
| sigma schedule, exponential dynamic shift | 7.7e-08 | FlowMatchEulerDiscreteScheduler |
|
| 80 |
+
| position ids, latent packing, patchify | bit-exact | pipeline helpers |
|
| 81 |
+
|
| 82 |
+
A Swift MLX implementation was then checked against the Python one:
|
| 83 |
+
|
| 84 |
+
| module | rel_max | note |
|
| 85 |
+
|------------------------------|-----------------|------------------------------------------------|
|
| 86 |
+
| tokenization, both pipelines | bit-exact | same ids |
|
| 87 |
+
| Qwen3 encoder, bf16 | 2.2e-04 | Z-Image branch |
|
| 88 |
+
| Qwen3 encoder, bf16 | 4.6e-03 | klein branch, 512 tokens with mask |
|
| 89 |
+
| Qwen3 encoder, 4-bit | 3.0e-04 | proves the quantized layout is rebuilt exactly |
|
| 90 |
+
| Z-Image DiT, float32 | 3.7e-07 | 8 layers |
|
| 91 |
+
| Flux2 DiT, float32 | 1.0e-06 | 2 double + 2 single blocks |
|
| 92 |
+
| VAE decode / tiled / encode | 1e-05 or better | both VAE classes |
|
| 93 |
+
| VAE BatchNorm statistics | 0 | exact |
|
| 94 |
+
|
| 95 |
+
The bf16 figures for a full-depth DiT are larger - 2e-02 for both models - and that is rounding order, not a defect: at float32 the same code agrees to 1e-06, and the deviation grows with depth from a bf16-level 1e-04 per block. Comparing two bf16 implementations below 1e-02 is not meaningful for a 30-block network.
|
| 96 |
+
|
| 97 |
+
</details>
|
| 98 |
+
|
| 99 |
+
<details>
|
| 100 |
+
<summary><b>Performance on an M1 Max</b></summary>
|
| 101 |
+
|
| 102 |
+
Mac Studio, Apple M1 Max, 10-core CPU, 24-core GPU, 32 GB, macOS 26.5. Release build, 4-bit weights, batch 1, warm page cache.
|
| 103 |
+
|
| 104 |
+
| run | per step | VAE decode |
|
| 105 |
+
|--------------------------------------|----------|------------|
|
| 106 |
+
| Z-Image, 512 px, 4 steps | 2.6 s | 1.5 s |
|
| 107 |
+
| klein, 512 px, 4 steps | 2.3 s | 0.9 s |
|
| 108 |
+
| klein, 1024 px, 4 steps | 8.0 s | 0.2 s |
|
| 109 |
+
| klein, 512 px + one 512 px reference | 4.1 s | 0.9 s |
|
| 110 |
+
|
| 111 |
+
Reading the weights dominates a cold run and depends entirely on the storage: on an external volume delivering about 0.2 GiB/s, the 4-bit encoder took 22-32 s and a 4-bit DiT 18-35 s, against 60-107 s for the same weights in bf16. On internal storage expect these to be several times shorter. Compute is unaffected: the numbers above are steady state.
|
| 112 |
+
|
| 113 |
+
Stages run one at a time and release their weights, so peak memory is set by the largest single component rather than their sum. The 4-bit DiT is 2.03 GiB for klein and 3.40 GiB for Z-Image; activations at 1024 px add to that, and no separate peak measurement was made.
|
| 114 |
+
|
| 115 |
+
</details>
|
| 116 |
+
|
| 117 |
+
<details>
|
| 118 |
+
<summary><b>Limitations</b></summary>
|
| 119 |
+
|
| 120 |
+
- Seeds are not compatible with the upstream pipelines, which use `torch.Generator`. The same prompt gives comparable images, never the same file.
|
| 121 |
+
- Both models are distilled: CFG is not applied, and step counts are low by design (8 for Z-Image, 4 for klein).
|
| 122 |
+
- Z-Image's Omni mode is not covered: it needs a SigLIP encoder that is not part of this bundle.
|
| 123 |
+
- Tiled VAE decode is available for high resolutions and gives a result that differs slightly from a single pass, exactly as it does in diffusers.
|
| 124 |
+
- Batches larger than one are not implemented.
|
| 125 |
+
|
| 126 |
+
</details>
|
| 127 |
+
|
| 128 |
+
## Licence
|
| 129 |
+
|
| 130 |
+
[Apache-2.0](https://www.apache.org/licenses/LICENSE-2.0), the licence of the upstream model. Conversion tooling and this card: MIT. The upstream card states usage restrictions that redistribution does not repeal.
|
z-image-turbo/README.md
ADDED
|
@@ -0,0 +1,134 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: apache-2.0
|
| 3 |
+
library_name: mlx
|
| 4 |
+
tags:
|
| 5 |
+
- mlx
|
| 6 |
+
- mlx-swift
|
| 7 |
+
- apple-silicon
|
| 8 |
+
base_model:
|
| 9 |
+
- Tongyi-MAI/Z-Image-Turbo
|
| 10 |
+
---
|
| 11 |
+
|
| 12 |
+
<div align="center">
|
| 13 |
+
|
| 14 |
+
# Z-Image-Turbo for MLX, bf16
|
| 15 |
+
|
| 16 |
+
**Text to image, 6B, 8 steps**
|
| 17 |
+
|
| 18 |
+
  [](https://www.apache.org/licenses/LICENSE-2.0) [](https://huggingface.co/kruatech/studio-mlx)
|
| 19 |
+
|
| 20 |
+
</div>
|
| 21 |
+
|
| 22 |
+
6B single-stream DiT (S3-DiT) plus the flux-dev VAE, text to image, 8 steps, no CFG.
|
| 23 |
+
|
| 24 |
+
Converted from [Tongyi-MAI/Z-Image-Turbo](https://huggingface.co/Tongyi-MAI/Z-Image-Turbo) for native MLX inference on Apple silicon.
|
| 25 |
+
|
| 26 |
+
## Download
|
| 27 |
+
|
| 28 |
+
Also download `shared/qwen3-4b-text-encoder`. This folder has no `tokenizer/`: the tokenizer lives with the encoder, and without it a prompt cannot be built.
|
| 29 |
+
|
| 30 |
+
```bash
|
| 31 |
+
pip install "huggingface_hub[hf_xet]"
|
| 32 |
+
|
| 33 |
+
hf download kruatech/studio-mlx --local-dir bundles \
|
| 34 |
+
--include "shared/qwen3-4b-text-encoder/*" "z-image-turbo/*"
|
| 35 |
+
```
|
| 36 |
+
|
| 37 |
+
## Files
|
| 38 |
+
|
| 39 |
+
| component | class | precision | tensors | size |
|
| 40 |
+
|-------------|--------------------------|-----------|---------|-----------|
|
| 41 |
+
| transformer | ZImageTransformer2DModel | bf16 | 521 | 11.46 GiB |
|
| 42 |
+
| vae | AutoencoderKL | keep | 244 | 160 MiB |
|
| 43 |
+
|
| 44 |
+
Folder total: **11.62 GiB**. `manifest.json` carries SHA-256, byte sizes and tensor counts for every file.
|
| 45 |
+
|
| 46 |
+
## How this model works in MLX
|
| 47 |
+
|
| 48 |
+
- 30 main layers, 2 noise-refiner and 2 context-refiner layers, dim 3840, 30 heads.
|
| 49 |
+
- Rotary embedding over three axes with dims 32/48/48 and theta 256, applied to adjacent pairs rather than split halves.
|
| 50 |
+
- Image and caption tokens are padded to a multiple of 32 with learned pad tokens, and the sequence is `[image, caption]`.
|
| 51 |
+
- The scheduler uses the static shift of 3.0 from `scheduler_config.json`, and the prediction is negated before the Euler step.
|
| 52 |
+
- Latents scale by `1/0.3611` and shift by `0.1159` before decoding.
|
| 53 |
+
|
| 54 |
+
<details>
|
| 55 |
+
<summary><b>Tensor layout</b></summary>
|
| 56 |
+
|
| 57 |
+
MLX convolutions expect channels last, PyTorch expects channels first, so convolution weights are permuted during conversion. Everything else keeps its upstream shape.
|
| 58 |
+
|
| 59 |
+
| kind | PyTorch | MLX |
|
| 60 |
+
|-----------------------------|---------------------|---------------------|
|
| 61 |
+
| `Conv2d.weight` | `(out, in, kH, kW)` | `(out, kH, kW, in)` |
|
| 62 |
+
| `Linear`, norms, embeddings | `(out, in)` | unchanged |
|
| 63 |
+
|
| 64 |
+
No `weight_norm` and no 3-D convolutions appear in these models, so no folding was needed. Parameter names match the upstream checkpoints, so weights load without remapping.
|
| 65 |
+
|
| 66 |
+
</details>
|
| 67 |
+
|
| 68 |
+
<details>
|
| 69 |
+
<summary><b>Verification numbers</b></summary>
|
| 70 |
+
|
| 71 |
+
Every module was compared against the upstream reference on fixed inputs in float32. The metric is `rel_max = max|a-b| / max|a|`.
|
| 72 |
+
|
| 73 |
+
| module | rel_max | reference |
|
| 74 |
+
|-------------------------------------------------|-----------|---------------------------------|
|
| 75 |
+
| Qwen3 encoder, hidden_states[-2] | 2.7e-07 | transformers |
|
| 76 |
+
| Qwen3 encoder, layers 9/18/27 with padding mask | 6.0e-07 | transformers |
|
| 77 |
+
| Z-Image DiT | 4.2e-06 | diffusers |
|
| 78 |
+
| Flux2 DiT | 3.7e-07 | diffusers |
|
| 79 |
+
| VAE decoder | 1.3e-05 | diffusers |
|
| 80 |
+
| VAE encoder | 5.3e-06 | diffusers |
|
| 81 |
+
| VAE tiled decode | 5.1e-06 | diffusers tiled_decode |
|
| 82 |
+
| sigma schedule, static shift | 3.2e-08 | FlowMatchEulerDiscreteScheduler |
|
| 83 |
+
| sigma schedule, exponential dynamic shift | 7.7e-08 | FlowMatchEulerDiscreteScheduler |
|
| 84 |
+
| position ids, latent packing, patchify | bit-exact | pipeline helpers |
|
| 85 |
+
|
| 86 |
+
A Swift MLX implementation was then checked against the Python one:
|
| 87 |
+
|
| 88 |
+
| module | rel_max | note |
|
| 89 |
+
|------------------------------|-----------------|------------------------------------------------|
|
| 90 |
+
| tokenization, both pipelines | bit-exact | same ids |
|
| 91 |
+
| Qwen3 encoder, bf16 | 2.2e-04 | Z-Image branch |
|
| 92 |
+
| Qwen3 encoder, bf16 | 4.6e-03 | klein branch, 512 tokens with mask |
|
| 93 |
+
| Qwen3 encoder, 4-bit | 3.0e-04 | proves the quantized layout is rebuilt exactly |
|
| 94 |
+
| Z-Image DiT, float32 | 3.7e-07 | 8 layers |
|
| 95 |
+
| Flux2 DiT, float32 | 1.0e-06 | 2 double + 2 single blocks |
|
| 96 |
+
| VAE decode / tiled / encode | 1e-05 or better | both VAE classes |
|
| 97 |
+
| VAE BatchNorm statistics | 0 | exact |
|
| 98 |
+
|
| 99 |
+
The bf16 figures for a full-depth DiT are larger - 2e-02 for both models - and that is rounding order, not a defect: at float32 the same code agrees to 1e-06, and the deviation grows with depth from a bf16-level 1e-04 per block. Comparing two bf16 implementations below 1e-02 is not meaningful for a 30-block network.
|
| 100 |
+
|
| 101 |
+
</details>
|
| 102 |
+
|
| 103 |
+
<details>
|
| 104 |
+
<summary><b>Performance on an M1 Max</b></summary>
|
| 105 |
+
|
| 106 |
+
Mac Studio, Apple M1 Max, 10-core CPU, 24-core GPU, 32 GB, macOS 26.5. Release build, 4-bit weights, batch 1, warm page cache.
|
| 107 |
+
|
| 108 |
+
| run | per step | VAE decode |
|
| 109 |
+
|--------------------------------------|----------|------------|
|
| 110 |
+
| Z-Image, 512 px, 4 steps | 2.6 s | 1.5 s |
|
| 111 |
+
| klein, 512 px, 4 steps | 2.3 s | 0.9 s |
|
| 112 |
+
| klein, 1024 px, 4 steps | 8.0 s | 0.2 s |
|
| 113 |
+
| klein, 512 px + one 512 px reference | 4.1 s | 0.9 s |
|
| 114 |
+
|
| 115 |
+
Reading the weights dominates a cold run and depends entirely on the storage: on an external volume delivering about 0.2 GiB/s, the 4-bit encoder took 22-32 s and a 4-bit DiT 18-35 s, against 60-107 s for the same weights in bf16. On internal storage expect these to be several times shorter. Compute is unaffected: the numbers above are steady state.
|
| 116 |
+
|
| 117 |
+
Stages run one at a time and release their weights, so peak memory is set by the largest single component rather than their sum. The 4-bit DiT is 2.03 GiB for klein and 3.40 GiB for Z-Image; activations at 1024 px add to that, and no separate peak measurement was made.
|
| 118 |
+
|
| 119 |
+
</details>
|
| 120 |
+
|
| 121 |
+
<details>
|
| 122 |
+
<summary><b>Limitations</b></summary>
|
| 123 |
+
|
| 124 |
+
- Seeds are not compatible with the upstream pipelines, which use `torch.Generator`. The same prompt gives comparable images, never the same file.
|
| 125 |
+
- Both models are distilled: CFG is not applied, and step counts are low by design (8 for Z-Image, 4 for klein).
|
| 126 |
+
- Z-Image's Omni mode is not covered: it needs a SigLIP encoder that is not part of this bundle.
|
| 127 |
+
- Tiled VAE decode is available for high resolutions and gives a result that differs slightly from a single pass, exactly as it does in diffusers.
|
| 128 |
+
- Batches larger than one are not implemented.
|
| 129 |
+
|
| 130 |
+
</details>
|
| 131 |
+
|
| 132 |
+
## Licence
|
| 133 |
+
|
| 134 |
+
[Apache-2.0](https://www.apache.org/licenses/LICENSE-2.0), the licence of the upstream model. Conversion tooling and this card: MIT. The upstream card states usage restrictions that redistribution does not repeal.
|