Instructions to use mlx-community/Ming-Image-0.1-Design-bf16 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use mlx-community/Ming-Image-0.1-Design-bf16 with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] huggingface-cli download --local-dir Ming-Image-0.1-Design-bf16 mlx-community/Ming-Image-0.1-Design-bf16
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
mlx-community/Ming-Image-0.1-Design-bf16
MLX (bf16) snapshot of inclusionAI/Ming-Image-0.1-Design (MIT) for Apple Silicon, loaded by the Swift/MLX port ming-image-swift.
Ming-Image-0.1-Design is a design-focused text-to-image model for posters, signage backings, cards, UI layouts and typography, and it generates RGBA images. A BailingMM2 MLLM (Ling-mini-2.0 MoE) conditions a Z-Image-style S3-DiT through learnable query tokens and a direct hidden-state stream; a 4-channel Qwen-Image VAE decodes colour and alpha.
Sibling: mlx-community/Ming-Image-0.1-Design-Layer-bf16, which decomposes a flat design into editable RGBA layers.
Quantized tiers of this model:
- 8-bit: 28.3 GB, matches bf16 on every quality gate, 48 GB Macs;
- 4-bit: 20.5 GB, a small measured cost, 36 GB Macs.
What's in this repo
The upstream diffusers-style tree, with each component stored at the precision the port runs it:
| Component | Stored | vs upstream |
|---|---|---|
mllm/: Ling-mini-2.0 MoE MLLM, Qwen2.5 ViT, tokenizer |
bf16, 34.0 GB | byte-identical |
connector/: Qwen2-1.5B non-causal connector |
bf16, 3.1 GB | cast from f32 |
mlp/: query tokens and projection heads |
f32, 0.12 GB | byte-identical |
transformer/: S3-DiT |
bf16, 12.3 GB | byte-identical |
vae/: RGBA VAE |
bf16, 0.25 GB | byte-identical |
Total: 49.8 GB. The one cast is MLX's own f32 → bf16 conversion, the same one the port applies when it loads the upstream checkpoint, so the port runs bit-identical parameters from either repo (checked tensor by tensor).
Parity (Swift port vs the PyTorch reference)
Components on the fp32 CPU parity lane, against reference goldens:
| Component | Result |
|---|---|
| S3-DiT at 256² and 1024² | relL2 ≤ 1.5e-5 |
| VAE decode / encode | relL2 9.1e-7 / 3.5e-6 |
| Scheduler | bit-exact |
| Tokenizer, chat template, 3-D position ids | exact |
| MoE hidden states (with the reference's expert routing; ties make unforced routing implementation-defined) | ≤ 3.3e-6 |
| Conditioning outputs (same routing) | ≤ 3.6e-5 |
End to end:
- 2-step fp32 generation: final latents relL2 3.5e-6, and the RGBA output is within 1 LSB.
- Production bf16 on the GPU lands 41.9 dB from the fp32-truth render at 1024².
Native alpha
Transparency is prompt-driven. The Swift package exposes it as a request field (background: .transparent, MLXEngine
contract 1.48.0) and runs the recipe we measured on this port:
- the phrase "transparent canvas, not white, not checkerboard" first;
- a retry at the same seed with "不要白底,不要棋盘格,只要透明通道" when the alpha comes back opaque;
- then an alpha floor (α < 0.1 → 0).
That recipe gave usable alpha on 5 of the 6 test subjects. Glass and other transparent materials fail: they come back opaque, or with a checkerboard painted into the alpha. Matte those instead.
Memory and speed (M5 Max)
Measured as process phys_footprint, with MLX's buffer cache capped at 2 GB (MLXEngine's default).
- Post-load resident: 13.2 GB (the DiT and the VAE).
- Peak: 49.9 GB at 1024², 2048² and 2560×1440 alike. That peak is the conditioning stage: the 34 GB MLLM is loaded, conditions the prompt, and is released before denoising. The 2048²-class VAE decode is bounded, using a chunked mid-block attention and a tiled up path, and is exact.
- Under MLXEngine: it declares 13.5 GB resident plus 44.1 GB activation. Against the engine's 0.7× memory budget, that needs a 96 GB Mac. On 48 GB use the 8-bit tier, and on 36 GB the 4-bit tier.
- Speed at 12 steps, including conditioning: 39 s at 1024², 245 s at 2048².
Use (Swift / MLXEngine)
import MLXMingImage
import MLXToolKit
let package = MingImageT2IPackage(configuration: MingImageConfiguration(snapshotPath: "<this repo, downloaded>"))
try await package.load()
let response = try await package.run(T2IRequest(
prompt: "A modern tech conference poster titled 'MLX SUMMIT 2026', bold geometric shapes",
width: 1024, height: 1024, seed: 42)) as! T2IResponse
// response.image: an RGBA PNG with straight alpha; pass background: .transparent for a cut-out asset
Under MLXEngine, register MingImageT2IPackage.registration with
MingImageConfiguration(repo: "mlx-community/Ming-Image-0.1-Design"), and the engine downloads this repo on first
use.
License
MIT, as the upstream weights. The upstream LICENSE is included.
Quantized
Model tree for mlx-community/Ming-Image-0.1-Design-bf16
Base model
inclusionAI/Ming-Image-0.1-Design