mlx-community/Ming-Image-0.1-Design-bf16

MLX (bf16) snapshot of inclusionAI/Ming-Image-0.1-Design (MIT) for Apple Silicon, loaded by the Swift/MLX port ming-image-swift.

Ming-Image-0.1-Design is a design-focused text-to-image model for posters, signage backings, cards, UI layouts and typography, and it generates RGBA images. A BailingMM2 MLLM (Ling-mini-2.0 MoE) conditions a Z-Image-style S3-DiT through learnable query tokens and a direct hidden-state stream; a 4-channel Qwen-Image VAE decodes colour and alpha.

Sibling: mlx-community/Ming-Image-0.1-Design-Layer-bf16, which decomposes a flat design into editable RGBA layers.

Quantized tiers of this model:

  • 8-bit: 28.3 GB, matches bf16 on every quality gate, 48 GB Macs;
  • 4-bit: 20.5 GB, a small measured cost, 36 GB Macs.

What's in this repo

The upstream diffusers-style tree, with each component stored at the precision the port runs it:

Component Stored vs upstream
mllm/: Ling-mini-2.0 MoE MLLM, Qwen2.5 ViT, tokenizer bf16, 34.0 GB byte-identical
connector/: Qwen2-1.5B non-causal connector bf16, 3.1 GB cast from f32
mlp/: query tokens and projection heads f32, 0.12 GB byte-identical
transformer/: S3-DiT bf16, 12.3 GB byte-identical
vae/: RGBA VAE bf16, 0.25 GB byte-identical

Total: 49.8 GB. The one cast is MLX's own f32 → bf16 conversion, the same one the port applies when it loads the upstream checkpoint, so the port runs bit-identical parameters from either repo (checked tensor by tensor).

Parity (Swift port vs the PyTorch reference)

Components on the fp32 CPU parity lane, against reference goldens:

Component Result
S3-DiT at 256² and 1024² relL2 ≤ 1.5e-5
VAE decode / encode relL2 9.1e-7 / 3.5e-6
Scheduler bit-exact
Tokenizer, chat template, 3-D position ids exact
MoE hidden states (with the reference's expert routing; ties make unforced routing implementation-defined) ≤ 3.3e-6
Conditioning outputs (same routing) ≤ 3.6e-5

End to end:

  • 2-step fp32 generation: final latents relL2 3.5e-6, and the RGBA output is within 1 LSB.
  • Production bf16 on the GPU lands 41.9 dB from the fp32-truth render at 1024².

Native alpha

Transparency is prompt-driven. The Swift package exposes it as a request field (background: .transparent, MLXEngine contract 1.48.0) and runs the recipe we measured on this port:

  • the phrase "transparent canvas, not white, not checkerboard" first;
  • a retry at the same seed with "不要白底,不要棋盘格,只要透明通道" when the alpha comes back opaque;
  • then an alpha floor (α < 0.1 → 0).

That recipe gave usable alpha on 5 of the 6 test subjects. Glass and other transparent materials fail: they come back opaque, or with a checkerboard painted into the alpha. Matte those instead.

Memory and speed (M5 Max)

Measured as process phys_footprint, with MLX's buffer cache capped at 2 GB (MLXEngine's default).

  • Post-load resident: 13.2 GB (the DiT and the VAE).
  • Peak: 49.9 GB at 1024², 2048² and 2560×1440 alike. That peak is the conditioning stage: the 34 GB MLLM is loaded, conditions the prompt, and is released before denoising. The 2048²-class VAE decode is bounded, using a chunked mid-block attention and a tiled up path, and is exact.
  • Under MLXEngine: it declares 13.5 GB resident plus 44.1 GB activation. Against the engine's 0.7× memory budget, that needs a 96 GB Mac. On 48 GB use the 8-bit tier, and on 36 GB the 4-bit tier.
  • Speed at 12 steps, including conditioning: 39 s at 1024², 245 s at 2048².

Use (Swift / MLXEngine)

import MLXMingImage
import MLXToolKit

let package = MingImageT2IPackage(configuration: MingImageConfiguration(snapshotPath: "<this repo, downloaded>"))
try await package.load()
let response = try await package.run(T2IRequest(
    prompt: "A modern tech conference poster titled 'MLX SUMMIT 2026', bold geometric shapes",
    width: 1024, height: 1024, seed: 42)) as! T2IResponse
// response.image: an RGBA PNG with straight alpha; pass background: .transparent for a cut-out asset

Under MLXEngine, register MingImageT2IPackage.registration with MingImageConfiguration(repo: "mlx-community/Ming-Image-0.1-Design"), and the engine downloads this repo on first use.

License

MIT, as the upstream weights. The upstream LICENSE is included.

Downloads last month

-

Downloads are not tracked for this model. How to track
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for mlx-community/Ming-Image-0.1-Design-bf16

Finetuned
(15)
this model

Collection including mlx-community/Ming-Image-0.1-Design-bf16