xocialize's picture
README: int8 tier β€” measured quality, footprints, and why there is no 4-bit tier
1b13810 verified
|
Raw
History Blame Contribute Delete
4.42 kB
---
license: other
license_name: nvidia-open-model-license
license_link: https://www.nvidia.com/en-us/agreements/enterprise-software/nvidia-open-model-license
base_model:
- nvidia/Qwen-Image-Flash
pipeline_tag: text-to-image
library_name: mlx
tags:
- mlx
- mlx-swift
- text-to-image
- qwen-image
- dmd2
- few-step
- nvidia
- quantized
---
# Qwen-Image-Flash-8bit (MLX)
An **int8** MLX snapshot of [nvidia/Qwen-Image-Flash](https://huggingface.co/nvidia/Qwen-Image-Flash) β€”
NVIDIA's DMD2 four-step distillation of [Qwen/Qwen-Image](https://huggingface.co/Qwen/Qwen-Image).
28 GB total, versus 57.7 GB for
[the bf16 snapshot](https://huggingface.co/mlx-community/Qwen-Image-Flash-bf16).
This is the tier that makes a 20B-parameter image model reachable on ordinary Apple Silicon: a
**30.0 GB peak** at 1024Β²/4 steps instead of bf16's 57.4 GB, and it is **4Γ— faster**.
## Contents
| file | precision | size |
|---|---|---|
| `transformer/model-int8.safetensors` | int8 attention + feed-forward + modulation, group 64 | 21.75 GB |
| `text_encoder/model-int8.safetensors` | int8 Qwen2.5-VL-7B **language model only** | 7.51 GB |
| `vae/` | unquantized | 0.25 GB |
The `img_in` / `txt_in` / `time_text_embed` / `norm_out` / `proj_out` projections are left at full
precision, and the VAE is never quantized β€” decode is where precision loss shows up as visible
colour and banding artifacts.
The text encoder carries **no vision tower**: text-to-image conditions on text alone, so the ViT is
dead weight in this pipeline and is omitted entirely.
## Measured quality
Against PyTorch fp32 goldens (diffusers 0.37.1), on identical injected inputs:
| | int8 | bf16 | fp32 oracle |
|---|---|---|---|
| DiT step-0 cosine | **0.9973** | 0.99836 | 1.0 |
| VL-7B prompt-embed cosine | **0.99992** | 0.9999926 | 1.0 |
| 1024Β²/4-step render | 19.8 s | 83.3 s | β€” |
| load | 2.3 s | ~60 s | β€” |
| peak memory | **30.0 GB** | 57.4 GB | β€” |
Renders at this tier are visually indistinguishable from bf16 at the same seed.
### Why there is no 4-bit tier
int4 was built and measured, not skipped: DiT step-0 cosine **0.9623** at group 64 and **0.9659** at
group 32, with the VL encoder at 0.9845. The 1024Β² render came out visibly soft and washed out, with
fine fur and snow detail gone. Finer scale groups did not rescue it β€” this DiT is intrinsically lossy
at 4 bits β€” so no 4-bit snapshot is published rather than shipping one that looks like that.
## Inference notes that are easy to get wrong
The distillation **internalized CFG 4.0**, and the packaged scheduler is **static shift-3**
(`use_dynamic_shifting: false`):
- `num_inference_steps = 4`, `true_cfg_scale = 1.0` β€” applying CFG again double-counts guidance the
student already absorbed, and doubles the transformer evaluations per step for nothing.
- The four-step trajectory is sigmas `[1.0, 0.9, 0.75, 0.5, 0.0]`.
- Tested at 1024 Γ— 1024. Use width/height divisible by 16.
## Use from Swift (MLXEngine)
```swift
import MLXQwenImageFlash
import MLXToolKit
let package = QwenImageFlashPackage(configuration: .init(quant: .int8))
try await package.load()
let response = try await package.run(T2IRequest(
prompt: "A red fox in a snowy pine forest at golden hour, photorealistic",
width: 1024, height: 1024, seed: 42)) as! T2IResponse
```
Port: [xocialize/qwen-image-edit-swift](https://github.com/xocialize/qwen-image-edit-swift) (MIT).
The package also selects this tier automatically when the engine's memory governor reports a budget
that cannot seat bf16 β€” and it resolves that *before* downloading, so a constrained machine fetches
this 28 GB snapshot rather than 41 GB of bf16 it could never load.
These are **pre-quantized** weights. Consumers never materialize bf16 at any point, which is the
difference between a tier that runs on a 32–48 GB machine and one that merely claims to: quantizing
at load would require holding the 41 GB bf16 transformer first.
## License
Governing terms: **[NVIDIA Open Model License Agreement](https://www.nvidia.com/en-us/agreements/enterprise-software/nvidia-open-model-license)**.
Additional information: Apache License 2.0 (`LICENSE`).
> Licensed by NVIDIA Corporation under the NVIDIA Open Model License
Commercial use, derivative models, and redistribution are permitted; if you redistribute these
weights you must pass on the Agreement and this notice (Β§3.1). See `NOTICE`.