File size: 11,027 Bytes
e0d5e12
c029384
e0d5e12
c029384
 
 
 
ee26c6d
c029384
 
 
 
 
 
 
 
e0d5e12
c029384
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
b861069
 
c029384
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
---
library_name: diffusers
license: apache-2.0
license_link: https://huggingface.co/Qwen/Qwen-Image-2512/blob/main/LICENSE
pipeline_tag: text-to-image
base_model:
- Qwen/Qwen-Image-2512
base_model_relation: quantized
tags:
- qwen-image
- text-to-image
- image-generation
- vllm
- vllm-omni
- humming
- byteshape
---

# Qwen-Image-2512 Humming (ShapeLearn Quantized, vLLM-Omni)

This is a quantized version of **Qwen-Image-2512** for the **[vLLM-Omni](https://github.com/vllm-project/vllm-omni)** backend with optimized **Humming kernels**, the fastest way we ship this model. Quantization was produced with **ByteShape's ShapeLearn**, which learns the optimal datatype per tensor to maintain high quality even at very low bitlengths.

> **Experimental.** We tested this flow with vLLM 0.24.0, vLLM-Omni 0.24.0, `vllm-omni-humming` 0.3.0, and `humming-kernels` 0.1.11 on NVIDIA 30-, 40-, and 50-series and RTX Pro 6000 GPUs. **Linux and NVIDIA only** (SM75+, driver ≥ 575).

Six sizes are available, from **17.29 GB (6.77 bpw)** down to **7.83 GB (3.07 bpw)**. All variants generate high-quality images.

Each package bundles the text encoder, VAE, scheduler, and tokenizer, so there are no separate component downloads: you pick a size and serve.

> Prefer ComfyUI or a wider hardware range? We also publish GGUF builds of the same model for ComfyUI and `stable-diffusion.cpp`: [byteshape/Qwen-Image-2512-GGUF](https://huggingface.co/byteshape/Qwen-Image-2512-GGUF).

To learn more about ShapeLearn, and for the full walkthrough with screenshots and a downloadable ComfyUI workflow, visit our [blog](https://byteshape.com/blogs/Qwen-Image-2512/).

If you have questions or want to share feedback, reach us on [Reddit](https://www.reddit.com/r/ByteShape/).

## A Note on Speed

Unlike the decoding stage of an LLM, diffusion inference is **not** heavily constrained by memory bandwidth. Compressing the model therefore does not necessarily make image generation faster. Quantization here buys you **VRAM headroom**, not throughput. Backend optimization is what moves the needle.

This backend uses optimized Humming kernels and is roughly **2-3x faster** than the GGUF path. On an RTX 5090 each step takes around **0.5 seconds**, so a 20-step image samples in about **10 seconds**.

| Backend | Time per step (RTX 5090) | Time for 20 steps |
|---|---|---|
| **vLLM-Omni** (this repo) | \~0.5 s | \~10 s |
| [GGUF](https://huggingface.co/byteshape/Qwen-Image-2512-GGUF) | \~1.2 s | \~24 s |

For measured peak VRAM and end-to-end times for **every size at four output resolutions**, with and without CPU offload, see the [VRAM and speed table](https://byteshape.com/blogs/Qwen-Image-2512/#requirements) in the blog post.

## What Do the Sizes Look Like?

![One prompt rendered by BF16 and all six Humming sizes, with a matching close-up crop of the ring face under each variant](img/collage.jpg)

Quality does not fall off a cliff as the bits come off. BF16, 6.77 and 5.33 bpw are near-interchangeable on this prompt, down to the same scratch fan and specular streak; from 4.70 bpw a soft grey haze spreads across the ring face and the surviving scratches grow thicker and fewer at each step, and at 3.07 bpw the composition itself drifts, with a narrower band sitting higher in the frame; even there the image is still a clean macro shot. Differences like these are subtle and highly prompt-dependent, and the row above is one prompt out of 24.

Judge for yourself: the [blog post](https://byteshape.com/blogs/Qwen-Image-2512/) walks through the full set, and the [interactive comparison explorer](https://byteshape.com/blogs/Qwen-Image-2512/comparison/) lets you A/B any two variants across 24 curated prompts at full resolution, with a slider, zoom, and a curator note for each prompt. If you read the outputs differently, tell us on [Reddit](https://www.reddit.com/r/ByteShape/).

## Repository Layout

The package is a shared **shell** (the pipeline) plus a pool of **transformers**, one per size:

```
Qwen-Image-2512/          # the shell: text encoder, VAE, scheduler, tokenizer
quantized_models/
  6.77bpw/transformer/
  5.33bpw/transformer/
  4.70bpw/transformer/
  4.18bpw/transformer/
  3.57bpw/transformer/
  3.07bpw/transformer/
```

You download the shell once, then drop the transformer of your choice into it.

## Available Models

| Model ID | Bits/Weight | Transformer Size | Folder |
|---|---|---|---|
| VLLM-1 | 3.07 | 7.83 GB | `quantized_models/3.07bpw` |
| VLLM-2 | 3.57 | 9.11 GB | `quantized_models/3.57bpw` |
| VLLM-3 | 4.18 | 10.68 GB | `quantized_models/4.18bpw` |
| VLLM-4 | 4.70 | 12.00 GB | `quantized_models/4.70bpw` |
| VLLM-5 | 5.33 | 13.60 GB | `quantized_models/5.33bpw` |
| VLLM-6 | 6.77 | 17.29 GB | `quantized_models/6.77bpw` |

**On VRAM.** The size above is the diffusion transformer only. The shell adds a **BF16 text encoder (\~16.6 GB)** and the VAE (\~0.25 GB), so keeping everything resident on the GPU needs roughly `transformer + 17 GB`: about 25 GB for the smallest size and 34 GB for the largest.

- **≤ 32 GB cards (including the RTX 5090):** serve with `--enable-cpu-offload`. All six sizes work; the text encoder lives in system RAM.
- **48 GB and larger (e.g. RTX Pro 6000):** drop `--enable-cpu-offload` for full speed.

**Selection rule:** take the largest model that comfortably fits your card. Since compression does not speed up diffusion, there is little reason to go smaller than your VRAM allows.

## Quick Start

### 1. Set up the environment

```bash
conda create -y -n qwen-image-humming python=3.12
conda activate qwen-image-humming

# torch trio (cu129)
pip install torch==2.11.0 torchvision==0.26.0 torchaudio==2.11.0 \
  --index-url https://download.pytorch.org/whl/cu129

# vLLM 0.24.0 (cu129 wheel)
printf 'torch==2.11.0+cu129\ntorchvision==0.26.0+cu129\ntorchaudio==2.11.0+cu129\n' > constraints.txt
PIP_CONSTRAINT=constraints.txt pip install \
  https://github.com/vllm-project/vllm/releases/download/v0.24.0/vllm-0.24.0+cu129-cp38-abi3-manylinux_2_28_x86_64.whl

# vLLM-Omni + our loader plugin
PIP_CONSTRAINT=constraints.txt pip install vllm-omni==0.24.0
PIP_CONSTRAINT=constraints.txt pip install vllm-omni-humming==0.3.0

# humming kernels: must be >= 0.1.11 (older versions produce corrupted images)
pip install --no-deps "humming-kernels>=0.1.11"
```

`vllm-omni-humming` is our own plugin, and it is the one required extra piece for this path: stock vLLM-Omni cannot read a pre-quantized Humming checkpoint on its own. The plugin registers itself and detects our models from their `transformer/config.json`.

### 2. Download the model

Grab the shell and the size you want:

```bash
pip install -U huggingface_hub

# the shell + one transformer (here the 3.07 bpw size)
hf download byteshape/Qwen-Image-2512-Humming \
  --include "Qwen-Image-2512/*" --include "quantized_models/3.07bpw/*" \
  --local-dir Qwen-Image-2512-Humming
cd Qwen-Image-2512-Humming
```

Then drop your chosen transformer into the shell. The folder is already named `transformer`, so it lands with the correct name:

```bash
# symlink it in (instant, no extra disk) ...
ln -sfn ../quantized_models/3.07bpw/transformer Qwen-Image-2512/transformer

# ... or move it in (a real, self-contained model dir)
mv quantized_models/3.07bpw/transformer Qwen-Image-2512/transformer
```

To try a different size later, swap the symlink, or move one transformer out and another in. The folders are named by bit-width (`6.77bpw`, `4.18bpw`, `3.07bpw`, and so on).

### 3. Serve

From the folder root:

```bash
vllm-omni serve Qwen-Image-2512 \
  --omni --served-model-name Qwen/Qwen-Image-2512 \
  --enable-cpu-offload --port 8124
```

Keep `--enable-cpu-offload` on cards with ≤ 32 GB (an RTX 5090 needs it); drop it on larger GPUs for full speed.

Do **not** pass `--diffusion-quantization-config` or set any `VLLM_HUMMING_*` variables. The plugin drives everything from the checkpoint.

The first serve on a new GPU spends a few minutes compiling kernels for your card, then caches them.

### 4. Generate

The server exposes the OpenAI images API:

```bash
curl -s http://127.0.0.1:8124/v1/images/generations \
  -H "Content-Type: application/json" \
  -d '{
    "model": "Qwen/Qwen-Image-2512",
    "prompt": "A red panda on a mossy log at dawn",
    "size": "1024x1024",
    "num_inference_steps": 20,
    "true_cfg_scale": 2.5,
    "negative_prompt": " ",
    "seed": 1234
  }' \
  | python3 -c "import sys,json,base64; d=json.load(sys.stdin); open('out.png','wb').write(base64.b64decode(d['data'][0]['b64_json']))"
```

**Always include a `negative_prompt`, even a single space.** vLLM-Omni applies classifier-free guidance only when a negative prompt is present with `true_cfg_scale > 1`; with none, guidance is skipped and the image comes out flatter and lower-contrast.

## Driving the Server from ComfyUI

If you prefer a GUI over `curl`, the `ComfyUI-vLLM-Omni` custom node sends generation requests to a running vLLM-Omni server, local or remote. ComfyUI is only the front end; vLLM still does the work.

Install the node into your ComfyUI installation:

```bash
cd /path/to/ComfyUI/custom_nodes
git clone https://github.com/vllm-project/vllm-omni.git
cp -r vllm-omni/apps/ComfyUI-vLLM-Omni ComfyUI-vLLM-Omni
rm -rf vllm-omni
```

Restart ComfyUI (`python main.py`), then open a browser and go to **http://127.0.0.1:8188** to reach the ComfyUI interface. A new **vLLM-Omni** folder appears in the **Node Library** sidebar. No extra Python packages are needed beyond ComfyUI's own.

Keep the vLLM-Omni server from step 3 running. ComfyUI reaches it at `http://127.0.0.1:8124/v1`. If ComfyUI runs on a different machine, use the server's address instead of `127.0.0.1`.

For the workflow, download `qwen-image-humming-vllm-omni.json` from our [blog post](https://byteshape.com/blogs/Qwen-Image-2512/) and load it with **`Ctrl+O`**, or start from ComfyUI's built-in template at **sidebar → Templates → vLLM-Omni → vLLM-Omni Image Generation**.

In the **Generate Image** node, set:

- **url** → `http://127.0.0.1:8124/v1`
- **model** → `Qwen/Qwen-Image-2512`
- **prompt** / **negative_prompt** → your text (negative: at least a space)
- **width** / **height** → e.g. 1024 x 1024

In the attached **Diffusion Sampling Params** node, set the seed, `num_inference_steps` (e.g. 20), and `true_cfg_scale` (e.g. 2.5), then run the workflow.

## Troubleshooting

| Symptom | Cause and fix |
|---|---|
| `HummingConfig ... unexpected keyword argument`, or a shape mismatch on `img_mod.1` / `txt_mod.1` at load | The plugin is not installed in the active environment. Check with `pip show vllm-omni-humming`. |
| Grainy or speckled images | The Humming kernels are too old. Reinstall with `pip install --no-deps "humming-kernels>=0.1.11"`. |
| Flat, low-contrast images | No `negative_prompt` was sent, so CFG was skipped. Pass one; a single space is enough. |
| Out of memory on serve | Add `--enable-cpu-offload`, or pick a smaller bit-width from the table above. |