--- library_name: diffusers license: apache-2.0 license_link: https://huggingface.co/Qwen/Qwen-Image-2512/blob/main/LICENSE pipeline_tag: text-to-image base_model: - Qwen/Qwen-Image-2512 base_model_relation: quantized tags: - qwen-image - text-to-image - image-generation - vllm - vllm-omni - humming - byteshape --- # Qwen-Image-2512 Humming (ShapeLearn Quantized, vLLM-Omni) This is a quantized version of **Qwen-Image-2512** for the **[vLLM-Omni](https://github.com/vllm-project/vllm-omni)** backend with optimized **Humming kernels**, the fastest way we ship this model. Quantization was produced with **ByteShape's ShapeLearn**, which learns the optimal datatype per tensor to maintain high quality even at very low bitlengths. > **Experimental.** We tested this flow with vLLM 0.24.0, vLLM-Omni 0.24.0, `vllm-omni-humming` 0.3.0, and `humming-kernels` 0.1.11 on NVIDIA 30-, 40-, and 50-series and RTX Pro 6000 GPUs. **Linux and NVIDIA only** (SM75+, driver ≥ 575). Six sizes are available, from **17.29 GB (6.77 bpw)** down to **7.83 GB (3.07 bpw)**. All variants generate high-quality images. Each package bundles the text encoder, VAE, scheduler, and tokenizer, so there are no separate component downloads: you pick a size and serve. > Prefer ComfyUI or a wider hardware range? We also publish GGUF builds of the same model for ComfyUI and `stable-diffusion.cpp`: [byteshape/Qwen-Image-2512-GGUF](https://huggingface.co/byteshape/Qwen-Image-2512-GGUF). To learn more about ShapeLearn, and for the full walkthrough with screenshots and a downloadable ComfyUI workflow, visit our [blog](https://byteshape.com/blogs/Qwen-Image-2512/). If you have questions or want to share feedback, reach us on [Reddit](https://www.reddit.com/r/ByteShape/). ## A Note on Speed Unlike the decoding stage of an LLM, diffusion inference is **not** heavily constrained by memory bandwidth. Compressing the model therefore does not necessarily make image generation faster. Quantization here buys you **VRAM headroom**, not throughput. Backend optimization is what moves the needle. This backend uses optimized Humming kernels and is roughly **2-3x faster** than the GGUF path. On an RTX 5090 each step takes around **0.5 seconds**, so a 20-step image samples in about **10 seconds**. | Backend | Time per step (RTX 5090) | Time for 20 steps | |---|---|---| | **vLLM-Omni** (this repo) | \~0.5 s | \~10 s | | [GGUF](https://huggingface.co/byteshape/Qwen-Image-2512-GGUF) | \~1.2 s | \~24 s | For measured peak VRAM and end-to-end times for **every size at four output resolutions**, with and without CPU offload, see the [VRAM and speed table](https://byteshape.com/blogs/Qwen-Image-2512/#requirements) in the blog post. ## What Do the Sizes Look Like? ![One prompt rendered by BF16 and all six Humming sizes, with a matching close-up crop of the ring face under each variant](img/collage.jpg) Quality does not fall off a cliff as the bits come off. BF16, 6.77 and 5.33 bpw are near-interchangeable on this prompt, down to the same scratch fan and specular streak; from 4.70 bpw a soft grey haze spreads across the ring face and the surviving scratches grow thicker and fewer at each step, and at 3.07 bpw the composition itself drifts, with a narrower band sitting higher in the frame; even there the image is still a clean macro shot. Differences like these are subtle and highly prompt-dependent, and the row above is one prompt out of 24. Judge for yourself: the [blog post](https://byteshape.com/blogs/Qwen-Image-2512/) walks through the full set, and the [interactive comparison explorer](https://byteshape.com/blogs/Qwen-Image-2512/comparison/) lets you A/B any two variants across 24 curated prompts at full resolution, with a slider, zoom, and a curator note for each prompt. If you read the outputs differently, tell us on [Reddit](https://www.reddit.com/r/ByteShape/). ## Repository Layout The package is a shared **shell** (the pipeline) plus a pool of **transformers**, one per size: ``` Qwen-Image-2512/ # the shell: text encoder, VAE, scheduler, tokenizer quantized_models/ 6.77bpw/transformer/ 5.33bpw/transformer/ 4.70bpw/transformer/ 4.18bpw/transformer/ 3.57bpw/transformer/ 3.07bpw/transformer/ ``` You download the shell once, then drop the transformer of your choice into it. ## Available Models | Model ID | Bits/Weight | Transformer Size | Folder | |---|---|---|---| | VLLM-1 | 3.07 | 7.83 GB | `quantized_models/3.07bpw` | | VLLM-2 | 3.57 | 9.11 GB | `quantized_models/3.57bpw` | | VLLM-3 | 4.18 | 10.68 GB | `quantized_models/4.18bpw` | | VLLM-4 | 4.70 | 12.00 GB | `quantized_models/4.70bpw` | | VLLM-5 | 5.33 | 13.60 GB | `quantized_models/5.33bpw` | | VLLM-6 | 6.77 | 17.29 GB | `quantized_models/6.77bpw` | **On VRAM.** The size above is the diffusion transformer only. The shell adds a **BF16 text encoder (\~16.6 GB)** and the VAE (\~0.25 GB), so keeping everything resident on the GPU needs roughly `transformer + 17 GB`: about 25 GB for the smallest size and 34 GB for the largest. - **≤ 32 GB cards (including the RTX 5090):** serve with `--enable-cpu-offload`. All six sizes work; the text encoder lives in system RAM. - **48 GB and larger (e.g. RTX Pro 6000):** drop `--enable-cpu-offload` for full speed. **Selection rule:** take the largest model that comfortably fits your card. Since compression does not speed up diffusion, there is little reason to go smaller than your VRAM allows. ## Quick Start ### 1. Set up the environment ```bash conda create -y -n qwen-image-humming python=3.12 conda activate qwen-image-humming # torch trio (cu129) pip install torch==2.11.0 torchvision==0.26.0 torchaudio==2.11.0 \ --index-url https://download.pytorch.org/whl/cu129 # vLLM 0.24.0 (cu129 wheel) printf 'torch==2.11.0+cu129\ntorchvision==0.26.0+cu129\ntorchaudio==2.11.0+cu129\n' > constraints.txt PIP_CONSTRAINT=constraints.txt pip install \ https://github.com/vllm-project/vllm/releases/download/v0.24.0/vllm-0.24.0+cu129-cp38-abi3-manylinux_2_28_x86_64.whl # vLLM-Omni + our loader plugin PIP_CONSTRAINT=constraints.txt pip install vllm-omni==0.24.0 PIP_CONSTRAINT=constraints.txt pip install vllm-omni-humming==0.3.0 # humming kernels: must be >= 0.1.11 (older versions produce corrupted images) pip install --no-deps "humming-kernels>=0.1.11" ``` `vllm-omni-humming` is our own plugin, and it is the one required extra piece for this path: stock vLLM-Omni cannot read a pre-quantized Humming checkpoint on its own. The plugin registers itself and detects our models from their `transformer/config.json`. ### 2. Download the model Grab the shell and the size you want: ```bash pip install -U huggingface_hub # the shell + one transformer (here the 3.07 bpw size) hf download byteshape/Qwen-Image-2512-Humming \ --include "Qwen-Image-2512/*" --include "quantized_models/3.07bpw/*" \ --local-dir Qwen-Image-2512-Humming cd Qwen-Image-2512-Humming ``` Then drop your chosen transformer into the shell. The folder is already named `transformer`, so it lands with the correct name: ```bash # symlink it in (instant, no extra disk) ... ln -sfn ../quantized_models/3.07bpw/transformer Qwen-Image-2512/transformer # ... or move it in (a real, self-contained model dir) mv quantized_models/3.07bpw/transformer Qwen-Image-2512/transformer ``` To try a different size later, swap the symlink, or move one transformer out and another in. The folders are named by bit-width (`6.77bpw`, `4.18bpw`, `3.07bpw`, and so on). ### 3. Serve From the folder root: ```bash vllm-omni serve Qwen-Image-2512 \ --omni --served-model-name Qwen/Qwen-Image-2512 \ --enable-cpu-offload --port 8124 ``` Keep `--enable-cpu-offload` on cards with ≤ 32 GB (an RTX 5090 needs it); drop it on larger GPUs for full speed. Do **not** pass `--diffusion-quantization-config` or set any `VLLM_HUMMING_*` variables. The plugin drives everything from the checkpoint. The first serve on a new GPU spends a few minutes compiling kernels for your card, then caches them. ### 4. Generate The server exposes the OpenAI images API: ```bash curl -s http://127.0.0.1:8124/v1/images/generations \ -H "Content-Type: application/json" \ -d '{ "model": "Qwen/Qwen-Image-2512", "prompt": "A red panda on a mossy log at dawn", "size": "1024x1024", "num_inference_steps": 20, "true_cfg_scale": 2.5, "negative_prompt": " ", "seed": 1234 }' \ | python3 -c "import sys,json,base64; d=json.load(sys.stdin); open('out.png','wb').write(base64.b64decode(d['data'][0]['b64_json']))" ``` **Always include a `negative_prompt`, even a single space.** vLLM-Omni applies classifier-free guidance only when a negative prompt is present with `true_cfg_scale > 1`; with none, guidance is skipped and the image comes out flatter and lower-contrast. ## Driving the Server from ComfyUI If you prefer a GUI over `curl`, the `ComfyUI-vLLM-Omni` custom node sends generation requests to a running vLLM-Omni server, local or remote. ComfyUI is only the front end; vLLM still does the work. Install the node into your ComfyUI installation: ```bash cd /path/to/ComfyUI/custom_nodes git clone https://github.com/vllm-project/vllm-omni.git cp -r vllm-omni/apps/ComfyUI-vLLM-Omni ComfyUI-vLLM-Omni rm -rf vllm-omni ``` Restart ComfyUI (`python main.py`), then open a browser and go to **http://127.0.0.1:8188** to reach the ComfyUI interface. A new **vLLM-Omni** folder appears in the **Node Library** sidebar. No extra Python packages are needed beyond ComfyUI's own. Keep the vLLM-Omni server from step 3 running. ComfyUI reaches it at `http://127.0.0.1:8124/v1`. If ComfyUI runs on a different machine, use the server's address instead of `127.0.0.1`. For the workflow, download `qwen-image-humming-vllm-omni.json` from our [blog post](https://byteshape.com/blogs/Qwen-Image-2512/) and load it with **`Ctrl+O`**, or start from ComfyUI's built-in template at **sidebar → Templates → vLLM-Omni → vLLM-Omni Image Generation**. In the **Generate Image** node, set: - **url** → `http://127.0.0.1:8124/v1` - **model** → `Qwen/Qwen-Image-2512` - **prompt** / **negative_prompt** → your text (negative: at least a space) - **width** / **height** → e.g. 1024 x 1024 In the attached **Diffusion Sampling Params** node, set the seed, `num_inference_steps` (e.g. 20), and `true_cfg_scale` (e.g. 2.5), then run the workflow. ## Troubleshooting | Symptom | Cause and fix | |---|---| | `HummingConfig ... unexpected keyword argument`, or a shape mismatch on `img_mod.1` / `txt_mod.1` at load | The plugin is not installed in the active environment. Check with `pip show vllm-omni-humming`. | | Grainy or speckled images | The Humming kernels are too old. Reinstall with `pip install --no-deps "humming-kernels>=0.1.11"`. | | Flat, low-contrast images | No `negative_prompt` was sent, so CFG was skipped. Pass one; a single space is enough. | | Out of memory on serve | Add `--enable-cpu-offload`, or pick a smaller bit-width from the table above. |