| --- |
| license: creativeml-openrail-m |
| base_model: OnomaAIResearch/Illustrious-XL-v2.0 |
| tags: |
| - mnn |
| - stable-diffusion-xl |
| - text-to-image |
| - illustration |
| - on-device |
| - quantized |
| - android |
| --- |
| |
| # Illustrious-XL-v2.0 (MNN) |
|
|
| MNN format conversion of [OnomaAIResearch/Illustrious-XL-v2.0](https://huggingface.co/OnomaAIResearch/Illustrious-XL-v2.0), |
| for on-device image generation in [nezumi-ai](https://github.com/mouse0329/nezumi-ai) — a private, fully offline |
| AI chat app for **Android**. The MNN image-generation engine in this model is one of several |
| on-device inference backends used by the app. |
|
|
| A Windows CLI (`nezumi-ai-sd-cli`) is also provided for testing/debugging on desktop, but the |
| primary target platform is Android. |
|
|
| > **Platform note**: `nezumi-ai-sd-cli` currently builds for **Windows only** (`.exe`). |
| > A Linux build is planned. The Android app is the main way to use this model. |
|
|
| > **Requirements note**: unlike the app's general 6GB RAM minimum, **SDXL / Illustrious models |
| > require 8GB RAM minimum on Android** (verified by testing). This is higher than the LLM-only |
| > requirement due to SDXL's larger UNet and dual text encoders. 8GB+ recommended for comfortable use. |
|
|
| > ⚠️ **Important: use a lower CFG scale with this quantized model.** See |
| > [CFG scale — please read](#cfg-scale--please-read) below. The default CLI value of `--cfg 7.0` |
| > is tuned for the original fp16 model and is **too high for the quantized UNet here** — at |
| > CFG 7 you will likely get flat, oversaturated, "washed out" colors or broken line art even |
| > though the model itself loaded and ran correctly. Use `--cfg 4.0` to `--cfg 5.0` instead. |
|
|
| ## Variants |
|
|
| | File | UNet quantization | Size | Notes | |
| |---|---|---|---| |
| | `Illustrious-XL-v2.0-diffusers-mnn-int4-block32.zip` | 4-bit, block size 32 | 3.01 GB | Smaller / faster, more quality loss — recommended for 8GB devices | |
| | `Illustrious-XL-v2.0-diffusers-mnn-int8-block128.zip` | 8-bit, block size 128 | 4.08 GB | Larger, closer to original quality — for 12GB+ devices | |
|
|
| Both share the same CLIP/VAE settings (see below). Both variants need the lower CFG scale |
| described below — this is not specific to one bit-width. |
|
|
| ## CFG scale — please read |
|
|
| During testing, this quantized UNet produced badly corrupted output (flat solid-color |
| fills, blown-out contrast, broken line art) at the "normal" SDXL default of `cfg=7.0`, |
| even though every individual component (CLIP text encoders, VAE, tokenizer, the UNet |
| weights themselves) checked out fine in isolation. The same quantization settings applied |
| to a smaller/distilled SDXL UNet (SSD-1B) showed no such problem at `cfg=7.0`. |
|
|
| The cause: classifier-free guidance amplifies the *difference* between the conditional and |
| unconditional UNet predictions (`pred = uncond + cfg * (cond - uncond)`). Quantizing a large, |
| full-depth SDXL UNet (Illustrious keeps the full ~2.6B-parameter SDXL-Base architecture, |
| unlike distilled variants such as SSD-1B) introduces a small amount of per-step numerical |
| error. That error doesn't cancel out in the `cond - uncond` subtraction, and a high CFG |
| scale re-amplifies it on every step, which is what produced the corrupted output. |
|
|
| **Fix: lower the CFG scale.** In our tests: |
|
|
| | CFG scale | Result | |
| |---|---| |
| | `7.0` (CLI default) | Broken — flat colors, oversaturated, line art may survive but shading collapses | |
| | `5.0` | Clearly improved, still slightly flatter than reference | |
| | `4.0` | Clean, detailed output (highlights, hair, fabric folds render correctly) — no quantization artifacts visible | |
|
|
| We recommend starting at **`--cfg 4.0`** and adjusting up toward `5.0` if you want stronger |
| prompt adherence and can tolerate slightly flatter shading. This applies to both the int4 and |
| int8 variants above. If you need the full `cfg=7.0`-style output fidelity, you would need an |
| unquantized (fp16) UNet export instead, which is much larger (~9-10GB) and not distributed here. |
|
|
| This is a property of quantizing this particular (full-size, non-distilled) SDXL UNet, not a |
| bug in the CLIP/VAE conversion or in the MNN runtime — both were independently verified to be |
| correct during debugging. |
|
|
| ## Model Provenance |
|
|
| | Field | Value | |
| |---|---| |
| | Base model | [OnomaAIResearch/Illustrious-XL-v2.0](https://huggingface.co/OnomaAIResearch/Illustrious-XL-v2.0) | |
| | Format | MNN (`clip1.mnn`, `clip2.mnn`, `unet.mnn`, `vae_decoder_fp16.mnn`, tokenizers) | |
| | Conversion tool | [`convert_hf_to_mnn_sdxl.py`](https://github.com/mouse0329/nezumi-ai/blob/main/mnn-sd-engine/conversion/convert_hf_to_mnn_sdxl.py) (nezumi-ai) | |
|
|
| ### Conversion settings |
|
|
| ```bash |
| # int4-block32 |
| python convert_hf_to_mnn_sdxl.py --model OnomaAIResearch/Illustrious-XL-v2.0 \ |
| --out ./out/int4-block32 --size 1024 --clip-skip 2 \ |
| --unet-bits 4 --unet-block 32 --clip-bits 8 --vae-bits 8 |
| |
| # int8-block128 |
| python convert_hf_to_mnn_sdxl.py --model OnomaAIResearch/Illustrious-XL-v2.0 \ |
| --out ./out/int8-block128 --size 1024 --clip-skip 2 \ |
| --unet-bits 8 --unet-block 128 --clip-bits 8 --vae-bits 8 |
| ``` |
|
|
| No fine-tuning or retraining was performed — weights are unchanged from the original |
| checkpoint aside from format conversion and the quantization above. See |
| [CFG scale — please read](#cfg-scale--please-read) above for the recommended sampling |
| settings to use with this quantized checkpoint. |
|
|
| ## License |
|
|
| - **Original model license**: [CreativeML Open RAIL-M](https://huggingface.co/spaces/CompVis/stable-diffusion-license) |
| — all credit for the weights and training goes to OnomaAI Research. |
| - **Redistribution**: 可(元モデルのライセンス上、量子化・派生物の再配布は許可されています) |
| - **Commercial use**: RAIL-Mの条件内で可 |
| - **Attribution**: 要(本README内に明記) |
|
|
| This checkpoint inherits the original model's use-based restrictions in full |
| (see Attachment A of the [full license text](https://huggingface.co/spaces/CompVis/stable-diffusion-license)), |
| including prohibitions on use for exploiting minors, generating disinformation, harassment, |
| discrimination, unauthorized medical advice, and law-enforcement/immigration profiling. |
| On Android, note that nezumi-ai includes an `ImageSafetyChecker` (NSFW detection with |
| auto-block/blur), which complements — but does not replace — compliance with these restrictions. |
|
|
| > Note: the conversion script itself is part of the nezumi-ai project and licensed separately |
| > under LGPL v3 / a commercial license (see [LICENSE.md](https://github.com/mouse0329/nezumi-ai/blob/main/LICENSE.md)). |
| > That license applies to the *code*, not to this model checkpoint. |
|
|
| ## Requirements (Android) |
|
|
| | Item | Minimum | Recommended | |
| |---|---|---| |
| | Android Version | 12 (API 30) | 14+ (API 34+) | |
| | RAM (SDXL / Illustrious) | **8GB** | 12GB+ | |
| | Storage | 4GB (int4) / 5GB (int8), plus space for other models | 8GB+ | |
| | GPU/NPU | Optional | Snapdragon / Mali / Adreno (OpenCL) | |
|
|
| ## Usage |
|
|
| ### Android (primary) |
|
|
| Used automatically by the [nezumi-ai](https://github.com/mouse0329/nezumi-ai) app's image-generation |
| feature (MNN backend, GPU/OpenCL → CPU fallback). Download/select this model from within the app; |
| manual extraction is not required on Android. **Set CFG scale to 4.0–5.0** in the generation |
| settings (see [CFG scale — please read](#cfg-scale--please-read) above) — the app's own default |
| may still assume the unquantized-model value of 7.0. |
|
|
| ### Windows CLI (testing/debugging) |
|
|
| Distributed as zip archives — extract before use: |
|
|
| ```bash |
| unzip Illustrious-XL-v2.0-diffusers-mnn-int8-block128.zip -d C:\sdxl-model |
| ``` |
|
|
| ```bat |
| nezumi-ai-sd-cli "C:\sdxl-model" "1girl, cute, cat ears" --steps 20 --width 1024 --height 1024 --cfg 4.0 --backend cpu --out out.png |
| ``` |
|
|
| Note the explicit `--cfg 4.0` above — see [CFG scale — please read](#cfg-scale--please-read). |
|
|
| #### Options |
|
|
| | Option | Description | Default | |
| |---|---|---| |
| | `<model_path>` | Path to the extracted MNN model folder | — | |
| | `<prompt>` | Text prompt | — | |
| | `--negative <text>` | Negative prompt | empty | |
| | `--width <n>` / `--height <n>` | Image size | 512 / 512 | |
| | `--steps <n>` | Sampling steps | 20 | |
| | `--cfg <f>` | CFG scale | 7.0 (⚠️ use `4.0`–`5.0` for this model — see above) | |
| | `--seed <n>` | Seed (negative = random) | -1 | |
| | `--scheduler <name>` | `euler`\|`ddim`\|`dpm`\|`dpm++2m`\|`dpm++2m-karras`\|`lcm`\|`eulera`\|`unipc` | `dpm++2m` | |
| | `--backend <name>` | `cpu`\|`opencl` | `cpu` | |
| | `--out <path>` | `.ppm` always works; `.png` needs `stb_image_write.h` | — | |
|
|
| > SDXL/Illustrious系では`--width 1024 --height 1024`を推奨します(デフォルトの512は非対応解像度のため画質が崩れます)。 |
| > また、このモデル(量子化UNet)では`--cfg 7.0`のデフォルト値は強すぎるため、`--cfg 4.0`〜`5.0`を推奨します(詳細は上記「CFG scale — please read」を参照)。 |
|
|
| ## Roadmap |
|
|
| - [ ] Linux build of `nezumi-ai-sd-cli` |
|
|
| ## Disclaimer |
|
|
| This is an unofficial, community conversion and is not affiliated with or endorsed |
| by OnomaAI Research. |