--- license: creativeml-openrail-m base_model: OnomaAIResearch/Illustrious-XL-v2.0 tags: - mnn - stable-diffusion-xl - text-to-image - illustration - on-device - quantized - android --- # Illustrious-XL-v2.0 (MNN) MNN format conversion of [OnomaAIResearch/Illustrious-XL-v2.0](https://huggingface.co/OnomaAIResearch/Illustrious-XL-v2.0), for on-device image generation in [nezumi-ai](https://github.com/mouse0329/nezumi-ai) — a private, fully offline AI chat app for **Android**. The MNN image-generation engine in this model is one of several on-device inference backends used by the app. A Windows CLI (`nezumi-ai-sd-cli`) is also provided for testing/debugging on desktop, but the primary target platform is Android. > **Platform note**: `nezumi-ai-sd-cli` currently builds for **Windows only** (`.exe`). > A Linux build is planned. The Android app is the main way to use this model. > **Requirements note**: unlike the app's general 6GB RAM minimum, **SDXL / Illustrious models > require 8GB RAM minimum on Android** (verified by testing). This is higher than the LLM-only > requirement due to SDXL's larger UNet and dual text encoders. 8GB+ recommended for comfortable use. > ⚠️ **Important: use a lower CFG scale with this quantized model.** See > [CFG scale — please read](#cfg-scale--please-read) below. The default CLI value of `--cfg 7.0` > is tuned for the original fp16 model and is **too high for the quantized UNet here** — at > CFG 7 you will likely get flat, oversaturated, "washed out" colors or broken line art even > though the model itself loaded and ran correctly. Use `--cfg 4.0` to `--cfg 5.0` instead. ## Variants | File | UNet quantization | Size | Notes | |---|---|---|---| | `Illustrious-XL-v2.0-diffusers-mnn-int4-block32.zip` | 4-bit, block size 32 | 3.01 GB | Smaller / faster, more quality loss — recommended for 8GB devices | | `Illustrious-XL-v2.0-diffusers-mnn-int8-block128.zip` | 8-bit, block size 128 | 4.08 GB | Larger, closer to original quality — for 12GB+ devices | Both share the same CLIP/VAE settings (see below). Both variants need the lower CFG scale described below — this is not specific to one bit-width. ## CFG scale — please read During testing, this quantized UNet produced badly corrupted output (flat solid-color fills, blown-out contrast, broken line art) at the "normal" SDXL default of `cfg=7.0`, even though every individual component (CLIP text encoders, VAE, tokenizer, the UNet weights themselves) checked out fine in isolation. The same quantization settings applied to a smaller/distilled SDXL UNet (SSD-1B) showed no such problem at `cfg=7.0`. The cause: classifier-free guidance amplifies the *difference* between the conditional and unconditional UNet predictions (`pred = uncond + cfg * (cond - uncond)`). Quantizing a large, full-depth SDXL UNet (Illustrious keeps the full ~2.6B-parameter SDXL-Base architecture, unlike distilled variants such as SSD-1B) introduces a small amount of per-step numerical error. That error doesn't cancel out in the `cond - uncond` subtraction, and a high CFG scale re-amplifies it on every step, which is what produced the corrupted output. **Fix: lower the CFG scale.** In our tests: | CFG scale | Result | |---|---| | `7.0` (CLI default) | Broken — flat colors, oversaturated, line art may survive but shading collapses | | `5.0` | Clearly improved, still slightly flatter than reference | | `4.0` | Clean, detailed output (highlights, hair, fabric folds render correctly) — no quantization artifacts visible | We recommend starting at **`--cfg 4.0`** and adjusting up toward `5.0` if you want stronger prompt adherence and can tolerate slightly flatter shading. This applies to both the int4 and int8 variants above. If you need the full `cfg=7.0`-style output fidelity, you would need an unquantized (fp16) UNet export instead, which is much larger (~9-10GB) and not distributed here. This is a property of quantizing this particular (full-size, non-distilled) SDXL UNet, not a bug in the CLIP/VAE conversion or in the MNN runtime — both were independently verified to be correct during debugging. ## Model Provenance | Field | Value | |---|---| | Base model | [OnomaAIResearch/Illustrious-XL-v2.0](https://huggingface.co/OnomaAIResearch/Illustrious-XL-v2.0) | | Format | MNN (`clip1.mnn`, `clip2.mnn`, `unet.mnn`, `vae_decoder_fp16.mnn`, tokenizers) | | Conversion tool | [`convert_hf_to_mnn_sdxl.py`](https://github.com/mouse0329/nezumi-ai/blob/main/mnn-sd-engine/conversion/convert_hf_to_mnn_sdxl.py) (nezumi-ai) | ### Conversion settings ```bash # int4-block32 python convert_hf_to_mnn_sdxl.py --model OnomaAIResearch/Illustrious-XL-v2.0 \ --out ./out/int4-block32 --size 1024 --clip-skip 2 \ --unet-bits 4 --unet-block 32 --clip-bits 8 --vae-bits 8 # int8-block128 python convert_hf_to_mnn_sdxl.py --model OnomaAIResearch/Illustrious-XL-v2.0 \ --out ./out/int8-block128 --size 1024 --clip-skip 2 \ --unet-bits 8 --unet-block 128 --clip-bits 8 --vae-bits 8 ``` No fine-tuning or retraining was performed — weights are unchanged from the original checkpoint aside from format conversion and the quantization above. See [CFG scale — please read](#cfg-scale--please-read) above for the recommended sampling settings to use with this quantized checkpoint. ## License - **Original model license**: [CreativeML Open RAIL-M](https://huggingface.co/spaces/CompVis/stable-diffusion-license) — all credit for the weights and training goes to OnomaAI Research. - **Redistribution**: 可(元モデルのライセンス上、量子化・派生物の再配布は許可されています) - **Commercial use**: RAIL-Mの条件内で可 - **Attribution**: 要(本README内に明記) This checkpoint inherits the original model's use-based restrictions in full (see Attachment A of the [full license text](https://huggingface.co/spaces/CompVis/stable-diffusion-license)), including prohibitions on use for exploiting minors, generating disinformation, harassment, discrimination, unauthorized medical advice, and law-enforcement/immigration profiling. On Android, note that nezumi-ai includes an `ImageSafetyChecker` (NSFW detection with auto-block/blur), which complements — but does not replace — compliance with these restrictions. > Note: the conversion script itself is part of the nezumi-ai project and licensed separately > under LGPL v3 / a commercial license (see [LICENSE.md](https://github.com/mouse0329/nezumi-ai/blob/main/LICENSE.md)). > That license applies to the *code*, not to this model checkpoint. ## Requirements (Android) | Item | Minimum | Recommended | |---|---|---| | Android Version | 12 (API 30) | 14+ (API 34+) | | RAM (SDXL / Illustrious) | **8GB** | 12GB+ | | Storage | 4GB (int4) / 5GB (int8), plus space for other models | 8GB+ | | GPU/NPU | Optional | Snapdragon / Mali / Adreno (OpenCL) | ## Usage ### Android (primary) Used automatically by the [nezumi-ai](https://github.com/mouse0329/nezumi-ai) app's image-generation feature (MNN backend, GPU/OpenCL → CPU fallback). Download/select this model from within the app; manual extraction is not required on Android. **Set CFG scale to 4.0–5.0** in the generation settings (see [CFG scale — please read](#cfg-scale--please-read) above) — the app's own default may still assume the unquantized-model value of 7.0. ### Windows CLI (testing/debugging) Distributed as zip archives — extract before use: ```bash unzip Illustrious-XL-v2.0-diffusers-mnn-int8-block128.zip -d C:\sdxl-model ``` ```bat nezumi-ai-sd-cli "C:\sdxl-model" "1girl, cute, cat ears" --steps 20 --width 1024 --height 1024 --cfg 4.0 --backend cpu --out out.png ``` Note the explicit `--cfg 4.0` above — see [CFG scale — please read](#cfg-scale--please-read). #### Options | Option | Description | Default | |---|---|---| | `` | Path to the extracted MNN model folder | — | | `` | Text prompt | — | | `--negative ` | Negative prompt | empty | | `--width ` / `--height ` | Image size | 512 / 512 | | `--steps ` | Sampling steps | 20 | | `--cfg ` | CFG scale | 7.0 (⚠️ use `4.0`–`5.0` for this model — see above) | | `--seed ` | Seed (negative = random) | -1 | | `--scheduler ` | `euler`\|`ddim`\|`dpm`\|`dpm++2m`\|`dpm++2m-karras`\|`lcm`\|`eulera`\|`unipc` | `dpm++2m` | | `--backend ` | `cpu`\|`opencl` | `cpu` | | `--out ` | `.ppm` always works; `.png` needs `stb_image_write.h` | — | > SDXL/Illustrious系では`--width 1024 --height 1024`を推奨します(デフォルトの512は非対応解像度のため画質が崩れます)。 > また、このモデル(量子化UNet)では`--cfg 7.0`のデフォルト値は強すぎるため、`--cfg 4.0`〜`5.0`を推奨します(詳細は上記「CFG scale — please read」を参照)。 ## Roadmap - [ ] Linux build of `nezumi-ai-sd-cli` ## Disclaimer This is an unofficial, community conversion and is not affiliated with or endorsed by OnomaAI Research.