Nova XL LCM — OpenVINO INT4

將 sca255/nova_xl_lcm(SDXL LCM 動漫模型)轉換為 OpenVINO INT4 權重,並以 純 CPU 實測 1024×1024 8-step 出圖。

optimum-intel 匯出 → NNCF weight-only INT4 量化 → 6.5 GB → 2.0 GB(↓ ~69%), 在 192 vCPU 的 Xeon 上 平均 30.6 s / 張(1024px, 8 steps)。


目錄


快速資訊

項目 內容
基座模型 sca255/nova_xl_lcm(SDXL + LCM 蒸餾,動漫向)
Pipeline StableDiffusionXLPipeline(diffusers 格式)
Scheduler LCMScheduler
量化格式 UNet / text_encoder / text_encoder_2 → INT4;VAE → INT8
模型大小 FP16 6.5 GB → INT4 2.0 GB(↓ ~69%)
推論裝置 CPU(OpenVINO CPU plugin,無需 GPU)
推薦參數 num_inference_steps=8、guidance_scale=1.5、1024×1024
轉換工具 optimum 2.3.0 / optimum-intel 2.2.0 / OpenVINO 2026.4.0 / NNCF 3.4.0

特色

  • **體積 ↓ 69%**:FP16 6.5 GB → INT4 2.0 GB,CPU 記憶體需求大幅下降。
  • 真·8 step:LCM 蒸餾模型,guidance_scale=1.5 即可出圖,無需 CFG 拉高。
  • 純 CPU 可用:不依賴 GPU / CUDA,OpenVINO 自動使用全部核心。
  • diffusers 目錄結構:可直接 from_pretrained(),與原模型使用方式一致。
  • 附 20 張實測圖:10 組 prompt × (1024px + 512px 對照組),seed 固定可重現。
  • 附完整 benchmark:逐 step 耗時、系統資訊、量化設定全部記錄於 REPORT.md 與 examples/benchmark.json。

安裝

pip install diffusers==0.37.1 transformers==4.57.6 tokenizers==0.22.0 \
            huggingface-hub==0.35.1 optimum==2.3.0 optimum-intel==2.2.0 \
            openvino==2026.4.0 nncf==3.4.0 torch pillow psutil

僅推論時其實不需要 nncf;nncf 為重新量化所需。


快速開始

import torch
from optimum.intel import OVDiffusionPipeline

pipe = OVDiffusionPipeline.from_pretrained(
    "HelloSun/nova_xl_lcm-OpenVINO-INT4",
    compile=True,
)

image = pipe(
    prompt="1girl, anime style, masterpiece, best quality, detailed face",
    height=1024,
    width=1024,
    num_inference_steps=8,
    guidance_scale=1.5,
    generator=torch.Generator().manual_seed(42),
).images[0]

image.save("out.png")

完整可執行範例:inference_int4.py 批次生成 10 組 + benchmark:generate5.py


推論參數建議

參數 建議值 說明
num_inference_steps 8 LCM 蒸餾步數。請勿增加,8 步已足夠且是此模型的設計值。
guidance_scale 1.5 LCM 模型的標準 CFG 值,調高反而容易劣化。
height / width 1024 SDXL 原生解析度。
compile=True 開啟 編譯模型以取得較佳效能(代價為編譯時間)。

範例結果

全部為 8 steps / guidance 1.5 / CPU,seed 固定,可完全重現。 每組另附 512px 對照組,點擊圖片可看原始尺寸。

01 · hanfu — seed 42 — 32.34 s

1girl, red hanfu, intricate embroidery, golden phoenix headdress, soft night lighting,
tiered pagoda background, lantern lights, anime style, masterpiece, best quality, ultra detailed

01 hanfu

02 · astronaut — seed 43 — 28.75 s

1girl astronaut in lush jungle, cold color palette, detailed foliage, cinematic lighting,
anime style, masterpiece, best quality, 8k

02 astronaut

03 · taipei — seed 44 — 34.20 s

1girl in cyberpunk Taipei at night, heavy rain, neon signs 'TAIPEI' and '台北',
wet asphalt reflections, crowded night market, anime style, masterpiece, best quality

03 taipei

04 · shiba — seed 45 — 31.90 s

cute shiba inu wearing tiny astronaut helmet, sunflower field under starry sky,
dreamy illustration, vibrant colors, anime style, masterpiece, best quality

04 shiba

05 · ink — seed 46 — 29.66 s

traditional chinese ink wash landscape, misty mountains, small pagoda on cliff,
cranes flying, minimalist, elegant, anime style, masterpiece, best quality

05 ink

06 · ghibli style — seed 47 — 29.74 s

studio ghibli style, 1girl in floating castle, magical clouds, whimsical atmosphere,
detailed background, masterpiece, best quality

06 ghibli style

07 · makoto shinkai — seed 48 — 29.97 s

makoto shinkai style, 1girl under comet sky, breathtaking lighting, detailed clouds,
emotional atmosphere, anime style, masterpiece, 8k

07 makoto shinkai

08 · flat vector — seed 49 — 30.13 s

flat vector illustration, 1girl in modern city, clean geometric shapes, bold colors,
minimalist, commercial anime art style, behance trending

08 flat vector

09 · retro 80s — seed 50 — 29.89 s

1980s retro anime, 1girl with neon city background, VHS aesthetic, synthwave colors,
cel shaded, nostalgic anime style, masterpiece

09 retro 80s

10 · impasto thick — seed 51 — 29.52 s

thick impasto anime concept art, 1girl with heavy brushstrokes, expressive texture,
dramatic lighting, artstation masterpiece, detailed anime style

10 impasto thick


效能實測摘要

完整逐 step 數據見 REPORT.md 與 examples/benchmark.json。

測試環境

項目 內容
CPU Intel Xeon Platinum 8559C,2 socket × 48 core × 2 thread = 192 vCPU(KVM)
RAM 2.0 TiB
執行環境 OpenVINO CPU only,openvino 2026.4.0
設定 num_inference_steps=8、guidance_scale=1.5

結果

指標 1024×1024 512×512
平均總耗時 30.61 s / 張 9.75 s / 張
平均單步耗時 3.60 s 1.15 s
模型載入 + 編譯 12.45 s(僅一次) —

1024px 約為 512px 的 3.1× 時間。 首張圖的第一個 step 明顯較慢(01 為 6.02 s,後續皆約 3.4–3.5 s), 主因是 text_encoder / text_encoder_2 編碼與 UNet warmup。 文字編碼與 VAE decode 的時間已包含在總耗時內。

逐張耗時(1024px / 8 steps)

# Prompt Seed 總耗時 平均單步
01 hanfu 42 32.34 s 3.80 s
02 astronaut 43 28.75 s 3.39 s
03 taipei 44 34.20 s 4.01 s
04 shiba 45 31.90 s 3.74 s
05 ink 46 29.66 s 3.48 s
06 ghibli_style 47 29.74 s 3.49 s
07 makoto_shinkai 48 29.97 s 3.52 s
08 flat_vector 49 30.13 s 3.54 s
09 retro_80s 50 29.89 s 3.51 s
10 impasto_thick 51 29.52 s 3.47 s
平均 30.61 s 3.60 s

模型大小

元件 FP16 INT4 降幅
unet 4.8 GB 1.5 GB ~69%
text_encoder_2 1.3 GB 377 MB ~71%
text_encoder 236 MB 81 MB ~66%
vae_decoder 95 MB 48 MB ~49%
vae_encoder 66 MB 33 MB ~50%
合計 6.5 GB 2.0 GB ~69%

檔案結構

.
├── README.md                  # 本文件
├── REPORT.md                  # 完整轉換 + 實測報告
├── model_index.json           # diffusers pipeline 索引
├── openvino_config.json       # OpenVINO 執行期設定
├── inference_int4.py          # 單張推論範例
├── generate5.py               # 10 組 prompt 批次生成 + benchmark
├── quantize_int4.py           # FP16 OV → INT4 OV 量化腳本
├── unet/                      # INT4 UNet
├── vae_encoder/               # INT8 VAE encoder
├── vae_decoder/               # INT8 VAE decoder
├── text_encoder/              # INT4 CLIP text encoder
├── text_encoder_2/            # INT4 CLIP text encoder (projection)
├── tokenizer/                 # CLIP tokenizer
├── tokenizer_2/               # CLIP tokenizer 2
├── scheduler/                 # LCMScheduler 設定
└── examples/                  # 20 張實測圖 + benchmark.json + prompts.txt
    ├── *_1024.png             # 主測組
    ├── *_512.png              # 對照組
    ├── benchmark.json         # 逐 step 耗時 + 系統資訊
    └── prompts.txt            # 10 組 prompt 與 seed

從零復現

# 1. 匯出 FP16 OpenVINO 模型
optimum-cli export openvino \
  -m sca255/nova_xl_lcm \
  --task text-to-image \
  --library diffusers \
  --weight-format fp16 \
  ./nova_xl_lcm-ov-fp16

# 2. NNCF weight-only INT4 量化(unet + 兩個 text encoder),其餘預設 INT8
python quantize_int4.py --fp16-dir ./nova_xl_lcm-ov-fp16 --int4-dir ./nova_xl_lcm-ov-int4

# 3. 單張推論
python inference_int4.py

# 4. 批次生成 10 組 + 512px 對照 + benchmark
python generate5.py

量化設定(quantize_int4.py):

int4_config = OVWeightQuantizationConfig(
    bits=4, sym=False, group_size=128, group_size_fallback="adjust", ratio=1.0,
)
pipeline_config = OVPipelineQuantizationConfig(
    quantization_configs={"unet": int4_config,
                          "text_encoder": int4_config,
                          "text_encoder_2": int4_config},
    default_config=OVWeightQuantizationConfig(bits=8, sym=True),
)

已知限制

  • 僅為 weight-only 量化:啟動時仍需即時量化權重,首次載入 + 編譯約 12 s。
  • CPU 專用:本 repo 為 OpenVINO IR 格式,若要使用 GPU 請改用原模型 sca255/nova_xl_lcm 或自行轉換 OpenVINO GPU / NNCF。
  • LCM 不宜增加步數:設定 num_inference_steps > 8 不會變好,反而可能劣化。
  • INT4 為非對稱量化(sym=False),品質損失在本模型上實測不明顯,但與 FP16 原圖仍可能有細微差異。
  • 實測僅涵蓋 192 vCPU 伺服器:一般桌上型 CPU 的絕對數字會較慢,比例關係大致相同。

授權與出處

  • 來源模型:sca255/nova_xl_lcm
  • 授權:openrail++(沿用來源模型授權)
  • 轉換:僅做格式轉換與權重量化,模型權重來自來源模型

使用本模型時請一併遵守來源模型的授權條款與 OpenRAIL++ 使用政策。


Made with OpenVINO + optimum-intel + NNCF

Downloads last month
16
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for HelloSun/nova_xl_lcm-OpenVINO-INT4

Finetuned
(1)
this model