Text-to-Image
Diffusers
OpenVINO
ZImagePipeline
int4
optimum
optimum-intel
nncf
z-image
transformer
cpu
Instructions to use HelloSun/Z-Image-Turbo-OpenVINO-INT4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use HelloSun/Z-Image-Turbo-OpenVINO-INT4 with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("HelloSun/Z-Image-Turbo-OpenVINO-INT4", dtype=torch.bfloat16, device_map="cuda") prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k" image = pipe(prompt).images[0] - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- Draw Things
- DiffusionBee
Download quantize_int4.py from HelloSun/Z-Image-Turbo-OpenVINO-INT4: direct link, hf CLI and curl.
- Browser
- Download file 1.86 kB
-
https://huggingface.co/HelloSun/Z-Image-Turbo-OpenVINO-INT4/resolve/main/quantize_int4.py
- Command line
-
hf download hf://HelloSun/Z-Image-Turbo-OpenVINO-INT4/quantize_int4.py
-
curl -L -o quantize_int4.py https://huggingface.co/HelloSun/Z-Image-Turbo-OpenVINO-INT4/resolve/main/quantize_int4.py
1.86 kB
| """Quantize Z-Image-Turbo OV FP16 -> INT4 using optimum (NNCF weight-only).""" | |
| from pathlib import Path | |
| from optimum.intel import OVDiffusionPipeline, OVQuantizer | |
| from optimum.intel.openvino.configuration import ( | |
| OVConfig, | |
| OVWeightQuantizationConfig, | |
| OVPipelineQuantizationConfig, | |
| ) | |
| fp16_dir = "./z-image-turbo-ov-fp16" | |
| int4_dir = "./z-image-turbo-ov-int4" | |
| # INT4 config for transformer + text_encoder (main weights) | |
| # group_size=128, sym=False is NNCF default sweet spot for quality/size | |
| int4_config = OVWeightQuantizationConfig( | |
| bits=4, | |
| sym=False, | |
| group_size=128, | |
| group_size_fallback="adjust", | |
| ratio=1.0, | |
| ) | |
| # Keep VAE in FP16 (don't quantize) by using default FP32? Actually set bits=8 for others, | |
| # but we will explicitly only quantize transformer+text_encoder and copy rest as-is. | |
| # Using OVPipelineQuantizationConfig: | |
| pipeline_config = OVPipelineQuantizationConfig( | |
| quantization_configs={ | |
| "transformer": int4_config, | |
| "text_encoder": int4_config, | |
| }, | |
| # default: 8-bit weight-only for any other submodel that gets quantized, | |
| # but we skip vae by not including it? OVQuantizer quantizes only listed? | |
| # To be safe, set default 8-bit; vae will be copied if not quantized? | |
| default_config=OVWeightQuantizationConfig(bits=8, sym=True), | |
| ) | |
| print(f"Loading FP16 OV pipeline from {fp16_dir} ...") | |
| pipe = OVDiffusionPipeline.from_pretrained(fp16_dir) | |
| print(f"Components: {list(pipe.components.keys())}") | |
| print(f"OV submodels: {list(pipe.ov_submodels) if hasattr(pipe, 'ov_submodels') else 'n/a'}") | |
| quantizer = OVQuantizer.from_pretrained(pipe) | |
| print("Quantizing transformer+text_encoder to INT4 (NNCF, optimum)...") | |
| ov_config = OVConfig(quantization_config=pipeline_config) | |
| quantizer.quantize( | |
| save_directory=int4_dir, | |
| ov_config=ov_config, | |
| ) | |
| print(f"Saved INT4 model to {int4_dir}") | |