Iris-3B
Pixel-space generation & general vision learner.
A 3-billion-parameter model that paints every pixel directly β no VAE, no latent space β and, fine-tuned, estimates depth and restores and upscales images.
Generative priors are a promising foundation for downstream vision tasks. In this project we explore pixel-space generative models as an alternative to vision foundation models such as DINOv2.
Most image generators work in a compressed "latent" space and rely on a separate decoder to turn that into an image. Iris-3B skips that step: the network itself outputs every pixel, so nothing is lost to a lossy, texture-biased latent. You type a prompt, and you get a 1024-pixel image.
Iris-3B is not only an image generator. The same pixel-space prior, fine-tuned with no architectural change, is a general vision learner for dense tasks where detail matters: monocular depth estimation and image restoration and upscaling ship in this repo (see A general vision learner below).
Examples
Generated at native aspect ratios of about one megapixel with the default settings (CFG 3, 100 steps).
![]() Studio portrait of a Maasai elder wearing vibrant beaded jewelry, deep red cloth, dark backdrop, Rembrandt lighting, ultra detailed skin texture |
![]() Close-up portrait of a young woman with silver glitter freckles and iridescent makeup, soft pastel background, high fashion beauty photography |
![]() A glass sculpture of a heart filled with flowers, caustics and reflections, 3D render |
![]() Black and white portrait of a fisherman with a thick grey beard and deep wrinkles, piercing eyes, overcast light, fine grain film photograph |
![]() A Byzantine-style mosaic of a peacock made of tiny gold and turquoise tiles, shimmering texture |
![]() A bronze sculpture of a horse in motion, patina, dramatic museum spotlight |
![]() The ancient city of Petra with the Treasury carved into pink sandstone, morning light |
![]() A tree with lightbulbs instead of fruits glowing at dusk, surreal concept art |
![]() Thousands of sky lanterns rising into the night sky at Yi Peng festival in Chiang Mai |
![]() The aurora borealis swirling green and violet over a snowy Lofoten fishing village with red wooden cabins, reflections in a calm fjord, night photograph |
![]() Volcanic eruption at night in Iceland, rivers of glowing lava flowing across black fields, plumes of steam lit orange, long exposure |
Get started
1. Install the code
git clone https://github.com/speridlabs/iris-3b.git
cd iris-3b
pip install -e . # Python 3.11+, PyTorch 2.7.1+
2. Download the weights (about 12 GB for text-to-image; the depth/ and upscaler/ folders add about 12 GB each)
hf download speridlabs/iris-3b --local-dir iris-3b --exclude "depth/*" "upscaler/*"
3. Generate an image
python scripts/sample.py --checkpoint iris-3b \
--prompt "a red fox sleeping in fresh snow, golden hour"
The text encoder (Qwen3-VL-4B-Instruct) downloads automatically on the first run. You need an NVIDIA GPU with CUDA.
Tips
| Setting | Default | What it does |
|---|---|---|
--cfg-scale |
3 | How strictly the image follows the prompt. Higher values follow the prompt more literally but can look harsher. |
--steps |
100 | Number of denoising steps. Fewer steps are faster but lose some detail. |
--seed |
β | Fix it to get the same image again. |
--negative-prompt |
β | Things you don't want in the image. |
--txt-file |
β | A text file with one prompt per line, for batches. |
Default output is 1024Γ1024. Write prompts as plain descriptive English sentences.
A general vision learner: depth, image restoration and upscaling
The same backbone, fine-tuned per task: Iris-3B estimates monocular depth and restores and upscales images. Both run the whole model in a single forward pass with the empty prompt, so no text encoder is needed. The weights download on first use.
Monocular depth β photo β relative depth:
![]() Photo |
![]() Iris-3B depth |
![]() Photo |
![]() Iris-3B depth |
Image restoration and upscaling β degraded input β restored image:
![]() Input |
![]() Iris-3B |
![]() Input |
![]() Iris-3B |
Try both interactively in the demo, or from the command line:
python scripts/depth.py photo.jpg --out depth_out # <name>.npy + colorized <name>.png
python scripts/upscale.py photo.jpg --out upscaled # <name>_x4.png
| Task | Training | Result |
|---|---|---|
| Depth | Direct regression of relative log depth, 10K steps | AbsRel 0.071, Ξ΄1 0.946 (mean of NYUv2, KITTI, ETH3D, ScanNet, DIODE; Marigold V2 protocol, zero-shot) |
| Image restoration and upscaling | One-step adversarial restoration (HYPIR recipe), 10K steps, EMA weights | LPIPS 0.292, PSNR 20.68 (DIV2K validation, 4Γ) |
Depth is affine-invariant, not metric: fit a scale and shift in log space to compare with ground truth. Restoration tiles large outputs into overlapping 1024Γ1024 windows.
What's in this repo
| File | Contents |
|---|---|
model.safetensors |
Text-to-image weights (EMA, float32; run with bfloat16 autocast) |
config.yaml |
Architecture and sampling settings read by scripts/sample.py |
depth/ |
Depth model: model.safetensors, empty_prompt.safetensors, config.yaml |
upscaler/ |
Image restoration and upscaling model: same three files |
About the model
| Size | 3B parameters (plus a frozen 4B text encoder) |
| Architecture | Diffusion transformer: 8 dual-stream + 16 single-stream blocks, then a 4-block pixel head that turns each 16Γ16 patch back into pixels |
| Text encoder | Qwen3-VL-4B-Instruct (frozen) |
| Training | Rectified flow, trained from scratch at 256 β 512 β 1024 px, then supervised fine-tuning at 1024 px (665K steps total) |
The full training recipe is in the code repository.
Limitations
- Like every image generator, Iris-3B can produce inaccurate, biased or unsafe content. Review outputs before using them.
- Text rendering inside images and exact object counts are not always reliable.
- Images are generated at about one megapixel; for larger prints, use an upscaler.
License
Released under the Apache License 2.0. The text encoder, Qwen3-VL-4B-Instruct, is downloaded from its publisher and stays under its own license (Apache 2.0).
Citation
@article{licai2026iris,
title = {Iris-3B: Going Beyond the Latent with Pixel-Space
Diffusion Training, Conversion and Fine-Tuning},
author = {Li Cai, Hanqiu and Garabito, Chema},
journal = {arXiv preprint arXiv:2610.09450},
year = {2026},
url = {https://arxiv.org/abs/2610.09450}
}
Made with β€οΈ by speridlabs.com
- Downloads last month
- 18


















