Iris-3B

Pixel-space generation & general vision learner.
A 3-billion-parameter model that paints every pixel directly β€” no VAE, no latent space β€” and, fine-tuned, estimates depth and restores and upscales images.

Paper Project Page Code Demo License

Generative priors are a promising foundation for downstream vision tasks. In this project we explore pixel-space generative models as an alternative to vision foundation models such as DINOv2.

Most image generators work in a compressed "latent" space and rely on a separate decoder to turn that into an image. Iris-3B skips that step: the network itself outputs every pixel, so nothing is lost to a lossy, texture-biased latent. You type a prompt, and you get a 1024-pixel image.

Iris-3B is not only an image generator. The same pixel-space prior, fine-tuned with no architectural change, is a general vision learner for dense tasks where detail matters: monocular depth estimation and image restoration and upscaling ship in this repo (see A general vision learner below).

Examples

Generated at native aspect ratios of about one megapixel with the default settings (CFG 3, 100 steps).


Studio portrait of a Maasai elder wearing vibrant beaded jewelry, deep red cloth, dark backdrop, Rembrandt lighting, ultra detailed skin texture

Close-up portrait of a young woman with silver glitter freckles and iridescent makeup, soft pastel background, high fashion beauty photography

A glass sculpture of a heart filled with flowers, caustics and reflections, 3D render

Black and white portrait of a fisherman with a thick grey beard and deep wrinkles, piercing eyes, overcast light, fine grain film photograph

A Byzantine-style mosaic of a peacock made of tiny gold and turquoise tiles, shimmering texture

A bronze sculpture of a horse in motion, patina, dramatic museum spotlight

The ancient city of Petra with the Treasury carved into pink sandstone, morning light

A tree with lightbulbs instead of fruits glowing at dusk, surreal concept art

Thousands of sky lanterns rising into the night sky at Yi Peng festival in Chiang Mai

The aurora borealis swirling green and violet over a snowy Lofoten fishing village with red wooden cabins, reflections in a calm fjord, night photograph

Volcanic eruption at night in Iceland, rivers of glowing lava flowing across black fields, plumes of steam lit orange, long exposure

Get started

1. Install the code

git clone https://github.com/speridlabs/iris-3b.git
cd iris-3b
pip install -e .          # Python 3.11+, PyTorch 2.7.1+

2. Download the weights (about 12 GB for text-to-image; the depth/ and upscaler/ folders add about 12 GB each)

hf download speridlabs/iris-3b --local-dir iris-3b --exclude "depth/*" "upscaler/*"

3. Generate an image

python scripts/sample.py --checkpoint iris-3b \
    --prompt "a red fox sleeping in fresh snow, golden hour"

The text encoder (Qwen3-VL-4B-Instruct) downloads automatically on the first run. You need an NVIDIA GPU with CUDA.

Tips

Setting Default What it does
--cfg-scale 3 How strictly the image follows the prompt. Higher values follow the prompt more literally but can look harsher.
--steps 100 Number of denoising steps. Fewer steps are faster but lose some detail.
--seed β€” Fix it to get the same image again.
--negative-prompt β€” Things you don't want in the image.
--txt-file β€” A text file with one prompt per line, for batches.

Default output is 1024Γ—1024. Write prompts as plain descriptive English sentences.

A general vision learner: depth, image restoration and upscaling

The same backbone, fine-tuned per task: Iris-3B estimates monocular depth and restores and upscales images. Both run the whole model in a single forward pass with the empty prompt, so no text encoder is needed. The weights download on first use.

Monocular depth β€” photo β†’ relative depth:


Photo

Iris-3B depth

Photo

Iris-3B depth

Image restoration and upscaling β€” degraded input β†’ restored image:


Input

Iris-3B

Input

Iris-3B

Try both interactively in the demo, or from the command line:

python scripts/depth.py photo.jpg --out depth_out      # <name>.npy + colorized <name>.png
python scripts/upscale.py photo.jpg --out upscaled     # <name>_x4.png
Task Training Result
Depth Direct regression of relative log depth, 10K steps AbsRel 0.071, Ξ΄1 0.946 (mean of NYUv2, KITTI, ETH3D, ScanNet, DIODE; Marigold V2 protocol, zero-shot)
Image restoration and upscaling One-step adversarial restoration (HYPIR recipe), 10K steps, EMA weights LPIPS 0.292, PSNR 20.68 (DIV2K validation, 4Γ—)

Depth is affine-invariant, not metric: fit a scale and shift in log space to compare with ground truth. Restoration tiles large outputs into overlapping 1024Γ—1024 windows.

What's in this repo

File Contents
model.safetensors Text-to-image weights (EMA, float32; run with bfloat16 autocast)
config.yaml Architecture and sampling settings read by scripts/sample.py
depth/ Depth model: model.safetensors, empty_prompt.safetensors, config.yaml
upscaler/ Image restoration and upscaling model: same three files

About the model

Size 3B parameters (plus a frozen 4B text encoder)
Architecture Diffusion transformer: 8 dual-stream + 16 single-stream blocks, then a 4-block pixel head that turns each 16Γ—16 patch back into pixels
Text encoder Qwen3-VL-4B-Instruct (frozen)
Training Rectified flow, trained from scratch at 256 β†’ 512 β†’ 1024 px, then supervised fine-tuning at 1024 px (665K steps total)

The full training recipe is in the code repository.

Limitations

  • Like every image generator, Iris-3B can produce inaccurate, biased or unsafe content. Review outputs before using them.
  • Text rendering inside images and exact object counts are not always reliable.
  • Images are generated at about one megapixel; for larger prints, use an upscaler.

License

Released under the Apache License 2.0. The text encoder, Qwen3-VL-4B-Instruct, is downloaded from its publisher and stays under its own license (Apache 2.0).

Citation

@article{licai2026iris,
  title   = {Iris-3B: Going Beyond the Latent with Pixel-Space
             Diffusion Training, Conversion and Fine-Tuning},
  author  = {Li Cai, Hanqiu and Garabito, Chema},
  journal = {arXiv preprint arXiv:2610.09450},
  year    = {2026},
  url     = {https://arxiv.org/abs/2610.09450}
}

Made with ❀️ by speridlabs.com

Downloads last month
18
Safetensors
Model size
3B params
Tensor type
F32
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Space using speridlabs/iris-3b 1

Paper for speridlabs/iris-3b