Iris-3B: Going Beyond the Latent with Pixel-Space Diffusion Training, Conversion and Fine-Tuning
Abstract
Pixel-space diffusion models avoid the lossy VAE of latent models, which suggests an advantage on downstream tasks where fine-grained detail matters. We test this claim along both routes to a pixel-space backbone. We pretrain Iris-3B, a 3B-parameter pixel-space text-to-image transformer, from scratch through a 256to512to1024 curriculum, after first ablating the prediction target and representation alignment at 256^2 to decide what to scale. We also convert a pretrained latent model, FLUX.2 Klein base 4B, to pixel space. We fine-tune both families for monocular depth estimation and for image restoration/super-resolution. We find no significant improvement from using a pixel-space generative prior. Fine-tuned for depth with one matched direct-regression recipe, Iris-3B is level with the latent FLUX.2 Klein and the converted pixel FLUX.2 Klein falls behind it, and on 4times DIV2K restoration neither pixel model beats a latent FLUX.2 Klein fine-tune, the converted one trailing it slightly. We document the recipes, the failure modes and the remaining confounds behind this negative result. Nevertheless, Iris-3B shows that pixel-space pretraining with the pixel-transformer (PiT) head of PixelDiT scales to 3B parameters and to text-to-image quality competitive with latent models, matching Qwen-Image on OneIG under the official evaluators at 1024^2. We release its weights and training code in the hope that they help pave the way for further work on pixel-space generation.
Community
Meet Iris-3B: a pixel-space T2I model and general vision learner that generates every pixel directly. No VAE.
Pre-trained from scratch, 3B, open weights.
We also explore its generative prior on detail-critical vision tasks: depth estimation and image restoration.
Fully open: weights, code and demo (Apache 2.0)
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- PixelDiT2: Representation-Grounded Pixel Diffusion Transformers (2026)
- EagleDepth: Efficient Fine-Grained Depth Estimation via Pixel Diffusion Decoder (2026)
- PixRestore: Unified Image Restoration via Pixel Diffusion Transformer (2026)
- When Latents Forget Pixels: Restoring Fidelity in Diffusion Transformer Super-Resolution (2026)
- PixelDense: Dense Prediction as Representation Alignment for Pixel Diffusion (2026)
- Adversarial Training for Pixel Diffusion (2026)
- PixelART: Image-to-Layer Decomposition without Latents or Text-to-Image Pretraining (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2610.09450 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 2
suryatmodulus/iris-3b
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 4
Collections including this paper 0
No Collection including this paper