license: mit
pipeline_tag: image-to-image
ReChannel: Pixel-Space Dense Prediction with Text-to-Image Models
This repository contains the weights for ReChannel presented in From RGB Generation to Dense Field Readout: Pixel-Space Dense Prediction with Text-to-Image Models.
ReChannel reads dense prediction targets (such as depth, surface normals, matting, and referring segmentation) out of a FLUX-Klein text-to-image DiT with lightweight per-task LoRA adapters on the (otherwise frozen) backbone.
- Code: GitHub Repository
Quick Start (Inference)
To run the model locally, please clone the GitHub repository and install the dependencies:
git clone https://github.com/xmz111/ReChannel.git
cd ReChannel
pip install -r requirements.txt
You can then run dense prediction tasks on a single image:
python infer.py --image assets/demo_input.jpg \
--tasks depth,normal,matting,refseg \
--phrase "the right couch" \
--out out.png
--tasks: any subset ofdepth, normal, matting, refseg.--phrase: the referring expression used byrefseg(text-conditioned).
How it works
RGB --VAE encoder--> latent tokens --DiT (frozen θ + task LoRA Δt, σ=0)--> token field Z_t
-- Ŷ = reshape( W_t · z_ij + b_t ) ∈ R^{p×p×K} --tile over the plane--> dense field
The backbone is frozen; only a lightweight per-task LoRA adapter and the token-local linear head are trained. The head has no spatial mixing — all spatial structure comes from the adapted token field, not the head.
Citation
@article{wang2026rechannel,
title={From RGB Generation to Dense Field Readout: Pixel-Space Dense Prediction with Text-to-Image Models},
author={Wang, Zanyi and Lin, Xin and Li, Haodong and Jiang, Dengyang and Li, Yijiang},
journal={arXiv preprint arXiv:2607.06553},
year={2026}
}