ReChannel / README.md
nielsr's picture
nielsr HF Staff
Add model card, link to paper and code
16afaf5 verified
|
Raw
History Blame
1.96 kB
metadata
license: mit
pipeline_tag: image-to-image

ReChannel: Pixel-Space Dense Prediction with Text-to-Image Models

This repository contains the weights for ReChannel presented in From RGB Generation to Dense Field Readout: Pixel-Space Dense Prediction with Text-to-Image Models.

ReChannel reads dense prediction targets (such as depth, surface normals, matting, and referring segmentation) out of a FLUX-Klein text-to-image DiT with lightweight per-task LoRA adapters on the (otherwise frozen) backbone.

Quick Start (Inference)

To run the model locally, please clone the GitHub repository and install the dependencies:

git clone https://github.com/xmz111/ReChannel.git
cd ReChannel
pip install -r requirements.txt

You can then run dense prediction tasks on a single image:

python infer.py --image assets/demo_input.jpg \
    --tasks depth,normal,matting,refseg \
    --phrase "the right couch" \
    --out out.png
  • --tasks: any subset of depth, normal, matting, refseg.
  • --phrase: the referring expression used by refseg (text-conditioned).

How it works

RGB --VAE encoder--> latent tokens --DiT (frozen θ + task LoRA Δt, σ=0)--> token field Z_t
    --  Ŷ = reshape( W_t · z_ij + b_t ) ∈ R^{p×p×K}  --tile over the plane-->  dense field

The backbone is frozen; only a lightweight per-task LoRA adapter and the token-local linear head are trained. The head has no spatial mixing — all spatial structure comes from the adapted token field, not the head.

Citation

@article{wang2026rechannel,
  title={From RGB Generation to Dense Field Readout: Pixel-Space Dense Prediction with Text-to-Image Models},
  author={Wang, Zanyi and Lin, Xin and Li, Haodong and Jiang, Dengyang and Li, Yijiang},
  journal={arXiv preprint arXiv:2607.06553},
  year={2026}
}