Flux LRM Architecture
Data Path
- Dataset returns pairwise examples:
- caption
- image_0, image_1
- preference labels label_0, label_1
- timestep tensor for each image in pair
- Tokenization:
- tokenizer -> input_ids for CLIP text encoder
- tokenizer_2 -> input_ids_2 for T5 text encoder
- Image preprocessing:
- Resize/crop/normalize to model input resolution.
Model Path
- Text branch
- CLIP text encoder produces pooled prompt embeddings.
- T5 encoder produces sequence prompt embeddings.
- Pooled CLIP output is projected to reward text embedding space.
- Image branch
- VAE encodes images into latent tensors.
- Noise is sampled and mixed with latents using flow-style sigma schedule.
- Latents are packed into Flux token format.
- Flux transformer predicts token outputs conditioned on text embeddings.
- Token outputs are pooled and projected to reward image embedding space.
- Reward scoring
- Normalize text/image embeddings.
- Pairwise logits are computed with learnable temperature (logit_scale).
- Criterion computes pairwise preference loss.
Training Control
- Task orchestrates train/eval loops and metric reporting.
- Accelerator manages distributed setup, mixed precision, checkpointing.
- Hydra config composes model/dataset/criterion/task/optimizer/scheduler groups.
Notes
- Flux schnell guidance is typically 0.0.
- Dual text encoders are trainable by default (unless explicitly frozen).