all_code_base / lrm /flux /docs /architecture.md
aryadomain's picture
Add files using upload-large-folder tool
b4efe93 verified
|
Raw
History Blame Contribute Delete
1.44 kB

Flux LRM Architecture

Data Path

  1. Dataset returns pairwise examples:
  • caption
  • image_0, image_1
  • preference labels label_0, label_1
  • timestep tensor for each image in pair
  1. Tokenization:
  • tokenizer -> input_ids for CLIP text encoder
  • tokenizer_2 -> input_ids_2 for T5 text encoder
  1. Image preprocessing:
  • Resize/crop/normalize to model input resolution.

Model Path

  1. Text branch
  • CLIP text encoder produces pooled prompt embeddings.
  • T5 encoder produces sequence prompt embeddings.
  • Pooled CLIP output is projected to reward text embedding space.
  1. Image branch
  • VAE encodes images into latent tensors.
  • Noise is sampled and mixed with latents using flow-style sigma schedule.
  • Latents are packed into Flux token format.
  • Flux transformer predicts token outputs conditioned on text embeddings.
  • Token outputs are pooled and projected to reward image embedding space.
  1. Reward scoring
  • Normalize text/image embeddings.
  • Pairwise logits are computed with learnable temperature (logit_scale).
  • Criterion computes pairwise preference loss.

Training Control

  • Task orchestrates train/eval loops and metric reporting.
  • Accelerator manages distributed setup, mixed precision, checkpointing.
  • Hydra config composes model/dataset/criterion/task/optimizer/scheduler groups.

Notes

  • Flux schnell guidance is typically 0.0.
  • Dual text encoders are trainable by default (unless explicitly frozen).