FLAT: Resampling Image and Text into 1D Flexible-Length Aligned Transmodal Tokens for Retrieval and Generation Paper • 2609.16591 • Published 3 days ago • 13
RepFusion: Leveraging Multimodal Priors for Denoising in Representation Space Paper • 2606.14700 • Published Jun 12 • 18
DREAM: Where Visual Understanding Meets Text-to-Image Generation Paper • 2603.02667 • Published Mar 3 • 6