EnvACE: Internalizing Environment Dynamics via World Rehearsal for Agentic Reinforcement Learning Paper • 2608.06197 • Published 3 days ago • 35
Can Text-to-Image Models Draw from the Right Frame of Reference? Paper • 2608.03357 • Published 5 days ago • 23
JLT Collection JLT: Clean-Latent Prediction in Latent Diffusion Transformers • 2 items • Updated May 30 • 1
JLT: Clean-Latent Prediction in Latent Diffusion Transformers Paper • 2605.27102 • Published May 26 • 33
Mitigating Modality Prior-Induced Hallucinations in Multimodal Large Language Models via Deciphering Attention Causality Paper • 2410.04780 • Published Oct 7, 2024 • 1
VisionFoundry: Teaching VLMs Visual Perception with Synthetic Images Paper • 2604.09531 • Published Apr 10 • 10
Modality Gap-Driven Subspace Alignment Training Paradigm For Multimodal Large Language Models Paper • 2602.07026 • Published Feb 2 • 141
Does Understanding Inform Generation in Unified Multimodal Models? From Analysis to Path Forward Paper • 2511.20561 • Published Nov 25, 2025 • 33
UltraFlux: Data-Model Co-Design for High-quality Native 4K Text-to-Image Generation across Diverse Aspect Ratios Paper • 2511.18050 • Published Nov 22, 2025 • 38