OmniTaskonomy: When Does Visual Generation Improve Visual Understanding? Paper • 2609.38079 • Published 1 day ago • 28
WorldAttention: An Efficient Attention Architecture for Interactive Video World Models Paper • 2609.34606 • Published 3 days ago • 13
Lucida: Parse, Generate, and Place for Composable Real-to-Sim Scene Modeling Paper • 2608.30821 • Published about 1 month ago • 69
Zero-WAM: In-Context World-Action Modeling from Human Videos for Open-Ended Task Generalization Paper • 2608.26103 • Published Aug 26 • 26
ReViV: Reconstructing the Viewer and the View in 4D from Monocular Egocentric Video Paper • 2607.17790 • Published Jul 20 • 5
Economy of Minds: Emerging Multi-Agent Intelligence with Economic Interactions Paper • 2606.02859 • Published Jun 1 • 11
Crafter: A Multi-Agent Harness for Editable Scientific Figure Generation from Diverse Inputs Paper • 2605.30611 • Published May 28 • 65
When Do Diffusion Models learn to Generate Multiple Objects? Paper • 2605.00273 • Published Apr 30 • 9
CityRAG: Stepping Into a City via Spatially-Grounded Video Generation Paper • 2604.19741 • Published Apr 21 • 18