Multi-Agent Egocentric World Model with Fine-Grained Embodied Interaction Paper • 2610.12299 • Published 4 days ago • 52
GRACE: Generation-aware latent compression for efficient video generation Paper • 2610.10524 • Published 5 days ago • 76
Tetris3D: 3D Scene Generation With Objects That Fit Together Paper • 2610.10539 • Published 5 days ago • 45
World Observer: Joint Actor-Observer Generation for Persistent World Modeling Paper • 2610.02162 • Published 11 days ago • 90
Imagine3D-LLM: Teaching MLLMs to Imagine 3D Scenes Before Answering Paper • 2609.38177 • Published 13 days ago • 74
ReDesign: Recovering Editable Design Structures from Images via Agentic Decomposition Paper • 2607.25565 • Published Jul 28 • 65
Geometry-Aware Representation Denoising for Robust Multi-view 3D Reconstruction Paper • 2605.26230 • Published May 25 • 38
TrackCraft3R: Repurposing Video Diffusion Transformers for Dense 3D Tracking Paper • 2605.12587 • Published May 12 • 37
WorldCam: Interactive Autoregressive 3D Gaming Worlds with Camera Pose as a Unifying Geometric Representation Paper • 2603.16871 • Published Mar 17 • 61
Grounding World Simulation Models in a Real-World Metropolis Paper • 2603.15583 • Published Mar 16 • 156
Lookahead Anchoring: Preserving Character Identity in Audio-Driven Human Animation Paper • 2510.23581 • Published Oct 27, 2025 • 42
Exploring Conditions for Diffusion models in Robotic Control Paper • 2510.15510 • Published Oct 17, 2025 • 40
TAG:Tangential Amplifying Guidance for Hallucination-Resistant Diffusion Sampling Paper • 2510.04533 • Published Oct 6, 2025 • 48
MATRIX: Mask Track Alignment for Interaction-aware Video Generation Paper • 2510.07310 • Published Oct 8, 2025 • 36
Visual Representation Alignment for Multimodal Large Language Models Paper • 2509.07979 • Published Sep 9, 2025 • 84