Uncovering Understanding-Generation Synergy in Native Unified Multimodal Models: From Representation, Task to System Paper • 2609.01607 • Published 25 days ago • 24
VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning Paper • 2608.26105 • Published Aug 26 • 190
V-Rubrics: Visual Faithfulness via Rubric-Based Reinforcement Learning Paper • 2608.25580 • Published Aug 26 • 16
Zetta ζ: An Efficient Closed-Loop Embodied Harness for Self-Evolving Physical Intelligence Paper • 2608.16590 • Published Aug 17 • 152
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Paper • 2607.17423 • Published Jul 19 • 98
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Paper • 2607.14935 • Published Jul 16 • 120
S-Agent: Spatial Tool-Use Elicits Reasoning for Spatial Intelligence Paper • 2606.20515 • Published Jun 18 • 42
Show the Signal, Hide the Noise: Spectral Forcing for Pixel-Space Diffusion Paper • 2606.15236 • Published Jun 16 • 22
On the Scaling of PEFT: Towards Million Personal Models of Trillion Parameters Paper • 2606.02437 • Published Jun 1 • 146
Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments Paper • 2605.30280 • Published May 28 • 146
Function2Scene: 3D Indoor Scene Layout from Functional Specifications Paper • 2605.30819 • Published May 29 • 39
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Paper • 2605.12500 • Published May 12 • 199
MolmoAct2: Action Reasoning Models for Real-world Deployment Paper • 2605.02881 • Published May 4 • 356
SpatialBench: Is Your Spatial Foundation Model an All-Round Player? Paper • 2605.27367 • Published May 26 • 70
PhysX-Omni: Unified Simulation-Ready Physical 3D Generation for Rigid, Deformable, and Articulated Objects Paper • 2605.21572 • Published May 20 • 54
VisionFoundry: Teaching VLMs Visual Perception with Synthetic Images Paper • 2604.09531 • Published Apr 10 • 10