NeoMME: A Single-Tower Multimodal-Native Multilingual Foundation Encoder for Efficient Fine-Tuning and Inference Paper • 2609.01657 • Published 23 days ago • 34
Beyond Retrieval: Progressive Latent Memory Evolution for Streaming Video Understanding Paper • 2609.04131 • Published 20 days ago • 31
Knowing When Not to Reuse: Conditional Experience Transfer in Autonomous LLM Post-Training Paper • 2608.26730 • Published 27 days ago • 155
LightNav-0: Eliciting VLM Spatial Intelligence for Generalist Embodied Navigation Paper • 2608.30935 • Published 23 days ago • 32
Revisiting Local Context for Long-Horizon Streaming 3D Reconstruction Paper • 2608.27529 • Published 27 days ago • 32
EgoSuite-Open100K Collection The largest fully-annotated open egocentric human dataset. 100,000 hours across 15,000+ tasks and scenes. • 3 items • Updated Aug 19 • 59
ARIS: Autonomous Research via Adversarial Multi-Agent Collaboration Paper • 2605.03042 • Published May 4 • 152
VibeWorlding: Can Multimodal Agents Construct 3D Open Worlds End-to-End? Paper • 2608.15265 • Published Aug 15 • 60
Zetta ζ: An Efficient Closed-Loop Embodied Harness for Self-Evolving Physical Intelligence Paper • 2608.16590 • Published Aug 17 • 152
MatrAIx: Simulating the World with 8.3 Billion Persona Agents Paper • 2608.04205 • Published Aug 4 • 54
HarnessEval-W: Agentifying the Evaluation of Visual Worlds Paper • 2608.16859 • Published Aug 17 • 122