MotorMind: Scaffolding General Vision Language Models for Zero-Shot Robot Manipulation Paper • 2609.38078 • Published 7 days ago • 83
VGI-Bench: Probing Visual Intelligence in Video Generation Models Paper • 2608.19583 • Published Aug 26 • 337
Watch Before You Answer: Learning from Visually Grounded Post-Training Paper • 2604.05117 • Published Apr 6 • 162
RenderFormer-V2: Neural Rendering with Heterogeneous Scene Primitives Paper • 2609.05738 • Published Sep 4 • 6
Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills Paper • 2609.02749 • Published Sep 2 • 408
HarnessEval-W: Agentifying the Evaluation of Visual Worlds Paper • 2608.16859 • Published Aug 17 • 122