From RGB Generation to Dense Field Readout: Pixel-Space Dense Prediction with Text-to-Image Models Paper • 2607.06553 • Published 23 days ago • 20
SPACENUM: Revisiting Spatial Numerical Understanding in VLMs Paper • 2605.23898 • Published May 22 • 7
Reward-Decomposed Reinforcement Learning for Immersive Video Role-Playing Paper • 2605.04733 • Published Jun 3
On-Policy Distillation with Best-of-N Teacher Rollout Selection Paper • 2605.09725 • Published May 13
ThinkJEPA: Empowering Latent World Models with Large Vision-Language Reasoning Model Paper • 2603.22281 • Published Mar 23 • 20
Vision Language Models Cannot Reason About Physical Transformation Paper • 2603.07109 • Published Mar 7 • 2
SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks Paper • 2602.12670 • Published Feb 13 • 65
Less is Enough: Synthesizing Diverse Data in Feature Space of LLMs Paper • 2602.10388 • Published Feb 11 • 246
Ref-NeuS: Ambiguity-Reduced Neural Implicit Surface Learning for Multi-View Reconstruction with Reflection Paper • 2303.10840 • Published Mar 20, 2023 • 1
GaussianProperty: Integrating Physical Properties to 3D Gaussians with LMMs Paper • 2412.11258 • Published Dec 15, 2024 • 13
AccidentBench: Benchmarking Multimodal Understanding and Reasoning in Vehicle Accidents and Beyond Paper • 2509.26636 • Published Sep 30, 2025 • 1
See-Control: A Multimodal Agent Framework for Smartphone Interaction with a Robotic Arm Paper • 2512.08629 • Published Dec 9, 2025 • 1
LLM-Optic: Unveiling the Capabilities of Large Language Models for Universal Visual Grounding Paper • 2405.17104 • Published May 27, 2024
CoMAS: Co-Evolving Multi-Agent Systems via Interaction Rewards Paper • 2510.08529 • Published Oct 9, 2025 • 19