DistilVDR: A Compact End-to-End Visual Document Retriever via Dual-Student Distillation Paper • 2608.10636 • Published 1 day ago • 4
SimWAM: A Simple World Action Model for End-to-End Autonomous Driving Paper • 2608.07468 • Published 6 days ago • 101
ReToken: One Token to Improve Vision-Language Models for Visual Retrieval Paper • 2607.28627 • Published 14 days ago • 9
ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU Paper • 2607.19191 • Published 22 days ago • 311
HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enchancement Paper • 2607.18217 • Published 24 days ago • 61
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Paper • 2607.17423 • Published 25 days ago • 167
MultAttnAttrib: Training-Free Multimodal Attribution in Long Document Question Answering Paper • 2607.01420 • Published Jul 1 • 13
OmniOpt: Taxonomy, Geometry, and Benchmarking of Modern Optimizers Paper • 2607.04033 • Published Jul 4 • 76
ABot-M0.5: Unified Mobility-and-Manipulation World Action Model Paper • 2607.00678 • Published Jul 1 • 20
Parallel Rollout Approximation for Pixel-Space Autoregressive Image Generation Paper • 2606.27978 • Published Jun 26 • 6
SkillHone: A Harness for Continual Agent Skill Evolution Through Persistent Decision History Paper • 2606.08671 • Published Jun 23 • 47
Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models Paper • 2606.11324 • Published Jun 9 • 172
InterleaveThinker: Reinforcing Agentic Interleaved Generation Paper • 2606.13679 • Published Jun 11 • 84
Benchmarks are Not Enough: RAMP for Runtime Assessing of Agentic Models in Production Systems Paper • 2605.27492 • Published May 26 • 25
How and What to Imagine? Visual Thinking in Unified Multimodal Models for Cross-View Spatial Reasoning Paper • 2605.27310 • Published May 26 • 20