SWE-Touch: Benchmarking Coding Agents When Users Touch the Code Paper • 2608.02499 • Published 3 days ago • 23
The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation Paper • 2607.24720 • Published 10 days ago • 26
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Paper • 2607.15330 • Published 21 days ago • 73
Why Multi-Step Tool-Use Reinforcement Learning Collapses and How Supervisory Signals Fix It Paper • 2606.26027 • Published Jun 24 • 18
Look Light, Think Heavy: What Multimodal Chain-of-Thought Reasoning Can and Cannot Do Paper • 2606.22565 • Published Jun 21 • 9
Joint Training of Multi-Token Prediction in Reinforcement Learning via Optimal Coefficient Calibration Paper • 2605.28184 • Published May 27 • 6
Joint Training of Multi-Token Prediction in Reinforcement Learning via Optimal Coefficient Calibration Paper • 2605.28184 • Published May 27 • 6
Coloring the Noise: Adversarial Sobolev Alignment for Faithful Image Super Resolution Paper • 2605.23264 • Published May 22 • 7
Uncovering Entity Identity Confusion in Multimodal Knowledge Editing Paper • 2605.06096 • Published May 7 • 1
ResRL: Boosting LLM Reasoning via Negative Sample Projection Residual Reinforcement Learning Paper • 2605.00380 • Published May 1 • 7
From P(y|x) to P(y): Investigating Reinforcement Learning in Pre-train Space Paper • 2604.14142 • Published Apr 15 • 30
PLUME: Latent Reasoning Based Universal Multimodal Embedding Paper • 2604.02073 • Published Apr 2 • 15
CLIPO: Contrastive Learning in Policy Optimization Generalizes RLVR Paper • 2603.10101 • Published Mar 10 • 6
CC-VQA: Conflict- and Correlation-Aware Method for Mitigating Knowledge Conflict in Knowledge-Based Visual Question Answering Paper • 2602.23952 • Published Feb 27 • 3
CC-VQA: Conflict- and Correlation-Aware Method for Mitigating Knowledge Conflict in Knowledge-Based Visual Question Answering Paper • 2602.23952 • Published Feb 27 • 3
Flexible Entropy Control in RLVR with Gradient-Preserving Perspective Paper • 2602.09782 • Published Feb 10 • 3
IF-Bench: Benchmarking and Enhancing MLLMs for Infrared Images with Generative Visual Prompting Paper • 2512.09663 • Published Dec 10, 2025 • 4
IF-Bench: Benchmarking and Enhancing MLLMs for Infrared Images with Generative Visual Prompting Paper • 2512.09663 • Published Dec 10, 2025 • 4