From Generic Correlation to Input-Specific Credit in On-Policy Self Distillation Paper • 2605.11613 • Published May 12 • 3
Anti-Self-Distillation for Reasoning RL via Pointwise Mutual Information Paper • 2605.11609 • Published May 12 • 196
VibeLifeBench: Can Your Life Agent Be Proactive and Persistent in a Living World? Paper • 2608.10875 • Published 1 day ago • 11
Bootstrapping Exploration with Group-Level Natural Language Feedback in Reinforcement Learning Paper • 2603.04597 • Published Mar 4 • 211
REDSearcher: A Scalable and Cost-Efficient Framework for Long-Horizon Search Agents Paper • 2602.14234 • Published Feb 15 • 28
VESPO: Variational Sequence-Level Soft Policy Optimization for Stable Off-Policy LLM Training Paper • 2602.10693 • Published Feb 11 • 222
DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning Paper • 2505.14362 • Published May 20, 2025 • 6
REDSearcher: A Scalable and Cost-Efficient Framework for Long-Horizon Search Agents Paper • 2602.14234 • Published Feb 15 • 28
VESPO: Variational Sequence-Level Soft Policy Optimization for Stable Off-Policy LLM Training Paper • 2602.10693 • Published Feb 11 • 222
DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning Paper • 2505.14362 • Published May 20, 2025 • 6