Skill Self-Play: Pushing the Frontier of LLM Capability with Co-Evolving Skills Paper • 2607.22529 • Published 24 days ago • 48
Learning from Your Own Mistakes: Constructing Learnable Micro-Reflective Trajectories for Self-Distillation Paper • 2606.18844 • Published Jun 17 • 20
ADHint: Adaptive Hints with Difficulty Priors for Reinforcement Learning Paper • 2512.13095 • Published Dec 15, 2025 • 2
RAVE: Re-Allocating Visual Attention in Large Multimodal Models Paper • 2605.18359 • Published May 26 • 1
Revisiting Reinforcement Learning with Verifiable Rewards from a Contrastive Perspective Paper • 2605.12969 • Published May 30 • 1
Towards Flash Thinking via Decoupled Advantage Policy Optimization Paper • 2510.15374 • Published Oct 17, 2025
Learning from Your Own Mistakes: Constructing Learnable Micro-Reflective Trajectories for Self-Distillation Paper • 2606.18844 • Published Jun 17 • 20