WarpSAC: Towards the Pinnacle of Scalable Off-policy RL by Rethinking Exploration and Exploitation Paper • 2608.24479 • Published 3 days ago • 134
The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning Paper • 2606.29526 • Published Jun 28 • 170
Trust-Region Behavior Blending for On-Policy Distillation Paper • 2605.31159 • Published May 29 • 69
Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models Paper • 2606.11324 • Published Jun 9 • 172