RRSI: Regularized Recursive Self-Improvement of Agent Harnesses Paper • 2609.24972 • Published 2 days ago • 152
CodeMidas: Scaling Agentic Coding RL Environments from Code Itself Paper • 2609.22068 • Published 5 days ago • 106
MindZero: Learning Online Mental Reasoning With Zero Annotations Paper • 2606.00240 • Published May 29 • 4