β-OPSD: Deriving with Policy Optimization, Training with Self-Distillation Paper • 2607.28582 • Published 12 days ago • 23
You Only Need Minimal RLVR Training: Extrapolating LLMs via Rank-1 Trajectories Paper • 2605.21468 • Published May 20 • 51
Weak-Driven Learning: How Weak Agents make Strong Agents Stronger Paper • 2602.08222 • Published Feb 9 • 290
AutoTool: Dynamic Tool Selection and Integration for Agentic Reasoning Paper • 2512.13278 • Published Dec 15, 2025
Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory Paper • 2511.20857 • Published Nov 25, 2025 • 3