Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement Paper • 2608.31046 • Published 28 days ago • 97
HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness? Paper • 2609.01437 • Published 27 days ago • 221
Rethinking On-Policy Distillation of Large Language Models II: One Training Example Paper • 2609.04172 • Published 25 days ago • 102
SkillGate: Training In-Policy Skill Selection in Long-Horizon Agents Paper • 2608.18852 • Published Aug 19 • 9
AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Paper • 2608.05987 • Published Aug 6 • 103
Harness-R1: Learning to Edit Executable Runtime Harnesses from Agent Failure Trajectories Paper • 2608.02276 • Published Aug 3 • 4
VAD: Attributing Visual Evidence for Target Reconstruction in Multimodal On-Policy Distillation Paper • 2607.28590 • Published Jul 30 • 46
VAD: Attributing Visual Evidence for Target Reconstruction in Multimodal On-Policy Distillation Paper • 2607.28590 • Published Jul 30 • 46
SkillRise: Agentic Reinforcement Learning for Cross-Task Skill Evolution Paper • 2607.26784 • Published Jul 29 • 29
SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning Paper • 2607.14777 • Published Jul 16 • 104
Latent Reasoning in LLMs as a Vocabulary-Space Superposition Paper • 2510.15522 • Published Oct 17, 2025 • 7
SkillOpt: Executive Strategy for Self-Evolving Agent Skills Paper • 2605.23904 • Published May 22 • 264