Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation Paper • 2609.08798 • Published 2 days ago • 78
HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness? Paper • 2609.01437 • Published 9 days ago • 261
LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL Paper • 2607.04412 • Published Jul 5 • 35