SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent Harness Paper • 2609.20519 • Published 9 days ago • 131
PLC-DPO: Posterior Label Correction in Noisy and Ambiguous Preference Optimization Paper • 2608.30597 • Published 26 days ago • 27
Online Draft Co-Training for Speculative Decoding in Large-Scale, Long-Context RL Post-Training Paper • 2609.07108 • Published 19 days ago • 36
Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation Paper • 2609.08798 • Published 18 days ago • 83
WHALE: A Simple Recipe for Joint Harness-Weight Optimization Paper • 2609.00196 • Published 26 days ago • 37
SkillCoach: Self-Evolving Rubrics for Evaluating and Enhancing Agentic Skill-Use Paper • 2607.01874 • Published Jul 2 • 22
LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL Paper • 2607.04412 • Published Jul 5 • 36
LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL Paper • 2607.04412 • Published Jul 5 • 36
LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL Paper • 2607.04412 • Published Jul 5 • 36
Evolution Fine-Tuning: Learning to Discover Across 371 Optimization Tasks Paper • 2606.29082 • Published Jun 27 • 43
SoCRATES: Towards Reliable Automated Evaluation of Proactive LLM Mediation across Domains and Socio-cognitive Variations Paper • 2606.05563 • Published Jun 4 • 56