Not Worth Another Token: Marginal Value Estimation for Efficient Deep Research Agents Paper • 2608.08389 • Published 4 days ago • 11
Mendel Gödel Machine: Recursive Self-Improving Coding Agents via Comparative Evolution Paper • 2608.07645 • Published 6 days ago • 22
The Optimizer Is the Agent: Reasoning-Driven Search across Prompts, Programs, and ML Workflows Paper • 2608.06714 • Published 6 days ago • 9
Efficient Knowledge Distillation for LLMs: Offline Top-K Logits and a Fused Chunked KL Loss Paper • 2608.03796 • Published 9 days ago • 14
SFT Conflicts, RL Coexists: A Theoretical and Empirical Analysis of Multi-Task Learning for LLMs Paper • 2608.03573 • Published 7 days ago • 51
Agent Memory Distillation: Empowering Small LLM Agents with Hierarchical Teacher Memory Paper • 2608.07169 • Published 6 days ago • 47
Ouroboros: A Self-Developing Frontier Coding Agent with Reviewed Core Evolution Paper • 2608.08311 • Published 5 days ago • 79
Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA Paper • 2608.09819 • Published 3 days ago • 324
BDH-CQ: In-Context Learning with Recurrent Latent Reasoning Paper • 2608.09888 • Published 3 days ago • 555
Distill Where You Fail: Recovering Learning Signals of Negative RL-Groups from Adaptive Teacher Guidance Paper • 2608.00782 • Published 12 days ago • 16
GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks Paper • 2608.03764 • Published 9 days ago • 27
Progressive Agent Skill Generation via Reinforcement Learning Paper • 2608.01678 • Published 10 days ago • 58
SkillRise: Agentic Reinforcement Learning for Cross-Task Skill Evolution Paper • 2607.26784 • Published 15 days ago • 28
Multi-Head Latent Control: A Unified Interface for LLM Agent Decision Making Paper • 2607.14277 • Published 29 days ago • 10