Rethinking On-Policy Distillation of Large Language Models II: One Training Example Paper • 2609.04172 • Published 27 days ago • 104
Random Attention: Rethinking KV Cache Eviction for Efficient Reasoning Paper • 2609.03430 • Published 27 days ago • 187
Evaluating the Hidden Costs of Personalization in Large Language Models Paper • 2608.28833 • Published Aug 28 • 30
Evaluating the Hidden Costs of Personalization in Large Language Models Paper • 2608.28833 • Published Aug 28 • 30
Co-Evolution in Agentic Systems: Toward Self-Directed Evolution Beyond Human Design Paper • 2608.10299 • Published Aug 10 • 138
MemGuard: Preventing Memory Contamination in Long-Term Memory-Augmented Large Language Models Paper • 2605.28009 • Published May 27 • 3
AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses Paper • 2608.12307 • Published Aug 12 • 118
MemGuard: Preventing Memory Contamination in Long-Term Memory-Augmented Large Language Models Paper • 2605.28009 • Published May 27 • 3
Co-Evolution in Agentic Systems: Toward Self-Directed Evolution Beyond Human Design Paper • 2608.10299 • Published Aug 10 • 138
Explorative Modeling: Unlocking a Third Pretraining Axis and End-to-End Generation Paper • 2607.27372 • Published Jul 29 • 19
AgentDebugX: An Open-Source Toolkit for Failure Observability, Attribution, and Recovery in LLM Agents Paper • 2607.18754 • Published Jul 21 • 25
EvolvingWorld: An Open-Schema Framework for Co-Evolving Role-Play Agents and World Model in Interactive Literary World Paper • 2607.17250 • Published Jul 19 • 93
EvolvingWorld: An Open-Schema Framework for Co-Evolving Role-Play Agents and World Model in Interactive Literary World Paper • 2607.17250 • Published Jul 19 • 93
AdaPlanBench: Evaluating Adaptive Planning in Large Language Model Agents under World and User Constraints Paper • 2606.05622 • Published Jun 4 • 45
Trimming the Long-Tail of Visual World Modeling Evaluation Paper • 2606.24256 • Published Jun 23 • 41
GBC: Gradient-Based Connections for Optimizing Multi-Agent Systems Paper • 2606.28187 • Published Jun 26 • 14
BioInsight: Multi-Agent Orchestration for Interactive Biomedical Knowledge Discovery Paper • 2606.20997 • Published Jun 19 • 11
BioInsight: Multi-Agent Orchestration for Interactive Biomedical Knowledge Discovery Paper • 2606.20997 • Published Jun 19 • 11
PlanBench-XL: Evaluating Long-Horizon Planning of LLM Tool-Use Agents in Large-Scale Tool Ecosystems Paper • 2606.22388 • Published Jun 21 • 96