Looping Beyond Twice: A Scalable Recipe for Looped Mixture-of-Experts Paper • 2610.01153 • Published 9 days ago • 24
SelfSearch: Reward-Free Search for Self-Improving Agents Paper • 2609.37968 • Published 10 days ago • 3
Follow the Entities: A Corpus Map for Agentic Search Paper • 2609.37226 • Published 11 days ago • 100
VeriHarness: Scaling Agentic Verification for Long-Horizon Tasks Paper • 2610.00972 • Published 9 days ago • 59
PivotOPD: Learning to Recover from Pivotal Mistakes in Multi-Turn Agents Paper • 2609.40285 • Published 10 days ago • 23
WorldAuditBench: Interactive 3D World Auditing with Multimodal Agents Paper • 2609.40325 • Published 10 days ago • 104
EvoDuet: Bilevel Co-Evolution of Web Searching and Task Solving for Scientific Discovery Paper • 2609.40340 • Published 10 days ago • 111
Self-Evolving Coding Agents: From Digital Programs to Physical-World Intelligence Paper • 2609.35432 • Published 12 days ago • 109
Mid-Harness: Scaling Actions Between Model and Harness for Terminal Agents Paper • 2609.39982 • Published 10 days ago • 119
LEGO-Anything: Coding Agents for 3D Scene Reconstruction Paper • 2609.36380 • Published 12 days ago • 130
Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL Paper • 2609.32577 • Published 14 days ago • 106
AREX-2: Advancing Self-Improving Agents through Long-Horizon Reflective Tasks Paper • 2609.38288 • Published 11 days ago • 136
YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality Paper • 2609.33757 • Published 13 days ago • 234
False Frontiers: Diagnosing and Mitigating Co-Cheating in Self-Evolving Search Agents Paper • 2609.39102 • Published 10 days ago • 362
Raven: The Harness of Harnesses for Composable Agentic Intelligence Paper • 2609.33439 • Published 13 days ago • 672
Skill-based Agentic Evaluation for Real-time Data Science Tasks Paper • 2609.16487 • Published 25 days ago • 2
ExplorationBench: Measuring AI Systems' Exploration in Verifiable Alien Worlds Paper • 2609.30199 • Published 16 days ago • 30
JEPA-Anything: Learning Predictive Models across Different Worlds Paper • 2609.20800 • Published 23 days ago • 78
Language Models that Play Chess and Explain Their Moves Paper • 2610.03695 • Published 8 days ago • 37