MemBodied: Recurrent Associative Memory for Vision-Language-Action Models Paper • 2609.28256 • Published 3 days ago • 10
StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows Paper • 2608.17800 • Published Aug 18 • 12
AgentMercury: Your Agent Can Synthesize Verifiable Environments for Business Scenarios at scale Paper • 2608.20634 • Published Aug 21 • 13
SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent Harness Paper • 2609.20519 • Published 9 days ago • 130
HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness? Paper • 2609.01437 • Published 25 days ago • 220
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Paper • 2607.15330 • Published Jul 16 • 72
Evolution Fine-Tuning: Learning to Discover Across 371 Optimization Tasks Paper • 2606.29082 • Published Jun 27 • 43
EvoPolicyGym: Evaluating Autonomous Policy Evolution in Interactive Environments Paper • 2607.02440 • Published Jul 2 • 49
Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents Paper • 2606.19704 • Published Jun 18 • 43
EvoArena: Tracking Memory Evolution for Robust LLM Agents in Dynamic Environments Paper • 2606.13681 • Published Jun 11 • 146
VibeSearchBench: Benchmarking Long-horizon Proactive Search in the Wild Paper • 2605.27882 • Published May 27 • 17
Foundation Protocol: A Coordination Layer for Agentic Society Paper • 2605.23218 • Published May 22 • 80