The Tool Decathlon: Benchmarking Language Agents for Diverse, Realistic, and Long-Horizon Task Execution Paper • 2510.25726 • Published Oct 29, 2025 • 48
DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression Paper • 2609.19969 • Published 5 days ago • 159
Dream-RSI: Recursive Self-Improvement through Evolving Worlds Paper • 2609.14858 • Published 8 days ago • 242
The Last AI Built by Humans: Toward Genuine Recursive Self-Improvement Paper • 2609.11873 • Published 12 days ago • 91
SPADE: Self-Play in Adaptive Synthetic Executable Environments Paper • 2608.19197 • Published Aug 19 • 54
StarVLA: A Lego-like Codebase for Vision-Language-Action Model Developing Paper • 2604.05014 • Published Apr 6 • 1
Beyond Data Scaling: Representation-Centric Continued Pre-training for Vision-Language-Action Models Paper • 2608.27550 • Published 26 days ago • 84
Recursive Experiential-Working Memory Evolution for Long-Horizon Agent Harnesses Paper • 2608.24876 • Published 28 days ago • 30
SemComp-Bench: Benchmarking Semantic Task Completion in Video Generation Paper • 2608.17426 • Published Aug 18 • 160