RRSI: Regularized Recursive Self-Improvement of Agent Harnesses Paper • 2609.24972 • Published 5 days ago • 206
The Tool Decathlon: Benchmarking Language Agents for Diverse, Realistic, and Long-Horizon Task Execution Paper • 2510.25726 • Published Oct 29, 2025 • 48
DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression Paper • 2609.19969 • Published 9 days ago • 180
Dream-RSI: Recursive Self-Improvement through Evolving Worlds Paper • 2609.14858 • Published 12 days ago • 246
The Last AI Built by Humans: Toward Genuine Recursive Self-Improvement Paper • 2609.11873 • Published 16 days ago • 91
SPADE: Self-Play in Adaptive Synthetic Executable Environments Paper • 2608.19197 • Published Aug 19 • 54
StarVLA: A Lego-like Codebase for Vision-Language-Action Model Developing Paper • 2604.05014 • Published Apr 6 • 1
Beyond Data Scaling: Representation-Centric Continued Pre-training for Vision-Language-Action Models Paper • 2608.27550 • Published about 1 month ago • 84
Recursive Experiential-Working Memory Evolution for Long-Horizon Agent Harnesses Paper • 2608.24876 • Published Aug 25 • 30
SemComp-Bench: Benchmarking Semantic Task Completion in Video Generation Paper • 2608.17426 • Published Aug 18 • 160
view article Article Hugging Face and Cerebras bring Gemma 4 to real-time voice AI +2 A-Mahla, andito, lvwerra, vyassaurabh • Jul 1 • 105
view article Article Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident +2 hlarcher, XciD, raphael-gl, chris-rannou • Jul 27 • 506
LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks Paper • 2608.01964 • Published Aug 3 • 188
view article Article AI and the Future of Cybersecurity: Why Openness Matters +1 meg, yjernite, clem • Apr 21 • 52
S-Agent: Spatial Tool-Use Elicits Reasoning for Spatial Intelligence Paper • 2606.20515 • Published Jun 18 • 42
Apple-π: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence Paper • 2607.16401 • Published Jul 17 • 45