ModularRSI: Modular and Generalizable Recursive Harness Self-Improvement Paper • 2609.14857 • Published 14 days ago • 214
EarlyEval: Cheaper Agent Evaluation via Early Outcome Prediction Paper • 2609.02783 • Published 26 days ago • 62
CyberFactory: Scaling Cyber Security Capabilities with Instances from the Wild Paper • 2608.23181 • Published Aug 24 • 34
Second Thought: Reasoning in Parallel as LLM Agents Act and Observe Paper • 2608.13667 • Published Aug 13 • 17
How Do Practitioners Build SE Agents? Insights from a Mixed-Methods Study Paper • 2607.10856 • Published Jul 12 • 7
Are Performance-Optimization Benchmarks Reliably Measuring Coding Agents? Paper • 2607.01211 • Published Jul 1 • 12
Running on CPU Upgrade 14.1k Open LLM Leaderboard 🏆 14.1k Track, rank and evaluate open LLMs and chatbots
Running Agents 1.52k Big Code Models Leaderboard 📈 1.52k Explore code model rankings and submit evaluation requests