Schrödinger's Code Repository: Have LLMs Learned SWE-bench or Memorized It? Paper • 2609.27891 • Published Aug 21 • 16
Schrödinger's Code Repository: Have LLMs Learned SWE-bench or Memorized It? Paper • 2609.27891 • Published Aug 21 • 16
ModularRSI: Modular and Generalizable Recursive Harness Self-Improvement Paper • 2609.14857 • Published 12 days ago • 214
SRPO: Self-Reflective Policy Optimization for Long-Horizon Reasoning Paper • 2608.23493 • Published Aug 24 • 1
EarlyEval: Cheaper Agent Evaluation via Early Outcome Prediction Paper • 2609.02783 • Published 24 days ago • 62
view post Post 74 You don't have to run an agent evaluation to the end. An agent's final outcome is usually visible from its early behavior, so we just stop the run once it's clear. EarlyEval cuts 13% to 26% of steps and up to 44% of input tokens, and resolve rates move by only 1 to 2 points. EarlyEval: Cheaper Agent Evaluation via Early Outcome Prediction (2609.02783)https://github.com/inphotoo/earlyeval See translation 🔥 1 1 + Reply
SRPO: Self-Reflective Policy Optimization for Long-Horizon Reasoning Paper • 2608.23493 • Published Aug 24 • 1
EarlyEval: Cheaper Agent Evaluation via Early Outcome Prediction Paper • 2609.02783 • Published 24 days ago • 62
EarlyEval: Cheaper Agent Evaluation via Early Outcome Prediction Paper • 2609.02783 • Published 24 days ago • 62
CyberFactory: Scaling Cyber Security Capabilities with Instances from the Wild Paper • 2608.23181 • Published Aug 24 • 34
ParaTempo: Efficient Parallel Reasoning via Temporal Confidence Paper • 2608.16425 • Published Aug 17 • 40
ParaTempo: Efficient Parallel Reasoning via Temporal Confidence Paper • 2608.16425 • Published Aug 17 • 40
ParaTempo: Efficient Parallel Reasoning via Temporal Confidence Paper • 2608.16425 • Published Aug 17 • 40
SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring Paper • 2608.09802 • Published Aug 10 • 84
Rethinking Code Complexity Through the Lens of Large Language Models Paper • 2602.07882 • Published May 27
SkillForge: Self-Distilling Agents for Project-Specific Issue Resolution Paper • 2608.18933 • Published Aug 19 • 13
SkillForge: Self-Distilling Agents for Project-Specific Issue Resolution Paper • 2608.18933 • Published Aug 19 • 13
Second Thought: Reasoning in Parallel as LLM Agents Act and Observe Paper • 2608.13667 • Published Aug 13 • 17