SWE Refactor Bench: Can Coding Agents Complete a Long-Horizon, Whole-Repository Stack Migration? Paper • 2608.23564 • Published 6 days ago • 14
SWE Refactor Bench: Can Coding Agents Complete a Long-Horizon, Whole-Repository Stack Migration? Paper • 2608.23564 • Published 6 days ago • 14
Reasoning to Rank: An End-to-End Solution for Exploiting Large Language Models for Recommendation Paper • 2602.12530 • Published Feb 13
The User Asks, Platforms Compete: How Agentic Recommendation Markets Take Shape Paper • 2607.25253 • Published Jul 29
SWE Refactor Bench: Can Coding Agents Complete a Long-Horizon, Whole-Repository Stack Migration? Paper • 2608.23564 • Published 6 days ago • 14
AI4AI-Bench: Benchmarking LLM Agents in Algorithmic Design for Recursive Self-Improvement Paper • 2608.20318 • Published 10 days ago • 2
Frontier-Eng: Benchmarking Self-Evolving Agents on Real-World Engineering Tasks with Generative Optimization Paper • 2604.12290 • Published Apr 14 • 16
Weak-to-Strong Generalization via Direct On-Policy Distillation Paper • 2607.05394 • Published Jul 8 • 150