SimpleMemVLA: A Simple but Effective Native-Video Memory for Vision-Language-Action Models Paper • 2609.05533 • Published 18 days ago • 15
StudyBench: Can Self-Evolution Squeeze Textbooks for Olympiad Capability? Paper • 2609.00787 • Published 19 days ago • 19
Φ-Bench: Can Large Language Models Engineer the Infrastructure That Powers Them? Paper • 2609.10226 • Published 11 days ago • 19
SWE-Bench Pro Verified: A Reliable Benchmark for Software Engineering Agents Paper • 2609.08149 • Published 12 days ago • 28
T1: Terminal Agent Reinforcement Learning for Long-Horizon Tasks Paper • 2609.11042 • Published 10 days ago • 61
Rethinking On-Policy Distillation of Large Language Models II: One Training Example Paper • 2609.04172 • Published 17 days ago • 98
Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Requirements Paper • 2608.17310 • Published Aug 18 • 108
MinT: Managed Infrastructure for Training and Serving Millions of LLMs Paper • 2605.13779 • Published May 13 • 225
Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients Paper • 2606.18216 • Published Jun 16 • 65
LoopCoder-v2: Only Loop Once for Efficient Test-Time Computation Scaling Paper • 2606.18023 • Published Jun 16 • 161
MobileGym: A Verifiable and Highly Parallel Simulation Platform for Mobile GUI Agent Research Paper • 2605.26114 • Published May 25 • 63
TerminalWorld: Benchmarking Agents on Real-World Terminal Tasks Paper • 2605.22535 • Published May 21 • 9