Demystifying Agent Skills: Why They Work-Until They Don't Paper • 2608.14036 • Published 11 days ago • 163
Harness the Memory: A Holistic Evaluation of Memory Substrates in Memory Agents Paper • 2608.15008 • Published 10 days ago • 15
Oracle Agent Memory as an Enterprise Memory Substrate for Long-Horizon AI Agents Paper • 2607.13157 • Published Jul 14
What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents Paper • 2607.08032 • Published Jul 9
Memory as a Controlled Process: Learned Adaptive Memory Management for LLM Agents Paper • 2607.13591 • Published Jul 15
SkillForge: Self-Distilling Agents for Project-Specific Issue Resolution Paper • 2608.18933 • Published 6 days ago • 11
Know Before Fix: QA-Driven Repository Knowledge Acquisition for Software Issue Resolution Paper • 2607.11111 • Published Jul 13 • 25
FastContext: Training Efficient Repository Explorer for Coding Agents Paper • 2606.14066 • Published Jun 12 • 96
StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows Paper • 2608.17800 • Published 7 days ago • 9
TUA-Bench: A Benchmark for General-Purpose Terminal-Use Agents Paper • 2606.28480 • Published Jun 26 • 48
Managing Procedural Memory in LLM Agents: Control, Adaptation, and Evaluation Paper • 2606.23127 • Published Jun 22 • 26
SkillAdaptor: Self-Adapting Skills for LLM Agents from Trajectories Paper • 2606.01311 • Published May 31 • 37
SkillX: Automatically Constructing Skill Knowledge Bases for Agents Paper • 2604.04804 • Published Apr 6 • 35
GUI vs. CLI: Execution Bottlenecks in Screen-Only and Skill-Mediated Computer-Use Agents Paper • 2606.24551 • Published Jun 22 • 28
TerminalWorld: Benchmarking Agents on Real-World Terminal Tasks Paper • 2605.22535 • Published May 21 • 11
Agent Skills Can Be Harmful: An Empirical Study of Skill-Induced Failures in LLM Agents Paper • 2608.11888 • Published 13 days ago • 1
Evo-Harness: Context-to-Harness Skill Compilation for Self-Evolving Agents Paper • 2608.15071 • Published 10 days ago
Linear Attention Architectures: Mechanisms, Trade-offs, and Cross-Layer Routing Paper • 2607.07953 • Published Jul 8 • 16
Efficient Reinforcement Learning for Long-Horizon Tool-Use Agentic Tasks Paper • 2608.10357 • Published 14 days ago
Harness-R1: Learning to Edit Executable Runtime Harnesses from Agent Failure Trajectories Paper • 2608.02276 • Published 22 days ago • 4