Abstract
AI agents increasingly operate in environments where they can diagnose failures and improve through experience, yet existing evaluations largely measure what an agent can do at a fixed point in time rather than how effectively it learns. Evaluating self-improvement requires answering three questions: does future performance improve and generalize beyond the interactions that enabled learning; how efficiently are new capabilities acquired; and where does the self-improvement process break down? To answer these questions, we study self-improvement in a controlled setting where agents amortize past experience into reusable artifacts that are inherited by future instances. At each checkpoint, we measure performance on training and held-out environment interactions while accounting for learning cost. We introduce agent plasticity, the efficiency with which an agent converts experience into gains in future held-out performance. Across multiple environments, frontier models exhibit sharply different improvement trajectories despite comparable opportunities to learn. Some achieve substantial and persistent gains, while others remain near or below their initial performance, and gains within the training regime often transfer only partially to out-of-distribution conditions. Endpoint capability and acquisition efficiency also diverge: the agent that ultimately performs best need not be the one that improves most efficiently. Tracing failures through the improvement loop further reveals different candidate bottlenecks. Agents with low plasticity often fail to reuse relevant artifacts, whereas more plastic agents may still fail despite reusing relevant artifacts, pointing to limitations in artifact quality, generalization, or application. Evaluating self-improving agents requires measuring not only what they can do, but how effectively they become better through experience.
Community
How well do AI agents actually learn from experience?
We study self-improving agents that turn past interactions into reusable artifacts, such as tools, skills, and memory, which are inherited by future instances.
We introduce Agent Plasticity, a measure of how efficiently an agent converts learning cost into improvements in future held-out performance.
Across Chess, Go, Hex, and NetHack, we find that:
- Frontier models exhibit very different improvement trajectories despite comparable opportunities to learn.
- The best-performing agent is not necessarily the most efficient learner.
- Improvements on the training distribution often transfer only partially out of distribution.
- Effective self-improvement depends not just on retrieving past experience and created text artifacts, but also on producing useful artifacts and applying them correctly.
Overall, our results suggest that evaluating self-improving agents requires measuring not only endpoint capability, but also how efficiently and reliably agents improve through experience.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development (2026)
- The Right Lesson at the Right Step: Deriving Control Updates for Self-Evolving Agents (2026)
- Demystifying Agent Skills: Why They Work-Until They Don't (2026)
- Self-Evolving Skills via Surrogate-Guided Solve-and-Reproduce (2026)
- Better, Faster, Stronger: Programmatic Skill Learning Best Reduces Agent Cost (2026)
- ServeLearnBench: How Well Can Agents Self-Improve from Serving Experience? (2026)
- Aspire: Can Models Self-Evolve from Vague Goals? (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2610.08902 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper