Papers
arxiv:2609.39306

ReSAIL: Mitigating Collapse in Iterative Agent Self-Distillation

Published on Sep 30
· Submitted by
Shengjie Jin
on Oct 8
Authors:
,
,
,

Abstract

Iterative self-distillation enables LLM agents to learn from successive deployments, offering a path toward recursive self-improvement (RSI). Yet our experiments with existing methods reveal a collapse in deployment performance across cycles, while task performance with privileged information (PI) also declines. We address this collapse by prioritizing informative interaction steps for distillation and preserving PI-conditioned behavior as the student becomes the next teacher. We introduce Retentive and Selective Augmentation for Iterative Self-Distillation (ReSAIL), a plug-in augmentation for iterative PI-based self-distillation. ReSAIL selects interaction steps where PI most strongly changes the teacher's predictions and balances the resulting distillation losses across trajectories. It also regularizes the student's PI-conditioned output distributions toward those of the frozen teacher at selected and unselected steps to preserve PI-conditioned behavior for supervision in the next cycle. On ALFWorld and TextCraft, ReSAIL sustains substantial gains across model scales over three cycles, with an average absolute gain of 22.5% in final-cycle success rates when added to self-distillation baselines. Sensitivity-guided selection of offline data also improves action prediction accuracy for multimodal GUI agents on AITZ. These findings provide the first evidence that a more robust learning mechanism can effectively mitigate performance collapse in iterative agent self-distillation over deployment trajectories.

Community

Paper author Paper submitter

Can agents keep improving by distilling from their own deployment experience? In our experiments, existing self-distillation methods can lose both deployment performance and task competence with privileged information (PI) across cycles. This matters because today's student becomes tomorrow's teacher.

ReSAIL is a plug-in augmentation with two complementary components:

  • Trajectory-balanced selective distillation: prioritize interaction steps where PI most changes the teacher's predictions, and balance the losses across trajectories.
  • Privileged retention: preserve the student's PI-conditioned behavior to support supervision in the next cycle.

With Qwen3-4B and Qwen3-8B, ReSAIL sustains gains over three deployment cycles on ALFWorld and TextCraft. At cycle 3, it improves success rates by 22.5 percentage points on average over the corresponding SDPO/OEL baselines, across two model scales and three evaluation settings (ALFWorld ID/OOD and TextCraft).

Code, training/evaluation scripts, and model checkpoints are publicly available. Happy to discuss the method, results, and limitations!

Sign up or log in to comment

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2609.39306 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2609.39306 in a dataset README.md to link it from this page.

Spaces citing this paper 1

Collections including this paper 1