Abstract
Iterative self-distillation enables LLM agents to learn from successive deployments, offering a path toward recursive self-improvement (RSI). Yet our experiments with existing methods reveal a collapse in deployment performance across cycles, while task performance with privileged information (PI) also declines. We address this collapse by prioritizing informative interaction steps for distillation and preserving PI-conditioned behavior as the student becomes the next teacher. We introduce Retentive and Selective Augmentation for Iterative Self-Distillation (ReSAIL), a plug-in augmentation for iterative PI-based self-distillation. ReSAIL selects interaction steps where PI most strongly changes the teacher's predictions and balances the resulting distillation losses across trajectories. It also regularizes the student's PI-conditioned output distributions toward those of the frozen teacher at selected and unselected steps to preserve PI-conditioned behavior for supervision in the next cycle. On ALFWorld and TextCraft, ReSAIL sustains substantial gains across model scales over three cycles, with an average absolute gain of 22.5% in final-cycle success rates when added to self-distillation baselines. Sensitivity-guided selection of offline data also improves action prediction accuracy for multimodal GUI agents on AITZ. These findings provide the first evidence that a more robust learning mechanism can effectively mitigate performance collapse in iterative agent self-distillation over deployment trajectories.
Community
Can agents keep improving by distilling from their own deployment experience? In our experiments, existing self-distillation methods can lose both deployment performance and task competence with privileged information (PI) across cycles. This matters because today's student becomes tomorrow's teacher.
ReSAIL is a plug-in augmentation with two complementary components:
- Trajectory-balanced selective distillation: prioritize interaction steps where PI most changes the teacher's predictions, and balance the losses across trajectories.
- Privileged retention: preserve the student's PI-conditioned behavior to support supervision in the next cycle.
With Qwen3-4B and Qwen3-8B, ReSAIL sustains gains over three deployment cycles on ALFWorld and TextCraft. At cycle 3, it improves success rates by 22.5 percentage points on average over the corresponding SDPO/OEL baselines, across two model scales and three evaluation settings (ALFWorld ID/OOD and TextCraft).
Code, training/evaluation scripts, and model checkpoints are publicly available. Happy to discuss the method, results, and limitations!
Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper