Self-Generated Feedback Destabilizes Test-Time Training: A Causal Decomposition of Long-Horizon Adaptation
Abstract
Test-time training (TTT) lets a model store information in its weights during inference. When the model learns from its own output, however, each update also changes the model that generates the next training example. Across 128K-token streams, retaining generated-text updates worsens prediction on independent human-written text with three TTT-E2E model configurations (labeled 125M, 760M, and 3B). The same failure occurs when Adam updates Qwen3-4B's existing weights. The same update mechanisms can improve on real text, so writing itself is not the failure. Three matched comparisons trace the causal pathway. Fixed Generation removes over 98% of the damage at 125M and 760M by using a frozen model to generate training chunks. Recorded Replay separates the loss caused by reading degraded text from the additional loss stored by updating on it. A paired one-update comparison then shows the local conflict: an update predicts its source better but new real text worse. This cost grows after Closed Loop adaptation, with a few trajectories accounting for most large failures. Finally, Settlement evaluates the candidate state on independent real text before commitment. It leaves mean endpoint gaps of 0.07 and -0.02 nats at 125M and 760M while retaining real-text adaptation. These results motivate checking prediction on independent evidence before retaining an update.
Community
Hi everyone! Author here 👋
What happens when a model keeps learning from its own outputs during inference?
We study this feedback loop over 128K-token streams with TTT-E2E models from 125M to 3B and Adam-based adaptation of Qwen3-4B.
Our key findings:
Feedback matters: using a frozen model to generate the training text removes over 98% of the damage at 125M and 760M, even though the learner still updates on generated text.
Better source fit can mean worse transfer: an update can improve prediction of its training passage while hurting independent real text.
Validate before committing: Settlement tests candidate updates on independent text before retaining them, substantially reducing damage while preserving real-text adaptation.
Code: https://github.com/lingjivoo/ttt-ouroboros
We’d love to hear your thoughts on how continuously adapting models should decide which updates to keep!
Get this paper in your agent:
hf papers read 2610.05076 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper