Why Self-Training Helps and Hurts: Denoising vs. Signal Forgetting
Paper • 2602.14029 • Published
Small CNN (MediumCNN(4 conv blocks 64/128/256/512 + GAP, ~4.69M params); not ResNet-50) trained as iterate t=0 in a self-distillation trajectory reproducing the denoising-vs-forgetting trade-off from Wu, Yang & Sun, Why Self-Training Helps and Hurts (arXiv:2602.14029).
This is a budget-conscious qualitative reproduction (NOT paper-scale ResNet-50). The full trajectory exhibits a U-shaped test-error curve (denoise then forget). Full per-iterate trajectory CSV: see evalstate/synthetic-selftrain-denoising-forgetting.