size-40k model variants (epoch ablation, eval_results §6.7)
All: latent=8, ce4.0/al2.0/emph2.0, eff batch 128. Weights only (optimizer omitted).
- (repo root model + checkpoint-250/500) = stage2 baseline (2 epoch)
- stage3/ = baseline stage3, s2=2ep / s3=2ep, ckpt-626 (the §6.7 baseline column; best size-40k model)
- stage2-ep1/ = s2e1: stage2 trained 1 epoch (ckpt-313)
- stage3-s2e1/checkpoint-313 = s2e1 stage3, s2=1ep / s3=1ep
- stage3-s2e1/checkpoint-626 = s2e1 stage3, s2=1ep / s3=2ep
- epoch1/ (already uploaded) = s3e1: s2=2ep / s3=1ep (stage3 1 epoch)