# Training report ## Lineage - Source weights: the selected one-epoch Mossez-100M-Coder-Base. - Source model SHA-256: `aba529bf10ad9f3acb5294c8bc2b4c93d20d25c6cff3a235a8659503b9ac1837`. - Final model SHA-256: `0aade7d070122633abccd70cc5500e5bc36c7f69fafeadd9f5ee0b5a3e0766bf`. - The general Mossez-100M-Instruct supplied no source weights. ## Run - Objective: assistant-only supervised fine-tuning. - One bounded epoch: 660 contiguous finite optimizer steps. - Examples: 2,640 train; 330 validation; 330 test. - Rendered training tokens: 323,960; assistant target tokens: 105,100. - Precision: BF16; fused AdamW; micro-batch 1; gradient accumulation 4. - Validation loss: 1.991049 at step 0 to 0.031888 at step 660. - Early stopping: not triggered; validation regression count: 0. - Final checkpoint: best validation checkpoint, step 660. ## Preflight and integrity - Assistant-only masks were asserted at corpus load and before each forward pass. - A real CUDA memory probe passed through micro-batch 8; training used micro-batch 1. - A 10-step smoke passed, followed by a bit-exact cross-process resume test. - An independent 100-step pilot improved overall and all per-task held-out losses. - Corpus manifest SHA-256: `aca68b911788a1a7d93671bc27835b3d279cd612c02e8e3add91e2d3b2079084`. - Tokenizer manifest SHA-256: `f24c2e10522e866a889177e3a2258d84f1d8a863d784894a72974e391c1abdec`.