Mossez-100M-Coder-Instruct / TRAINING_REPORT.md
mossez-systems's picture
Publish Mossez-100M-Coder-Instruct
8dc5e05 verified
|
Raw
History Blame Contribute Delete
1.39 kB
# Training report
## Lineage
- Source weights: the selected one-epoch Mossez-100M-Coder-Base.
- Source model SHA-256: `aba529bf10ad9f3acb5294c8bc2b4c93d20d25c6cff3a235a8659503b9ac1837`.
- Final model SHA-256: `0aade7d070122633abccd70cc5500e5bc36c7f69fafeadd9f5ee0b5a3e0766bf`.
- The general Mossez-100M-Instruct supplied no source weights.
## Run
- Objective: assistant-only supervised fine-tuning.
- One bounded epoch: 660 contiguous finite optimizer steps.
- Examples: 2,640 train; 330 validation; 330 test.
- Rendered training tokens: 323,960; assistant target tokens: 105,100.
- Precision: BF16; fused AdamW; micro-batch 1; gradient accumulation 4.
- Validation loss: 1.991049 at step 0 to 0.031888 at step 660.
- Early stopping: not triggered; validation regression count: 0.
- Final checkpoint: best validation checkpoint, step 660.
## Preflight and integrity
- Assistant-only masks were asserted at corpus load and before each forward pass.
- A real CUDA memory probe passed through micro-batch 8; training used micro-batch 1.
- A 10-step smoke passed, followed by a bit-exact cross-process resume test.
- An independent 100-step pilot improved overall and all per-task held-out losses.
- Corpus manifest SHA-256: `aca68b911788a1a7d93671bc27835b3d279cd612c02e8e3add91e2d3b2079084`.
- Tokenizer manifest SHA-256: `f24c2e10522e866a889177e3a2258d84f1d8a863d784894a72974e391c1abdec`.