Mossez-100M-Coder-Instruct / TRAINING_REPORT.md
mossez-systems's picture
Publish Mossez-100M-Coder-Instruct
8dc5e05 verified
|
Raw
History Blame Contribute Delete
1.39 kB

Training report

Lineage

  • Source weights: the selected one-epoch Mossez-100M-Coder-Base.
  • Source model SHA-256: aba529bf10ad9f3acb5294c8bc2b4c93d20d25c6cff3a235a8659503b9ac1837.
  • Final model SHA-256: 0aade7d070122633abccd70cc5500e5bc36c7f69fafeadd9f5ee0b5a3e0766bf.
  • The general Mossez-100M-Instruct supplied no source weights.

Run

  • Objective: assistant-only supervised fine-tuning.
  • One bounded epoch: 660 contiguous finite optimizer steps.
  • Examples: 2,640 train; 330 validation; 330 test.
  • Rendered training tokens: 323,960; assistant target tokens: 105,100.
  • Precision: BF16; fused AdamW; micro-batch 1; gradient accumulation 4.
  • Validation loss: 1.991049 at step 0 to 0.031888 at step 660.
  • Early stopping: not triggered; validation regression count: 0.
  • Final checkpoint: best validation checkpoint, step 660.

Preflight and integrity

  • Assistant-only masks were asserted at corpus load and before each forward pass.
  • A real CUDA memory probe passed through micro-batch 8; training used micro-batch 1.
  • A 10-step smoke passed, followed by a bit-exact cross-process resume test.
  • An independent 100-step pilot improved overall and all per-task held-out losses.
  • Corpus manifest SHA-256: aca68b911788a1a7d93671bc27835b3d279cd612c02e8e3add91e2d3b2079084.
  • Tokenizer manifest SHA-256: f24c2e10522e866a889177e3a2258d84f1d8a863d784894a72974e391c1abdec.