Checkpoint stores an untied lm_head: 154.9M params on disk vs 119.4M on the card
Hi! I verified this model's parameter claim against the actual artifact and found a small discrepancy worth flagging.
What the card says: ~119.4M params (weights tied) / ~120M.
What the checkpoint stores: I parsed the safetensors header of model.safetensors:
- 173 tensors, all F32, 154,875,776 stored elements
- Both
lm_head.weightandtransformer.wte.weightare present, each[50257, 704]= 35,380,928 elements - File size checks out exactly: 8 + 17,656 (header) + 154,875,776 Γ 4 = 619,520,768 bytes
So the checkpoint on disk holds 154.9M parameters, which is 29.6% more than the card's 119.4M. The gap is exactly one vocab Γ n_embd (35,380,928) β a redundant, untied language-model head.
Why the card is still right about the model: model.py does tie the weights (wte.weight = lm_head.weight), so the intended model is 119,494,848 params β 119.5M, matching the card. The checkpoint just carries a second, unused copy of the head.
Suggested fixes (either works):
- Drop the redundant
lm_headfrom the export β the file would go from ~620 MB to ~478 MB, and the on-disk count would match the card. - Or clarify on the card that the checkpoint stores an untied head (154.9M on disk) while the model is 119.4M when tied.
Happy to open a PR with the slimmed checkpoint if that's useful. No urgency β just a heads-up!
i dont give a fuck