# Training a small LLM from scratch — an operating guide For the next agent session that has to build, train, publish and evaluate a ~100M-parameter language model on rented, capped, per-week GPU. It is the distilled version of one real project: 106,194,240 params, 999,817,216 tokens, 2×Tesla T4 on Kaggle, storage on the Hugging Face Hub, eight academic benchmarks, and a public report — plus 54 written-up failure records spanning E-001…E-056 that are the actual source of these rules. Read in order once; then use `06-checklists.md` as the working page. | File | What it covers | |---|---| | `01-before-you-start.md` | The constitution: what to fix in writing, and why memory files are the real state | | `02-data.md` | Mix design, tokenisation, sharding, the contamination audit, publishing reproducibly | | `03-preflight.md` | The evidence ladder: free rehearsal → cheap smoke → measurement probe → run | | `04-the-run.md` | Checkpoints, resume, session sizing, quota arithmetic, what to do when it breaks | | `05-publish-and-evaluate.md` | Model card, clean-room load, benchmark protocol, honest reporting | | `06-checklists.md` | Copy-paste gates for each phase, and the recurring failure patterns | | `07-platform-notes.md` | Kaggle and HF mechanics measured rather than believed, with the numbers | ## The five sentences that matter most 1. **Every claim about a tool is a hypothesis until you have executed it** — the expensive mistakes in this project were all "documented behaviour" that the installed version did not have. 2. **Verify bytes, never listings.** A commit that returned, a file that is listed, and bytes that download are three different facts, and only the third authorises deleting your copy. 3. **Design for interruption; it is not an exception.** Every interval must end on a verified, pointed-at checkpoint on shared storage, and the only resume test that counts deletes the local disk first. 4. **Free compute is a strategy, not a fallback.** A CPU rehearsal that costs nine minutes buys back a five-hour session; run everything that does not need the accelerator there, forever. 5. **Write the report before the results exist.** Frozen choices, pre-registered targets and a contamination statement composed after you can see the scores are a different kind of document. ## What "done" means A public model repo that a stranger loads with `from_pretrained` from an empty cache; a dataset repo whose manifest, hashes and exclusion masks let someone rebuild that exact mix; a benchmark table where every row names its metric, its shot count and its split, and can be re-run from published config; and a report that says what failed. Model quality is not the deliverable — a trustworthy number is.