ounce100m-code / guide /README.md
Cion-lab's picture
publish guide/: the operating manual distilled from 54 failure records, for future agent sessions
6302710 verified
|
Raw History Blame Contribute Delete
2.77 kB

Training a small LLM from scratch — an operating guide

For the next agent session that has to build, train, publish and evaluate a ~100M-parameter language model on rented, capped, per-week GPU. It is the distilled version of one real project: 106,194,240 params, 999,817,216 tokens, 2×Tesla T4 on Kaggle, storage on the Hugging Face Hub, eight academic benchmarks, and a public report — plus 54 written-up failure records spanning E-001…E-056 that are the actual source of these rules.

Read in order once; then use 06-checklists.md as the working page.

File What it covers
01-before-you-start.md The constitution: what to fix in writing, and why memory files are the real state
02-data.md Mix design, tokenisation, sharding, the contamination audit, publishing reproducibly
03-preflight.md The evidence ladder: free rehearsal → cheap smoke → measurement probe → run
04-the-run.md Checkpoints, resume, session sizing, quota arithmetic, what to do when it breaks
05-publish-and-evaluate.md Model card, clean-room load, benchmark protocol, honest reporting
06-checklists.md Copy-paste gates for each phase, and the recurring failure patterns
07-platform-notes.md Kaggle and HF mechanics measured rather than believed, with the numbers

The five sentences that matter most

  1. Every claim about a tool is a hypothesis until you have executed it — the expensive mistakes in this project were all "documented behaviour" that the installed version did not have.
  2. Verify bytes, never listings. A commit that returned, a file that is listed, and bytes that download are three different facts, and only the third authorises deleting your copy.
  3. Design for interruption; it is not an exception. Every interval must end on a verified, pointed-at checkpoint on shared storage, and the only resume test that counts deletes the local disk first.
  4. Free compute is a strategy, not a fallback. A CPU rehearsal that costs nine minutes buys back a five-hour session; run everything that does not need the accelerator there, forever.
  5. Write the report before the results exist. Frozen choices, pre-registered targets and a contamination statement composed after you can see the scores are a different kind of document.

What "done" means

A public model repo that a stranger loads with from_pretrained from an empty cache; a dataset repo whose manifest, hashes and exclusion masks let someone rebuild that exact mix; a benchmark table where every row names its metric, its shot count and its split, and can be re-run from published config; and a report that says what failed. Model quality is not the deliverable — a trustworthy number is.