ounce100m-code / guide /README.md
Cion-lab's picture
publish guide/: the operating manual distilled from 54 failure records, for future agent sessions
6302710 verified
|
Raw History Blame Contribute Delete
2.77 kB
# Training a small LLM from scratch — an operating guide
For the next agent session that has to build, train, publish and evaluate a ~100M-parameter language model
on rented, capped, per-week GPU. It is the distilled version of one real project: 106,194,240 params,
999,817,216 tokens, 2×Tesla T4 on Kaggle, storage on the Hugging Face Hub, eight academic benchmarks, and a
public report — plus 54 written-up failure records spanning E-001…E-056 that are the actual source of these rules.
Read in order once; then use `06-checklists.md` as the working page.
| File | What it covers |
|---|---|
| `01-before-you-start.md` | The constitution: what to fix in writing, and why memory files are the real state |
| `02-data.md` | Mix design, tokenisation, sharding, the contamination audit, publishing reproducibly |
| `03-preflight.md` | The evidence ladder: free rehearsal → cheap smoke → measurement probe → run |
| `04-the-run.md` | Checkpoints, resume, session sizing, quota arithmetic, what to do when it breaks |
| `05-publish-and-evaluate.md` | Model card, clean-room load, benchmark protocol, honest reporting |
| `06-checklists.md` | Copy-paste gates for each phase, and the recurring failure patterns |
| `07-platform-notes.md` | Kaggle and HF mechanics measured rather than believed, with the numbers |
## The five sentences that matter most
1. **Every claim about a tool is a hypothesis until you have executed it** — the expensive mistakes in this
project were all "documented behaviour" that the installed version did not have.
2. **Verify bytes, never listings.** A commit that returned, a file that is listed, and bytes that
download are three different facts, and only the third authorises deleting your copy.
3. **Design for interruption; it is not an exception.** Every interval must end on a verified, pointed-at
checkpoint on shared storage, and the only resume test that counts deletes the local disk first.
4. **Free compute is a strategy, not a fallback.** A CPU rehearsal that costs nine minutes buys back a
five-hour session; run everything that does not need the accelerator there, forever.
5. **Write the report before the results exist.** Frozen choices, pre-registered targets and a contamination
statement composed after you can see the scores are a different kind of document.
## What "done" means
A public model repo that a stranger loads with `from_pretrained` from an empty cache; a dataset repo whose
manifest, hashes and exclusion masks let someone rebuild that exact mix; a benchmark table where every row
names its metric, its shot count and its split, and can be re-run from published config; and a report that
says what failed. Model quality is not the deliverable — a trustworthy number is.