Download guide/README.md from Cion-lab/ounce100m-code: direct link, hf CLI and curl.
- Browser
- Download file 2.77 kB
-
https://huggingface.co/Cion-lab/ounce100m-code/resolve/main/guide/README.md
- Command line
-
hf download hf://Cion-lab/ounce100m-code/guide/README.md
-
curl -L -o README.md https://huggingface.co/Cion-lab/ounce100m-code/resolve/main/guide/README.md
Training a small LLM from scratch — an operating guide
For the next agent session that has to build, train, publish and evaluate a ~100M-parameter language model on rented, capped, per-week GPU. It is the distilled version of one real project: 106,194,240 params, 999,817,216 tokens, 2×Tesla T4 on Kaggle, storage on the Hugging Face Hub, eight academic benchmarks, and a public report — plus 54 written-up failure records spanning E-001…E-056 that are the actual source of these rules.
Read in order once; then use 06-checklists.md as the working page.
| File | What it covers |
|---|---|
01-before-you-start.md |
The constitution: what to fix in writing, and why memory files are the real state |
02-data.md |
Mix design, tokenisation, sharding, the contamination audit, publishing reproducibly |
03-preflight.md |
The evidence ladder: free rehearsal → cheap smoke → measurement probe → run |
04-the-run.md |
Checkpoints, resume, session sizing, quota arithmetic, what to do when it breaks |
05-publish-and-evaluate.md |
Model card, clean-room load, benchmark protocol, honest reporting |
06-checklists.md |
Copy-paste gates for each phase, and the recurring failure patterns |
07-platform-notes.md |
Kaggle and HF mechanics measured rather than believed, with the numbers |
The five sentences that matter most
- Every claim about a tool is a hypothesis until you have executed it — the expensive mistakes in this project were all "documented behaviour" that the installed version did not have.
- Verify bytes, never listings. A commit that returned, a file that is listed, and bytes that download are three different facts, and only the third authorises deleting your copy.
- Design for interruption; it is not an exception. Every interval must end on a verified, pointed-at checkpoint on shared storage, and the only resume test that counts deletes the local disk first.
- Free compute is a strategy, not a fallback. A CPU rehearsal that costs nine minutes buys back a five-hour session; run everything that does not need the accelerator there, forever.
- Write the report before the results exist. Frozen choices, pre-registered targets and a contamination statement composed after you can see the scores are a different kind of document.
What "done" means
A public model repo that a stranger loads with from_pretrained from an empty cache; a dataset repo whose
manifest, hashes and exclusion masks let someone rebuild that exact mix; a benchmark table where every row
names its metric, its shot count and its split, and can be re-run from published config; and a report that
says what failed. Model quality is not the deliverable — a trustworthy number is.