ounce100m-code / README.md
Cion-lab's picture
README: name the three kept repos; record what the pruning removed and why a fresh run is unaffected
f7968e1 verified
|
Raw History Blame Contribute Delete
7.32 kB
# ounce100m-code
Canonical mirror of the scripts for project `ounce100m`, so a Kaggle instance can fetch its own source
instead of having it pasted into a kernel. **Kept public deliberately: this is the training code, and it is
intended to be reused for a different model without reading the project's history.**
No credential ever appears in this repo. Anything that needs the Hugging Face token reads it at runtime
through `ounce100m_credentials.install()`, which fetches a private Kaggle dataset (D-006). The local
workspace's `MASTER_PROMPT.md` is the constitution and is **never** uploaded here.
## Layout
| path | what it is |
|---|---|
| `train/train_ounce100m.py` | the trainer: model build, fp16 autocast + fp32 master + GradScaler, DDP, cursor/checkpoint save, resume, `--dry-run` |
| `train/hubckpt.py` | the storage cycle: push → **verify bytes** → roll pointer → read pointer back → `prune_verified` (refuses to delete unless verification passed) |
| `train/shard_dataset.py` | **a library, not a script**: the map-style reader that makes resumption an index advance instead of a replay, plus `cursor_dict` / `save_cursor` / `load_cursor` / `perm_sha` / `mix_fingerprint`. The trainer imports it; tokenising and sharding is done by `build/build_mix.py` |
| `train/preflight.py` | Gate 3 stages, incl. the throughput cells that decide a config |
| `build/build_mix.py` | the corpus builder: sources, dedup, exclusion masks, held-out `val/` split |
| `build/audit_contamination.py` | the mechanical 13-gram overlap audit (counts only; **validation/dev only**, never test) |
| `build/publish_mix.py`, `build/verify_mix.py`, `build/hubsync.py` | dataset publication, anonymous re-hash verification, workspace↔Hub sync |
| `build/publish_model.py` | Phase 5: model + tokenizer + generated card, then `--verify-only` = the Gate 5 clean room |
| `eval/run_benchmarks.py` | Phase 6: lm-eval driver — smoke pass first, per-task incremental save, push after every task |
| `kernels/phase4_session.py` | the session planner: reads the pointer, tiles steps onto the session/quota grid, emits the `torchrun` line |
| `kernels/phase4_bootstrap.py` | 21-line GPU kernel body: fetch launcher at a pinned commit sha, sha256-assert, exec |
| `kernels/p4_stop_probe.py`, `kernels/phase4_prep.py`, `kernels/p4_plan_sweep.py` | the pre-launch evidence chain (GPU smoke/soak probe, free-CPU rehearsal, quota sweep) |
| `kernels/p6_evals_bootstrap.py` | benchmark launcher **with a heartbeat**: pushes its own log to the Hub every 120 s, so "queued / running / dead" is distinguishable without kernel-log access |
| `probes/` | Phase 0-2 platform measurements: GPU/DDP behaviour, filesystem, ingest rate, credential paths, push bandwidth, source inventory |
| `config/param_count.py` | parameter arithmetic used to prove the model is inside its allowed range |
## Training a new model
The trainer takes its geometry from flags, so a new model is normally a flag change and a new mix — not a new
script. Everything below ran on Kaggle 2×T4; nothing here is Kaggle-specific except `kernels/*`, which exist
because a kernel cannot be passed environment variables.
```bash
# 1. data. Assemble with headroom above the token target (dedup, masks and val/ all subtract),
# then audit and publish. Run each with --help for the full flag set.
python build/build_mix.py --root <mixroot> --budget-tokens 1310000000 # sources are the SOURCES table
python build/audit_contamination.py # 13-gram counts vs validation/dev only
python build/publish_mix.py --root <mixroot> --repo <you>/my-mix --dry-run
python build/verify_mix.py # anonymous re-hash against the manifest
# 2. architecture + schedule: the trainer's geometry flags, all defaulting to ounce100m's frozen values
torchrun --nproc_per_node=2 train/train_ounce100m.py \
--layers 22 --hidden 576 --seq-len 1024 --tokens 1000000000 \
--micro-batch 4 --accum 32 --lr 6e-4 --warmup-frac 0.02 --decay-frac 0.8 \
--attn eager --optim adamw_torch --no-grad-ckpt \
--root <mixroot> --hub-repo <you>/my-ckpt --push-every-steps 127 \
--max-steps 3814 --seed 20260919 --data-seed 20260919
```
Flags worth knowing about: `--dry-run` (build the model, print parameter counts, exit — no data, no GPU
needed), `--resume auto` (read the Hub pointer; never falls back to step 0 if the pointer exists and the
bytes do not), `--stop-after-steps N` (end a segment **on a Hub-verified checkpoint**, which is what makes a
capped session safe), `--prune` (delete local only after byte verification), `--grad-ckpt`, `--torch-compile`.
`train/hubckpt.py` and `train/preflight.py` are independently runnable.
### The four invariants that make a resumable run
Breaking any of them loses work silently, so they are asserted in code, not documented as convention:
1. **A sample index is a pure function.** `sample k` depends only on `(dataset fingerprint, shuffle_seed, k)`,
so the cursor is a position, not a state. Persisted with every checkpoint: `seq_len`, `samples_consumed`,
`shuffle_seed`, `shuffle_perm_sha`, `dataset_files_sha`, `step`, `tokens_consumed`. A resume **raises** if a
hash differs — it never re-shuffles quietly.
2. **Write → push → verify bytes → point → read pointer back → prune.** `prune_verified()` takes the
verification result and refuses to delete otherwise. Listing is not proof; an HTTP 200 on an LFS pointer is
not proof.
3. **Cadence divides the session length.** With a 127-step cadence and 762-step sessions, an interruption
costs ≤45 min. The trainer rejects a cadence that is not a whole multiple of its checkpoint grid.
4. **Cold resume is the only resume test that counts.** Delete the local run dir, start a fresh instance, and
recover from the Hub alone.
## Reproducing a run exactly
Every kernel fetches its script **at a commit sha and asserts the sha256 after download**, so a re-run names
the code it used. The `ounce100m` run is at REV `c237b478` (trainer `dcb0ac19…`, `hubckpt` `d4b50ed0…`,
launcher `5fb14bc3…`); Phase 5 published from REV `55f23ef5`. The kept repos are exactly three:
`Cion-lab/ounce100m-v1` (model, tokenizer, card, `cursor.json`, `final/run_summary.json`),
`Cion-lab/ounce100m-mix-v1` (mix, manifest, `audit.json`, exclusion masks) and this code repo. The run's
31 intermediate checkpoints, the preflight probe repo and the audit's cached reference tables were deleted on
2026-09-21 once their findings were recorded: `docs/04-run-log.md` carries the per-checkpoint loss history,
`docs/03-preflight-report.md` the probe verdicts, and `audit.json` the contamination counts. Re-running the
audit only needs the public benchmark datasets again, and a **new** run recreates its own checkpoint repo
(`hubckpt` calls `create_repo(exist_ok=True)`; a missing `latest.json` reads as `not_found` → step 0).
The operational lessons — the evidence ladder, the checkpoint ordering, the contamination rules, the quota
arithmetic, and 55 numbered failures — are written up in `guide/` in this repo and summarised in
`docs/final-report.md`, also in this repo: https://huggingface.co/Cion-lab/ounce100m-code
Nothing in this repository depends on a file that lives only on one machine.