# ounce100m-code Canonical mirror of the scripts for project `ounce100m`, so a Kaggle instance can fetch its own source instead of having it pasted into a kernel. **Kept public deliberately: this is the training code, and it is intended to be reused for a different model without reading the project's history.** No credential ever appears in this repo. Anything that needs the Hugging Face token reads it at runtime through `ounce100m_credentials.install()`, which fetches a private Kaggle dataset (D-006). The local workspace's `MASTER_PROMPT.md` is the constitution and is **never** uploaded here. ## Layout | path | what it is | |---|---| | `train/train_ounce100m.py` | the trainer: model build, fp16 autocast + fp32 master + GradScaler, DDP, cursor/checkpoint save, resume, `--dry-run` | | `train/hubckpt.py` | the storage cycle: push → **verify bytes** → roll pointer → read pointer back → `prune_verified` (refuses to delete unless verification passed) | | `train/shard_dataset.py` | **a library, not a script**: the map-style reader that makes resumption an index advance instead of a replay, plus `cursor_dict` / `save_cursor` / `load_cursor` / `perm_sha` / `mix_fingerprint`. The trainer imports it; tokenising and sharding is done by `build/build_mix.py` | | `train/preflight.py` | Gate 3 stages, incl. the throughput cells that decide a config | | `build/build_mix.py` | the corpus builder: sources, dedup, exclusion masks, held-out `val/` split | | `build/audit_contamination.py` | the mechanical 13-gram overlap audit (counts only; **validation/dev only**, never test) | | `build/publish_mix.py`, `build/verify_mix.py`, `build/hubsync.py` | dataset publication, anonymous re-hash verification, workspace↔Hub sync | | `build/publish_model.py` | Phase 5: model + tokenizer + generated card, then `--verify-only` = the Gate 5 clean room | | `eval/run_benchmarks.py` | Phase 6: lm-eval driver — smoke pass first, per-task incremental save, push after every task | | `kernels/phase4_session.py` | the session planner: reads the pointer, tiles steps onto the session/quota grid, emits the `torchrun` line | | `kernels/phase4_bootstrap.py` | 21-line GPU kernel body: fetch launcher at a pinned commit sha, sha256-assert, exec | | `kernels/p4_stop_probe.py`, `kernels/phase4_prep.py`, `kernels/p4_plan_sweep.py` | the pre-launch evidence chain (GPU smoke/soak probe, free-CPU rehearsal, quota sweep) | | `kernels/p6_evals_bootstrap.py` | benchmark launcher **with a heartbeat**: pushes its own log to the Hub every 120 s, so "queued / running / dead" is distinguishable without kernel-log access | | `probes/` | Phase 0-2 platform measurements: GPU/DDP behaviour, filesystem, ingest rate, credential paths, push bandwidth, source inventory | | `config/param_count.py` | parameter arithmetic used to prove the model is inside its allowed range | ## Training a new model The trainer takes its geometry from flags, so a new model is normally a flag change and a new mix — not a new script. Everything below ran on Kaggle 2×T4; nothing here is Kaggle-specific except `kernels/*`, which exist because a kernel cannot be passed environment variables. ```bash # 1. data. Assemble with headroom above the token target (dedup, masks and val/ all subtract), # then audit and publish. Run each with --help for the full flag set. python build/build_mix.py --root --budget-tokens 1310000000 # sources are the SOURCES table python build/audit_contamination.py # 13-gram counts vs validation/dev only python build/publish_mix.py --root --repo /my-mix --dry-run python build/verify_mix.py # anonymous re-hash against the manifest # 2. architecture + schedule: the trainer's geometry flags, all defaulting to ounce100m's frozen values torchrun --nproc_per_node=2 train/train_ounce100m.py \ --layers 22 --hidden 576 --seq-len 1024 --tokens 1000000000 \ --micro-batch 4 --accum 32 --lr 6e-4 --warmup-frac 0.02 --decay-frac 0.8 \ --attn eager --optim adamw_torch --no-grad-ckpt \ --root --hub-repo /my-ckpt --push-every-steps 127 \ --max-steps 3814 --seed 20260919 --data-seed 20260919 ``` Flags worth knowing about: `--dry-run` (build the model, print parameter counts, exit — no data, no GPU needed), `--resume auto` (read the Hub pointer; never falls back to step 0 if the pointer exists and the bytes do not), `--stop-after-steps N` (end a segment **on a Hub-verified checkpoint**, which is what makes a capped session safe), `--prune` (delete local only after byte verification), `--grad-ckpt`, `--torch-compile`. `train/hubckpt.py` and `train/preflight.py` are independently runnable. ### The four invariants that make a resumable run Breaking any of them loses work silently, so they are asserted in code, not documented as convention: 1. **A sample index is a pure function.** `sample k` depends only on `(dataset fingerprint, shuffle_seed, k)`, so the cursor is a position, not a state. Persisted with every checkpoint: `seq_len`, `samples_consumed`, `shuffle_seed`, `shuffle_perm_sha`, `dataset_files_sha`, `step`, `tokens_consumed`. A resume **raises** if a hash differs — it never re-shuffles quietly. 2. **Write → push → verify bytes → point → read pointer back → prune.** `prune_verified()` takes the verification result and refuses to delete otherwise. Listing is not proof; an HTTP 200 on an LFS pointer is not proof. 3. **Cadence divides the session length.** With a 127-step cadence and 762-step sessions, an interruption costs ≤45 min. The trainer rejects a cadence that is not a whole multiple of its checkpoint grid. 4. **Cold resume is the only resume test that counts.** Delete the local run dir, start a fresh instance, and recover from the Hub alone. ## Reproducing a run exactly Every kernel fetches its script **at a commit sha and asserts the sha256 after download**, so a re-run names the code it used. The `ounce100m` run is at REV `c237b478` (trainer `dcb0ac19…`, `hubckpt` `d4b50ed0…`, launcher `5fb14bc3…`); Phase 5 published from REV `55f23ef5`. The kept repos are exactly three: `Cion-lab/ounce100m-v1` (model, tokenizer, card, `cursor.json`, `final/run_summary.json`), `Cion-lab/ounce100m-mix-v1` (mix, manifest, `audit.json`, exclusion masks) and this code repo. The run's 31 intermediate checkpoints, the preflight probe repo and the audit's cached reference tables were deleted on 2026-09-21 once their findings were recorded: `docs/04-run-log.md` carries the per-checkpoint loss history, `docs/03-preflight-report.md` the probe verdicts, and `audit.json` the contamination counts. Re-running the audit only needs the public benchmark datasets again, and a **new** run recreates its own checkpoint repo (`hubckpt` calls `create_repo(exist_ok=True)`; a missing `latest.json` reads as `not_found` → step 0). The operational lessons — the evidence ladder, the checkpoint ordering, the contamination rules, the quota arithmetic, and 55 numbered failures — are written up in `guide/` in this repo and summarised in `docs/final-report.md`, also in this repo: https://huggingface.co/Cion-lab/ounce100m-code Nothing in this repository depends on a file that lives only on one machine.