|
Download README.md from Cion-lab/ounce100m-code: direct link, hf CLI and curl.
- Browser
- Download file 7.32 kB
-
https://huggingface.co/Cion-lab/ounce100m-code/resolve/main/README.md
- Command line
-
hf download hf://Cion-lab/ounce100m-code/README.md
-
curl -L -o README.md https://huggingface.co/Cion-lab/ounce100m-code/resolve/main/README.md
7.32 kB
| # ounce100m-code | |
| Canonical mirror of the scripts for project `ounce100m`, so a Kaggle instance can fetch its own source | |
| instead of having it pasted into a kernel. **Kept public deliberately: this is the training code, and it is | |
| intended to be reused for a different model without reading the project's history.** | |
| No credential ever appears in this repo. Anything that needs the Hugging Face token reads it at runtime | |
| through `ounce100m_credentials.install()`, which fetches a private Kaggle dataset (D-006). The local | |
| workspace's `MASTER_PROMPT.md` is the constitution and is **never** uploaded here. | |
| ## Layout | |
| | path | what it is | | |
| |---|---| | |
| | `train/train_ounce100m.py` | the trainer: model build, fp16 autocast + fp32 master + GradScaler, DDP, cursor/checkpoint save, resume, `--dry-run` | | |
| | `train/hubckpt.py` | the storage cycle: push → **verify bytes** → roll pointer → read pointer back → `prune_verified` (refuses to delete unless verification passed) | | |
| | `train/shard_dataset.py` | **a library, not a script**: the map-style reader that makes resumption an index advance instead of a replay, plus `cursor_dict` / `save_cursor` / `load_cursor` / `perm_sha` / `mix_fingerprint`. The trainer imports it; tokenising and sharding is done by `build/build_mix.py` | | |
| | `train/preflight.py` | Gate 3 stages, incl. the throughput cells that decide a config | | |
| | `build/build_mix.py` | the corpus builder: sources, dedup, exclusion masks, held-out `val/` split | | |
| | `build/audit_contamination.py` | the mechanical 13-gram overlap audit (counts only; **validation/dev only**, never test) | | |
| | `build/publish_mix.py`, `build/verify_mix.py`, `build/hubsync.py` | dataset publication, anonymous re-hash verification, workspace↔Hub sync | | |
| | `build/publish_model.py` | Phase 5: model + tokenizer + generated card, then `--verify-only` = the Gate 5 clean room | | |
| | `eval/run_benchmarks.py` | Phase 6: lm-eval driver — smoke pass first, per-task incremental save, push after every task | | |
| | `kernels/phase4_session.py` | the session planner: reads the pointer, tiles steps onto the session/quota grid, emits the `torchrun` line | | |
| | `kernels/phase4_bootstrap.py` | 21-line GPU kernel body: fetch launcher at a pinned commit sha, sha256-assert, exec | | |
| | `kernels/p4_stop_probe.py`, `kernels/phase4_prep.py`, `kernels/p4_plan_sweep.py` | the pre-launch evidence chain (GPU smoke/soak probe, free-CPU rehearsal, quota sweep) | | |
| | `kernels/p6_evals_bootstrap.py` | benchmark launcher **with a heartbeat**: pushes its own log to the Hub every 120 s, so "queued / running / dead" is distinguishable without kernel-log access | | |
| | `probes/` | Phase 0-2 platform measurements: GPU/DDP behaviour, filesystem, ingest rate, credential paths, push bandwidth, source inventory | | |
| | `config/param_count.py` | parameter arithmetic used to prove the model is inside its allowed range | | |
| ## Training a new model | |
| The trainer takes its geometry from flags, so a new model is normally a flag change and a new mix — not a new | |
| script. Everything below ran on Kaggle 2×T4; nothing here is Kaggle-specific except `kernels/*`, which exist | |
| because a kernel cannot be passed environment variables. | |
| ```bash | |
| # 1. data. Assemble with headroom above the token target (dedup, masks and val/ all subtract), | |
| # then audit and publish. Run each with --help for the full flag set. | |
| python build/build_mix.py --root <mixroot> --budget-tokens 1310000000 # sources are the SOURCES table | |
| python build/audit_contamination.py # 13-gram counts vs validation/dev only | |
| python build/publish_mix.py --root <mixroot> --repo <you>/my-mix --dry-run | |
| python build/verify_mix.py # anonymous re-hash against the manifest | |
| # 2. architecture + schedule: the trainer's geometry flags, all defaulting to ounce100m's frozen values | |
| torchrun --nproc_per_node=2 train/train_ounce100m.py \ | |
| --layers 22 --hidden 576 --seq-len 1024 --tokens 1000000000 \ | |
| --micro-batch 4 --accum 32 --lr 6e-4 --warmup-frac 0.02 --decay-frac 0.8 \ | |
| --attn eager --optim adamw_torch --no-grad-ckpt \ | |
| --root <mixroot> --hub-repo <you>/my-ckpt --push-every-steps 127 \ | |
| --max-steps 3814 --seed 20260919 --data-seed 20260919 | |
| ``` | |
| Flags worth knowing about: `--dry-run` (build the model, print parameter counts, exit — no data, no GPU | |
| needed), `--resume auto` (read the Hub pointer; never falls back to step 0 if the pointer exists and the | |
| bytes do not), `--stop-after-steps N` (end a segment **on a Hub-verified checkpoint**, which is what makes a | |
| capped session safe), `--prune` (delete local only after byte verification), `--grad-ckpt`, `--torch-compile`. | |
| `train/hubckpt.py` and `train/preflight.py` are independently runnable. | |
| ### The four invariants that make a resumable run | |
| Breaking any of them loses work silently, so they are asserted in code, not documented as convention: | |
| 1. **A sample index is a pure function.** `sample k` depends only on `(dataset fingerprint, shuffle_seed, k)`, | |
| so the cursor is a position, not a state. Persisted with every checkpoint: `seq_len`, `samples_consumed`, | |
| `shuffle_seed`, `shuffle_perm_sha`, `dataset_files_sha`, `step`, `tokens_consumed`. A resume **raises** if a | |
| hash differs — it never re-shuffles quietly. | |
| 2. **Write → push → verify bytes → point → read pointer back → prune.** `prune_verified()` takes the | |
| verification result and refuses to delete otherwise. Listing is not proof; an HTTP 200 on an LFS pointer is | |
| not proof. | |
| 3. **Cadence divides the session length.** With a 127-step cadence and 762-step sessions, an interruption | |
| costs ≤45 min. The trainer rejects a cadence that is not a whole multiple of its checkpoint grid. | |
| 4. **Cold resume is the only resume test that counts.** Delete the local run dir, start a fresh instance, and | |
| recover from the Hub alone. | |
| ## Reproducing a run exactly | |
| Every kernel fetches its script **at a commit sha and asserts the sha256 after download**, so a re-run names | |
| the code it used. The `ounce100m` run is at REV `c237b478` (trainer `dcb0ac19…`, `hubckpt` `d4b50ed0…`, | |
| launcher `5fb14bc3…`); Phase 5 published from REV `55f23ef5`. The kept repos are exactly three: | |
| `Cion-lab/ounce100m-v1` (model, tokenizer, card, `cursor.json`, `final/run_summary.json`), | |
| `Cion-lab/ounce100m-mix-v1` (mix, manifest, `audit.json`, exclusion masks) and this code repo. The run's | |
| 31 intermediate checkpoints, the preflight probe repo and the audit's cached reference tables were deleted on | |
| 2026-09-21 once their findings were recorded: `docs/04-run-log.md` carries the per-checkpoint loss history, | |
| `docs/03-preflight-report.md` the probe verdicts, and `audit.json` the contamination counts. Re-running the | |
| audit only needs the public benchmark datasets again, and a **new** run recreates its own checkpoint repo | |
| (`hubckpt` calls `create_repo(exist_ok=True)`; a missing `latest.json` reads as `not_found` → step 0). | |
| The operational lessons — the evidence ladder, the checkpoint ordering, the contamination rules, the quota | |
| arithmetic, and 55 numbered failures — are written up in `guide/` in this repo and summarised in | |
| `docs/final-report.md`, also in this repo: https://huggingface.co/Cion-lab/ounce100m-code | |
| Nothing in this repository depends on a file that lives only on one machine. | |