YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
ounce100m-code
Canonical mirror of the scripts for project ounce100m, so a Kaggle instance can fetch its own source
instead of having it pasted into a kernel. Kept public deliberately: this is the training code, and it is
intended to be reused for a different model without reading the project's history.
No credential ever appears in this repo. Anything that needs the Hugging Face token reads it at runtime
through ounce100m_credentials.install(), which fetches a private Kaggle dataset (D-006). The local
workspace's MASTER_PROMPT.md is the constitution and is never uploaded here.
Layout
| path | what it is |
|---|---|
train/train_ounce100m.py |
the trainer: model build, fp16 autocast + fp32 master + GradScaler, DDP, cursor/checkpoint save, resume, --dry-run |
train/hubckpt.py |
the storage cycle: push → verify bytes → roll pointer → read pointer back → prune_verified (refuses to delete unless verification passed) |
train/shard_dataset.py |
a library, not a script: the map-style reader that makes resumption an index advance instead of a replay, plus cursor_dict / save_cursor / load_cursor / perm_sha / mix_fingerprint. The trainer imports it; tokenising and sharding is done by build/build_mix.py |
train/preflight.py |
Gate 3 stages, incl. the throughput cells that decide a config |
build/build_mix.py |
the corpus builder: sources, dedup, exclusion masks, held-out val/ split |
build/audit_contamination.py |
the mechanical 13-gram overlap audit (counts only; validation/dev only, never test) |
build/publish_mix.py, build/verify_mix.py, build/hubsync.py |
dataset publication, anonymous re-hash verification, workspace↔Hub sync |
build/publish_model.py |
Phase 5: model + tokenizer + generated card, then --verify-only = the Gate 5 clean room |
eval/run_benchmarks.py |
Phase 6: lm-eval driver — smoke pass first, per-task incremental save, push after every task |
kernels/phase4_session.py |
the session planner: reads the pointer, tiles steps onto the session/quota grid, emits the torchrun line |
kernels/phase4_bootstrap.py |
21-line GPU kernel body: fetch launcher at a pinned commit sha, sha256-assert, exec |
kernels/p4_stop_probe.py, kernels/phase4_prep.py, kernels/p4_plan_sweep.py |
the pre-launch evidence chain (GPU smoke/soak probe, free-CPU rehearsal, quota sweep) |
kernels/p6_evals_bootstrap.py |
benchmark launcher with a heartbeat: pushes its own log to the Hub every 120 s, so "queued / running / dead" is distinguishable without kernel-log access |
probes/ |
Phase 0-2 platform measurements: GPU/DDP behaviour, filesystem, ingest rate, credential paths, push bandwidth, source inventory |
config/param_count.py |
parameter arithmetic used to prove the model is inside its allowed range |
Training a new model
The trainer takes its geometry from flags, so a new model is normally a flag change and a new mix — not a new
script. Everything below ran on Kaggle 2×T4; nothing here is Kaggle-specific except kernels/*, which exist
because a kernel cannot be passed environment variables.
# 1. data. Assemble with headroom above the token target (dedup, masks and val/ all subtract),
# then audit and publish. Run each with --help for the full flag set.
python build/build_mix.py --root <mixroot> --budget-tokens 1310000000 # sources are the SOURCES table
python build/audit_contamination.py # 13-gram counts vs validation/dev only
python build/publish_mix.py --root <mixroot> --repo <you>/my-mix --dry-run
python build/verify_mix.py # anonymous re-hash against the manifest
# 2. architecture + schedule: the trainer's geometry flags, all defaulting to ounce100m's frozen values
torchrun --nproc_per_node=2 train/train_ounce100m.py \
--layers 22 --hidden 576 --seq-len 1024 --tokens 1000000000 \
--micro-batch 4 --accum 32 --lr 6e-4 --warmup-frac 0.02 --decay-frac 0.8 \
--attn eager --optim adamw_torch --no-grad-ckpt \
--root <mixroot> --hub-repo <you>/my-ckpt --push-every-steps 127 \
--max-steps 3814 --seed 20260919 --data-seed 20260919
Flags worth knowing about: --dry-run (build the model, print parameter counts, exit — no data, no GPU
needed), --resume auto (read the Hub pointer; never falls back to step 0 if the pointer exists and the
bytes do not), --stop-after-steps N (end a segment on a Hub-verified checkpoint, which is what makes a
capped session safe), --prune (delete local only after byte verification), --grad-ckpt, --torch-compile.
train/hubckpt.py and train/preflight.py are independently runnable.
The four invariants that make a resumable run
Breaking any of them loses work silently, so they are asserted in code, not documented as convention:
- A sample index is a pure function.
sample kdepends only on(dataset fingerprint, shuffle_seed, k), so the cursor is a position, not a state. Persisted with every checkpoint:seq_len,samples_consumed,shuffle_seed,shuffle_perm_sha,dataset_files_sha,step,tokens_consumed. A resume raises if a hash differs — it never re-shuffles quietly. - Write → push → verify bytes → point → read pointer back → prune.
prune_verified()takes the verification result and refuses to delete otherwise. Listing is not proof; an HTTP 200 on an LFS pointer is not proof. - Cadence divides the session length. With a 127-step cadence and 762-step sessions, an interruption costs ≤45 min. The trainer rejects a cadence that is not a whole multiple of its checkpoint grid.
- Cold resume is the only resume test that counts. Delete the local run dir, start a fresh instance, and recover from the Hub alone.
Reproducing a run exactly
Every kernel fetches its script at a commit sha and asserts the sha256 after download, so a re-run names
the code it used. The ounce100m run is at REV c237b478 (trainer dcb0ac19…, hubckpt d4b50ed0…,
launcher 5fb14bc3…); Phase 5 published from REV 55f23ef5. The kept repos are exactly three:
Cion-lab/ounce100m-v1 (model, tokenizer, card, cursor.json, final/run_summary.json),
Cion-lab/ounce100m-mix-v1 (mix, manifest, audit.json, exclusion masks) and this code repo. The run's
31 intermediate checkpoints, the preflight probe repo and the audit's cached reference tables were deleted on
2026-09-21 once their findings were recorded: docs/04-run-log.md carries the per-checkpoint loss history,
docs/03-preflight-report.md the probe verdicts, and audit.json the contamination counts. Re-running the
audit only needs the public benchmark datasets again, and a new run recreates its own checkpoint repo
(hubckpt calls create_repo(exist_ok=True); a missing latest.json reads as not_found → step 0).
The operational lessons — the evidence ladder, the checkpoint ordering, the contamination rules, the quota
arithmetic, and 55 numbered failures — are written up in guide/ in this repo and summarised in
docs/final-report.md, also in this repo: https://huggingface.co/Cion-lab/ounce100m-code
Nothing in this repository depends on a file that lives only on one machine.