YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

ounce100m-code

Canonical mirror of the scripts for project ounce100m, so a Kaggle instance can fetch its own source instead of having it pasted into a kernel. Kept public deliberately: this is the training code, and it is intended to be reused for a different model without reading the project's history.

No credential ever appears in this repo. Anything that needs the Hugging Face token reads it at runtime through ounce100m_credentials.install(), which fetches a private Kaggle dataset (D-006). The local workspace's MASTER_PROMPT.md is the constitution and is never uploaded here.

Layout

path what it is
train/train_ounce100m.py the trainer: model build, fp16 autocast + fp32 master + GradScaler, DDP, cursor/checkpoint save, resume, --dry-run
train/hubckpt.py the storage cycle: push → verify bytes → roll pointer → read pointer back → prune_verified (refuses to delete unless verification passed)
train/shard_dataset.py a library, not a script: the map-style reader that makes resumption an index advance instead of a replay, plus cursor_dict / save_cursor / load_cursor / perm_sha / mix_fingerprint. The trainer imports it; tokenising and sharding is done by build/build_mix.py
train/preflight.py Gate 3 stages, incl. the throughput cells that decide a config
build/build_mix.py the corpus builder: sources, dedup, exclusion masks, held-out val/ split
build/audit_contamination.py the mechanical 13-gram overlap audit (counts only; validation/dev only, never test)
build/publish_mix.py, build/verify_mix.py, build/hubsync.py dataset publication, anonymous re-hash verification, workspace↔Hub sync
build/publish_model.py Phase 5: model + tokenizer + generated card, then --verify-only = the Gate 5 clean room
eval/run_benchmarks.py Phase 6: lm-eval driver — smoke pass first, per-task incremental save, push after every task
kernels/phase4_session.py the session planner: reads the pointer, tiles steps onto the session/quota grid, emits the torchrun line
kernels/phase4_bootstrap.py 21-line GPU kernel body: fetch launcher at a pinned commit sha, sha256-assert, exec
kernels/p4_stop_probe.py, kernels/phase4_prep.py, kernels/p4_plan_sweep.py the pre-launch evidence chain (GPU smoke/soak probe, free-CPU rehearsal, quota sweep)
kernels/p6_evals_bootstrap.py benchmark launcher with a heartbeat: pushes its own log to the Hub every 120 s, so "queued / running / dead" is distinguishable without kernel-log access
probes/ Phase 0-2 platform measurements: GPU/DDP behaviour, filesystem, ingest rate, credential paths, push bandwidth, source inventory
config/param_count.py parameter arithmetic used to prove the model is inside its allowed range

Training a new model

The trainer takes its geometry from flags, so a new model is normally a flag change and a new mix — not a new script. Everything below ran on Kaggle 2×T4; nothing here is Kaggle-specific except kernels/*, which exist because a kernel cannot be passed environment variables.

# 1. data. Assemble with headroom above the token target (dedup, masks and val/ all subtract),
#    then audit and publish. Run each with --help for the full flag set.
python build/build_mix.py --root <mixroot> --budget-tokens 1310000000   # sources are the SOURCES table
python build/audit_contamination.py                                     # 13-gram counts vs validation/dev only
python build/publish_mix.py --root <mixroot> --repo <you>/my-mix --dry-run
python build/verify_mix.py                                              # anonymous re-hash against the manifest

# 2. architecture + schedule: the trainer's geometry flags, all defaulting to ounce100m's frozen values
torchrun --nproc_per_node=2 train/train_ounce100m.py \
  --layers 22 --hidden 576 --seq-len 1024 --tokens 1000000000 \
  --micro-batch 4 --accum 32 --lr 6e-4 --warmup-frac 0.02 --decay-frac 0.8 \
  --attn eager --optim adamw_torch --no-grad-ckpt \
  --root <mixroot> --hub-repo <you>/my-ckpt --push-every-steps 127 \
  --max-steps 3814 --seed 20260919 --data-seed 20260919

Flags worth knowing about: --dry-run (build the model, print parameter counts, exit — no data, no GPU needed), --resume auto (read the Hub pointer; never falls back to step 0 if the pointer exists and the bytes do not), --stop-after-steps N (end a segment on a Hub-verified checkpoint, which is what makes a capped session safe), --prune (delete local only after byte verification), --grad-ckpt, --torch-compile. train/hubckpt.py and train/preflight.py are independently runnable.

The four invariants that make a resumable run

Breaking any of them loses work silently, so they are asserted in code, not documented as convention:

  1. A sample index is a pure function. sample k depends only on (dataset fingerprint, shuffle_seed, k), so the cursor is a position, not a state. Persisted with every checkpoint: seq_len, samples_consumed, shuffle_seed, shuffle_perm_sha, dataset_files_sha, step, tokens_consumed. A resume raises if a hash differs — it never re-shuffles quietly.
  2. Write → push → verify bytes → point → read pointer back → prune. prune_verified() takes the verification result and refuses to delete otherwise. Listing is not proof; an HTTP 200 on an LFS pointer is not proof.
  3. Cadence divides the session length. With a 127-step cadence and 762-step sessions, an interruption costs ≤45 min. The trainer rejects a cadence that is not a whole multiple of its checkpoint grid.
  4. Cold resume is the only resume test that counts. Delete the local run dir, start a fresh instance, and recover from the Hub alone.

Reproducing a run exactly

Every kernel fetches its script at a commit sha and asserts the sha256 after download, so a re-run names the code it used. The ounce100m run is at REV c237b478 (trainer dcb0ac19…, hubckpt d4b50ed0…, launcher 5fb14bc3…); Phase 5 published from REV 55f23ef5. The kept repos are exactly three: Cion-lab/ounce100m-v1 (model, tokenizer, card, cursor.json, final/run_summary.json), Cion-lab/ounce100m-mix-v1 (mix, manifest, audit.json, exclusion masks) and this code repo. The run's 31 intermediate checkpoints, the preflight probe repo and the audit's cached reference tables were deleted on 2026-09-21 once their findings were recorded: docs/04-run-log.md carries the per-checkpoint loss history, docs/03-preflight-report.md the probe verdicts, and audit.json the contamination counts. Re-running the audit only needs the public benchmark datasets again, and a new run recreates its own checkpoint repo (hubckpt calls create_repo(exist_ok=True); a missing latest.json reads as not_found → step 0).

The operational lessons — the evidence ladder, the checkpoint ordering, the contamination rules, the quota arithmetic, and 55 numbered failures — are written up in guide/ in this repo and summarised in docs/final-report.md, also in this repo: https://huggingface.co/Cion-lab/ounce100m-code Nothing in this repository depends on a file that lives only on one machine.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support