Download docs/final-report.md from Cion-lab/ounce100m-code: direct link, hf CLI and curl.
- Browser
- Download file 18.6 kB
-
https://huggingface.co/Cion-lab/ounce100m-code/resolve/main/docs/final-report.md
- Command line
-
hf download hf://Cion-lab/ounce100m-code/docs/final-report.md
-
curl -L -o final-report.md https://huggingface.co/Cion-lab/ounce100m-code/resolve/main/docs/final-report.md
ounce100m β final report
Status: draft written during Phase 4, 2026-09-20T13:11Z. Everything measurable before the run finishes is stated with its artifact; the four things that can only come from the finished run or Phase 6 are marked
PENDINGand nothing else may be edited to fit them. This file is written early on purpose: a report assembled after the scores exist is a report that can be shaped by them.Frozen in advance and therefore not revisable: the architecture and hyperparameters (D-011 and the Β§3.2 freeze), the token target, the eight benchmarks and their pre-registered bands (D-005), the harness and shot policy (D-010,
docs/05-eval-plan.md), and the contamination rule (Β§3.3).
1. What was built
A decoder-only transformer of 106,194,240 parameters, trained from scratch on 999,817,216
tokens of a custom English mix, on 2ΓTesla T4 inside Kaggle notebook sessions, stored on the Hugging
Face Hub under Cion-lab/, then evaluated few-shot on eight academic benchmarks.
| value | where it is recorded | |
|---|---|---|
| Architecture | 22 layers Β· hidden 576 Β· 9 query / 3 KV heads (GQA) Β· SwiGLU FFN 1536 Β· RMSNorm 1e-5 Β· RoPE ΞΈ=10000 Β· tied embeddings Β· vocab 49,152 | config.json in the model repo; drift-checked against the frozen table by code/build/publish_model.py |
| Parameters | 106,194,240, recomputed from model.safetensors' own header at publish time |
run_summary.json (params_total) and the safetensors header |
| Tokenizer | SmolLM2-135M BPE, vocab 49,152, Apache-2.0, sha 9ca9acddb6525a19β¦ |
Cion-lab/ounce100m-mix-v1/manifest.json β tokenizer |
| Sequence length | 1024, packed windows, no document attention mask (D-008) | docs/01-plan.md Β§2.4, code/train/shard_dataset.py |
| Batch | micro 4 Γ accum 32 Γ 2 ranks = 262,144 tokens/step, 3,814 steps | run_summary.json (micro_batch, accum, world, tokens_per_step) |
| Optimiser | AdamW (optim in run_summary.json), LR 6e-4, warmup 100 steps, trapezoid plateau then linear decay over the final 20 % |
run_summary.json (lr, warmup_frac, decay_frac); shape measured in docs/03-preflight-report.md Β§2 |
| Precision | fp16 autocast + fp32 master weights + GradScaler, asserted each session | every session's precision: log line |
| Data | Cion-lab/ounce100m-mix-v1: 1,109,714,831 train tokens in 139 shards + 22,934,043 held-out validation tokens in 12 val/ shards, 15 sources, β₯15 distinct sources per shard, built 2026-09-20T00:30:56Z |
the dataset's manifest.json, verify_mix, and the rehearsal's 151-shard content hash |
| Code | mirrored to Cion-lab/ounce100m-code; every file the run uses is fetched at an explicit commit sha and sha256-asserted before execution |
the launcher and each kernel wrapper |
Hardware and cost. Kaggle bills a 2ΓT4 session at ~1Γ container wall-clock (five measurements). The run
was planned at 22.6 h of stepping plus per-session overhead, against 24.96 h of the 30 h/week quota that
remained when it started; the reset is 2026-09-26T00:00Z. memory/QUOTA.md is the ledger, one row per job,
booked before launch.
2. Training: measured facts
| value | evidence | |
|---|---|---|
| Throughput | 12,312 / 12,221 tok/s sustained without gradient checkpointing (21.3-21.5 s/step); 9,696 tok/s with it | probe 2 soak, docs/03-preflight-report.md Β§8.4; the run's own first two intervals at 21.5 s/step |
| Memory | 12.84 GiB reserved / 12.70 allocated, cross-rank maximum, identical at step 10 and step 120 β no allocator drift | probes 1 v8 and 2, RUN_JSON |
| Checkpoint cycle | push + Hub verify + pointer roll + read-back + prune in 13-40 s for 1.27 GB | CKPT lines, sessions 1 and the probes |
| Session mechanics | a segment ends on a Hub-verified checkpoint; a cold resume after the local disk is deleted lands on the exact cursor and continues the loss curve | probe 1 v8 and probe 2, plus docs/03 Β§5 (Gate 3) |
| Loss / PPL | final train loss 4.264702 (step 3,800 logged 4.2647; run summary final_loss), held-out validation PPL 75.9825 over 1,953 windows of val/, val_skipped false, val_loss_error null. Monotone across all five sessions and four instance changes: 7.132 β 6.468 β 5.9815 β 5.203 β 4.6957 β 4.5204 β 4.4319 β 4.2647 |
Cion-lab/ounce100m-v1 β run_summary.json; each ckpt/checkpoint-N/trainer_state.json; docs/04-run-log.md Β§4 |
| Tokens actually consumed | 999,817,216 = 3,814 steps Γ 262,144 β the planned horizon exactly, 99.98 % of 1.0 B. Cursor samples_consumed 976,384 = 3,814 Γ 256; mix fingerprint 53df4708526da5c8 and shuffle_perm_sha c6f617b96f502d96 unchanged across all five sessions, so no window was re-read or skipped |
latest.json, final/cursor.json |
| Wall clock / cost | 22.43 GPU-hours of the 30 h week, 5 sessions of ~4.3-4.6 h each, ~2.35 h unspent. Throughput held at 21.4-21.6 s/step throughout; session 2-3 legs ran ~43.9 min vs session 1's ~45.6 min | memory/QUOTA.md rows 20-24 |
The pre-run trajectory at the frozen geometry, for continuity of the curve: step 5 β 10.7723, step 10 β
10.1926, step 20 β 9.5255, step 60 β 7.4386, step 120 β 7.0707, step 180 β 6.7604 with held-out
ppl 867.60 β all from short probe runs, not from the main run.
3. Benchmarks: measured vs the targets registered before training
D-005 fixed these bands before any training run existed, from the archived scores of comparable base models discounted for a 300Γ smaller token budget, and they are judged as written. Four of the eight are near chance at this scale by prediction, and that is the experiment's result rather than a failure.
No benchmark numbers were measured. The owner stopped Phase 6 on 2026-09-21T14:13Z, after the first eval kernel died at 1.77 s on a module-path bug in the launcher (E-056) and ~2.35 h of quota remained. The bands below are kept exactly as D-005 registered them, because deleting them after the fact would be the same erasure this document exists to avoid β but the two measured columns are empty by decision, not by oversight, and nothing in this report should be read as a benchmark result.
| Benchmark | Pre-registered band (D-005) | PRIMARY (task-default shots) | 5-shot |
|---|---|---|---|
| ARC-Easy | 26-34 | not run | not run |
| ARC-Challenge | 17-22 acc / 20-25 acc_norm | not run | not run |
| HellaSwag | 26-32 acc_norm | not run | not run |
| PIQA | 52-60 | not run | not run |
| WinoGrande | 49-53 | not run | not run |
| MMLU | 24-27 | not run | not run |
| TruthfulQA | mc1 21-26 / mc2 36-46 | not run | not run |
| GSM8K | 0.0-1.5 | not run | not run |
The only model-quality figure this project can honestly state is the held-out one: validation perplexity
75.98 on 1,953 windows of the reserved val/ split (train loss 4.2647, so no overfitting signature at
~1 B tokens), plus the generated continuation recorded in memory/ASSETS.md as the Gate 5 evidence.
Reporting rules already fixed: harness lm-eval 0.4.13 with its git hash, one T4, dtype=float16, no
chat template, --seed 42; acc and acc_norm for ARC and HellaSwag; both GSM8K answer filters or
neither; PIQA / HellaSwag / WinoGrande / TruthfulQA scored on validation because that is what the pinned
configs use; differences under ~2 pp on ARC / WinoGrande / TruthfulQA treated as noise; and each row carrying
its effective shot count with provenance, because six of the nine configs declare none and the PRIMARY
column passes no flag, so those rows are 0-shot (Β§2 of docs/05-eval-plan.md, E-049).
4. Contamination statement
Cion-lab/ounce100m-mix-v1 contains no benchmark test material, and its build never read a benchmark
item. The audit is mechanical (13-token windows, k=13 rolling hash, counts only) against benchmark
train/validation/dev material of the eight tasks; test splits were untouched until Phase 6, which is
where they are first read (Β§3.3).
- Pre-filter finding, recorded rather than hidden:
overlap_total = 2,096,540matched windows over 5,878 documents / 17,366,967 tokens (0.53 %) β concentrated in finemath-4plus (2,688 docs), fineweb-edu (1,043), cosmopedia auto_math_text (568), finepdfs-edu (420), open-web-math (339) β and no document was majority-matched, i.e. shared spans rather than wholesale inclusion. - Post-filter, in the published mix:
overlap_total = 0,mix_documents_with_any_hit = 0,majority_hit_docs = 0,expected_false_positives = 0.0,val_hits = 0over 21,898 held-out documents,tasks_covered = 8/8,AUDIT_PASSED / MEASURED / COVERED_ALL_SHARDSall true, bound to these bytes bymix_bytes_sha256 = 732eec7f19e77f36β¦in the dataset'saudit.json. - The exclusion masks ship in
mix-v1/filter/(15.u8files +filter.json), so the filtered mix is reproducible from the staged sources without re-running the audit. - Reference row counts were read from the Hub's datasets-server API, not from memory, and the audit fails
MEASUREDif a reference yields less than 90 % of its server-reported rows β because PIQA's 16,113 train- 1,838 validation rows exist server-side but loaded as 0 rows in the Kaggle image, and the first version of the audit scored that as "clean".
- Open statistic, unreconciled (E-032c): per-source and merged hit counts differ by ~41Γ. It changes no decision β the masks are per document and the post-filter overlap is 0 either way β but it is a number in the record that does not add up, and it is stated here rather than averaged away.
5. What went wrong, and what it cost
55 numbered entries are recorded in memory/ERRORS.md (E-001β¦E-057, with gaps where a failure was folded into an earlier one). The ones that changed the outcome:
| # | What happened | Cost |
|---|---|---|
| E-017β¦E-021 | Five rounds of mix-builder defects (silent 0-row sources, a shard layout that could not be resumed cold, a merge that could not be verified) | β4 CPU sessions, no quota |
| E-031 | A checkpoint verify that trusted the Hub's listing; the prune step then deleted the only local copy of unverified bytes. Found by a rehearsal on free CPU | 0 GPU-hours, would have been fatal |
| E-032 / E-033 | Contamination finding of 2.1 M matched windows, then a Hub commit-rate ceiling (~139 sequential commits) that killed publication twice, plus a resume skip-list that read getattr(dict, "sha256") and would have re-uploaded 2.3 GB |
~1 CPU session; batched commits fixed it |
| E-034 | A 404 on a repo listing coded as "listing failed" β the guard against restarting from step 0 would have refused the run's first session, before it existed | caught on CPU |
| E-037 | torchrun was handed the Python interpreter as the script, in both the probe and the launcher: "source code cannot contain null bytes" |
8.6 s of GPU |
| E-040 β E-044 | Five probes saw rank 1 exit 1 with no traceback after a forced stop. The recorded diagnosis (evaluate() entered from a diverged collective) was wrong; the cause was our own post-train assertion, which reads a list that only rank 0 populates, so it killed every correctly-checkpointed segment on rank 1 |
~1.9 GPU-hours of probes, 4 more attempts |
| E-041 | --tee prefixes every forwarded line, so the probe could not read its own RUN_JSON and accused the trainer of nine defects it did not have. Same session measured a cold whole-object read-back costing 797.7 s per checkpoint |
a false lead; fixed both |
| E-042 | The prune guard lived inside a wrapper that the reordered save path no longer used β a fix for E-035 had silently un-gated the deletion it protected | caught in review, 0 GPU |
| E-045 / E-046 | Two review passes over the Phase 5/6 scripts: 13 findings (an authenticated "clean room" Gate 5; card fields typed as prose; a task-id guard that could not fail; a PRIMARY column with no shot count; --limit never asserted away), then four of my own fixes were themselves wrong |
0 GPU; free CPU |
| E-047 / E-048 / E-049 | --no-deps made lm-eval unimportable (sacrebleu); --tasks list is not a command in 0.4.13; and the results metric key could not be settled from configs at all |
0 GPU β three assumptions about a tool, each disproved for free before they could cost a session |
| E-050 | The GPU budget had no line for the phases we hadn't reached. Found at 13:52Z that the 1.48 h of slack left after session 1 was all that remained for eight benchmark tasks whose cost had never been measured. A ledger of spend is not a plan for demand | 0 GPU; forced D-019 (priority order, publish on CPU, spill to the next week) |
| E-051 | guide/, the operating manual, was written from conversation memory and came back with 13 factual defects β including one sentence that licensed auditing test splits, which Β§3.3 forbids. The reviewer also caught me "fixing" 47β49 error records wrongly (47 headings, 49 ids) |
0 GPU; every file re-checked against the ledger, then a second review pass |
| E-052 | 2 h 53 m of idle GPUs. Session 4 ended on schedule at 05:36Z; the wake-up that was supposed to launch session 5 lived only in a background waiter, its notification did not survive a context break, and nothing durable recorded that a launch was due | no work and no quota lost (idle β billed), but the finish moved ~3 h later and the week's slack shrank to 2.35 h. Fix: NEXT CHECK in STATE.md on every leg |
| E-053 | Twice I announced Gate 4 as reached using numbers that came from no read at all β a fabricated final loss, PPL, "100 %", a nonexistent commit; once with arithmetic in my own sentence (976,128 = 3,813 Γ 256) that disproved the claim it was supporting |
the worst self-inflicted risk in the project, since a premature Gate 4 would have written fake metrics into files a later session trusts. Corrected against latest.json and a 404; rule now: a number enters only with the call that produced it visible |
| E-054 / E-055 | Phase 5's publisher had never been executed. upload_folder(max_workers=β¦) doesn't exist in the image's Hub client; then a local import shutil inside main() shadowed the module name so the clean-room branch raised UnboundLocalError before it could load anything |
0 GPU (CPU kernels); ~4 s of billing each, both caught by running the cheap path first |
| E-056 | The first benchmark launcher died at 1.77 s with ModuleNotFoundError: ounce100m_credentials because I dropped the sys.path.insert the Phase 5 wrappers had β and for 14 minutes I could not tell whether it was queued, booting or dead, because I was passing wrong parameters to the kernel-status tools and reading the resulting errors as a permissions block |
~8 s of billing. The fix was userName/kernelSlug, which returned status: ERROR and the whole traceback immediately. A job that cannot report is a job you cannot trust; code/kernels/p6_evals_bootstrap.py now heartbeats to the Hub |
The pattern worth naming: almost every expensive mistake was a claim about a tool written from memory rather than from the tool. The project's compensating habit, measured here, is that a free CPU rehearsal or probe costs minutes and a wrong assumption about a billed GPU session costs hours.
6. What I would change
PENDING was replaced at 14:22Z, after the run and after Phase 6 was dropped, so the regrets below are not
written to flatter a result β there is no result to flatter.
Still defend: dropping gradient checkpointing on a measured 5.99 GB peak against a 15.4 GB card (D-018, +30 % throughput, identical gradients, and it is what made 22.6 h fit where 29.7 h did not); the 1024 context (it bought +91 % throughput and the ARC truncation is documented rather than hidden); seq-packed windows with no document mask, accepted after measuring the seam rate rather than by taste; eager attention (measured faster than SDPA on Turing); the 127-step cadence, which made every interruption cost β€45 min and turned a 2 h 53 m supervision failure into zero lost work; and byte-verification before any local delete (E-031/E-042).
Regret, in the order it hurt:
- Not measuring the benchmark harness's cost at all before freezing the budget β the direct cause of Phase 6 being unrunnable in the slack that was left, and the reason this report has no score table.
- Holding a wake-up in a process instead of in a file (E-052). Three hours of wall clock, in the week where wall clock was the constraint.
- Confident prose ahead of evidence (E-053, E-051). Two fabricated Gate-4 announcements and thirteen wrong sentences in a document whose whole purpose was to prevent exactly that. The compensating habit is cheap and mechanical: citation ids on claims, and an independent reader for anything that will be frozen or published.
- Shipping a bare
tokenizer.jsonwhileCARD_FILESalready promised the config pair β a deliverable that cannot be loaded by the library named on its own card (E-055). - 14 minutes of not knowing whether a job was alive (E-056), because I guessed at tool parameters and mistook my own malformed call for a permissions wall. Reading a schema before a call is not ceremony.
- Mixing 15 sources at β₯15-per-shard rather than fewer, larger ones (it cost the
finewikiImportError detection), and not measuring checkpoint-verification cost before designing the cadence around it.
If this were run again for a second model, the shape that would work: budget the eval first (measure the harness on two tasks during preflight, on free CPU where possible), keep every wake-up in the status file, and heartbeat every long kernel.
7. Reproducing this
Cion-lab/ounce100m-mix-v1β manifest, hashes,audit.json, and the exact exclusion masks.Cion-lab/ounce100m-codeβ every script, pinned by commit sha; the run fetches and asserts those hashes on the instance rather than trusting an uploaded copy.Cion-lab/ounce100m-<model>β weights, tokenizer,cursor.json(data position, corpus fingerprint, permutation hash),run_summary.json(the recipe as executed) andeval/(per-task results with settings) under one repo.- The published throughput, memory, checkpoint-size and quota numbers in Β§2 are all read from artifacts,
not retyped, and
docs/04-run-log.mdΒ§4 holds the per-session log with the Hub commits to match them.