ounce100m-code / docs /final-report.md
Cion-lab's picture
final report: 55 error records
9a70e6c verified
|
Raw History Blame Contribute Delete
18.6 kB

ounce100m β€” final report

Status: draft written during Phase 4, 2026-09-20T13:11Z. Everything measurable before the run finishes is stated with its artifact; the four things that can only come from the finished run or Phase 6 are marked PENDING and nothing else may be edited to fit them. This file is written early on purpose: a report assembled after the scores exist is a report that can be shaped by them.

Frozen in advance and therefore not revisable: the architecture and hyperparameters (D-011 and the Β§3.2 freeze), the token target, the eight benchmarks and their pre-registered bands (D-005), the harness and shot policy (D-010, docs/05-eval-plan.md), and the contamination rule (Β§3.3).

1. What was built

A decoder-only transformer of 106,194,240 parameters, trained from scratch on 999,817,216 tokens of a custom English mix, on 2Γ—Tesla T4 inside Kaggle notebook sessions, stored on the Hugging Face Hub under Cion-lab/, then evaluated few-shot on eight academic benchmarks.

value where it is recorded
Architecture 22 layers Β· hidden 576 Β· 9 query / 3 KV heads (GQA) Β· SwiGLU FFN 1536 Β· RMSNorm 1e-5 Β· RoPE ΞΈ=10000 Β· tied embeddings Β· vocab 49,152 config.json in the model repo; drift-checked against the frozen table by code/build/publish_model.py
Parameters 106,194,240, recomputed from model.safetensors' own header at publish time run_summary.json (params_total) and the safetensors header
Tokenizer SmolLM2-135M BPE, vocab 49,152, Apache-2.0, sha 9ca9acddb6525a19… Cion-lab/ounce100m-mix-v1/manifest.json β†’ tokenizer
Sequence length 1024, packed windows, no document attention mask (D-008) docs/01-plan.md Β§2.4, code/train/shard_dataset.py
Batch micro 4 Γ— accum 32 Γ— 2 ranks = 262,144 tokens/step, 3,814 steps run_summary.json (micro_batch, accum, world, tokens_per_step)
Optimiser AdamW (optim in run_summary.json), LR 6e-4, warmup 100 steps, trapezoid plateau then linear decay over the final 20 % run_summary.json (lr, warmup_frac, decay_frac); shape measured in docs/03-preflight-report.md Β§2
Precision fp16 autocast + fp32 master weights + GradScaler, asserted each session every session's precision: log line
Data Cion-lab/ounce100m-mix-v1: 1,109,714,831 train tokens in 139 shards + 22,934,043 held-out validation tokens in 12 val/ shards, 15 sources, β‰₯15 distinct sources per shard, built 2026-09-20T00:30:56Z the dataset's manifest.json, verify_mix, and the rehearsal's 151-shard content hash
Code mirrored to Cion-lab/ounce100m-code; every file the run uses is fetched at an explicit commit sha and sha256-asserted before execution the launcher and each kernel wrapper

Hardware and cost. Kaggle bills a 2Γ—T4 session at ~1Γ— container wall-clock (five measurements). The run was planned at 22.6 h of stepping plus per-session overhead, against 24.96 h of the 30 h/week quota that remained when it started; the reset is 2026-09-26T00:00Z. memory/QUOTA.md is the ledger, one row per job, booked before launch.

2. Training: measured facts

value evidence
Throughput 12,312 / 12,221 tok/s sustained without gradient checkpointing (21.3-21.5 s/step); 9,696 tok/s with it probe 2 soak, docs/03-preflight-report.md Β§8.4; the run's own first two intervals at 21.5 s/step
Memory 12.84 GiB reserved / 12.70 allocated, cross-rank maximum, identical at step 10 and step 120 β€” no allocator drift probes 1 v8 and 2, RUN_JSON
Checkpoint cycle push + Hub verify + pointer roll + read-back + prune in 13-40 s for 1.27 GB CKPT lines, sessions 1 and the probes
Session mechanics a segment ends on a Hub-verified checkpoint; a cold resume after the local disk is deleted lands on the exact cursor and continues the loss curve probe 1 v8 and probe 2, plus docs/03 Β§5 (Gate 3)
Loss / PPL final train loss 4.264702 (step 3,800 logged 4.2647; run summary final_loss), held-out validation PPL 75.9825 over 1,953 windows of val/, val_skipped false, val_loss_error null. Monotone across all five sessions and four instance changes: 7.132 β†’ 6.468 β†’ 5.9815 β†’ 5.203 β†’ 4.6957 β†’ 4.5204 β†’ 4.4319 β†’ 4.2647 Cion-lab/ounce100m-v1 β†’ run_summary.json; each ckpt/checkpoint-N/trainer_state.json; docs/04-run-log.md Β§4
Tokens actually consumed 999,817,216 = 3,814 steps Γ— 262,144 β€” the planned horizon exactly, 99.98 % of 1.0 B. Cursor samples_consumed 976,384 = 3,814 Γ— 256; mix fingerprint 53df4708526da5c8 and shuffle_perm_sha c6f617b96f502d96 unchanged across all five sessions, so no window was re-read or skipped latest.json, final/cursor.json
Wall clock / cost 22.43 GPU-hours of the 30 h week, 5 sessions of ~4.3-4.6 h each, ~2.35 h unspent. Throughput held at 21.4-21.6 s/step throughout; session 2-3 legs ran ~43.9 min vs session 1's ~45.6 min memory/QUOTA.md rows 20-24

The pre-run trajectory at the frozen geometry, for continuity of the curve: step 5 β†’ 10.7723, step 10 β†’ 10.1926, step 20 β†’ 9.5255, step 60 β†’ 7.4386, step 120 β†’ 7.0707, step 180 β†’ 6.7604 with held-out ppl 867.60 β€” all from short probe runs, not from the main run.

3. Benchmarks: measured vs the targets registered before training

D-005 fixed these bands before any training run existed, from the archived scores of comparable base models discounted for a 300Γ— smaller token budget, and they are judged as written. Four of the eight are near chance at this scale by prediction, and that is the experiment's result rather than a failure.

No benchmark numbers were measured. The owner stopped Phase 6 on 2026-09-21T14:13Z, after the first eval kernel died at 1.77 s on a module-path bug in the launcher (E-056) and ~2.35 h of quota remained. The bands below are kept exactly as D-005 registered them, because deleting them after the fact would be the same erasure this document exists to avoid β€” but the two measured columns are empty by decision, not by oversight, and nothing in this report should be read as a benchmark result.

Benchmark Pre-registered band (D-005) PRIMARY (task-default shots) 5-shot
ARC-Easy 26-34 not run not run
ARC-Challenge 17-22 acc / 20-25 acc_norm not run not run
HellaSwag 26-32 acc_norm not run not run
PIQA 52-60 not run not run
WinoGrande 49-53 not run not run
MMLU 24-27 not run not run
TruthfulQA mc1 21-26 / mc2 36-46 not run not run
GSM8K 0.0-1.5 not run not run

The only model-quality figure this project can honestly state is the held-out one: validation perplexity 75.98 on 1,953 windows of the reserved val/ split (train loss 4.2647, so no overfitting signature at ~1 B tokens), plus the generated continuation recorded in memory/ASSETS.md as the Gate 5 evidence.

Reporting rules already fixed: harness lm-eval 0.4.13 with its git hash, one T4, dtype=float16, no chat template, --seed 42; acc and acc_norm for ARC and HellaSwag; both GSM8K answer filters or neither; PIQA / HellaSwag / WinoGrande / TruthfulQA scored on validation because that is what the pinned configs use; differences under ~2 pp on ARC / WinoGrande / TruthfulQA treated as noise; and each row carrying its effective shot count with provenance, because six of the nine configs declare none and the PRIMARY column passes no flag, so those rows are 0-shot (Β§2 of docs/05-eval-plan.md, E-049).

4. Contamination statement

Cion-lab/ounce100m-mix-v1 contains no benchmark test material, and its build never read a benchmark item. The audit is mechanical (13-token windows, k=13 rolling hash, counts only) against benchmark train/validation/dev material of the eight tasks; test splits were untouched until Phase 6, which is where they are first read (Β§3.3).

  • Pre-filter finding, recorded rather than hidden: overlap_total = 2,096,540 matched windows over 5,878 documents / 17,366,967 tokens (0.53 %) β€” concentrated in finemath-4plus (2,688 docs), fineweb-edu (1,043), cosmopedia auto_math_text (568), finepdfs-edu (420), open-web-math (339) β€” and no document was majority-matched, i.e. shared spans rather than wholesale inclusion.
  • Post-filter, in the published mix: overlap_total = 0, mix_documents_with_any_hit = 0, majority_hit_docs = 0, expected_false_positives = 0.0, val_hits = 0 over 21,898 held-out documents, tasks_covered = 8/8, AUDIT_PASSED / MEASURED / COVERED_ALL_SHARDS all true, bound to these bytes by mix_bytes_sha256 = 732eec7f19e77f36… in the dataset's audit.json.
  • The exclusion masks ship in mix-v1/filter/ (15 .u8 files + filter.json), so the filtered mix is reproducible from the staged sources without re-running the audit.
  • Reference row counts were read from the Hub's datasets-server API, not from memory, and the audit fails MEASURED if a reference yields less than 90 % of its server-reported rows β€” because PIQA's 16,113 train
    • 1,838 validation rows exist server-side but loaded as 0 rows in the Kaggle image, and the first version of the audit scored that as "clean".
  • Open statistic, unreconciled (E-032c): per-source and merged hit counts differ by ~41Γ—. It changes no decision β€” the masks are per document and the post-filter overlap is 0 either way β€” but it is a number in the record that does not add up, and it is stated here rather than averaged away.

5. What went wrong, and what it cost

55 numbered entries are recorded in memory/ERRORS.md (E-001…E-057, with gaps where a failure was folded into an earlier one). The ones that changed the outcome:

# What happened Cost
E-017…E-021 Five rounds of mix-builder defects (silent 0-row sources, a shard layout that could not be resumed cold, a merge that could not be verified) β‰ˆ4 CPU sessions, no quota
E-031 A checkpoint verify that trusted the Hub's listing; the prune step then deleted the only local copy of unverified bytes. Found by a rehearsal on free CPU 0 GPU-hours, would have been fatal
E-032 / E-033 Contamination finding of 2.1 M matched windows, then a Hub commit-rate ceiling (~139 sequential commits) that killed publication twice, plus a resume skip-list that read getattr(dict, "sha256") and would have re-uploaded 2.3 GB ~1 CPU session; batched commits fixed it
E-034 A 404 on a repo listing coded as "listing failed" β€” the guard against restarting from step 0 would have refused the run's first session, before it existed caught on CPU
E-037 torchrun was handed the Python interpreter as the script, in both the probe and the launcher: "source code cannot contain null bytes" 8.6 s of GPU
E-040 β†’ E-044 Five probes saw rank 1 exit 1 with no traceback after a forced stop. The recorded diagnosis (evaluate() entered from a diverged collective) was wrong; the cause was our own post-train assertion, which reads a list that only rank 0 populates, so it killed every correctly-checkpointed segment on rank 1 ~1.9 GPU-hours of probes, 4 more attempts
E-041 --tee prefixes every forwarded line, so the probe could not read its own RUN_JSON and accused the trainer of nine defects it did not have. Same session measured a cold whole-object read-back costing 797.7 s per checkpoint a false lead; fixed both
E-042 The prune guard lived inside a wrapper that the reordered save path no longer used β€” a fix for E-035 had silently un-gated the deletion it protected caught in review, 0 GPU
E-045 / E-046 Two review passes over the Phase 5/6 scripts: 13 findings (an authenticated "clean room" Gate 5; card fields typed as prose; a task-id guard that could not fail; a PRIMARY column with no shot count; --limit never asserted away), then four of my own fixes were themselves wrong 0 GPU; free CPU
E-047 / E-048 / E-049 --no-deps made lm-eval unimportable (sacrebleu); --tasks list is not a command in 0.4.13; and the results metric key could not be settled from configs at all 0 GPU β€” three assumptions about a tool, each disproved for free before they could cost a session
E-050 The GPU budget had no line for the phases we hadn't reached. Found at 13:52Z that the 1.48 h of slack left after session 1 was all that remained for eight benchmark tasks whose cost had never been measured. A ledger of spend is not a plan for demand 0 GPU; forced D-019 (priority order, publish on CPU, spill to the next week)
E-051 guide/, the operating manual, was written from conversation memory and came back with 13 factual defects β€” including one sentence that licensed auditing test splits, which Β§3.3 forbids. The reviewer also caught me "fixing" 47β†’49 error records wrongly (47 headings, 49 ids) 0 GPU; every file re-checked against the ledger, then a second review pass
E-052 2 h 53 m of idle GPUs. Session 4 ended on schedule at 05:36Z; the wake-up that was supposed to launch session 5 lived only in a background waiter, its notification did not survive a context break, and nothing durable recorded that a launch was due no work and no quota lost (idle β‰  billed), but the finish moved ~3 h later and the week's slack shrank to 2.35 h. Fix: NEXT CHECK in STATE.md on every leg
E-053 Twice I announced Gate 4 as reached using numbers that came from no read at all β€” a fabricated final loss, PPL, "100 %", a nonexistent commit; once with arithmetic in my own sentence (976,128 = 3,813 Γ— 256) that disproved the claim it was supporting the worst self-inflicted risk in the project, since a premature Gate 4 would have written fake metrics into files a later session trusts. Corrected against latest.json and a 404; rule now: a number enters only with the call that produced it visible
E-054 / E-055 Phase 5's publisher had never been executed. upload_folder(max_workers=…) doesn't exist in the image's Hub client; then a local import shutil inside main() shadowed the module name so the clean-room branch raised UnboundLocalError before it could load anything 0 GPU (CPU kernels); ~4 s of billing each, both caught by running the cheap path first
E-056 The first benchmark launcher died at 1.77 s with ModuleNotFoundError: ounce100m_credentials because I dropped the sys.path.insert the Phase 5 wrappers had β€” and for 14 minutes I could not tell whether it was queued, booting or dead, because I was passing wrong parameters to the kernel-status tools and reading the resulting errors as a permissions block ~8 s of billing. The fix was userName/kernelSlug, which returned status: ERROR and the whole traceback immediately. A job that cannot report is a job you cannot trust; code/kernels/p6_evals_bootstrap.py now heartbeats to the Hub

The pattern worth naming: almost every expensive mistake was a claim about a tool written from memory rather than from the tool. The project's compensating habit, measured here, is that a free CPU rehearsal or probe costs minutes and a wrong assumption about a billed GPU session costs hours.

6. What I would change

PENDING was replaced at 14:22Z, after the run and after Phase 6 was dropped, so the regrets below are not written to flatter a result β€” there is no result to flatter.

Still defend: dropping gradient checkpointing on a measured 5.99 GB peak against a 15.4 GB card (D-018, +30 % throughput, identical gradients, and it is what made 22.6 h fit where 29.7 h did not); the 1024 context (it bought +91 % throughput and the ARC truncation is documented rather than hidden); seq-packed windows with no document mask, accepted after measuring the seam rate rather than by taste; eager attention (measured faster than SDPA on Turing); the 127-step cadence, which made every interruption cost ≀45 min and turned a 2 h 53 m supervision failure into zero lost work; and byte-verification before any local delete (E-031/E-042).

Regret, in the order it hurt:

  1. Not measuring the benchmark harness's cost at all before freezing the budget β€” the direct cause of Phase 6 being unrunnable in the slack that was left, and the reason this report has no score table.
  2. Holding a wake-up in a process instead of in a file (E-052). Three hours of wall clock, in the week where wall clock was the constraint.
  3. Confident prose ahead of evidence (E-053, E-051). Two fabricated Gate-4 announcements and thirteen wrong sentences in a document whose whole purpose was to prevent exactly that. The compensating habit is cheap and mechanical: citation ids on claims, and an independent reader for anything that will be frozen or published.
  4. Shipping a bare tokenizer.json while CARD_FILES already promised the config pair β€” a deliverable that cannot be loaded by the library named on its own card (E-055).
  5. 14 minutes of not knowing whether a job was alive (E-056), because I guessed at tool parameters and mistook my own malformed call for a permissions wall. Reading a schema before a call is not ceremony.
  6. Mixing 15 sources at β‰₯15-per-shard rather than fewer, larger ones (it cost the finewiki ImportError detection), and not measuring checkpoint-verification cost before designing the cadence around it.

If this were run again for a second model, the shape that would work: budget the eval first (measure the harness on two tasks during preflight, on free CPU where possible), keep every wake-up in the status file, and heartbeat every long kernel.

7. Reproducing this

  1. Cion-lab/ounce100m-mix-v1 β€” manifest, hashes, audit.json, and the exact exclusion masks.
  2. Cion-lab/ounce100m-code β€” every script, pinned by commit sha; the run fetches and asserts those hashes on the instance rather than trusting an uploaded copy.
  3. Cion-lab/ounce100m-<model> β€” weights, tokenizer, cursor.json (data position, corpus fingerprint, permutation hash), run_summary.json (the recipe as executed) and eval/ (per-task results with settings) under one repo.
  4. The published throughput, memory, checkpoint-size and quota numbers in Β§2 are all read from artifacts, not retyped, and docs/04-run-log.md Β§4 holds the per-session log with the Hub commits to match them.