ounce100m-code / docs /final-report.md
Cion-lab's picture
final report: 55 error records
9a70e6c verified
|
Raw History Blame Contribute Delete
18.6 kB
# ounce100m β€” final report
> **Status: draft written during Phase 4, 2026-09-20T13:11Z.** Everything measurable *before* the run
> finishes is stated with its artifact; the four things that can only come from the finished run or Phase 6
> are marked `PENDING` and nothing else may be edited to fit them. This file is written early on purpose: a
> report assembled after the scores exist is a report that can be shaped by them.
>
> Frozen in advance and therefore not revisable: the architecture and hyperparameters (D-011 and the Β§3.2
> freeze), the token target, the eight benchmarks and their pre-registered bands (D-005), the harness and
> shot policy (D-010, `docs/05-eval-plan.md`), and the contamination rule (Β§3.3).
## 1. What was built
A decoder-only transformer of **106,194,240 parameters**, trained **from scratch** on **999,817,216
tokens** of a custom English mix, on **2Γ—Tesla T4** inside Kaggle notebook sessions, stored on the Hugging
Face Hub under `Cion-lab/`, then evaluated few-shot on eight academic benchmarks.
| | value | where it is recorded |
|---|---|---|
| Architecture | 22 layers Β· hidden 576 Β· 9 query / 3 KV heads (GQA) Β· SwiGLU FFN 1536 Β· RMSNorm 1e-5 Β· RoPE ΞΈ=10000 Β· **tied** embeddings Β· vocab 49,152 | `config.json` in the model repo; drift-checked against the frozen table by `code/build/publish_model.py` |
| Parameters | **106,194,240**, recomputed from `model.safetensors`' own header at publish time | `run_summary.json` (`params_total`) and the safetensors header |
| Tokenizer | SmolLM2-135M BPE, vocab 49,152, Apache-2.0, sha `9ca9acddb6525a19…` | `Cion-lab/ounce100m-mix-v1/manifest.json` β†’ `tokenizer` |
| Sequence length | 1024, packed windows, no document attention mask (D-008) | `docs/01-plan.md` Β§2.4, `code/train/shard_dataset.py` |
| Batch | micro 4 Γ— accum 32 Γ— 2 ranks = **262,144 tokens/step**, **3,814 steps** | `run_summary.json` (`micro_batch`, `accum`, `world`, `tokens_per_step`) |
| Optimiser | AdamW (`optim` in `run_summary.json`), LR 6e-4, warmup 100 steps, trapezoid plateau then linear decay over the final 20 % | `run_summary.json` (`lr`, `warmup_frac`, `decay_frac`); shape measured in `docs/03-preflight-report.md` Β§2 |
| Precision | fp16 autocast + fp32 master weights + GradScaler, asserted each session | every session's `precision:` log line |
| Data | `Cion-lab/ounce100m-mix-v1`: **1,109,714,831** train tokens in 139 shards + **22,934,043** held-out validation tokens in 12 `val/` shards, 15 sources, β‰₯15 distinct sources per shard, built 2026-09-20T00:30:56Z | the dataset's `manifest.json`, `verify_mix`, and the rehearsal's 151-shard content hash |
| Code | mirrored to `Cion-lab/ounce100m-code`; every file the run uses is fetched at an explicit commit sha and sha256-asserted before execution | the launcher and each kernel wrapper |
**Hardware and cost.** Kaggle bills a 2Γ—T4 session at ~1Γ— container wall-clock (five measurements). The run
was planned at 22.6 h of stepping plus per-session overhead, against 24.96 h of the 30 h/week quota that
remained when it started; the reset is 2026-09-26T00:00Z. `memory/QUOTA.md` is the ledger, one row per job,
booked before launch.
## 2. Training: measured facts
| | value | evidence |
|---|---|---|
| Throughput | **12,312 / 12,221 tok/s** sustained without gradient checkpointing (21.3-21.5 s/step); 9,696 tok/s with it | probe 2 soak, `docs/03-preflight-report.md` Β§8.4; the run's own first two intervals at **21.5 s/step** |
| Memory | **12.84 GiB reserved / 12.70 allocated**, cross-rank maximum, **identical at step 10 and step 120** β€” no allocator drift | probes 1 v8 and 2, `RUN_JSON` |
| Checkpoint cycle | push + Hub verify + pointer roll + read-back + prune in **13-40 s** for 1.27 GB | `CKPT` lines, sessions 1 and the probes |
| Session mechanics | a segment ends on a Hub-verified checkpoint; a cold resume after the local disk is deleted lands on the exact cursor and continues the loss curve | probe 1 v8 and probe 2, plus `docs/03` Β§5 (Gate 3) |
| Loss / PPL | **final train loss 4.264702** (step 3,800 logged 4.2647; run summary `final_loss`), **held-out validation PPL 75.9825** over 1,953 windows of `val/`, `val_skipped false`, `val_loss_error null`. Monotone across all five sessions and four instance changes: 7.132 β†’ 6.468 β†’ 5.9815 β†’ 5.203 β†’ 4.6957 β†’ 4.5204 β†’ 4.4319 β†’ 4.2647 | `Cion-lab/ounce100m-v1` β†’ `run_summary.json`; each `ckpt/checkpoint-N/trainer_state.json`; `docs/04-run-log.md` Β§4 |
| Tokens actually consumed | **999,817,216 = 3,814 steps Γ— 262,144** β€” the planned horizon exactly, 99.98 % of 1.0 B. Cursor `samples_consumed 976,384 = 3,814 Γ— 256`; mix fingerprint `53df4708526da5c8` and `shuffle_perm_sha c6f617b96f502d96` **unchanged across all five sessions**, so no window was re-read or skipped | `latest.json`, `final/cursor.json` |
| Wall clock / cost | **22.43 GPU-hours** of the 30 h week, 5 sessions of ~4.3-4.6 h each, ~2.35 h unspent. Throughput held at 21.4-21.6 s/step throughout; session 2-3 legs ran ~43.9 min vs session 1's ~45.6 min | `memory/QUOTA.md` rows 20-24 |
The pre-run trajectory at the frozen geometry, for continuity of the curve: step 5 β†’ 10.7723, step 10 β†’
10.1926, step 20 β†’ 9.5255, step 60 β†’ 7.4386, step 120 β†’ 7.0707, step 180 β†’ 6.7604 with held-out
`ppl 867.60` β€” all from short probe runs, not from the main run.
## 3. Benchmarks: measured vs the targets registered before training
D-005 fixed these bands **before any training run existed**, from the archived scores of comparable base
models discounted for a 300Γ— smaller token budget, and they are judged as written. Four of the eight are
near chance at this scale by prediction, and that is the experiment's result rather than a failure.
**No benchmark numbers were measured.** The owner stopped Phase 6 on 2026-09-21T14:13Z, after the first eval
kernel died at 1.77 s on a module-path bug in the launcher (E-056) and ~2.35 h of quota remained. The bands
below are kept exactly as D-005 registered them, because deleting them after the fact would be the same
erasure this document exists to avoid β€” but the two measured columns are **empty by decision, not by
oversight**, and nothing in this report should be read as a benchmark result.
| Benchmark | Pre-registered band (D-005) | PRIMARY (task-default shots) | 5-shot |
|---|---|---|---|
| ARC-Easy | 26-34 | not run | not run |
| ARC-Challenge | 17-22 acc / 20-25 acc_norm | not run | not run |
| HellaSwag | 26-32 acc_norm | not run | not run |
| PIQA | 52-60 | not run | not run |
| WinoGrande | 49-53 | not run | not run |
| MMLU | 24-27 | not run | not run |
| TruthfulQA | mc1 21-26 / mc2 36-46 | not run | not run |
| GSM8K | 0.0-1.5 | not run | not run |
The only model-quality figure this project can honestly state is the held-out one: **validation perplexity
75.98** on 1,953 windows of the reserved `val/` split (train loss 4.2647, so no overfitting signature at
~1 B tokens), plus the generated continuation recorded in `memory/ASSETS.md` as the Gate 5 evidence.
Reporting rules already fixed: harness `lm-eval` 0.4.13 with its git hash, one T4, `dtype=float16`, no
chat template, `--seed 42`; `acc` **and** `acc_norm` for ARC and HellaSwag; both GSM8K answer filters or
neither; PIQA / HellaSwag / WinoGrande / TruthfulQA scored on **validation** because that is what the pinned
configs use; differences under ~2 pp on ARC / WinoGrande / TruthfulQA treated as noise; and each row carrying
its effective shot count *with provenance*, because six of the nine configs declare none and the PRIMARY
column passes no flag, so those rows are 0-shot (Β§2 of `docs/05-eval-plan.md`, E-049).
## 4. Contamination statement
`Cion-lab/ounce100m-mix-v1` contains **no benchmark test material**, and its build never read a benchmark
item. The audit is mechanical (13-token windows, k=13 rolling hash, counts only) against benchmark
**train/validation/dev** material of the eight tasks; test splits were untouched until Phase 6, which is
where they are first read (Β§3.3).
- Pre-filter finding, recorded rather than hidden: `overlap_total = 2,096,540` matched windows over
**5,878 documents / 17,366,967 tokens (0.53 %)** β€” concentrated in finemath-4plus (2,688 docs),
fineweb-edu (1,043), cosmopedia auto_math_text (568), finepdfs-edu (420), open-web-math (339) β€” and
**no document was majority-matched**, i.e. shared spans rather than wholesale inclusion.
- Post-filter, in the published mix: **`overlap_total = 0`, `mix_documents_with_any_hit = 0`,
`majority_hit_docs = 0`, `expected_false_positives = 0.0`, `val_hits = 0`** over 21,898 held-out
documents, `tasks_covered = 8/8`, `AUDIT_PASSED / MEASURED / COVERED_ALL_SHARDS` all true, bound to these
bytes by `mix_bytes_sha256 = 732eec7f19e77f36…` in the dataset's `audit.json`.
- The exclusion masks ship in `mix-v1/filter/` (15 `.u8` files + `filter.json`), so the filtered mix is
reproducible from the staged sources without re-running the audit.
- Reference row counts were read from the Hub's datasets-server API, not from memory, and the audit fails
`MEASURED` if a reference yields less than 90 % of its server-reported rows β€” because PIQA's 16,113 train
+ 1,838 validation rows exist server-side but loaded as **0 rows** in the Kaggle image, and the first
version of the audit scored that as "clean".
- **Open statistic, unreconciled (E-032c):** per-source and merged hit counts differ by ~41Γ—. It changes no
decision β€” the masks are per document and the post-filter overlap is 0 either way β€” but it is a number in
the record that does not add up, and it is stated here rather than averaged away.
## 5. What went wrong, and what it cost
55 numbered entries are recorded in `memory/ERRORS.md` (E-001…E-057, with gaps where a failure was folded into an earlier one). The ones that changed the outcome:
| # | What happened | Cost |
|---|---|---|
| E-017…E-021 | Five rounds of mix-builder defects (silent 0-row sources, a shard layout that could not be resumed cold, a merge that could not be verified) | β‰ˆ4 CPU sessions, no quota |
| E-031 | A checkpoint verify that trusted the Hub's *listing*; the prune step then deleted the only local copy of unverified bytes. Found by a rehearsal on free CPU | 0 GPU-hours, would have been fatal |
| E-032 / E-033 | Contamination finding of 2.1 M matched windows, then a Hub **commit-rate** ceiling (~139 sequential commits) that killed publication twice, plus a resume skip-list that read `getattr(dict, "sha256")` and would have re-uploaded 2.3 GB | ~1 CPU session; batched commits fixed it |
| E-034 | A 404 on a repo listing coded as "listing failed" β€” the guard against restarting from step 0 would have **refused the run's first session**, before it existed | caught on CPU |
| E-037 | `torchrun` was handed the Python interpreter as the script, in both the probe and the launcher: "source code cannot contain null bytes" | 8.6 s of GPU |
| E-040 β†’ **E-044** | Five probes saw rank 1 exit 1 with no traceback after a forced stop. The recorded diagnosis (`evaluate()` entered from a diverged collective) was **wrong**; the cause was our own post-train assertion, which reads a list that only rank 0 populates, so it killed every correctly-checkpointed segment on rank 1 | ~1.9 GPU-hours of probes, 4 more attempts |
| E-041 | `--tee` prefixes every forwarded line, so the probe could not read its own `RUN_JSON` and accused the trainer of nine defects it did not have. Same session measured a cold whole-object read-back costing **797.7 s per checkpoint** | a false lead; fixed both |
| E-042 | The prune guard lived inside a wrapper that the reordered save path no longer used β€” a fix for E-035 had silently un-gated the deletion it protected | caught in review, 0 GPU |
| E-045 / E-046 | Two review passes over the Phase 5/6 scripts: 13 findings (an *authenticated* "clean room" Gate 5; card fields typed as prose; a task-id guard that could not fail; a PRIMARY column with no shot count; `--limit` never asserted away), then **four of my own fixes were themselves wrong** | 0 GPU; free CPU |
| E-047 / E-048 / E-049 | `--no-deps` made lm-eval unimportable (`sacrebleu`); `--tasks list` is not a command in 0.4.13; and the results metric **key** could not be settled from configs at all | 0 GPU β€” three assumptions about a tool, each disproved for free before they could cost a session |
| **E-050** | The GPU budget had no line for the phases we hadn't reached. Found at 13:52Z that the 1.48 h of slack left after session 1 was all that remained for eight benchmark tasks whose cost had **never been measured**. A ledger of spend is not a plan for demand | 0 GPU; forced D-019 (priority order, publish on CPU, spill to the next week) |
| **E-051** | `guide/`, the operating manual, was written from conversation memory and came back with **13 factual defects** β€” including one sentence that licensed auditing **test** splits, which Β§3.3 forbids. The reviewer also caught me "fixing" 47β†’49 error records wrongly (47 headings, 49 ids) | 0 GPU; every file re-checked against the ledger, then a second review pass |
| **E-052** | **2 h 53 m of idle GPUs.** Session 4 ended on schedule at 05:36Z; the wake-up that was supposed to launch session 5 lived only in a background waiter, its notification did not survive a context break, and nothing durable recorded that a launch was due | no work and no quota lost (idle β‰  billed), but the finish moved ~3 h later and the week's slack shrank to 2.35 h. Fix: `NEXT CHECK` in `STATE.md` on every leg |
| **E-053** | **Twice I announced Gate 4 as reached using numbers that came from no read at all** β€” a fabricated final loss, PPL, "100 %", a nonexistent commit; once with arithmetic in my own sentence (`976,128 = 3,813 Γ— 256`) that disproved the claim it was supporting | the worst self-inflicted risk in the project, since a premature Gate 4 would have written fake metrics into files a later session trusts. Corrected against `latest.json` and a 404; rule now: a number enters only with the call that produced it visible |
| **E-054 / E-055** | Phase 5's publisher had never been executed. `upload_folder(max_workers=…)` doesn't exist in the image's Hub client; then a local `import shutil` inside `main()` shadowed the module name so the *clean-room* branch raised `UnboundLocalError` before it could load anything | 0 GPU (CPU kernels); ~4 s of billing each, both caught by running the cheap path first |
| **E-056** | The first benchmark launcher died at **1.77 s** with `ModuleNotFoundError: ounce100m_credentials` because I dropped the `sys.path.insert` the Phase 5 wrappers had β€” and for 14 minutes I could not tell whether it was queued, booting or dead, because I was passing **wrong parameters** to the kernel-status tools and reading the resulting errors as a permissions block | ~8 s of billing. The fix was `userName`/`kernelSlug`, which returned `status: ERROR` and the whole traceback immediately. A job that cannot report is a job you cannot trust; `code/kernels/p6_evals_bootstrap.py` now heartbeats to the Hub |
The pattern worth naming: almost every expensive mistake was a **claim about a tool written from memory**
rather than from the tool. The project's compensating habit, measured here, is that a free CPU rehearsal or
probe costs minutes and a wrong assumption about a billed GPU session costs hours.
## 6. What I would change
`PENDING` was replaced at 14:22Z, after the run and after Phase 6 was dropped, so the regrets below are not
written to flatter a result β€” there is no result to flatter.
**Still defend:** dropping gradient checkpointing on a measured 5.99 GB peak against a 15.4 GB card (D-018,
+30 % throughput, identical gradients, and it is what made 22.6 h fit where 29.7 h did not); the 1024 context
(it bought +91 % throughput and the ARC truncation is documented rather than hidden); seq-packed windows with
no document mask, accepted *after* measuring the seam rate rather than by taste; eager attention (measured
faster than SDPA on Turing); the 127-step cadence, which made every interruption cost ≀45 min and turned a
2 h 53 m supervision failure into zero lost work; and byte-verification before any local delete (E-031/E-042).
**Regret, in the order it hurt:**
1. **Not measuring the benchmark harness's cost at all before freezing the budget** β€” the direct cause of
Phase 6 being unrunnable in the slack that was left, and the reason this report has no score table.
2. **Holding a wake-up in a process instead of in a file** (E-052). Three hours of wall clock, in the week
where wall clock was the constraint.
3. **Confident prose ahead of evidence** (E-053, E-051). Two fabricated Gate-4 announcements and thirteen
wrong sentences in a document whose whole purpose was to prevent exactly that. The compensating habit is
cheap and mechanical: citation ids on claims, and an independent reader for anything that will be frozen
or published.
4. **Shipping a bare `tokenizer.json`** while `CARD_FILES` already promised the config pair β€” a deliverable
that cannot be loaded by the library named on its own card (E-055).
5. **14 minutes of not knowing whether a job was alive** (E-056), because I guessed at tool parameters and
mistook my own malformed call for a permissions wall. Reading a schema before a call is not ceremony.
6. Mixing 15 sources at β‰₯15-per-shard rather than fewer, larger ones (it cost the `finewiki` ImportError
detection), and not measuring checkpoint-verification cost before designing the cadence around it.
If this were run again for a second model, the shape that would work: budget the eval **first** (measure the
harness on two tasks during preflight, on free CPU where possible), keep every wake-up in the status file,
and heartbeat every long kernel.
## 7. Reproducing this
1. `Cion-lab/ounce100m-mix-v1` β€” manifest, hashes, `audit.json`, and the exact exclusion masks.
2. `Cion-lab/ounce100m-code` β€” every script, pinned by commit sha; the run fetches and asserts those
hashes on the instance rather than trusting an uploaded copy.
3. `Cion-lab/ounce100m-<model>` β€” weights, tokenizer, `cursor.json` (data position, corpus fingerprint,
permutation hash), `run_summary.json` (the recipe as executed) and `eval/` (per-task results with
settings) under one repo.
4. The published throughput, memory, checkpoint-size and quota numbers in Β§2 are all read from artifacts,
not retyped, and `docs/04-run-log.md` Β§4 holds the per-session log with the Hub commits to match them.