File size: 18,614 Bytes
f345921
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
9a70e6c
f345921
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
# ounce100m β€” final report

> **Status: draft written during Phase 4, 2026-09-20T13:11Z.** Everything measurable *before* the run
> finishes is stated with its artifact; the four things that can only come from the finished run or Phase 6
> are marked `PENDING` and nothing else may be edited to fit them. This file is written early on purpose: a
> report assembled after the scores exist is a report that can be shaped by them.
>
> Frozen in advance and therefore not revisable: the architecture and hyperparameters (D-011 and the Β§3.2
> freeze), the token target, the eight benchmarks and their pre-registered bands (D-005), the harness and
> shot policy (D-010, `docs/05-eval-plan.md`), and the contamination rule (Β§3.3).

## 1. What was built

A decoder-only transformer of **106,194,240 parameters**, trained **from scratch** on **999,817,216
tokens** of a custom English mix, on **2Γ—Tesla T4** inside Kaggle notebook sessions, stored on the Hugging
Face Hub under `Cion-lab/`, then evaluated few-shot on eight academic benchmarks.

| | value | where it is recorded |
|---|---|---|
| Architecture | 22 layers Β· hidden 576 Β· 9 query / 3 KV heads (GQA) Β· SwiGLU FFN 1536 Β· RMSNorm 1e-5 Β· RoPE ΞΈ=10000 Β· **tied** embeddings Β· vocab 49,152 | `config.json` in the model repo; drift-checked against the frozen table by `code/build/publish_model.py` |
| Parameters | **106,194,240**, recomputed from `model.safetensors`' own header at publish time | `run_summary.json` (`params_total`) and the safetensors header |
| Tokenizer | SmolLM2-135M BPE, vocab 49,152, Apache-2.0, sha `9ca9acddb6525a19…` | `Cion-lab/ounce100m-mix-v1/manifest.json` β†’ `tokenizer` |
| Sequence length | 1024, packed windows, no document attention mask (D-008) | `docs/01-plan.md` Β§2.4, `code/train/shard_dataset.py` |
| Batch | micro 4 Γ— accum 32 Γ— 2 ranks = **262,144 tokens/step**, **3,814 steps** | `run_summary.json` (`micro_batch`, `accum`, `world`, `tokens_per_step`) |
| Optimiser | AdamW (`optim` in `run_summary.json`), LR 6e-4, warmup 100 steps, trapezoid plateau then linear decay over the final 20 % | `run_summary.json` (`lr`, `warmup_frac`, `decay_frac`); shape measured in `docs/03-preflight-report.md` Β§2 |
| Precision | fp16 autocast + fp32 master weights + GradScaler, asserted each session | every session's `precision:` log line |
| Data | `Cion-lab/ounce100m-mix-v1`: **1,109,714,831** train tokens in 139 shards + **22,934,043** held-out validation tokens in 12 `val/` shards, 15 sources, β‰₯15 distinct sources per shard, built 2026-09-20T00:30:56Z | the dataset's `manifest.json`, `verify_mix`, and the rehearsal's 151-shard content hash |
| Code | mirrored to `Cion-lab/ounce100m-code`; every file the run uses is fetched at an explicit commit sha and sha256-asserted before execution | the launcher and each kernel wrapper |

**Hardware and cost.** Kaggle bills a 2Γ—T4 session at ~1Γ— container wall-clock (five measurements). The run
was planned at 22.6 h of stepping plus per-session overhead, against 24.96 h of the 30 h/week quota that
remained when it started; the reset is 2026-09-26T00:00Z. `memory/QUOTA.md` is the ledger, one row per job,
booked before launch.

## 2. Training: measured facts

| | value | evidence |
|---|---|---|
| Throughput | **12,312 / 12,221 tok/s** sustained without gradient checkpointing (21.3-21.5 s/step); 9,696 tok/s with it | probe 2 soak, `docs/03-preflight-report.md` Β§8.4; the run's own first two intervals at **21.5 s/step** |
| Memory | **12.84 GiB reserved / 12.70 allocated**, cross-rank maximum, **identical at step 10 and step 120** β€” no allocator drift | probes 1 v8 and 2, `RUN_JSON` |
| Checkpoint cycle | push + Hub verify + pointer roll + read-back + prune in **13-40 s** for 1.27 GB | `CKPT` lines, sessions 1 and the probes |
| Session mechanics | a segment ends on a Hub-verified checkpoint; a cold resume after the local disk is deleted lands on the exact cursor and continues the loss curve | probe 1 v8 and probe 2, plus `docs/03` Β§5 (Gate 3) |
| Loss / PPL | **final train loss 4.264702** (step 3,800 logged 4.2647; run summary `final_loss`), **held-out validation PPL 75.9825** over 1,953 windows of `val/`, `val_skipped false`, `val_loss_error null`. Monotone across all five sessions and four instance changes: 7.132 β†’ 6.468 β†’ 5.9815 β†’ 5.203 β†’ 4.6957 β†’ 4.5204 β†’ 4.4319 β†’ 4.2647 | `Cion-lab/ounce100m-v1` β†’ `run_summary.json`; each `ckpt/checkpoint-N/trainer_state.json`; `docs/04-run-log.md` Β§4 |
| Tokens actually consumed | **999,817,216 = 3,814 steps Γ— 262,144** β€” the planned horizon exactly, 99.98 % of 1.0 B. Cursor `samples_consumed 976,384 = 3,814 Γ— 256`; mix fingerprint `53df4708526da5c8` and `shuffle_perm_sha c6f617b96f502d96` **unchanged across all five sessions**, so no window was re-read or skipped | `latest.json`, `final/cursor.json` |
| Wall clock / cost | **22.43 GPU-hours** of the 30 h week, 5 sessions of ~4.3-4.6 h each, ~2.35 h unspent. Throughput held at 21.4-21.6 s/step throughout; session 2-3 legs ran ~43.9 min vs session 1's ~45.6 min | `memory/QUOTA.md` rows 20-24 |

The pre-run trajectory at the frozen geometry, for continuity of the curve: step 5 β†’ 10.7723, step 10 β†’
10.1926, step 20 β†’ 9.5255, step 60 β†’ 7.4386, step 120 β†’ 7.0707, step 180 β†’ 6.7604 with held-out
`ppl 867.60` β€” all from short probe runs, not from the main run.

## 3. Benchmarks: measured vs the targets registered before training

D-005 fixed these bands **before any training run existed**, from the archived scores of comparable base
models discounted for a 300Γ— smaller token budget, and they are judged as written. Four of the eight are
near chance at this scale by prediction, and that is the experiment's result rather than a failure.

**No benchmark numbers were measured.** The owner stopped Phase 6 on 2026-09-21T14:13Z, after the first eval
kernel died at 1.77 s on a module-path bug in the launcher (E-056) and ~2.35 h of quota remained. The bands
below are kept exactly as D-005 registered them, because deleting them after the fact would be the same
erasure this document exists to avoid β€” but the two measured columns are **empty by decision, not by
oversight**, and nothing in this report should be read as a benchmark result.

| Benchmark | Pre-registered band (D-005) | PRIMARY (task-default shots) | 5-shot |
|---|---|---|---|
| ARC-Easy | 26-34 | not run | not run |
| ARC-Challenge | 17-22 acc / 20-25 acc_norm | not run | not run |
| HellaSwag | 26-32 acc_norm | not run | not run |
| PIQA | 52-60 | not run | not run |
| WinoGrande | 49-53 | not run | not run |
| MMLU | 24-27 | not run | not run |
| TruthfulQA | mc1 21-26 / mc2 36-46 | not run | not run |
| GSM8K | 0.0-1.5 | not run | not run |

The only model-quality figure this project can honestly state is the held-out one: **validation perplexity
75.98** on 1,953 windows of the reserved `val/` split (train loss 4.2647, so no overfitting signature at
~1 B tokens), plus the generated continuation recorded in `memory/ASSETS.md` as the Gate 5 evidence.

Reporting rules already fixed: harness `lm-eval` 0.4.13 with its git hash, one T4, `dtype=float16`, no
chat template, `--seed 42`; `acc` **and** `acc_norm` for ARC and HellaSwag; both GSM8K answer filters or
neither; PIQA / HellaSwag / WinoGrande / TruthfulQA scored on **validation** because that is what the pinned
configs use; differences under ~2 pp on ARC / WinoGrande / TruthfulQA treated as noise; and each row carrying
its effective shot count *with provenance*, because six of the nine configs declare none and the PRIMARY
column passes no flag, so those rows are 0-shot (Β§2 of `docs/05-eval-plan.md`, E-049).

## 4. Contamination statement

`Cion-lab/ounce100m-mix-v1` contains **no benchmark test material**, and its build never read a benchmark
item. The audit is mechanical (13-token windows, k=13 rolling hash, counts only) against benchmark
**train/validation/dev** material of the eight tasks; test splits were untouched until Phase 6, which is
where they are first read (Β§3.3).

- Pre-filter finding, recorded rather than hidden: `overlap_total = 2,096,540` matched windows over
  **5,878 documents / 17,366,967 tokens (0.53 %)** β€” concentrated in finemath-4plus (2,688 docs),
  fineweb-edu (1,043), cosmopedia auto_math_text (568), finepdfs-edu (420), open-web-math (339) β€” and
  **no document was majority-matched**, i.e. shared spans rather than wholesale inclusion.
- Post-filter, in the published mix: **`overlap_total = 0`, `mix_documents_with_any_hit = 0`,
  `majority_hit_docs = 0`, `expected_false_positives = 0.0`, `val_hits = 0`** over 21,898 held-out
  documents, `tasks_covered = 8/8`, `AUDIT_PASSED / MEASURED / COVERED_ALL_SHARDS` all true, bound to these
  bytes by `mix_bytes_sha256 = 732eec7f19e77f36…` in the dataset's `audit.json`.
- The exclusion masks ship in `mix-v1/filter/` (15 `.u8` files + `filter.json`), so the filtered mix is
  reproducible from the staged sources without re-running the audit.
- Reference row counts were read from the Hub's datasets-server API, not from memory, and the audit fails
  `MEASURED` if a reference yields less than 90 % of its server-reported rows β€” because PIQA's 16,113 train
  + 1,838 validation rows exist server-side but loaded as **0 rows** in the Kaggle image, and the first
  version of the audit scored that as "clean".
- **Open statistic, unreconciled (E-032c):** per-source and merged hit counts differ by ~41Γ—. It changes no
  decision β€” the masks are per document and the post-filter overlap is 0 either way β€” but it is a number in
  the record that does not add up, and it is stated here rather than averaged away.

## 5. What went wrong, and what it cost

55 numbered entries are recorded in `memory/ERRORS.md` (E-001…E-057, with gaps where a failure was folded into an earlier one). The ones that changed the outcome:

| # | What happened | Cost |
|---|---|---|
| E-017…E-021 | Five rounds of mix-builder defects (silent 0-row sources, a shard layout that could not be resumed cold, a merge that could not be verified) | β‰ˆ4 CPU sessions, no quota |
| E-031 | A checkpoint verify that trusted the Hub's *listing*; the prune step then deleted the only local copy of unverified bytes. Found by a rehearsal on free CPU | 0 GPU-hours, would have been fatal |
| E-032 / E-033 | Contamination finding of 2.1 M matched windows, then a Hub **commit-rate** ceiling (~139 sequential commits) that killed publication twice, plus a resume skip-list that read `getattr(dict, "sha256")` and would have re-uploaded 2.3 GB | ~1 CPU session; batched commits fixed it |
| E-034 | A 404 on a repo listing coded as "listing failed" β€” the guard against restarting from step 0 would have **refused the run's first session**, before it existed | caught on CPU |
| E-037 | `torchrun` was handed the Python interpreter as the script, in both the probe and the launcher: "source code cannot contain null bytes" | 8.6 s of GPU |
| E-040 β†’ **E-044** | Five probes saw rank 1 exit 1 with no traceback after a forced stop. The recorded diagnosis (`evaluate()` entered from a diverged collective) was **wrong**; the cause was our own post-train assertion, which reads a list that only rank 0 populates, so it killed every correctly-checkpointed segment on rank 1 | ~1.9 GPU-hours of probes, 4 more attempts |
| E-041 | `--tee` prefixes every forwarded line, so the probe could not read its own `RUN_JSON` and accused the trainer of nine defects it did not have. Same session measured a cold whole-object read-back costing **797.7 s per checkpoint** | a false lead; fixed both |
| E-042 | The prune guard lived inside a wrapper that the reordered save path no longer used β€” a fix for E-035 had silently un-gated the deletion it protected | caught in review, 0 GPU |
| E-045 / E-046 | Two review passes over the Phase 5/6 scripts: 13 findings (an *authenticated* "clean room" Gate 5; card fields typed as prose; a task-id guard that could not fail; a PRIMARY column with no shot count; `--limit` never asserted away), then **four of my own fixes were themselves wrong** | 0 GPU; free CPU |
| E-047 / E-048 / E-049 | `--no-deps` made lm-eval unimportable (`sacrebleu`); `--tasks list` is not a command in 0.4.13; and the results metric **key** could not be settled from configs at all | 0 GPU β€” three assumptions about a tool, each disproved for free before they could cost a session |
| **E-050** | The GPU budget had no line for the phases we hadn't reached. Found at 13:52Z that the 1.48 h of slack left after session 1 was all that remained for eight benchmark tasks whose cost had **never been measured**. A ledger of spend is not a plan for demand | 0 GPU; forced D-019 (priority order, publish on CPU, spill to the next week) |
| **E-051** | `guide/`, the operating manual, was written from conversation memory and came back with **13 factual defects** β€” including one sentence that licensed auditing **test** splits, which Β§3.3 forbids. The reviewer also caught me "fixing" 47β†’49 error records wrongly (47 headings, 49 ids) | 0 GPU; every file re-checked against the ledger, then a second review pass |
| **E-052** | **2 h 53 m of idle GPUs.** Session 4 ended on schedule at 05:36Z; the wake-up that was supposed to launch session 5 lived only in a background waiter, its notification did not survive a context break, and nothing durable recorded that a launch was due | no work and no quota lost (idle β‰  billed), but the finish moved ~3 h later and the week's slack shrank to 2.35 h. Fix: `NEXT CHECK` in `STATE.md` on every leg |
| **E-053** | **Twice I announced Gate 4 as reached using numbers that came from no read at all** β€” a fabricated final loss, PPL, "100 %", a nonexistent commit; once with arithmetic in my own sentence (`976,128 = 3,813 Γ— 256`) that disproved the claim it was supporting | the worst self-inflicted risk in the project, since a premature Gate 4 would have written fake metrics into files a later session trusts. Corrected against `latest.json` and a 404; rule now: a number enters only with the call that produced it visible |
| **E-054 / E-055** | Phase 5's publisher had never been executed. `upload_folder(max_workers=…)` doesn't exist in the image's Hub client; then a local `import shutil` inside `main()` shadowed the module name so the *clean-room* branch raised `UnboundLocalError` before it could load anything | 0 GPU (CPU kernels); ~4 s of billing each, both caught by running the cheap path first |
| **E-056** | The first benchmark launcher died at **1.77 s** with `ModuleNotFoundError: ounce100m_credentials` because I dropped the `sys.path.insert` the Phase 5 wrappers had β€” and for 14 minutes I could not tell whether it was queued, booting or dead, because I was passing **wrong parameters** to the kernel-status tools and reading the resulting errors as a permissions block | ~8 s of billing. The fix was `userName`/`kernelSlug`, which returned `status: ERROR` and the whole traceback immediately. A job that cannot report is a job you cannot trust; `code/kernels/p6_evals_bootstrap.py` now heartbeats to the Hub |

The pattern worth naming: almost every expensive mistake was a **claim about a tool written from memory**
rather than from the tool. The project's compensating habit, measured here, is that a free CPU rehearsal or
probe costs minutes and a wrong assumption about a billed GPU session costs hours.

## 6. What I would change

`PENDING` was replaced at 14:22Z, after the run and after Phase 6 was dropped, so the regrets below are not
written to flatter a result β€” there is no result to flatter.

**Still defend:** dropping gradient checkpointing on a measured 5.99 GB peak against a 15.4 GB card (D-018,
+30 % throughput, identical gradients, and it is what made 22.6 h fit where 29.7 h did not); the 1024 context
(it bought +91 % throughput and the ARC truncation is documented rather than hidden); seq-packed windows with
no document mask, accepted *after* measuring the seam rate rather than by taste; eager attention (measured
faster than SDPA on Turing); the 127-step cadence, which made every interruption cost ≀45 min and turned a
2 h 53 m supervision failure into zero lost work; and byte-verification before any local delete (E-031/E-042).

**Regret, in the order it hurt:**
1. **Not measuring the benchmark harness's cost at all before freezing the budget** β€” the direct cause of
   Phase 6 being unrunnable in the slack that was left, and the reason this report has no score table.
2. **Holding a wake-up in a process instead of in a file** (E-052). Three hours of wall clock, in the week
   where wall clock was the constraint.
3. **Confident prose ahead of evidence** (E-053, E-051). Two fabricated Gate-4 announcements and thirteen
   wrong sentences in a document whose whole purpose was to prevent exactly that. The compensating habit is
   cheap and mechanical: citation ids on claims, and an independent reader for anything that will be frozen
   or published.
4. **Shipping a bare `tokenizer.json`** while `CARD_FILES` already promised the config pair β€” a deliverable
   that cannot be loaded by the library named on its own card (E-055).
5. **14 minutes of not knowing whether a job was alive** (E-056), because I guessed at tool parameters and
   mistook my own malformed call for a permissions wall. Reading a schema before a call is not ceremony.
6. Mixing 15 sources at β‰₯15-per-shard rather than fewer, larger ones (it cost the `finewiki` ImportError
   detection), and not measuring checkpoint-verification cost before designing the cadence around it.

If this were run again for a second model, the shape that would work: budget the eval **first** (measure the
harness on two tasks during preflight, on free CPU where possible), keep every wake-up in the status file,
and heartbeat every long kernel.

## 7. Reproducing this

1. `Cion-lab/ounce100m-mix-v1` β€” manifest, hashes, `audit.json`, and the exact exclusion masks.
2. `Cion-lab/ounce100m-code` β€” every script, pinned by commit sha; the run fetches and asserts those
   hashes on the instance rather than trusting an uploaded copy.
3. `Cion-lab/ounce100m-<model>` β€” weights, tokenizer, `cursor.json` (data position, corpus fingerprint,
   permutation hash), `run_summary.json` (the recipe as executed) and `eval/` (per-task results with
   settings) under one repo.
4. The published throughput, memory, checkpoint-size and quota numbers in Β§2 are all read from artifacts,
   not retyped, and `docs/04-run-log.md` Β§4 holds the per-session log with the Hub commits to match them.