# 01 — Plan: architecture, hyperparameters, throughput, targets Phase 1 deliverable. Every choice below cites a source actually inspected on 2026-09-19, as Gate 1 requires. Items marked **→ test** are decisions whose *numbers* Phase 3 must validate on the real pipeline before the run freezes; the design itself is settled. --- ## 1. The envelope the design has to fit Measured in Phase 0 (`docs/00-platform-notes.md`), not assumed: | Constraint | Value | Consequence | |---|---|---| | Accelerator | 2x Tesla T4, cc 7.5, **14.56 GiB usable each** | fp16 only; no bf16 tensor cores; no flash-attention (sm80+ only) | | GPU quota | 108,000 s/week, billed at **~1x container wall-clock** (3 samples) | ≈30 wall-clock h/week, both cards | | Test budget | ≤6 GPU-h lifetime | 0.073 h spent; **5.93 h remains for all of Phase 3** | | Working disk in a job | **19.5 GB** `/kaggle/working` | shards consumed a few at a time, never whole (§3.13) | | RAM / CPU | 30 GiB cgroup, 4 vCPU, **no swap** | input pipeline memory-bounded by design | | Image | torch 2.10.0+cu128, **transformers 5.0.0**, datasets 5.0.0, accelerate 1.13.0, triton 3.6.0; **no** trl/deepspeed/flash-attn/xformers/bnb | prefer the preinstalled stack; each install is risk + round trip | | Hub from a job | reads anonymous & fine, 35–89 MB/s; writes **401** until given a token | credential path: private Kaggle dataset mount (§8) | **Parameter convention.** §2's 90–110M is *including embeddings* — the Pythia convention (Pythia was explicitly renamed to include embedding + unembedding, [pythia README](https://raw.githubusercontent.com/pythia/main/README.md)). This matters more at 100M than anywhere else: the tied embedding is 21–37 % of the model across the candidates below. `code/config/param_count.py` computes it two ways and both were checked to agree exactly against transformers 5.0.0 in `dodosoomro/ounce100m-p1-param-count`. ## 2. Architecture ### 2.1 The shape: deep-thin, tied, GQA, SwiGLU, pre-norm **Choice: `hidden 576, layers 22, 9 Q heads / 3 KV heads (GQA 3:1), SwiGLU intermediate 1536, RMSNorm pre-norm eps 1e-5, attention_bias false, RoPE theta 10,000, tied embeddings, vocab 49,152, seq 2048`.** Evidence, and it is convergent rather than a single citation: - **Deep-thin beats wide at this size.** Meta's MobileLLM ablated exactly this at a fixed 135M: 30 layers × d512 scored **44.8 avg vs 43.9 for 12 × d768**, with embedding-sharing removing 13 % of parameters at ~equal accuracy ([arXiv 2402.14905](https://arxiv.org/abs/2402.14905), HIGH). Its shipped 125M model is **30 layers, d576, 9Q/3KV GQA, SwiGLU, tied**, 124.6M total. - **The same shape is what HF's SmolLM2-135M uses**: 30 layers, d576, 9Q/**3KV**, `intermediate_size` 1536 (2.67×d), silu, `attention_bias=false`, **`tie_word_embeddings=true`**, RMSNorm 1e-5, [model config](https://huggingface.co/HuggingFaceTB/SmolLM2-135M/blob/main/config.json) + [arXiv 2502.02737 §6](https://arxiv.org/abs/2502.02737). Two independent labs landing on d576 + GQA 3:1 + SwiGLU at ~130M is the strongest available signal. - Our budget is 110M, not 135M, so the same d576/GQA3:1/SwiGLU-1536 stack is cut to **22 layers** rather than 30 — verified count below — and RoPE θ drops to 10,000 to match the shorter context (SmolLM **v1** used θ=10,000 at ctx 2048; SmolLM2's 100,000 accompanies ctx 8192, which we cannot afford and do not need for these benchmarks). - **GQA at 3:1 is a parameter choice as much as an attention choice**, since k/v are `d_head × n_kv` wide; MobileLLM reports GQA-plus-enlarged-d as +0.4 avg. **Not chosen: Mamba/SSM or hybrid.** The official kernels *do* build for `sm_75` ([state-spaces/mamba setup.py](https://github.com/state-spaces/mamba)), which corrects the common assumption that they are Ampere-only — but they are absent from the image (source install only), and the research found **no published 50–200M SSM/hybrid with released hyperparameters** to imitate. That is research risk, not a default, and §4 forbids the kind of custom code where autonomous projects die. ### 2.2 Verified parameter counts Closed form and `sum(p.numel())` agreed exactly on all eight shapes, in the target environment. Tying confirmed genuinely tied, not merely declared. | candidate | H | L | heads | kv | FFN | vocab | **total params** | emb share | in 90–110M | |---|---|---|---|---|---|---|---|---|---| | Phase-0 probe shape | 768 | 12 | 12 | 3 | 2048 | 50257 | 112,934,400 | 34.2 % | ✗ | | A | 768 | 11 | 12 | 3 | 2048 | 50257 | 106,739,712 | 36.2 % | ✓ | | B | 768 | 12 | 12 | 3 | 1792 | 50257 | 105,856,512 | 36.5 % | ✓ | | C | 768 | 12 | 12 | 2 | 1824 | 50257 | 105,561,600 | 36.6 % | ✓ | | D | 640 | 16 | 10 | 4 | 2048 | 50257 | 113,450,240 | 28.4 % | ✗ | | E | 768 | 12 | 12 | 3 | 2048 | 32768 | 99,502,848 | 25.3 % | ✓ | | F | 896 | 9 | 14 | 2 | 2432 | 50257 | 120,397,312 | 37.4 % | ✗ | | G | 1024 | 8 | 8 | 2 | 2816 | 50257 | 141,658,112 | 36.3 % | ✗ | **The binding insight: the tokenizer decides how much model you are allowed to have.** At 50,257 vocab and d768 the tied embedding is 38.6M — a third of the entire budget — which is why every 50k candidate clusters at 105–113M and must trim depth or FFN. MobileLLM's own ablation reached the same conclusion (embedding sharing: −13 % parameters at ~equal accuracy). Choosing a smaller `hidden` is therefore not a weakening, it is how the embedding tax gets paid for. **→ test:** the frozen shape's exact count was re-derived from the constructed model — see §2.3. ### 2.3 The frozen shape, counted to the unit `dodosoomro/ounce100m-p1-param-count` v2 built each candidate as a real `LlamaForCausalLM` under transformers 5.0.0. Closed form and `sum(p.numel())` agree exactly on every row. | candidate | H | L | Q/KV | FFN | vocab | **total params** | emb share | in 90–110M | |---|---|---|---|---|---|---|---|---| | H | 576 | 20 | 9/3 | 1536 | 49,152 | 99,114,048 | 28.6 % | ✓ | | **I — CHOSEN** | **576** | **22** | **9/3** | **1536** | **49,152** | **106,194,240** | **26.7 %** | ✓ | | J | 576 | 24 | 9/3 | 1536 | 49,152 | 113,274,432 | 25.0 % | ✗ over | | K | 576 | 26 | 9/3 | 1536 | 49,152 | 120,354,624 | 23.5 % | ✗ over | | L | 640 | 20 | 10/4 | 1728 | 49,152 | 120,776,320 | 26.0 % | ✗ over | **Frozen: 106,194,240 parameters (embeddings included, tied), counted as `vocab×h` once + `L × (attn + SwiGLU)` + norms + final norm, and confirmed by constructing the model.** The counting method is stated here because §2 requires the number *and* how it was counted. Two consequences worth recording: - **Depth is capped at 22 by the budget, not by taste.** Each layer at d576 costs 3,538,944 parameters, so 24 layers is 113.3M — over. That is a hard stop, and it means the Phase-3 "deep-thin vs wide" throughput test (T1) must compare **H (20 layers, 99.1M) against I (22 layers, 106.2M)**, both in-budget, rather than the 12-layer d768 probe shape which cannot fit the SmolLM2 vocab without trimming elsewhere. - **Widening is not a free alternative to deepening.** Candidate L at d640 × 20 layers is *over* budget (120.8M) despite having fewer layers than I, because attention and MLP cost scale with `h²`/`h·ff`. So the deep-thin choice is forced from two directions: it is better per parameter (MobileLLM's ablation) and it is the only way to spend the 78M non-embedding budget on 22 layers. ## 3. Tokenizer **Choice: reuse the SmolLM2 BPE tokenizer** — vocab 49,152, byte-level, **Apache-2.0**, from [`HuggingFaceTB/SmolLM2-135M`](https://huggingface.co/HuggingFaceTB/SmolLM2-135M)'s `tokenizer.json`; trained on SmolCorpus per [arXiv 2502.02737](https://arxiv.org/abs/2502.02737). Not trained by us. §4 says maximum reuse, minimal custom code, and training a tokenizer from scratch is precisely the kind of optional machinery that adds failure surface without adding capability. Candidates checked today: SmolLM2 **49,152 / Apache-2.0**; `gpt2` 50,257 / MIT; `pythia-160m` 50,304 / Apache-2.0; `t5-v1_1-small` 32,128 Unigram (not byte-level, carries `extra_ids` sentinels); `Qwen3-0.6B` 151,936 (disqualifying — its embedding alone would be ~78M at d512); `TinyLlama` 32,000 but Llama-2-derived and **redistribution status unverified** → excluded. SmolLM2 wins on three counts: licence is clean and permissive, it is byte-level and English-curated at exactly this model scale, and sharing it with a family of released 135M models gives the benchmark numbers somebody else's token statistics to be compared against. Its 49,152 × 576 tied embedding is 28.3M = **21 %** of budget, versus 34–37 % for the 50k/d768 shapes. ## 4. Optimizer, batch, schedule, precision | | choice | justification | |---|---|---| | Optimizer | AdamW, β=(0.9, 0.95), wd 0.1, clip 1.0 | Pythia-160M and SmolLM2 both use β₂=0.95 at small scale ([2304.01373](https://arxiv.org/abs/2304.01373), [2502.02737](https://arxiv.org/abs/2502.02737)) | | Peak LR | **6e-4** | bracketed from three directions: Cerebras-111M 6.0e-4, Pythia-160M 6e-4, and the 150M low-tokens/param run of Sardana et al. at 4.6e-4 ([2304.03208](https://arxiv.org/abs/2304.03208), [2401.00448](https://arxiv.org/abs/2401.00448)). The 2e-3/3e-3 figures that appear in small-model recipes belong to 1–2 T-token runs, which this is not. | | Warmup | 2 % of steps, floor 100 steps | Pythia used 1 %; SmolLM2 used 2,000 steps at 2M tok/batch (~0.2 %). Warmer than 1 % because **fp16 without bf16 needs the ramp** — Phase 0 showed an unramped/unscaled fp16 run going straight to NaN (E-006). | | Schedule | **trapezoid / WSD: constant, then linear decay to 0 over the final 20 %** — *not* cosine-to-10 % | SmolLM1 used a trapezoid with 20 % cooldown; SmolLM2 used WSD ([blog/smollm](https://huggingface.co/blog/smollm), [2502.02737 App. A](https://arxiv.org/abs/2502.02737)). *Straight to Zero* finds linear-decay-to-zero beats cosine-to-10 % when peak LR is optimal, across sizes/batch/data ([2502.15938](https://arxiv.org/abs/2502.15938)); COLT 2026 theory says decay *shape* barely matters but **overly slow terminal decay causes schedule-induced capacity saturation** ([2602.06797](https://arxiv.org/abs/2602.06797)). Its caveat applies: their evidence is at/above compute-optimal tokens/param, we are at half. | | Global batch | **≈262 k tokens/step** (~3,815 steps for 1 B tokens) | Critical batch size fits `B* = 621.341·N^0.087`, i.e. **nearly independent of model size** ([2410.21676](https://arxiv.org/abs/2410.21676)) — no reason to chase 2–4M-token batches. Sardana et al. deliberately used *smaller* batches for smaller models "so that low-token-count training runs see enough steps". At 10 tokens/param we are such a run. | | Micro-batch | **2 × 2048 per GPU × 2 GPUs × 32 accumulation** | Set by memory, not taste: Phase 0 measured bs4×1024 at 11.5–13.8 GB of 14.56 GB and **bs8 OOMed outright**. → test | | Precision | **fp16 autocast + fp32 master weights + `GradScaler` + grad clip** | Only mixed-precision option on Turing; transformers v5 documents this path by name ("fall back to fp16 on older hardware like V100 or T4") and warns **load the model in fp32 or autocast is a no-op** ([v5 mixed_precision docs](https://huggingface.co/docs/transformers/perf_training)). Verified finite in Phase 0 probe C; bf16 *runs* but is emulated → must be asserted off, not defaulted. | | Seed / RNG | single fixed seed, recorded | §3.1 requires "same RNG semantics" across resumes | **Undertrained by design, and honestly so.** 1 B tokens on ~100M ≈ **10 tokens/param**, about half Chinchilla-optimal (Cerebras-111M's compute-optimal point was 2.2 B tokens = 20 t/p). That is mildly undertrained, not pathological — the genuine anomaly is 2025-26 practice, which overtrains tiny models by three orders of magnitude (SmolLM2-135M at 2 T tokens ≈ 14,800 t/p). Sardana et al. trained a 150M/d768/12L model across 3→10,000 t/p and quality kept improving; Gadre et al. show scaling laws extrapolate across over-training ([2403.08540](https://arxiv.org/abs/2403.08540)). We are on the left-hand side of that curve because the quota puts us there, and the report should say so. ## 5. Training stack **Choice: `torchrun --nproc_per_node=2` + Hugging Face `Trainer` / `TrainingArguments`, `fp16=True`, on a pre-tokenised **map-style** `Dataset`.** Minimal glue, no framework authored here (§4). - Everything needed is already in the image: transformers 5.0.0, accelerate 1.13.0, datasets 5.0.0. - v5 `Trainer` restores **RNG (`_save_rng_state`/`_load_rng_load`), optimizer, LR scheduler and the fp16 GradScaler** — the exact set §3.1 demands — via `Accelerator.save_state`. Verified against v5.17.0 source; the image's 5.0.0 is ~8 patch releases behind but the same major. **→ test at Gate 3.** - **The resume trap, stated plainly:** v5 recovers data position with `skip_first_batches`, which **replays** discarded batches — O(steps) wasted work on an `IterableDataset`, up to 3,815 steps of pure I/O after every interruption. Mitigation is architectural, not a flag: a **map-style, concatenated token store** makes skipping an index advance rather than a decode. A `dataset_shard` layout with an explicit recorded cursor is what Gate 3 must prove, and §Phase 2's "exact positional resumption" requirement is the reason it is designed that way up front. - **v5 migration traps to check in preflight** ([MIGRATION_GUIDE_V5.md](https://github.com/huggingface/transformers/blob/main/MIGRATION_GUIDE_V5.md), HF *Transformers v5* blog): PyTorch-only backend; slow tokenizers gone (irrelevant, we use `tokenizers`); several `TrainingArguments` removed **without a deprecation cycle**; attention moved to `AttentionInterface`; `remove_unused_columns` **still defaults True** (silently drops dataset columns a loss function needs); `lr_scheduler_type` default is still `"linear"`, *not* cosine — so an unstated schedule is a linear decay to zero over the whole run; new `train_sampling_strategy` replaces the `group_by_length` bool; `restore_callback_states_from_checkpoint` is required to restore scheduler state properly. - **Disqualified:** **torchtitan** — `training_dtype` is typed `Literal["bfloat16","float32"]`, i.e. **no fp16 path at all**, plus Hopper-oriented CI → unusable on T4 ([torchtitan config](https://github.com/pytorch/torchtitan)). **unsloth** — fine-tuning/GRPO only, no from-scratch pretraining path. **nanoGPT** — dormant since 2025-11, no HF export, no distributed input pipeline. **FSDP** — ~100M params in fp32 master + grads + AdamW is ~1.8 GB; DDP on 2×16 GB does not need sharding, and sharding would add failure surface for nothing. **DeepSpeed / flash-attn / xformers / trl** — absent from the image; installing them is risk without a capability we lack. ## 6. Throughput, and the token target it implies Measured on a main-run-shaped model, `seq 1024`, bs4/card, DDP over both T4s: **11,062 tok/s aggregate** (5,531/rank; single-GPU 6,478 SDPA / 7,239 eager; **85 % DDP efficiency**). Sanity-checking that against compute: `6ND ≈ 6 × 1.13e8 × 11,062 ≈ 7.5e12 FLOP/s`, which is **~6 % of a 2×T4 fp16 tensor peak** (65 TFLOPS/card — an *unverified* datasheet prior; the spec page could not be retrieved). Independent corroboration: a documented 163M / ctx-1024 run reached ~19.9k tok/s on a single RTX 3090 ([gilesthomas.com](https://gilesthomas.com/2025/12/llm-from-scratch-28-training-a-base-model-from-scratch)), and Pythia-160M's 1,030 A100-hours for 300 B tokens implies ~81k tok/s per A100. So ~11k tok/s on two Turing cards is the right order of magnitude, and **the low implied MFU means this is a floor with headroom, not a ceiling** — which is the opposite of the risk I most wanted to rule out. **The tension to resolve by measurement, not argument:** deep-thin is better per *parameter* (MobileLLM's ablation) but worse per *second* on 2018 hardware, because 22 small layers issue more sequential kernel launches than 12 fat ones, and we are wall-clock-bound rather than parameter-bound. Phase 3 therefore measures both shapes. | token target | at 11.1k tok/s (measured) | at 9k tok/s (planning floor) | |---|---|---| | 1.10 B | 27.6 h | 34.0 h | | **1.00 B** | **25.1 h** | **30.9 h** | | 0.90 B | 22.6 h | 27.8 h | **Decision: target 1.0 B tokens, with a pre-registered fallback to 0.9 B if Gate 3's end-to-end rate on the frozen config and real input pipeline is below ~10.5k tok/s.** Both are inside §2's band, and the fallback is chosen *now* rather than later so it cannot become a post-hoc excuse. Even the optimistic column leaves under 5 h of slack in a 30 h week, which the interruption budget in §3.1 will eat; so the plan assumes **a two-week run crossing the 2026-09-26 quota reset**, with a checkpoint at every 10 % and a rolling `latest`, exactly as §2 requires. ## 6.1 Amendment after Gate 3's measurements — the fallback trigger fired, and here is what was done about it **Read §6 above first: it is kept verbatim, including its numbers, which are now known to be wrong.** The 11,062 tok/s figure came from p0c's *raw* training loop at seq 1024 on a 113M model — no `Trainer`, no real input pipeline, no fp16-master/GradScaler bookkeeping, no seq-2048 activations. Gate 3 measured the frozen 106,194,240-parameter shape inside the actual stack on 2×T4 (ledger rows 10-14, `memory/QUOTA.md`): | config | tok/s | peak GB | 1.0 B tokens | |---|---|---|---| | sdpa @2048 micro 2 (the §6 plan-of-record) | 4,071 | 14.22 | **68.2 h** | | eager + grad-ckpt @2048 micro 2 | 4,679 | 10.04 | 59.4 h | | eager @1024 micro 4 | 7,732 | 12.25 | 35.9 h | | eager + grad-ckpt @1024 micro 8 | 7,828 | 7.12 | 35.5 h | | eager @2048 micro 2 | OOM | — | — | Three consequences, in the order that matters: 1. **Sequence length changed to 1024** — that is **D-011**, decided on these numbers and documented in `memory/DECISIONS.md`. Micro-batch 4 × accum 32 × 2 cards keeps tokens/step at **262,144**, so the batch size in tokens, the ~3,815 steps, the LR and the trapezoid are all unchanged. §2.3's parameter table is unaffected — it never depended on the measurement. 2. **The pre-registered fallback condition in §6 fired.** "Below ~10.5k tok/s" is satisfied: every measured seq-1024 cell is ~7.7-7.8k. Mechanically that means 0.9 B. 3. **The premise of that trigger is gone.** 10.5k was chosen so 1.0 B would fit *one* 30 h week with enough slack for interruption overhead. At the measured rate neither 1.0 B (35.9 h) nor 0.9 B (32.3 h) fits week 1, so the trigger no longer discriminates between the two — it only buys 3.6 GPU-hours for a 10 % cut in training tokens. Against the two-week schedule in `docs/04-run-log.md` §2 (60 h of quota, 35.9 h of training, **~22 h of slack**) the 1.0 B target is affordable. **Therefore, re-registered before launch and before any checkpoint or score exists (D-012): the token target stays 1.0 B, and the fallback trigger is restated against measured throughput rather than against p0c's number.** New mechanical rule, same purpose as the old one: **if preflight P3's end-to-end rate at the frozen geometry (22L/576, seq 1024, eager, grad-ckpt on, micro 4, accum 32) is below 5,150 tok/s — i.e. 1 B tokens needs more than ~54 h of the ~60 h available across the two weeks after allowing 6 h for interruption overhead and Phase 6's GPU evaluation — the target drops to 0.9 B at launch, not later.** 5,150 tok/s is 33 % below the planning rate, so this is a genuine tripwire, not a formality. What did *not* change: D-005's benchmark bands, the harness pin (D-010), the mix, the model shape, the optimiser, the schedule, or the one-run rule. If the 0.9 B fallback fires, it fires as a documented arithmetic consequence, and this section is where a reader can check it. ## 7. Benchmark targets — pre-registered before the main run (§4) Recorded now, before any training, so the goalposts cannot move. Full anchor table, protocol and noise floor in §7.1–7.3. Sources: archived Open LLM Leaderboard run metadata and Pythia Appendix G ([arXiv 2304.01373](https://arxiv.org/pdf/2304.01373)), read 2026-09-19. **The dominant fact: every published ≥100M base anchor was trained on ~300 B tokens.** Our budget is ~1 B — a **300× gap**. Those anchors are therefore ceilings, not expectations. No harness score for any 50–350M base trained at ~1–3 B tokens could be found published anywhere, so these bands are extrapolations from the low end of the size curve and are labelled as such. | Task | metric | band | chance | |---|---|---|---| | ARC-Easy | `acc` | **26–34** | 25 | | ARC-Challenge | `acc` / `acc_norm` | **17–22** / **20–25** | 25 | | HellaSwag | `acc_norm` (report `acc` too) | **26–32** (acc 25–30) | 25 | | PIQA | `acc` | **52–60** | 50 | | WinoGrande | `acc` | **49–53** | 50 | | MMLU | `acc` | **24–27** | 25 | | TruthfulQA | `mc2` / `mc1` | **36–46** / **21–26** | ≈38 / ≈22 | | GSM8K | `exact_match,strict-match` | **0.0–1.5** | ≈0 | - **MMLU, WinoGrande, GSM8K and ARC-Challenge are expected to be indistinguishable from chance.** MMLU measures flat 24.9–27.3 from 70M to 1.4B — OPT-1.3B scores *lower* than OPT-125M — so it has almost no discriminative power at this scale. GSM8K anchors: 0.23 (OPT-125M), 0.68 (GPT-2, Pythia-410M), 1.52 (Pythia-1.4B), 1.4 even for SmolLM2-135M at 11.2 T tokens. Multi-step arithmetic does not emerge from a 100M base at 1 B tokens; predicting otherwise would be the inflation §3.12 forbids. - **The real signals are ARC-Easy, HellaSwag and PIQA** — the three where anchors clear chance by a visible margin. Failing to beat chance there is a finding about the mix or the run, not the benchmarks. - **Noise floor:** SE ≈ 1.3 pp ARC-C (n=1172), 1.4 pp WinoGrande (1267), 1.2 pp PIQA (1838), 1.5 pp TruthfulQA (817), 0.45 pp HellaSwag (10042). **Differences under ~2 pp on ARC/WinoGrande/TruthfulQA are not results**, and no shot count or prompt format gets chosen per-task after seeing scores (§3.3). ### 7.1 Protocol, pinned for Phase 6 Harness **EleutherAI `lm-evaluation-harness` v0.4.13** (2026-08-31) — still what published anchors report against; Lighteval v0.13.0 is the active HF alternative but switching would break comparability. Record the git SHA per run; cadence is real (v0.4.10 2026-01 → v0.4.13 2026-08). | Task | dataset / config | split | shots | |---|---|---|---| | `arc_easy`, `arc_challenge` | `allenai/ai2_arc` | **test** (2376 / 1172) | **25** | | `hellaswag` | `Rowan/hellaswag` | **validation** (10042; no test split) | **10** | | `piqa` | `baber/piqa` default | **validation** (1838) | **10** | | `winogrande` | `allenai/winogrande`, **`winogrande_xl`** | **validation** (1267) | **5** | | `mmlu` | group of 57 `mmlu_` | test | **5** | | `truthfulqa_mc1`, `_mc2` | `sylinrl/TruthfulQA` | validation | **0** (pinned) | | `gsm8k` | `openai/gsm8k` `main` | test | **5** | ### 7.2 Traps that would silently falsify the numbers 1. **Shots are NOT pinned in the ARC/HellaSwag/PIQA/WinoGrande/MMLU YAMLs** — an omitted `--num_fewshot` silently yields 0-shot. Pass shots explicitly for every task. 2. **No chat templates.** `--apply_chat_template` / `--fewshot_as_multiturn` are for fine-tuned models (§3.9 makes this a base model). `gsm8k_cot_llama` requires them → not used. 3. Everything but GSM8K is `output_type: multiple_choice` = loglikelihood ranking; the model never generates, so only GSM8K is sensitive to generation settings. 4. **MMLU aggregation:** the `mmlu` group uses `weight_by_size: True`, while archived leaderboard numbers were an *unweighted* subject mean. Compute and state which was used. 5. **TruthfulQA mc2 changed definition** 2024-03-11 (PR #2768); pre-April-2024 mc2 is not comparable. mc1/mc2 differ by ~20 pp and both report under the metric name `acc`. 6. **Harness PIQA is mean per-item accuracy, not AI2's `p_win`** — the harness never computes `p_win`. 7. **Winogrande**: `winogrande_xl`, single fold, validation; the official leaderboard averages 5 folds, and `winogrande_debiased` is a different number. 8. **GSM8K**: `temperature 0`, `do_sample false`, `max_gen_toks` inherits the global 256; tiny bases loop repetitively and never emit `#### `, so strict-match ≈0 while flexible-extract looks inflated. Report both filters or neither; do **not** add `repetition_penalty`, which silently deviates from everyone. 9. **Splits are mixed** — PIQA/HellaSwag/WinoGrande/TruthfulQA run on *validation*. Never call them test. 10. **Parameter-count conventions differ between suites**; Pythia's includes embeddings. Quote ours with the convention attached. ## 8. Still open before Gate 1 closes - ~~Deep-thin-vs-wide throughput result~~ → scheduled as Phase 3 preflight test T1. - Session wall-clock cap, from the still-running `p0e-session-cap`: it sets how much progress one session can make and therefore how often the checkpoint cycle interrupts training. **Known so far: a CPU session was still alive at 51 min** with no cap hit, so short sessions are not the failure mode; the ceiling is somewhere above that. - **Closed this session:** the credential path (§1 last row) — jobs retrieve the HF token from the account's own private Kaggle dataset via `code/ounce100m_credentials.py`; a Hub write from inside a job was verified by anonymous readback (D-006). And the checkpoint arithmetic below. ### 8.1 Checkpoint size and cadence — now measured, and it is cheap A `latest`-quality exact-resume checkpoint for a ~100M model is **~1.6 GB** (fp32 weights 400 MB + two fp32 Adam moments 800 MB + master/grad copies + scheduler/scaler/RNG state, which are all negligible next to the optimizer). Pushed from a Kaggle session at **42.7 MB/s → 37.4 s**, and pulled back in **13.8 s**, with byte-exact anonymous readback at 200 MB / 800 MB / 1.6 GB (`dodosoomro/ounce100m-p1-push-bench`, D-007). Consequences that change design choices rather than just confirming them: - 10 checkpoints + `latest` ≈ **7 min total, ~0.12 GPU-h of a 30 h week**. Checkpoint cadence is not a budget problem, so there is no reason to economise on it — and rolling `latest` *more* often than the required 10 % is nearly free, which is worth doing since every interruption costs at most one roll. - A cold resume with an empty disk costs seconds of transfer, not minutes. The expensive part of a resume is the `skip_first_batches` replay discussed in §5, which is why the input format in `docs/02-mix-plan.md` §5 step 6 is designed to make it an index advance. - Peak disk: 1.6 GB written + 1.6 GB staged elsewhere on the 19.5 GB volume is comfortable, and the §3.13 prune-after-verify step keeps at most one copy resident.