ounce100m-code / docs /01-plan.md
Cion-lab's picture
publish docs/: plan, mix rationale, preflight report, run log, frozen eval protocol, final report
f345921 verified
|
Raw History Blame Contribute Delete
26.9 kB

01 β€” Plan: architecture, hyperparameters, throughput, targets

Phase 1 deliverable. Every choice below cites a source actually inspected on 2026-09-19, as Gate 1 requires. Items marked β†’ test are decisions whose numbers Phase 3 must validate on the real pipeline before the run freezes; the design itself is settled.


1. The envelope the design has to fit

Measured in Phase 0 (docs/00-platform-notes.md), not assumed:

Constraint Value Consequence
Accelerator 2x Tesla T4, cc 7.5, 14.56 GiB usable each fp16 only; no bf16 tensor cores; no flash-attention (sm80+ only)
GPU quota 108,000 s/week, billed at ~1x container wall-clock (3 samples) β‰ˆ30 wall-clock h/week, both cards
Test budget ≀6 GPU-h lifetime 0.073 h spent; 5.93 h remains for all of Phase 3
Working disk in a job 19.5 GB /kaggle/working shards consumed a few at a time, never whole (Β§3.13)
RAM / CPU 30 GiB cgroup, 4 vCPU, no swap input pipeline memory-bounded by design
Image torch 2.10.0+cu128, transformers 5.0.0, datasets 5.0.0, accelerate 1.13.0, triton 3.6.0; no trl/deepspeed/flash-attn/xformers/bnb prefer the preinstalled stack; each install is risk + round trip
Hub from a job reads anonymous & fine, 35–89 MB/s; writes 401 until given a token credential path: private Kaggle dataset mount (Β§8)

Parameter convention. Β§2's 90–110M is including embeddings β€” the Pythia convention (Pythia was explicitly renamed to include embedding + unembedding, pythia README). This matters more at 100M than anywhere else: the tied embedding is 21–37 % of the model across the candidates below. code/config/param_count.py computes it two ways and both were checked to agree exactly against transformers 5.0.0 in dodosoomro/ounce100m-p1-param-count.

2. Architecture

2.1 The shape: deep-thin, tied, GQA, SwiGLU, pre-norm

Choice: hidden 576, layers 22, 9 Q heads / 3 KV heads (GQA 3:1), SwiGLU intermediate 1536, RMSNorm pre-norm eps 1e-5, attention_bias false, RoPE theta 10,000, tied embeddings, vocab 49,152, seq 2048.

Evidence, and it is convergent rather than a single citation:

  • Deep-thin beats wide at this size. Meta's MobileLLM ablated exactly this at a fixed 135M: 30 layers Γ— d512 scored 44.8 avg vs 43.9 for 12 Γ— d768, with embedding-sharing removing 13 % of parameters at ~equal accuracy (arXiv 2402.14905, HIGH). Its shipped 125M model is 30 layers, d576, 9Q/3KV GQA, SwiGLU, tied, 124.6M total.
  • The same shape is what HF's SmolLM2-135M uses: 30 layers, d576, 9Q/3KV, intermediate_size 1536 (2.67Γ—d), silu, attention_bias=false, tie_word_embeddings=true, RMSNorm 1e-5, model config + arXiv 2502.02737 Β§6. Two independent labs landing on d576 + GQA 3:1 + SwiGLU at ~130M is the strongest available signal.
  • Our budget is 110M, not 135M, so the same d576/GQA3:1/SwiGLU-1536 stack is cut to 22 layers rather than 30 β€” verified count below β€” and RoPE ΞΈ drops to 10,000 to match the shorter context (SmolLM v1 used ΞΈ=10,000 at ctx 2048; SmolLM2's 100,000 accompanies ctx 8192, which we cannot afford and do not need for these benchmarks).
  • GQA at 3:1 is a parameter choice as much as an attention choice, since k/v are d_head Γ— n_kv wide; MobileLLM reports GQA-plus-enlarged-d as +0.4 avg.

Not chosen: Mamba/SSM or hybrid. The official kernels do build for sm_75 (state-spaces/mamba setup.py), which corrects the common assumption that they are Ampere-only β€” but they are absent from the image (source install only), and the research found no published 50–200M SSM/hybrid with released hyperparameters to imitate. That is research risk, not a default, and Β§4 forbids the kind of custom code where autonomous projects die.

2.2 Verified parameter counts

Closed form and sum(p.numel()) agreed exactly on all eight shapes, in the target environment. Tying confirmed genuinely tied, not merely declared.

candidate H L heads kv FFN vocab total params emb share in 90–110M
Phase-0 probe shape 768 12 12 3 2048 50257 112,934,400 34.2 % βœ—
A 768 11 12 3 2048 50257 106,739,712 36.2 % βœ“
B 768 12 12 3 1792 50257 105,856,512 36.5 % βœ“
C 768 12 12 2 1824 50257 105,561,600 36.6 % βœ“
D 640 16 10 4 2048 50257 113,450,240 28.4 % βœ—
E 768 12 12 3 2048 32768 99,502,848 25.3 % βœ“
F 896 9 14 2 2432 50257 120,397,312 37.4 % βœ—
G 1024 8 8 2 2816 50257 141,658,112 36.3 % βœ—

The binding insight: the tokenizer decides how much model you are allowed to have. At 50,257 vocab and d768 the tied embedding is 38.6M β€” a third of the entire budget β€” which is why every 50k candidate clusters at 105–113M and must trim depth or FFN. MobileLLM's own ablation reached the same conclusion (embedding sharing: βˆ’13 % parameters at ~equal accuracy). Choosing a smaller hidden is therefore not a weakening, it is how the embedding tax gets paid for.

β†’ test: the frozen shape's exact count was re-derived from the constructed model β€” see Β§2.3.

2.3 The frozen shape, counted to the unit

dodosoomro/ounce100m-p1-param-count v2 built each candidate as a real LlamaForCausalLM under transformers 5.0.0. Closed form and sum(p.numel()) agree exactly on every row.

candidate H L Q/KV FFN vocab total params emb share in 90–110M
H 576 20 9/3 1536 49,152 99,114,048 28.6 % βœ“
I β€” CHOSEN 576 22 9/3 1536 49,152 106,194,240 26.7 % βœ“
J 576 24 9/3 1536 49,152 113,274,432 25.0 % βœ— over
K 576 26 9/3 1536 49,152 120,354,624 23.5 % βœ— over
L 640 20 10/4 1728 49,152 120,776,320 26.0 % βœ— over

Frozen: 106,194,240 parameters (embeddings included, tied), counted as vocabΓ—h once + L Γ— (attn + SwiGLU) + norms + final norm, and confirmed by constructing the model. The counting method is stated here because Β§2 requires the number and how it was counted.

Two consequences worth recording:

  • Depth is capped at 22 by the budget, not by taste. Each layer at d576 costs 3,538,944 parameters, so 24 layers is 113.3M β€” over. That is a hard stop, and it means the Phase-3 "deep-thin vs wide" throughput test (T1) must compare H (20 layers, 99.1M) against I (22 layers, 106.2M), both in-budget, rather than the 12-layer d768 probe shape which cannot fit the SmolLM2 vocab without trimming elsewhere.
  • Widening is not a free alternative to deepening. Candidate L at d640 Γ— 20 layers is over budget (120.8M) despite having fewer layers than I, because attention and MLP cost scale with hΒ²/hΒ·ff. So the deep-thin choice is forced from two directions: it is better per parameter (MobileLLM's ablation) and it is the only way to spend the 78M non-embedding budget on 22 layers.

3. Tokenizer

Choice: reuse the SmolLM2 BPE tokenizer β€” vocab 49,152, byte-level, Apache-2.0, from HuggingFaceTB/SmolLM2-135M's tokenizer.json; trained on SmolCorpus per arXiv 2502.02737. Not trained by us.

Β§4 says maximum reuse, minimal custom code, and training a tokenizer from scratch is precisely the kind of optional machinery that adds failure surface without adding capability. Candidates checked today: SmolLM2 49,152 / Apache-2.0; gpt2 50,257 / MIT; pythia-160m 50,304 / Apache-2.0; t5-v1_1-small 32,128 Unigram (not byte-level, carries extra_ids sentinels); Qwen3-0.6B 151,936 (disqualifying β€” its embedding alone would be ~78M at d512); TinyLlama 32,000 but Llama-2-derived and redistribution status unverified β†’ excluded.

SmolLM2 wins on three counts: licence is clean and permissive, it is byte-level and English-curated at exactly this model scale, and sharing it with a family of released 135M models gives the benchmark numbers somebody else's token statistics to be compared against. Its 49,152 Γ— 576 tied embedding is 28.3M = 21 % of budget, versus 34–37 % for the 50k/d768 shapes.

4. Optimizer, batch, schedule, precision

choice justification
Optimizer AdamW, Ξ²=(0.9, 0.95), wd 0.1, clip 1.0 Pythia-160M and SmolLM2 both use Ξ²β‚‚=0.95 at small scale (2304.01373, 2502.02737)
Peak LR 6e-4 bracketed from three directions: Cerebras-111M 6.0e-4, Pythia-160M 6e-4, and the 150M low-tokens/param run of Sardana et al. at 4.6e-4 (2304.03208, 2401.00448). The 2e-3/3e-3 figures that appear in small-model recipes belong to 1–2 T-token runs, which this is not.
Warmup 2 % of steps, floor 100 steps Pythia used 1 %; SmolLM2 used 2,000 steps at 2M tok/batch (~0.2 %). Warmer than 1 % because fp16 without bf16 needs the ramp β€” Phase 0 showed an unramped/unscaled fp16 run going straight to NaN (E-006).
Schedule trapezoid / WSD: constant, then linear decay to 0 over the final 20 % β€” not cosine-to-10 % SmolLM1 used a trapezoid with 20 % cooldown; SmolLM2 used WSD (blog/smollm, 2502.02737 App. A). Straight to Zero finds linear-decay-to-zero beats cosine-to-10 % when peak LR is optimal, across sizes/batch/data (2502.15938); COLT 2026 theory says decay shape barely matters but overly slow terminal decay causes schedule-induced capacity saturation (2602.06797). Its caveat applies: their evidence is at/above compute-optimal tokens/param, we are at half.
Global batch β‰ˆ262 k tokens/step (~3,815 steps for 1 B tokens) Critical batch size fits B* = 621.341Β·N^0.087, i.e. nearly independent of model size (2410.21676) β€” no reason to chase 2–4M-token batches. Sardana et al. deliberately used smaller batches for smaller models "so that low-token-count training runs see enough steps". At 10 tokens/param we are such a run.
Micro-batch 2 Γ— 2048 per GPU Γ— 2 GPUs Γ— 32 accumulation Set by memory, not taste: Phase 0 measured bs4Γ—1024 at 11.5–13.8 GB of 14.56 GB and bs8 OOMed outright. β†’ test
Precision fp16 autocast + fp32 master weights + GradScaler + grad clip Only mixed-precision option on Turing; transformers v5 documents this path by name ("fall back to fp16 on older hardware like V100 or T4") and warns load the model in fp32 or autocast is a no-op (v5 mixed_precision docs). Verified finite in Phase 0 probe C; bf16 runs but is emulated β†’ must be asserted off, not defaulted.
Seed / RNG single fixed seed, recorded Β§3.1 requires "same RNG semantics" across resumes

Undertrained by design, and honestly so. 1 B tokens on ~100M β‰ˆ 10 tokens/param, about half Chinchilla-optimal (Cerebras-111M's compute-optimal point was 2.2 B tokens = 20 t/p). That is mildly undertrained, not pathological β€” the genuine anomaly is 2025-26 practice, which overtrains tiny models by three orders of magnitude (SmolLM2-135M at 2 T tokens β‰ˆ 14,800 t/p). Sardana et al. trained a 150M/d768/12L model across 3β†’10,000 t/p and quality kept improving; Gadre et al. show scaling laws extrapolate across over-training (2403.08540). We are on the left-hand side of that curve because the quota puts us there, and the report should say so.

5. Training stack

Choice: torchrun --nproc_per_node=2 + Hugging Face Trainer / TrainingArguments, fp16=True, on a pre-tokenised map-style Dataset. Minimal glue, no framework authored here (Β§4).

  • Everything needed is already in the image: transformers 5.0.0, accelerate 1.13.0, datasets 5.0.0.
  • v5 Trainer restores RNG (_save_rng_state/_load_rng_load), optimizer, LR scheduler and the fp16 GradScaler β€” the exact set Β§3.1 demands β€” via Accelerator.save_state. Verified against v5.17.0 source; the image's 5.0.0 is ~8 patch releases behind but the same major. β†’ test at Gate 3.
  • The resume trap, stated plainly: v5 recovers data position with skip_first_batches, which replays discarded batches β€” O(steps) wasted work on an IterableDataset, up to 3,815 steps of pure I/O after every interruption. Mitigation is architectural, not a flag: a map-style, concatenated token store makes skipping an index advance rather than a decode. A dataset_shard layout with an explicit recorded cursor is what Gate 3 must prove, and Β§Phase 2's "exact positional resumption" requirement is the reason it is designed that way up front.
  • v5 migration traps to check in preflight (MIGRATION_GUIDE_V5.md, HF Transformers v5 blog): PyTorch-only backend; slow tokenizers gone (irrelevant, we use tokenizers); several TrainingArguments removed without a deprecation cycle; attention moved to AttentionInterface; remove_unused_columns still defaults True (silently drops dataset columns a loss function needs); lr_scheduler_type default is still "linear", not cosine β€” so an unstated schedule is a linear decay to zero over the whole run; new train_sampling_strategy replaces the group_by_length bool; restore_callback_states_from_checkpoint is required to restore scheduler state properly.
  • Disqualified: torchtitan β€” training_dtype is typed Literal["bfloat16","float32"], i.e. no fp16 path at all, plus Hopper-oriented CI β†’ unusable on T4 (torchtitan config). unsloth β€” fine-tuning/GRPO only, no from-scratch pretraining path. nanoGPT β€” dormant since 2025-11, no HF export, no distributed input pipeline. FSDP β€” ~100M params in fp32 master + grads + AdamW is ~1.8 GB; DDP on 2Γ—16 GB does not need sharding, and sharding would add failure surface for nothing. DeepSpeed / flash-attn / xformers / trl β€” absent from the image; installing them is risk without a capability we lack.

6. Throughput, and the token target it implies

Measured on a main-run-shaped model, seq 1024, bs4/card, DDP over both T4s: 11,062 tok/s aggregate (5,531/rank; single-GPU 6,478 SDPA / 7,239 eager; 85 % DDP efficiency).

Sanity-checking that against compute: 6ND β‰ˆ 6 Γ— 1.13e8 Γ— 11,062 β‰ˆ 7.5e12 FLOP/s, which is ~6 % of a 2Γ—T4 fp16 tensor peak (65 TFLOPS/card β€” an unverified datasheet prior; the spec page could not be retrieved). Independent corroboration: a documented 163M / ctx-1024 run reached ~19.9k tok/s on a single RTX 3090 (gilesthomas.com), and Pythia-160M's 1,030 A100-hours for 300 B tokens implies ~81k tok/s per A100. So ~11k tok/s on two Turing cards is the right order of magnitude, and the low implied MFU means this is a floor with headroom, not a ceiling β€” which is the opposite of the risk I most wanted to rule out.

The tension to resolve by measurement, not argument: deep-thin is better per parameter (MobileLLM's ablation) but worse per second on 2018 hardware, because 22 small layers issue more sequential kernel launches than 12 fat ones, and we are wall-clock-bound rather than parameter-bound. Phase 3 therefore measures both shapes.

token target at 11.1k tok/s (measured) at 9k tok/s (planning floor)
1.10 B 27.6 h 34.0 h
1.00 B 25.1 h 30.9 h
0.90 B 22.6 h 27.8 h

Decision: target 1.0 B tokens, with a pre-registered fallback to 0.9 B if Gate 3's end-to-end rate on the frozen config and real input pipeline is below ~10.5k tok/s. Both are inside Β§2's band, and the fallback is chosen now rather than later so it cannot become a post-hoc excuse. Even the optimistic column leaves under 5 h of slack in a 30 h week, which the interruption budget in Β§3.1 will eat; so the plan assumes a two-week run crossing the 2026-09-26 quota reset, with a checkpoint at every 10 % and a rolling latest, exactly as Β§2 requires.

6.1 Amendment after Gate 3's measurements β€” the fallback trigger fired, and here is what was done about it

Read Β§6 above first: it is kept verbatim, including its numbers, which are now known to be wrong. The 11,062 tok/s figure came from p0c's raw training loop at seq 1024 on a 113M model β€” no Trainer, no real input pipeline, no fp16-master/GradScaler bookkeeping, no seq-2048 activations. Gate 3 measured the frozen 106,194,240-parameter shape inside the actual stack on 2Γ—T4 (ledger rows 10-14, memory/QUOTA.md):

config tok/s peak GB 1.0 B tokens
sdpa @2048 micro 2 (the Β§6 plan-of-record) 4,071 14.22 68.2 h
eager + grad-ckpt @2048 micro 2 4,679 10.04 59.4 h
eager @1024 micro 4 7,732 12.25 35.9 h
eager + grad-ckpt @1024 micro 8 7,828 7.12 35.5 h
eager @2048 micro 2 OOM β€” β€”

Three consequences, in the order that matters:

  1. Sequence length changed to 1024 β€” that is D-011, decided on these numbers and documented in memory/DECISIONS.md. Micro-batch 4 Γ— accum 32 Γ— 2 cards keeps tokens/step at 262,144, so the batch size in tokens, the ~3,815 steps, the LR and the trapezoid are all unchanged. Β§2.3's parameter table is unaffected β€” it never depended on the measurement.
  2. The pre-registered fallback condition in Β§6 fired. "Below ~10.5k tok/s" is satisfied: every measured seq-1024 cell is ~7.7-7.8k. Mechanically that means 0.9 B.
  3. The premise of that trigger is gone. 10.5k was chosen so 1.0 B would fit one 30 h week with enough slack for interruption overhead. At the measured rate neither 1.0 B (35.9 h) nor 0.9 B (32.3 h) fits week 1, so the trigger no longer discriminates between the two β€” it only buys 3.6 GPU-hours for a 10 % cut in training tokens. Against the two-week schedule in docs/04-run-log.md Β§2 (60 h of quota, 35.9 h of training, ~22 h of slack) the 1.0 B target is affordable.

Therefore, re-registered before launch and before any checkpoint or score exists (D-012): the token target stays 1.0 B, and the fallback trigger is restated against measured throughput rather than against p0c's number. New mechanical rule, same purpose as the old one: if preflight P3's end-to-end rate at the frozen geometry (22L/576, seq 1024, eager, grad-ckpt on, micro 4, accum 32) is below 5,150 tok/s β€” i.e. 1 B tokens needs more than ~54 h of the ~60 h available across the two weeks after allowing 6 h for interruption overhead and Phase 6's GPU evaluation β€” the target drops to 0.9 B at launch, not later. 5,150 tok/s is 33 % below the planning rate, so this is a genuine tripwire, not a formality. What did not change: D-005's benchmark bands, the harness pin (D-010), the mix, the model shape, the optimiser, the schedule, or the one-run rule. If the 0.9 B fallback fires, it fires as a documented arithmetic consequence, and this section is where a reader can check it.

7. Benchmark targets β€” pre-registered before the main run (Β§4)

Recorded now, before any training, so the goalposts cannot move. Full anchor table, protocol and noise floor in Β§7.1–7.3. Sources: archived Open LLM Leaderboard run metadata and Pythia Appendix G (arXiv 2304.01373), read 2026-09-19.

The dominant fact: every published β‰₯100M base anchor was trained on ~300 B tokens. Our budget is ~1 B β€” a 300Γ— gap. Those anchors are therefore ceilings, not expectations. No harness score for any 50–350M base trained at ~1–3 B tokens could be found published anywhere, so these bands are extrapolations from the low end of the size curve and are labelled as such.

Task metric band chance
ARC-Easy acc 26–34 25
ARC-Challenge acc / acc_norm 17–22 / 20–25 25
HellaSwag acc_norm (report acc too) 26–32 (acc 25–30) 25
PIQA acc 52–60 50
WinoGrande acc 49–53 50
MMLU acc 24–27 25
TruthfulQA mc2 / mc1 36–46 / 21–26 β‰ˆ38 / β‰ˆ22
GSM8K exact_match,strict-match 0.0–1.5 β‰ˆ0
  • MMLU, WinoGrande, GSM8K and ARC-Challenge are expected to be indistinguishable from chance. MMLU measures flat 24.9–27.3 from 70M to 1.4B β€” OPT-1.3B scores lower than OPT-125M β€” so it has almost no discriminative power at this scale. GSM8K anchors: 0.23 (OPT-125M), 0.68 (GPT-2, Pythia-410M), 1.52 (Pythia-1.4B), 1.4 even for SmolLM2-135M at 11.2 T tokens. Multi-step arithmetic does not emerge from a 100M base at 1 B tokens; predicting otherwise would be the inflation Β§3.12 forbids.
  • The real signals are ARC-Easy, HellaSwag and PIQA β€” the three where anchors clear chance by a visible margin. Failing to beat chance there is a finding about the mix or the run, not the benchmarks.
  • Noise floor: SE β‰ˆ 1.3 pp ARC-C (n=1172), 1.4 pp WinoGrande (1267), 1.2 pp PIQA (1838), 1.5 pp TruthfulQA (817), 0.45 pp HellaSwag (10042). Differences under ~2 pp on ARC/WinoGrande/TruthfulQA are not results, and no shot count or prompt format gets chosen per-task after seeing scores (Β§3.3).

7.1 Protocol, pinned for Phase 6

Harness EleutherAI lm-evaluation-harness v0.4.13 (2026-08-31) β€” still what published anchors report against; Lighteval v0.13.0 is the active HF alternative but switching would break comparability. Record the git SHA per run; cadence is real (v0.4.10 2026-01 β†’ v0.4.13 2026-08).

Task dataset / config split shots
arc_easy, arc_challenge allenai/ai2_arc test (2376 / 1172) 25
hellaswag Rowan/hellaswag validation (10042; no test split) 10
piqa baber/piqa default validation (1838) 10
winogrande allenai/winogrande, winogrande_xl validation (1267) 5
mmlu group of 57 mmlu_<subject> test 5
truthfulqa_mc1, _mc2 sylinrl/TruthfulQA validation 0 (pinned)
gsm8k openai/gsm8k main test 5

7.2 Traps that would silently falsify the numbers

  1. Shots are NOT pinned in the ARC/HellaSwag/PIQA/WinoGrande/MMLU YAMLs β€” an omitted --num_fewshot silently yields 0-shot. Pass shots explicitly for every task.
  2. No chat templates. --apply_chat_template / --fewshot_as_multiturn are for fine-tuned models (Β§3.9 makes this a base model). gsm8k_cot_llama requires them β†’ not used.
  3. Everything but GSM8K is output_type: multiple_choice = loglikelihood ranking; the model never generates, so only GSM8K is sensitive to generation settings.
  4. MMLU aggregation: the mmlu group uses weight_by_size: True, while archived leaderboard numbers were an unweighted subject mean. Compute and state which was used.
  5. TruthfulQA mc2 changed definition 2024-03-11 (PR #2768); pre-April-2024 mc2 is not comparable. mc1/mc2 differ by ~20 pp and both report under the metric name acc.
  6. Harness PIQA is mean per-item accuracy, not AI2's p_win β€” the harness never computes p_win.
  7. Winogrande: winogrande_xl, single fold, validation; the official leaderboard averages 5 folds, and winogrande_debiased is a different number.
  8. GSM8K: temperature 0, do_sample false, max_gen_toks inherits the global 256; tiny bases loop repetitively and never emit #### , so strict-match β‰ˆ0 while flexible-extract looks inflated. Report both filters or neither; do not add repetition_penalty, which silently deviates from everyone.
  9. Splits are mixed β€” PIQA/HellaSwag/WinoGrande/TruthfulQA run on validation. Never call them test.
  10. Parameter-count conventions differ between suites; Pythia's includes embeddings. Quote ours with the convention attached.

8. Still open before Gate 1 closes

  • Deep-thin-vs-wide throughput result β†’ scheduled as Phase 3 preflight test T1.
  • Session wall-clock cap, from the still-running p0e-session-cap: it sets how much progress one session can make and therefore how often the checkpoint cycle interrupts training. Known so far: a CPU session was still alive at 51 min with no cap hit, so short sessions are not the failure mode; the ceiling is somewhere above that.
  • Closed this session: the credential path (Β§1 last row) β€” jobs retrieve the HF token from the account's own private Kaggle dataset via code/ounce100m_credentials.py; a Hub write from inside a job was verified by anonymous readback (D-006). And the checkpoint arithmetic below.

8.1 Checkpoint size and cadence β€” now measured, and it is cheap

A latest-quality exact-resume checkpoint for a 100M model is **1.6 GB** (fp32 weights 400 MB + two fp32 Adam moments 800 MB + master/grad copies + scheduler/scaler/RNG state, which are all negligible next to the optimizer). Pushed from a Kaggle session at 42.7 MB/s β†’ 37.4 s, and pulled back in 13.8 s, with byte-exact anonymous readback at 200 MB / 800 MB / 1.6 GB (dodosoomro/ounce100m-p1-push-bench, D-007).

Consequences that change design choices rather than just confirming them:

  • 10 checkpoints + latest β‰ˆ 7 min total, ~0.12 GPU-h of a 30 h week. Checkpoint cadence is not a budget problem, so there is no reason to economise on it β€” and rolling latest more often than the required 10 % is nearly free, which is worth doing since every interruption costs at most one roll.
  • A cold resume with an empty disk costs seconds of transfer, not minutes. The expensive part of a resume is the skip_first_batches replay discussed in Β§5, which is why the input format in docs/02-mix-plan.md Β§5 step 6 is designed to make it an index advance.
  • Peak disk: 1.6 GB written + 1.6 GB staged elsewhere on the 19.5 GB volume is comfortable, and the Β§3.13 prune-after-verify step keeps at most one copy resident.