Download docs/01-plan.md from Cion-lab/ounce100m-code: direct link, hf CLI and curl.
- Browser
- Download file 26.9 kB
-
https://huggingface.co/Cion-lab/ounce100m-code/resolve/main/docs/01-plan.md
- Command line
-
hf download hf://Cion-lab/ounce100m-code/docs/01-plan.md
-
curl -L -o 01-plan.md https://huggingface.co/Cion-lab/ounce100m-code/resolve/main/docs/01-plan.md
01 β Plan: architecture, hyperparameters, throughput, targets
Phase 1 deliverable. Every choice below cites a source actually inspected on 2026-09-19, as Gate 1 requires. Items marked β test are decisions whose numbers Phase 3 must validate on the real pipeline before the run freezes; the design itself is settled.
1. The envelope the design has to fit
Measured in Phase 0 (docs/00-platform-notes.md), not assumed:
| Constraint | Value | Consequence |
|---|---|---|
| Accelerator | 2x Tesla T4, cc 7.5, 14.56 GiB usable each | fp16 only; no bf16 tensor cores; no flash-attention (sm80+ only) |
| GPU quota | 108,000 s/week, billed at ~1x container wall-clock (3 samples) | β30 wall-clock h/week, both cards |
| Test budget | β€6 GPU-h lifetime | 0.073 h spent; 5.93 h remains for all of Phase 3 |
| Working disk in a job | 19.5 GB /kaggle/working |
shards consumed a few at a time, never whole (Β§3.13) |
| RAM / CPU | 30 GiB cgroup, 4 vCPU, no swap | input pipeline memory-bounded by design |
| Image | torch 2.10.0+cu128, transformers 5.0.0, datasets 5.0.0, accelerate 1.13.0, triton 3.6.0; no trl/deepspeed/flash-attn/xformers/bnb | prefer the preinstalled stack; each install is risk + round trip |
| Hub from a job | reads anonymous & fine, 35β89 MB/s; writes 401 until given a token | credential path: private Kaggle dataset mount (Β§8) |
Parameter convention. Β§2's 90β110M is including embeddings β the Pythia convention (Pythia was
explicitly renamed to include embedding + unembedding, pythia README).
This matters more at 100M than anywhere else: the tied embedding is 21β37 % of the model across the
candidates below. code/config/param_count.py computes it two ways and both were checked to agree
exactly against transformers 5.0.0 in dodosoomro/ounce100m-p1-param-count.
2. Architecture
2.1 The shape: deep-thin, tied, GQA, SwiGLU, pre-norm
Choice: hidden 576, layers 22, 9 Q heads / 3 KV heads (GQA 3:1), SwiGLU intermediate 1536, RMSNorm pre-norm eps 1e-5, attention_bias false, RoPE theta 10,000, tied embeddings, vocab 49,152, seq 2048.
Evidence, and it is convergent rather than a single citation:
- Deep-thin beats wide at this size. Meta's MobileLLM ablated exactly this at a fixed 135M: 30 layers Γ d512 scored 44.8 avg vs 43.9 for 12 Γ d768, with embedding-sharing removing 13 % of parameters at ~equal accuracy (arXiv 2402.14905, HIGH). Its shipped 125M model is 30 layers, d576, 9Q/3KV GQA, SwiGLU, tied, 124.6M total.
- The same shape is what HF's SmolLM2-135M uses: 30 layers, d576, 9Q/3KV,
intermediate_size1536 (2.67Γd), silu,attention_bias=false,tie_word_embeddings=true, RMSNorm 1e-5, model config + arXiv 2502.02737 Β§6. Two independent labs landing on d576 + GQA 3:1 + SwiGLU at ~130M is the strongest available signal. - Our budget is 110M, not 135M, so the same d576/GQA3:1/SwiGLU-1536 stack is cut to 22 layers rather than 30 β verified count below β and RoPE ΞΈ drops to 10,000 to match the shorter context (SmolLM v1 used ΞΈ=10,000 at ctx 2048; SmolLM2's 100,000 accompanies ctx 8192, which we cannot afford and do not need for these benchmarks).
- GQA at 3:1 is a parameter choice as much as an attention choice, since k/v are
d_head Γ n_kvwide; MobileLLM reports GQA-plus-enlarged-d as +0.4 avg.
Not chosen: Mamba/SSM or hybrid. The official kernels do build for sm_75
(state-spaces/mamba setup.py), which corrects the common
assumption that they are Ampere-only β but they are absent from the image (source install only), and
the research found no published 50β200M SSM/hybrid with released hyperparameters to imitate. That is
research risk, not a default, and Β§4 forbids the kind of custom code where autonomous projects die.
2.2 Verified parameter counts
Closed form and sum(p.numel()) agreed exactly on all eight shapes, in the target environment.
Tying confirmed genuinely tied, not merely declared.
| candidate | H | L | heads | kv | FFN | vocab | total params | emb share | in 90β110M |
|---|---|---|---|---|---|---|---|---|---|
| Phase-0 probe shape | 768 | 12 | 12 | 3 | 2048 | 50257 | 112,934,400 | 34.2 % | β |
| A | 768 | 11 | 12 | 3 | 2048 | 50257 | 106,739,712 | 36.2 % | β |
| B | 768 | 12 | 12 | 3 | 1792 | 50257 | 105,856,512 | 36.5 % | β |
| C | 768 | 12 | 12 | 2 | 1824 | 50257 | 105,561,600 | 36.6 % | β |
| D | 640 | 16 | 10 | 4 | 2048 | 50257 | 113,450,240 | 28.4 % | β |
| E | 768 | 12 | 12 | 3 | 2048 | 32768 | 99,502,848 | 25.3 % | β |
| F | 896 | 9 | 14 | 2 | 2432 | 50257 | 120,397,312 | 37.4 % | β |
| G | 1024 | 8 | 8 | 2 | 2816 | 50257 | 141,658,112 | 36.3 % | β |
The binding insight: the tokenizer decides how much model you are allowed to have. At 50,257 vocab
and d768 the tied embedding is 38.6M β a third of the entire budget β which is why every 50k candidate
clusters at 105β113M and must trim depth or FFN. MobileLLM's own ablation reached the same conclusion
(embedding sharing: β13 % parameters at ~equal accuracy). Choosing a smaller hidden is therefore not
a weakening, it is how the embedding tax gets paid for.
β test: the frozen shape's exact count was re-derived from the constructed model β see Β§2.3.
2.3 The frozen shape, counted to the unit
dodosoomro/ounce100m-p1-param-count v2 built each candidate as a real LlamaForCausalLM under
transformers 5.0.0. Closed form and sum(p.numel()) agree exactly on every row.
| candidate | H | L | Q/KV | FFN | vocab | total params | emb share | in 90β110M |
|---|---|---|---|---|---|---|---|---|
| H | 576 | 20 | 9/3 | 1536 | 49,152 | 99,114,048 | 28.6 % | β |
| I β CHOSEN | 576 | 22 | 9/3 | 1536 | 49,152 | 106,194,240 | 26.7 % | β |
| J | 576 | 24 | 9/3 | 1536 | 49,152 | 113,274,432 | 25.0 % | β over |
| K | 576 | 26 | 9/3 | 1536 | 49,152 | 120,354,624 | 23.5 % | β over |
| L | 640 | 20 | 10/4 | 1728 | 49,152 | 120,776,320 | 26.0 % | β over |
Frozen: 106,194,240 parameters (embeddings included, tied), counted as
vocabΓh once + L Γ (attn + SwiGLU) + norms + final norm, and confirmed by constructing the model.
The counting method is stated here because Β§2 requires the number and how it was counted.
Two consequences worth recording:
- Depth is capped at 22 by the budget, not by taste. Each layer at d576 costs 3,538,944 parameters, so 24 layers is 113.3M β over. That is a hard stop, and it means the Phase-3 "deep-thin vs wide" throughput test (T1) must compare H (20 layers, 99.1M) against I (22 layers, 106.2M), both in-budget, rather than the 12-layer d768 probe shape which cannot fit the SmolLM2 vocab without trimming elsewhere.
- Widening is not a free alternative to deepening. Candidate L at d640 Γ 20 layers is over budget
(120.8M) despite having fewer layers than I, because attention and MLP cost scale with
hΒ²/hΒ·ff. So the deep-thin choice is forced from two directions: it is better per parameter (MobileLLM's ablation) and it is the only way to spend the 78M non-embedding budget on 22 layers.
3. Tokenizer
Choice: reuse the SmolLM2 BPE tokenizer β vocab 49,152, byte-level, Apache-2.0, from
HuggingFaceTB/SmolLM2-135M's tokenizer.json;
trained on SmolCorpus per arXiv 2502.02737. Not trained by us.
Β§4 says maximum reuse, minimal custom code, and training a tokenizer from scratch is precisely the
kind of optional machinery that adds failure surface without adding capability. Candidates checked today:
SmolLM2 49,152 / Apache-2.0; gpt2 50,257 / MIT; pythia-160m 50,304 / Apache-2.0;
t5-v1_1-small 32,128 Unigram (not byte-level, carries extra_ids sentinels); Qwen3-0.6B 151,936
(disqualifying β its embedding alone would be ~78M at d512); TinyLlama 32,000 but Llama-2-derived and
redistribution status unverified β excluded.
SmolLM2 wins on three counts: licence is clean and permissive, it is byte-level and English-curated at exactly this model scale, and sharing it with a family of released 135M models gives the benchmark numbers somebody else's token statistics to be compared against. Its 49,152 Γ 576 tied embedding is 28.3M = 21 % of budget, versus 34β37 % for the 50k/d768 shapes.
4. Optimizer, batch, schedule, precision
| choice | justification | |
|---|---|---|
| Optimizer | AdamW, Ξ²=(0.9, 0.95), wd 0.1, clip 1.0 | Pythia-160M and SmolLM2 both use Ξ²β=0.95 at small scale (2304.01373, 2502.02737) |
| Peak LR | 6e-4 | bracketed from three directions: Cerebras-111M 6.0e-4, Pythia-160M 6e-4, and the 150M low-tokens/param run of Sardana et al. at 4.6e-4 (2304.03208, 2401.00448). The 2e-3/3e-3 figures that appear in small-model recipes belong to 1β2 T-token runs, which this is not. |
| Warmup | 2 % of steps, floor 100 steps | Pythia used 1 %; SmolLM2 used 2,000 steps at 2M tok/batch (~0.2 %). Warmer than 1 % because fp16 without bf16 needs the ramp β Phase 0 showed an unramped/unscaled fp16 run going straight to NaN (E-006). |
| Schedule | trapezoid / WSD: constant, then linear decay to 0 over the final 20 % β not cosine-to-10 % | SmolLM1 used a trapezoid with 20 % cooldown; SmolLM2 used WSD (blog/smollm, 2502.02737 App. A). Straight to Zero finds linear-decay-to-zero beats cosine-to-10 % when peak LR is optimal, across sizes/batch/data (2502.15938); COLT 2026 theory says decay shape barely matters but overly slow terminal decay causes schedule-induced capacity saturation (2602.06797). Its caveat applies: their evidence is at/above compute-optimal tokens/param, we are at half. |
| Global batch | β262 k tokens/step (~3,815 steps for 1 B tokens) | Critical batch size fits B* = 621.341Β·N^0.087, i.e. nearly independent of model size (2410.21676) β no reason to chase 2β4M-token batches. Sardana et al. deliberately used smaller batches for smaller models "so that low-token-count training runs see enough steps". At 10 tokens/param we are such a run. |
| Micro-batch | 2 Γ 2048 per GPU Γ 2 GPUs Γ 32 accumulation | Set by memory, not taste: Phase 0 measured bs4Γ1024 at 11.5β13.8 GB of 14.56 GB and bs8 OOMed outright. β test |
| Precision | fp16 autocast + fp32 master weights + GradScaler + grad clip |
Only mixed-precision option on Turing; transformers v5 documents this path by name ("fall back to fp16 on older hardware like V100 or T4") and warns load the model in fp32 or autocast is a no-op (v5 mixed_precision docs). Verified finite in Phase 0 probe C; bf16 runs but is emulated β must be asserted off, not defaulted. |
| Seed / RNG | single fixed seed, recorded | Β§3.1 requires "same RNG semantics" across resumes |
Undertrained by design, and honestly so. 1 B tokens on ~100M β 10 tokens/param, about half Chinchilla-optimal (Cerebras-111M's compute-optimal point was 2.2 B tokens = 20 t/p). That is mildly undertrained, not pathological β the genuine anomaly is 2025-26 practice, which overtrains tiny models by three orders of magnitude (SmolLM2-135M at 2 T tokens β 14,800 t/p). Sardana et al. trained a 150M/d768/12L model across 3β10,000 t/p and quality kept improving; Gadre et al. show scaling laws extrapolate across over-training (2403.08540). We are on the left-hand side of that curve because the quota puts us there, and the report should say so.
5. Training stack
Choice: torchrun --nproc_per_node=2 + Hugging Face Trainer / TrainingArguments, fp16=True,
on a pre-tokenised map-style Dataset. Minimal glue, no framework authored here (Β§4).
- Everything needed is already in the image: transformers 5.0.0, accelerate 1.13.0, datasets 5.0.0.
- v5
Trainerrestores RNG (_save_rng_state/_load_rng_load), optimizer, LR scheduler and the fp16 GradScaler β the exact set Β§3.1 demands β viaAccelerator.save_state. Verified against v5.17.0 source; the image's 5.0.0 is ~8 patch releases behind but the same major. β test at Gate 3. - The resume trap, stated plainly: v5 recovers data position with
skip_first_batches, which replays discarded batches β O(steps) wasted work on anIterableDataset, up to 3,815 steps of pure I/O after every interruption. Mitigation is architectural, not a flag: a map-style, concatenated token store makes skipping an index advance rather than a decode. Adataset_shardlayout with an explicit recorded cursor is what Gate 3 must prove, and Β§Phase 2's "exact positional resumption" requirement is the reason it is designed that way up front. - v5 migration traps to check in preflight (MIGRATION_GUIDE_V5.md, HF Transformers v5 blog): PyTorch-only backend;
slow tokenizers gone (irrelevant, we use
tokenizers); severalTrainingArgumentsremoved without a deprecation cycle; attention moved toAttentionInterface;remove_unused_columnsstill defaults True (silently drops dataset columns a loss function needs);lr_scheduler_typedefault is still"linear", not cosine β so an unstated schedule is a linear decay to zero over the whole run; newtrain_sampling_strategyreplaces thegroup_by_lengthbool;restore_callback_states_from_checkpointis required to restore scheduler state properly. - Disqualified: torchtitan β
training_dtypeis typedLiteral["bfloat16","float32"], i.e. no fp16 path at all, plus Hopper-oriented CI β unusable on T4 (torchtitan config). unsloth β fine-tuning/GRPO only, no from-scratch pretraining path. nanoGPT β dormant since 2025-11, no HF export, no distributed input pipeline. FSDP β ~100M params in fp32 master + grads + AdamW is ~1.8 GB; DDP on 2Γ16 GB does not need sharding, and sharding would add failure surface for nothing. DeepSpeed / flash-attn / xformers / trl β absent from the image; installing them is risk without a capability we lack.
6. Throughput, and the token target it implies
Measured on a main-run-shaped model, seq 1024, bs4/card, DDP over both T4s: 11,062 tok/s
aggregate (5,531/rank; single-GPU 6,478 SDPA / 7,239 eager; 85 % DDP efficiency).
Sanity-checking that against compute: 6ND β 6 Γ 1.13e8 Γ 11,062 β 7.5e12 FLOP/s, which is ~6 % of a
2ΓT4 fp16 tensor peak (65 TFLOPS/card β an unverified datasheet prior; the spec page could not be
retrieved). Independent corroboration: a documented 163M / ctx-1024 run reached ~19.9k tok/s on a single
RTX 3090 (gilesthomas.com),
and Pythia-160M's 1,030 A100-hours for 300 B tokens implies ~81k tok/s per A100. So ~11k tok/s on two
Turing cards is the right order of magnitude, and the low implied MFU means this is a floor with headroom,
not a ceiling β which is the opposite of the risk I most wanted to rule out.
The tension to resolve by measurement, not argument: deep-thin is better per parameter (MobileLLM's ablation) but worse per second on 2018 hardware, because 22 small layers issue more sequential kernel launches than 12 fat ones, and we are wall-clock-bound rather than parameter-bound. Phase 3 therefore measures both shapes.
| token target | at 11.1k tok/s (measured) | at 9k tok/s (planning floor) |
|---|---|---|
| 1.10 B | 27.6 h | 34.0 h |
| 1.00 B | 25.1 h | 30.9 h |
| 0.90 B | 22.6 h | 27.8 h |
Decision: target 1.0 B tokens, with a pre-registered fallback to 0.9 B if Gate 3's end-to-end rate on
the frozen config and real input pipeline is below ~10.5k tok/s. Both are inside Β§2's band, and the
fallback is chosen now rather than later so it cannot become a post-hoc excuse. Even the optimistic
column leaves under 5 h of slack in a 30 h week, which the interruption budget in Β§3.1 will eat; so the
plan assumes a two-week run crossing the 2026-09-26 quota reset, with a checkpoint at every 10 % and
a rolling latest, exactly as Β§2 requires.
6.1 Amendment after Gate 3's measurements β the fallback trigger fired, and here is what was done about it
Read Β§6 above first: it is kept verbatim, including its numbers, which are now known to be wrong. The
11,062 tok/s figure came from p0c's raw training loop at seq 1024 on a 113M model β no Trainer, no
real input pipeline, no fp16-master/GradScaler bookkeeping, no seq-2048 activations. Gate 3 measured the
frozen 106,194,240-parameter shape inside the actual stack on 2ΓT4 (ledger rows 10-14, memory/QUOTA.md):
| config | tok/s | peak GB | 1.0 B tokens |
|---|---|---|---|
| sdpa @2048 micro 2 (the Β§6 plan-of-record) | 4,071 | 14.22 | 68.2 h |
| eager + grad-ckpt @2048 micro 2 | 4,679 | 10.04 | 59.4 h |
| eager @1024 micro 4 | 7,732 | 12.25 | 35.9 h |
| eager + grad-ckpt @1024 micro 8 | 7,828 | 7.12 | 35.5 h |
| eager @2048 micro 2 | OOM | β | β |
Three consequences, in the order that matters:
- Sequence length changed to 1024 β that is D-011, decided on these numbers and documented in
memory/DECISIONS.md. Micro-batch 4 Γ accum 32 Γ 2 cards keeps tokens/step at 262,144, so the batch size in tokens, the ~3,815 steps, the LR and the trapezoid are all unchanged. Β§2.3's parameter table is unaffected β it never depended on the measurement. - The pre-registered fallback condition in Β§6 fired. "Below ~10.5k tok/s" is satisfied: every measured seq-1024 cell is ~7.7-7.8k. Mechanically that means 0.9 B.
- The premise of that trigger is gone. 10.5k was chosen so 1.0 B would fit one 30 h week with
enough slack for interruption overhead. At the measured rate neither 1.0 B (35.9 h) nor 0.9 B (32.3 h)
fits week 1, so the trigger no longer discriminates between the two β it only buys 3.6 GPU-hours for a
10 % cut in training tokens. Against the two-week schedule in
docs/04-run-log.mdΒ§2 (60 h of quota, 35.9 h of training, ~22 h of slack) the 1.0 B target is affordable.
Therefore, re-registered before launch and before any checkpoint or score exists (D-012): the token target stays 1.0 B, and the fallback trigger is restated against measured throughput rather than against p0c's number. New mechanical rule, same purpose as the old one: if preflight P3's end-to-end rate at the frozen geometry (22L/576, seq 1024, eager, grad-ckpt on, micro 4, accum 32) is below 5,150 tok/s β i.e. 1 B tokens needs more than ~54 h of the ~60 h available across the two weeks after allowing 6 h for interruption overhead and Phase 6's GPU evaluation β the target drops to 0.9 B at launch, not later. 5,150 tok/s is 33 % below the planning rate, so this is a genuine tripwire, not a formality. What did not change: D-005's benchmark bands, the harness pin (D-010), the mix, the model shape, the optimiser, the schedule, or the one-run rule. If the 0.9 B fallback fires, it fires as a documented arithmetic consequence, and this section is where a reader can check it.
7. Benchmark targets β pre-registered before the main run (Β§4)
Recorded now, before any training, so the goalposts cannot move. Full anchor table, protocol and noise floor in Β§7.1β7.3. Sources: archived Open LLM Leaderboard run metadata and Pythia Appendix G (arXiv 2304.01373), read 2026-09-19.
The dominant fact: every published β₯100M base anchor was trained on ~300 B tokens. Our budget is ~1 B β a 300Γ gap. Those anchors are therefore ceilings, not expectations. No harness score for any 50β350M base trained at ~1β3 B tokens could be found published anywhere, so these bands are extrapolations from the low end of the size curve and are labelled as such.
| Task | metric | band | chance |
|---|---|---|---|
| ARC-Easy | acc |
26β34 | 25 |
| ARC-Challenge | acc / acc_norm |
17β22 / 20β25 | 25 |
| HellaSwag | acc_norm (report acc too) |
26β32 (acc 25β30) | 25 |
| PIQA | acc |
52β60 | 50 |
| WinoGrande | acc |
49β53 | 50 |
| MMLU | acc |
24β27 | 25 |
| TruthfulQA | mc2 / mc1 |
36β46 / 21β26 | β38 / β22 |
| GSM8K | exact_match,strict-match |
0.0β1.5 | β0 |
- MMLU, WinoGrande, GSM8K and ARC-Challenge are expected to be indistinguishable from chance. MMLU measures flat 24.9β27.3 from 70M to 1.4B β OPT-1.3B scores lower than OPT-125M β so it has almost no discriminative power at this scale. GSM8K anchors: 0.23 (OPT-125M), 0.68 (GPT-2, Pythia-410M), 1.52 (Pythia-1.4B), 1.4 even for SmolLM2-135M at 11.2 T tokens. Multi-step arithmetic does not emerge from a 100M base at 1 B tokens; predicting otherwise would be the inflation Β§3.12 forbids.
- The real signals are ARC-Easy, HellaSwag and PIQA β the three where anchors clear chance by a visible margin. Failing to beat chance there is a finding about the mix or the run, not the benchmarks.
- Noise floor: SE β 1.3 pp ARC-C (n=1172), 1.4 pp WinoGrande (1267), 1.2 pp PIQA (1838), 1.5 pp TruthfulQA (817), 0.45 pp HellaSwag (10042). Differences under ~2 pp on ARC/WinoGrande/TruthfulQA are not results, and no shot count or prompt format gets chosen per-task after seeing scores (Β§3.3).
7.1 Protocol, pinned for Phase 6
Harness EleutherAI lm-evaluation-harness v0.4.13 (2026-08-31) β still what published anchors
report against; Lighteval v0.13.0 is the active HF alternative but switching would break comparability.
Record the git SHA per run; cadence is real (v0.4.10 2026-01 β v0.4.13 2026-08).
| Task | dataset / config | split | shots |
|---|---|---|---|
arc_easy, arc_challenge |
allenai/ai2_arc |
test (2376 / 1172) | 25 |
hellaswag |
Rowan/hellaswag |
validation (10042; no test split) | 10 |
piqa |
baber/piqa default |
validation (1838) | 10 |
winogrande |
allenai/winogrande, winogrande_xl |
validation (1267) | 5 |
mmlu |
group of 57 mmlu_<subject> |
test | 5 |
truthfulqa_mc1, _mc2 |
sylinrl/TruthfulQA |
validation | 0 (pinned) |
gsm8k |
openai/gsm8k main |
test | 5 |
7.2 Traps that would silently falsify the numbers
- Shots are NOT pinned in the ARC/HellaSwag/PIQA/WinoGrande/MMLU YAMLs β an omitted
--num_fewshotsilently yields 0-shot. Pass shots explicitly for every task. - No chat templates.
--apply_chat_template/--fewshot_as_multiturnare for fine-tuned models (Β§3.9 makes this a base model).gsm8k_cot_llamarequires them β not used. - Everything but GSM8K is
output_type: multiple_choice= loglikelihood ranking; the model never generates, so only GSM8K is sensitive to generation settings. - MMLU aggregation: the
mmlugroup usesweight_by_size: True, while archived leaderboard numbers were an unweighted subject mean. Compute and state which was used. - TruthfulQA mc2 changed definition 2024-03-11 (PR #2768); pre-April-2024 mc2 is not comparable.
mc1/mc2 differ by ~20 pp and both report under the metric name
acc. - Harness PIQA is mean per-item accuracy, not AI2's
p_winβ the harness never computesp_win. - Winogrande:
winogrande_xl, single fold, validation; the official leaderboard averages 5 folds, andwinogrande_debiasedis a different number. - GSM8K:
temperature 0,do_sample false,max_gen_toksinherits the global 256; tiny bases loop repetitively and never emit####, so strict-match β0 while flexible-extract looks inflated. Report both filters or neither; do not addrepetition_penalty, which silently deviates from everyone. - Splits are mixed β PIQA/HellaSwag/WinoGrande/TruthfulQA run on validation. Never call them test.
- Parameter-count conventions differ between suites; Pythia's includes embeddings. Quote ours with the convention attached.
8. Still open before Gate 1 closes
Deep-thin-vs-wide throughput resultβ scheduled as Phase 3 preflight test T1.- Session wall-clock cap, from the still-running
p0e-session-cap: it sets how much progress one session can make and therefore how often the checkpoint cycle interrupts training. Known so far: a CPU session was still alive at 51 min with no cap hit, so short sessions are not the failure mode; the ceiling is somewhere above that. - Closed this session: the credential path (Β§1 last row) β jobs retrieve the HF token from the
account's own private Kaggle dataset via
code/ounce100m_credentials.py; a Hub write from inside a job was verified by anonymous readback (D-006). And the checkpoint arithmetic below.
8.1 Checkpoint size and cadence β now measured, and it is cheap
A latest-quality exact-resume checkpoint for a 100M model is **1.6 GB** (fp32 weights 400 MB + two
fp32 Adam moments 800 MB + master/grad copies + scheduler/scaler/RNG state, which are all negligible
next to the optimizer). Pushed from a Kaggle session at 42.7 MB/s β 37.4 s, and pulled back in
13.8 s, with byte-exact anonymous readback at 200 MB / 800 MB / 1.6 GB
(dodosoomro/ounce100m-p1-push-bench, D-007).
Consequences that change design choices rather than just confirming them:
- 10 checkpoints +
latestβ 7 min total, ~0.12 GPU-h of a 30 h week. Checkpoint cadence is not a budget problem, so there is no reason to economise on it β and rollinglatestmore often than the required 10 % is nearly free, which is worth doing since every interruption costs at most one roll. - A cold resume with an empty disk costs seconds of transfer, not minutes. The expensive part of a resume
is the
skip_first_batchesreplay discussed in Β§5, which is why the input format indocs/02-mix-plan.mdΒ§5 step 6 is designed to make it an index advance. - Peak disk: 1.6 GB written + 1.6 GB staged elsewhere on the 19.5 GB volume is comfortable, and the Β§3.13 prune-after-verify step keeps at most one copy resident.