| # vanilla-small-20B-valmatched — dense baselines for the small 20B cell |
|
|
| 12L/d=1024 GPT, `sparsity_mode=none`, trained on 20B tokens of FineWeb-Edu. Produces |
| **both** arms of the small-cell comparison in one run: the val-matched dense |
| baselines (early stops) and the token-matched dense baseline (full budget). |
|
|
| - **Best val_loss: 2.8586** (iter 12200); final at 20B **2.8664** |
| - 12,716 iters, 1,572,864 tok/iter → 20,000,538,624 tokens |
| - 2× NVIDIA H200 (143 GB), wall clock 8h32m (12:29 → 21:01 PDT, 2026-07-24) |
| - Wandb: [`9iezcto0`](https://wandb.ai/markhenrysoftware/sparse-nanogpt/runs/9iezcto0) |
| - HF: [markhenry/vanilla-small-20B-valmatched](https://huggingface.co/markhenry/vanilla-small-20B-valmatched) |
| |
| ## Headline numbers |
| |
| | Run | Val loss | Δ vs dense@20B | |
| |---|---|---| |
| | `vanilla-small-20B` (this, iter 12200) | **2.8586** | — | |
| | `vanilla-small-20B` (this, final 20B) | 2.8664 | +0.008 (noise) | |
| | `cayley-small-3L-mlp_in-20B` | 3.1330 | +0.274 (worse) | |
| | `cayley-small-2L-mlp_in-20B` | 3.1584 | +0.300 (worse) | |
| |
| At matched 20B tokens the dense model is **0.274–0.300 nats better**. That is the |
| alignment tax in loss terms at this scale, and it is the expected direction: the |
| CayleySAE bottleneck buys disentanglement by spending capability. |
| |
| ## Alignment tax (token efficiency) |
| |
| Dense reaches each sibling's final val far inside the budget: |
| |
| | Sibling | its val | dense first eval ≤ | interpolated crossing | tax | |
| |---|---|---|---|---| |
| | 2L | 3.1584 | iter 1900 (3.1536) | ~iter 1840, 2.89B | **~6.9×** | |
| | 3L | 3.1330 | iter 2300 (3.1162) | ~iter 2215, 3.48B | **~5.7×** | |
| |
| Both well inside the <15× target. Crossings are linearly interpolated between |
| adjacent evals; eval spacing (100 iters = 157M tokens) caps resolution at ~±5%, |
| so treat these as one significant figure. Quoting the *saved-checkpoint* iter |
| instead would understate the tax (6.7× / 5.5×). |
| |
| ## The save-at-val ladder, and what it taught us |
| |
| Four rungs were requested: each sibling's final val plus a companion 0.01 nats |
| below (3.1584 / 3.1484 / 3.1330 / 3.1230). All four fired inside 400 iters. |
| |
| | rung | iter | tokens | val at trigger | overshoot | |
| |---|---|---|---|---| |
| | 3.1584 | 1900 | 2.99B | 3.1536 | 0.005 | |
| | 3.1484 | 2100 | 3.30B | 3.1354 | 0.013 | |
| | 3.1330 | 2300 | 3.62B | 3.1162 | 0.017 | |
| | 3.1230 | 2300 | 3.62B | 3.1162 | — | |
| |
| **The 3L pair collapsed into one model.** 3.1330 and 3.1230 fired on the same |
| eval; weights are bit-identical (md5 `4d7c75acd798` over the sorted state dict). |
| Only 3 distinct checkpoints exist, and only 3 were uploaded. |
| |
| Root cause is *eval spacing, not noise*: dense descends ~0.017–0.020 nats per |
| 100-iter window in this region, wider than the 0.01-nat companion gap. A |
| companion can only be distinct if the per-window descent is smaller than the |
| gap. **For the large sibling, either drop companions or tighten |
| `--eval-interval` to ~25–50 in the crossing region.** |
| |
| The 2L pair did stay distinct, and earns its keep for a different reason than |
| designed. `ckpt_val_3.1584.pt` beats 2L by only 0.0048 nats — a statistical tie |
| at 4.9M eval tokens, so it does not support "dense was at least as good." |
| `ckpt_val_3.1484.pt` beats it by 0.0230 (~5× the noise floor) and does. The tight |
| rung is for the tax crossing; the companion is for capability parity. |
| |
| ## Eval noise floor |
| |
| The run-end sequence: 2.8603, 2.8676, 2.8665, 2.8652, 2.8641, 2.8664 — spread |
| 0.0073 with no trend, so **~0.005–0.008 nats** at 150 eval iters × bs 32 × 1024. |
| In-crossing region the same wobble shows as 3.1354 → **3.1359** → 3.1162. |
| |
| Note in-run eval is master-rank only (`train.py:899`), so eval tokens do **not** |
| scale with world size: 150 × 32 × 1024 = **4.9M**, not the 47M the parent |
| script's header claimed (it multiplied by 8 ranks). Any threshold decision from |
| in-run evals inherits the 4.9M noise floor. |
| |
| ## Training health |
| |
| - **Throughput** 688–697k tok/s from iter 10 to the end; no drift, no warmup |
| pathology (compile warmup cost only iter 0 at 77k). Exactly half the 4×H200 |
| 1,389k of report 37, i.e. clean linear scaling in world size. |
| - **grad_norm** fell to ~0.20 by iter 3400 and ~0.09 at the end; monotone, no |
| spikes. Peak VRAM 27.6 GB / 143 GB per GPU — bs could go far higher, but was |
| held at 32 to keep tok/iter identical to the recorded recipe. |
| - **Loss** monotone modulo the noise floor throughout. |
|
|
| ## Node-shape delta (2× H200, not 4×) |
|
|
| Recipe unchanged from `train_vanilla_small_20B_valmatched_4xh200.sh`. Only |
| `--nproc_per_node` 4→2. Because `ga` is the **global** micro-step count, tok/iter |
| stays 1,572,864 and the LR schedule occupies the same 12,716-iter space; each rank |
| runs 24 micro-steps instead of 12. Wall clock doubled, nothing else changed. |
| Script: `mh/scripts/train_vanilla_small_20B_valmatched_2xh200.sh`. |
|
|
| ## Two defects found in the existing scripts |
|
|
| **1. The vanilla-small script family cannot launch.** `train_vanilla_small_20B.sh`, |
| `..._valmatched.sh`, and `..._valmatched_4xh200.sh` all pass `--n-embd 1024` |
| without `--n-head`. `train.py` applies *mode-dependent* arch defaults |
| (`train.py:1131`): `cayley` → `n_head=8`, else `n_head=12`. With |
| `sparsity_mode=none` that resolves to 12, and `assert n_embd % n_head == 0` fails |
| instantly since 1024 % 12 ≠0. The cayley-small siblings silently took the cayley |
| branch and got `n_head=8`, which is why only the vanilla side is broken. Fixed in |
| the 2xh200 script with an explicit `--n-head 8` (verified 202.53M params, matching |
| the siblings exactly). **The three parents remain un-runnable.** |
|
|
| **2. `quick_eval` silently skipped for the entire run** — 6 attempts, all |
| `import failed (No module named 'evals'); skipping`. The script does `cd deeptopk` |
| then `uv run`, which resolves to `deeptopk/.venv`; the `evals` package is declared |
| by the *root* `pyproject.toml` and is only present in the root `.venv`. So no |
| pile/hellaswag/lambada during training. Harmless here (val_loss unaffected, |
| report 42 does these post-hoc) but it means the quick-eval columns are absent |
| from this wandb run. |
| |
| ## Artifacts |
| |
| Five checkpoints in one HF repo, all also at |
| `sparse-nanogpt-private/out/vanilla-small-20B/` (md5-verified against `/dev/shm`): |
| |
| | file | iter | tokens | val | |
| |---|---|---|---| |
| | `ckpt_val_3.1584.pt` | 1900 | 2.99B | 3.1536 | |
| | `ckpt_val_3.1484.pt` | 2100 | 3.30B | 3.1354 | |
| | `ckpt_val_3.1330.pt` | 2300 | 3.62B | 3.1162 | |
| | `ckpt_final_20B.pt` | 12716 | 20.001B | 2.8664 | |
| | `ckpt_best.pt` | 12200 | 19.19B | 2.8586 | |
| |
| ## Recommended use |
| |
| - **Interpretability comparison vs 2L** → `ckpt_val_3.1484.pt` (unambiguously |
| better than the sibling, so no handicapped-baseline objection). |
| - **Interpretability comparison vs 3L** → `ckpt_val_3.1330.pt`. |
| - **Single comparator for both** → `ckpt_val_3.1330.pt` (val 3.1162 beats 2L by |
| 0.042 and 3L by 0.017). |
| - **Equal-budget tax in nats** → `ckpt_final_20B.pt`. |
| - **Token-efficiency tax** → interpolated crossings above, *not* the saved iters. |
| |
| All val numbers above are in-run trigger values at 4.9M eval tokens. Re-evaluate |
| with a large eval before quoting any as a val-matched number; the purpose of that |
| pass is to confirm the dense ≥ cayley inequality still holds, not to find the |
| closest match. |
| |
| ## Data note |
| |
| `fineweb-edu-25B` was re-tokenized on this node from `sample-100BT` |
| (24,975,629,888 train / 24,433,651 val tokens, 40 streams, ~6 min). Same subset |
| and distribution as the siblings, but a different stream count means a different |
| shard→token layout, so this is not a byte-identical data replay. Val comparisons |
| hold; exact-trajectory comparisons do not. |
| |