File size: 6,230 Bytes
465cd90 4b57a0e 465cd90 4b57a0e 465cd90 4b57a0e 465cd90 4b57a0e 465cd90 4b57a0e e47adb4 465cd90 4b57a0e 33cf988 4b57a0e e47adb4 4b57a0e 15d9fdd 33cf988 4b57a0e 33cf988 4b57a0e 33cf988 e47adb4 4b57a0e e47adb4 4b57a0e e47adb4 33cf988 e47adb4 15d9fdd 33cf988 e47adb4 15d9fdd e47adb4 4b57a0e 15d9fdd e47adb4 33cf988 4b57a0e e47adb4 4b57a0e e47adb4 4b57a0e e47adb4 33cf988 e47adb4 4b57a0e 15d9fdd 33cf988 4b57a0e 15d9fdd 4b57a0e e47adb4 465cd90 15d9fdd e47adb4 15d9fdd e47adb4 465cd90 4b57a0e e47adb4 4b57a0e 465cd90 4b57a0e 465cd90 4b57a0e 465cd90 e47adb4 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 | ---
license: mit
datasets:
- HuggingFaceFW/fineweb-edu
language:
- en
tags:
- gpt
- baseline
- interpretability
- paritytransformer
---
# vanilla-large-20B β dense baselines (24L / d=2048, 1.31B)
Dense (`sparsity_mode=none`) counterparts to the ParityTransformer large-chassis
flagship [`markhenry/cayley-large-2L-mlp_in-20B`](https://huggingface.co/markhenry/cayley-large-2L-mlp_in-20B).
All checkpoints come from **one 20B-token run**, so both baseline definitions
share a single trajectory.
## Checkpoints
Files are named by **iter**. (They were previously named by the `--save-at-val`
threshold that fired; those names misstated the actual loss β `ckpt_val_2.8000.pt`
held a model at 2.7950 β so they were renamed. Two thresholds fired on the same
eval and produced byte-identical files; only one copy is kept.)
| file | iter | tokens | **measured val_loss** | vs flagship |
|---|---|---|---|---|
| `ckpt_val_3400.pt` | 3400 | 5.35B | **2.7981** | +0.0023 worse (7Ο) |
| `ckpt_val_3800.pt` | 3800 | 5.98B | **2.7764** | β0.0194 better (60Ο) |
| `ckpt_final_20B.pt` | 12716 | 20.001B | **2.5419** | β0.2538 better (338Ο) |
Reference, measured in the same pass on the same batches:
| reference | measured val_loss | published (in-run) |
|---|---|---|
| `cayley-large-2L-mlp_in-20B` (flagship) | **2.7957** | 2.8081 |
| `cayley-flagship-1.3b-20b` (2L + aux losses) | **2.8309** | 2.842 |
`ckpt_val_3400.pt` fired a `--save-at-val` trigger but is **0.0023 nats worse**
than the flagship (7Ο) β it does *not* clear it. It does clear the aux-loss
variant by 0.0328 (86Ο). `ckpt_val_3800.pt` clears both.
`ckpt_val_3800.pt` being *better* than the flagship is the useful direction: a
dense model that beats the ParityTransformer on val and is still less
interpretable cannot be dismissed as a handicapped baseline.
`ckpt_best.pt` (iter 12400) is **not** uploaded β `ckpt_final_20B.pt` measures better, so the "best" label was a noise artifact of the 3.28M-token
in-run eval and the file carries no information the final one lacks.
## How these numbers were measured
`evals/language_modeling/eval_val_loss.py` over the **entire** FineWeb-Edu val
split β contiguous non-overlapping windows, bf16 autocast, batch 8 Γ 1024,
`--data fineweb-edu-25B`, seed 0.
| quantity | value |
|---|---|
| tokens per model | **25,161,728** (whole val split) |
| batches per model | 3,072, **byte-identical across all models** |
| absolute SEM (single model) | Β±0.0047 |
| paired SEM (difference of two models) | **Β±0.0003** |
The paired SEM is ~17Γ tighter because shared passage difficulty cancels. Quote
the paired differences, not error bars added in quadrature.
Two reasons the run log disagrees, both of which bite if you mix sources:
| source of disagreement | effect |
|---|---|
| In-run evals are master-rank only, `eval_iters Γ bs Γ seq_len` = **3.28M tokens** regardless of world size; adjacent evals wobble 0.005β0.017 nats. | `--save-at-val` triggers and "best" labels inherit that noise. |
| `best_val_loss` inside a checkpoint is the **run's** best, not that checkpoint's val. | Differencing it against a reference `ckpt.pt` saved at another iter is meaningless. |
The flagship's published **2.8081** is an in-run number; on this protocol it is
**2.7957**. Use 2.7957 against these checkpoints and do not mix the two.
## Alignment tax
**Equal budget (20B tokens both sides):**
| pair | dense | cayley | gap |
|---|---|---|---|
| vs `cayley-large-2L` | 2.5419 | 2.7957 | **0.254 nats** |
| vs `cayley-large-2L`+aux | 2.5419 | 2.8309 | **0.289 nats** |
**Token efficiency** β where dense first reaches the flagship's measured val,
interpolated between the two bracketing measured points:
| sibling | its measured val | dense crossing | tokens | tax |
|---|---|---|---|---|
| `cayley-large-2L` | 2.7957 | iter ~3444 | 5.42B | **~3.7Γ** |
Caveats before this goes in a paper table:
| caveat | detail |
|---|---|
| not scale-invariant | the small chassis gives **6.3β7.5Γ** on the same construction ([`vanilla-small-20B`](https://huggingface.co/markhenry/vanilla-small-20B)); the tax shrinks as the backbone grows |
| not comparable to `vanilla-large-parity-3B`'s 6.7Γ | that run had a 3B budget and was *warming down into* the target, which buys val per token; this one passes the crossing at flat peak LR (warmdown starts iter 6358) |
| resolution | eval spacing (100 iters = 157M tokens) caps it; two significant figures |
## Recipe: identical to the flagship except `sparsity_mode`
| setting | value |
|---|---|
| arch | 24L / 16H / d=2048, learned abs. pos enc, bf16 |
| muon lr | 8e-3 β 1e-4 |
| adamw lr | 3e-4 β 1e-5 |
| schedule | `linear_warmdown`, wf=0.5, warmup 200 |
| tokens/iter | 32 Γ 1024 Γ 48 = **1,572,864** (12,716 iters) |
| data | FineWeb-Edu `sample-100BT`, GPT-2 tok, 24.97B train tokens |
| run | 4Γ H200, 23h57m, ~252k tok/s, peak 67.8 GB/GPU |
| wandb | [`700d4i1i`](https://wandb.ai/markhenrysoftware/sparse-nanogpt/runs/700d4i1i) |
Deltas vs the flagship:
| delta | note |
|---|---|
| `--sparsity-mode none` | drops the five cayley flags β the only intended difference |
| `--eval-interval 100` (vs 500) | pins the val-matched crossing |
| `--save-at-val` | writes the first-crossing checkpoints |
| 4Γ H200 instead of 8 | **no effect on optimization** β `--gradient-accumulation-steps` is the *global* micro-step count, so ga stays 48 and tok/iter is identical |
## Loading
```python
import torch
ck = torch.load("ckpt_val_3800.pt", map_location="cpu", weights_only=False)
ck["iter_num"] # 3800
cfg, sd = ck["model_config"], ck["model"] # GPTConfig fields, 148 tensors
```
Checkpoints carry `optimizer_states` (Muon momentum + AdamW), so they are
resumable, not inference-only.
## Siblings
- [`markhenry/cayley-large-2L-mlp_in-20B`](https://huggingface.co/markhenry/cayley-large-2L-mlp_in-20B) β the flagship this is matched to
- [`markhenry/vanilla-small-20B`](https://huggingface.co/markhenry/vanilla-small-20B) β same construction, small chassis
---
*Measurements and card by Claude Code, from the run log and a paired
re-evaluation of every checkpoint over the whole val split.*
|