File size: 15,812 Bytes
cadb374 5f6be9b cadb374 8773aa4 5f6be9b 9d1cf74 cadb374 5f6be9b cadb374 5f6be9b 9d1cf74 5f6be9b 9d1cf74 5f6be9b 9d1cf74 5f6be9b 9d1cf74 5f6be9b cadb374 5f6be9b cadb374 5f6be9b cadb374 5f6be9b 9d1cf74 5f6be9b cadb374 5f6be9b cadb374 5f6be9b cadb374 5f6be9b 9d1cf74 5f6be9b 8773aa4 5f6be9b cadb374 5f6be9b 9d1cf74 5f6be9b 9d1cf74 5f6be9b 9d1cf74 5f6be9b 9d1cf74 5f6be9b cadb374 5f6be9b cadb374 5f6be9b cadb374 5f6be9b cadb374 5f6be9b cadb374 5f6be9b 9d1cf74 5f6be9b | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 | ---
license: mit
tags:
- nanogpt
- ablation
- attention
- conditional-computation
- gating
- research-artifact
- negative-results
library_name: pytorch
---
# QGate β checkpoints and results from a 19-phase attention-gating ablation
Training artifacts from an independent research program on **query-conditioned attention
gating** in GPT-2-class transformers. 33 checkpoints, the code that produced them, and the
complete result tables β including the runs that did **not** support the hypothesis.
**Code, per-phase result tables, and the full experiment history:**
[github.com/briantkolb/qgate](https://github.com/briantkolb/qgate)
This repository is an archive first and a model release second. Nothing here is intended
for downstream use as a general-purpose language model.
> ### Revision notice β 2026-08-15
> This card was **substantially revised** the day after first publication. The initial
> version claimed the gate "produces a real improvement over an ungated baseline,"
> replicated six ways. That claim was too strong and is retracted. Three findings already
> in the project record had not been incorporated:
> 1. a 20-run controlled matrix on a hardened baseline where the learned gate **harms**;
> 2. a warmup sweep where **every gate tested loses to the ungated baseline**;
> 3. the per-dimension bands were reported ~3Γ too tight, and the "flat loss direction"
> mechanism is contradicted by this repository's own per-layer results.
>
> The corrected thesis is below. Where two files in this corpus disagree, both are now
> shown rather than one being chosen.
## What the data supports
**The gate's benefit is a function of how much fixed structure the baseline already has.**
On a plain GPT-2-class 124M model at 20% warmup it is a large, reproducible improvement. At
1B, on a modern stack, it is marginal and its statistical support is under review. On a
short warmup schedule it is *worse than no gate at all*. On a heavily-optimized speedrun
baseline it is decisively harmful. That ordering is the result; the 124M number alone is not.
A second, narrower finding concerns parameterization: the same blend quantity that a
12-parameter per-layer gate learns a clear profile for is one that a 768-parameter
per-dimension gate fails to relocate from its initialization at this token budget. Adding
capacity made it less learnable, not more.
## The intervention
A single gate applied at the **G1 seam** β post-SDPA, pre-output-projection:
```python
hs = n_embd // n_head # head_dim
self.qa_gate_proj = nn.Linear(hs * 2, hs, bias=False)
gate = torch.sigmoid(self.qa_gate_proj(torch.cat([q, y], dim=-1)))
y = y * gate # before c_proj (W_O)
```
At the 124M scale this costs **98,304 parameters (+0.079%)**. In this corpus the `y` term is
called **A** (the attention output); `code/model_cheap_qa_minimal.py` is a self-contained
single-variant reference implementation with the port surface marked by `G1 SEAM` banners.
## Where the effect is large
**nanoGPT 124M, OpenWebText, 10k iters, warmup 2000, seeds 42/1337/123** (val CE).
Source: `results/csv/priority1_runs_20260322_1355.csv`.
| variant | added params | val CE | vs baseline |
|---|---|---|---|
| baseline | β | 4.4708 Β± 0.0074 | β |
| **cheap_qa** (gate on `cat(q, y)`) | +98,304 (+0.079%) | **4.3897 Β± 0.0051** | **β0.0811 (β1.81%)** |
| full_x (gate on `x`, the published G1 form) | +7,077,888 (+5.71%) | 4.3950 Β± 0.0116 | β0.0758 |
| midtier_q (per-head MLP on `q`) | +196,608 (+0.159%) | 4.4135 Β± 0.0086 | β0.0573 |
| cheap_q (linear gate on `q`) | +49,152 (+0.040%) | 4.4198 Β± 0.0185 | β0.0510 |
The one comparison here that separates cleanly separates on **cost**, not loss: `cheap_qa`
and `full_x` differ by 0.0053 CE β inside seed noise β and by **72Γ in added parameters**.
Parameter counts are exact; loss deltas are estimates from three seeds.
## Where the effect goes away, and reverses
**1. Short warmup β every gate tested loses to the baseline.**
Source: `results/csv/warmup500_bundle_runs_20260324_1937.csv`,
`results/csv/phase6b_grok_qafollowups_w500_runs_20260328_0028.csv`,
`results/csv/phase5_warmup500_toptiers_runs_20260326_1722.csv` (warmup 500, 3 seeds each).
| variant | warmup 500 | warmup 2000 | penalty |
|---|---|---|---|
| **baseline** | **4.5663** | 4.4708 | +0.0955 |
| qa_normed | 4.5902 | 4.3926 | +0.1976 |
| cheap_qa | 4.6130 | 4.3897 | +0.2233 |
| qa_lowrank | 4.6289 | 4.4068 | +0.2221 |
| mlp128_qa | 4.6332 | 4.4032 | +0.2300 |
| full_x | 4.8518 | 4.3950 | +0.4568 |
Short warmup hurts everything, but it hurts every gate roughly twice as much as the
baseline, and at warmup 500 the ungated model wins outright. **The headline improvement is
conditional on a 20% warmup schedule.**
**2. A hardened baseline β the learned gate harms, decisively.**
A 20-run matrix (4 arms Γ 5 seeds) against the `2025-01-04_SoftCap` modded-nanoGPT record on
8ΓH100, August 2β5 2026. Published noise floor for that record: 3.2791 Β± 0.0019 over 80 runs.
| arm | mean val loss | vs baseline |
|---|---|---|
| B β baseline | 3.27790 | β |
| S β static gate (0.5) | 3.27960 | +0.0017 |
| Z β zero-init learned gate | 3.30788 | +0.0300 |
| F β learned gate | 3.30796 | +0.0301 |
**F β S = +0.0284**, 95% CI [+0.0262, +0.0305], t(4) β 36.5, all five seeds positive, **zero
distributional overlap** (worst ungated 3.2822, best gated 3.3049). Gated arms are also ~7%
slower.
The decomposition matters more than the sign: **static scaling costs +0.0017; making the
gate learnable costs +0.0284 β 16Γ more.** The damage comes from *learning* the gate, not
from gating. `F β Z = +0.00008` β random and zero init reach the same solution.
This is the best-powered experiment in the corpus, and it is negative. Primary logs
(`qgate_run3_MATRIX.tar.gz`) are not in this repository; the figures above are from the
August 2β5 session record.
**3. Scale β the ordering among gates does not survive.**
OLMo-2 1B, OpenWebText, 3 seeds, H100 (`rope`, `rmsnorm`, `swiglu`, QK-norm, post-norm):
| variant | mean ppl | sd |
|---|---|---|
| baseline | 44.137 | 0.179 |
| cheap_qa | 43.703 | 0.124 |
| midtier_q | 43.683 | 0.159 |
`cheap_qa` and `midtier_q` differ by 0.02 ppl β p = 0.87, a tie β despite clear separation at
124M. **See the caveats below before using the 1B result for anything.**
## What is contested inside this corpus
Two or more files here disagree. These are shown rather than resolved.
**The A10 comparison.** Two boards, same hardware, same recipe, different conclusions:
| source | baseline | cheap_qa | midtier_q | ordering |
|---|---|---|---|---|
| `results/summaries/priority1_summary_20260322_1355.txt` (A10 ref block) | 4.4675 | 4.4151 | **4.4039** | midtier_q wins |
| `results/phase8_warmup2000_bulk_a10_runs_20260329_0650.csv` (19 variants, 57 runs) | 4.4676 | **4.4155** | 4.4183 | cheap_qa wins |
Baseline and cheap_qa agree to 0.0004 across both. **`midtier_q` differs by 0.0144, and that
single number is what the "ordering inverts on A10" claim rests on.** The larger and later
board does not reproduce the inversion. Additionally, the A10 lane ran a different software
stack (torch 2.7.0+cu128) from the canonical 3060 board (2.5.1+cu121), and the project record
explicitly blocks a hardware-only interpretation until a version-matched rerun exists.
What *is* consistent across both A10 boards: `cheap_qa` beats baseline, and it drops from
1st on the 3060 to 7th of 19 on the A10 (behind `x_full_headspec` at 4.3858).
**The 1B result.** An August 4 review found: the gate-norm "stability" claim retracted
(post-training βWβ_F is statistically indistinguishable from an untouched init draw,
P = 0.494); the win record mislabelled (`cheap_qa` alone is 3/3, sign-test **p = 0.125 β not
significant**; the gated family pooled is 6/6, p = 0.0156); seed 42 resumed from a
`step1536` checkpoint in all three arms and carries the largest margin (dropping it moves the
mean delta 0.433 β 0.285); the advantage is late-emerging (baseline *leads* on 2/3 seeds at
step 500); `midtier_q` logged no gate proxy, so there is no cross-variant control. **Whether
the OLMo gate trained at all is not established** β a frozen random projection is not
excluded. The checkpoints were lost with the rented instance, so this may stay open.
**Whether the 1B run left LR warmup.** One record says the schedule was truncated but
`t_warmup` fixed 200Mβ40M with `expected_max_steps: 1526`; another says warmup was never
exited. The rendered per-run config was never recovered. Unresolved.
**The WSL lane.** Two docs report baseline/cheap_qa as 4.4759/4.3902 and 4.4732/4.3800. The
WSL baseline also shows late-training spikes attributed to WSL2 memory management β
contamination in the direction that widens the gap. This lane also ran torch 2.7.0.
**Seed sensitivity.** On out-of-band seeds 0/1/2 (`results/phase9_...20260328_2327.csv`),
`cheap_qa` is 4.4139 Β± 0.0094 rather than 4.3897, and `qa_normed` (4.4103) beats it. The
Β±0.0051 above is a within-seed-set figure for 42/1337/123.
## The per-dimension result
**Corrected 2026-08-15.** Earlier versions of this card reported far tighter bands. Those
were per-*layer mean* bands presented as individual-value bands. Full vectors, all 12 layers,
3 seeds β 4,608 values per condition:
| phase | init | individual values | per-layer means | mean | val CE |
|---|---|---|---|---|---|
| 19b | 0.50 | β | q 0.4587β0.5202 / y 0.4687β0.5766 | 0.481 / 0.504 | 4.4015 |
| 19d free | ~0.50 | **0.3593β0.6821** | 0.4558β0.5718 | **0.4924** | 4.4085 |
| 19c informed | 0.74 | **0.5745β0.8457** | 0.6995β0.7826 | **0.7317** | 4.3950 |
**What holds:** the *mean* does not move from its initialization β 0.74 β 0.7317, 0.50 β
0.4924. Initialization sets where the distribution sits.
**What does not hold:** "the scales do not move." They spread substantially β 19c covers a
0.27 range, 19d covers 0.32. Individual dimensions differentiate; the center does not shift.
**The mechanism claimed earlier β a flat loss direction β is contradicted by this repository.**
The same blend quantity moves decisively under coarser parameterization:
| parameterization | params | init | converged |
|---|---|---|---|
| phase 14, single scalar | 1 | 0.5 | **0.7391β0.7443** |
| phase 17, per-layer | 12 | 0.5 | q 0.3803β0.5551 (mean 0.4481) / y 0.4126β0.6529 (mean 0.5144) |
| phase 19b/c/d, per-dimension | 768 | 0.50 / 0.74 | mean stays at init |
A scalar moves +0.24 from its initialization. Twelve per-layer values differentiate across a
0.27 range and reproduce a consistent profile. Seven hundred sixty-eight per-dimension values
do not relocate their mean. **The loss is not flat along this axis β the fine parameterization
is not identified at this budget.** Gradient dilution across 1,536 parameters and simple
undertraining are both live explanations and this corpus cannot separate them.
## Honest limitations
1. **The headline 124M result is conditional on warmup 2000.** At warmup 500 the ungated
baseline beats every gate tested.
2. **The best-powered experiment in this corpus is negative** (speedrun matrix, above).
3. **The 1B evidence is weak and partly under review** β see the contested section. Do not
cite it as scale validation.
4. **`cheap_qa`'s margin is seed-set dependent** β 4.3897 on seeds 42/1337/123, 4.4139 on
seeds 0/1/2.
5. **Only 1 of 3 `cheap_qa` seeds, and 0 of 3 baseline seeds, ended with their best
validation loss at the final evaluation.** Most runs were drifting upward late.
6. **Elaborations (phases 14β19d) tie with the simple gate.** Seven variants sit within
0.0032 CE against a within-variant seed sd of 0.005β0.012; phase 14 (4.3865) and 17
(4.3890) are nominally ahead (p = 0.64, 0.94). Off-recipe, warmup 3000 reaches 4.3761.
7. **Standard deviations are population sd at 124M, sample sd at 1B.** A 3-sample sd carries
roughly 52% relative standard error either way.
8. **Head-specific vs head-shared is confounded** by gate input width and is not claimed.
## Traps in the result files
β οΈ `results/phase19b_canonical_*_0709.*` is a **failed run** β six seeds, `returncode=2`,
`best=nan`, no checkpoint. It sits beside the real six-seed data (`..._0711`).
β οΈ `results/summaries/p2_rerun_summary_20260323_1714.txt` ranks `full_x` first at 4.3870 on a
**two-seed partial**. `results/summaries/full_x_seed123_result_20260324_1712.txt` (4.3950,
3 seeds) supersedes it.
β οΈ `results/summaries/a10_replication_summary_20260322_0639.txt` ends with an auto-generated
`PAPER STATEMENTS` block asserting hardware independence and `|delta| < 0.02`. Its own table
forty lines above shows deltas to β0.1257 and prints `Hardware consistency: INVESTIGATE β`.
Trust the tables, not the prose blocks.
β οΈ **Era warning.** Results predating the `beta2=0.99` fix land in the 4.5β4.6 CE band and are
not comparable to clean-recipe results in the 4.38β4.47 band.
β οΈ `checkpoints/canon/out-shakespeare-char/` is the upstream nanoGPT demo, not part of this
study.
## Relation to published work
Gating the SDPA output at G1 is established. **This is a replication and mechanism study, not
a novelty claim.**
- [Qiu et al. 2025](https://arxiv.org/abs/2505.06708) (NeurIPS 2025 Oral) β the G1 gate at
15B-MoE / 1.7B-dense scale, conditioned on **X** (layer input) in all fifteen reported
variants. `full_x` here is that formulation ported to nanoGPT.
- [Bu et al. 2025](https://arxiv.org/abs/2510.09017) β moves the gate input from X to **V**.
- [Zhou et al. 2026](https://arxiv.org/abs/2608.11805) (Tencent Hunyuan, 12 Aug 2026) β adds
an **H-gate** computed from the SDPA output. Up to notation, their `H` is the `y`/A term in
`cat(q, y)`. Subsequent and independent; their 5B/500B result reaches a regime this corpus
cannot.
Two of Qiu et al.'s qualitative findings reproduce here at ~1/10,000 of their token budget:
G1 gating beats baseline (6.026 β 5.761 PPL there; 4.4708 β 4.3897 CE here, at warmup 2000),
and input-*dependent* beats input-*independent* (5.917 β 5.761; 4.4483 β 4.3897). The
input-independent control is an exact structural match β a zero-initialized learnable
`(n_head Γ head_dim)` parameter through a sigmoid in both cases.
`cheap_qa` was confirmed 2026-03-20
(`results/csv/cheap_qa_confirmation_runs_20260320_1854.csv`), before the prior-art search that
located Qiu et al.; the project was then reclassified from a novelty claim to a replication
study. **This establishes no priority** β these results were private until August 2026.
## Layout
```
checkpoints/
canon/ 21 x ckpt.pt base run + phases 14-19d (RTX 3060, Windows)
wsl_bridge/ 12 x ckpt.pt phases 10-12 (RTX 3060, WSL)
code/ model/train for both lanes + the minimal champion module
results/ phase 7-19d run CSVs, summaries and logs
SHA256SUMS.txt integrity manifest for all 33 checkpoints
```
nanoGPT-format `ckpt.pt` at fp32, ~1.4 GiB each, optimizer state included.
```bash
sha256sum -c SHA256SUMS.txt
```
## Open questions
- Does the benefit really track baseline hardness, or is the three-point arc a coincidence of
three different codebases?
- Why does making the gate *learnable* cost 16Γ what the gate itself costs?
- Is the per-dimension non-identifiability gradient dilution or undertraining? A longer run at
one initialization would separate them.
- Did the OLMo gate train at all?
- Why does `cheap_qa` fall from 1st to 7th on the A10 while still beating baseline?
## Attribution
Derived from [nanoGPT](https://github.com/karpathy/nanoGPT) by Andrej Karpathy (MIT).
See `LICENSE`.
|