File size: 15,812 Bytes
cadb374
 
 
 
 
 
 
 
 
5f6be9b
cadb374
 
 
 
 
 
 
 
 
8773aa4
 
 
5f6be9b
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
9d1cf74
cadb374
 
 
 
 
 
 
 
 
 
 
 
5f6be9b
 
 
cadb374
5f6be9b
9d1cf74
5f6be9b
 
9d1cf74
5f6be9b
 
 
 
 
 
 
9d1cf74
5f6be9b
 
 
9d1cf74
5f6be9b
cadb374
5f6be9b
 
 
 
cadb374
5f6be9b
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
cadb374
5f6be9b
 
 
 
 
 
 
9d1cf74
5f6be9b
 
 
cadb374
5f6be9b
 
cadb374
 
 
 
 
 
 
5f6be9b
 
cadb374
5f6be9b
9d1cf74
5f6be9b
8773aa4
5f6be9b
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
cadb374
5f6be9b
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
9d1cf74
 
 
5f6be9b
 
 
 
 
 
 
 
 
 
 
9d1cf74
5f6be9b
 
9d1cf74
5f6be9b
 
9d1cf74
5f6be9b
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
cadb374
 
 
 
 
 
5f6be9b
cadb374
5f6be9b
cadb374
 
 
5f6be9b
cadb374
 
 
 
 
5f6be9b
cadb374
5f6be9b
 
 
 
 
 
 
9d1cf74
 
 
 
5f6be9b
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
---
license: mit
tags:
  - nanogpt
  - ablation
  - attention
  - conditional-computation
  - gating
  - research-artifact
  - negative-results
library_name: pytorch
---

# QGate β€” checkpoints and results from a 19-phase attention-gating ablation

Training artifacts from an independent research program on **query-conditioned attention
gating** in GPT-2-class transformers. 33 checkpoints, the code that produced them, and the
complete result tables β€” including the runs that did **not** support the hypothesis.

**Code, per-phase result tables, and the full experiment history:**
[github.com/briantkolb/qgate](https://github.com/briantkolb/qgate)

This repository is an archive first and a model release second. Nothing here is intended
for downstream use as a general-purpose language model.

> ### Revision notice β€” 2026-08-15
> This card was **substantially revised** the day after first publication. The initial
> version claimed the gate "produces a real improvement over an ungated baseline,"
> replicated six ways. That claim was too strong and is retracted. Three findings already
> in the project record had not been incorporated:
> 1. a 20-run controlled matrix on a hardened baseline where the learned gate **harms**;
> 2. a warmup sweep where **every gate tested loses to the ungated baseline**;
> 3. the per-dimension bands were reported ~3Γ— too tight, and the "flat loss direction"
>    mechanism is contradicted by this repository's own per-layer results.
>
> The corrected thesis is below. Where two files in this corpus disagree, both are now
> shown rather than one being chosen.

## What the data supports

**The gate's benefit is a function of how much fixed structure the baseline already has.**
On a plain GPT-2-class 124M model at 20% warmup it is a large, reproducible improvement. At
1B, on a modern stack, it is marginal and its statistical support is under review. On a
short warmup schedule it is *worse than no gate at all*. On a heavily-optimized speedrun
baseline it is decisively harmful. That ordering is the result; the 124M number alone is not.

A second, narrower finding concerns parameterization: the same blend quantity that a
12-parameter per-layer gate learns a clear profile for is one that a 768-parameter
per-dimension gate fails to relocate from its initialization at this token budget. Adding
capacity made it less learnable, not more.

## The intervention

A single gate applied at the **G1 seam** β€” post-SDPA, pre-output-projection:

```python
hs = n_embd // n_head                       # head_dim
self.qa_gate_proj = nn.Linear(hs * 2, hs, bias=False)

gate = torch.sigmoid(self.qa_gate_proj(torch.cat([q, y], dim=-1)))
y = y * gate                                # before c_proj (W_O)
```

At the 124M scale this costs **98,304 parameters (+0.079%)**. In this corpus the `y` term is
called **A** (the attention output); `code/model_cheap_qa_minimal.py` is a self-contained
single-variant reference implementation with the port surface marked by `G1 SEAM` banners.

## Where the effect is large

**nanoGPT 124M, OpenWebText, 10k iters, warmup 2000, seeds 42/1337/123** (val CE).
Source: `results/csv/priority1_runs_20260322_1355.csv`.

| variant | added params | val CE | vs baseline |
|---|---|---|---|
| baseline | β€” | 4.4708 Β± 0.0074 | β€” |
| **cheap_qa** (gate on `cat(q, y)`) | +98,304 (+0.079%) | **4.3897 Β± 0.0051** | **βˆ’0.0811 (βˆ’1.81%)** |
| full_x (gate on `x`, the published G1 form) | +7,077,888 (+5.71%) | 4.3950 Β± 0.0116 | βˆ’0.0758 |
| midtier_q (per-head MLP on `q`) | +196,608 (+0.159%) | 4.4135 Β± 0.0086 | βˆ’0.0573 |
| cheap_q (linear gate on `q`) | +49,152 (+0.040%) | 4.4198 Β± 0.0185 | βˆ’0.0510 |

The one comparison here that separates cleanly separates on **cost**, not loss: `cheap_qa`
and `full_x` differ by 0.0053 CE β€” inside seed noise β€” and by **72Γ— in added parameters**.
Parameter counts are exact; loss deltas are estimates from three seeds.

## Where the effect goes away, and reverses

**1. Short warmup β€” every gate tested loses to the baseline.**
Source: `results/csv/warmup500_bundle_runs_20260324_1937.csv`,
`results/csv/phase6b_grok_qafollowups_w500_runs_20260328_0028.csv`,
`results/csv/phase5_warmup500_toptiers_runs_20260326_1722.csv` (warmup 500, 3 seeds each).

| variant | warmup 500 | warmup 2000 | penalty |
|---|---|---|---|
| **baseline** | **4.5663** | 4.4708 | +0.0955 |
| qa_normed | 4.5902 | 4.3926 | +0.1976 |
| cheap_qa | 4.6130 | 4.3897 | +0.2233 |
| qa_lowrank | 4.6289 | 4.4068 | +0.2221 |
| mlp128_qa | 4.6332 | 4.4032 | +0.2300 |
| full_x | 4.8518 | 4.3950 | +0.4568 |

Short warmup hurts everything, but it hurts every gate roughly twice as much as the
baseline, and at warmup 500 the ungated model wins outright. **The headline improvement is
conditional on a 20% warmup schedule.**

**2. A hardened baseline β€” the learned gate harms, decisively.**
A 20-run matrix (4 arms Γ— 5 seeds) against the `2025-01-04_SoftCap` modded-nanoGPT record on
8Γ—H100, August 2–5 2026. Published noise floor for that record: 3.2791 Β± 0.0019 over 80 runs.

| arm | mean val loss | vs baseline |
|---|---|---|
| B β€” baseline | 3.27790 | β€” |
| S β€” static gate (0.5) | 3.27960 | +0.0017 |
| Z β€” zero-init learned gate | 3.30788 | +0.0300 |
| F β€” learned gate | 3.30796 | +0.0301 |

**F βˆ’ S = +0.0284**, 95% CI [+0.0262, +0.0305], t(4) β‰ˆ 36.5, all five seeds positive, **zero
distributional overlap** (worst ungated 3.2822, best gated 3.3049). Gated arms are also ~7%
slower.

The decomposition matters more than the sign: **static scaling costs +0.0017; making the
gate learnable costs +0.0284 β€” 16Γ— more.** The damage comes from *learning* the gate, not
from gating. `F βˆ’ Z = +0.00008` β€” random and zero init reach the same solution.

This is the best-powered experiment in the corpus, and it is negative. Primary logs
(`qgate_run3_MATRIX.tar.gz`) are not in this repository; the figures above are from the
August 2–5 session record.

**3. Scale β€” the ordering among gates does not survive.**
OLMo-2 1B, OpenWebText, 3 seeds, H100 (`rope`, `rmsnorm`, `swiglu`, QK-norm, post-norm):

| variant | mean ppl | sd |
|---|---|---|
| baseline | 44.137 | 0.179 |
| cheap_qa | 43.703 | 0.124 |
| midtier_q | 43.683 | 0.159 |

`cheap_qa` and `midtier_q` differ by 0.02 ppl β€” p = 0.87, a tie β€” despite clear separation at
124M. **See the caveats below before using the 1B result for anything.**

## What is contested inside this corpus

Two or more files here disagree. These are shown rather than resolved.

**The A10 comparison.** Two boards, same hardware, same recipe, different conclusions:

| source | baseline | cheap_qa | midtier_q | ordering |
|---|---|---|---|---|
| `results/summaries/priority1_summary_20260322_1355.txt` (A10 ref block) | 4.4675 | 4.4151 | **4.4039** | midtier_q wins |
| `results/phase8_warmup2000_bulk_a10_runs_20260329_0650.csv` (19 variants, 57 runs) | 4.4676 | **4.4155** | 4.4183 | cheap_qa wins |

Baseline and cheap_qa agree to 0.0004 across both. **`midtier_q` differs by 0.0144, and that
single number is what the "ordering inverts on A10" claim rests on.** The larger and later
board does not reproduce the inversion. Additionally, the A10 lane ran a different software
stack (torch 2.7.0+cu128) from the canonical 3060 board (2.5.1+cu121), and the project record
explicitly blocks a hardware-only interpretation until a version-matched rerun exists.

What *is* consistent across both A10 boards: `cheap_qa` beats baseline, and it drops from
1st on the 3060 to 7th of 19 on the A10 (behind `x_full_headspec` at 4.3858).

**The 1B result.** An August 4 review found: the gate-norm "stability" claim retracted
(post-training β€–Wβ€–_F is statistically indistinguishable from an untouched init draw,
P = 0.494); the win record mislabelled (`cheap_qa` alone is 3/3, sign-test **p = 0.125 β€” not
significant**; the gated family pooled is 6/6, p = 0.0156); seed 42 resumed from a
`step1536` checkpoint in all three arms and carries the largest margin (dropping it moves the
mean delta 0.433 β†’ 0.285); the advantage is late-emerging (baseline *leads* on 2/3 seeds at
step 500); `midtier_q` logged no gate proxy, so there is no cross-variant control. **Whether
the OLMo gate trained at all is not established** β€” a frozen random projection is not
excluded. The checkpoints were lost with the rented instance, so this may stay open.

**Whether the 1B run left LR warmup.** One record says the schedule was truncated but
`t_warmup` fixed 200M→40M with `expected_max_steps: 1526`; another says warmup was never
exited. The rendered per-run config was never recovered. Unresolved.

**The WSL lane.** Two docs report baseline/cheap_qa as 4.4759/4.3902 and 4.4732/4.3800. The
WSL baseline also shows late-training spikes attributed to WSL2 memory management β€”
contamination in the direction that widens the gap. This lane also ran torch 2.7.0.

**Seed sensitivity.** On out-of-band seeds 0/1/2 (`results/phase9_...20260328_2327.csv`),
`cheap_qa` is 4.4139 Β± 0.0094 rather than 4.3897, and `qa_normed` (4.4103) beats it. The
Β±0.0051 above is a within-seed-set figure for 42/1337/123.

## The per-dimension result

**Corrected 2026-08-15.** Earlier versions of this card reported far tighter bands. Those
were per-*layer mean* bands presented as individual-value bands. Full vectors, all 12 layers,
3 seeds β€” 4,608 values per condition:

| phase | init | individual values | per-layer means | mean | val CE |
|---|---|---|---|---|---|
| 19b | 0.50 | β€” | q 0.4587–0.5202 / y 0.4687–0.5766 | 0.481 / 0.504 | 4.4015 |
| 19d free | ~0.50 | **0.3593–0.6821** | 0.4558–0.5718 | **0.4924** | 4.4085 |
| 19c informed | 0.74 | **0.5745–0.8457** | 0.6995–0.7826 | **0.7317** | 4.3950 |

**What holds:** the *mean* does not move from its initialization β€” 0.74 β†’ 0.7317, 0.50 β†’
0.4924. Initialization sets where the distribution sits.

**What does not hold:** "the scales do not move." They spread substantially β€” 19c covers a
0.27 range, 19d covers 0.32. Individual dimensions differentiate; the center does not shift.

**The mechanism claimed earlier β€” a flat loss direction β€” is contradicted by this repository.**
The same blend quantity moves decisively under coarser parameterization:

| parameterization | params | init | converged |
|---|---|---|---|
| phase 14, single scalar | 1 | 0.5 | **0.7391–0.7443** |
| phase 17, per-layer | 12 | 0.5 | q 0.3803–0.5551 (mean 0.4481) / y 0.4126–0.6529 (mean 0.5144) |
| phase 19b/c/d, per-dimension | 768 | 0.50 / 0.74 | mean stays at init |

A scalar moves +0.24 from its initialization. Twelve per-layer values differentiate across a
0.27 range and reproduce a consistent profile. Seven hundred sixty-eight per-dimension values
do not relocate their mean. **The loss is not flat along this axis β€” the fine parameterization
is not identified at this budget.** Gradient dilution across 1,536 parameters and simple
undertraining are both live explanations and this corpus cannot separate them.

## Honest limitations

1. **The headline 124M result is conditional on warmup 2000.** At warmup 500 the ungated
   baseline beats every gate tested.
2. **The best-powered experiment in this corpus is negative** (speedrun matrix, above).
3. **The 1B evidence is weak and partly under review** β€” see the contested section. Do not
   cite it as scale validation.
4. **`cheap_qa`'s margin is seed-set dependent** β€” 4.3897 on seeds 42/1337/123, 4.4139 on
   seeds 0/1/2.
5. **Only 1 of 3 `cheap_qa` seeds, and 0 of 3 baseline seeds, ended with their best
   validation loss at the final evaluation.** Most runs were drifting upward late.
6. **Elaborations (phases 14–19d) tie with the simple gate.** Seven variants sit within
   0.0032 CE against a within-variant seed sd of 0.005–0.012; phase 14 (4.3865) and 17
   (4.3890) are nominally ahead (p = 0.64, 0.94). Off-recipe, warmup 3000 reaches 4.3761.
7. **Standard deviations are population sd at 124M, sample sd at 1B.** A 3-sample sd carries
   roughly 52% relative standard error either way.
8. **Head-specific vs head-shared is confounded** by gate input width and is not claimed.

## Traps in the result files

⚠️ `results/phase19b_canonical_*_0709.*` is a **failed run** β€” six seeds, `returncode=2`,
`best=nan`, no checkpoint. It sits beside the real six-seed data (`..._0711`).

⚠️ `results/summaries/p2_rerun_summary_20260323_1714.txt` ranks `full_x` first at 4.3870 on a
**two-seed partial**. `results/summaries/full_x_seed123_result_20260324_1712.txt` (4.3950,
3 seeds) supersedes it.

⚠️ `results/summaries/a10_replication_summary_20260322_0639.txt` ends with an auto-generated
`PAPER STATEMENTS` block asserting hardware independence and `|delta| < 0.02`. Its own table
forty lines above shows deltas to βˆ’0.1257 and prints `Hardware consistency: INVESTIGATE βœ—`.
Trust the tables, not the prose blocks.

⚠️ **Era warning.** Results predating the `beta2=0.99` fix land in the 4.5–4.6 CE band and are
not comparable to clean-recipe results in the 4.38–4.47 band.

⚠️ `checkpoints/canon/out-shakespeare-char/` is the upstream nanoGPT demo, not part of this
study.

## Relation to published work

Gating the SDPA output at G1 is established. **This is a replication and mechanism study, not
a novelty claim.**

- [Qiu et al. 2025](https://arxiv.org/abs/2505.06708) (NeurIPS 2025 Oral) β€” the G1 gate at
  15B-MoE / 1.7B-dense scale, conditioned on **X** (layer input) in all fifteen reported
  variants. `full_x` here is that formulation ported to nanoGPT.
- [Bu et al. 2025](https://arxiv.org/abs/2510.09017) β€” moves the gate input from X to **V**.
- [Zhou et al. 2026](https://arxiv.org/abs/2608.11805) (Tencent Hunyuan, 12 Aug 2026) β€” adds
  an **H-gate** computed from the SDPA output. Up to notation, their `H` is the `y`/A term in
  `cat(q, y)`. Subsequent and independent; their 5B/500B result reaches a regime this corpus
  cannot.

Two of Qiu et al.'s qualitative findings reproduce here at ~1/10,000 of their token budget:
G1 gating beats baseline (6.026 β†’ 5.761 PPL there; 4.4708 β†’ 4.3897 CE here, at warmup 2000),
and input-*dependent* beats input-*independent* (5.917 β†’ 5.761; 4.4483 β†’ 4.3897). The
input-independent control is an exact structural match β€” a zero-initialized learnable
`(n_head Γ— head_dim)` parameter through a sigmoid in both cases.

`cheap_qa` was confirmed 2026-03-20
(`results/csv/cheap_qa_confirmation_runs_20260320_1854.csv`), before the prior-art search that
located Qiu et al.; the project was then reclassified from a novelty claim to a replication
study. **This establishes no priority** β€” these results were private until August 2026.

## Layout

```
checkpoints/
  canon/       21 x ckpt.pt   base run + phases 14-19d      (RTX 3060, Windows)
  wsl_bridge/  12 x ckpt.pt   phases 10-12                  (RTX 3060, WSL)
code/          model/train for both lanes + the minimal champion module
results/       phase 7-19d run CSVs, summaries and logs
SHA256SUMS.txt integrity manifest for all 33 checkpoints
```

nanoGPT-format `ckpt.pt` at fp32, ~1.4 GiB each, optimizer state included.

```bash
sha256sum -c SHA256SUMS.txt
```

## Open questions

- Does the benefit really track baseline hardness, or is the three-point arc a coincidence of
  three different codebases?
- Why does making the gate *learnable* cost 16Γ— what the gate itself costs?
- Is the per-dimension non-identifiability gradient dilution or undertraining? A longer run at
  one initialization would separate them.
- Did the OLMo gate train at all?
- Why does `cheap_qa` fall from 1st to 7th on the A10 while still beating baseline?

## Attribution

Derived from [nanoGPT](https://github.com/karpathy/nanoGPT) by Andrej Karpathy (MIT).
See `LICENSE`.