--- license: mit tags: - nanogpt - ablation - attention - conditional-computation - gating - research-artifact - negative-results library_name: pytorch --- # QGate — checkpoints and results from a 19-phase attention-gating ablation Training artifacts from an independent research program on **query-conditioned attention gating** in GPT-2-class transformers. 33 checkpoints, the code that produced them, and the complete result tables — including the runs that did **not** support the hypothesis. **Code, per-phase result tables, and the full experiment history:** [github.com/briantkolb/qgate](https://github.com/briantkolb/qgate) This repository is an archive first and a model release second. Nothing here is intended for downstream use as a general-purpose language model. > ### Revision notice — 2026-08-15 > This card was **substantially revised** the day after first publication. The initial > version claimed the gate "produces a real improvement over an ungated baseline," > replicated six ways. That claim was too strong and is retracted. Three findings already > in the project record had not been incorporated: > 1. a 20-run controlled matrix on a hardened baseline where the learned gate **harms**; > 2. a warmup sweep where **every gate tested loses to the ungated baseline**; > 3. the per-dimension bands were reported ~3× too tight, and the "flat loss direction" > mechanism is contradicted by this repository's own per-layer results. > > The corrected thesis is below. Where two files in this corpus disagree, both are now > shown rather than one being chosen. ## What the data supports **The gate's benefit is a function of how much fixed structure the baseline already has.** On a plain GPT-2-class 124M model at 20% warmup it is a large, reproducible improvement. At 1B, on a modern stack, it is marginal and its statistical support is under review. On a short warmup schedule it is *worse than no gate at all*. On a heavily-optimized speedrun baseline it is decisively harmful. That ordering is the result; the 124M number alone is not. A second, narrower finding concerns parameterization: the same blend quantity that a 12-parameter per-layer gate learns a clear profile for is one that a 768-parameter per-dimension gate fails to relocate from its initialization at this token budget. Adding capacity made it less learnable, not more. ## The intervention A single gate applied at the **G1 seam** — post-SDPA, pre-output-projection: ```python hs = n_embd // n_head # head_dim self.qa_gate_proj = nn.Linear(hs * 2, hs, bias=False) gate = torch.sigmoid(self.qa_gate_proj(torch.cat([q, y], dim=-1))) y = y * gate # before c_proj (W_O) ``` At the 124M scale this costs **98,304 parameters (+0.079%)**. In this corpus the `y` term is called **A** (the attention output); `code/model_cheap_qa_minimal.py` is a self-contained single-variant reference implementation with the port surface marked by `G1 SEAM` banners. ## Where the effect is large **nanoGPT 124M, OpenWebText, 10k iters, warmup 2000, seeds 42/1337/123** (val CE). Source: `results/csv/priority1_runs_20260322_1355.csv`. | variant | added params | val CE | vs baseline | |---|---|---|---| | baseline | — | 4.4708 ± 0.0074 | — | | **cheap_qa** (gate on `cat(q, y)`) | +98,304 (+0.079%) | **4.3897 ± 0.0051** | **−0.0811 (−1.81%)** | | full_x (gate on `x`, the published G1 form) | +7,077,888 (+5.71%) | 4.3950 ± 0.0116 | −0.0758 | | midtier_q (per-head MLP on `q`) | +196,608 (+0.159%) | 4.4135 ± 0.0086 | −0.0573 | | cheap_q (linear gate on `q`) | +49,152 (+0.040%) | 4.4198 ± 0.0185 | −0.0510 | The one comparison here that separates cleanly separates on **cost**, not loss: `cheap_qa` and `full_x` differ by 0.0053 CE — inside seed noise — and by **72× in added parameters**. Parameter counts are exact; loss deltas are estimates from three seeds. ## Where the effect goes away, and reverses **1. Short warmup — every gate tested loses to the baseline.** Source: `results/csv/warmup500_bundle_runs_20260324_1937.csv`, `results/csv/phase6b_grok_qafollowups_w500_runs_20260328_0028.csv`, `results/csv/phase5_warmup500_toptiers_runs_20260326_1722.csv` (warmup 500, 3 seeds each). | variant | warmup 500 | warmup 2000 | penalty | |---|---|---|---| | **baseline** | **4.5663** | 4.4708 | +0.0955 | | qa_normed | 4.5902 | 4.3926 | +0.1976 | | cheap_qa | 4.6130 | 4.3897 | +0.2233 | | qa_lowrank | 4.6289 | 4.4068 | +0.2221 | | mlp128_qa | 4.6332 | 4.4032 | +0.2300 | | full_x | 4.8518 | 4.3950 | +0.4568 | Short warmup hurts everything, but it hurts every gate roughly twice as much as the baseline, and at warmup 500 the ungated model wins outright. **The headline improvement is conditional on a 20% warmup schedule.** **2. A hardened baseline — the learned gate harms, decisively.** A 20-run matrix (4 arms × 5 seeds) against the `2025-01-04_SoftCap` modded-nanoGPT record on 8×H100, August 2–5 2026. Published noise floor for that record: 3.2791 ± 0.0019 over 80 runs. | arm | mean val loss | vs baseline | |---|---|---| | B — baseline | 3.27790 | — | | S — static gate (0.5) | 3.27960 | +0.0017 | | Z — zero-init learned gate | 3.30788 | +0.0300 | | F — learned gate | 3.30796 | +0.0301 | **F − S = +0.0284**, 95% CI [+0.0262, +0.0305], t(4) ≈ 36.5, all five seeds positive, **zero distributional overlap** (worst ungated 3.2822, best gated 3.3049). Gated arms are also ~7% slower. The decomposition matters more than the sign: **static scaling costs +0.0017; making the gate learnable costs +0.0284 — 16× more.** The damage comes from *learning* the gate, not from gating. `F − Z = +0.00008` — random and zero init reach the same solution. This is the best-powered experiment in the corpus, and it is negative. Primary logs (`qgate_run3_MATRIX.tar.gz`) are not in this repository; the figures above are from the August 2–5 session record. **3. Scale — the ordering among gates does not survive.** OLMo-2 1B, OpenWebText, 3 seeds, H100 (`rope`, `rmsnorm`, `swiglu`, QK-norm, post-norm): | variant | mean ppl | sd | |---|---|---| | baseline | 44.137 | 0.179 | | cheap_qa | 43.703 | 0.124 | | midtier_q | 43.683 | 0.159 | `cheap_qa` and `midtier_q` differ by 0.02 ppl — p = 0.87, a tie — despite clear separation at 124M. **See the caveats below before using the 1B result for anything.** ## What is contested inside this corpus Two or more files here disagree. These are shown rather than resolved. **The A10 comparison.** Two boards, same hardware, same recipe, different conclusions: | source | baseline | cheap_qa | midtier_q | ordering | |---|---|---|---|---| | `results/summaries/priority1_summary_20260322_1355.txt` (A10 ref block) | 4.4675 | 4.4151 | **4.4039** | midtier_q wins | | `results/phase8_warmup2000_bulk_a10_runs_20260329_0650.csv` (19 variants, 57 runs) | 4.4676 | **4.4155** | 4.4183 | cheap_qa wins | Baseline and cheap_qa agree to 0.0004 across both. **`midtier_q` differs by 0.0144, and that single number is what the "ordering inverts on A10" claim rests on.** The larger and later board does not reproduce the inversion. Additionally, the A10 lane ran a different software stack (torch 2.7.0+cu128) from the canonical 3060 board (2.5.1+cu121), and the project record explicitly blocks a hardware-only interpretation until a version-matched rerun exists. What *is* consistent across both A10 boards: `cheap_qa` beats baseline, and it drops from 1st on the 3060 to 7th of 19 on the A10 (behind `x_full_headspec` at 4.3858). **The 1B result.** An August 4 review found: the gate-norm "stability" claim retracted (post-training ‖W‖_F is statistically indistinguishable from an untouched init draw, P = 0.494); the win record mislabelled (`cheap_qa` alone is 3/3, sign-test **p = 0.125 — not significant**; the gated family pooled is 6/6, p = 0.0156); seed 42 resumed from a `step1536` checkpoint in all three arms and carries the largest margin (dropping it moves the mean delta 0.433 → 0.285); the advantage is late-emerging (baseline *leads* on 2/3 seeds at step 500); `midtier_q` logged no gate proxy, so there is no cross-variant control. **Whether the OLMo gate trained at all is not established** — a frozen random projection is not excluded. The checkpoints were lost with the rented instance, so this may stay open. **Whether the 1B run left LR warmup.** One record says the schedule was truncated but `t_warmup` fixed 200M→40M with `expected_max_steps: 1526`; another says warmup was never exited. The rendered per-run config was never recovered. Unresolved. **The WSL lane.** Two docs report baseline/cheap_qa as 4.4759/4.3902 and 4.4732/4.3800. The WSL baseline also shows late-training spikes attributed to WSL2 memory management — contamination in the direction that widens the gap. This lane also ran torch 2.7.0. **Seed sensitivity.** On out-of-band seeds 0/1/2 (`results/phase9_...20260328_2327.csv`), `cheap_qa` is 4.4139 ± 0.0094 rather than 4.3897, and `qa_normed` (4.4103) beats it. The ±0.0051 above is a within-seed-set figure for 42/1337/123. ## The per-dimension result **Corrected 2026-08-15.** Earlier versions of this card reported far tighter bands. Those were per-*layer mean* bands presented as individual-value bands. Full vectors, all 12 layers, 3 seeds — 4,608 values per condition: | phase | init | individual values | per-layer means | mean | val CE | |---|---|---|---|---|---| | 19b | 0.50 | — | q 0.4587–0.5202 / y 0.4687–0.5766 | 0.481 / 0.504 | 4.4015 | | 19d free | ~0.50 | **0.3593–0.6821** | 0.4558–0.5718 | **0.4924** | 4.4085 | | 19c informed | 0.74 | **0.5745–0.8457** | 0.6995–0.7826 | **0.7317** | 4.3950 | **What holds:** the *mean* does not move from its initialization — 0.74 → 0.7317, 0.50 → 0.4924. Initialization sets where the distribution sits. **What does not hold:** "the scales do not move." They spread substantially — 19c covers a 0.27 range, 19d covers 0.32. Individual dimensions differentiate; the center does not shift. **The mechanism claimed earlier — a flat loss direction — is contradicted by this repository.** The same blend quantity moves decisively under coarser parameterization: | parameterization | params | init | converged | |---|---|---|---| | phase 14, single scalar | 1 | 0.5 | **0.7391–0.7443** | | phase 17, per-layer | 12 | 0.5 | q 0.3803–0.5551 (mean 0.4481) / y 0.4126–0.6529 (mean 0.5144) | | phase 19b/c/d, per-dimension | 768 | 0.50 / 0.74 | mean stays at init | A scalar moves +0.24 from its initialization. Twelve per-layer values differentiate across a 0.27 range and reproduce a consistent profile. Seven hundred sixty-eight per-dimension values do not relocate their mean. **The loss is not flat along this axis — the fine parameterization is not identified at this budget.** Gradient dilution across 1,536 parameters and simple undertraining are both live explanations and this corpus cannot separate them. ## Honest limitations 1. **The headline 124M result is conditional on warmup 2000.** At warmup 500 the ungated baseline beats every gate tested. 2. **The best-powered experiment in this corpus is negative** (speedrun matrix, above). 3. **The 1B evidence is weak and partly under review** — see the contested section. Do not cite it as scale validation. 4. **`cheap_qa`'s margin is seed-set dependent** — 4.3897 on seeds 42/1337/123, 4.4139 on seeds 0/1/2. 5. **Only 1 of 3 `cheap_qa` seeds, and 0 of 3 baseline seeds, ended with their best validation loss at the final evaluation.** Most runs were drifting upward late. 6. **Elaborations (phases 14–19d) tie with the simple gate.** Seven variants sit within 0.0032 CE against a within-variant seed sd of 0.005–0.012; phase 14 (4.3865) and 17 (4.3890) are nominally ahead (p = 0.64, 0.94). Off-recipe, warmup 3000 reaches 4.3761. 7. **Standard deviations are population sd at 124M, sample sd at 1B.** A 3-sample sd carries roughly 52% relative standard error either way. 8. **Head-specific vs head-shared is confounded** by gate input width and is not claimed. ## Traps in the result files ⚠️ `results/phase19b_canonical_*_0709.*` is a **failed run** — six seeds, `returncode=2`, `best=nan`, no checkpoint. It sits beside the real six-seed data (`..._0711`). ⚠️ `results/summaries/p2_rerun_summary_20260323_1714.txt` ranks `full_x` first at 4.3870 on a **two-seed partial**. `results/summaries/full_x_seed123_result_20260324_1712.txt` (4.3950, 3 seeds) supersedes it. ⚠️ `results/summaries/a10_replication_summary_20260322_0639.txt` ends with an auto-generated `PAPER STATEMENTS` block asserting hardware independence and `|delta| < 0.02`. Its own table forty lines above shows deltas to −0.1257 and prints `Hardware consistency: INVESTIGATE ✗`. Trust the tables, not the prose blocks. ⚠️ **Era warning.** Results predating the `beta2=0.99` fix land in the 4.5–4.6 CE band and are not comparable to clean-recipe results in the 4.38–4.47 band. ⚠️ `checkpoints/canon/out-shakespeare-char/` is the upstream nanoGPT demo, not part of this study. ## Relation to published work Gating the SDPA output at G1 is established. **This is a replication and mechanism study, not a novelty claim.** - [Qiu et al. 2025](https://arxiv.org/abs/2505.06708) (NeurIPS 2025 Oral) — the G1 gate at 15B-MoE / 1.7B-dense scale, conditioned on **X** (layer input) in all fifteen reported variants. `full_x` here is that formulation ported to nanoGPT. - [Bu et al. 2025](https://arxiv.org/abs/2510.09017) — moves the gate input from X to **V**. - [Zhou et al. 2026](https://arxiv.org/abs/2608.11805) (Tencent Hunyuan, 12 Aug 2026) — adds an **H-gate** computed from the SDPA output. Up to notation, their `H` is the `y`/A term in `cat(q, y)`. Subsequent and independent; their 5B/500B result reaches a regime this corpus cannot. Two of Qiu et al.'s qualitative findings reproduce here at ~1/10,000 of their token budget: G1 gating beats baseline (6.026 → 5.761 PPL there; 4.4708 → 4.3897 CE here, at warmup 2000), and input-*dependent* beats input-*independent* (5.917 → 5.761; 4.4483 → 4.3897). The input-independent control is an exact structural match — a zero-initialized learnable `(n_head × head_dim)` parameter through a sigmoid in both cases. `cheap_qa` was confirmed 2026-03-20 (`results/csv/cheap_qa_confirmation_runs_20260320_1854.csv`), before the prior-art search that located Qiu et al.; the project was then reclassified from a novelty claim to a replication study. **This establishes no priority** — these results were private until August 2026. ## Layout ``` checkpoints/ canon/ 21 x ckpt.pt base run + phases 14-19d (RTX 3060, Windows) wsl_bridge/ 12 x ckpt.pt phases 10-12 (RTX 3060, WSL) code/ model/train for both lanes + the minimal champion module results/ phase 7-19d run CSVs, summaries and logs SHA256SUMS.txt integrity manifest for all 33 checkpoints ``` nanoGPT-format `ckpt.pt` at fp32, ~1.4 GiB each, optimizer state included. ```bash sha256sum -c SHA256SUMS.txt ``` ## Open questions - Does the benefit really track baseline hardness, or is the three-point arc a coincidence of three different codebases? - Why does making the gate *learnable* cost 16× what the gate itself costs? - Is the per-dimension non-identifiability gradient dilution or undertraining? A longer run at one initialization would separate them. - Did the OLMo gate train at all? - Why does `cheap_qa` fall from 1st to 7th on the A10 while still beating baseline? ## Attribution Derived from [nanoGPT](https://github.com/karpathy/nanoGPT) by Andrej Karpathy (MIT). See `LICENSE`.