wpferrell commited on
Commit
91fd355
Β·
verified Β·
1 Parent(s): 833aa87

PLAIN_SUMMARY.md

Browse files
Files changed (1) hide show
  1. paper/PLAIN_SUMMARY.md +98 -0
paper/PLAIN_SUMMARY.md ADDED
@@ -0,0 +1,98 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # PLAIN_SUMMARY.md β€” the paper in one page (for Will, 2026-07-06)
2
+
3
+ **What the paper claims.** We took one production language model β€” GPT-2 β€” and built what the
4
+ paper claims in its locked form: the first complete, certified, bidirectional decode of an entire
5
+ production language model. We accounted for its ENTIRE internal state, at every one of
6
+ its 13 layer checkpoints, in three kinds of text, against a pass bar the model itself sets. Not
7
+ "here are some interesting features" β€” an account: throw away the model's whole hidden state at a
8
+ checkpoint, rebuild it from only the things our decoder can read, and the model's behavior stays
9
+ inside its own noise floor at 39 out of 39 boundaryΓ—regime cells β€” 13 checkpoints Γ— 3 kinds of
10
+ text (36/39 on the older, stricter ruler, and
11
+ those 3 are priced ruler-geometry, not model behavior). The amount of behavior we could NOT
12
+ explain went from 11.2 units of surprise to exactly zero, in public, dated steps.
13
+
14
+ **And then we translated it β€” the BABEL codec** (your ratified name; the paper's long form, used
15
+ once: "a certified bidirectional codec for GPT-2's residual-stream language"). Every one of the
16
+ decoder's 351 channels was put on trial: 54% carry
17
+ an explicit English meaning (a "naval/warship" field, a "clause boundary" word, an "operator"
18
+ anchor); the other 46% are PROVEN to carry no word β€” proven, because random directions of the same
19
+ loudness move the model more. How the language composes from layer to layer is certified linear at
20
+ all 36 layer-seam tests. We built the exact inverse (English β†’ state) and proved the round trip is
21
+ behaviorally invisible at all 39 cells. And the model obeys hand edits: turn the "naval"
22
+ field up and GPT-2 starts predicting *amphib, sunk, ashore, reefs, sailed, submarine*. Three of
23
+ the four axes we hand-edited steer the model in their own vocabulary; the fourth (the "rung", an
24
+ executable formula rather than a static dial) is now CERTIFIED UNUSABLE AS A STEERING LEVER: a
25
+ same-day follow-up (L5) probed it through the channel matched to what it IS β€” the repetition
26
+ behavior itself β€” plus two internal readouts, and its tiny effect cleared none of the internal
27
+ readouts, while a genuine onset direction did move the behavior. A second follow-up (L6) tested
28
+ it against an honest 20-draw random null at both doses: at Β±6 it does not separate from a random
29
+ nudge of the same size, and at Β±3 its tiny (~0.4%-probability) effect is indistinguishable from
30
+ the honest random floor itself (two pre-registered 20-draw nulls straddle it β€” one lands just
31
+ below the effect, one above), growing SUB-linearly with dose β€” pushing harder buys nothing. A gauge
32
+ you can read, not a lever you can pull, at both doses we tested. One new fact came out of the
33
+ hunt: this bounds what can be written INTO the rung, not what reads out of it β€” one candidate
34
+ certified channel (the naval/warship field, one layer later; a thin 2.8% margin over its
35
+ multiplicity null) consumes the rung's output. Perfectly readable, not a steering handle.
36
+
37
+ **Why anyone should believe it.** This is the paper's real weapon. Every experiment had its pass
38
+ bar written down and locked BEFORE the data, with a stated bet; the completeness verdict came back
39
+ NOT-YET six consecutive times and we published the gap table and stopped each time; the meter was
40
+ recalibrated once, under pre-registered sanity gates, and both meters are reported everywhere,
41
+ forever. The hardest object in the model β€” a repetition-keeping computation at layer 5 β€” got three
42
+ certified impossibility results (can't be read, can't be looked up, can't be forged by its own
43
+ circuit) before the thing that finally worked: we taught a tiny LINEAR student to compute it from
44
+ readable inputs, and the test was built so that bigger students who merely memorize get caught β€”
45
+ and they were caught (they ace training, fail on never-seen repeat periods; the linear one passes).
46
+ One instrument bug happened all week; our own replay gate caught it. Every number in the paper
47
+ traces to a frozen, hash-stamped artifact (Appendix A maps each one), and everything ran on the
48
+ one A4500 in this machine β€” any skeptic can re-run any row.
49
+
50
+ **What we do NOT claim.** Not the first "activations β†’ English" concept β€” Anthropic's Natural
51
+ Language Autoencoders (May 2026) and the Cycle-Consistent Activation Oracles published that idea
52
+ first, and the paper says so generously. Our claim is the locked one β€” the first complete,
53
+ certified, bidirectional decode of an entire production language model β€” on four axes they don't
54
+ touch: all boundaries priced (they read middle-to-late-layer
55
+ samples); falsifiable verdicts incl. certified-word-less channels (they score plausible glosses by
56
+ reconstruction); a decoder built from the model's own certified channels with an algebraic inverse
57
+ (theirs is a trained external translator); and a round trip scored in BEHAVIOR against
58
+ matched-random nulls β€” edit the English, the model obeys (theirs is scored in activation space,
59
+ with a qualitative steering demo on top and no matched-null certification β€” the paper credits the
60
+ demo). Also honestly scoped: one model, one grain,
61
+ three regimes; the 46% word-less fraction and the un-steerable rung are named, not hidden.
62
+
63
+ **The three honest percentages** (the paper states all three, always together): 100% of the
64
+ pre-registered definition met (certified-no-word counts as an answer); behavioral round trip
65
+ 100% / 94.7% / 3-of-4 (reconstruct / meaning-transplant / human-edit; the transplant number is
66
+ 16 prose pairs at one mid-stack checkpoint); 53.6% of channels carry an
67
+ actual English word.
68
+
69
+ **The two loose ends are now finished β€” as certified negatives (L5, same day).** Neither headline
70
+ number moved, and both favorite bets lost. The missing 5.3% of transplantable meaning is NOT
71
+ hiding in the certified door channels: adding the certified door read (summarized or in full)
72
+ moves the model exactly 0% further β€”
73
+ it is certified to live in genuinely un-charted dark mass outside the whole dictionary. It
74
+ resisted every translation method we tried: it transfers only as its exact raw configuration,
75
+ never through any compressed or named form. And editing the
76
+ rung cleared none of the four channels we can read it through (the honest dose story is above).
77
+ Both remainders are
78
+ closed properties now, not open wounds; the honest claim is sharper, not bigger.
79
+
80
+ **And the questions those answers opened are finished too (L6, same day).** We hunted the 5.3% to
81
+ ground. It is DIFFUSE: smeared across the entire 329-dimension dark space β€” no hidden low-rank
82
+ concept (we tried every rank up to 256; none captures 80% of the gap, and a RANDOM 256-dim slice
83
+ of the dark does nearly as well as the best hand-picked one). And it is mostly word-less: of its
84
+ 8 biggest directions, only 2 faintly cross the naming bar, through side channels, with no stable
85
+ meaning (they're recorded as provisional dark signatures, NOT dictionary entries). The rung stays
86
+ steering-unusable at double dose (above). And the one L4 oddity we'd never explained β€” the "operator"
87
+ axis appearing to steer on-manifold when it was predicted inert β€” dissolves: it was read-out ECHO
88
+ (~94% of the response is the injected word-vector riding straight through to the output; the
89
+ computed part is inside the noise), so no new capability, and the L4 result's own control always
90
+ held. L6's scoreboard: 2 favorite bets hit, 3 lost β€” every loss logged as a certified finding.
91
+ No headline number moved, in any of it.
92
+
93
+ **Deliverables ready for your review:** PAPER_DRAFT_V1.md (+ 6 figures in paper_figs/, all
94
+ regenerable from the frozen JSONs by one script; L5 verdicts in Β§6.6, L6 verdicts in Β§6.7, new
95
+ Fig. 6 = the dark-mass rank ladder) and the outreach kit (AF post, 4 emails, Zenodo + GitHub
96
+ checklists, release sequence, and the new GitHub-front-door README_DRAFT.md you asked for) in
97
+ outreach_kit/, all updated post-L6 and standardized to the ratified name "the BABEL codec".
98
+ Nothing has been sent, posted, uploaded, or committed anywhere β€” every send is yours to fire.