PLAIN_SUMMARY.md
Browse files- paper/PLAIN_SUMMARY.md +98 -0
paper/PLAIN_SUMMARY.md
ADDED
|
@@ -0,0 +1,98 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# PLAIN_SUMMARY.md β the paper in one page (for Will, 2026-07-06)
|
| 2 |
+
|
| 3 |
+
**What the paper claims.** We took one production language model β GPT-2 β and built what the
|
| 4 |
+
paper claims in its locked form: the first complete, certified, bidirectional decode of an entire
|
| 5 |
+
production language model. We accounted for its ENTIRE internal state, at every one of
|
| 6 |
+
its 13 layer checkpoints, in three kinds of text, against a pass bar the model itself sets. Not
|
| 7 |
+
"here are some interesting features" β an account: throw away the model's whole hidden state at a
|
| 8 |
+
checkpoint, rebuild it from only the things our decoder can read, and the model's behavior stays
|
| 9 |
+
inside its own noise floor at 39 out of 39 boundaryΓregime cells β 13 checkpoints Γ 3 kinds of
|
| 10 |
+
text (36/39 on the older, stricter ruler, and
|
| 11 |
+
those 3 are priced ruler-geometry, not model behavior). The amount of behavior we could NOT
|
| 12 |
+
explain went from 11.2 units of surprise to exactly zero, in public, dated steps.
|
| 13 |
+
|
| 14 |
+
**And then we translated it β the BABEL codec** (your ratified name; the paper's long form, used
|
| 15 |
+
once: "a certified bidirectional codec for GPT-2's residual-stream language"). Every one of the
|
| 16 |
+
decoder's 351 channels was put on trial: 54% carry
|
| 17 |
+
an explicit English meaning (a "naval/warship" field, a "clause boundary" word, an "operator"
|
| 18 |
+
anchor); the other 46% are PROVEN to carry no word β proven, because random directions of the same
|
| 19 |
+
loudness move the model more. How the language composes from layer to layer is certified linear at
|
| 20 |
+
all 36 layer-seam tests. We built the exact inverse (English β state) and proved the round trip is
|
| 21 |
+
behaviorally invisible at all 39 cells. And the model obeys hand edits: turn the "naval"
|
| 22 |
+
field up and GPT-2 starts predicting *amphib, sunk, ashore, reefs, sailed, submarine*. Three of
|
| 23 |
+
the four axes we hand-edited steer the model in their own vocabulary; the fourth (the "rung", an
|
| 24 |
+
executable formula rather than a static dial) is now CERTIFIED UNUSABLE AS A STEERING LEVER: a
|
| 25 |
+
same-day follow-up (L5) probed it through the channel matched to what it IS β the repetition
|
| 26 |
+
behavior itself β plus two internal readouts, and its tiny effect cleared none of the internal
|
| 27 |
+
readouts, while a genuine onset direction did move the behavior. A second follow-up (L6) tested
|
| 28 |
+
it against an honest 20-draw random null at both doses: at Β±6 it does not separate from a random
|
| 29 |
+
nudge of the same size, and at Β±3 its tiny (~0.4%-probability) effect is indistinguishable from
|
| 30 |
+
the honest random floor itself (two pre-registered 20-draw nulls straddle it β one lands just
|
| 31 |
+
below the effect, one above), growing SUB-linearly with dose β pushing harder buys nothing. A gauge
|
| 32 |
+
you can read, not a lever you can pull, at both doses we tested. One new fact came out of the
|
| 33 |
+
hunt: this bounds what can be written INTO the rung, not what reads out of it β one candidate
|
| 34 |
+
certified channel (the naval/warship field, one layer later; a thin 2.8% margin over its
|
| 35 |
+
multiplicity null) consumes the rung's output. Perfectly readable, not a steering handle.
|
| 36 |
+
|
| 37 |
+
**Why anyone should believe it.** This is the paper's real weapon. Every experiment had its pass
|
| 38 |
+
bar written down and locked BEFORE the data, with a stated bet; the completeness verdict came back
|
| 39 |
+
NOT-YET six consecutive times and we published the gap table and stopped each time; the meter was
|
| 40 |
+
recalibrated once, under pre-registered sanity gates, and both meters are reported everywhere,
|
| 41 |
+
forever. The hardest object in the model β a repetition-keeping computation at layer 5 β got three
|
| 42 |
+
certified impossibility results (can't be read, can't be looked up, can't be forged by its own
|
| 43 |
+
circuit) before the thing that finally worked: we taught a tiny LINEAR student to compute it from
|
| 44 |
+
readable inputs, and the test was built so that bigger students who merely memorize get caught β
|
| 45 |
+
and they were caught (they ace training, fail on never-seen repeat periods; the linear one passes).
|
| 46 |
+
One instrument bug happened all week; our own replay gate caught it. Every number in the paper
|
| 47 |
+
traces to a frozen, hash-stamped artifact (Appendix A maps each one), and everything ran on the
|
| 48 |
+
one A4500 in this machine β any skeptic can re-run any row.
|
| 49 |
+
|
| 50 |
+
**What we do NOT claim.** Not the first "activations β English" concept β Anthropic's Natural
|
| 51 |
+
Language Autoencoders (May 2026) and the Cycle-Consistent Activation Oracles published that idea
|
| 52 |
+
first, and the paper says so generously. Our claim is the locked one β the first complete,
|
| 53 |
+
certified, bidirectional decode of an entire production language model β on four axes they don't
|
| 54 |
+
touch: all boundaries priced (they read middle-to-late-layer
|
| 55 |
+
samples); falsifiable verdicts incl. certified-word-less channels (they score plausible glosses by
|
| 56 |
+
reconstruction); a decoder built from the model's own certified channels with an algebraic inverse
|
| 57 |
+
(theirs is a trained external translator); and a round trip scored in BEHAVIOR against
|
| 58 |
+
matched-random nulls β edit the English, the model obeys (theirs is scored in activation space,
|
| 59 |
+
with a qualitative steering demo on top and no matched-null certification β the paper credits the
|
| 60 |
+
demo). Also honestly scoped: one model, one grain,
|
| 61 |
+
three regimes; the 46% word-less fraction and the un-steerable rung are named, not hidden.
|
| 62 |
+
|
| 63 |
+
**The three honest percentages** (the paper states all three, always together): 100% of the
|
| 64 |
+
pre-registered definition met (certified-no-word counts as an answer); behavioral round trip
|
| 65 |
+
100% / 94.7% / 3-of-4 (reconstruct / meaning-transplant / human-edit; the transplant number is
|
| 66 |
+
16 prose pairs at one mid-stack checkpoint); 53.6% of channels carry an
|
| 67 |
+
actual English word.
|
| 68 |
+
|
| 69 |
+
**The two loose ends are now finished β as certified negatives (L5, same day).** Neither headline
|
| 70 |
+
number moved, and both favorite bets lost. The missing 5.3% of transplantable meaning is NOT
|
| 71 |
+
hiding in the certified door channels: adding the certified door read (summarized or in full)
|
| 72 |
+
moves the model exactly 0% further β
|
| 73 |
+
it is certified to live in genuinely un-charted dark mass outside the whole dictionary. It
|
| 74 |
+
resisted every translation method we tried: it transfers only as its exact raw configuration,
|
| 75 |
+
never through any compressed or named form. And editing the
|
| 76 |
+
rung cleared none of the four channels we can read it through (the honest dose story is above).
|
| 77 |
+
Both remainders are
|
| 78 |
+
closed properties now, not open wounds; the honest claim is sharper, not bigger.
|
| 79 |
+
|
| 80 |
+
**And the questions those answers opened are finished too (L6, same day).** We hunted the 5.3% to
|
| 81 |
+
ground. It is DIFFUSE: smeared across the entire 329-dimension dark space β no hidden low-rank
|
| 82 |
+
concept (we tried every rank up to 256; none captures 80% of the gap, and a RANDOM 256-dim slice
|
| 83 |
+
of the dark does nearly as well as the best hand-picked one). And it is mostly word-less: of its
|
| 84 |
+
8 biggest directions, only 2 faintly cross the naming bar, through side channels, with no stable
|
| 85 |
+
meaning (they're recorded as provisional dark signatures, NOT dictionary entries). The rung stays
|
| 86 |
+
steering-unusable at double dose (above). And the one L4 oddity we'd never explained β the "operator"
|
| 87 |
+
axis appearing to steer on-manifold when it was predicted inert β dissolves: it was read-out ECHO
|
| 88 |
+
(~94% of the response is the injected word-vector riding straight through to the output; the
|
| 89 |
+
computed part is inside the noise), so no new capability, and the L4 result's own control always
|
| 90 |
+
held. L6's scoreboard: 2 favorite bets hit, 3 lost β every loss logged as a certified finding.
|
| 91 |
+
No headline number moved, in any of it.
|
| 92 |
+
|
| 93 |
+
**Deliverables ready for your review:** PAPER_DRAFT_V1.md (+ 6 figures in paper_figs/, all
|
| 94 |
+
regenerable from the frozen JSONs by one script; L5 verdicts in Β§6.6, L6 verdicts in Β§6.7, new
|
| 95 |
+
Fig. 6 = the dark-mass rank ladder) and the outreach kit (AF post, 4 emails, Zenodo + GitHub
|
| 96 |
+
checklists, release sequence, and the new GitHub-front-door README_DRAFT.md you asked for) in
|
| 97 |
+
outreach_kit/, all updated post-L6 and standardized to the ratified name "the BABEL codec".
|
| 98 |
+
Nothing has been sent, posted, uploaded, or committed anywhere β every send is yours to fire.
|