Upload INTERP_RESULTS.md with huggingface_hub
Browse files- INTERP_RESULTS.md +1 -1
INTERP_RESULTS.md
CHANGED
|
@@ -3439,4 +3439,4 @@ Folding wave-6/7 into §18: §18 established c_i is WHERE decodable meaning live
|
|
| 3439 |
|
| 3440 |
---
|
| 3441 |
|
| 3442 |
-
*This §19 is the wave-6 'reading c_i' chapter (appended after §18), refined by wave-7 (W1/W2/W4/W6/W8/W9), wave-8 (W10–W18), wave-9 (W19/W21/W22/W23/W24), wave-10 cross-model generality (W30: the whole profile — content-store + no-abstract-binding — holds across MiniLM/MPNet/e5; no encoder reads thematic role, §19.7b), and a **natural-text validation (W31)**: on *real* mined active/passive pairs the two load-bearing claims REPLICATE off-template — **no-binding** (cross-construction thematic-role AUC z **0.009** vs surface **0.991**, z_bag role-blind 0.50) and **content-store robustness** (z cross-paraphrase **1.0**); W31's reading-coverage arm was a ridge-fit failure (bag below the mean-only floor), **since FIXED by W41**: with a proper intercept-ridge, natural-text reading-coverage REPLICATES the synthetic R9 ladder rung-for-rung (ceiling **0.895**, bag-ridge **0.471**, +SAE-features **0.644** vs synthetic 0.901/0.499/0.706; bag-ridge ≥ mean-only floor 0.063, fit-sanity passes) — so **all three load-bearing claims (no-binding, content-store, AND reading-coverage) are now validated on natural text**, with no meaningful length/difficulty penalty. The pooled-field SAE per-feature analysis is COMPLETE (R1: 273 atoms / 16% energy, §19.5), its atoms are validated as a real monosemantic/auditable READABLE dictionary, NOT causal levers (W20, §19.5: 93.5% pass held-out auto-interp, only 2.5% steer coherently — W15's n=2 "cleaner levers" does not generalize) and shown CROSS-LINGUAL (W23, §19.5); R8's steering arms are closed (W2). Owed before deployment: natural-text replication of the synthetic-template results (C2/C3/R5/W4/W6/W11/W13/W16/W17/W18/W19/W21/W22), pushing the atom fraction past ~16%, and a wider typological sweep
|
|
|
|
| 3439 |
|
| 3440 |
---
|
| 3441 |
|
| 3442 |
+
*This §19 is the wave-6 'reading c_i' chapter (appended after §18), refined by wave-7 (W1/W2/W4/W6/W8/W9), wave-8 (W10–W18), wave-9 (W19/W21/W22/W23/W24), wave-10 cross-model generality (W30: the whole profile — content-store + no-abstract-binding — holds across MiniLM/MPNet/e5; no encoder reads thematic role, §19.7b), and a **natural-text validation (W31)**: on *real* mined active/passive pairs the two load-bearing claims REPLICATE off-template — **no-binding** (cross-construction thematic-role AUC z **0.009** vs surface **0.991**, z_bag role-blind 0.50) and **content-store robustness** (z cross-paraphrase **1.0**); W31's reading-coverage arm was a ridge-fit failure (bag below the mean-only floor), **since FIXED by W41**: with a proper intercept-ridge, natural-text reading-coverage REPLICATES the synthetic R9 ladder rung-for-rung (ceiling **0.895**, bag-ridge **0.471**, +SAE-features **0.644** vs synthetic 0.901/0.499/0.706; bag-ridge ≥ mean-only floor 0.063, fit-sanity passes) — so **all three load-bearing claims (no-binding, content-store, AND reading-coverage) are now validated on natural text**, with no meaningful length/difficulty penalty. The pooled-field SAE per-feature analysis is COMPLETE (R1: 273 atoms / 16% energy, §19.5), its atoms are validated as a real monosemantic/auditable READABLE dictionary, NOT causal levers (W20, §19.5: 93.5% pass held-out auto-interp, only 2.5% steer coherently — W15's n=2 "cleaner levers" does not generalize) and shown CROSS-LINGUAL (W23, §19.5); R8's steering arms are closed (W2). Owed before deployment: natural-text replication of the synthetic-template results (C2/C3/R5/W4/W6/W11/W13/W16/W17/W18/W19/W21/W22), pushing the atom fraction past ~16% (W40 in progress), and a wider typological sweep — **now done by W42**: no-binding holds across Japanese (SOV, が/を case particles), Hindi, Turkish, Russian, and Arabic (cross-construction thematic AUC 0.02–0.17, surface 0.83–0.98); decisively, even Japanese's *explicit readable case-particle tokens* yield ZERO abstract role binding (cross-construction 0.017, z_bag 0.499 — the particles aren't even read), so no-binding is typologically *deep*, not a surface-marker artifact. The genitive (W19, a nominal possession relation) remains the lone abstract-bound relation cross-typologically; thematic agency is never bound regardless of morphological case-marking. The corrected headlines: **z has essentially no abstract relational structure — wave-9's W19 CORRECTS wave-8's W13, showing possession's apparent binding is a genitive-morpheme readout (collapses 1.0→0.043 when neutralized), §19.4**; numbers read exactly (W16); the readability ceiling is capped ~0.71 and external-encoder-irreducible (W14); and readable≠causal holds at the atom level (W20: R1's atoms are a real monosemantic/auditable dictionary but NOT causal levers — only 2.5% steer; W15's n=2 "cleaner levers" does not generalize). The safety split: z is a paraphrase-robust CONTENT auditor (W22) but unsafe for relational/directional-harm auditing (W6/W21 — confirmed on real threat-source/perpetrator/incitement, fixable only by a dependency-parse agent-slot, §19.8). Mechanistically c_i is the post-final-LayerNorm residual, rank-expanded ~53× (W8/W12, §19.5).*
|