v160: wscollapse scope audit closes the stack. Tabs are 15.4% of the gain from 4.9% of runs - but 91.0% of that is ONE challenge, so AN AGGREGATE CAN BE ONE DATA POINT. Both mechanism hypotheses rejected by measurement. Seventh isolated-probe failure: both classes save 0 tokens standalone while the arms save 673. New defect class: A UNITTEST SKIP EXITS 0. 1296 tests, 40/40 mutations fire
v158: the free narrowing SURVIVES in the stack - identical tokens, component, retention, clamped count and clamp crossings from 150 fewer rewrites, confirmed on two chains. Reconciling against the record exposed that MY harness omitted the decoder legend (10.00 tok/challenge x 150 = the exact 1500-token gap), and the legend is what makes pathprefix lossless by restoring 'user'. Closed a new tautology class: assertIn on an empty sigil. 1271 tests, 37/37 mutations fire
v157: THE FIRST FREE NARROWING. rolemark's 150 U-markers save EXACTLY ZERO tokens (unanimously, all 150, one identified cause) while carrying 36.3% of the transform's sigil ambiguity - 13.3x more ambiguous per substitution than the A-marker that holds 100% of the value. Refines v149 rather than contradicting it: the U-marker is not a trade because it has no gain to trade. Scope audit now complete for all four stack members. 1259 tests, 33/33 mutations fire
v156: I REIMPLEMENTED relabel and nearly reported a defect in a booked gain - the booked number was right. Found a SIXTH proxy blind spot, the first with the WRONG SIGN: relabel emitted the label 'ran' where the answer says 'ran without errors', manufacturing an answer term the proxy scored as a GAIN. v145's additivity generalisation falsified (interaction 66x v151's, 100% unclamped). The aggregate hid a two-signed population: 5 challenges outweigh 71. 1248 tests, 41/41 mutations fire
v155: narrowing pathprefix's scope is a PRICED trade (17.3% of gain for 15.2x less retention loss), and my risk assignment was BACKWARDS - the 277 bare mentions are lossless while the 1714 continuations destroy an answer term on 41 challenges. Both mechanism hypotheses were rejected individually and both were partly right, covering only 45% of the gap. Corrected two of my own claims mid-turn. 1235 tests, 34/34 mutations fire
v153: the truncation marker's predicate is VACUOUS - at-cap bodies in 150/150 challenges, so there is no content gate for it. Last turn's 'content-triggered repair' was the marker riding on the tilde's unrelated predicate, and adversely: it marked the 2.03x most at-cap-dense challenges. Separated, the tilde repair buys full coverage for -0.0000211 while the marker costs 108x more and is unpriceable by any instrument I have. 1212 tests, 26/26 mutations fire
v152: the elegant clamp-gated repair policy is WRONG ON THE MERITS - it cuts cost 38.7x but 3 of 4 answer-bearing tilde references sit exactly where it declines to repair. Fragment count was the wrong axis. Content-triggered repair dominates: full coverage for 8.6x less than ungated. Caught an impossible-sign attribution error and 5 guard gaps incl. a new class - a magnitude comparison cannot detect a vanished quantity. 1202 tests
v151: v145's closed form does NOT extend from gains to costs. Tokens still add exactly but the component interaction is +0.0000028396 - negligible vs the total cost, yet 11.8% of the smaller cost. The clamp is a ONE-WAY DOOR: gains lose value crossing INTO it, costs lose penalty sitting ABOVE it. Also found BOTH marker figures wrong (my arm marked 755 never-truncated bodies; v141 used the retracted population) and a real guard gap my own mutations exposed. 1187 tests
v150: audited every sigil for distinguishability. All my variants are clean; the SHIPPED compressor silently corrupts raw tildes to hyphens because '~' is BOTH an escape target and the thinking delimiter - Django's ~Q negation becomes -Q across 269 tildes in 47 of 150 challenges, and 4 QA answers reference that construct. The proxy cannot see it: 0 of 21074 answer terms contain a tilde. Repair priced at -0.0000240323 (0.1% of the stack). 1173 tests
v149: found a REAL defect in v137's booked role-marker gain while trying to bound flip risk - the '>' and '<' sigils are indistinguishable from Python REPL prompts and object reprs on 146 lines in 34 of 150 challenges, exactly the 34 that fail inversion. The repair fails: unambiguous sigils save EXACTLY 0 tokens, so the gain and the ambiguity are the same property. Invertibility bounds information loss, NOT flip risk - promotion stays blocked. 1158 tests
v148: prose surface CLOSED. Lossless prose dedup reaches only 4.2% of v147's truncation bound. My pre-registered decodability trap fired: the retention proxy read EXACTLY 0.0 while 16 of 150 challenges failed round-trip - I had reintroduced the v140 indentation defect by keying on stripped lines. Fixed for 172 tokens: 150/150 exact. Also caught a 6.2x corpus-wide census overcount and rejected generalised path discovery at 41x the rejected retention loss. 1143 tests
v147: the worst-compressing challenge is 87% a pasted GitHub issue - the compressor is TAG-BLIND, truncating tool_results while preserving identical prose verbatim. That blind spot is worth 2.26x the whole stack but requires deleting the task statement, so it is a BOUND on the approach. Caught a bundling error via an implausible tokens-per-line ratio; separated a free frame rewrite (+0.0021945122, retention 0.0) from a declined lossy deletion. 1130 tests
v145: four-way stack +0.0233454772 (-19931 tokens, retention exactly 0.0), 2.11x the previous stack. THE FAMILY HAS A CLOSED FORM: residual interaction exactly 0 again at 2x scale, so any combination can be priced from token deltas alone. Transforms commute. Harness validation caught my own relabel reimplementation defect first. Corrected a narration error where unequal scale steps made a decaying marginal look rising. 1101 tests
v144: LARGEST SINGLE GAIN TO DATE +0.0137724979 from path-prefix substitution, found on the surface v130 wrongly foreclosed. 2.06x the best single, 1.24x the honest stack, retention exactly 0.0. Boundary is principled: a deeper prefix saves MORE tokens but destroys a path-valued answer term. Legend must be emitted conditionally. Blank lines closed three ways. 1085 tests
v137: first gain OUTSIDE the span machinery - 52.67% of tokens live there and I had never looked. Role markers rewritten bijectively save 3489 tokens, component +0.0039295098 at retention exactly 0.0. Found the emitter by enumerating literals after three failed greps. Dropping the colon saves EXACTLY 0 because ':\n' is already one token. 996 tests
v136: reopened the v125 delimiter family (closed for the wrong reason) and closed it properly. Characters equally cheap IN ISOLATION are not equally cheap IN CONTEXT - pilcrow has 0 collisions and is 1 token alone yet costs +6762. Elision gains 2025 tokens but 36.65% of span bodies contain newlines, so it is undecodable. 982 tests
v135: LARGEST FREE GAIN ON RECORD - span labels re-encoded as letters save 5608 tokens for component +0.0066846686 at retention EXACTLY 0.0 (1.74x ws+tab collapse). 'r0' is 2 tokens, 'ra' is 1. A direct resolution test caught a real collision ('labelcolor=red') the arm's own checks missed; fixed by keeping digits on the 179 reference targets. Closed a real guard gap. 969 tests