File size: 2,442 Bytes
d61821a | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 | # Study 4 frozen result: protocol-normalized fresh-task retrieval
**Main experiment:** E10, 180/180 cells at `cd660ada4b98af2072d453e2a6412107e6f76813`
**Ancillaries:** E11/E12, 30/30 cells at `34cc57691dd757c736d0893226d1c06a315301db`
**E10 raw manifest/final-metric digest:**
`883c0c2176e79647c8f566d45ba05091cee5f1def2836060d4a909efd84b9186`
E10 used 20 newly mined hidden-test-validated tasks and the E09 resolution-blind model/interface
gate. H000 and H007 exposed the same search tool and differed only in fixed retrieval behavior.
## Confirmatory result
H007 resolved 1/60 model-task pairs and H000 resolved 0/60. The paired risk difference was
`+0.0167`, with a 20,000-draw task-cluster bootstrap 95% interval of `[0, 0.05]` and exact
`2^20` task-cluster sign-flip `p=1`. Nineteen of 20 task-cluster effects were zero. The result does
not support hybrid superiority.
All 21 prespecified within-model secondary tests had Holm-adjusted `p=1`. H007 did not improve
gold-file retrieval or reading before the first accepted edit for any model. Compared with H000,
hybrid accepted-edit counts fell from 34 to 26 and mean latency rose from 31.4 to 73.0 seconds for
M002, 34.8 to 71.8 for M003, and 50.7 to 90.8 for M004.
H018 oracle file names resolved 4/60, versus 1/60 H007 and 0/60 H000. This is a localization
control, not an oracle repair: 56/60 oracle cells still failed. E10 terminal stages were 75 empty
patches, 11 patch-application failures, 89 test failures, and 5 resolutions.
## Reliability and context sensitivity
E11 repeated six outcome-independent balanced groups across temperature-0.2 seeds 0/1/2. All
18 responses failed. All six groups were unanimous on resolution, but none had an identical
trajectory and only three had an identical final patch. Binary unanimity therefore does not imply
behavioral determinism.
E12 paired 16,384 and 65,536 loaded context for M002 on H000/H007 and one SHA-selected task per
repository. Larger context increased accepted/applicable edits in one of six pairs and changed
resolution in zero. Both contexts resolved 0/6. These compact sensitivities are descriptive and
are never pooled with E10's primary estimand.
Machine-readable evidence is under `results/derived/study4/` and
`results/derived/study4_ancillary/`. Preflights verify official LM Studio API identity, exclusive
model loads, context profiles, tool behavior, and cleanup. No Torch or cloud workload was used.
|