| # Study 4 frozen result: protocol-normalized fresh-task retrieval |
|
|
| **Main experiment:** E10, 180/180 cells at `cd660ada4b98af2072d453e2a6412107e6f76813` |
| **Ancillaries:** E11/E12, 30/30 cells at `34cc57691dd757c736d0893226d1c06a315301db` |
| **E10 raw manifest/final-metric digest:** |
| `883c0c2176e79647c8f566d45ba05091cee5f1def2836060d4a909efd84b9186` |
|
|
| E10 used 20 newly mined hidden-test-validated tasks and the E09 resolution-blind model/interface |
| gate. H000 and H007 exposed the same search tool and differed only in fixed retrieval behavior. |
|
|
| ## Confirmatory result |
|
|
| H007 resolved 1/60 model-task pairs and H000 resolved 0/60. The paired risk difference was |
| `+0.0167`, with a 20,000-draw task-cluster bootstrap 95% interval of `[0, 0.05]` and exact |
| `2^20` task-cluster sign-flip `p=1`. Nineteen of 20 task-cluster effects were zero. The result does |
| not support hybrid superiority. |
|
|
| All 21 prespecified within-model secondary tests had Holm-adjusted `p=1`. H007 did not improve |
| gold-file retrieval or reading before the first accepted edit for any model. Compared with H000, |
| hybrid accepted-edit counts fell from 34 to 26 and mean latency rose from 31.4 to 73.0 seconds for |
| M002, 34.8 to 71.8 for M003, and 50.7 to 90.8 for M004. |
|
|
| H018 oracle file names resolved 4/60, versus 1/60 H007 and 0/60 H000. This is a localization |
| control, not an oracle repair: 56/60 oracle cells still failed. E10 terminal stages were 75 empty |
| patches, 11 patch-application failures, 89 test failures, and 5 resolutions. |
|
|
| ## Reliability and context sensitivity |
|
|
| E11 repeated six outcome-independent balanced groups across temperature-0.2 seeds 0/1/2. All |
| 18 responses failed. All six groups were unanimous on resolution, but none had an identical |
| trajectory and only three had an identical final patch. Binary unanimity therefore does not imply |
| behavioral determinism. |
|
|
| E12 paired 16,384 and 65,536 loaded context for M002 on H000/H007 and one SHA-selected task per |
| repository. Larger context increased accepted/applicable edits in one of six pairs and changed |
| resolution in zero. Both contexts resolved 0/6. These compact sensitivities are descriptive and |
| are never pooled with E10's primary estimand. |
|
|
| Machine-readable evidence is under `results/derived/study4/` and |
| `results/derived/study4_ancillary/`. Preflights verify official LM Studio API identity, exclusive |
| model loads, context profiles, tool behavior, and cleanup. No Torch or cloud workload was used. |
|
|