agent-harness / docs /STUDY4_RESULTS.md
cuber12's picture
Publish agent harness research code and paper artifacts
d61821a verified
|
Raw
History Blame Contribute Delete
2.44 kB

Study 4 frozen result: protocol-normalized fresh-task retrieval

Main experiment: E10, 180/180 cells at cd660ada4b98af2072d453e2a6412107e6f76813
Ancillaries: E11/E12, 30/30 cells at 34cc57691dd757c736d0893226d1c06a315301db
E10 raw manifest/final-metric digest: 883c0c2176e79647c8f566d45ba05091cee5f1def2836060d4a909efd84b9186

E10 used 20 newly mined hidden-test-validated tasks and the E09 resolution-blind model/interface gate. H000 and H007 exposed the same search tool and differed only in fixed retrieval behavior.

Confirmatory result

H007 resolved 1/60 model-task pairs and H000 resolved 0/60. The paired risk difference was +0.0167, with a 20,000-draw task-cluster bootstrap 95% interval of [0, 0.05] and exact 2^20 task-cluster sign-flip p=1. Nineteen of 20 task-cluster effects were zero. The result does not support hybrid superiority.

All 21 prespecified within-model secondary tests had Holm-adjusted p=1. H007 did not improve gold-file retrieval or reading before the first accepted edit for any model. Compared with H000, hybrid accepted-edit counts fell from 34 to 26 and mean latency rose from 31.4 to 73.0 seconds for M002, 34.8 to 71.8 for M003, and 50.7 to 90.8 for M004.

H018 oracle file names resolved 4/60, versus 1/60 H007 and 0/60 H000. This is a localization control, not an oracle repair: 56/60 oracle cells still failed. E10 terminal stages were 75 empty patches, 11 patch-application failures, 89 test failures, and 5 resolutions.

Reliability and context sensitivity

E11 repeated six outcome-independent balanced groups across temperature-0.2 seeds 0/1/2. All 18 responses failed. All six groups were unanimous on resolution, but none had an identical trajectory and only three had an identical final patch. Binary unanimity therefore does not imply behavioral determinism.

E12 paired 16,384 and 65,536 loaded context for M002 on H000/H007 and one SHA-selected task per repository. Larger context increased accepted/applicable edits in one of six pairs and changed resolution in zero. Both contexts resolved 0/6. These compact sensitivities are descriptive and are never pooled with E10's primary estimand.

Machine-readable evidence is under results/derived/study4/ and results/derived/study4_ancillary/. Preflights verify official LM Studio API identity, exclusive model loads, context profiles, tool behavior, and cleanup. No Torch or cloud workload was used.