File size: 5,766 Bytes
d61821a
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
# Implementation status

## Completed

- immutable Python dataclasses, TOML validation, and one separate configuration
  file for each H000-H020 harness;
- strict official-API discovery, loading, unloading, and identity checks for
  Qwen3.6-35B-A3B, GPT-OSS-20B, Qwen3-Coder-30B, and Qwen3 Embedding 0.6B;
- phase-exclusive LM Studio model residency controlled through `lms server`
  plus native REST model lifecycle endpoints;
- frozen Git snapshots, exact/regex search, BM25/fuzzy retrieval, Tree-sitter
  symbols, dense retrieval, rank fusion, graph expansion, context packers, and
  unified/specialized interfaces;
- worktree-isolated patch application, hidden fail-to-pass/pass-to-pass test
  execution, terminal failure staging, and append-only trajectories;
- FAISS FlatIP, FAISS HNSW, and sqlite-vec backend evaluation;
- 965 Study 1 cells, including 70 live Qwen tool-loop cells;
- a prospectively frozen Study 2 with 60 real-change tasks across GitLab Runner,
  GitLab CLI, and python-gitlab; seven treatments; two local models; 840 main
  cells; and 72 repeated-seed cells;
- artifact identity audits, paired exact tests, 20,000-sample task-cluster
  bootstrap intervals, Holm correction, and Bayesian binomial task-random-
  intercept analysis;
- a prospectively frozen Study 3 crossing 60 tasks, three local models, and
  three edit interfaces in 540 real LLM cells, with an outcome-blind
  compatibility gate for the fresh retrieval replication;
- a prospectively frozen Study 4 with 20 newly mined executable tasks, three
  gated models, exact/hybrid/oracle retrieval, 180 main cells, 18 stochastic
  reliability cells, and 12 context-scarcity cells;
- a prospectively frozen Study 5 with 1,440 retrieval-component factorial
  cells, 540 retrieval-by-action cells, 540 navigation/packing cells, and a
  separately frozen 306-cell held-out validation on 17 fresh tasks;
- a LaTeX manuscript, bibliography, generated evidence, reproducibility
  manifests, and rendered-PDF visual QA;
- public-release metadata, dual licensing, third-party notices, a standalone
  reproduction guide, and a fail-closed repository/Zenodo audit.

## Evidence inventory

| Study | Experiment | Cells | Terminal scope |
|---|---|---:|---|
| 1 | E01 | 150 | Static retrieval and graph depth |
| 1 | E02 | 105 | Model-selected localization, interface, query, packing |
| 1 | E03 | 100 | Strict single-generation repair |
| 1 | E04 | 135 | Stale-index and nested-distractor robustness |
| 1 | E05 | 405 | Dense vector backends and timing repetitions |
| 1 | E07 | 70 | Live Qwen tool loop with hidden-test evaluation |
| 2 | E08 main | 840 | 60 tasks x 7 treatments x 2 models |
| 2 | E08 reliability | 72 | 24 groups x 3 stochastic seeds |
| 3 | E09 | 540 | Model-by-edit-interface compatibility |
| 4 | E10 | 180 | Fresh protocol-normalized retrieval replication |
| 4 | E11 | 18 | Six balanced groups x three stochastic seeds |
| 4 | E12 | 12 | Qwen exact/hybrid x 16K/65K context |
| 5 | E13 | 1,440 | Lexical/syntax/dense factorial x three models |
| 5 | E14 | 540 | Retrieval-by-edit-interface factorial |
| 5 | E15 | 540 | Graph/query/interface/packing paired blocks |
| 5 | E16 | 306 | Six archetypes x 17 fresh tasks x three models |
| **Total** | | **5,453** | Audited unique scored run IDs |

Development-only artifacts remain excluded: 15 E00 pilot cells, three
superseded E02 cells, and one E06 smoke cell. Frozen analysis is in
`results/derived/confirmatory_analysis_174ce71bcbce`, `results/derived/e07`, and
`results/derived/study2`, `results/derived/study3`, `results/derived/study4`,
`results/derived/study4_ancillary`, and `results/derived/study5`. The stable paper is
`output/pdf/dissecting_repository_scale_code_agent_harnesses.pdf`.

## Scientific outcome

Study 2 did not establish a winning retrieval harness. The registered H007 vs
H000/Qwen contrast was 2/60 vs 2/60 (RD 0.000, CI [-0.050, 0.050], p=1), and
all corrected secondary tests were null. Study 3 then causally established
model-specific action-protocol effects. Its registered GPT-OSS contrast was
27/60 accepted edit cells under exact replacement versus 0/60 under strict
unified diff (paired RD +0.450, 95% CI [0.333, 0.583], exact p=1.49e-8).
The outcome-blind gate selected P002 for M002 and P003 for M003/M004. This
separates executor compatibility from retrieval quality in the next fresh-task
replication.

Study 4 then tested retrieval on 20 fresh tasks under the gate. H007 resolved
1/60 model-task pairs versus 0/60 for H000 (RD +0.0167, task-cluster 95% CI
[0, 0.05], exact sign-flip p=1). H007 improved no registered localization
stage and was slower. H018 resolved 4/60, showing that file knowledge helps
occasionally but does not remove the semantic repair bottleneck. E11 found
unanimous failure but zero unanimous trajectories across six groups; E12 found
no resolution difference between 16K and 65K context.

Study 5 then tested the harness design space directly. E13 found no corrected
lexical, syntax, or dense component main effect. E14 found no corrected
retrieval-by-action interaction. E15 showed that iteration and whole-file
packing increase systems cost without a demonstrated applicable-patch gain.
E16 quality ranks transferred weakly to 17 fresh tasks, while token-cost rank
transferred strongly; only one E16 cell resolved, so no universal winner is
claimed.

## Deliberately outside this paper

- official, unmodified Agentless/SWE-agent/OpenHands implementations;
- larger repository and programming-language samples that independently
  identify language effects;
- protocol-adaptive patch normalization as a confirmatory treatment;
- additional model sizes, quantizations, and context budgets;
- hidden-test-feedback retries or production call/diff repair.