# Implementation status ## Completed - immutable Python dataclasses, TOML validation, and one separate configuration file for each H000-H020 harness; - strict official-API discovery, loading, unloading, and identity checks for Qwen3.6-35B-A3B, GPT-OSS-20B, Qwen3-Coder-30B, and Qwen3 Embedding 0.6B; - phase-exclusive LM Studio model residency controlled through `lms server` plus native REST model lifecycle endpoints; - frozen Git snapshots, exact/regex search, BM25/fuzzy retrieval, Tree-sitter symbols, dense retrieval, rank fusion, graph expansion, context packers, and unified/specialized interfaces; - worktree-isolated patch application, hidden fail-to-pass/pass-to-pass test execution, terminal failure staging, and append-only trajectories; - FAISS FlatIP, FAISS HNSW, and sqlite-vec backend evaluation; - 965 Study 1 cells, including 70 live Qwen tool-loop cells; - a prospectively frozen Study 2 with 60 real-change tasks across GitLab Runner, GitLab CLI, and python-gitlab; seven treatments; two local models; 840 main cells; and 72 repeated-seed cells; - artifact identity audits, paired exact tests, 20,000-sample task-cluster bootstrap intervals, Holm correction, and Bayesian binomial task-random- intercept analysis; - a prospectively frozen Study 3 crossing 60 tasks, three local models, and three edit interfaces in 540 real LLM cells, with an outcome-blind compatibility gate for the fresh retrieval replication; - a prospectively frozen Study 4 with 20 newly mined executable tasks, three gated models, exact/hybrid/oracle retrieval, 180 main cells, 18 stochastic reliability cells, and 12 context-scarcity cells; - a prospectively frozen Study 5 with 1,440 retrieval-component factorial cells, 540 retrieval-by-action cells, 540 navigation/packing cells, and a separately frozen 306-cell held-out validation on 17 fresh tasks; - a LaTeX manuscript, bibliography, generated evidence, reproducibility manifests, and rendered-PDF visual QA; - public-release metadata, dual licensing, third-party notices, a standalone reproduction guide, and a fail-closed repository/Zenodo audit. ## Evidence inventory | Study | Experiment | Cells | Terminal scope | |---|---|---:|---| | 1 | E01 | 150 | Static retrieval and graph depth | | 1 | E02 | 105 | Model-selected localization, interface, query, packing | | 1 | E03 | 100 | Strict single-generation repair | | 1 | E04 | 135 | Stale-index and nested-distractor robustness | | 1 | E05 | 405 | Dense vector backends and timing repetitions | | 1 | E07 | 70 | Live Qwen tool loop with hidden-test evaluation | | 2 | E08 main | 840 | 60 tasks x 7 treatments x 2 models | | 2 | E08 reliability | 72 | 24 groups x 3 stochastic seeds | | 3 | E09 | 540 | Model-by-edit-interface compatibility | | 4 | E10 | 180 | Fresh protocol-normalized retrieval replication | | 4 | E11 | 18 | Six balanced groups x three stochastic seeds | | 4 | E12 | 12 | Qwen exact/hybrid x 16K/65K context | | 5 | E13 | 1,440 | Lexical/syntax/dense factorial x three models | | 5 | E14 | 540 | Retrieval-by-edit-interface factorial | | 5 | E15 | 540 | Graph/query/interface/packing paired blocks | | 5 | E16 | 306 | Six archetypes x 17 fresh tasks x three models | | **Total** | | **5,453** | Audited unique scored run IDs | Development-only artifacts remain excluded: 15 E00 pilot cells, three superseded E02 cells, and one E06 smoke cell. Frozen analysis is in `results/derived/confirmatory_analysis_174ce71bcbce`, `results/derived/e07`, and `results/derived/study2`, `results/derived/study3`, `results/derived/study4`, `results/derived/study4_ancillary`, and `results/derived/study5`. The stable paper is `output/pdf/dissecting_repository_scale_code_agent_harnesses.pdf`. ## Scientific outcome Study 2 did not establish a winning retrieval harness. The registered H007 vs H000/Qwen contrast was 2/60 vs 2/60 (RD 0.000, CI [-0.050, 0.050], p=1), and all corrected secondary tests were null. Study 3 then causally established model-specific action-protocol effects. Its registered GPT-OSS contrast was 27/60 accepted edit cells under exact replacement versus 0/60 under strict unified diff (paired RD +0.450, 95% CI [0.333, 0.583], exact p=1.49e-8). The outcome-blind gate selected P002 for M002 and P003 for M003/M004. This separates executor compatibility from retrieval quality in the next fresh-task replication. Study 4 then tested retrieval on 20 fresh tasks under the gate. H007 resolved 1/60 model-task pairs versus 0/60 for H000 (RD +0.0167, task-cluster 95% CI [0, 0.05], exact sign-flip p=1). H007 improved no registered localization stage and was slower. H018 resolved 4/60, showing that file knowledge helps occasionally but does not remove the semantic repair bottleneck. E11 found unanimous failure but zero unanimous trajectories across six groups; E12 found no resolution difference between 16K and 65K context. Study 5 then tested the harness design space directly. E13 found no corrected lexical, syntax, or dense component main effect. E14 found no corrected retrieval-by-action interaction. E15 showed that iteration and whole-file packing increase systems cost without a demonstrated applicable-patch gain. E16 quality ranks transferred weakly to 17 fresh tasks, while token-cost rank transferred strongly; only one E16 cell resolved, so no universal winner is claimed. ## Deliberately outside this paper - official, unmodified Agentless/SWE-agent/OpenHands implementations; - larger repository and programming-language samples that independently identify language effects; - protocol-adaptive patch normalization as a confirmatory treatment; - additional model sizes, quantizations, and context budgets; - hidden-test-feedback retries or production call/diff repair.