Study 2 Prospective Preregistration
Frozen-design date: 2026-07-18
Outcome-data status at freeze: no Study 2 LLM repair outcome has been generated or inspected.
Purpose: test whether repository-navigation harness effects replicate across repositories,
languages, and model families, and compare the component treatments with controlled adaptations
of established coding-agent orchestration systems.
Research questions and claims
The confirmatory primary question is whether full hybrid retrieval (H007) improves end-to-end
repair success over exact/regular-expression search (H000) for the originally specified local
Qwen3.6-35B-A3B agent. The smallest effect of scientific interest is a 20 percentage-point paired
improvement. All other treatment and model-family conclusions are prespecified secondary claims.
The study will distinguish three claim types:
- Component causality: comparisons among immutable harness configurations.
- System comparison: controlled local reimplementations of published orchestration designs.
- Generalization: variation across repositories, languages, and model families.
A001 and A002 must be described as controlled adaptations, not as results from the authors'
official implementations. The design is informed by Agentless, SWE-agent, OpenHands, and
Multi-SWE-bench, but does not claim implementation identity with those projects.
Experimental units and benchmark construction
The frozen target sample is 60 real, test-backed software changes: 20 from each repository below. All repositories are public GitLab checkouts, and every task is pinned to a full base and gold SHA.
| ID | Repository | Language | Head observed before sampling |
|---|---|---|---|
| R001 | gitlab-org/gitlab-runner |
Go | c2831b75a3ff0782dca8f64498cbc6f71c76819e |
| R002 | gitlab-org/cli |
Go | 757294c01143360b466a70daf3fbef869fc3a41b |
| R003 | python-gitlab/python-gitlab |
Python | f62dda7fa44e3bc46f03bd6402eba3f641f365eb |
Ten previously frozen and validated R001 tasks may be reused. The remaining tasks are selected by a deterministic, outcome-blind pipeline. Starting from the observed repository head and walking first-parent history newest to oldest (not earlier than 2023-01-01), retain a change only if it:
- changes at least one production source file and at least one test file;
- changes no more than eight production files and eight test files;
- has no more than 400 changed production lines (
added + deleted); - is not dependency-only, formatting-only, generated-code-only, documentation-only, or a revert;
- yields disjoint production and hidden-test patches that both apply at the pinned base;
- has passing selected public tests at the base;
- has at least one held-out test that fails with only the test patch applied; and
- passes the same tests after applying the gold production patch.
Selection continues until 20 valid tasks exist for a repository. Rejections and their first failed criterion are retained in a machine-readable audit. Task statements use the public issue/MR text when available; otherwise they use the non-solution commit title/body. Statements exclude gold paths, gold symbols, diff text, and test assertions. No LLM outcome may influence task selection.
The complete split, patches, validation logs, repository sizes, language counts, and task-difficulty descriptors are frozen in Git before inference begins.
Models and local runtime
Both agent models are served locally on port 1234 through LM Studio. lms server is the sole server
lifecycle controller. The native REST endpoints /api/v1/models, /api/v1/models/load, and
/api/v1/models/unload are the sole model residency controls. At most one model is resident at any
time.
| ID | Frozen model | LM Studio variant | Context | Reasoning |
|---|---|---|---|---|
| M002 | Qwen3.6-35B-A3B | qwen/qwen3.6-35b-a3b@4bit (MLX 4-bit) |
65,536 | on |
| M003 | GPT-OSS-20B | openai/gpt-oss-20b@mxfp4 (MLX MXFP4) |
65,536 | low |
Generation parameters are temperature 0, top-p 1, seed 0, and at most 8,192 completion tokens per turn. The Qwen3 Embedding 0.6B Q8_0 model remains the fixed dense retriever. Embedding construction and query inference occur under exclusive residency; the embedding model is unloaded before an agent model is loaded. A wrong model key, variant, quantization, context, or reasoning mode is fatal.
Treatments and execution budget
The component block contains five immutable harnesses:
| Treatment | Role |
|---|---|
| H000 | exact/substring/regex raw-file baseline |
| H003 | H000 plus dense code embeddings |
| H007 | exact, BM25/fuzzy, Tree-sitter, and dense retrieval with RRF |
| H011 | specialized retrieval tools plus one graph hop |
| H018 | oracle-file upper-bound control; file names only, never gold code |
The separate system block contains A001, a three-stage Agentless-style controlled adaptation,
and A002, a SWE-agent-style controlled interactive search/read/edit/test adaptation. Every
treatment uses the same frozen model build, task statement, context cap, patch restrictions, and
hidden-test evaluator. Interactive treatments receive at most 12 model turns, 12 tool calls, and
two public test runs. A001 receives its native fixed three model stages and one validation run;
resource differences are reported rather than concealed.
The main matrix is (5 harnesses + 2 systems) × 60 tasks × 2 models = 840 unique cells. Treatment
order is cyclically counterbalanced within repository/model blocks. Model order alternates by task.
An additional reliability audit repeats 24 deterministically selected non-oracle cells at seeds 1
and 2, balanced across the two models, three repositories, and six non-oracle treatments. This adds
48 cells, for 888 planned live cells total. Selection of repeat cells is frozen before primary
outcomes are read.
Outcomes
The primary endpoint is resolved_at_1: the generated production patch applies, all held-out
fail-to-pass tests pass, and all pass-to-pass tests pass. No human patch repair is allowed.
Prespecified secondary outcomes are fail-to-pass success, pass-to-pass success, gold-file recall in search/read/edit actions, exact modified-file match, patch applicability, empty-patch rate, model turns, tool calls, test runs, latency, prompt/completion tokens, model-switch count/time, and protocol-violation rate. Repository size in bytes, tracked source files, source lines, and estimated tokens is reported to verify that full-repository prompting exceeds the 65,536-token treatment cap.
Power and statistical analysis
The sole confirmatory primary contrast is H007 versus H000 on the 60 Qwen (M002) paired tasks.
It uses a two-sided exact McNemar test at alpha 0.05. Under preregistered discordance probabilities
P(H007 only resolves)=0.25 and P(H000 only resolves)=0.05, the exact enumerated power is 0.797.
This calculation is reproduced by:
python3 scripts/power_study2.py --tasks 60 --new-only 0.25 --baseline-only 0.05 --alpha 0.05
The paired risk difference is reported with a task-cluster bootstrap 95% interval. The GPT-OSS replication of H007 versus H000, the H003 contribution, H011 comparison, oracle gap, and each controlled-system comparison are secondary. Exact paired p-values are Holm-adjusted within the declared secondary family. Estimates and intervals are emphasized even when tests are nonsignificant.
A hierarchical logistic model includes fixed effects for treatment, model, repository/language, and treatment-by-model interaction with a random task intercept. If that model is singular or fails to converge, the preregistered fallback is a task-clustered GEE; if unavailable, a stratified task-cluster bootstrap is reported without asymptotic model p-values. Reliability uses within-cell agreement, variance of binary success, and intraclass agreement across the three seeds.
Missingness, retries, exclusions, and stopping
Gold-validation failures are handled only before the task split is frozen. After freeze, a model stop, malformed tool call, invalid patch, empty patch, timeout, or exhausted budget is an observed failure, not missing data. An inference transport failure may be retried once only when no valid assistant response was observed; the recovery is logged. Infrastructure failures that prevent any cell observation are rerun under the identical immutable identity and reported. Cells are never excluded based on success, patch quality, or treatment ranking.
Execution stops only after all 888 cells and declared analyses complete, or for a documented safety condition such as insufficient disk, unavailable LM Studio runtime, irrecoverable repository test environment, or unavailable MPS for any newly introduced Torch operation. This study does not use Torch and does not invoke cloud infrastructure.
Amendment policy and references
Any change after this document's first Git commit but before outcome inspection is recorded in
docs/PROTOCOL_AMENDMENTS.md, with its reason and whether it changes a claim. Any change after an
outcome is observed is labeled exploratory. The execution revision and dirty-tree check are written
to every run manifest.
Primary design references:
- Agentless: https://arxiv.org/abs/2407.01489
- SWE-agent: https://arxiv.org/abs/2405.15793
- OpenHands: https://arxiv.org/abs/2407.16741
- Multi-SWE-bench: https://arxiv.org/abs/2504.02605