| # Study 2 Prospective Preregistration |
|
|
| **Frozen-design date:** 2026-07-18 |
| **Outcome-data status at freeze:** no Study 2 LLM repair outcome has been generated or inspected. |
| **Purpose:** test whether repository-navigation harness effects replicate across repositories, |
| languages, and model families, and compare the component treatments with controlled adaptations |
| of established coding-agent orchestration systems. |
|
|
| ## Research questions and claims |
|
|
| The confirmatory primary question is whether full hybrid retrieval (`H007`) improves end-to-end |
| repair success over exact/regular-expression search (`H000`) for the originally specified local |
| Qwen3.6-35B-A3B agent. The smallest effect of scientific interest is a 20 percentage-point paired |
| improvement. All other treatment and model-family conclusions are prespecified secondary claims. |
|
|
| The study will distinguish three claim types: |
|
|
| 1. **Component causality:** comparisons among immutable harness configurations. |
| 2. **System comparison:** controlled local reimplementations of published orchestration designs. |
| 3. **Generalization:** variation across repositories, languages, and model families. |
|
|
| `A001` and `A002` must be described as controlled adaptations, not as results from the authors' |
| official implementations. The design is informed by Agentless, SWE-agent, OpenHands, and |
| Multi-SWE-bench, but does not claim implementation identity with those projects. |
|
|
| ## Experimental units and benchmark construction |
|
|
| The frozen target sample is 60 real, test-backed software changes: 20 from each repository below. |
| All repositories are public GitLab checkouts, and every task is pinned to a full base and gold SHA. |
|
|
| | ID | Repository | Language | Head observed before sampling | |
| |---|---|---:|---| |
| | R001 | `gitlab-org/gitlab-runner` | Go | `c2831b75a3ff0782dca8f64498cbc6f71c76819e` | |
| | R002 | `gitlab-org/cli` | Go | `757294c01143360b466a70daf3fbef869fc3a41b` | |
| | R003 | `python-gitlab/python-gitlab` | Python | `f62dda7fa44e3bc46f03bd6402eba3f641f365eb` | |
|
|
| Ten previously frozen and validated R001 tasks may be reused. The remaining tasks are selected by |
| a deterministic, outcome-blind pipeline. Starting from the observed repository head and walking |
| first-parent history newest to oldest (not earlier than 2023-01-01), retain a change only if it: |
|
|
| - changes at least one production source file and at least one test file; |
| - changes no more than eight production files and eight test files; |
| - has no more than 400 changed production lines (`added + deleted`); |
| - is not dependency-only, formatting-only, generated-code-only, documentation-only, or a revert; |
| - yields disjoint production and hidden-test patches that both apply at the pinned base; |
| - has passing selected public tests at the base; |
| - has at least one held-out test that fails with only the test patch applied; and |
| - passes the same tests after applying the gold production patch. |
|
|
| Selection continues until 20 valid tasks exist for a repository. Rejections and their first failed |
| criterion are retained in a machine-readable audit. Task statements use the public issue/MR text |
| when available; otherwise they use the non-solution commit title/body. Statements exclude gold |
| paths, gold symbols, diff text, and test assertions. No LLM outcome may influence task selection. |
|
|
| The complete split, patches, validation logs, repository sizes, language counts, and task-difficulty |
| descriptors are frozen in Git before inference begins. |
|
|
| ## Models and local runtime |
|
|
| Both agent models are served locally on port 1234 through LM Studio. `lms server` is the sole server |
| lifecycle controller. The native REST endpoints `/api/v1/models`, `/api/v1/models/load`, and |
| `/api/v1/models/unload` are the sole model residency controls. At most one model is resident at any |
| time. |
|
|
| | ID | Frozen model | LM Studio variant | Context | Reasoning | |
| |---|---|---|---:|---| |
| | M002 | Qwen3.6-35B-A3B | `qwen/qwen3.6-35b-a3b@4bit` (MLX 4-bit) | 65,536 | on | |
| | M003 | GPT-OSS-20B | `openai/gpt-oss-20b@mxfp4` (MLX MXFP4) | 65,536 | low | |
|
|
| Generation parameters are temperature 0, top-p 1, seed 0, and at most 8,192 completion tokens per |
| turn. The Qwen3 Embedding 0.6B Q8_0 model remains the fixed dense retriever. Embedding construction |
| and query inference occur under exclusive residency; the embedding model is unloaded before an |
| agent model is loaded. A wrong model key, variant, quantization, context, or reasoning mode is fatal. |
| |
| ## Treatments and execution budget |
| |
| The component block contains five immutable harnesses: |
| |
| | Treatment | Role | |
| |---|---| |
| | H000 | exact/substring/regex raw-file baseline | |
| | H003 | H000 plus dense code embeddings | |
| | H007 | exact, BM25/fuzzy, Tree-sitter, and dense retrieval with RRF | |
| | H011 | specialized retrieval tools plus one graph hop | |
| | H018 | oracle-file upper-bound control; file names only, never gold code | |
| |
| The separate system block contains `A001`, a three-stage Agentless-style controlled adaptation, |
| and `A002`, a SWE-agent-style controlled interactive search/read/edit/test adaptation. Every |
| treatment uses the same frozen model build, task statement, context cap, patch restrictions, and |
| hidden-test evaluator. Interactive treatments receive at most 12 model turns, 12 tool calls, and |
| two public test runs. `A001` receives its native fixed three model stages and one validation run; |
| resource differences are reported rather than concealed. |
| |
| The main matrix is `(5 harnesses + 2 systems) × 60 tasks × 2 models = 840` unique cells. Treatment |
| order is cyclically counterbalanced within repository/model blocks. Model order alternates by task. |
| An additional reliability audit repeats 24 deterministically selected non-oracle cells at seeds 1 |
| and 2, balanced across the two models, three repositories, and six non-oracle treatments. This adds |
| 48 cells, for 888 planned live cells total. Selection of repeat cells is frozen before primary |
| outcomes are read. |
| |
| ## Outcomes |
| |
| The primary endpoint is `resolved_at_1`: the generated production patch applies, all held-out |
| fail-to-pass tests pass, and all pass-to-pass tests pass. No human patch repair is allowed. |
| |
| Prespecified secondary outcomes are fail-to-pass success, pass-to-pass success, gold-file recall in |
| search/read/edit actions, exact modified-file match, patch applicability, empty-patch rate, model |
| turns, tool calls, test runs, latency, prompt/completion tokens, model-switch count/time, and |
| protocol-violation rate. Repository size in bytes, tracked source files, source lines, and estimated |
| tokens is reported to verify that full-repository prompting exceeds the 65,536-token treatment cap. |
| |
| ## Power and statistical analysis |
| |
| The sole confirmatory primary contrast is H007 versus H000 on the 60 Qwen (`M002`) paired tasks. |
| It uses a two-sided exact McNemar test at alpha 0.05. Under preregistered discordance probabilities |
| `P(H007 only resolves)=0.25` and `P(H000 only resolves)=0.05`, the exact enumerated power is 0.797. |
| This calculation is reproduced by: |
| |
| ```bash |
| python3 scripts/power_study2.py --tasks 60 --new-only 0.25 --baseline-only 0.05 --alpha 0.05 |
| ``` |
| |
| The paired risk difference is reported with a task-cluster bootstrap 95% interval. The GPT-OSS |
| replication of H007 versus H000, the H003 contribution, H011 comparison, oracle gap, and each |
| controlled-system comparison are secondary. Exact paired p-values are Holm-adjusted within the |
| declared secondary family. Estimates and intervals are emphasized even when tests are nonsignificant. |
| |
| A hierarchical logistic model includes fixed effects for treatment, model, repository/language, |
| and treatment-by-model interaction with a random task intercept. If that model is singular or fails |
| to converge, the preregistered fallback is a task-clustered GEE; if unavailable, a stratified |
| task-cluster bootstrap is reported without asymptotic model p-values. Reliability uses within-cell |
| agreement, variance of binary success, and intraclass agreement across the three seeds. |
| |
| ## Missingness, retries, exclusions, and stopping |
| |
| Gold-validation failures are handled only before the task split is frozen. After freeze, a model |
| stop, malformed tool call, invalid patch, empty patch, timeout, or exhausted budget is an observed |
| failure, not missing data. An inference transport failure may be retried once only when no valid |
| assistant response was observed; the recovery is logged. Infrastructure failures that prevent any |
| cell observation are rerun under the identical immutable identity and reported. Cells are never |
| excluded based on success, patch quality, or treatment ranking. |
| |
| Execution stops only after all 888 cells and declared analyses complete, or for a documented safety |
| condition such as insufficient disk, unavailable LM Studio runtime, irrecoverable repository test |
| environment, or unavailable MPS for any newly introduced Torch operation. This study does not use |
| Torch and does not invoke cloud infrastructure. |
| |
| ## Amendment policy and references |
| |
| Any change after this document's first Git commit but before outcome inspection is recorded in |
| `docs/PROTOCOL_AMENDMENTS.md`, with its reason and whether it changes a claim. Any change after an |
| outcome is observed is labeled exploratory. The execution revision and dirty-tree check are written |
| to every run manifest. |
| |
| Primary design references: |
| |
| - Agentless: https://arxiv.org/abs/2407.01489 |
| - SWE-agent: https://arxiv.org/abs/2405.15793 |
| - OpenHands: https://arxiv.org/abs/2407.16741 |
| - Multi-SWE-bench: https://arxiv.org/abs/2504.02605 |
| |