# Study 3 Prospective Preregistration: Model-by-Edit-Interface Compatibility **Design date:** 2026-07-19 **Outcome status at first freeze:** no E09 model response or endpoint exists. **Relationship to prior evidence:** Study 2's post-outcome audit found that 877/887 GPT-OSS patch attempts violated a strict raw-unified-diff contract. E09 is a prospective causal follow-up; it does not relabel or replace any E08 endpoint. ## Research question and estimand E09 asks whether the edit-action protocol changes a model's ability to produce an executor-accepted repository modification when localization is held constant. All arms receive the same issue text, oracle production-file names, bounded reads, public tests, hidden evaluator, model context, and interaction budgets. The randomized/counterbalanced factor is the edit tool: - P001: strict raw unified diff; - P002: exact old-text to new-text replacement; and - P003: complete existing-file rewrite. The sole confirmatory contrast is P002 versus P001 for M003 (GPT-OSS-20B). The primary endpoint is `accepted_edit_cell`: at least one edit call is accepted by the executor and the final worktree diff is nonempty. This endpoint isolates action compatibility before patch correctness. The primary effect is the paired risk difference over 60 tasks and a two-sided exact McNemar test at alpha 0.05. `resolved_at_1` remains the principal end-to-end secondary endpoint: the final production diff applies after the hidden test patch, all fail-to-pass tests pass, and all pass-to-pass tests pass. Additional outcomes are any edit attempt, executor acceptance per attempt, applicable final patch, protocol-error type, modified-file match, tests reached, model/tool calls, tokens, latency, and finish reason. ## Hypotheses 1. **H1 (confirmatory):** P002 increases M003 accepted-edit cells over P001. 2. **H2:** edit-interface effects interact with model family; they are not transportable constants. 3. **H3:** protocol acceptance mediates part of the path from model response to resolution. 4. **H4:** whole-file rewriting is more token-intensive and may be less applicable on long files. Only H1 is confirmatory. H2--H4 and every resolution comparison form one prespecified secondary family where discrete binary p-values are Holm adjusted. Mechanism decompositions are secondary; any post-outcome parsing or new grouping is labeled exploratory. ## Experimental units and matrix The task unit is each of the 60 frozen, test-backed E08 changes: 20 each from GitLab Runner (Go), GitLab CLI (Go), and python-gitlab (Python). Reuse is deliberate: the new causal factor was absent from E08, task validation is already sealed, and the study tests an interface mechanism rather than claiming an independent benchmark estimate. Task statements, base/gold commits, production/test patches, public commands, and hidden evaluation are unchanged. Three local LM Studio models are crossed with three edit interfaces: | ID | Model | Variant | Context | Reasoning metadata | |---|---|---|---:|---| | M002 | Qwen3.6-35B-A3B | `qwen/qwen3.6-35b-a3b@4bit` | 65,536 | on | | M003 | GPT-OSS-20B | `openai/gpt-oss-20b@mxfp4` | 65,536 | low | | M004 | Qwen3-Coder-30B-A3B-Instruct | `qwen/qwen3-coder-30b@4bit` | 65,536 | none exposed | The main matrix is 60 tasks x 3 models x 3 interfaces = **540 unique cells**. Temperature is 0, top-p is 1, seed is 0, maximum completion is 8,192 tokens, and each cell permits 12 model calls, 12 total tool calls, two public tests, and 1,800 seconds. Model and interface order rotate by task so each appears equally in every ordinal position within each repository block. ## Shared interaction contract The user prompt exposes gold production-file names only, never gold code, symbols, line ranges, patches, or tests. Tools shared by every arm are bounded `read_file`, frozen-allowlist `run_tests`, and `finish`. Exactly one edit tool is visible. The agent must read before editing, cannot create or delete files, cannot edit tests, and receives tool errors as ordinary observations. Normal assistant text never mutates the repository. P001 accepts only a standard raw unified diff with `a/` and `b/` paths. P002 requires `old_text` to match exactly once in the named tracked file before atomically replacing it. P003 replaces the full contents of one existing tracked file. Every successful structured edit is converted from the actual worktree state to a standard final unified diff before the identical hidden evaluator runs. ## Power Under the smallest preregistered effect of interest, `P(P002 only accepted)=0.25` and `P(P001 only accepted)=0.05`, 60 pairs give exact McNemar power 0.797 at two-sided alpha 0.05. This is reproduced by: ```bash python3 scripts/power_study2.py --tasks 60 --new-only 0.25 --baseline-only 0.05 --alpha 0.05 ``` The result will be reported with a 20,000-draw task-bootstrap 95% interval regardless of significance. Failure to reject is not evidence of equivalence unless a separately declared equivalence margin is tested; no such test is preregistered. ## Compatibility gate for the later retrieval replication The gate is applied per model without inspecting future fresh-task outcomes. An interface qualifies when at least 12/60 cells contain an executor-accepted edit and at least half of its edit attempts are accepted. For each model, choose the qualifying interface with the highest accepted-edit cell rate; ties are broken by applicable-final-patch rate, then lower mean tokens, then lexicographic interface ID. A model with no qualifying interface is excluded from the later retrieval experiment and that exclusion is itself reported. Selection never uses hidden-test resolution. The later retrieval study receives a separate preregistration, fresh tasks, new identities, and no retroactive E09 modification. ## Statistical model and multiplicity The primary is paired task-level exact inference. Secondary paired binary contrasts use exact McNemar tests with Holm correction within the declared family and task-bootstrap intervals. A hierarchical binomial model includes interface, model, repository, interface-by-model interaction, and a task random intercept. With sparse events, exact estimates and intervals remain primary; a failed or unstable hierarchical fit is reported rather than interpreted. Repeated calls within a cell are process observations, not independent samples. Repositories and languages are descriptive strata; language is not separately identifiable because R003 is the only Python repository. ## Missingness, amendments, and stopping Malformed tool calls, unknown tools, assistant stops, invalid paths, ambiguous replacements, rejected diffs, timeouts after a valid response, and exhausted budgets are observed failures. A transport failure may be retried only if no valid assistant response was received. Incomplete infrastructure attempts are retained outside scored raw cells before an identical-identity retry. No cell is excluded for outcome quality. Implementation and deterministic conformance tests are committed before inference. Any subsequent pre-outcome correction is dated in `docs/PROTOCOL_AMENDMENTS.md`; any change after the first E09 response is exploratory. Execution stops only at 540 finalized cells or a documented safety/runtime blocker. The study uses no Torch and no cloud resources. LM Studio server lifecycle uses `lms` only; native REST endpoints exclusively load, inspect, and unload one model at a time.