File size: 11,692 Bytes
d61821a | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 | # Scientific protocol
## Study objective
Measure how repository-navigation harness design changes the ability of a fixed
LLM coding agent to localize relevant code and produce correct patches when the
repository cannot fit in the available context window.
The causal object is the **harness**, not the model. Every confirmatory run uses
the same task statement, repository commit, model, prompt policy, context
budget, execution limits, and validation environment. Only the treatment
encoded by the harness specification changes.
## Research questions
1. How much do lexical, structural, and dense retrieval improve localization
and end-to-end repair over exact/regex search?
2. Are retrieval sources complementary, or does their fusion introduce mostly
redundant context?
3. Does structural graph expansion improve multi-file localization, and when
does a second hop create harmful context noise?
4. Do iterative query reformulation and specialized tools improve outcomes
enough to justify their token and latency costs?
5. Which context-packing policy best converts retrieved evidence into correct
edits under a fixed token budget?
6. How robust are the results to stale indexes, plausible distractors, task
difficulty, repository scale, and run-to-run variation?
7. How do the same retrieval/interface choices behave when the evaluated Qwen
model itself chooses live searches, reads, edits, and public tests?
## Fixed agent model and serving policy
The agent model is fixed to **Qwen3.6-35B-A3B**, the mixture-of-experts model
with 35B total parameters and 3B active parameters identified by the
[official Qwen release](https://qwen.ai/blog?id=qwen3.6-35b-a3b) and
[official model card](https://huggingface.co/Qwen/Qwen3.6-35B-A3B). It is served
locally by LM Studio at `http://127.0.0.1:1234`.
Before any run, the runner must:
1. query both `/v1/models` and `/api/v1/models`;
2. find exactly one inference-visible record matching `Qwen3.6-35B-A3B`;
3. record and verify the exposed model key, selected variant, format,
quantization, loaded context length, and reasoning default;
4. abort on no match or ambiguous matches; and
5. store the model-config hash in the immutable run manifest.
There is no automatic fallback to another local or remote model. The initial
sampling policy is temperature `0.0`, top-p `1.0`, seed `0`, and reasoning mode
`on`. The observed M001 runtime is the MLX `4bit` variant with a 262,144-token
loaded context. Model revision, LM Studio version, and remaining inference
settings must be frozen before confirmatory runs. Because local inference can remain
nondeterministic, a pilot must estimate run-to-run variance and determine the
number of repetitions used in the confirmatory manifests.
E07 pins a second runtime profile, M002, for the same weights and 4-bit variant
at 65,536 tokens. The `lms` CLI is the sole server lifecycle controller. Native
REST endpoints inspect/load/unload models. E07 verifies exclusive residency at
every transition; a dense query unloads Qwen, loads EMB001, embeds the query,
unloads EMB001, and reloads Qwen before continuing the saved conversation.
No PyTorch workload is required by this scaffold. If a future component uses
PyTorch on this machine, it must verify and use MPS; CPU and CUDA experiments
are outside the allowed protocol.
## Fixed embedding model
Dense treatments use **Qwen3 Embedding 0.6B** through the same local LM Studio
server, but it remains scientifically distinct from the Qwen3.6 coding agent.
EMB001 pins the LM Studio key `text-embedding-qwen3-embedding-0.6b`, GGUF Q8_0
format, 8,192-token loaded context, 32,768-token maximum context, and
1,024-dimensional L2-normalized output.
Every dense run must perform discovery and an inference preflight before index
construction. Chunking, query instructions, fusion, and index parameters are
separate experimental choices and must not vary silently across harnesses.
## Harness catalog and controlled contrasts
The catalog contains 21 independently addressable treatments. Each treatment
has a separate TOML file and canonical hash.
| Family | Harnesses | Controlled question |
|---|---:|---|
| Retrieval factorial | H000-H007 | All `2 x 2 x 2` combinations of lexical, syntax, and dense retrieval |
| Graph depth | H007-H009 | Zero, one, or two graph-expansion hops with retrieval fixed |
| Query/interface | H008, H010-H012 | `2 x 2` query policy by tool-interface design |
| Context packing | H008, H013-H015 | Snippets, skeletons, whole files, or role summaries |
| Controls | H016-H019 | No search, random context, oracle files, and oracle functions |
| Adaptive policy | H020 | Whether dynamic source selection beats a fixed full stack |
This is intentionally not the unrestricted Cartesian product of all seven
dimensions (`2 x 2 x 2 x 3 x 2 x 2 x 4 = 384`). The 21-treatment catalog uses
focused factorial blocks so that each contrast is interpretable and the study
is computationally feasible. New combinations receive new identifiers; an
existing treatment is never overwritten after it has generated a run.
## Staged experiments
### E00: development pilot
Run H000, H001, and H003 on five retrieval-ready GitLab Runner tasks. E00 exists
to validate task snapshots, telemetry, chunking, embedding cache behavior, and
metric computation. It is explicitly non-confirmatory and must not be pooled
with E01-E05 paper results.
### E01: retrieval factorial
Evaluate H000-H009 without allowing editing. Given the task statement, each
harness produces a ranked evidence set under the same context budget. Primary
outcomes are file recall@10, function recall@10, MRR, NDCG@10, and whether all
gold locations fit in the packed context. This isolates retrieval quality from
the model's editing ability.
### E02: interaction and packing
Evaluate query policy, tool interface, and packing on localization tasks. Log
every query, candidate, score, rank, graph edge, returned token, and model tool
decision. Compare both localization accuracy and cost to reach the first gold
location. Responses that violate the frozen tool-output schema or select paths
outside the supplied evidence are retained and scored as empty selections, not
dropped as missing observations. See `PROTOCOL_AMENDMENTS.md`.
### E03: end-to-end repair
Allow repository reads, edits, and bounded test execution. The primary outcome
is `resolved_at_1`: the generated patch applies, all fail-to-pass tests pass,
and no pass-to-pass regression occurs. Secondary outcomes cover localization,
patch validity, tool use, tokens, latency, and failure stage.
### E04: robustness
Re-run selected baselines with a stale index and with controlled plausible
distractors. Perturbations must preserve the repository's build and tests so
that the only intended change is navigation difficulty.
### E05: dense-index systems study
Hold chunks, embeddings, queries, and rankings constant where possible while
comparing exact FAISS, approximate FAISS, and a pinned persistent vector
database. Report recall against exact search, build/update latency, query p50
and p95, RAM, and disk size. This experiment cannot start until the embedding
model and backend versions are frozen.
### E07: live-agent repair
Evaluate H000, H003, H007, H008, H011, H016, and H018 on the ten frozen
end-to-end-ready tasks. Qwen must issue the actual search/read/apply/test/finish
calls. Hidden tests and gold changes are unavailable during interaction; H018
receives file names only. The complete paired grid has 70 cells, cyclically
counterbalanced treatment order, 12 model calls, 12 tool calls, two public test
calls, and a 1,800-second ceiling per cell. The primary endpoint is hidden-test
`resolved_at_1`. The frozen details and declared contrasts are in
`E07_LIVE_AGENT_PROTOCOL.md`.
## Repository and task construction
Use one or more version-pinned GitLab repositories whose checked-out source,
tests, and build metadata exceed the model's usable context after tokenization.
Repository size must be reported as files, source lines, bytes, and tokens under
the exact tokenizer used for budgeting; bytes or Git history alone do not prove
that the code exceeds context.
Tasks should be reproducible from public commits or independently authored
mutations. Each task manifest contains the base commit, visible problem
statement, hidden gold patch, gold files and symbols, fail-to-pass tests,
pass-to-pass tests, language, provenance, and difficulty annotations. Remove
tasks that are flaky, cannot be built in isolation, leak the patch through
generated artifacts, or require unavailable external services.
Task difficulty should include at least:
- number of gold files and symbols;
- shortest dependency distance from obvious query matches to gold code;
- lexical overlap between task statement and gold code;
- number of plausible distractors;
- patch size and whether edits cross modules; and
- baseline retrieval rank of the first gold location.
Split tasks before confirmatory evaluation. Use development tasks for prompt,
budget, chunking, and hyperparameter choices; never tune these choices on the
held-out confirmatory split.
## Execution controls
- Start every task from the exact clean base commit in an isolated worktree or
disposable copy.
- Prebuild indexes from the same base commit except in the stale-index scenario.
- Give all non-control harnesses the same exact read/write/test primitives.
- Apply identical context, tool-call, test-run, token, and wall-clock limits.
- Do not reveal gold files, symbols, patches, or hidden tests except to the
explicitly labeled oracle controls.
- Randomize harness execution order within each task and record that order.
- Cache immutable indexes, but never share model conversation state between
runs.
- Record failures and timeouts as outcomes; do not silently retry only failed
cells.
## Statistical analysis plan
The task is the principal sampling unit and every task is evaluated under every
eligible harness, producing paired observations.
- Report point estimates with 95% task-cluster bootstrap confidence intervals.
- For binary repair success, report paired risk differences and risk ratios;
use a task-blocked logistic model or McNemar test for prespecified pairwise
contrasts.
- For ranks, tokens, tool calls, and latency, report medians and paired
bootstrap differences in addition to means.
- Fit a hierarchical model with a task random intercept when estimating the
effects and interactions of L, S, and D across the factorial block.
- Correct the prespecified family of pairwise tests with Holm's method. Treat
all unregistered follow-up analyses as exploratory.
- Report effect sizes and uncertainty, not only p-values.
Run a pilot on development tasks to estimate baseline resolution and paired
discordance. Use those values for simulation-based power analysis; do not claim
an a priori task count without those empirical inputs. The confirmatory sample
size, exclusions, primary contrasts, and analysis code should then be
preregistered and frozen.
## Validity threats
The initial study uses one agent model and may not generalize to other models,
languages, repositories, or context sizes. Public issue tasks can contain
training-data contamination. Oracle annotations can be incomplete. Tree-sitter
support differs by language. Dense retrieval results are sensitive to the
embedding model, query instruction, chunking, and index parameters. A local quantized model can
also differ from the official checkpoint. These factors must be documented,
and the strongest claims should remain within the evaluated model, task
population, and implementation versions.
|