# Paper plan Status: completed through E07. This file preserves the pre-execution plan; the realized design, deviations, evidence inventory, and release gates are recorded in `paper/main.tex` and `paper/reproducibility_manifest.md`. ## Working title **Beyond Grep: A Controlled Study of Repository-Navigation Harnesses for Local LLM Coding Agents** ## Candidate contribution The paper should make a narrower, defensible claim than “better harnesses make agents better.” The intended contribution is a controlled empirical decomposition of navigation architecture—retrieval source, structural expansion, interaction policy, tool interface, and context packing—under a fixed local coding model on repositories larger than the model's usable context. Potential publishable outputs are: 1. an open, immutable harness-treatment catalog with focused factorial blocks; 2. paired evidence connecting localization quality to end-to-end repair; 3. accuracy/latency/token trade-off curves rather than success alone; 4. robustness results for stale indexes and plausible lexical distractors; and 5. a reproducible local inference and telemetry protocol. Novelty depends on the completed literature review and should not be asserted until adjacent agent, code-retrieval, repository-level repair, and tool-use benchmarks have been systematically compared. ## Prespecified hypotheses - H1: lexical, syntax, and dense retrieval each improve file recall over exact search, but their effects are not purely additive. - H2: one-hop graph expansion improves multi-file task localization; two hops increase noise and cost more often than they improve resolution. - H3: iterative querying improves hard-task localization but consumes more tokens and wall time. - H4: structural skeleton packing outperforms whole-file packing at small fixed context budgets. - H5: improved file/function recall mediates part, but not all, of the effect on repair success. - H6: adaptive routing can approach full-stack accuracy with lower retrieval and context cost. These hypotheses must be revised or preregistered before looking at the held-out confirmatory results. ## Manuscript structure 1. **Abstract:** problem, controlled design, task/model scope, main effect sizes. 2. **Introduction:** why context overflow turns navigation into a systems and reasoning bottleneck; research questions and contributions. 3. **Related work:** code search and embeddings, repository-level program repair, LLM tool use, agent benchmarks, context selection, code graphs. 4. **Method:** treatment dimensions, fixed model, task construction, execution controls, metrics, and statistical analysis. 5. **Experiments:** repository/task statistics, E01-E05, compute and local serving details. 6. **Results:** prespecified contrasts, confidence intervals, cost-quality frontiers, subgroup and robustness analyses. 7. **Discussion:** mechanisms, practical recommendations, failure cases, and where complex harnesses do not pay for themselves. 8. **Threats and ethics:** contamination, generalization, annotation quality, local hardware/quantization effects, licensing, and energy/compute reporting. 9. **Reproducibility statement:** configs, commits, prompts, manifests, raw event schemas, analysis code, exclusions, and environment lockfiles. ## Milestones and release gates 1. Complete executable baseline H000 and validate task isolation. 2. Select repositories and construct a clean development task set. 3. Implement and unit-test retrieval sources, fusion, packing, and telemetry. 4. Freeze embedding query/chunk policies, index profiles, model variant, prompts, and budgets. 5. Run a small pilot; perform failure analysis and power simulation. 6. Freeze protocol, task split, configs, analysis, and exclusion rules. 7. Run confirmatory cells without treatment-specific intervention. 8. Reproduce tables from raw artifacts, then write and internally audit the paper.