agent-harness / docs /PAPER_PLAN.md
cuber12's picture
Publish agent harness research code and paper artifacts
d61821a verified
|
Raw
History Blame Contribute Delete
3.98 kB

Paper plan

Status: completed through E07. This file preserves the pre-execution plan; the realized design, deviations, evidence inventory, and release gates are recorded in paper/main.tex and paper/reproducibility_manifest.md.

Working title

Beyond Grep: A Controlled Study of Repository-Navigation Harnesses for Local LLM Coding Agents

Candidate contribution

The paper should make a narrower, defensible claim than “better harnesses make agents better.” The intended contribution is a controlled empirical decomposition of navigation architecture—retrieval source, structural expansion, interaction policy, tool interface, and context packing—under a fixed local coding model on repositories larger than the model's usable context.

Potential publishable outputs are:

  1. an open, immutable harness-treatment catalog with focused factorial blocks;
  2. paired evidence connecting localization quality to end-to-end repair;
  3. accuracy/latency/token trade-off curves rather than success alone;
  4. robustness results for stale indexes and plausible lexical distractors; and
  5. a reproducible local inference and telemetry protocol.

Novelty depends on the completed literature review and should not be asserted until adjacent agent, code-retrieval, repository-level repair, and tool-use benchmarks have been systematically compared.

Prespecified hypotheses

  • H1: lexical, syntax, and dense retrieval each improve file recall over exact search, but their effects are not purely additive.
  • H2: one-hop graph expansion improves multi-file task localization; two hops increase noise and cost more often than they improve resolution.
  • H3: iterative querying improves hard-task localization but consumes more tokens and wall time.
  • H4: structural skeleton packing outperforms whole-file packing at small fixed context budgets.
  • H5: improved file/function recall mediates part, but not all, of the effect on repair success.
  • H6: adaptive routing can approach full-stack accuracy with lower retrieval and context cost.

These hypotheses must be revised or preregistered before looking at the held-out confirmatory results.

Manuscript structure

  1. Abstract: problem, controlled design, task/model scope, main effect sizes.
  2. Introduction: why context overflow turns navigation into a systems and reasoning bottleneck; research questions and contributions.
  3. Related work: code search and embeddings, repository-level program repair, LLM tool use, agent benchmarks, context selection, code graphs.
  4. Method: treatment dimensions, fixed model, task construction, execution controls, metrics, and statistical analysis.
  5. Experiments: repository/task statistics, E01-E05, compute and local serving details.
  6. Results: prespecified contrasts, confidence intervals, cost-quality frontiers, subgroup and robustness analyses.
  7. Discussion: mechanisms, practical recommendations, failure cases, and where complex harnesses do not pay for themselves.
  8. Threats and ethics: contamination, generalization, annotation quality, local hardware/quantization effects, licensing, and energy/compute reporting.
  9. Reproducibility statement: configs, commits, prompts, manifests, raw event schemas, analysis code, exclusions, and environment lockfiles.

Milestones and release gates

  1. Complete executable baseline H000 and validate task isolation.
  2. Select repositories and construct a clean development task set.
  3. Implement and unit-test retrieval sources, fusion, packing, and telemetry.
  4. Freeze embedding query/chunk policies, index profiles, model variant, prompts, and budgets.
  5. Run a small pilot; perform failure analysis and power simulation.
  6. Freeze protocol, task split, configs, analysis, and exclusion rules.
  7. Run confirmatory cells without treatment-specific intervention.
  8. Reproduce tables from raw artifacts, then write and internally audit the paper.