| --- |
| license: other |
| tags: |
| - code-agents |
| - software-engineering |
| - llm-evaluation |
| - reproducibility |
| --- |
| |
| # Agent Harness Research |
|
|
| [](https://doi.org/10.5281/zenodo.21781711) |
|
|
| **Author:** Mandeep Sidhu, Independent Researcher |
|
|
| **Release:** 1.0.0 preprint candidate |
|
|
| **Paper:** [Dissecting Repository-Scale Code-Agent Harnesses](output/pdf/dissecting_repository_scale_code_agent_harnesses.pdf) |
|
|
| **Archival DOI:** [10.5281/zenodo.21781711](https://doi.org/10.5281/zenodo.21781711) |
|
|
| This repository is an experimental platform for a scientific study of how |
| repository-navigation harnesses affect an LLM coding agent on codebases that do |
| not fit in the model context window. |
|
|
| Study 1 fixes **Qwen3.6-35B-A3B** (35B total parameters, 3B active parameters). |
| The completed Studies 2--5 evaluate that model, **GPT-OSS-20B**, and |
| **Qwen3-Coder-30B-A3B-Instruct**, |
| served locally by LM Studio on port `1234`. Experiments abort if the exposed |
| model, variant, quantization, context, or reasoning default does not match. The |
| canonical model is documented by the [Qwen release](https://qwen.ai/blog?id=qwen3.6-35b-a3b) |
| and [official model repository](https://huggingface.co/Qwen/Qwen3.6-35B-A3B). |
|
|
| LM Studio exposes model discovery through `/api/v1/models` and OpenAI-compatible |
| inference and tool calling through `/v1/chat/completions`. See the |
| [LM Studio server documentation](https://lmstudio.ai/docs/developer/core/server), |
| [model listing documentation](https://lmstudio.ai/docs/developer/rest/list), and |
| [tool-use documentation](https://lmstudio.ai/docs/developer/openai-compat/tools). |
|
|
| ## Completed research program |
|
|
| - 21 immutable harness specifications (`H000` through `H020`) |
| - Fifteen completed experiment families (`E01`--`E05` and `E07`--`E16`) |
| - 5,453 audited cells: 965 Study 1, 912 Study 2, 540 Study 3, 210 Study 4, |
| and 2,826 Study 5 |
| - Separate robustness scenarios and vector-index backend profiles |
| - Strict local-model discovery with no silent model fallback |
| - Append-only JSONL telemetry and deterministic run identities |
| - Confirmatory statistical analysis with paired exact tests, task bootstrap |
| intervals, Holm correction, and a mixed-effects factorial model |
| - A reproducible research paper with generated figures/tables, checksums, and |
| page-by-page rendered-PDF quality assurance |
|
|
| ## Powered Study 2 |
|
|
| - 60 validated real-change tasks across GitLab Runner, GitLab CLI, and python-gitlab |
| - Go and Python repositories whose full source exceeds the 65,536-token treatment cap |
| - Five component harnesses plus controlled Agentless-style and SWE-agent-style baselines |
| - Qwen3.6-35B-A3B and GPT-OSS-20B under identical local-runtime controls |
| - 840 deterministic main cells plus 72 preregistered stochastic reliability cells |
| - Exact paired testing, task-cluster intervals, multiplicity control, and hierarchical analysis |
| - Primary result: H007 hybrid and H000 lexical both resolved 2/60 Qwen tasks |
| (risk difference 0.000, 95% task-bootstrap CI [-0.050, 0.050], exact p=1) |
| - No secondary contrast survived correction; GPT-OSS produced no accepted |
| repair and exposed a strong model/tool-protocol compatibility limitation |
|
|
| ## Protocol-normalized Studies 3--4 |
|
|
| - E09: 60 tasks x 3 models x 3 edit interfaces = 540 real local-LLM cells |
| - GPT-OSS exact replacement accepted edits on 27/60 tasks versus 0/60 for raw |
| diff (RD +0.450, 95% CI [0.333, 0.583], exact p=1.49e-8) |
| - Resolution-blind gate: P002 for Qwen3.6; P003 for GPT-OSS and Qwen3-Coder |
| - E10: 20 fresh tasks x 3 gated models x exact/hybrid/oracle retrieval = 180 |
| - Hybrid resolved 1/60 model-task pairs versus exact 0/60 (RD +0.0167, |
| task-cluster CI [0, 0.05], exact p=1); oracle file names resolved 4/60 |
| - E11/E12: 30 prespecified stochastic and context-scarcity sensitivity cells |
|
|
| ## Harness-design Study 5 |
|
|
| - E13: 60 tasks x 8 lexical/syntax/dense combinations x 3 models = 1,440 cells |
| - E14: 20 tasks x 3 retrievers x 3 edit interfaces x 3 models = 540 cells |
| - E15: 20 tasks x 9 graph/query/interface/packing harnesses x 3 models = 540 cells |
| - E16: six deterministically selected archetypes x 17 fresh tasks x 3 models = 306 cells |
| - No E13 component main effect survives correction; E14 finds no adjusted |
| retrieval-by-action interaction |
| - Iteration and whole-file packing reliably increase cost without a demonstrated |
| applicable-patch benefit |
| - E16 quality-rank transfer is weak (accepted Spearman rho 0.04; applicable |
| rho -0.21), while token-cost rank is stable (rho 0.94) |
| - Only one E16 cell resolves, so the study reports stage-aware trends and |
| Pareto candidates rather than a universal winning harness |
|
|
| ## Repository layout |
|
|
| ```text |
| configs/ |
| harnesses/ one immutable file per H000-H020 treatment |
| experiments/ staged E00-E16 experiment manifests |
| models/ fixed Qwen/LM Studio runtime policy |
| embeddings/ independently pinned dense-retrieval model |
| scenarios/ robustness perturbations |
| backends/ dense-index systems profiles |
| src/agent_harness/ schemas, LM Studio client, interfaces, metrics, telemetry |
| tasks/ task manifests and frozen split lists |
| results/ immutable local raw artifacts, analyses, and reports |
| docs/ protocol, catalog, metrics, runtime, and paper plan |
| tests/ standard-library unit tests |
| paper/ LaTeX manuscript and bibliography |
| output/pdf/ stable rendered research paper |
| ``` |
|
|
| Dense retrieval uses the separately pinned **Qwen3 Embedding 0.6B** model |
| (`text-embedding-qwen3-embedding-0.6b`) through LM Studio. The local GGUF Q8_0 |
| model produces 1,024-dimensional normalized vectors and exposes a 32,768-token |
| maximum context. The generative Qwen agent and embedding model remain separate |
| experimental components. |
| |
| ## Quick start |
| |
| Create a clean Python 3.11 environment and install the harness plus its pinned |
| research dependencies: |
| |
| ```bash |
| python3.11 -m venv .venv |
| .venv/bin/python -m pip install --upgrade pip |
| .venv/bin/python -m pip install -r requirements.lock |
| .venv/bin/python -m pip install -e . --no-deps |
| ``` |
| |
| For the saved-evidence reproduction path, LM Studio and model weights are not |
| required. See [`REPRODUCING.md`](REPRODUCING.md) for the exact artifact and |
| inference workflows. |
| |
| ```bash |
| PYTHONPATH=src python3 -m agent_harness.cli validate |
| PYTHONPATH=src python3 -m agent_harness.cli list-harnesses |
| PYTHONPATH=src python3 -m agent_harness.cli plan --experiment E03 |
| PYTHONPATH=src python3 -m agent_harness.cli probe-model |
| PYTHONPATH=src python3 -m agent_harness.cli probe-embedding --infer |
| PYTHONPATH=src python3 -m agent_harness.cli run-pilot --experiment E00 |
| PYTHONPATH=src .venv/bin/python -m agent_harness.cli run-live-agent --experiment E07 |
| PYTHONPATH=src .venv/bin/python scripts/preflight_study2.py |
| PYTHONPATH=src .venv/bin/python -m agent_harness.cli run-study2 |
| PYTHONPATH=src .venv/bin/python -m agent_harness.cli run-study2-reliability |
| PYTHONPATH=src .venv/bin/python scripts/analyze_confirmatory.py |
| MPLCONFIGDIR=/tmp/agent-harness-mpl PYTHONPATH=src .venv/bin/python scripts/analyze_live_agent.py |
| MPLCONFIGDIR=/tmp/agent-harness-mpl PYTHONPATH=src .venv/bin/python scripts/analyze_study2.py |
| MPLCONFIGDIR=/tmp/agent-harness-mpl PYTHONPATH=src .venv/bin/python scripts/analyze_study3.py |
| MPLCONFIGDIR=/tmp/agent-harness-mpl PYTHONPATH=src .venv/bin/python scripts/analyze_study4.py |
| PYTHONPATH=src .venv/bin/python scripts/analyze_study4_ancillary.py |
| MPLBACKEND=Agg MPLCONFIGDIR=/tmp/agent-harness-mpl PYTHONPATH=src .venv/bin/python scripts/analyze_study5.py |
| PYTHONPATH=src .venv/bin/python scripts/check_study5_paper_consistency.py |
| PYTHONPATH=src .venv/bin/python -m pytest -q |
| ``` |
| |
| `probe-model` discovers and validates the exact model key, MLX 4-bit variant, |
| 262,144-token loaded context, and reasoning default. Add `--infer` to perform a |
| small completion and verify the response semantics after discovery. |
| |
| ## Scientific protocol |
| |
| Start with the [scientific protocol](docs/SCIENTIFIC_PROTOCOL.md), |
| [harness catalog](docs/HARNESS_CATALOG.md), [LM Studio policy](docs/LM_STUDIO.md), |
| [metric and artifact schema](docs/METRICS_AND_ARTIFACTS.md), observed |
| [runtime baseline](docs/RUNTIME_BASELINE.md), [pilot dataset](docs/PILOT_DATASET.md), |
| [pilot results](docs/PILOT_RESULTS.md), [memory-safety protocol](docs/MEMORY_SAFETY.md), |
| and [paper plan](docs/PAPER_PLAN.md). |
| Study 2 is governed by its [preregistration](docs/STUDY2_PREREGISTRATION.md), |
| [implementation contract](docs/STUDY2_IMPLEMENTATION.md), |
| [design audit](docs/STUDY2_DESIGN_AUDIT.json), and |
| [preflight procedure](docs/STUDY2_PREFLIGHT.md). |
| Study 3 and 4 use [Study 3 preregistration](docs/STUDY3_PREREGISTRATION.md), |
| [Study 4 preregistration](docs/STUDY4_PREREGISTRATION.md), and their frozen |
| design audits, compatibility gate, preflights, and ancillary manifest audit. |
| Study 5 uses its [prospective preregistration](docs/STUDY5_PREREGISTRATION.md), |
| frozen cell manifests, and sequential E16 selection ledger. |
| Configuration files are part of the experimental treatment and become immutable |
| once a recorded run uses them. Any treatment change receives a new harness ID |
| or experiment version. |
| |
| The completed manuscript is |
| [`output/pdf/dissecting_repository_scale_code_agent_harnesses.pdf`](output/pdf/dissecting_repository_scale_code_agent_harnesses.pdf). |
| Its source and build instructions are under [`paper/`](paper/). The immutable |
| local analysis used by the paper is |
| `results/derived/confirmatory_analysis_174ce71bcbce` plus |
| `results/derived/e07`, `results/derived/study2`, `results/derived/study3`, |
| `results/derived/study4`, `results/derived/study4_ancillary`, and |
| `results/derived/study5`. |
|
|
| ## Citation and release |
|
|
| Citation metadata is provided in [`CITATION.cff`](CITATION.cff), and the exact |
| Zenodo record fields are in [`.zenodo.json`](.zenodo.json). The canonical |
| archival DOI is |
| [10.5281/zenodo.21781711](https://doi.org/10.5281/zenodo.21781711). |
| Publication upload instructions are in |
| [`README_ZENODO.md`](README_ZENODO.md). |
|
|
| The repository intentionally contains one canonical PDF. Raw trajectories, |
| repository checkouts, index caches, and model weights are excluded from Git and |
| from the compact release; raw trajectories belong in a separate |
| checksum-addressed Zenodo archive. |
|
|
| ## Licensing |
|
|
| Original software is MIT licensed under [`LICENSE`](LICENSE). The manuscript, |
| documentation, figures, tables, and original derived research artifacts are CC |
| BY 4.0 under [`LICENSE-DATA`](LICENSE-DATA). Benchmark patches retain their |
| upstream licenses; see [`THIRD_PARTY_NOTICES.md`](THIRD_PARTY_NOTICES.md). |
|
|