File size: 3,592 Bytes
d61821a
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
# Reproducing the agent-harness study

This repository separates inexpensive artifact verification from expensive
model-in-the-loop reruns. Start with the saved-evidence path; it validates the
paper without loading an LLM or embedding model.

## Requirements

- macOS on the recorded Apple Silicon host for an exact runtime replication,
  or Python 3.11+ on another platform for analysis-only reproduction.
- Git and Go 1.26.5 for reconstructing and validating benchmark tasks.
- LM Studio 0.4.16+2 and the exact model variants in `docs/LM_STUDIO.md` only
  when rerunning inference.
- No PyTorch workload is used. Any future Torch experiment must use MPS; CPU
  and CUDA runs are outside the repository's execution policy.

Create an analysis environment:

```bash
python3.11 -m venv .venv
.venv/bin/python -m pip install --upgrade pip
.venv/bin/python -m pip install -e .
.venv/bin/python -m pip install -r requirements-analysis.txt
```

For exact dependency versions, install `requirements.lock` instead of the two
package-install commands above.

## Verify the archived evidence

The compact Zenodo archive contains all derived inputs required by these
commands:

```bash
PYTHONPATH=src .venv/bin/python -m agent_harness.cli validate
PYTHONPATH=src .venv/bin/python -m pytest -q
PYTHONPATH=src .venv/bin/python scripts/check_study2_paper_consistency.py
PYTHONPATH=src .venv/bin/python scripts/check_study5_paper_consistency.py
MPLBACKEND=Agg MPLCONFIGDIR=/tmp/agent-harness-mpl \
  PYTHONPATH=src .venv/bin/python scripts/analyze_study5.py
```

Earlier-study analysis commands and frozen revisions are listed in
`paper/reproducibility_manifest.md`. The fail-closed Study 5 analysis verifies
the exact E13--E16 cell grids, revisions, raw hashes, model-residency records,
required artifacts, and final outcome funnels before emitting results.

## Build the paper

```bash
mkdir -p tmp/pdfs
cd paper
latexmk -pdf -interaction=nonstopmode -halt-on-error \
  -outdir=../tmp/pdfs main.tex
cp ../tmp/pdfs/main.pdf \
  ../output/pdf/dissecting_repository_scale_code_agent_harnesses.pdf
```

Then verify the PDF against `output/pdf/SHA256SUMS` and inspect a PNG rendering
of every page. The publication package is rebuilt only after committing:

```bash
PYTHONPATH=src .venv/bin/python scripts/build_publication_package.py
```

## Rerun inference

Inference requires the ignored repository checkouts under `data/repos`, the
ignored indexes, and the exact local LM Studio models documented in
`docs/LM_STUDIO.md`. Run the matching preflight before every experiment family.
For Study 5:

```bash
PYTHONPATH=src .venv/bin/python scripts/preflight_study5.py
PYTHONPATH=src .venv/bin/python -m agent_harness.cli run-study5 --experiment E13
PYTHONPATH=src .venv/bin/python -m agent_harness.cli run-study5 --experiment E14
PYTHONPATH=src .venv/bin/python -m agent_harness.cli run-study5 --experiment E15
PYTHONPATH=src .venv/bin/python -m agent_harness.cli run-study5 --experiment E16
```

The runner uses `lms` for server lifecycle and official LM Studio APIs for
model inspection, load, and unload. Generative and embedding models are never
intentionally co-resident. Durable run identities make interruption recovery
idempotent.

## Artifact boundary

The compact archive excludes raw trajectories, repository checkouts, index
caches, and model weights. Raw trajectories should be uploaded as a separate
checksum-addressed Zenodo file or linked deposit. Some immutable provenance
records retain machine-local paths; these identify the original environment
and are not required paths for reproduction.