File size: 3,464 Bytes
d61821a
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
# Metrics and artifacts

## Outcome hierarchy

The primary end-to-end metric is `resolved_at_1`: one bounded agent run produces
a patch that applies, passes all fail-to-pass tests, and introduces no failure
in the pass-to-pass suite. The primary retrieval metric is file recall@10;
function recall@10, MRR, and NDCG@10 are secondary retrieval outcomes.

Metric groups are registered in `src/agent_harness/metrics.py`:

| Group | Examples |
|---|---|
| Correctness | resolution, fail-to-pass, pass-to-pass, regressions |
| Localization | recall@k, MRR, NDCG, gold evidence within budget |
| Navigation | time/tokens/tool calls to first gold, repeated reads, churn |
| Editing | patch application, syntax, lint/build, patch size |
| Efficiency | input/output/tool tokens, calls, tests, wall time, throughput |
| Index | build/update time, query p50/p95, RAM, disk |
| Reliability | timeout, tool/parser/edit failures, failure stage |

All rates require explicit numerators, denominators, and exclusion counts in a
paper table. Token usage must distinguish model input, model output, cached
input when reported by the server, and tool-result text. A missing usage field
is missing data, not zero.

## Run identity and lineage

Every run has a deterministic ID derived from:

- experiment, task, and harness IDs;
- complete harness and model-config hashes;
- resolved LM Studio model key;
- context budget, seed, and repetition;
- repository commit and research-code revision.

Artifacts live under:

```text
results/raw/<experiment>/<harness>/<task>/<run_id>/
  run_manifest.json
  trajectory.jsonl
  patch.diff
  test_results.json
  final_metrics.json
```

The raw run directory is created exclusively and must never be overwritten.
`trajectory.jsonl` is append-only and contains monotonically sequenced events.
The manifest stores the resolved treatment and model metadata so a result never
depends only on a filename.

## Required trajectory events

- `run_started`: limits, environment, task metadata, execution order
- `model_call`: request hash, sampling, token usage, latency, stop reason
- `tool_call`: tool name, sanitized arguments, result hash, size, latency
- `retrieval_candidate`: source, query, file/symbol span, score, rank, fusion
- `file_read`: path, span, content hash, token count, relevance label after run
- `edit`: patch hash, target paths, application result
- `test_run`: command identity, exit status, duration, structured test counts
- `resource_sample`: process memory and relevant local-system measurements
- `run_finished`: outcome, failure stage, totals, artifact hashes

Secrets, environment contents, and unrelated user files must never enter
telemetry. Large source snippets are identified by repository SHA and content
hash; record only the content needed to audit the model-visible context.

## Derived datasets and tables

Raw JSONL is the source of truth. Deterministic analysis code should produce a
versioned tabular dataset with one row per run and separate candidate/query
tables. Planned paper tables are:

1. task and repository characteristics;
2. retrieval factorial effects and interactions;
3. query/interface and packing ablations;
4. end-to-end correctness and efficiency;
5. robustness by scenario and difficulty stratum;
6. dense-index systems results; and
7. failure taxonomy and exclusions.

Every derived artifact should record the input run IDs, analysis-code revision,
schema version, and creation timestamp.