Metrics and artifacts
Outcome hierarchy
The primary end-to-end metric is resolved_at_1: one bounded agent run produces
a patch that applies, passes all fail-to-pass tests, and introduces no failure
in the pass-to-pass suite. The primary retrieval metric is file recall@10;
function recall@10, MRR, and NDCG@10 are secondary retrieval outcomes.
Metric groups are registered in src/agent_harness/metrics.py:
| Group | Examples |
|---|---|
| Correctness | resolution, fail-to-pass, pass-to-pass, regressions |
| Localization | recall@k, MRR, NDCG, gold evidence within budget |
| Navigation | time/tokens/tool calls to first gold, repeated reads, churn |
| Editing | patch application, syntax, lint/build, patch size |
| Efficiency | input/output/tool tokens, calls, tests, wall time, throughput |
| Index | build/update time, query p50/p95, RAM, disk |
| Reliability | timeout, tool/parser/edit failures, failure stage |
All rates require explicit numerators, denominators, and exclusion counts in a paper table. Token usage must distinguish model input, model output, cached input when reported by the server, and tool-result text. A missing usage field is missing data, not zero.
Run identity and lineage
Every run has a deterministic ID derived from:
- experiment, task, and harness IDs;
- complete harness and model-config hashes;
- resolved LM Studio model key;
- context budget, seed, and repetition;
- repository commit and research-code revision.
Artifacts live under:
results/raw/<experiment>/<harness>/<task>/<run_id>/
run_manifest.json
trajectory.jsonl
patch.diff
test_results.json
final_metrics.json
The raw run directory is created exclusively and must never be overwritten.
trajectory.jsonl is append-only and contains monotonically sequenced events.
The manifest stores the resolved treatment and model metadata so a result never
depends only on a filename.
Required trajectory events
run_started: limits, environment, task metadata, execution ordermodel_call: request hash, sampling, token usage, latency, stop reasontool_call: tool name, sanitized arguments, result hash, size, latencyretrieval_candidate: source, query, file/symbol span, score, rank, fusionfile_read: path, span, content hash, token count, relevance label after runedit: patch hash, target paths, application resulttest_run: command identity, exit status, duration, structured test countsresource_sample: process memory and relevant local-system measurementsrun_finished: outcome, failure stage, totals, artifact hashes
Secrets, environment contents, and unrelated user files must never enter telemetry. Large source snippets are identified by repository SHA and content hash; record only the content needed to audit the model-visible context.
Derived datasets and tables
Raw JSONL is the source of truth. Deterministic analysis code should produce a versioned tabular dataset with one row per run and separate candidate/query tables. Planned paper tables are:
- task and repository characteristics;
- retrieval factorial effects and interactions;
- query/interface and packing ablations;
- end-to-end correctness and efficiency;
- robustness by scenario and difficulty stratum;
- dense-index systems results; and
- failure taxonomy and exclusions.
Every derived artifact should record the input run IDs, analysis-code revision, schema version, and creation timestamp.