agent-harness / docs /PILOT_DATASET.md
cuber12's picture
Publish agent harness research code and paper artifacts
d61821a verified
|
Raw
History Blame Contribute Delete
2.6 kB

Development pilot dataset

The E00 development pilot uses the official GitLab Runner repository, a Go project hosted on GitLab. This dataset is for pipeline debugging and pilot effect-size estimation only; it is not a confirmatory benchmark split.

Pinned checkout

Field Value
Repository https://gitlab.com/gitlab-org/gitlab-runner.git
Release v19.0.1
Commit c2831b75a3ff0782dca8f64498cbc6f71c76819e
Tracked files 1,229
Readable tracked bytes approximately 13.35 MB
Go/Proto files 845
Go/Proto lines 211,912
Go/Proto bytes 6,153,499

The model is loaded with a 262,144-token context. The Go/Proto subset alone would fit only if the tokenizer averaged more than 23 bytes per token. This is a conservative size sanity check, not the final tokenizer measurement. Before confirmatory evaluation, repository size must be measured with the exact local Qwen tokenizer file and the measurement script/version must be recorded.

Task construction

Five real fixes are converted into localization tasks. For every task, the searchable snapshot is the parent of the fix commit. The visible statement is an issue-style description that does not name the gold file or symbol. Changed production locations that already exist at the parent commit form the hidden retrieval gold.

Task Gold commit Topic Difficulty
TASK_GR_001 991cf3c0 proactive authentication during checkout lazy fetch medium, two files
TASK_GR_002 fd33f65d proxy-exec inherited-pipe hang medium, one file
TASK_GR_003 9968f744 adaptive concurrency under network retries hard, cross-module
TASK_GR_004 6cd0dce8 S3 checksum defaults for custom endpoints medium, domain terminology
TASK_GR_005 8c5b8c70 autoscaler context-cause classification medium, semantic

These tasks have validation_status = "retrieval_ready". They must not be used for end-to-end repair claims until hidden regression tests are extracted, verified to fail on the base commit, verified to pass on the gold commit, and checked for flakiness.

E00 treatments

  • H000: deterministic literal occurrence ranking over issue-derived terms.
  • H001: in-memory BM25 with a bounded fuzzy path-token bonus.
  • H003: Qwen3 Embedding 0.6B with exact flat cosine ranking.

All treatments use the same 120-line chunks, 20-line overlap, and 16,000-character ceiling in this pilot. This common chunking is a deliberate control. E00 has 15 cells: five tasks by three harnesses.