# Development pilot dataset The E00 development pilot uses the official [GitLab Runner repository](https://gitlab.com/gitlab-org/gitlab-runner/), a Go project hosted on GitLab. This dataset is for pipeline debugging and pilot effect-size estimation only; it is not a confirmatory benchmark split. ## Pinned checkout | Field | Value | |---|---| | Repository | `https://gitlab.com/gitlab-org/gitlab-runner.git` | | Release | `v19.0.1` | | Commit | `c2831b75a3ff0782dca8f64498cbc6f71c76819e` | | Tracked files | 1,229 | | Readable tracked bytes | approximately 13.35 MB | | Go/Proto files | 845 | | Go/Proto lines | 211,912 | | Go/Proto bytes | 6,153,499 | The model is loaded with a 262,144-token context. The Go/Proto subset alone would fit only if the tokenizer averaged more than 23 bytes per token. This is a conservative size sanity check, not the final tokenizer measurement. Before confirmatory evaluation, repository size must be measured with the exact local Qwen tokenizer file and the measurement script/version must be recorded. ## Task construction Five real fixes are converted into localization tasks. For every task, the searchable snapshot is the parent of the fix commit. The visible statement is an issue-style description that does not name the gold file or symbol. Changed production locations that already exist at the parent commit form the hidden retrieval gold. | Task | Gold commit | Topic | Difficulty | |---|---|---|---| | TASK_GR_001 | `991cf3c0` | proactive authentication during checkout lazy fetch | medium, two files | | TASK_GR_002 | `fd33f65d` | proxy-exec inherited-pipe hang | medium, one file | | TASK_GR_003 | `9968f744` | adaptive concurrency under network retries | hard, cross-module | | TASK_GR_004 | `6cd0dce8` | S3 checksum defaults for custom endpoints | medium, domain terminology | | TASK_GR_005 | `8c5b8c70` | autoscaler context-cause classification | medium, semantic | These tasks have `validation_status = "retrieval_ready"`. They must not be used for end-to-end repair claims until hidden regression tests are extracted, verified to fail on the base commit, verified to pass on the gold commit, and checked for flakiness. ## E00 treatments - H000: deterministic literal occurrence ranking over issue-derived terms. - H001: in-memory BM25 with a bounded fuzzy path-token bonus. - H003: Qwen3 Embedding 0.6B with exact flat cosine ranking. All treatments use the same 120-line chunks, 20-line overlap, and 16,000-character ceiling in this pilot. This common chunking is a deliberate control. E00 has 15 cells: five tasks by three harnesses.