kiyam's picture
Split perturbations by eval split (msmarco-dev/trec-dl-2019/trec-dl-2020), add TSV format with original+perturbed query text, add data_download scripts for MS MARCO/TREC-DL originals
391165c verified
|
Raw
History Blame Contribute Delete
3.62 kB
# Lost in Decoding — Evaluation Artifacts
**Provenance:** derived evaluation artifacts created from the official MS
MARCO Dev queries and official TREC Deep Learning 2019/2020 topics, for this
paper (*Lost in Decoding? Reproducing and Stress-Testing the Look-Ahead Prior
in Generative Retrieval*, SIGIR 2026,
[DOI: 10.1145/3805712.3808567](https://doi.org/10.1145/3805712.3808567) ·
[arXiv:2604.23396](https://arxiv.org/abs/2604.23396)), not part of the
original PAG release.
The generation code and metadata here are Apache-2.0-licensed. The perturbed
query **text** is a derived artifact of the original MS MARCO / TREC-DL query
text and remains subject to the same provenance constraints as that source
text — see [`../LICENSE.md`](../LICENSE.md).
## `msmarco-dev/`, `trec-dl-2019/`, `trec-dl-2020/`
Query variations (misspelling, reordering, synonym replacement, paraphrasing,
naturality normalization) used for the RQ2 robustness evaluation, one
directory per evaluation split:
| Directory | Split | Queries |
|---|---|---|
| `msmarco-dev/` | MS MARCO Passage Ranking Dev | 6,980 |
| `trec-dl-2019/` | TREC Deep Learning 2019 passage-ranking topics | 43 |
| `trec-dl-2020/` | TREC Deep Learning 2020 passage-ranking topics | 54 |
Each directory contains one TSV file per perturbation type:
`misspelling.tsv`, `reordering.tsv`, `synonym.tsv`, `paraphrase.tsv`,
`naturality.tsv`.
**Columns:** `qid`, `original_query`, `perturbed_query`, `perturbation`, `seed`
```text
qid original_query perturbed_query perturbation seed
1037798 who is robert gray who is robrt gray misspelling 42
```
- **Seeded perturbations** (`misspelling`, `reordering`, `synonym`,
`paraphrase`): five seeds — `1999`, `5`, `27`, `2016`, `2026` — stacked as
separate rows per `qid` within the same file.
- **`naturality`:** a deterministic style-normalization rule rather than a
randomized perturbation, so it produces identical output for every seed.
Represented as a single row per `qid` with `seed=deterministic` rather than
five duplicate rows.
`msmarco-dev/original_qids.txt` lists the 6,980 Dev query IDs covered by
these files (a subset of the full MS MARCO Dev qrel set).
Implemented in
[`robustness/query_variations/penha/`](https://github.com/kidist-amde/lost-in-decoding/tree/main/robustness/query_variations/penha)
in the code repository.
## `metadata/`
- `generation_config.json`: perturbation types, split sizes, file format, and provenance.
- `seeds.json`: seed values and which perturbations use them.
## `data_download/`
The original MS MARCO passage collection, MS MARCO Dev queries/qrels, and
TREC-DL 2019/2020 topics/qrels are **not** redistributed in this repository.
Scripts to fetch them directly from their official sources (Microsoft,
NIST) are provided at the repository root:
```bash
git clone https://huggingface.co/kiyam/lost-in-decoding-pag
cd lost-in-decoding-pag
bash data_download/prepare_evaluation_data.sh
```
See [`../data_download/`](../data_download/) for individual per-dataset
scripts and checksums.
## Citation
If you use these artifacts, please cite:
```bibtex
@inproceedings{mekonnen2026lost,
title = {Lost in Decoding? Reproducing and Stress-Testing the Look-Ahead Prior in Generative Retrieval},
author = {Mekonnen, Kidist Amde and Li, Yongkang and Tang, Yubao and Lupart, Simon and de Rijke, Maarten},
booktitle = {Proceedings of the 49th International ACM SIGIR Conference on Research and Development in Information Retrieval},
pages = {2994--3005},
year = {2026},
doi = {10.1145/3805712.3808567}
}
```