kiyam's picture
Correct provenance, licensing, and citation details: identifiers/README.md
0aa8154 verified
|
Raw
History Blame Contribute Delete
1.77 kB
# Document Identifiers
**Provenance:** created by the PAG authors (Zeng, Luo, & Zamani, 2024) and
released in the [upstream PAG repository](https://github.com/HansiZeng/PAG).
Re-hosted here unmodified. See [`../LICENSE.md`](../LICENSE.md) for licensing
notes on these third-party files.
Both files are JSON objects mapping `docid (str) -> list[int]`, where the list
is the sequence of token IDs that make up that document's generative
identifier under the PAG tokenizer/vocabulary.
## `rq_docids.json`
Sequential document identifiers produced by PAG's residual-quantization (AQ)
stage. Corresponds to `aq_smtid/docid_to_tokenids.json` in the upstream
release. Used by the sequential constrained-decoding stage of the PAG
pipeline (Stage 2).
## `set_docids.json`
Set-based (bag-of-words) document identifiers produced by PAG's lexical/SPLADE
planning stage. Corresponds to `top_bow/docid_to_tokenids.json` in the
upstream release. Used by the lexical planning stage of the PAG pipeline
(Stage 1).
## Usage
See the two-stage constrained-decoding pipeline in the
[Lost-in-Decoding repository](https://github.com/kidist-amde/lost-in-decoding)
(`t5_pretrainer/evaluate.py`) for how these identifier files are consumed
during retrieval.
## Citation
If you use these identifiers, please cite the original PAG paper:
```bibtex
@inproceedings{zeng2024planning,
title = {Planning Ahead in Generative Retrieval: Guiding Autoregressive Generation through Simultaneous Decoding},
author = {Zeng, Hansi and Luo, Chen and Zamani, Hamed},
booktitle = {Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval},
pages = {469--480},
year = {2024},
doi = {10.1145/3626772.3657746}
}
```