# Document Identifiers **Provenance:** created by the PAG authors (Zeng, Luo, & Zamani, 2024) and released in the [upstream PAG repository](https://github.com/HansiZeng/PAG). Re-hosted here unmodified. See [`../LICENSE.md`](../LICENSE.md) for licensing notes on these third-party files. Both files are JSON objects mapping `docid (str) -> list[int]`, where the list is the sequence of token IDs that make up that document's generative identifier under the PAG tokenizer/vocabulary. ## `rq_docids.json` Sequential document identifiers produced by PAG's residual-quantization (AQ) stage. Corresponds to `aq_smtid/docid_to_tokenids.json` in the upstream release. Used by the sequential constrained-decoding stage of the PAG pipeline (Stage 2). ## `set_docids.json` Set-based (bag-of-words) document identifiers produced by PAG's lexical/SPLADE planning stage. Corresponds to `top_bow/docid_to_tokenids.json` in the upstream release. Used by the lexical planning stage of the PAG pipeline (Stage 1). ## Usage See the two-stage constrained-decoding pipeline in the [Lost-in-Decoding repository](https://github.com/kidist-amde/lost-in-decoding) (`t5_pretrainer/evaluate.py`) for how these identifier files are consumed during retrieval. ## Citation If you use these identifiers, please cite the original PAG paper: ```bibtex @inproceedings{zeng2024planning, title = {Planning Ahead in Generative Retrieval: Guiding Autoregressive Generation through Simultaneous Decoding}, author = {Zeng, Hansi and Luo, Chen and Zamani, Hamed}, booktitle = {Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval}, pages = {469--480}, year = {2024}, doi = {10.1145/3626772.3657746} } ```