kiyam's picture
Correct provenance, licensing, and citation details: identifiers/README.md
0aa8154 verified
|
Raw
History Blame Contribute Delete
1.77 kB

Document Identifiers

Provenance: created by the PAG authors (Zeng, Luo, & Zamani, 2024) and released in the upstream PAG repository. Re-hosted here unmodified. See ../LICENSE.md for licensing notes on these third-party files.

Both files are JSON objects mapping docid (str) -> list[int], where the list is the sequence of token IDs that make up that document's generative identifier under the PAG tokenizer/vocabulary.

rq_docids.json

Sequential document identifiers produced by PAG's residual-quantization (AQ) stage. Corresponds to aq_smtid/docid_to_tokenids.json in the upstream release. Used by the sequential constrained-decoding stage of the PAG pipeline (Stage 2).

set_docids.json

Set-based (bag-of-words) document identifiers produced by PAG's lexical/SPLADE planning stage. Corresponds to top_bow/docid_to_tokenids.json in the upstream release. Used by the lexical planning stage of the PAG pipeline (Stage 1).

Usage

See the two-stage constrained-decoding pipeline in the Lost-in-Decoding repository (t5_pretrainer/evaluate.py) for how these identifier files are consumed during retrieval.

Citation

If you use these identifiers, please cite the original PAG paper:

@inproceedings{zeng2024planning,
  title     = {Planning Ahead in Generative Retrieval: Guiding Autoregressive Generation through Simultaneous Decoding},
  author    = {Zeng, Hansi and Luo, Chen and Zamani, Hamed},
  booktitle = {Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval},
  pages     = {469--480},
  year      = {2024},
  doi       = {10.1145/3626772.3657746}
}