title: ACL Anthology Search API
emoji: π
sdk: docker
app_port: 7860
pinned: false
ACL Anthology Search API
FastAPI service providing keyword (SQLite FTS5) and embedding-similarity (FAISS)
search over the full ACL Anthology corpus (~128k papers). Runs entirely on free
infrastructure (HF Spaces + HF Hub dataset repo + GitHub Actions). See
docs/superpowers/specs/2026-07-03-acl-anthology-search-api-design.md for the
full design.
API reference
All endpoints return JSON. The service answers 503 on every route except
/health until the index/DB snapshot has finished loading on startup.
GET /health
Liveness + readiness probe.
Query: none.
Response: 200
{
"status": "ok",
"index_size": 128618,
"last_synced_at": "2026-07-04T15:06:00.654970+00:00"
}
Returns 503 when the snapshot is not yet loaded. last_synced_at is null
until the first sync has run.
GET /search/keyword?q=<query>&limit=<n>
Keyword search over paper titles and abstracts via SQLite FTS5 (Porter-stemmed, ranked by BM25).
Query parameters:
| Param | Type | Default | Notes |
|---|---|---|---|
q |
str | β | Required. Raw FTS5 query (see below). |
limit |
int | 10 |
Max number of results. |
FTS5 query format: space-separated terms are ANDed. Supports "exact phrase",
prefix*, term1 OR term2, term1 NOT term2, and parentheses. Examples:
q=transformer attention, q="neural machine", q=translat*, q=(attention OR transformer) NOT survey.
Response: 200 β array of PaperResult:
[
{
"id": "1993.tc-1.6",
"title": "Terminology and the computer - attention shifts to the micro",
"abstract": "β¦",
"authors": "Warren Weaver",
"venue": "TC",
"year": 1993,
"url": "https://aclanthology.org/1993.tc-1.6/",
"bibtex": "@inproceedings{weaver-1993-terminology,\n title = \"Terminology and the computer - attention shifts to the micro\",\n author = \"Weaver, Warren\",\n booktitle = \"Proceedings of the β¦\",\n year = \"1993\",\n url = \"https://aclanthology.org/1993.tc-1.6/\"\n}",
"pdf_url": "https://aclanthology.org/1993.tc-1.6.pdf"
}
]
year is null when the anthology has no year for a paper. bibtex is the
canonical ACL Anthology BibTeX entry for the paper (empty when the anthology
has no citable entry, e.g. some frontmatter). pdf_url is the direct link to
the paper's PDF (empty when the anthology has no PDF reference). Inactive
(retracted) papers are excluded.
GET /search/similarity?q=<query>&k=<n>
Embedding-similarity search. The query is embedded via the external embedding API, then matched against the FAISS index (cosine similarity over normalized vectors). Results are sorted by descending similarity.
Query parameters:
| Param | Type | Default | Notes |
|---|---|---|---|
q |
str | β | Required. Free-text query, no FTS5 syntax. |
k |
int | 10 |
Number of nearest neighbors to return. |
Response: 200 β array of SimilarityResult (a PaperResult plus a
similarity score):
[
{
"id": "N18-3011",
"title": "β¦",
"abstract": "β¦",
"authors": "β¦",
"venue": "β¦",
"year": 2018,
"url": "https://aclanthology.org/N18-3011/",
"bibtex": "@inproceedings{β¦}",
"pdf_url": "https://aclanthology.org/N18-3011.pdf",
"score": 0.8123
}
]
score is cosine similarity in [-1, 1] (higher is more similar).
GET /paper/{paper_id}
Metadata lookup by ACL Anthology paper id (e.g. N18-3011, 2023.acl-long.123).
Response: 200 β a single PaperResult (shape as in /search/keyword),
including the canonical bibtex entry for citing the paper and its pdf_url.
Returns 404 if the id is unknown or the paper is inactive.
Examples
# health
curl "http://localhost:8000/health"
# keyword search
curl "http://localhost:8000/search/keyword?q=transformer%20attention&limit=5"
# similarity search
curl "http://localhost:8000/search/similarity?q=attention%20mechanism&k=5"
# paper lookup
curl "http://localhost:8000/paper/1993.tc-1.6"
Configuration
Settings are read from the environment (and, for local dev, from a .env file
in the working directory). Required:
HF_REPO_IDβ dataset repo holdingindex.faiss/papers.db/state.jsonHF_TOKENβ token with read access to that dataset repoEMBEDDING_BASE_URLβ university embedding API base URLEMBEDDING_API_KEYβ university embedding API key
Optional:
data_dirβ local directory for the downloaded snapshot (default./data)
Local development
pip install -e ".[dev]"
python -m app # serves on http://0.0.0.0:8000 with reload
On startup the service downloads the snapshot from the HF Hub dataset repo into
data_dir and loads it into memory (~1.3 GB index + ~300 MB DB). The only
outbound call at request time is to the embedding API, to embed
similarity-search queries.
Required Space secrets
HF_REPO_IDβ dataset repo holdingindex.faiss/papers.db/state.jsonHF_TOKENβ token with read access to that dataset repoEMBEDDING_BASE_URLβ university embedding API base URLEMBEDDING_API_KEYβ university embedding API key
Required GitHub Actions secrets
HF_REPO_ID,HF_TOKEN,EMBEDDING_BASE_URL,EMBEDDING_API_KEYβ same as the Space secrets aboveHF_SPACE_IDβ e.g.your-org/acl-anthology-search-api, used to trigger a restart after sync