ResearchIT / docs /phases /PHASE8-Search-And-Recommendation-Design.md
siddhm11
deploy: search latency, metadata sidecar, durable user data
b861f87
|
Raw
History Blame Contribute Delete
8.06 kB
# PHASE 8 β€” Search and Recommendation: the exact design
**Status:** Partly shipped Β· **Updated:** 2026-07-30
The definitive specification of both pipelines: every stage, every constant,
where it lives, and whether it is live or pending. Latency figures are measured
against the deployed Space, not estimated.
Legend: **[LIVE]** merged to main Β· **[PENDING]** designed, not built
---
## 1. Search
`GET /search?q=…` β†’ `app/routers/search.py` β†’ `app/hybrid_search_svc.py:search()`
```
query
β”‚
β”œβ”€[1]─► Groq rewrite concurrent, non-blocking [LIVE]
β”œβ”€[2]─► BGE-M3 encode in executor, max_length=512 [LIVE]
β”‚
β”œβ”€[3a]β–Ί Qdrant dense limit=60, rescore=True, oversampling=1 [LIVE]
└─[3b]β–Ί lexical Zilliz sparse β†’ FTS5 [PENDING]
β”‚
[4] RRF fusion, k=60 [LIVE]
β”‚
[5] cross-encoder, top 50, full abstracts [LIVE / PENDING]
β”‚
[6] title-match + citation boost, top 50 [LIVE]
β”‚
[7] return top 10
```
### Stage detail
| # | Stage | Constants | Measured |
|---|---|---|---|
| 1 | Groq rewrite | `groq_svc.rewrite`, skipped if ≀2 words or `_looks_academic` | 161–329 ms, **overlapped** |
| 2 | BGE-M3 encode | `max_length=512`, LRU 128, `run_in_executor` | 285–642 ms |
| 3a | Qdrant dense | `limit = 10 Γ— SEARCH_FETCH_K_MULTIPLIER(6) = 60` | **~190 ms** |
| 3b | Zilliz sparse | same limit | ~network |
| 4 | RRF | `SEARCH_RRF_K = 60`, `1/(k+rank)` summed over all lists | ~0 ms |
| 5 | Cross-encoder | `SEARCH_RERANK_TOP_N = 50`, sigmoid β†’ [0,1] | 139–182 ms at n=10 |
| 6 | Boosts | exact 2.0 Β· substring 1.0 Β· coverage β‰₯0.8β†’1.0 / β‰₯0.5β†’0.5 Β· citation cap 0.2 | ~1 ms |
| 7 | Return | `ARXIV_MAX_RESULTS = 10` | |
### Rules that must not change
- **RRF is correct for search.** Many retrievers, one query β€” rank-based fusion
needs no score calibration. Do not replace it with quota (that is the
recommendation-side answer to a different problem).
- **Rerank window must exceed the result count.** At `SEARCH_RERANK_TOP_N = 10`
the stage re-ordered exactly the set already being returned and could never
promote a better paper from the retrieved pool. 50 of 60 is the current
setting; below 10 it is worse than useless because it still costs CPU.
- **Rescore stays on.** `rescore=False` is 615 ms vs 608 ms β€” no faster β€” and
drops recall@10 from 100% to 57%. Binary codes alone cannot rank 1024-dim
vectors.
- **Title boost uses the ORIGINAL query, never the rewrite.** The user's literal
text is what should match a title.
- **Stage timings no longer sum.** `groq_time_ms` overlaps `encode_time_ms`;
`search_meta.groq_overlapped` flags this.
### Pending
1. **Lexical from FTS5, not Zilliz** β€” same RRF input shape, no network, one
fewer vendor, ~2 GB less storage. A/B against Zilliz before switching:
BGE-M3 sparse weights are learned, BM25 is not.
2. **Full abstracts to the cross-encoder** β€” 90.0% of stored abstracts are
truncated at 500 chars while the retriever saw 1024. The precision stage
currently has less information than the stage it refines. Biggest remaining
search-quality item.
---
## 2. Recommendations
`GET /api/recommendations` β†’ `app/routers/recommendations.py`
Cascading tiers, first non-empty wins.
```
Tier 1 β‰₯5 saves clustering + quota fusion ← the actual product
Tier 2 β‰₯3 saves EWMA long-term vector
Tier 3 β‰₯1 save Qdrant Recommend (BEST_SCORE)
Tier 0 onboarded trending by category
```
### Tier 1 β€” the real pipeline
```
saved papers
└─[1] fetch vectors qdrant_svc.get_paper_vectors 1247 ms / 20 ← hot spot
└─[2] Ward clustering MIN_PAPERS_FOR_CLUSTERING = 5
└─[3] Hungarian stabilise against persisted medoids
└─[4] quota allocation total_slots=100, min_slots=3, by importance
└─[5] per-cluster ANN limit = quota Γ— _OVERSAMPLE(3), parallel
└─[6] short-term supplement _ST_SUPPLEMENT = 20
└─[7] scoring heuristic (default) | LightGBM [LIVE]
└─[8] category suppression β‰₯3 dismissals / 14 days, arXiv codes [LIVE]
└─[9] MMR diversity lambda_param = 0.6, top_k = 10
└─[10] exploration injection n_explore = 2 β†’ returns limit + 2
```
### Scoring β€” why the heuristic is the default
`RERANKER_MODE = "heuristic"` (`app/config.py`).
Parsing `reranker_v1.txt`: features 20–30 have **zero splits across all 141
trees** β€” every EWMA similarity, both cluster features, the suppression and
onboarding flags, all four interaction counts. A tree only reads features it
splits on, so a user's entire profile provably cannot change the output. Not a
statistical claim; structural.
`candidate_num_cited_by` additionally holds 65.2% of importance and is
hardcoded to 0 at serving time.
`heuristic_score()` reads features 20–22:
```
has_ewma: relevance = 0.40Β·lt_sim + 0.25Β·st_sim
otherwise: relevance = 0.65Β·qdrant_cosine
```
Demonstrated: opposite `ewma_longterm_similarity` produces rankings
`[0,1,2,3,4,5]` vs `[5,4,3,2,1,0]` β€” a full reversal from the profile alone.
Set `RERANKER_MODE=lightgbm` to compare once real engagement data exists.
`/healthz/reranker` reports `model_loaded` and `scoring_with` separately.
### Rules that must not change
- **Quota, not RRF, for recommendations.** Many queries (one per interest
cluster), one user. RRF would let the dominant cluster win on rank alone and
reintroduce interest collapse β€” the failure the whole product exists to
prevent.
- **Hungarian matching stays.** Without it a user's "NLP cluster" becomes
`cluster_7` after the next recluster and instrumentation loses continuity.
- **Suppression on arXiv codes, not `primary_topic`.** The latter has ~15 coarse
buckets; `AI/ML` alone is 20.2% of the corpus, so three dismissals used to
suppress a fifth of everything.
- **MMR Ξ»=0.6** β€” below ~0.5 relevance degrades visibly; above ~0.8 the feed
collapses toward one interest.
### Pending
1. **Candidate vectors local.** `get_paper_vectors` is 1,247 ms for 20; Tier 1
needs 100+. MMR and scoring both need real vectors, not ids. At int8, 1.9M
vectors is 1.95 GB β€” same sidecar pattern as metadata. Biggest remaining
recommendation latency item.
2. **Persistence to Turso.** Profiles, clusters and interactions live in SQLite
at `/tmp` and are destroyed on every rebuild. Gates everything below.
3. **Reranker retrain** on real interactions, restricted to features that are
non-zero at serving time. Requires 2.
---
## 3. Cold start
| Tier | Trigger | Source |
|---|---|---|
| 0 | onboarded, 0 saves | `fetch_trending_by_categories` |
| 3 | 1 save | Qdrant Recommend |
| 2 | 3 saves | EWMA long-term |
| 1 | 5 saves | full pipeline |
Trending ranks by citations **within a recency window** measured back from the
newest paper in the corpus, widening 24 β†’ 48 β†’ 96 months for thin categories,
with publication date decoded from the arXiv id (`update_date` is the revision
date, so 2017 classics looked new). `TRENDING_RECENCY_MONTHS = 24`.
**Known gap:** the onboarding wizard targets 5 seed saves and Tier 1 needs 5,
but nothing tells the user that 5 is a threshold. Save 4 and you silently get
Tier 3, the weakest path.
---
## 4. Build order
| # | Task | Depends on | Risk |
|---|---|---|---|
| 1 | Persistence β†’ Turso | β€” | none, additive |
| 2 | FTS5 in sidecar + A/B vs Zilliz | β€” | none until switched |
| 3 | Full abstracts (ingest + backfill) | Phase 7 | needs arXiv job |
| 4 | Local candidate vectors | index build | image size |
| 5 | Retire Zilliz | 2 | after A/B |
| 6 | Reranker retrain | 1 + months of data | β€” |
Shipped already: event-loop fix, Groq overlap, `RERANKER_MODE`, rerank window,
quantization search params, arXiv-code suppression, recency trending, metadata
sidecar.