ResearchIT / docs /phases /PHASE8-Search-And-Recommendation-Design.md
siddhm11
deploy: search latency, metadata sidecar, durable user data
b861f87
|
Raw
History Blame Contribute Delete
8.06 kB

PHASE 8 β€” Search and Recommendation: the exact design

Status: Partly shipped Β· Updated: 2026-07-30

The definitive specification of both pipelines: every stage, every constant, where it lives, and whether it is live or pending. Latency figures are measured against the deployed Space, not estimated.

Legend: [LIVE] merged to main Β· [PENDING] designed, not built


1. Search

GET /search?q=… β†’ app/routers/search.py β†’ app/hybrid_search_svc.py:search()

query
  β”‚
  β”œβ”€[1]─► Groq rewrite            concurrent, non-blocking        [LIVE]
  β”œβ”€[2]─► BGE-M3 encode           in executor, max_length=512     [LIVE]
  β”‚
  β”œβ”€[3a]β–Ί Qdrant dense    limit=60, rescore=True, oversampling=1  [LIVE]
  └─[3b]β–Ί lexical         Zilliz sparse β†’ FTS5                    [PENDING]
            β”‚
          [4] RRF fusion, k=60                                    [LIVE]
            β”‚
          [5] cross-encoder, top 50, full abstracts               [LIVE / PENDING]
            β”‚
          [6] title-match + citation boost, top 50                [LIVE]
            β”‚
          [7] return top 10

Stage detail

# Stage Constants Measured
1 Groq rewrite groq_svc.rewrite, skipped if ≀2 words or _looks_academic 161–329 ms, overlapped
2 BGE-M3 encode max_length=512, LRU 128, run_in_executor 285–642 ms
3a Qdrant dense limit = 10 Γ— SEARCH_FETCH_K_MULTIPLIER(6) = 60 ~190 ms
3b Zilliz sparse same limit ~network
4 RRF SEARCH_RRF_K = 60, 1/(k+rank) summed over all lists ~0 ms
5 Cross-encoder SEARCH_RERANK_TOP_N = 50, sigmoid β†’ [0,1] 139–182 ms at n=10
6 Boosts exact 2.0 Β· substring 1.0 Β· coverage β‰₯0.8β†’1.0 / β‰₯0.5β†’0.5 Β· citation cap 0.2 ~1 ms
7 Return ARXIV_MAX_RESULTS = 10

Rules that must not change

  • RRF is correct for search. Many retrievers, one query β€” rank-based fusion needs no score calibration. Do not replace it with quota (that is the recommendation-side answer to a different problem).
  • Rerank window must exceed the result count. At SEARCH_RERANK_TOP_N = 10 the stage re-ordered exactly the set already being returned and could never promote a better paper from the retrieved pool. 50 of 60 is the current setting; below 10 it is worse than useless because it still costs CPU.
  • Rescore stays on. rescore=False is 615 ms vs 608 ms β€” no faster β€” and drops recall@10 from 100% to 57%. Binary codes alone cannot rank 1024-dim vectors.
  • Title boost uses the ORIGINAL query, never the rewrite. The user's literal text is what should match a title.
  • Stage timings no longer sum. groq_time_ms overlaps encode_time_ms; search_meta.groq_overlapped flags this.

Pending

  1. Lexical from FTS5, not Zilliz β€” same RRF input shape, no network, one fewer vendor, ~2 GB less storage. A/B against Zilliz before switching: BGE-M3 sparse weights are learned, BM25 is not.
  2. Full abstracts to the cross-encoder β€” 90.0% of stored abstracts are truncated at 500 chars while the retriever saw 1024. The precision stage currently has less information than the stage it refines. Biggest remaining search-quality item.

2. Recommendations

GET /api/recommendations β†’ app/routers/recommendations.py

Cascading tiers, first non-empty wins.

Tier 1  β‰₯5 saves   clustering + quota fusion      ← the actual product
Tier 2  β‰₯3 saves   EWMA long-term vector
Tier 3  β‰₯1 save    Qdrant Recommend (BEST_SCORE)
Tier 0  onboarded  trending by category

Tier 1 β€” the real pipeline

saved papers
  └─[1] fetch vectors            qdrant_svc.get_paper_vectors    1247 ms / 20  ← hot spot
  └─[2] Ward clustering          MIN_PAPERS_FOR_CLUSTERING = 5
  └─[3] Hungarian stabilise      against persisted medoids
  └─[4] quota allocation         total_slots=100, min_slots=3, by importance
  └─[5] per-cluster ANN          limit = quota Γ— _OVERSAMPLE(3), parallel
  └─[6] short-term supplement    _ST_SUPPLEMENT = 20
  └─[7] scoring                  heuristic (default) | LightGBM      [LIVE]
  └─[8] category suppression     β‰₯3 dismissals / 14 days, arXiv codes [LIVE]
  └─[9] MMR diversity            lambda_param = 0.6, top_k = 10
  └─[10] exploration injection   n_explore = 2  β†’ returns limit + 2

Scoring β€” why the heuristic is the default

RERANKER_MODE = "heuristic" (app/config.py).

Parsing reranker_v1.txt: features 20–30 have zero splits across all 141 trees β€” every EWMA similarity, both cluster features, the suppression and onboarding flags, all four interaction counts. A tree only reads features it splits on, so a user's entire profile provably cannot change the output. Not a statistical claim; structural.

candidate_num_cited_by additionally holds 65.2% of importance and is hardcoded to 0 at serving time.

heuristic_score() reads features 20–22:

has_ewma:  relevance = 0.40Β·lt_sim + 0.25Β·st_sim
otherwise: relevance = 0.65Β·qdrant_cosine

Demonstrated: opposite ewma_longterm_similarity produces rankings [0,1,2,3,4,5] vs [5,4,3,2,1,0] β€” a full reversal from the profile alone.

Set RERANKER_MODE=lightgbm to compare once real engagement data exists. /healthz/reranker reports model_loaded and scoring_with separately.

Rules that must not change

  • Quota, not RRF, for recommendations. Many queries (one per interest cluster), one user. RRF would let the dominant cluster win on rank alone and reintroduce interest collapse β€” the failure the whole product exists to prevent.
  • Hungarian matching stays. Without it a user's "NLP cluster" becomes cluster_7 after the next recluster and instrumentation loses continuity.
  • Suppression on arXiv codes, not primary_topic. The latter has ~15 coarse buckets; AI/ML alone is 20.2% of the corpus, so three dismissals used to suppress a fifth of everything.
  • MMR Ξ»=0.6 β€” below ~0.5 relevance degrades visibly; above ~0.8 the feed collapses toward one interest.

Pending

  1. Candidate vectors local. get_paper_vectors is 1,247 ms for 20; Tier 1 needs 100+. MMR and scoring both need real vectors, not ids. At int8, 1.9M vectors is 1.95 GB β€” same sidecar pattern as metadata. Biggest remaining recommendation latency item.
  2. Persistence to Turso. Profiles, clusters and interactions live in SQLite at /tmp and are destroyed on every rebuild. Gates everything below.
  3. Reranker retrain on real interactions, restricted to features that are non-zero at serving time. Requires 2.

3. Cold start

Tier Trigger Source
0 onboarded, 0 saves fetch_trending_by_categories
3 1 save Qdrant Recommend
2 3 saves EWMA long-term
1 5 saves full pipeline

Trending ranks by citations within a recency window measured back from the newest paper in the corpus, widening 24 β†’ 48 β†’ 96 months for thin categories, with publication date decoded from the arXiv id (update_date is the revision date, so 2017 classics looked new). TRENDING_RECENCY_MONTHS = 24.

Known gap: the onboarding wizard targets 5 seed saves and Tier 1 needs 5, but nothing tells the user that 5 is a threshold. Save 4 and you silently get Tier 3, the weakest path.


4. Build order

# Task Depends on Risk
1 Persistence β†’ Turso β€” none, additive
2 FTS5 in sidecar + A/B vs Zilliz β€” none until switched
3 Full abstracts (ingest + backfill) Phase 7 needs arXiv job
4 Local candidate vectors index build image size
5 Retire Zilliz 2 after A/B
6 Reranker retrain 1 + months of data β€”

Shipped already: event-loop fix, Groq overlap, RERANKER_MODE, rerank window, quantization search params, arXiv-code suppression, recency trending, metadata sidecar.