Spaces:
Running
Running
| # PHASE 8 β Search and Recommendation: the exact design | |
| **Status:** Partly shipped Β· **Updated:** 2026-07-30 | |
| The definitive specification of both pipelines: every stage, every constant, | |
| where it lives, and whether it is live or pending. Latency figures are measured | |
| against the deployed Space, not estimated. | |
| Legend: **[LIVE]** merged to main Β· **[PENDING]** designed, not built | |
| --- | |
| ## 1. Search | |
| `GET /search?q=β¦` β `app/routers/search.py` β `app/hybrid_search_svc.py:search()` | |
| ``` | |
| query | |
| β | |
| ββ[1]ββΊ Groq rewrite concurrent, non-blocking [LIVE] | |
| ββ[2]ββΊ BGE-M3 encode in executor, max_length=512 [LIVE] | |
| β | |
| ββ[3a]βΊ Qdrant dense limit=60, rescore=True, oversampling=1 [LIVE] | |
| ββ[3b]βΊ lexical Zilliz sparse β FTS5 [PENDING] | |
| β | |
| [4] RRF fusion, k=60 [LIVE] | |
| β | |
| [5] cross-encoder, top 50, full abstracts [LIVE / PENDING] | |
| β | |
| [6] title-match + citation boost, top 50 [LIVE] | |
| β | |
| [7] return top 10 | |
| ``` | |
| ### Stage detail | |
| | # | Stage | Constants | Measured | | |
| |---|---|---|---| | |
| | 1 | Groq rewrite | `groq_svc.rewrite`, skipped if β€2 words or `_looks_academic` | 161β329 ms, **overlapped** | | |
| | 2 | BGE-M3 encode | `max_length=512`, LRU 128, `run_in_executor` | 285β642 ms | | |
| | 3a | Qdrant dense | `limit = 10 Γ SEARCH_FETCH_K_MULTIPLIER(6) = 60` | **~190 ms** | | |
| | 3b | Zilliz sparse | same limit | ~network | | |
| | 4 | RRF | `SEARCH_RRF_K = 60`, `1/(k+rank)` summed over all lists | ~0 ms | | |
| | 5 | Cross-encoder | `SEARCH_RERANK_TOP_N = 50`, sigmoid β [0,1] | 139β182 ms at n=10 | | |
| | 6 | Boosts | exact 2.0 Β· substring 1.0 Β· coverage β₯0.8β1.0 / β₯0.5β0.5 Β· citation cap 0.2 | ~1 ms | | |
| | 7 | Return | `ARXIV_MAX_RESULTS = 10` | | | |
| ### Rules that must not change | |
| - **RRF is correct for search.** Many retrievers, one query β rank-based fusion | |
| needs no score calibration. Do not replace it with quota (that is the | |
| recommendation-side answer to a different problem). | |
| - **Rerank window must exceed the result count.** At `SEARCH_RERANK_TOP_N = 10` | |
| the stage re-ordered exactly the set already being returned and could never | |
| promote a better paper from the retrieved pool. 50 of 60 is the current | |
| setting; below 10 it is worse than useless because it still costs CPU. | |
| - **Rescore stays on.** `rescore=False` is 615 ms vs 608 ms β no faster β and | |
| drops recall@10 from 100% to 57%. Binary codes alone cannot rank 1024-dim | |
| vectors. | |
| - **Title boost uses the ORIGINAL query, never the rewrite.** The user's literal | |
| text is what should match a title. | |
| - **Stage timings no longer sum.** `groq_time_ms` overlaps `encode_time_ms`; | |
| `search_meta.groq_overlapped` flags this. | |
| ### Pending | |
| 1. **Lexical from FTS5, not Zilliz** β same RRF input shape, no network, one | |
| fewer vendor, ~2 GB less storage. A/B against Zilliz before switching: | |
| BGE-M3 sparse weights are learned, BM25 is not. | |
| 2. **Full abstracts to the cross-encoder** β 90.0% of stored abstracts are | |
| truncated at 500 chars while the retriever saw 1024. The precision stage | |
| currently has less information than the stage it refines. Biggest remaining | |
| search-quality item. | |
| --- | |
| ## 2. Recommendations | |
| `GET /api/recommendations` β `app/routers/recommendations.py` | |
| Cascading tiers, first non-empty wins. | |
| ``` | |
| Tier 1 β₯5 saves clustering + quota fusion β the actual product | |
| Tier 2 β₯3 saves EWMA long-term vector | |
| Tier 3 β₯1 save Qdrant Recommend (BEST_SCORE) | |
| Tier 0 onboarded trending by category | |
| ``` | |
| ### Tier 1 β the real pipeline | |
| ``` | |
| saved papers | |
| ββ[1] fetch vectors qdrant_svc.get_paper_vectors 1247 ms / 20 β hot spot | |
| ββ[2] Ward clustering MIN_PAPERS_FOR_CLUSTERING = 5 | |
| ββ[3] Hungarian stabilise against persisted medoids | |
| ββ[4] quota allocation total_slots=100, min_slots=3, by importance | |
| ββ[5] per-cluster ANN limit = quota Γ _OVERSAMPLE(3), parallel | |
| ββ[6] short-term supplement _ST_SUPPLEMENT = 20 | |
| ββ[7] scoring heuristic (default) | LightGBM [LIVE] | |
| ββ[8] category suppression β₯3 dismissals / 14 days, arXiv codes [LIVE] | |
| ββ[9] MMR diversity lambda_param = 0.6, top_k = 10 | |
| ββ[10] exploration injection n_explore = 2 β returns limit + 2 | |
| ``` | |
| ### Scoring β why the heuristic is the default | |
| `RERANKER_MODE = "heuristic"` (`app/config.py`). | |
| Parsing `reranker_v1.txt`: features 20β30 have **zero splits across all 141 | |
| trees** β every EWMA similarity, both cluster features, the suppression and | |
| onboarding flags, all four interaction counts. A tree only reads features it | |
| splits on, so a user's entire profile provably cannot change the output. Not a | |
| statistical claim; structural. | |
| `candidate_num_cited_by` additionally holds 65.2% of importance and is | |
| hardcoded to 0 at serving time. | |
| `heuristic_score()` reads features 20β22: | |
| ``` | |
| has_ewma: relevance = 0.40Β·lt_sim + 0.25Β·st_sim | |
| otherwise: relevance = 0.65Β·qdrant_cosine | |
| ``` | |
| Demonstrated: opposite `ewma_longterm_similarity` produces rankings | |
| `[0,1,2,3,4,5]` vs `[5,4,3,2,1,0]` β a full reversal from the profile alone. | |
| Set `RERANKER_MODE=lightgbm` to compare once real engagement data exists. | |
| `/healthz/reranker` reports `model_loaded` and `scoring_with` separately. | |
| ### Rules that must not change | |
| - **Quota, not RRF, for recommendations.** Many queries (one per interest | |
| cluster), one user. RRF would let the dominant cluster win on rank alone and | |
| reintroduce interest collapse β the failure the whole product exists to | |
| prevent. | |
| - **Hungarian matching stays.** Without it a user's "NLP cluster" becomes | |
| `cluster_7` after the next recluster and instrumentation loses continuity. | |
| - **Suppression on arXiv codes, not `primary_topic`.** The latter has ~15 coarse | |
| buckets; `AI/ML` alone is 20.2% of the corpus, so three dismissals used to | |
| suppress a fifth of everything. | |
| - **MMR Ξ»=0.6** β below ~0.5 relevance degrades visibly; above ~0.8 the feed | |
| collapses toward one interest. | |
| ### Pending | |
| 1. **Candidate vectors local.** `get_paper_vectors` is 1,247 ms for 20; Tier 1 | |
| needs 100+. MMR and scoring both need real vectors, not ids. At int8, 1.9M | |
| vectors is 1.95 GB β same sidecar pattern as metadata. Biggest remaining | |
| recommendation latency item. | |
| 2. **Persistence to Turso.** Profiles, clusters and interactions live in SQLite | |
| at `/tmp` and are destroyed on every rebuild. Gates everything below. | |
| 3. **Reranker retrain** on real interactions, restricted to features that are | |
| non-zero at serving time. Requires 2. | |
| --- | |
| ## 3. Cold start | |
| | Tier | Trigger | Source | | |
| |---|---|---| | |
| | 0 | onboarded, 0 saves | `fetch_trending_by_categories` | | |
| | 3 | 1 save | Qdrant Recommend | | |
| | 2 | 3 saves | EWMA long-term | | |
| | 1 | 5 saves | full pipeline | | |
| Trending ranks by citations **within a recency window** measured back from the | |
| newest paper in the corpus, widening 24 β 48 β 96 months for thin categories, | |
| with publication date decoded from the arXiv id (`update_date` is the revision | |
| date, so 2017 classics looked new). `TRENDING_RECENCY_MONTHS = 24`. | |
| **Known gap:** the onboarding wizard targets 5 seed saves and Tier 1 needs 5, | |
| but nothing tells the user that 5 is a threshold. Save 4 and you silently get | |
| Tier 3, the weakest path. | |
| --- | |
| ## 4. Build order | |
| | # | Task | Depends on | Risk | | |
| |---|---|---|---| | |
| | 1 | Persistence β Turso | β | none, additive | | |
| | 2 | FTS5 in sidecar + A/B vs Zilliz | β | none until switched | | |
| | 3 | Full abstracts (ingest + backfill) | Phase 7 | needs arXiv job | | |
| | 4 | Local candidate vectors | index build | image size | | |
| | 5 | Retire Zilliz | 2 | after A/B | | |
| | 6 | Reranker retrain | 1 + months of data | β | | |
| Shipped already: event-loop fix, Groq overlap, `RERANKER_MODE`, rerank window, | |
| quantization search params, arXiv-code suppression, recency trending, metadata | |
| sidecar. | |