File size: 50,911 Bytes
65cee4e 95143a2 65cee4e 8ec920c 44e600b 05820e3 eca5609 d281212 84f9c55 20130d7 8c5326e 84f9c55 05820e3 65cee4e 44e600b 05820e3 84f9c55 10860d7 20130d7 8c5326e 95143a2 65cee4e 84f9c55 05820e3 65cee4e 95143a2 10860d7 95143a2 8ec920c 95143a2 05820e3 20130d7 84f9c55 95143a2 8c5326e 65cee4e d281212 84f9c55 d281212 95143a2 d281212 95143a2 8ec920c 5d9a6f1 65cee4e d281212 3f7febe eca5609 65cee4e 8ec920c 65cee4e 44e600b 65cee4e 84f9c55 65cee4e d281212 65cee4e 20130d7 65cee4e eca5609 65cee4e | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 390 391 392 393 394 395 396 397 398 399 400 401 402 403 404 405 406 407 408 409 410 411 412 413 414 415 416 417 418 419 420 421 422 423 424 425 426 427 428 429 430 431 432 433 434 435 436 437 438 439 440 441 442 443 444 445 446 447 448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 465 466 467 468 469 470 471 472 473 474 475 476 477 478 479 480 481 482 483 484 485 486 487 488 489 490 491 492 493 494 495 496 497 498 499 500 501 502 503 504 505 506 507 508 509 510 511 512 513 514 515 516 517 518 519 520 521 522 523 524 525 526 527 528 529 530 531 532 533 534 535 536 537 538 539 540 541 542 543 544 545 546 547 548 549 550 551 552 553 554 555 556 557 558 559 560 561 562 563 564 565 566 567 568 569 570 571 572 573 574 575 576 577 578 579 580 581 582 583 584 585 586 587 588 589 590 591 592 593 594 595 596 597 598 599 600 601 602 603 604 605 606 607 608 609 610 611 612 613 614 615 616 617 618 619 620 621 622 623 624 625 626 627 628 629 630 631 632 633 634 635 636 637 638 639 640 641 642 643 644 645 646 647 648 649 650 651 652 653 654 655 656 657 658 659 660 661 662 663 664 665 666 667 668 669 670 671 672 673 674 675 676 677 678 679 680 681 682 683 684 685 686 687 688 689 690 691 692 693 694 695 696 697 698 699 700 701 702 703 704 705 706 707 708 709 710 711 712 713 714 715 716 717 718 719 720 721 722 723 724 725 726 727 728 729 730 731 732 733 734 735 736 737 738 739 740 741 742 743 744 745 746 747 748 749 750 751 752 753 754 755 756 757 758 759 760 761 762 763 764 765 766 767 768 769 770 771 772 773 774 775 776 777 778 779 780 781 782 783 784 785 786 787 788 789 790 791 792 793 794 795 796 797 798 799 800 801 802 803 804 805 806 807 808 809 810 811 812 813 814 815 816 817 818 819 820 821 822 | # Prism β Project Evolution
> End-to-end record of what was broken at each stage, what was built to fix it, and what is planned next.
> Updated as the project evolves. Last updated: 2026-07-05 (Stage 18).
---
## Table of Contents
1. [Stage 0 β v1 Baseline](#stage-0--v1-baseline)
2. [Stage 1 β v2 Hybrid Retrieval Architecture](#stage-1--v2-hybrid-retrieval-architecture-2026-05-17)
3. [Stage 2 β Chain Scores + RAGAS Endpoint](#stage-2--chain-scores--ragas-endpoint-2026-05-23)
4. [Stage 3 β Groq Migration + Web Search](#stage-3--groq-migration--web-search-fixes-2026-05-24)
5. [Stage 4 β OOM Hell on Render](#stage-4--oom-hell-on-render-2026-05-24-four-sub-issues)
6. [Stage 5 β TinyBERT + RAGAS Removal + Benchmark JSON](#stage-5--tinybert--ragas-removal--benchmark-json-2026-05-30)
7. [Stage 6 β Multi-Workspace](#stage-6--multi-workspace-2026-06-early)
8. [Stage 7 β Singleton Cache + URL Guard](#stage-7--singleton-cache--url-guard-2026-06-14)
9. [Stage 8 β Eval Dashboard + Rigorous Metrics](#stage-8--eval-dashboard--rigorous-metrics-2026-06-17)
10. [Stage 9 β Multi-Query Retrieval](#stage-9--multi-query-retrieval-2026-06-19)
11. [Stage 10 β Contextual Retrieval (Eval)](#stage-10--contextual-retrieval-eval-2026-06-20)
14. [Stage 14 β Briefing Fix + HyDE Re-eval](#stage-14--briefing-fix--hyde-re-eval-2026-06-24)
12. [Stage 12 β HF Spaces Migration](#stage-12--hf-spaces-migration-2026-06-22)
12. [Stage 11 β Contextual Retrieval in Production + Dashboard Polish](#stage-11--contextual-retrieval-in-production--dashboard-polish-2026-06-20)
15. [Stage 15 β Semantic Chunking Ablation + Retrieval Stack Finalized](#stage-15--semantic-chunking-ablation--retrieval-stack-finalized-2026-06-26)
16. [Stage 16 β Citation Highlighting](#stage-16--citation-highlighting-2026-06-26)
17. [Stage 17 β Metadata Filtering](#stage-17--metadata-filtering-2026-06-27)
13. [Current State Snapshot](#current-state-snapshot)
13. [Roadmap β Retrieval & Answer Quality](#roadmap--retrieval--answer-quality)
14. [Roadmap β New Features](#roadmap--new-features)
---
## Stage 0 β v1 Baseline
### What existed
- Dense-only ChromaDB vector retrieval
- Single global document collection
- Basic chat with ConversationalRetrievalChain
- No evaluation framework
- No web search
- No logging
### What was wrong
| Problem | Impact |
|---------|--------|
| Dense-only retrieval | Misses exact keyword matches β regulatory text has section numbers, policy codes, specific terms that semantic search fails on |
| No evaluation | No way to measure if answers were correct or grounded |
| No web search | Static corpus only β cannot answer questions about current stock prices, recent news |
| Single collection | No topic isolation β all documents mixed in one retrieval pool |
| No logging | Impossible to debug production failures |
**This was the starting point. No fixes yet.**
---
## Stage 1 β v2 Hybrid Retrieval Architecture (2026-05-17)
### What was wrong before building
- `retriever.py` was dense-only ChromaDB β v2 was documented but not implemented
- `ragas_eval.py` missing entirely; RAGAS eval endpoint not wired
- No BM25, no reranker, no score visibility
### What we built
| File | What changed |
|------|-------------|
| `server/bm25_index.py` | BM25Okapi singleton; module-level (not `app.state`) so importable anywhere; rebuilt on startup + after upload |
| `server/reranker.py` | CrossEncoder singleton; pre-loaded at startup to avoid cold-start latency on first query |
| `server/retriever.py` | Full rewrite as `HybridRetriever(BaseRetriever)` β RRF fusion of dense (weight 0.7) + sparse (weight 0.3) |
| `server/main.py` | BM25 build + reranker load wired into lifespan startup |
| `server/routes/upload.py` | BM25 rebuild triggered after each upload |
### Key design decisions
- **`HybridRetriever` as `BaseRetriever` subclass** β `ConversationalRetrievalChain` expects a `BaseRetriever`; subclassing means `chain.py` needs zero changes
- **BM25 as module-level singleton** β avoids threading state through lifespan β constructor; `get_index()` importable anywhere
- **Reranker pre-loaded at startup** β ~0.5s load from disk cache; better to pay at startup than add latency to first user query
- **MiniLM-L-6-v2** chosen as reranker (~85MB) β best ranking quality available at the time
### What was still missing
- Chain score extraction (similarity/BM25/RRF/rerank not returned in API response)
- RAGAS eval endpoint
- Groq LLM (still on Euron)
---
## Stage 2 β Chain Scores + RAGAS Endpoint (2026-05-23)
### What was wrong
- API response had no retrieval scores β no way to show per-source similarity/BM25/RRF/rerank scores
- RAGAS eval endpoint not wired; `ragas_eval.py` missing
- `data/ground_truth/eval_pairs.json` had only keyword hints, no `ground_truth` answers β `context_precision` and `context_recall` always returned null
### What we built
| File | What changed |
|------|-------------|
| `server/chain.py` | Score extraction β similarity/bm25/rrf/rerank scores passed through to API response per source |
| `server/routes/chat.py` | Added `retrieval_method` field; stores contexts in `eval_log` for downstream RAGAS eval |
| `server/config.yaml` | Hybrid retrieval params: `dense_weight`, `sparse_weight`, `retrieve_k`, `rerank_k` |
| `requirements.txt` | Added `rank_bm25`, `sentence-transformers`, `ragas`, `datasets` |
| `server/eval/ragas_eval.py` | RAGAS faithfulness + answer_relevancy via `LangchainLLMWrapper` |
| `server/routes/eval.py` | `POST /api/eval/ragas` endpoint wired |
| `server/main.py` | `/health` endpoint added |
### What was still broken
- RAGAS not installed in venv (added to requirements.txt; installs at Docker build only)
- `eval_pairs.json` still had no ground_truth β 2 of 4 RAGAS metrics null
- LLM still on Euron gpt-4.1-mini
---
## Stage 3 β Groq Migration + Web Search Fixes (2026-05-24)
### What was wrong
| Problem | Root cause |
|---------|-----------|
| LLM on Euron (gpt-4.1-mini) | Closed model, slower, weaker interview story vs open-weight |
| Web search silently broken | `tavily-python` in requirements.txt but never pip-installed |
| Tavily returned shallow results | `search_depth="basic"` β not enough content from financial sites |
| Web context never reached LLM | `ConversationalRetrievalChain`'s condensation step rewrote the question and stripped prepended Tavily context before LLM ever saw it |
| Follow-up web queries returned garbage | Raw follow-up ("Is the price level good?") sent to Tavily with no chat history context |
| No request logging | Production failures undebuggable |
### What we built
| Component | Change |
|-----------|--------|
| LLM | Migrated Euron β Groq `llama-3.3-70b-versatile` via `langchain-groq`. Euron kept for embeddings (Groq has no embeddings endpoint) |
| `server/chain.py` | `run_query_with_web()` β bypasses chain condensation; direct LLM call with RAG + Tavily context + memory |
| `server/chain.py` | `condense_question()` β rewrites follow-up queries using chat history before Tavily search |
| Tavily | `search_depth="advanced"`, `max_results=3` (2Γ credits but richer content) |
| `server/main.py` | Request logging middleware β logs `METHOD /path STATUS Xms` per request |
| `server/utils.py` | Centralised logging β root logger + `logs/finrag.log` (5MBΓ3 rotation), noisy libs silenced |
| `frontend/.../MessageBubble.jsx` | `WebSourcesList` component β Tavily URLs as clickable green pill links |
| `.env.example` | Fixed β real keys had been committed; replaced with placeholders |
### Key discoveries
- `ConversationalRetrievalChain` condensation = silent context killer for web queries. Only fix: bypass the chain entirely for web path.
- Memory's `output_key="answer"` β `save_context` must use `{"answer": answer}` not `{"output": answer}` or KeyError.
---
## Stage 4 β OOM Hell on Render (2026-05-24, four sub-issues)
Render free tier: 512MB RAM. This stage was four separate OOM root causes discovered in sequence.
---
### 4a β CUDA torch OOM (startup crash)
**Problem:** `sentence-transformers` pulled CUDA torch (~2GB) by default. OOM before uvicorn bound to port β Render showed "No open ports detected" timeout. Zero server stdout β invisible failure.
**Diagnosis clue:** Build log showed `cuda-toolkit-13.0.2`, `nvidia-cublas` being installed. Port scan timeout = uvicorn crash at import time (not lifespan β lifespan runs *after* port bind).
**Fix:**
```dockerfile
# Install CPU-only torch BEFORE requirements.txt
RUN pip install torch --index-url https://download.pytorch.org/whl/cpu
RUN pip install -r requirements.txt
ENV HF_HUB_OFFLINE=1
ENV TRANSFORMERS_OFFLINE=1
```
---
### 4b β ragas startup ImportError
**Problem:** `ragas 0.4.3` imports `langchain_community.chat_models.vertexai` at package `__init__` level. That module was removed in `langchain-community 0.4.x`. Crash propagated: `eval.py` β `ragas_eval.py` β `ragas.__init__` β ImportError before uvicorn bound port.
**Fix:** All ragas imports moved inside `run_ragas_eval()` function body (lazy import).
---
### 4c β CrossEncoder OOM during chat
**Problem:** `CrossEncoder.predict(20 pairs)` = BERT forward pass on 20 pairs β ~200β400MB spike on top of base ~250MB β OOM on first web query.
**Fix:**
- `config.yaml`: `retrieve_k: 20 β 10`
- `reranker.py`: `model.predict(pairs, batch_size=4)` β limits how many pairs processed at once
---
### 4d β Web search OOM (post-GC headroom)
**Problem:** After first query, Python retained chain/LLM objects at ~490MB. Web query added ~9KB Tavily content + `condense_question` LLM call + `run_query_with_web` LLM call β OOM.
**Fix:**
- Tavily content truncated to 800 chars per result (was up to ~3000)
- `max_results`: 3 β 2
- `gc.collect()` after each chat request in `routes/chat.py`
---
## Stage 5 β TinyBERT + RAGAS Removal + Benchmark JSON (2026-05-30)
### What was wrong
| Problem | Root cause |
|---------|-----------|
| MiniLM-L-6-v2 (~85MB) still OOMing on web queries | Too close to 512MB ceiling even after GC |
| Live RAGAS eval always 500 on Render | `nest_asyncio.apply()` (called at ragas import time) cannot patch `uvloop` β the event loop uvicorn uses on Linux. `ValueError: Can't patch loop of type uvloop.Loop`. Permanently unfixable without replacing uvicorn's event loop. |
| Groq 70B exhausted 100k daily tokens in one RAGAS run | RAGAS makes ~10 LLM calls per sample for statement decomposition. 10 samples Γ 10 calls = 100k tokens gone. |
| context_precision and context_recall always null | `eval_pairs.json` had no `ground_truth` answers β only keyword hints |
### What we built
| Component | Change |
|-----------|--------|
| `server/reranker.py` | Switched to `cross-encoder/ms-marco-TinyBERT-L-2-v2` (~17MB vs 85MB). Saves 68MB permanently. |
| `server/routes/eval.py` | Removed `POST /api/eval/ragas` endpoint |
| `requirements.txt` | Removed `ragas` |
| `scripts/run_ragas_local.py` | Local RAGAS runner: ingest corpus β generate answers β run eval β write JSON. Uses `llama-3.1-8b-instant` as judge (500k TPD vs 70B's 100k TPD) |
| `frontend/src/data/ragas_benchmark.json` | Static scores β Vercel builds dashboard from file |
| `frontend/.../EvalPanel.jsx` | Replaced live run button with static benchmark panel |
| `data/ground_truth/eval_pairs.json` | Added `ground_truth` field to all 20 pairs β unlocked `context_precision` + `context_recall` |
| UI | i-button tooltips on faithfulness badge + all 4 RAGAS metric cards |
### Real scores committed
```
faithfulness: 1.0 (note: likely inflated β see below)
answer_relevancy: 0.90
context_precision: TBD (pending fresh run)
context_recall: TBD (pending fresh run)
```
### Key discoveries
- `results["metric_name"]` returns `None` in ragas 0.2.x β must use `results.to_pandas()["metric_name"].mean()`
- TinyBERT loads with harmless `UNEXPECTED key bert.embeddings.position_ids` warning
- **Faithfulness 1.0 is likely inflated** β eval queries were designed alongside the corpus, and 8B judge is lenient. Scores are directional, not absolute. Run on held-out queries for honest numbers.
---
## Stage 5.5 β Rebranding: FinRAG β Prism (2026-06-13)
### What changed
The project was originally named **FinRAG** β a fintech-specific RAG demo. As the architecture matured (multi-workspace, URL ingestion, domain-agnostic retrieval), it became clear the tool was no longer fintech-specific. Any corpus β legal, HR, medical, research β could be loaded and queried.
**Decision:** Rebrand to **Prism**. Name reflects the core idea: feed any document set in, get clear structured answers out. One engine, any domain.
| Before | After |
|--------|-------|
| FinRAG | Prism |
| Fintech-specific framing | Domain-agnostic positioning |
| `finrag-v2.onrender.com` | `prism.onrender.com` |
| README pitched at fintech analysts | README pitched at any knowledge-worker |
### What stayed the same
All retrieval architecture, eval framework, and deployment stack unchanged. Rebrand is naming and framing only β the engine is identical.
### What was wrong with the old name
- "FinRAG" implied fintech-only β narrowed the demo audience
- Interviewers at non-fintech MNCs (Adobe, Atlassian, Intuit) would dismiss it as domain-locked
- The actual retrieval engine is domain-agnostic β the name should match
---
## Stage 6 β Multi-Workspace (2026-06 early)
### What was wrong
- Single ChromaDB collection β no isolation between document sets
- Switching topics meant re-ingesting and overwriting previous docs
- `list_collections()` broke on chromadb β₯0.5.4 (returns `list[str]`, not `list[Collection]`)
- Non-web chat path used stale global chain's `source_documents` instead of workspace-specific retriever β wrong docs shown after workspace switch
### What we built
| File | Change |
|------|--------|
| `server/routes/workspaces.py` | Workspace CRUD β one ChromaDB collection per workspace |
| `server/routes/chat.py` | Always resolves workspace-specific retriever before branching on `web_search` |
| `server/routes/workspaces.py` | `list_collections()` normalised with `isinstance` check β works on chromadb β₯0.5.4 (`list[str]`) and <0.5 (`list[Collection]`) |
| `frontend/src/components/Sidebar.jsx` | Workspace switcher UI; per-workspace doc list |
| `frontend/src/App.jsx`, `api.js`, `ChatArea.jsx`, `FileUpload.jsx` | `workspace_id` passed on all requests |
### Key discovery
- Non-web path was relying on stale global chain's `source_documents` rather than workspace-specific retriever. After switching workspaces, the wrong collection's docs were being cited.
---
## Stage 7 β Singleton Cache + URL Guard (2026-06-14) β Current
### What was wrong
- Every `POST /api/chat` called `get_or_create_collection()` + built a new `HybridRetriever` = full embedding reload per request β OOM after 2β3 queries in the same workspace
- React component state (message list) persisted across workspace switch β showed previous workspace's chat history
- External URL ingestion had no size guard β large pages (news articles, regulatory filings) caused OOM during embed
### What we built
| File | Change |
|------|--------|
| `server/retriever.py` | Module-level `Dict[workspace_id, (vectorstore, retriever)]` cache. Cache invalidated after ingest. `routes/chat.py` reuses cached retriever. |
| `server/routes/chat.py` | Eliminated double retrieval on non-web path; fixed stray print statement |
| `server/url_loader.py` | Max content size guard before embedding external URL content |
| `frontend/src/App.jsx` | `key={workspaceId}` on `<ChatArea>` β remounts component on workspace switch β clears stale messages and state |
### Commits
```
529675f fix: singleton vectorstore/retriever cache to prevent OOM on repeated queries
6c9f809 fix: URL size guard for OOM prevention, eliminate double retrieval in chat
521a27a fix: remount ChatArea on workspace switch to clear stale messages
```
### Additional fix (2026-06-16) β HyDE (Hypothetical Document Embeddings)
**What:** Before dense ChromaDB search, LLM generates a hypothetical 2-sentence answer. That answer (not raw query) is embedded for ANN search. BM25 + reranker still use original query.
| File | Change |
|------|--------|
| `server/retriever.py` | `_hyde_expand()` method; `use_hyde: bool` field on `HybridRetriever`; dense path uses expanded query when enabled |
| `config.yaml` | `retrieval.hyde_enabled: false` β toggle without code change |
**Why off by default:** Adds one Groq call per query (~200ms). Enable to measure RAGAS context_recall lift, then decide.
**Commit:** `8945b43`
---
### Additional fix (2026-06-17) β Mandatory web search
**Problem:** Web search was opt-in toggle. Users querying corpus-only got hallucinated answers from irrelevant documents (e.g., Singapore visa question grounded in random passport-mentioning corpus doc, faithfulness 4/5).
**Fix:**
| File | Change |
|------|--------|
| `frontend/src/components/ChatArea.jsx` | Removed toggle button; `const webSearch = true` hardcoded; placeholder always says "docs + web" |
| `server/routes/chat.py` | `web_search: bool = True` as default in `ChatRequest` |
Every query now hits Tavily + RAG corpus. `run_query_with_web` always called with both rag_docs + web_sources.
---
## Stage 8 β Eval Dashboard + Rigorous Metrics (2026-06-17)
### What was wrong
- faithfulness 1.0 and context_precision 1.0 artificially inflated β eval pairs designed alongside corpus, 8B judge lenient. Meaningless scores.
- Per-message faithfulness badge cluttered user UI. Users don't care about LLM judge scores.
- 10 samples β not statistically meaningful.
- Single flat JSON, no versioning β no way to track metric evolution across architecture changes.
### What we built
| Component | Change |
|-----------|--------|
| `eval-dashboard/` | Separate Vite + React static site (own Vercel project). Reads versioned JSON run files. |
| `eval-dashboard/src/components/` | MetricCard (score + delta vs prev), EvolutionChart (Recharts line chart across versions), RunTable (per-query expandable rows with answer vs ground_truth), LatencyStats (p50/p95 bars) |
| `eval-dashboard/public/data/index.json` | Run registry β list of all versioned eval runs |
| `scripts/run_eval_versioned.py` | New eval script. Args: `--version`, `--tag`, `--n`. Computes answer_correctness (LLM judge vs ground_truth), answer_relevancy + context_recall (RAGAS), precision@5, latency p50/p95/p99. Writes versioned JSON + updates index. |
| `data/ground_truth/eval_pairs.json` | Expanded 20 β 50 pairs. Added multi-hop, comparative, negative, numeric, and edge-case questions. |
| `frontend/src/components/MessageBubble.jsx` | Removed FaithfulnessBadge component and rendering block. |
| `server/routes/chat.py` | Removed `score_faithfulness()` call. One fewer Groq API call per query β faster responses. |
### Metrics before vs after
| Metric | Before | After |
|--------|--------|-------|
| faithfulness | 1.0 (inflated) | Removed from prod path |
| context_precision | 1.0 (inflated) | Replaced by answer_correctness (LLM judge vs ground_truth) |
| answer_relevancy | 0.88 | Kept (RAGAS) |
| context_recall | 0.83 | Kept (RAGAS) |
| precision@5 | tracked separately | Now in main eval dashboard |
| latency p50/p95 | not tracked | Now tracked per eval run |
| sample_count | 10 | 50 (5Γ improvement) |
---
## Stage 9 β Multi-Query Retrieval (2026-06-19)
### What was wrong
- context_recall = 0.51 in v1.0.0 Violet β retriever missed ~half the relevant chunks
- Single-phrasing retrieval only surfaces chunks whose vocabulary matches the query tokens
- Chunks expressing same concept with different words (e.g. "PSP ceiling" vs "merchant limit") never entered the candidate pool
### What we built
| File | Change |
|------|--------|
| `server/retriever.py` | `_multi_query_expand()` β Groq LLM generates 3 phrasings (temperature=0.3). `_get_relevant_documents()` iterates all phrasings, deduplicates by content key keeping best rank, RRF fuses pooled results, reranks with original query. |
| `config.yaml` | `retrieval.multi_query_enabled: false` toggle |
| `docs/learning.md` | Concept 17 β Multi-Query Retrieval |
### Key design decisions
- **Deduplication keeps best rank** β a chunk at rank 1 in one phrasing and rank 8 in another enters RRF at rank 1, not 8
- **Reranker uses original query** β phrasings widen the pool; the reranker judges relevance against what the user actually asked
- **Off by default** β adds one Groq call per query (~200ms). Enable β run v1.1.0 "Indigo" eval β measure delta β decide
- **retrieve_k cap maintained** β reranker input capped at `retrieve_k` even with wider pool, preserving RAM budget on Render
### Expected outcome
- context_recall: 0.51 β measurably higher (target: >0.65)
- P@5: ~0.89 (no degradation expected β reranker filters noise from wider pool)
- Latency: +200β300ms per query (one extra Groq call for phrasing generation)
- Next eval run: `scripts/run_eval_versioned.py --version v1.1.0 --tag "Indigo" --n 50`
---
## Stage 10 β Contextual Retrieval (Eval) (2026-06-20)
### What was wrong
Phase 1 (HyDE + Multi-Query) left context_recall at ~0.51. Root cause confirmed: fixed-size 500-char splits produce decontextualized chunks. `"The limit was revised to βΉ2 lakh."` has no document name, no section, no subject β weak embedding that misses ~half relevant content. Query-side techniques cannot fix bad chunk quality.
### What we built
| File | Change |
|------|--------|
| `server/ingest.py` | `contextualize_chunks(chunks, documents, model, sleep_between_calls)` β calls Groq 8B per chunk, prepends 2-sentence situating context to `page_content` before embedding. Fallback to original text on any failure. |
| `config.yaml` | `contextual_retrieval.enabled: false`, `contextual_retrieval.model: llama-3.1-8b-instant` |
| `scripts/run_eval_versioned.py` | `--contextual` flag + `--data-dir` arg. When set: clears `eval_ctx` collection, re-ingests with `contextualize_chunks`, evaluates against `eval_ctx`. Production upload untouched. |
| `tests/test_ingest.py` | 3 tests: context prepended, fallback on failure, empty chunks skipped |
### Results β v1.3.0 "Violet"
| Metric | v1.0.0 baseline | v1.3.0 contextual | Delta |
|--------|-----------------|-------------------|-------|
| context_recall | 0.510 | 0.601 | **+9.1pp (+18%)** |
| precision_at_5 | 0.890 | 0.956 | **+6.6pp** |
| answer_relevancy | 0.620 | 0.633 | +1.3pp |
| answer_correctness | 0.820 | 0.815 | -0.5pp (noise) |
| latency p50 | 2029ms | 4161ms | **+2Γ β οΈ** |
| latency p95 | β | 6122ms | β |
### Key discoveries
- Biggest single lift across all Phase 1+2 experiments: recall +9.1pp absolute
- Precision also improved significantly (0.890β0.956) β wider context gives reranker stronger signal
- Latency 2Γ because contextualized chunks are longer (~150 extra tokens per chunk) β LLM processes more tokens per answer generation call. Zero retrieval-time overhead (as designed), but query-time cost is real.
- Recall target was 0.65 β hit 0.60. Gap remains; next candidate is semantic chunking (Phase 2b)
- Production path: if shipping contextual retrieval, need FastAPI BackgroundTask for async contextualization at upload time (otherwise user waits 40s+ per doc upload)
---
## Stage 11 β Contextual Retrieval in Production + Dashboard Polish (2026-06-20)
### What was wrong
- Contextual retrieval proven in eval (recall +18%) but never shipped to production β users got non-contextual chunks
- eval-dashboard X-axis showed raw version strings (`v1.0.0`) with no dates
- No version badge visible in main Prism UI
- Encrypted PDFs caused 500 Internal Server Error instead of a clean user-facing message
- `max_concurrent=20` for parallel Groq calls β 20k token burst β 429 TPM limit on free tier (6000 TPM)
- Render free tier ephemeral filesystem: docs lost on every cold start (known limitation)
- Upload blocking for ~52s while Euron embedding API processes chunks sequentially
### What we built
| File | Change |
|------|--------|
| `server/ingest.py` | `contextualize_chunks_async()` β parallel Groq calls via `asyncio.gather` + `Semaphore(max_concurrent)`. ~10Γ faster than sequential. Retry parses suggested wait time from 429 error message. |
| `server/routes/upload.py` | Two-phase upload: sync non-contextual embed first (user queryable immediately), then `_contextual_refresh_bg()` BackgroundTask replaces non-contextual chunks with contextual versions. |
| `config.yaml` | `contextual_retrieval.enabled: true`, `max_concurrent: 3` (3 Γ ~1000 tokens = 3000 TPM β safe under 6000 limit) |
| `server/ingest.py` | `load_documents_from_paths()`: catches `FileNotDecryptedError` β raises `ValueError` with user-friendly message |
| `server/routes/upload.py` | Catches `ValueError` from loader β returns HTTP 422 instead of 500 |
| `eval-dashboard/src/components/EvolutionChart.jsx` | Custom `XAxisTick`: stacked version name + short date (e.g. `Violet (v1.3)` / `20 Jun 26`) |
| `eval-dashboard/src/App.jsx` | `VERSION_NOTES` constant with bullet notes per version; release notes panel shown below run meta |
| `frontend/src/components/Sidebar.jsx` | `Violet v1.3` badge (indigo pill) in sidebar footer |
| `frontend/src/config.js` | New file β `MAINTENANCE_MODE` + `MAINTENANCE_MESSAGE` config flags |
| `frontend/src/App.jsx` | Maintenance banner driven by `config.js`; hidden when `MAINTENANCE_MODE = false` |
### Key discoveries
- `asyncio.gather` with `Semaphore(3)` keeps burst under 3000 TPM β safe on Groq free tier (6000 TPM limit)
- Groq 429 errors include `"Please try again in X.Xs"` β parse this for accurate retry sleep instead of hardcoded 2s
- Render free tier: ephemeral filesystem. Every cold start wipes `./chroma_db`. Docs must be re-uploaded. Fix: Render persistent disk ($0.25/GB/month)
- Euron embedding API sequential calls: 30 chunks Γ ~1.7s/call = ~52s blocking upload. Next optimization: move embed to background too (return 202 immediately, notify when ready)
- `max_concurrent=20` was the OOM trigger in the previous session β 20 async coroutines each holding ~10MB response + retry state saturated 512MB
---
## Stage 15 β Semantic Chunking Ablation + Retrieval Stack Finalized (2026-06-26)
### What was wrong
Ablation study incomplete β semantic chunking (v1.4.0) was blocked by Groq rate limits in the prior session. Best production stack unconfirmed.
### What we built / ran
| Version | Config | recall | P@5 | relevancy | correctness | p50 |
|---------|--------|--------|-----|-----------|-------------|-----|
| v1.1.0 | HyDE | 0.721 | 0.911 | 0.845 | 0.750 | 4018ms |
| v1.2.0 | HyDE+MQ | 0.645 | 0.904 | 0.890 | 0.770 | 1812ms |
| v1.3.0 | HyDE+MQ+CTX | 0.768 | **0.984** | 0.799 | 0.780 | 2610ms |
| v1.4.0 | HyDE+MQ+CTX+Semantic | **0.861** | 0.711 | 0.885 | 0.750 | 12952ms |
### Decision: semantic chunking rejected
Semantic chunking raises recall +9.3pp (0.768β0.861) but P@5 collapses -27.3pp (0.984β0.711) and latency is 5Γ worse (2610msβ12952ms p50).
**Root cause of P@5 collapse:** SemanticChunker produces variable-size, topic-boundary chunks. These don't align with the fixed ground-truth keyword spans used for precision@5 scoring. The reranker receives a wider but noisier candidate pool β recall expands while precision degrades.
**Best stack confirmed: v1.3.0 β HyDE + Multi-Query + Contextual Retrieval.**
### Key discoveries
- MQ alone hurts recall (-7.6pp vs HyDE-only) but recovers fully when combined with CTX
- CTX is highest-leverage single addition: +8pp P@5, recall recovery, at 2Γ query latency cost
- Semantic chunking is a double-edged sword β better chunk boundaries for recall, worse alignment with precision evaluation
- Ablation study is the interview story: systematic metric-driven elimination of techniques
---
## Stage 16 β Citation Highlighting (2026-06-26)
### What was wrong
Sources listed below each answer as truncated 200-char snippets. LLM already outputs `[1]`, `[2]` inline citations but they rendered as plain unclickable text. Users couldn't see which passage in the answer corresponded to which source.
### What we built
| File | Change |
|------|--------|
| `frontend/src/components/CitationPopover.jsx` | New β viewport-aware popover (fixed-position). Shows: source type badge (pdf/web/file), filename/title, page badge, full chunk content (scrollable), rerank score, "Open page N β" for PDF / "Open source β" for web |
| `frontend/src/components/MessageBubble.jsx` | Parse `[N]` markers in answer text β clickable `<sup>` superscripts. `openCitation` state (`{ idx, rect } \| null`). Toggle on same click. `onMouseDown` stopPropagation fix (prevents document mousedown from immediately re-opening after close). |
| `frontend/src/components/SourceExpander.jsx` | Removed 200-char content truncation β full chunk text shown |
| `server/main.py` | Added `GET /api/files/{filename}` β `FileResponse` from `data/raw/`. Path traversal blocked via `is_relative_to()`. `UPLOAD_DIR` made absolute (`Path(__file__).resolve().parent.parent / "data" / "raw"`). |
| `server/routes/upload.py` | `UPLOAD_DIR` made absolute (`Path(__file__).resolve().parent.parent.parent / "data" / "raw"`) |
### Key discoveries
- `mousedown` on document fires before `click` β without `e.stopPropagation()` on the `<sup>` mousedown, clicking an open citation closes then immediately reopens it (toggle broken)
- `startswith()` on raw path strings has prefix-confusion bug (`/data/rawevil` passes `/data/raw` check) β replaced with `Path.is_relative_to()` (Python 3.9+)
- Relative `Path("data/raw")` resolves against process CWD β if uvicorn starts from non-project-root directory, file serving breaks. Absolute `__file__`-relative path fixes this.
- `anchorRect` captured at click time via `el.getBoundingClientRect()` β stored in state as plain object, no ref needed in popover
### Interview story
> "The LLM cites [1], [2] in its answer. Clicking one opens a popover showing the exact passage retrieved β full text, source file, page number, and rerank score. For PDFs it links directly to that page in the browser."
---
## Stage 17 β Metadata Filtering (2026-06-27)
### What was wrong
All documents in a workspace were always searched together. A user with 10 docs spanning 5 years had no way to scope a query to a specific doc or subset. Corpus-wide retrieval diluted precision when the relevant content was known to be in one file.
### What we built
| File | Change |
|------|--------|
| `server/ingest.py` | `source_type` metadata field (`pdf`/`txt`/`csv`) added to all chunks at load time via `SOURCE_TYPE_MAP`. Both `load_documents` and `load_documents_from_paths` patched. |
| `server/url_loader.py` | `source_type: "url"` added to URL-ingested doc metadata. |
| `server/bm25_index.py` | `BM25Index.search()` gets `filter_sources: set[str] \| None = None`. When set, scores using full-corpus BM25 index (stable IDF) but restricts candidate pool to matching docs. |
| `server/retriever.py` | `filter_docs: list[str] \| None = None` field on `HybridRetriever`. Wired into `_dense_retrieve` (ChromaDB `where={"source": {"$in": filter_docs}}`) and `_get_relevant_documents` (BM25 `filter_sources`). New `get_retriever_filtered(workspace_id, filter_docs)` helper β one-off instance reusing cached vectorstore, not added to singleton cache. |
| `server/routes/chat.py` | `filter_docs: list[str] \| None = None` on `ChatRequest`. Guard: empty list β None. When truthy: `get_retriever_filtered(workspace, active_filter)`. Log includes `filter=%s`. |
| `frontend/src/api.js` | `streamChat` gets `filterDocs = null` as 4th arg; sends `filter_docs: filterDocs?.length ? filterDocs : null`. |
| `frontend/src/App.jsx` | `filterDocs: string[]` state. `useEffect` resets to `[]` on workspace change. `handleFilterChange` toggles doc in/out. Props forwarded to Sidebar + ChatArea. |
| `frontend/src/components/Sidebar.jsx` | Doc list items clickable β toggle filter on click. Selected: `ring-2 ring-indigo-500 bg-indigo-50`. Unselected during active filter: `opacity-50`. "Clear filter" button in section header when any selected. Delete button: `e.stopPropagation()` + deselects deleted doc from filter. |
| `frontend/src/components/ChatArea.jsx` | Filter badge above input when `filterDocs.length > 0` (shows scoped doc names + Γ clear). Placeholder: "Searching N selected doc(s)..." when filter active. `streamChat` called with `filterDocs.length > 0 ? filterDocs : null`. |
| `tests/test_bm25_filter.py` | 6 tests: no-filter returns all, filter restricts by source, empty set returns empty, nonexistent source returns empty, multiple sources, unbuilt index returns empty. |
| `tests/test_source_type.py` | 4 tests: pdf/txt/csv/url each gets correct `source_type`. |
### Key design decisions
- **Full-corpus BM25 for filtering**: spec suggested rebuilding BM25 on filtered subset; implementation uses full-corpus index to score + restricts candidate pool by source. Stable IDF β correct IR semantics. Accepted as superior to spec.
- **New retriever instance per filtered request**: `get_retriever_filtered()` creates a one-off `HybridRetriever`; singleton cache (`_retriever_cache`) untouched. Thread-safe: heavy vectorstore stays cached, lightweight retriever is cheap.
- **Empty filter = no filter**: backend guard `body.filter_docs if body.filter_docs else None` β empty array from frontend treated as no filter.
- **Filter resets on workspace switch**: `useEffect(() => setFilterDocs([]), [currentWorkspace])` β stale filter from workspace A doesn't carry to workspace B.
- **Delete deselects**: `handleDelete` calls `onFilterChange(docName)` if deleted doc was selected β prevents badge showing "Scoped to: [deleted]" with zero results.
### Interview story
> "Within a workspace, users can click any doc chip in the sidebar to scope retrieval. Dense retrieval passes `where={"source": {"$in": selected_docs}}` to ChromaDB; BM25 pre-filters its candidate pool. Zero selection = full-corpus behavior unchanged. Filter badge above the input makes the scope visible."
---
## Stage 18 β Free-Tier Stability + App Restored to Live (2026-07-05)
### What was wrong
- Maintenance banner left ON after Cerebras migration attempt (2026-06-29) failed and was reverted
- Config.yaml still had `hyde_enabled: true` + `multi_query_enabled: true` β each query burned 4 Groq calls
- Free tier limit: 6000 TPM β 429 storms under concurrent use with HyDE + MQ + contextual all on
- Eval dashboard had no indication which version is live or why best stack (v1.3.0) isn't deployed
### What we built
| File | Change |
|------|--------|
| `config.yaml` | `hyde_enabled: false`, `multi_query_enabled: false` β reduces query-time Groq calls 4 β 1β2 |
| `frontend/src/config.js` | `MAINTENANCE_MODE: false` β app live |
| `eval-dashboard/public/data/index.json` | `is_live: true` + `live_note` on v1.1.0 (closest proxy); `blocked_by` constraint on v1.3.0 + v1.4.0 |
| `eval-dashboard/src/App.jsx` | Green LIVE badge + prod config note on v1.1.0; amber "not in production" warning on v1.3.0/v1.4.0 |
### Key design decisions
- **Contextual retrieval kept ON** β uses `openai/gpt-oss-20b` via Euron API, zero Groq TPM impact at query time. Ingest-time only.
- **HyDE + MQ disabled, not removed** β toggles in config.yaml; re-enable instantly when on paid tier
- **v1.1.0 as live proxy in eval dashboard** β no eval run exists for "CTX-only, no HyDE, no MQ" config. v1.1.0 (recall=0.721) is an overestimate; actual live recall β 0.55β0.65 given contextual index without HyDE query expansion
- **Upgrade path documented in eval dashboard** β v1.3.0 blocked_by note explains exactly what to fix
### Groq call budget (current vs best)
| Config | Calls/query | TPM risk |
|--------|------------|----------|
| Current (CTX only) | 1β2 | Safe |
| v1.3.0 (HyDE+MQ+CTX) | 4 | 429 on free tier |
### Upgrade path to v1.3.0
1. Switch to paid Groq tier (or find higher-TPM free provider)
2. Set `hyde_enabled: true` + `multi_query_enabled: true` in `config.yaml`
3. Push β HF Spaces rebuilds β run `scripts/run_eval_versioned.py --version v1.3.1 --tag "Violet" --n 50` to confirm metrics
---
## Current State Snapshot
```
Retrieval: Hybrid BM25 (0.3) + ChromaDB dense (0.7) β RRF β TinyBERT rerank top-10β5
LLM: Groq llama-3.3-70b-versatile
Embeddings: Euron API text-embedding-3-small (sequential, ~1.7s/chunk β bottleneck)
Chunking: RecursiveCharacterTextSplitter 500-char, overlap 50
Memory: ConversationBufferWindowMemory k=10
Web search: Tavily advanced, 800-char truncation, max 2 results β MANDATORY (always on)
HyDE: DISABLED (hyde_enabled=false). Best measured: +21pp recall but costs 1 Groq call/query.
Re-enable when on paid Groq tier or higher-TPM provider.
Multi-Query: DISABLED (multi_query_enabled=false). Costs 1 Groq call/query β free tier cannot sustain.
Re-enable with HyDE together (v1.3.0 config) on paid tier.
Contextual: ENABLED (contextual_retrieval.enabled=true). Uses Euron model (openai/gpt-oss-20b) β
zero Groq TPM impact. Two-phase upload: sync non-contextual embed (<3s queryable),
BackgroundTask replaces with contextual chunks. max_concurrent=3, max_chunks=50 gate.
Semantic: DISABLED (semantic_enabled=false). Ablation showed recall +9.3pp but P@5 -27.3pp and 5Γ latency.
Rejected β v1.3.0 (HyDE+MQ+CTX) is the confirmed best stack when TPM allows.
Groq calls/query (current): 1β2 (condense_question if follow-up + answer). Safe under 6000 TPM free tier.
Groq calls/query (v1.3.0): 4 (condense + HyDE + Multi-Query + answer) β 429 storms on free tier.
Eval: Separate eval-dashboard/ static site β https://askprism-eval.vercel.app/
v1.1.0 marked LIVE (closest proxy). v1.3.0 and v1.4.0 show amber "not in production" warning.
Best measured: v1.3.0 recall=0.768, P@5=0.984, p50=2610ms
Versioning: MAJOR.MINOR.PATCH β name changes on MAJOR only (v1.x.x=Violet, v2.x.x=Indigo)
Citation: [N] markers in LLM answers β clickable <sup> β CitationPopover (fixed-position, viewport-aware).
Shows full chunk text, source name, page, rerank score. PDF: "Open page N β" link via GET /api/files/{filename}.
Web: "Open source β". Toggle, click-away, above/below flip at 60% viewport height.
SourceExpander: full content shown (200-char truncation removed).
Frontend: Violet v1.3 badge in sidebar footer. Maintenance banner config-driven (frontend/src/config.js).
MAINTENANCE_MODE=false β app is live as of 2026-07-05.
Filter: Sidebar doc chips toggleable. Selected: indigo ring. Badge above chat input shows scoped docs + clear Γ.
POST /api/chat accepts filter_docs: string[] | null. Empty = no filter. Resets on workspace switch.
Backend: get_retriever_filtered() creates one-off HybridRetriever; singleton cache untouched.
ChromaDB where={"source": {"$in": filter_docs}}. BM25 filters candidate pool, scores with full-corpus IDF.
Workspaces: Per-workspace ChromaDB collection, singleton retriever cache
Infra: HF Spaces CPU Basic (backend, 16GB RAM, ephemeral FS β re-upload required after cold start) +
https://askprism.vercel.app/ (frontend) + https://askprism-eval.vercel.app/ (eval)
Backend URL: https://benroshan-prism.hf.space
Known limits: Euron embed ~5s/chunk sequential β 30 chunks = ~150s total contextualization in background.
HF Spaces ephemeral FS: chroma_db lost on cold start. Fix: mount HF persistent storage bucket.
HyDE + MQ disabled for free-tier stability. Best stack (v1.3.0) needs paid Groq or alt provider.
Observability: LangSmith traces all LLM + retrieval calls (optional, env var)
Streaming: POST /api/chat returns SSE stream. token events per LLM chunk, done event with
sources + retrieval_method. Frontend streams tokens into pre-placed assistant
bubble. Bouncing dots while condense+search runs, blinking cursor during generation.
```
---
## Stage 12 β HF Spaces Migration (2026-06-22)
### What was wrong
Render free tier (512MB RAM) caused repeated OOM crashes under contextual retrieval:
- Base RSS after upload = 524MB (over the 512MB limit)
- `gc.collect()` had no effect β ChromaDB HNSW index + torch runtime held by native allocators, not Python heap
- Contextual refresh (3 async Groq coroutines) + simultaneous chat (Tavily + LLM + CrossEncoder) = peak exceeded 512MB
- Workarounds (RSS guard skipping contextual retrieval, web search suppression during refresh) negated the +18% recall improvement
### What we built
| File | Change |
|------|--------|
| `Dockerfile` | Port 8000 β 7860 (HF convention). Add `useradd -m -u 1000 user` + `chown -R user /app` (HF runs containers as UID 1000). Set `HF_HOME=/app/.cache/huggingface` BEFORE pre-download so user 1000 owns cached weights. Set `HF_HUB_OFFLINE=1` AFTER download. |
| `README.md` | Added HF Spaces frontmatter (`sdk: docker`, `app_port: 7860`). Updated deploy instructions. |
| `server/routes/chat.py` | Removed `is_contextualizing` web search suppression guard (Render-specific). |
| `server/routes/upload.py` | Removed `RSS > 460MB` contextual retrieval skip guard (Render-specific). |
| `docs/`, `decisions.md` | Render β HF Spaces across all infra references. |
### Key discoveries
- HF_HUB_OFFLINE must be set AFTER the pre-download RUN step β setting it before blocks the download itself
- Docker build runs pre-download as root by default; must `USER 1000` first then set `HF_HOME` under `/app` so runtime user 1000 can read the cached weights
- HF Spaces free CPU Basic: 2 vCPUs, 16GB RAM β resolves all Render OOM issues permanently
- Contextual retrieval now runs fully in production (was silently skipped by RSS guard on Render)
---
## Stage 13 β Async Embed Upload (2026-06-23)
### What was wrong
`embed_and_store()` blocked `POST /api/upload` for ~150s (30 chunks Γ ~5s/chunk via Euron API). User saw spinner, could not query, could not cancel. Upload timeout was 300s.
### What we built
| File | Change |
|------|--------|
| `server/main.py` | `app.state.upload_jobs = {}` initialized in lifespan |
| `server/routes/upload.py` | `POST /api/upload` returns 202 + `job_id` in <1s. Parse+chunk sync; embed+contextual in `_embed_and_contextualize_bg()` BackgroundTask. New `GET /api/upload/status/{job_id}` endpoint. |
| `frontend/src/api.js` | Added `getUploadStatus(jobId)`; reduced `uploadFiles` timeout 300s β 30s |
| `frontend/src/components/FileUpload.jsx` | Polls status every 2s; shows stage label under spinner; fires callbacks on ready. Defensive `|| []` guard on documents. |
### Key discoveries
- Old Vercel frontend receiving new 202 response before redeploy β `data.documents` undefined β React crash. Fix: defensive `docs?.documents || []` guard.
- Groq TPM 429s at `max_concurrent=3` still hit (~5/30 chunks fall back to original text) β some chunks are larger than average. Retry logic handles gracefully.
- Briefing fails with JSON parse error (pre-existing bug in `generate_briefing` β separate fix).
---
## Stage 14 β Briefing Fix + HyDE Re-eval (2026-06-24)
### What was wrong
- `generate_briefing()` crashed with `JSONDecodeError` when Groq LLM returned control characters (ASCII 0x00β0x1f) or Python dict syntax (single quotes) instead of valid JSON.
- Old eval runs (v1.0.0βv1.4.0) accumulated across multiple sessions; stale runs cluttered the dashboard.
- HyDE recall measurement from prior session (v1.1.0_20260619, recall=0.545) was based on 50-sample run that hit Groq 429s mid-run β partial results, unreliable numbers.
### What we built
| File | Change |
|------|--------|
| `server/briefing.py` | Strip control chars `[\x00-\x08\x0b\x0c\x0e-\x1f]` before JSON parse. Fall back to `ast.literal_eval()` on `JSONDecodeError` to handle Python dict syntax from LLM. Added `import ast`. |
| `config.yaml` | `hyde_enabled: true`, `multi_query_enabled: true`, `contextual_retrieval.enabled: false` (contextual off β 429s at 30-chunk scale even with Semaphore(3)) |
| `eval-dashboard/public/data/runs/` | Deleted stale runs (v1.0.0_20260618, v1.1.0_20260619, v1.2.0_20260619, v1.3.0_20260620, v1.3.0_20260623, v1.4.0_20260623). Added `v1.1.0_20260624.json` β fresh HyDE-only run. |
| `eval-dashboard/public/data/index.json` | Updated to single clean run registry. |
### HyDE re-eval results β v1.1.0_20260624 (18 samples, hyde=true, multi_query=false)
| Metric | v1.0.0 baseline | v1.1.0 HyDE | Delta |
|--------|-----------------|-------------|-------|
| answer_correctness | 0.820 | 0.750 | -0.070 |
| answer_relevancy | 0.620 | 0.845 | **+0.225** |
| context_recall | 0.510 | 0.721 | **+0.211** |
| precision_at_5 | 0.890 | 0.911 | +0.021 |
| latency p50 | 2029ms | 4018ms | +2Γ |
### Key discoveries
- HyDE gives **+21pp recall** (0.51β0.72) on this 18-sample run β much larger than previously measured (+3.5pp on 50 samples with 429s). Smaller sample set; repeat at 50 samples to confirm.
- answer_correctness flat at 0.75 for all 18 samples β 8B judge giving uniform score, not differentiating. May indicate judge calibration issue, not actual correctness plateau.
- Latency 2Γ (2029msβ4018ms) β HyDE adds one Groq call per query for hypothetical expansion.
- Briefing fix unblocks document upload β briefing flow end-to-end.
---
## Roadmap β Retrieval & Answer Quality
### Phase 1 β Quick wins (no infra change, measurable RAGAS lift)
#### ~~HyDE (Hypothetical Document Embeddings)~~ β
Done (Stage 7, commit 8945b43)
- Implemented in `server/retriever.py`. Toggle: `config.yaml hyde_enabled` (default: false).
- Enable + re-run eval to measure context_recall lift vs v2.0 baseline (0.70).
#### ~~Multi-Query Retrieval~~ β
Done (Stage 9, 2026-06-19)
- Implemented in `server/retriever.py`. Toggle: `config.yaml multi_query_enabled` (default: false).
- Enable + run `scripts/run_eval_versioned.py --version v1.1.0 --tag "Indigo" --n 50` to measure context_recall lift vs 0.51.
---
### Phase 2 β Ingest pipeline (requires re-ingest of all docs)
#### ~~Contextual Retrieval~~ β
Done + Shipped to Production (Stage 10+11, 2026-06-20)
- `contextualize_chunks_async()` + BackgroundTask in `routes/upload.py`. Two-phase: sync non-contextual embed (queryable <3s) β background contextual replacement.
- v1.3.0 results: recall 0.510β0.601 (+18%), P@5 0.890β0.956. Latency 2Γ at query time (longer chunks β more LLM tokens).
- `max_concurrent=3` in `config.yaml` β safe under Groq 6000 TPM limit.
#### Semantic Chunking
- **Problem:** Fixed 200-char splits cut mid-sentence, mid-table, mid-list. Embedding a truncated sentence returns a weak vector.
- **How:** Replace `RecursiveCharacterTextSplitter` with LangChain's `SemanticChunker` β splits at sentence boundaries where cosine similarity between adjacent sentences drops below a threshold (topic shift).
- **Effort:** Medium. Config change in `ingest.py` + re-ingest. Tune `breakpoint_threshold_type`.
- **Expected lift:** Fewer nonsensical chunks in top-5. Most noticeable on regulatory PDFs with section headers and numbered lists.
---
### Phase 3 β UX + trust
#### Streaming Responses
- **Problem:** User submits question β 8β15s wait β full answer appears. Feels broken even on fast hardware.
- **How:** Backend: `chain.astream_events()` β `StreamingResponse` yielding SSE tokens. Frontend: `EventSource` or `fetch` + `ReadableStream` β append tokens as they arrive. Faithfulness scoring runs as background task after full answer assembled.
- **Effort:** High β both backend and frontend change. `ConversationalRetrievalChain` supports `astream_events()` in LangChain β₯0.2.
- **Impact:** Perceived latency drops from 10s to ~1s. Single biggest UX improvement.
#### ~~Citation Highlighting~~ β
Done (Stage 16, 2026-06-26)
- `[N]` markers clickable β `CitationPopover` with full chunk text, page badge, rerank score. PDF "Open page N β" link. Zero new npm deps.
- Works for all source types: PDF, URL, TXT, CSV. No PDF viewer library needed β page link uses browser's built-in viewer.
---
### Phase 4 β Differentiation
#### Metadata Filtering
- **Problem:** Multi-workspace isolates by collection, but within a workspace (10 docs across 5 years) no way to scope retrieval to `year=2024` or `doc_type=rbi_circular`.
- **How:** Tag chunks with `{source_type, year, doc_name}` at ingest. Pass optional `filter` param in `/api/chat` request. ChromaDB `where` clause on dense retrieval; BM25 pre-filters corpus to matching chunk IDs.
- **Impact:** Precision boost on time-scoped or source-scoped queries.
#### Document Comparison Mode
- **Problem:** No way to ask "What changed between RBI circular 2023 and 2024?"
- **How:** Frontend sends two doc IDs + comparison query. Backend retrieves relevant chunks from each collection separately, synthesises a structured diff answer.
- **Impact:** Killer fintech feature. Unique demo moment. Differentiates from generic RAG.
#### Agentic Mode (LangGraph)
- **Problem:** Single-shot RAG cannot handle multi-step reasoning: retrieve β compute β web search β synthesise.
- **How:** Replace `ConversationalRetrievalChain` with a LangGraph graph. Nodes: retriever, web_search, calculator, synthesiser. LLM decides which tool to call.
- **Impact:** Separates Prism from basic RAG β becomes a research agent. Strongest interview story.
---
## Roadmap Priority Matrix
```
HIGH impact Γ LOW effort β Build first
HyDE
Multi-query retrieval
Metadata filtering
HIGH impact Γ MEDIUM effort β Build second
Contextual retrieval (+ re-ingest)
Semantic chunking (+ re-ingest)
Streaming responses
HIGH impact Γ HIGH effort β Build last
Citation highlighting
Document comparison
Agentic mode (LangGraph)
```
---
## Interview Story Arc
```
v1 β Dense-only retrieval. No eval. No baseline.
v2 β Hybrid BM25+dense, cross-encoder rerank. Measured with RAGAS.
β faithfulness=1.0, answer_relevancy=0.90 on 20-pair eval set.
+HyDE β context_recall 0.51β0.72 (+21pp). Hypothetical answer embedding closes vocabulary gap.
+Contextual β context_recall 0.60 (+18% vs baseline). Ingest-time LLM chunk augmentation.
+Agentic β Multi-step reasoning. Not RAG anymore β research agent.
```
Each step has a metric. That is the complete RAG engineering narrative for MNC DS interviews.
|