File size: 50,911 Bytes
65cee4e
 
 
95143a2
65cee4e
 
 
 
 
 
 
 
 
 
 
 
8ec920c
 
44e600b
05820e3
eca5609
d281212
84f9c55
20130d7
 
8c5326e
84f9c55
05820e3
 
65cee4e
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
44e600b
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
05820e3
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
84f9c55
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
10860d7
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
20130d7
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
8c5326e
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
95143a2
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
65cee4e
 
 
 
 
84f9c55
05820e3
65cee4e
 
95143a2
 
 
 
 
 
 
10860d7
95143a2
 
 
8ec920c
95143a2
 
05820e3
20130d7
 
 
 
84f9c55
95143a2
8c5326e
 
 
 
65cee4e
d281212
84f9c55
d281212
95143a2
d281212
95143a2
8ec920c
5d9a6f1
 
 
65cee4e
 
 
 
d281212
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
3f7febe
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
eca5609
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
65cee4e
 
 
 
8ec920c
 
 
65cee4e
44e600b
 
 
65cee4e
 
 
 
 
84f9c55
 
 
 
65cee4e
 
 
 
 
 
 
 
 
 
 
 
d281212
65cee4e
 
 
 
20130d7
 
 
65cee4e
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
eca5609
 
65cee4e
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
487
488
489
490
491
492
493
494
495
496
497
498
499
500
501
502
503
504
505
506
507
508
509
510
511
512
513
514
515
516
517
518
519
520
521
522
523
524
525
526
527
528
529
530
531
532
533
534
535
536
537
538
539
540
541
542
543
544
545
546
547
548
549
550
551
552
553
554
555
556
557
558
559
560
561
562
563
564
565
566
567
568
569
570
571
572
573
574
575
576
577
578
579
580
581
582
583
584
585
586
587
588
589
590
591
592
593
594
595
596
597
598
599
600
601
602
603
604
605
606
607
608
609
610
611
612
613
614
615
616
617
618
619
620
621
622
623
624
625
626
627
628
629
630
631
632
633
634
635
636
637
638
639
640
641
642
643
644
645
646
647
648
649
650
651
652
653
654
655
656
657
658
659
660
661
662
663
664
665
666
667
668
669
670
671
672
673
674
675
676
677
678
679
680
681
682
683
684
685
686
687
688
689
690
691
692
693
694
695
696
697
698
699
700
701
702
703
704
705
706
707
708
709
710
711
712
713
714
715
716
717
718
719
720
721
722
723
724
725
726
727
728
729
730
731
732
733
734
735
736
737
738
739
740
741
742
743
744
745
746
747
748
749
750
751
752
753
754
755
756
757
758
759
760
761
762
763
764
765
766
767
768
769
770
771
772
773
774
775
776
777
778
779
780
781
782
783
784
785
786
787
788
789
790
791
792
793
794
795
796
797
798
799
800
801
802
803
804
805
806
807
808
809
810
811
812
813
814
815
816
817
818
819
820
821
822
# Prism β€” Project Evolution

> End-to-end record of what was broken at each stage, what was built to fix it, and what is planned next.
> Updated as the project evolves. Last updated: 2026-07-05 (Stage 18).

---

## Table of Contents

1. [Stage 0 β€” v1 Baseline](#stage-0--v1-baseline)
2. [Stage 1 β€” v2 Hybrid Retrieval Architecture](#stage-1--v2-hybrid-retrieval-architecture-2026-05-17)
3. [Stage 2 β€” Chain Scores + RAGAS Endpoint](#stage-2--chain-scores--ragas-endpoint-2026-05-23)
4. [Stage 3 β€” Groq Migration + Web Search](#stage-3--groq-migration--web-search-fixes-2026-05-24)
5. [Stage 4 β€” OOM Hell on Render](#stage-4--oom-hell-on-render-2026-05-24-four-sub-issues)
6. [Stage 5 β€” TinyBERT + RAGAS Removal + Benchmark JSON](#stage-5--tinybert--ragas-removal--benchmark-json-2026-05-30)
7. [Stage 6 β€” Multi-Workspace](#stage-6--multi-workspace-2026-06-early)
8. [Stage 7 β€” Singleton Cache + URL Guard](#stage-7--singleton-cache--url-guard-2026-06-14)
9. [Stage 8 β€” Eval Dashboard + Rigorous Metrics](#stage-8--eval-dashboard--rigorous-metrics-2026-06-17)
10. [Stage 9 β€” Multi-Query Retrieval](#stage-9--multi-query-retrieval-2026-06-19)
11. [Stage 10 β€” Contextual Retrieval (Eval)](#stage-10--contextual-retrieval-eval-2026-06-20)
14. [Stage 14 β€” Briefing Fix + HyDE Re-eval](#stage-14--briefing-fix--hyde-re-eval-2026-06-24)
12. [Stage 12 β€” HF Spaces Migration](#stage-12--hf-spaces-migration-2026-06-22)
12. [Stage 11 β€” Contextual Retrieval in Production + Dashboard Polish](#stage-11--contextual-retrieval-in-production--dashboard-polish-2026-06-20)
15. [Stage 15 β€” Semantic Chunking Ablation + Retrieval Stack Finalized](#stage-15--semantic-chunking-ablation--retrieval-stack-finalized-2026-06-26)
16. [Stage 16 β€” Citation Highlighting](#stage-16--citation-highlighting-2026-06-26)
17. [Stage 17 β€” Metadata Filtering](#stage-17--metadata-filtering-2026-06-27)
13. [Current State Snapshot](#current-state-snapshot)
13. [Roadmap β€” Retrieval & Answer Quality](#roadmap--retrieval--answer-quality)
14. [Roadmap β€” New Features](#roadmap--new-features)

---

## Stage 0 β€” v1 Baseline

### What existed
- Dense-only ChromaDB vector retrieval
- Single global document collection
- Basic chat with ConversationalRetrievalChain
- No evaluation framework
- No web search
- No logging

### What was wrong

| Problem | Impact |
|---------|--------|
| Dense-only retrieval | Misses exact keyword matches β€” regulatory text has section numbers, policy codes, specific terms that semantic search fails on |
| No evaluation | No way to measure if answers were correct or grounded |
| No web search | Static corpus only β€” cannot answer questions about current stock prices, recent news |
| Single collection | No topic isolation β€” all documents mixed in one retrieval pool |
| No logging | Impossible to debug production failures |

**This was the starting point. No fixes yet.**

---

## Stage 1 β€” v2 Hybrid Retrieval Architecture (2026-05-17)

### What was wrong before building
- `retriever.py` was dense-only ChromaDB β€” v2 was documented but not implemented
- `ragas_eval.py` missing entirely; RAGAS eval endpoint not wired
- No BM25, no reranker, no score visibility

### What we built

| File | What changed |
|------|-------------|
| `server/bm25_index.py` | BM25Okapi singleton; module-level (not `app.state`) so importable anywhere; rebuilt on startup + after upload |
| `server/reranker.py` | CrossEncoder singleton; pre-loaded at startup to avoid cold-start latency on first query |
| `server/retriever.py` | Full rewrite as `HybridRetriever(BaseRetriever)` β€” RRF fusion of dense (weight 0.7) + sparse (weight 0.3) |
| `server/main.py` | BM25 build + reranker load wired into lifespan startup |
| `server/routes/upload.py` | BM25 rebuild triggered after each upload |

### Key design decisions

- **`HybridRetriever` as `BaseRetriever` subclass** β€” `ConversationalRetrievalChain` expects a `BaseRetriever`; subclassing means `chain.py` needs zero changes
- **BM25 as module-level singleton** β€” avoids threading state through lifespan β†’ constructor; `get_index()` importable anywhere
- **Reranker pre-loaded at startup** β€” ~0.5s load from disk cache; better to pay at startup than add latency to first user query
- **MiniLM-L-6-v2** chosen as reranker (~85MB) β€” best ranking quality available at the time

### What was still missing
- Chain score extraction (similarity/BM25/RRF/rerank not returned in API response)
- RAGAS eval endpoint
- Groq LLM (still on Euron)

---

## Stage 2 β€” Chain Scores + RAGAS Endpoint (2026-05-23)

### What was wrong
- API response had no retrieval scores β€” no way to show per-source similarity/BM25/RRF/rerank scores
- RAGAS eval endpoint not wired; `ragas_eval.py` missing
- `data/ground_truth/eval_pairs.json` had only keyword hints, no `ground_truth` answers β†’ `context_precision` and `context_recall` always returned null

### What we built

| File | What changed |
|------|-------------|
| `server/chain.py` | Score extraction β€” similarity/bm25/rrf/rerank scores passed through to API response per source |
| `server/routes/chat.py` | Added `retrieval_method` field; stores contexts in `eval_log` for downstream RAGAS eval |
| `server/config.yaml` | Hybrid retrieval params: `dense_weight`, `sparse_weight`, `retrieve_k`, `rerank_k` |
| `requirements.txt` | Added `rank_bm25`, `sentence-transformers`, `ragas`, `datasets` |
| `server/eval/ragas_eval.py` | RAGAS faithfulness + answer_relevancy via `LangchainLLMWrapper` |
| `server/routes/eval.py` | `POST /api/eval/ragas` endpoint wired |
| `server/main.py` | `/health` endpoint added |

### What was still broken
- RAGAS not installed in venv (added to requirements.txt; installs at Docker build only)
- `eval_pairs.json` still had no ground_truth β†’ 2 of 4 RAGAS metrics null
- LLM still on Euron gpt-4.1-mini

---

## Stage 3 β€” Groq Migration + Web Search Fixes (2026-05-24)

### What was wrong

| Problem | Root cause |
|---------|-----------|
| LLM on Euron (gpt-4.1-mini) | Closed model, slower, weaker interview story vs open-weight |
| Web search silently broken | `tavily-python` in requirements.txt but never pip-installed |
| Tavily returned shallow results | `search_depth="basic"` β€” not enough content from financial sites |
| Web context never reached LLM | `ConversationalRetrievalChain`'s condensation step rewrote the question and stripped prepended Tavily context before LLM ever saw it |
| Follow-up web queries returned garbage | Raw follow-up ("Is the price level good?") sent to Tavily with no chat history context |
| No request logging | Production failures undebuggable |

### What we built

| Component | Change |
|-----------|--------|
| LLM | Migrated Euron β†’ Groq `llama-3.3-70b-versatile` via `langchain-groq`. Euron kept for embeddings (Groq has no embeddings endpoint) |
| `server/chain.py` | `run_query_with_web()` β€” bypasses chain condensation; direct LLM call with RAG + Tavily context + memory |
| `server/chain.py` | `condense_question()` β€” rewrites follow-up queries using chat history before Tavily search |
| Tavily | `search_depth="advanced"`, `max_results=3` (2Γ— credits but richer content) |
| `server/main.py` | Request logging middleware β€” logs `METHOD /path STATUS Xms` per request |
| `server/utils.py` | Centralised logging β€” root logger + `logs/finrag.log` (5MBΓ—3 rotation), noisy libs silenced |
| `frontend/.../MessageBubble.jsx` | `WebSourcesList` component β€” Tavily URLs as clickable green pill links |
| `.env.example` | Fixed β€” real keys had been committed; replaced with placeholders |

### Key discoveries
- `ConversationalRetrievalChain` condensation = silent context killer for web queries. Only fix: bypass the chain entirely for web path.
- Memory's `output_key="answer"` β€” `save_context` must use `{"answer": answer}` not `{"output": answer}` or KeyError.

---

## Stage 4 β€” OOM Hell on Render (2026-05-24, four sub-issues)

Render free tier: 512MB RAM. This stage was four separate OOM root causes discovered in sequence.

---

### 4a β€” CUDA torch OOM (startup crash)

**Problem:** `sentence-transformers` pulled CUDA torch (~2GB) by default. OOM before uvicorn bound to port β†’ Render showed "No open ports detected" timeout. Zero server stdout β€” invisible failure.

**Diagnosis clue:** Build log showed `cuda-toolkit-13.0.2`, `nvidia-cublas` being installed. Port scan timeout = uvicorn crash at import time (not lifespan β€” lifespan runs *after* port bind).

**Fix:**
```dockerfile
# Install CPU-only torch BEFORE requirements.txt
RUN pip install torch --index-url https://download.pytorch.org/whl/cpu
RUN pip install -r requirements.txt

ENV HF_HUB_OFFLINE=1
ENV TRANSFORMERS_OFFLINE=1
```

---

### 4b β€” ragas startup ImportError

**Problem:** `ragas 0.4.3` imports `langchain_community.chat_models.vertexai` at package `__init__` level. That module was removed in `langchain-community 0.4.x`. Crash propagated: `eval.py` β†’ `ragas_eval.py` β†’ `ragas.__init__` β†’ ImportError before uvicorn bound port.

**Fix:** All ragas imports moved inside `run_ragas_eval()` function body (lazy import).

---

### 4c β€” CrossEncoder OOM during chat

**Problem:** `CrossEncoder.predict(20 pairs)` = BERT forward pass on 20 pairs β†’ ~200–400MB spike on top of base ~250MB β†’ OOM on first web query.

**Fix:**
- `config.yaml`: `retrieve_k: 20 β†’ 10`
- `reranker.py`: `model.predict(pairs, batch_size=4)` β€” limits how many pairs processed at once

---

### 4d β€” Web search OOM (post-GC headroom)

**Problem:** After first query, Python retained chain/LLM objects at ~490MB. Web query added ~9KB Tavily content + `condense_question` LLM call + `run_query_with_web` LLM call β†’ OOM.

**Fix:**
- Tavily content truncated to 800 chars per result (was up to ~3000)
- `max_results`: 3 β†’ 2
- `gc.collect()` after each chat request in `routes/chat.py`

---

## Stage 5 β€” TinyBERT + RAGAS Removal + Benchmark JSON (2026-05-30)

### What was wrong

| Problem | Root cause |
|---------|-----------|
| MiniLM-L-6-v2 (~85MB) still OOMing on web queries | Too close to 512MB ceiling even after GC |
| Live RAGAS eval always 500 on Render | `nest_asyncio.apply()` (called at ragas import time) cannot patch `uvloop` β€” the event loop uvicorn uses on Linux. `ValueError: Can't patch loop of type uvloop.Loop`. Permanently unfixable without replacing uvicorn's event loop. |
| Groq 70B exhausted 100k daily tokens in one RAGAS run | RAGAS makes ~10 LLM calls per sample for statement decomposition. 10 samples Γ— 10 calls = 100k tokens gone. |
| context_precision and context_recall always null | `eval_pairs.json` had no `ground_truth` answers β€” only keyword hints |

### What we built

| Component | Change |
|-----------|--------|
| `server/reranker.py` | Switched to `cross-encoder/ms-marco-TinyBERT-L-2-v2` (~17MB vs 85MB). Saves 68MB permanently. |
| `server/routes/eval.py` | Removed `POST /api/eval/ragas` endpoint |
| `requirements.txt` | Removed `ragas` |
| `scripts/run_ragas_local.py` | Local RAGAS runner: ingest corpus β†’ generate answers β†’ run eval β†’ write JSON. Uses `llama-3.1-8b-instant` as judge (500k TPD vs 70B's 100k TPD) |
| `frontend/src/data/ragas_benchmark.json` | Static scores β€” Vercel builds dashboard from file |
| `frontend/.../EvalPanel.jsx` | Replaced live run button with static benchmark panel |
| `data/ground_truth/eval_pairs.json` | Added `ground_truth` field to all 20 pairs β†’ unlocked `context_precision` + `context_recall` |
| UI | i-button tooltips on faithfulness badge + all 4 RAGAS metric cards |

### Real scores committed
```
faithfulness:      1.0   (note: likely inflated β€” see below)
answer_relevancy:  0.90
context_precision: TBD   (pending fresh run)
context_recall:    TBD   (pending fresh run)
```

### Key discoveries
- `results["metric_name"]` returns `None` in ragas 0.2.x β€” must use `results.to_pandas()["metric_name"].mean()`
- TinyBERT loads with harmless `UNEXPECTED key bert.embeddings.position_ids` warning
- **Faithfulness 1.0 is likely inflated** β€” eval queries were designed alongside the corpus, and 8B judge is lenient. Scores are directional, not absolute. Run on held-out queries for honest numbers.

---

## Stage 5.5 β€” Rebranding: FinRAG β†’ Prism (2026-06-13)

### What changed
The project was originally named **FinRAG** β€” a fintech-specific RAG demo. As the architecture matured (multi-workspace, URL ingestion, domain-agnostic retrieval), it became clear the tool was no longer fintech-specific. Any corpus β€” legal, HR, medical, research β€” could be loaded and queried.

**Decision:** Rebrand to **Prism**. Name reflects the core idea: feed any document set in, get clear structured answers out. One engine, any domain.

| Before | After |
|--------|-------|
| FinRAG | Prism |
| Fintech-specific framing | Domain-agnostic positioning |
| `finrag-v2.onrender.com` | `prism.onrender.com` |
| README pitched at fintech analysts | README pitched at any knowledge-worker |

### What stayed the same
All retrieval architecture, eval framework, and deployment stack unchanged. Rebrand is naming and framing only β€” the engine is identical.

### What was wrong with the old name
- "FinRAG" implied fintech-only β†’ narrowed the demo audience
- Interviewers at non-fintech MNCs (Adobe, Atlassian, Intuit) would dismiss it as domain-locked
- The actual retrieval engine is domain-agnostic β€” the name should match

---

## Stage 6 β€” Multi-Workspace (2026-06 early)

### What was wrong
- Single ChromaDB collection β€” no isolation between document sets
- Switching topics meant re-ingesting and overwriting previous docs
- `list_collections()` broke on chromadb β‰₯0.5.4 (returns `list[str]`, not `list[Collection]`)
- Non-web chat path used stale global chain's `source_documents` instead of workspace-specific retriever β†’ wrong docs shown after workspace switch

### What we built

| File | Change |
|------|--------|
| `server/routes/workspaces.py` | Workspace CRUD β€” one ChromaDB collection per workspace |
| `server/routes/chat.py` | Always resolves workspace-specific retriever before branching on `web_search` |
| `server/routes/workspaces.py` | `list_collections()` normalised with `isinstance` check β€” works on chromadb β‰₯0.5.4 (`list[str]`) and <0.5 (`list[Collection]`) |
| `frontend/src/components/Sidebar.jsx` | Workspace switcher UI; per-workspace doc list |
| `frontend/src/App.jsx`, `api.js`, `ChatArea.jsx`, `FileUpload.jsx` | `workspace_id` passed on all requests |

### Key discovery
- Non-web path was relying on stale global chain's `source_documents` rather than workspace-specific retriever. After switching workspaces, the wrong collection's docs were being cited.

---

## Stage 7 β€” Singleton Cache + URL Guard (2026-06-14) ← Current

### What was wrong
- Every `POST /api/chat` called `get_or_create_collection()` + built a new `HybridRetriever` = full embedding reload per request β†’ OOM after 2–3 queries in the same workspace
- React component state (message list) persisted across workspace switch β€” showed previous workspace's chat history
- External URL ingestion had no size guard β†’ large pages (news articles, regulatory filings) caused OOM during embed

### What we built

| File | Change |
|------|--------|
| `server/retriever.py` | Module-level `Dict[workspace_id, (vectorstore, retriever)]` cache. Cache invalidated after ingest. `routes/chat.py` reuses cached retriever. |
| `server/routes/chat.py` | Eliminated double retrieval on non-web path; fixed stray print statement |
| `server/url_loader.py` | Max content size guard before embedding external URL content |
| `frontend/src/App.jsx` | `key={workspaceId}` on `<ChatArea>` β†’ remounts component on workspace switch β†’ clears stale messages and state |

### Commits
```
529675f  fix: singleton vectorstore/retriever cache to prevent OOM on repeated queries
6c9f809  fix: URL size guard for OOM prevention, eliminate double retrieval in chat
521a27a  fix: remount ChatArea on workspace switch to clear stale messages
```

### Additional fix (2026-06-16) β€” HyDE (Hypothetical Document Embeddings)

**What:** Before dense ChromaDB search, LLM generates a hypothetical 2-sentence answer. That answer (not raw query) is embedded for ANN search. BM25 + reranker still use original query.

| File | Change |
|------|--------|
| `server/retriever.py` | `_hyde_expand()` method; `use_hyde: bool` field on `HybridRetriever`; dense path uses expanded query when enabled |
| `config.yaml` | `retrieval.hyde_enabled: false` β€” toggle without code change |

**Why off by default:** Adds one Groq call per query (~200ms). Enable to measure RAGAS context_recall lift, then decide.

**Commit:** `8945b43`

---

### Additional fix (2026-06-17) β€” Mandatory web search

**Problem:** Web search was opt-in toggle. Users querying corpus-only got hallucinated answers from irrelevant documents (e.g., Singapore visa question grounded in random passport-mentioning corpus doc, faithfulness 4/5).

**Fix:**

| File | Change |
|------|--------|
| `frontend/src/components/ChatArea.jsx` | Removed toggle button; `const webSearch = true` hardcoded; placeholder always says "docs + web" |
| `server/routes/chat.py` | `web_search: bool = True` as default in `ChatRequest` |

Every query now hits Tavily + RAG corpus. `run_query_with_web` always called with both rag_docs + web_sources.

---

## Stage 8 β€” Eval Dashboard + Rigorous Metrics (2026-06-17)

### What was wrong
- faithfulness 1.0 and context_precision 1.0 artificially inflated β€” eval pairs designed alongside corpus, 8B judge lenient. Meaningless scores.
- Per-message faithfulness badge cluttered user UI. Users don't care about LLM judge scores.
- 10 samples β€” not statistically meaningful.
- Single flat JSON, no versioning β€” no way to track metric evolution across architecture changes.

### What we built

| Component | Change |
|-----------|--------|
| `eval-dashboard/` | Separate Vite + React static site (own Vercel project). Reads versioned JSON run files. |
| `eval-dashboard/src/components/` | MetricCard (score + delta vs prev), EvolutionChart (Recharts line chart across versions), RunTable (per-query expandable rows with answer vs ground_truth), LatencyStats (p50/p95 bars) |
| `eval-dashboard/public/data/index.json` | Run registry β€” list of all versioned eval runs |
| `scripts/run_eval_versioned.py` | New eval script. Args: `--version`, `--tag`, `--n`. Computes answer_correctness (LLM judge vs ground_truth), answer_relevancy + context_recall (RAGAS), precision@5, latency p50/p95/p99. Writes versioned JSON + updates index. |
| `data/ground_truth/eval_pairs.json` | Expanded 20 β†’ 50 pairs. Added multi-hop, comparative, negative, numeric, and edge-case questions. |
| `frontend/src/components/MessageBubble.jsx` | Removed FaithfulnessBadge component and rendering block. |
| `server/routes/chat.py` | Removed `score_faithfulness()` call. One fewer Groq API call per query β†’ faster responses. |

### Metrics before vs after

| Metric | Before | After |
|--------|--------|-------|
| faithfulness | 1.0 (inflated) | Removed from prod path |
| context_precision | 1.0 (inflated) | Replaced by answer_correctness (LLM judge vs ground_truth) |
| answer_relevancy | 0.88 | Kept (RAGAS) |
| context_recall | 0.83 | Kept (RAGAS) |
| precision@5 | tracked separately | Now in main eval dashboard |
| latency p50/p95 | not tracked | Now tracked per eval run |
| sample_count | 10 | 50 (5Γ— improvement) |

---

## Stage 9 β€” Multi-Query Retrieval (2026-06-19)

### What was wrong
- context_recall = 0.51 in v1.0.0 Violet β€” retriever missed ~half the relevant chunks
- Single-phrasing retrieval only surfaces chunks whose vocabulary matches the query tokens
- Chunks expressing same concept with different words (e.g. "PSP ceiling" vs "merchant limit") never entered the candidate pool

### What we built

| File | Change |
|------|--------|
| `server/retriever.py` | `_multi_query_expand()` β€” Groq LLM generates 3 phrasings (temperature=0.3). `_get_relevant_documents()` iterates all phrasings, deduplicates by content key keeping best rank, RRF fuses pooled results, reranks with original query. |
| `config.yaml` | `retrieval.multi_query_enabled: false` toggle |
| `docs/learning.md` | Concept 17 β€” Multi-Query Retrieval |

### Key design decisions
- **Deduplication keeps best rank** β€” a chunk at rank 1 in one phrasing and rank 8 in another enters RRF at rank 1, not 8
- **Reranker uses original query** β€” phrasings widen the pool; the reranker judges relevance against what the user actually asked
- **Off by default** β€” adds one Groq call per query (~200ms). Enable β†’ run v1.1.0 "Indigo" eval β†’ measure delta β†’ decide
- **retrieve_k cap maintained** β€” reranker input capped at `retrieve_k` even with wider pool, preserving RAM budget on Render

### Expected outcome
- context_recall: 0.51 β†’ measurably higher (target: >0.65)
- P@5: ~0.89 (no degradation expected β€” reranker filters noise from wider pool)
- Latency: +200–300ms per query (one extra Groq call for phrasing generation)
- Next eval run: `scripts/run_eval_versioned.py --version v1.1.0 --tag "Indigo" --n 50`

---

## Stage 10 β€” Contextual Retrieval (Eval) (2026-06-20)

### What was wrong
Phase 1 (HyDE + Multi-Query) left context_recall at ~0.51. Root cause confirmed: fixed-size 500-char splits produce decontextualized chunks. `"The limit was revised to β‚Ή2 lakh."` has no document name, no section, no subject β€” weak embedding that misses ~half relevant content. Query-side techniques cannot fix bad chunk quality.

### What we built

| File | Change |
|------|--------|
| `server/ingest.py` | `contextualize_chunks(chunks, documents, model, sleep_between_calls)` β€” calls Groq 8B per chunk, prepends 2-sentence situating context to `page_content` before embedding. Fallback to original text on any failure. |
| `config.yaml` | `contextual_retrieval.enabled: false`, `contextual_retrieval.model: llama-3.1-8b-instant` |
| `scripts/run_eval_versioned.py` | `--contextual` flag + `--data-dir` arg. When set: clears `eval_ctx` collection, re-ingests with `contextualize_chunks`, evaluates against `eval_ctx`. Production upload untouched. |
| `tests/test_ingest.py` | 3 tests: context prepended, fallback on failure, empty chunks skipped |

### Results β€” v1.3.0 "Violet"

| Metric | v1.0.0 baseline | v1.3.0 contextual | Delta |
|--------|-----------------|-------------------|-------|
| context_recall | 0.510 | 0.601 | **+9.1pp (+18%)** |
| precision_at_5 | 0.890 | 0.956 | **+6.6pp** |
| answer_relevancy | 0.620 | 0.633 | +1.3pp |
| answer_correctness | 0.820 | 0.815 | -0.5pp (noise) |
| latency p50 | 2029ms | 4161ms | **+2Γ— ⚠️** |
| latency p95 | β€” | 6122ms | β€” |

### Key discoveries
- Biggest single lift across all Phase 1+2 experiments: recall +9.1pp absolute
- Precision also improved significantly (0.890β†’0.956) β€” wider context gives reranker stronger signal
- Latency 2Γ— because contextualized chunks are longer (~150 extra tokens per chunk) β†’ LLM processes more tokens per answer generation call. Zero retrieval-time overhead (as designed), but query-time cost is real.
- Recall target was 0.65 β€” hit 0.60. Gap remains; next candidate is semantic chunking (Phase 2b)
- Production path: if shipping contextual retrieval, need FastAPI BackgroundTask for async contextualization at upload time (otherwise user waits 40s+ per doc upload)

---

## Stage 11 β€” Contextual Retrieval in Production + Dashboard Polish (2026-06-20)

### What was wrong
- Contextual retrieval proven in eval (recall +18%) but never shipped to production β€” users got non-contextual chunks
- eval-dashboard X-axis showed raw version strings (`v1.0.0`) with no dates
- No version badge visible in main Prism UI
- Encrypted PDFs caused 500 Internal Server Error instead of a clean user-facing message
- `max_concurrent=20` for parallel Groq calls β†’ 20k token burst β†’ 429 TPM limit on free tier (6000 TPM)
- Render free tier ephemeral filesystem: docs lost on every cold start (known limitation)
- Upload blocking for ~52s while Euron embedding API processes chunks sequentially

### What we built

| File | Change |
|------|--------|
| `server/ingest.py` | `contextualize_chunks_async()` β€” parallel Groq calls via `asyncio.gather` + `Semaphore(max_concurrent)`. ~10Γ— faster than sequential. Retry parses suggested wait time from 429 error message. |
| `server/routes/upload.py` | Two-phase upload: sync non-contextual embed first (user queryable immediately), then `_contextual_refresh_bg()` BackgroundTask replaces non-contextual chunks with contextual versions. |
| `config.yaml` | `contextual_retrieval.enabled: true`, `max_concurrent: 3` (3 Γ— ~1000 tokens = 3000 TPM β€” safe under 6000 limit) |
| `server/ingest.py` | `load_documents_from_paths()`: catches `FileNotDecryptedError` β†’ raises `ValueError` with user-friendly message |
| `server/routes/upload.py` | Catches `ValueError` from loader β†’ returns HTTP 422 instead of 500 |
| `eval-dashboard/src/components/EvolutionChart.jsx` | Custom `XAxisTick`: stacked version name + short date (e.g. `Violet (v1.3)` / `20 Jun 26`) |
| `eval-dashboard/src/App.jsx` | `VERSION_NOTES` constant with bullet notes per version; release notes panel shown below run meta |
| `frontend/src/components/Sidebar.jsx` | `Violet v1.3` badge (indigo pill) in sidebar footer |
| `frontend/src/config.js` | New file β€” `MAINTENANCE_MODE` + `MAINTENANCE_MESSAGE` config flags |
| `frontend/src/App.jsx` | Maintenance banner driven by `config.js`; hidden when `MAINTENANCE_MODE = false` |

### Key discoveries
- `asyncio.gather` with `Semaphore(3)` keeps burst under 3000 TPM β€” safe on Groq free tier (6000 TPM limit)
- Groq 429 errors include `"Please try again in X.Xs"` β€” parse this for accurate retry sleep instead of hardcoded 2s
- Render free tier: ephemeral filesystem. Every cold start wipes `./chroma_db`. Docs must be re-uploaded. Fix: Render persistent disk ($0.25/GB/month)
- Euron embedding API sequential calls: 30 chunks Γ— ~1.7s/call = ~52s blocking upload. Next optimization: move embed to background too (return 202 immediately, notify when ready)
- `max_concurrent=20` was the OOM trigger in the previous session β€” 20 async coroutines each holding ~10MB response + retry state saturated 512MB

---

## Stage 15 β€” Semantic Chunking Ablation + Retrieval Stack Finalized (2026-06-26)

### What was wrong
Ablation study incomplete β€” semantic chunking (v1.4.0) was blocked by Groq rate limits in the prior session. Best production stack unconfirmed.

### What we built / ran

| Version | Config | recall | P@5 | relevancy | correctness | p50 |
|---------|--------|--------|-----|-----------|-------------|-----|
| v1.1.0 | HyDE | 0.721 | 0.911 | 0.845 | 0.750 | 4018ms |
| v1.2.0 | HyDE+MQ | 0.645 | 0.904 | 0.890 | 0.770 | 1812ms |
| v1.3.0 | HyDE+MQ+CTX | 0.768 | **0.984** | 0.799 | 0.780 | 2610ms |
| v1.4.0 | HyDE+MQ+CTX+Semantic | **0.861** | 0.711 | 0.885 | 0.750 | 12952ms |

### Decision: semantic chunking rejected

Semantic chunking raises recall +9.3pp (0.768→0.861) but P@5 collapses -27.3pp (0.984→0.711) and latency is 5× worse (2610ms→12952ms p50).

**Root cause of P@5 collapse:** SemanticChunker produces variable-size, topic-boundary chunks. These don't align with the fixed ground-truth keyword spans used for precision@5 scoring. The reranker receives a wider but noisier candidate pool β€” recall expands while precision degrades.

**Best stack confirmed: v1.3.0 β€” HyDE + Multi-Query + Contextual Retrieval.**

### Key discoveries
- MQ alone hurts recall (-7.6pp vs HyDE-only) but recovers fully when combined with CTX
- CTX is highest-leverage single addition: +8pp P@5, recall recovery, at 2Γ— query latency cost
- Semantic chunking is a double-edged sword β€” better chunk boundaries for recall, worse alignment with precision evaluation
- Ablation study is the interview story: systematic metric-driven elimination of techniques

---

## Stage 16 β€” Citation Highlighting (2026-06-26)

### What was wrong
Sources listed below each answer as truncated 200-char snippets. LLM already outputs `[1]`, `[2]` inline citations but they rendered as plain unclickable text. Users couldn't see which passage in the answer corresponded to which source.

### What we built

| File | Change |
|------|--------|
| `frontend/src/components/CitationPopover.jsx` | New β€” viewport-aware popover (fixed-position). Shows: source type badge (pdf/web/file), filename/title, page badge, full chunk content (scrollable), rerank score, "Open page N β†’" for PDF / "Open source β†’" for web |
| `frontend/src/components/MessageBubble.jsx` | Parse `[N]` markers in answer text β†’ clickable `<sup>` superscripts. `openCitation` state (`{ idx, rect } \| null`). Toggle on same click. `onMouseDown` stopPropagation fix (prevents document mousedown from immediately re-opening after close). |
| `frontend/src/components/SourceExpander.jsx` | Removed 200-char content truncation β€” full chunk text shown |
| `server/main.py` | Added `GET /api/files/{filename}` β†’ `FileResponse` from `data/raw/`. Path traversal blocked via `is_relative_to()`. `UPLOAD_DIR` made absolute (`Path(__file__).resolve().parent.parent / "data" / "raw"`). |
| `server/routes/upload.py` | `UPLOAD_DIR` made absolute (`Path(__file__).resolve().parent.parent.parent / "data" / "raw"`) |

### Key discoveries
- `mousedown` on document fires before `click` β€” without `e.stopPropagation()` on the `<sup>` mousedown, clicking an open citation closes then immediately reopens it (toggle broken)
- `startswith()` on raw path strings has prefix-confusion bug (`/data/rawevil` passes `/data/raw` check) β€” replaced with `Path.is_relative_to()` (Python 3.9+)
- Relative `Path("data/raw")` resolves against process CWD β€” if uvicorn starts from non-project-root directory, file serving breaks. Absolute `__file__`-relative path fixes this.
- `anchorRect` captured at click time via `el.getBoundingClientRect()` β€” stored in state as plain object, no ref needed in popover

### Interview story
> "The LLM cites [1], [2] in its answer. Clicking one opens a popover showing the exact passage retrieved β€” full text, source file, page number, and rerank score. For PDFs it links directly to that page in the browser."

---

## Stage 17 β€” Metadata Filtering (2026-06-27)

### What was wrong
All documents in a workspace were always searched together. A user with 10 docs spanning 5 years had no way to scope a query to a specific doc or subset. Corpus-wide retrieval diluted precision when the relevant content was known to be in one file.

### What we built

| File | Change |
|------|--------|
| `server/ingest.py` | `source_type` metadata field (`pdf`/`txt`/`csv`) added to all chunks at load time via `SOURCE_TYPE_MAP`. Both `load_documents` and `load_documents_from_paths` patched. |
| `server/url_loader.py` | `source_type: "url"` added to URL-ingested doc metadata. |
| `server/bm25_index.py` | `BM25Index.search()` gets `filter_sources: set[str] \| None = None`. When set, scores using full-corpus BM25 index (stable IDF) but restricts candidate pool to matching docs. |
| `server/retriever.py` | `filter_docs: list[str] \| None = None` field on `HybridRetriever`. Wired into `_dense_retrieve` (ChromaDB `where={"source": {"$in": filter_docs}}`) and `_get_relevant_documents` (BM25 `filter_sources`). New `get_retriever_filtered(workspace_id, filter_docs)` helper β€” one-off instance reusing cached vectorstore, not added to singleton cache. |
| `server/routes/chat.py` | `filter_docs: list[str] \| None = None` on `ChatRequest`. Guard: empty list β†’ None. When truthy: `get_retriever_filtered(workspace, active_filter)`. Log includes `filter=%s`. |
| `frontend/src/api.js` | `streamChat` gets `filterDocs = null` as 4th arg; sends `filter_docs: filterDocs?.length ? filterDocs : null`. |
| `frontend/src/App.jsx` | `filterDocs: string[]` state. `useEffect` resets to `[]` on workspace change. `handleFilterChange` toggles doc in/out. Props forwarded to Sidebar + ChatArea. |
| `frontend/src/components/Sidebar.jsx` | Doc list items clickable β€” toggle filter on click. Selected: `ring-2 ring-indigo-500 bg-indigo-50`. Unselected during active filter: `opacity-50`. "Clear filter" button in section header when any selected. Delete button: `e.stopPropagation()` + deselects deleted doc from filter. |
| `frontend/src/components/ChatArea.jsx` | Filter badge above input when `filterDocs.length > 0` (shows scoped doc names + Γ— clear). Placeholder: "Searching N selected doc(s)..." when filter active. `streamChat` called with `filterDocs.length > 0 ? filterDocs : null`. |
| `tests/test_bm25_filter.py` | 6 tests: no-filter returns all, filter restricts by source, empty set returns empty, nonexistent source returns empty, multiple sources, unbuilt index returns empty. |
| `tests/test_source_type.py` | 4 tests: pdf/txt/csv/url each gets correct `source_type`. |

### Key design decisions
- **Full-corpus BM25 for filtering**: spec suggested rebuilding BM25 on filtered subset; implementation uses full-corpus index to score + restricts candidate pool by source. Stable IDF β€” correct IR semantics. Accepted as superior to spec.
- **New retriever instance per filtered request**: `get_retriever_filtered()` creates a one-off `HybridRetriever`; singleton cache (`_retriever_cache`) untouched. Thread-safe: heavy vectorstore stays cached, lightweight retriever is cheap.
- **Empty filter = no filter**: backend guard `body.filter_docs if body.filter_docs else None` β€” empty array from frontend treated as no filter.
- **Filter resets on workspace switch**: `useEffect(() => setFilterDocs([]), [currentWorkspace])` β€” stale filter from workspace A doesn't carry to workspace B.
- **Delete deselects**: `handleDelete` calls `onFilterChange(docName)` if deleted doc was selected β€” prevents badge showing "Scoped to: [deleted]" with zero results.

### Interview story
> "Within a workspace, users can click any doc chip in the sidebar to scope retrieval. Dense retrieval passes `where={"source": {"$in": selected_docs}}` to ChromaDB; BM25 pre-filters its candidate pool. Zero selection = full-corpus behavior unchanged. Filter badge above the input makes the scope visible."

---

## Stage 18 β€” Free-Tier Stability + App Restored to Live (2026-07-05)

### What was wrong
- Maintenance banner left ON after Cerebras migration attempt (2026-06-29) failed and was reverted
- Config.yaml still had `hyde_enabled: true` + `multi_query_enabled: true` β€” each query burned 4 Groq calls
- Free tier limit: 6000 TPM β†’ 429 storms under concurrent use with HyDE + MQ + contextual all on
- Eval dashboard had no indication which version is live or why best stack (v1.3.0) isn't deployed

### What we built

| File | Change |
|------|--------|
| `config.yaml` | `hyde_enabled: false`, `multi_query_enabled: false` β€” reduces query-time Groq calls 4 β†’ 1–2 |
| `frontend/src/config.js` | `MAINTENANCE_MODE: false` β€” app live |
| `eval-dashboard/public/data/index.json` | `is_live: true` + `live_note` on v1.1.0 (closest proxy); `blocked_by` constraint on v1.3.0 + v1.4.0 |
| `eval-dashboard/src/App.jsx` | Green LIVE badge + prod config note on v1.1.0; amber "not in production" warning on v1.3.0/v1.4.0 |

### Key design decisions
- **Contextual retrieval kept ON** β€” uses `openai/gpt-oss-20b` via Euron API, zero Groq TPM impact at query time. Ingest-time only.
- **HyDE + MQ disabled, not removed** β€” toggles in config.yaml; re-enable instantly when on paid tier
- **v1.1.0 as live proxy in eval dashboard** β€” no eval run exists for "CTX-only, no HyDE, no MQ" config. v1.1.0 (recall=0.721) is an overestimate; actual live recall β‰ˆ 0.55–0.65 given contextual index without HyDE query expansion
- **Upgrade path documented in eval dashboard** β€” v1.3.0 blocked_by note explains exactly what to fix

### Groq call budget (current vs best)

| Config | Calls/query | TPM risk |
|--------|------------|----------|
| Current (CTX only) | 1–2 | Safe |
| v1.3.0 (HyDE+MQ+CTX) | 4 | 429 on free tier |

### Upgrade path to v1.3.0
1. Switch to paid Groq tier (or find higher-TPM free provider)
2. Set `hyde_enabled: true` + `multi_query_enabled: true` in `config.yaml`
3. Push β†’ HF Spaces rebuilds β†’ run `scripts/run_eval_versioned.py --version v1.3.1 --tag "Violet" --n 50` to confirm metrics

---

## Current State Snapshot

```
Retrieval:    Hybrid BM25 (0.3) + ChromaDB dense (0.7) β†’ RRF β†’ TinyBERT rerank top-10β†’5
LLM:          Groq llama-3.3-70b-versatile
Embeddings:   Euron API text-embedding-3-small (sequential, ~1.7s/chunk β€” bottleneck)
Chunking:     RecursiveCharacterTextSplitter 500-char, overlap 50
Memory:       ConversationBufferWindowMemory k=10
Web search:   Tavily advanced, 800-char truncation, max 2 results β€” MANDATORY (always on)
HyDE:         DISABLED (hyde_enabled=false). Best measured: +21pp recall but costs 1 Groq call/query.
              Re-enable when on paid Groq tier or higher-TPM provider.
Multi-Query:  DISABLED (multi_query_enabled=false). Costs 1 Groq call/query β€” free tier cannot sustain.
              Re-enable with HyDE together (v1.3.0 config) on paid tier.
Contextual:   ENABLED (contextual_retrieval.enabled=true). Uses Euron model (openai/gpt-oss-20b) β€”
              zero Groq TPM impact. Two-phase upload: sync non-contextual embed (<3s queryable),
              BackgroundTask replaces with contextual chunks. max_concurrent=3, max_chunks=50 gate.
Semantic:     DISABLED (semantic_enabled=false). Ablation showed recall +9.3pp but P@5 -27.3pp and 5Γ— latency.
              Rejected β€” v1.3.0 (HyDE+MQ+CTX) is the confirmed best stack when TPM allows.
Groq calls/query (current): 1–2 (condense_question if follow-up + answer). Safe under 6000 TPM free tier.
Groq calls/query (v1.3.0):  4 (condense + HyDE + Multi-Query + answer) β†’ 429 storms on free tier.
Eval:         Separate eval-dashboard/ static site β†’ https://askprism-eval.vercel.app/
              v1.1.0 marked LIVE (closest proxy). v1.3.0 and v1.4.0 show amber "not in production" warning.
              Best measured: v1.3.0 recall=0.768, P@5=0.984, p50=2610ms
              Versioning: MAJOR.MINOR.PATCH β€” name changes on MAJOR only (v1.x.x=Violet, v2.x.x=Indigo)
Citation:     [N] markers in LLM answers β†’ clickable <sup> β†’ CitationPopover (fixed-position, viewport-aware).
              Shows full chunk text, source name, page, rerank score. PDF: "Open page N β†’" link via GET /api/files/{filename}.
              Web: "Open source β†’". Toggle, click-away, above/below flip at 60% viewport height.
              SourceExpander: full content shown (200-char truncation removed).
Frontend:     Violet v1.3 badge in sidebar footer. Maintenance banner config-driven (frontend/src/config.js).
              MAINTENANCE_MODE=false β€” app is live as of 2026-07-05.
Filter:       Sidebar doc chips toggleable. Selected: indigo ring. Badge above chat input shows scoped docs + clear Γ—.
              POST /api/chat accepts filter_docs: string[] | null. Empty = no filter. Resets on workspace switch.
              Backend: get_retriever_filtered() creates one-off HybridRetriever; singleton cache untouched.
              ChromaDB where={"source": {"$in": filter_docs}}. BM25 filters candidate pool, scores with full-corpus IDF.
Workspaces:   Per-workspace ChromaDB collection, singleton retriever cache
Infra:        HF Spaces CPU Basic (backend, 16GB RAM, ephemeral FS β€” re-upload required after cold start) +
              https://askprism.vercel.app/ (frontend) + https://askprism-eval.vercel.app/ (eval)
              Backend URL: https://benroshan-prism.hf.space
Known limits: Euron embed ~5s/chunk sequential β€” 30 chunks = ~150s total contextualization in background.
              HF Spaces ephemeral FS: chroma_db lost on cold start. Fix: mount HF persistent storage bucket.
              HyDE + MQ disabled for free-tier stability. Best stack (v1.3.0) needs paid Groq or alt provider.
Observability: LangSmith traces all LLM + retrieval calls (optional, env var)
Streaming:    POST /api/chat returns SSE stream. token events per LLM chunk, done event with
              sources + retrieval_method. Frontend streams tokens into pre-placed assistant
              bubble. Bouncing dots while condense+search runs, blinking cursor during generation.
```

---

## Stage 12 β€” HF Spaces Migration (2026-06-22)

### What was wrong
Render free tier (512MB RAM) caused repeated OOM crashes under contextual retrieval:
- Base RSS after upload = 524MB (over the 512MB limit)
- `gc.collect()` had no effect β€” ChromaDB HNSW index + torch runtime held by native allocators, not Python heap
- Contextual refresh (3 async Groq coroutines) + simultaneous chat (Tavily + LLM + CrossEncoder) = peak exceeded 512MB
- Workarounds (RSS guard skipping contextual retrieval, web search suppression during refresh) negated the +18% recall improvement

### What we built

| File | Change |
|------|--------|
| `Dockerfile` | Port 8000 β†’ 7860 (HF convention). Add `useradd -m -u 1000 user` + `chown -R user /app` (HF runs containers as UID 1000). Set `HF_HOME=/app/.cache/huggingface` BEFORE pre-download so user 1000 owns cached weights. Set `HF_HUB_OFFLINE=1` AFTER download. |
| `README.md` | Added HF Spaces frontmatter (`sdk: docker`, `app_port: 7860`). Updated deploy instructions. |
| `server/routes/chat.py` | Removed `is_contextualizing` web search suppression guard (Render-specific). |
| `server/routes/upload.py` | Removed `RSS > 460MB` contextual retrieval skip guard (Render-specific). |
| `docs/`, `decisions.md` | Render β†’ HF Spaces across all infra references. |

### Key discoveries
- HF_HUB_OFFLINE must be set AFTER the pre-download RUN step β€” setting it before blocks the download itself
- Docker build runs pre-download as root by default; must `USER 1000` first then set `HF_HOME` under `/app` so runtime user 1000 can read the cached weights
- HF Spaces free CPU Basic: 2 vCPUs, 16GB RAM β€” resolves all Render OOM issues permanently
- Contextual retrieval now runs fully in production (was silently skipped by RSS guard on Render)

---

## Stage 13 β€” Async Embed Upload (2026-06-23)

### What was wrong
`embed_and_store()` blocked `POST /api/upload` for ~150s (30 chunks Γ— ~5s/chunk via Euron API). User saw spinner, could not query, could not cancel. Upload timeout was 300s.

### What we built

| File | Change |
|------|--------|
| `server/main.py` | `app.state.upload_jobs = {}` initialized in lifespan |
| `server/routes/upload.py` | `POST /api/upload` returns 202 + `job_id` in <1s. Parse+chunk sync; embed+contextual in `_embed_and_contextualize_bg()` BackgroundTask. New `GET /api/upload/status/{job_id}` endpoint. |
| `frontend/src/api.js` | Added `getUploadStatus(jobId)`; reduced `uploadFiles` timeout 300s β†’ 30s |
| `frontend/src/components/FileUpload.jsx` | Polls status every 2s; shows stage label under spinner; fires callbacks on ready. Defensive `|| []` guard on documents. |

### Key discoveries
- Old Vercel frontend receiving new 202 response before redeploy β†’ `data.documents` undefined β†’ React crash. Fix: defensive `docs?.documents || []` guard.
- Groq TPM 429s at `max_concurrent=3` still hit (~5/30 chunks fall back to original text) β€” some chunks are larger than average. Retry logic handles gracefully.
- Briefing fails with JSON parse error (pre-existing bug in `generate_briefing` β€” separate fix).

---

## Stage 14 β€” Briefing Fix + HyDE Re-eval (2026-06-24)

### What was wrong
- `generate_briefing()` crashed with `JSONDecodeError` when Groq LLM returned control characters (ASCII 0x00–0x1f) or Python dict syntax (single quotes) instead of valid JSON.
- Old eval runs (v1.0.0–v1.4.0) accumulated across multiple sessions; stale runs cluttered the dashboard.
- HyDE recall measurement from prior session (v1.1.0_20260619, recall=0.545) was based on 50-sample run that hit Groq 429s mid-run β€” partial results, unreliable numbers.

### What we built

| File | Change |
|------|--------|
| `server/briefing.py` | Strip control chars `[\x00-\x08\x0b\x0c\x0e-\x1f]` before JSON parse. Fall back to `ast.literal_eval()` on `JSONDecodeError` to handle Python dict syntax from LLM. Added `import ast`. |
| `config.yaml` | `hyde_enabled: true`, `multi_query_enabled: true`, `contextual_retrieval.enabled: false` (contextual off β€” 429s at 30-chunk scale even with Semaphore(3)) |
| `eval-dashboard/public/data/runs/` | Deleted stale runs (v1.0.0_20260618, v1.1.0_20260619, v1.2.0_20260619, v1.3.0_20260620, v1.3.0_20260623, v1.4.0_20260623). Added `v1.1.0_20260624.json` β€” fresh HyDE-only run. |
| `eval-dashboard/public/data/index.json` | Updated to single clean run registry. |

### HyDE re-eval results β€” v1.1.0_20260624 (18 samples, hyde=true, multi_query=false)

| Metric | v1.0.0 baseline | v1.1.0 HyDE | Delta |
|--------|-----------------|-------------|-------|
| answer_correctness | 0.820 | 0.750 | -0.070 |
| answer_relevancy | 0.620 | 0.845 | **+0.225** |
| context_recall | 0.510 | 0.721 | **+0.211** |
| precision_at_5 | 0.890 | 0.911 | +0.021 |
| latency p50 | 2029ms | 4018ms | +2Γ— |

### Key discoveries
- HyDE gives **+21pp recall** (0.51β†’0.72) on this 18-sample run β€” much larger than previously measured (+3.5pp on 50 samples with 429s). Smaller sample set; repeat at 50 samples to confirm.
- answer_correctness flat at 0.75 for all 18 samples β€” 8B judge giving uniform score, not differentiating. May indicate judge calibration issue, not actual correctness plateau.
- Latency 2× (2029ms→4018ms) — HyDE adds one Groq call per query for hypothetical expansion.
- Briefing fix unblocks document upload β†’ briefing flow end-to-end.

---

## Roadmap β€” Retrieval & Answer Quality

### Phase 1 β€” Quick wins (no infra change, measurable RAGAS lift)

#### ~~HyDE (Hypothetical Document Embeddings)~~ βœ… Done (Stage 7, commit 8945b43)
- Implemented in `server/retriever.py`. Toggle: `config.yaml hyde_enabled` (default: false).
- Enable + re-run eval to measure context_recall lift vs v2.0 baseline (0.70).

#### ~~Multi-Query Retrieval~~ βœ… Done (Stage 9, 2026-06-19)
- Implemented in `server/retriever.py`. Toggle: `config.yaml multi_query_enabled` (default: false).
- Enable + run `scripts/run_eval_versioned.py --version v1.1.0 --tag "Indigo" --n 50` to measure context_recall lift vs 0.51.

---

### Phase 2 β€” Ingest pipeline (requires re-ingest of all docs)

#### ~~Contextual Retrieval~~ βœ… Done + Shipped to Production (Stage 10+11, 2026-06-20)
- `contextualize_chunks_async()` + BackgroundTask in `routes/upload.py`. Two-phase: sync non-contextual embed (queryable <3s) β†’ background contextual replacement.
- v1.3.0 results: recall 0.510β†’0.601 (+18%), P@5 0.890β†’0.956. Latency 2Γ— at query time (longer chunks β†’ more LLM tokens).
- `max_concurrent=3` in `config.yaml` β€” safe under Groq 6000 TPM limit.

#### Semantic Chunking
- **Problem:** Fixed 200-char splits cut mid-sentence, mid-table, mid-list. Embedding a truncated sentence returns a weak vector.
- **How:** Replace `RecursiveCharacterTextSplitter` with LangChain's `SemanticChunker` β€” splits at sentence boundaries where cosine similarity between adjacent sentences drops below a threshold (topic shift).
- **Effort:** Medium. Config change in `ingest.py` + re-ingest. Tune `breakpoint_threshold_type`.
- **Expected lift:** Fewer nonsensical chunks in top-5. Most noticeable on regulatory PDFs with section headers and numbered lists.

---

### Phase 3 β€” UX + trust

#### Streaming Responses
- **Problem:** User submits question β†’ 8–15s wait β†’ full answer appears. Feels broken even on fast hardware.
- **How:** Backend: `chain.astream_events()` β†’ `StreamingResponse` yielding SSE tokens. Frontend: `EventSource` or `fetch` + `ReadableStream` β€” append tokens as they arrive. Faithfulness scoring runs as background task after full answer assembled.
- **Effort:** High β€” both backend and frontend change. `ConversationalRetrievalChain` supports `astream_events()` in LangChain β‰₯0.2.
- **Impact:** Perceived latency drops from 10s to ~1s. Single biggest UX improvement.

#### ~~Citation Highlighting~~ βœ… Done (Stage 16, 2026-06-26)
- `[N]` markers clickable β†’ `CitationPopover` with full chunk text, page badge, rerank score. PDF "Open page N β†’" link. Zero new npm deps.
- Works for all source types: PDF, URL, TXT, CSV. No PDF viewer library needed β€” page link uses browser's built-in viewer.

---

### Phase 4 β€” Differentiation

#### Metadata Filtering
- **Problem:** Multi-workspace isolates by collection, but within a workspace (10 docs across 5 years) no way to scope retrieval to `year=2024` or `doc_type=rbi_circular`.
- **How:** Tag chunks with `{source_type, year, doc_name}` at ingest. Pass optional `filter` param in `/api/chat` request. ChromaDB `where` clause on dense retrieval; BM25 pre-filters corpus to matching chunk IDs.
- **Impact:** Precision boost on time-scoped or source-scoped queries.

#### Document Comparison Mode
- **Problem:** No way to ask "What changed between RBI circular 2023 and 2024?"
- **How:** Frontend sends two doc IDs + comparison query. Backend retrieves relevant chunks from each collection separately, synthesises a structured diff answer.
- **Impact:** Killer fintech feature. Unique demo moment. Differentiates from generic RAG.

#### Agentic Mode (LangGraph)
- **Problem:** Single-shot RAG cannot handle multi-step reasoning: retrieve β†’ compute β†’ web search β†’ synthesise.
- **How:** Replace `ConversationalRetrievalChain` with a LangGraph graph. Nodes: retriever, web_search, calculator, synthesiser. LLM decides which tool to call.
- **Impact:** Separates Prism from basic RAG β€” becomes a research agent. Strongest interview story.

---

## Roadmap Priority Matrix

```
HIGH impact Γ— LOW effort  β†’ Build first
  HyDE
  Multi-query retrieval
  Metadata filtering

HIGH impact Γ— MEDIUM effort β†’ Build second
  Contextual retrieval (+ re-ingest)
  Semantic chunking (+ re-ingest)
  Streaming responses

HIGH impact Γ— HIGH effort β†’ Build last
  Citation highlighting
  Document comparison
  Agentic mode (LangGraph)
```

---

## Interview Story Arc

```
v1  β†’ Dense-only retrieval. No eval. No baseline.
v2  β†’ Hybrid BM25+dense, cross-encoder rerank. Measured with RAGAS.
     β†’ faithfulness=1.0, answer_relevancy=0.90 on 20-pair eval set.
+HyDE β†’ context_recall 0.51β†’0.72 (+21pp). Hypothetical answer embedding closes vocabulary gap.
+Contextual β†’ context_recall 0.60 (+18% vs baseline). Ingest-time LLM chunk augmentation.
+Agentic  β†’ Multi-step reasoning. Not RAG anymore β€” research agent.
```

Each step has a metric. That is the complete RAG engineering narrative for MNC DS interviews.