--- title: RAG Document Q&A emoji: πŸ“„ colorFrom: indigo colorTo: blue sdk: docker app_port: 7860 pinned: false --- # RAG Document Q&A System Upload PDF documents and get cited, grounded answers β€” with hybrid retrieval, cross-encoder re-ranking, and confidence-aware generation. ## Overview This project implements a production-grade Retrieval-Augmented Generation pipeline that goes beyond the typical "embed + nearest-neighbor + prompt" tutorial. It combines semantic and keyword search with cross-encoder re-ranking, applies Anthropic's contextual retrieval pattern to enrich embeddings with document metadata, and gates LLM calls on retrieval confidence to avoid hallucinated answers. The system ships with both a streaming Gradio UI and a FastAPI backend with Server-Sent Events, plus an evaluation harness that benchmarks chunking strategies and scores end-to-end answer quality using an LLM-as-judge. ## Architecture ``` β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚ INGESTION β”‚ β”‚ β”‚ PDF ──► pymupdf4llm ──► β”‚ Chunking (3 strategies) β”‚ (pdfplumber β”‚ β”œβ”€ fixed_size sliding window β”‚ fallback) β”‚ β”œβ”€ recursive_char LangChain splits β”‚ β”‚ └─ semantic embedding sim β”‚ β”‚ β”‚ β”‚ β”‚ β–Ό β”‚ β”‚ Contextual Enrichment β”‚ β”‚ "filename | Page N | Section\ntext" β”‚ β”‚ β”‚ β”‚ β”‚ β–Ό β”‚ β”‚ all-MiniLM-L6-v2 (384-dim, L2-norm) β”‚ β”‚ β”‚ β”‚ β”‚ β”Œβ”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β” β”‚ β”‚ β–Ό β–Ό β”‚ β”‚ FAISS BM25Okapi β”‚ β”‚ IndexFlatIP rank-bm25 β”‚ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚ QUERY β”‚ β”‚ β”‚ Question ──────────────►│ Adaptive Expansion (Groq) β”‚ β”‚ only if top_score < 0.45 β”‚ β”‚ β”‚ β”‚ β”‚ β–Ό β”‚ β”‚ Hybrid Search β”‚ β”‚ score = 0.7Β·semantic + 0.3Β·bm25 β”‚ β”‚ (top 20 candidates) β”‚ β”‚ β”‚ β”‚ β”‚ β–Ό β”‚ β”‚ Cross-Encoder Re-ranking β”‚ β”‚ ms-marco-MiniLM-L-6-v2 β”‚ β”‚ (20 β†’ top k) β”‚ β”‚ β”‚ β”‚ β”‚ β–Ό β”‚ β”‚ Confidence Gating β”‚ β”‚ cosine sim < 0.3 β†’ skip LLM β”‚ β”‚ β”‚ β”‚ β”‚ β–Ό β”‚ β”‚ Groq (Llama 3.3 70B) streaming β”‚ β”‚ β”‚ β”‚ β”‚ β–Ό β”‚ β”‚ Cited answer [Source N] notation β”‚ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚ EVALUATION β”‚ β”‚ β”‚ β”‚ Retrieval: Acc@1/3/5, MRR β”‚ β”‚ End-to-end: faithfulness + relevance β”‚ β”‚ (LLM-as-judge, temp=0.0) β”‚ β”‚ Config sweep: chunking strategy grid β”‚ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ ``` ## Key Features - **Cross-encoder re-ranking** β€” bi-encoder retrieval (FAISS) produces 20 candidates; a cross-encoder (`ms-marco-MiniLM-L-6-v2`) scores each `(query, chunk)` pair jointly and re-orders them. Cross-encoders are too slow for full-corpus search but dramatically improve precision over the shortlist. - **Hybrid search** β€” FAISS inner-product search and BM25 run in parallel. Both score sets are normalized to `[0, 1]` then fused (`0.7Β·semantic + 0.3Β·keyword`), catching exact-match terms that dense embeddings tend to smooth over. - **Adaptive query expansion** β€” the pipeline runs an initial search first. Only if the top cosine similarity falls below 0.45 does it call the LLM to generate 3 rephrased variants, search each, and merge results. High-confidence queries pay zero expansion cost. - **Contextual chunk enrichment** β€” each chunk is embedded as `" | Page N |
\n"` rather than raw text, following Anthropic's contextual retrieval pattern. The embedding captures document location and section context, not just lexical content. - **Confidence-based response gating** β€” `max_score` is taken from cosine similarity *before* re-ranking (cross-encoder scores on a βˆ’10 to +10 scale would break the threshold). If `max_score < 0.3`, the LLM is skipped entirely and the user receives an honest "insufficient context" message. Scores between 0.3–0.5 append a low-confidence warning. - **Streaming responses** β€” Gradio UI streams tokens via a generator; the FastAPI backend streams as Server-Sent Events (`text/event-stream`) with `{"text": chunk}` events and a terminal `{"done": true}`. - **Layout-aware PDF parsing** β€” `pymupdf4llm` extracts structured markdown with headings and tables preserved. `pdfplumber` serves as a fallback for PDFs that `pymupdf4llm` cannot parse cleanly. - **End-to-end evaluation** β€” the eval harness measures retrieval hit rate (Acc@K, MRR), then pipes results through the LLM-as-judge to score faithfulness (are claims grounded in context?) and relevance (does the answer address the question?) at temperature 0 for determinism. ## Tech Stack | Component | Technology | |---|---| | Embedding model | `all-MiniLM-L6-v2` (sentence-transformers, 384-dim) | | Vector index | FAISS `IndexFlatIP` (exact inner-product search) | | Keyword index | BM25 (`rank-bm25`, `BM25Okapi`) | | Re-ranker | `cross-encoder/ms-marco-MiniLM-L-6-v2` | | LLM | Llama 3.3 70B via Groq API | | PDF parsing | `pymupdf4llm` + `pdfplumber` fallback | | Text splitting | LangChain `RecursiveCharacterTextSplitter` | | Backend | FastAPI + Uvicorn | | Frontend | Gradio (Blocks, streaming) | | Containerization | Docker | | Cloud deployment | HuggingFace Spaces | ## Quick Start ```bash git clone cd rag-qa python -m venv venv && source venv/bin/activate pip install -r requirements.txt ``` Create a `.env` file: ``` GROQ_API_KEY=your_key_here ``` Get a free key at [console.groq.com](https://console.groq.com) β€” the free tier allows 14,400 requests/day. Launch the Gradio UI: ```bash python app_gradio.py # Opens at http://127.0.0.1:7860 ``` Or run the FastAPI server: ```bash uvicorn main:app --reload # API docs at http://127.0.0.1:8000/docs ``` ## API Endpoints | Method | Endpoint | Description | |---|---|---| | `GET` | `/health` | Liveness check | | `GET` | `/stats` | Index size, chunks ingested, model info | | `POST` | `/ingest` | Upload a PDF; returns chunk count and IDs | | `POST` | `/query` | Ask a question; returns answer + sources + confidence | | `POST` | `/query/stream` | Same as `/query` but streams as SSE | | `POST` | `/debug/chunks` | Return raw retrieved chunks without calling the LLM | All `POST` query endpoints accept `{"question": "...", "top_k": 5}`. The `/debug/chunks` endpoint is useful for diagnosing retrieval issues β€” it shows each chunk's rank, cosine similarity score, source, page, section header, and a 300-character text preview. ## Evaluation The evaluation framework has three layers: **Retrieval metrics** (`evaluation/eval.py`) β€” given a test set of `(question, expected_source, expected_page)` tuples, measures whether the correct chunk appears in the top-k results: | Metric | Definition | |---|---| | Accuracy@K | Fraction of queries where correct chunk is in top K | | MRR | Mean Reciprocal Rank of the first correct result | **End-to-end metrics** (`evaluation/eval_e2e.py`) β€” runs the full pipeline and scores each answer: | Metric | Scorer | |---|---| | Faithfulness | LLM judge: are all claims grounded in retrieved context? | | Relevance | LLM judge: does the answer address the question? | | Keyword coverage | Fraction of expected answer keywords found in output | **Chunking strategy comparison** (`evaluation/compare.py`) β€” grid search over chunk sizes, overlap values, and re-ranking on/off. Results table (fill in with your benchmark numbers): | Strategy | Chunk size | Overlap | Reranking | Acc@5 | MRR | Faithfulness | Relevance | |---|---|---|---|---|---|---|---| | recursive_char | 300 | 0 | No | β€” | β€” | β€” | β€” | | recursive_char | 500 | 50 | No | β€” | β€” | β€” | β€” | | recursive_char | 500 | 50 | Yes | β€” | β€” | β€” | β€” | | recursive_char | 1000 | 100 | Yes | β€” | β€” | β€” | β€” | | semantic | 500 | β€” | Yes | β€” | β€” | β€” | β€” | ## Project Structure ``` rag-qa/ β”œβ”€β”€ main.py # FastAPI server β”œβ”€β”€ app_gradio.py # Gradio web UI β”œβ”€β”€ requirements.txt β”œβ”€β”€ Dockerfile β”œβ”€β”€ .env.example β”‚ β”œβ”€β”€ ingestion/ β”‚ β”œβ”€β”€ pdf_reader.py # pymupdf4llm + pdfplumber extraction β”‚ β”œβ”€β”€ chunker.py # fixed_size, recursive_char, semantic strategies β”‚ β”œβ”€β”€ embedder.py # SentenceTransformer + contextual enrichment β”‚ └── pipeline.py # Orchestrates extract β†’ chunk β†’ embed β†’ index β”‚ β”œβ”€β”€ retrieval/ β”‚ β”œβ”€β”€ index.py # FAISS IndexFlatIP wrapper β”‚ β”œβ”€β”€ bm25_index.py # BM25Okapi wrapper β”‚ β”œβ”€β”€ reranker.py # Cross-encoder re-ranking β”‚ └── searcher.py # Hybrid search + expansion + confidence score β”‚ β”œβ”€β”€ generation/ β”‚ └── generator.py # Groq client, confidence gating, streaming β”‚ β”œβ”€β”€ evaluation/ β”‚ β”œβ”€β”€ judge.py # LLM-as-judge (faithfulness + relevance) β”‚ β”œβ”€β”€ eval.py # Retrieval metrics: Acc@K, MRR β”‚ β”œβ”€β”€ eval_e2e.py # End-to-end pipeline evaluation β”‚ └── compare.py # Configuration grid comparison β”‚ └── data/ # Sample PDFs for testing ``` ## How It Works **PDF Extraction** β€” `pymupdf4llm` converts each page to structured markdown, preserving headings and table layout. If it fails, `pdfplumber` extracts flat text. Pages are processed independently so that `page_num` metadata stays accurate for citations. **Chunking** β€” three strategies are available. `fixed_size` is a naive sliding window (fast baseline). `recursive_character` uses LangChain's splitter with markdown-aware separators (`\n## `, `\n\n`, `. `) and extracts the nearest heading above each chunk as `section_header` metadata. `semantic` embeds individual sentences and splits at cosine similarity drops below a threshold, grouping topically coherent content together. **Embedding** β€” before embedding, each chunk's text is prefixed with `" | Page N [| Section]\n"`. This means the stored vector encodes not just the chunk's content but where in the document it came from, which improves retrieval precision for location-specific queries. **Retrieval** β€” a query vector is searched against FAISS (semantic) and BM25 (keyword) simultaneously. Both result sets are min-max normalized to `[0, 1]` and fused by weighted sum. The top 20 fused candidates are then re-scored by the cross-encoder, which reads the full `(query, chunk)` string pair and produces a relevance score independent of embedding geometry. **Generation** β€” `max_score` (cosine similarity of the top-ranked chunk before re-ranking) determines whether to call the LLM. Below 0.3, the call is skipped. Between 0.3 and 0.5, the answer is appended with a confidence warning. Above 0.5, the model receives a numbered context block and is prompted to cite every claim with `[Source N]` notation and synthesize across sources. ## Deployment **Docker:** ```bash docker build -t rag-qa . docker run -p 7860:7860 -e GROQ_API_KEY=your_key rag-qa ``` **HuggingFace Spaces:** Set `GROQ_API_KEY` in your Space's Settings β†’ Repository secrets, then push the repository. The `Dockerfile` is picked up automatically by Spaces. ## Future Improvements - **OCR support** β€” integrate `pytesseract` or `surya-ocr` for scanned PDFs that contain no extractable text layer - **Persistent vector store** β€” replace the in-memory FAISS index with a disk-backed store (FAISS `write_index` with reload on startup, or a dedicated vector DB) so the index survives server restarts - **Multi-document cross-referencing** β€” surface when the same fact is corroborated across multiple uploaded documents rather than citing only one source - **GPU-accelerated embedding** β€” the sentence-transformers model runs on CPU by default; passing `device="cuda"` cuts batch embedding time significantly for large document sets