rag-document-qa / README.md
Amrita P
feat: add readme
3a4a51d
|
Raw
History Blame Contribute Delete
15 kB
metadata
title: RAG Document Q&A
emoji: πŸ“„
colorFrom: indigo
colorTo: blue
sdk: docker
app_port: 7860
pinned: false

RAG Document Q&A System

Upload PDF documents and get cited, grounded answers β€” with hybrid retrieval, cross-encoder re-ranking, and confidence-aware generation.

Overview

This project implements a production-grade Retrieval-Augmented Generation pipeline that goes beyond the typical "embed + nearest-neighbor + prompt" tutorial. It combines semantic and keyword search with cross-encoder re-ranking, applies Anthropic's contextual retrieval pattern to enrich embeddings with document metadata, and gates LLM calls on retrieval confidence to avoid hallucinated answers. The system ships with both a streaming Gradio UI and a FastAPI backend with Server-Sent Events, plus an evaluation harness that benchmarks chunking strategies and scores end-to-end answer quality using an LLM-as-judge.

Architecture

                          β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                          β”‚              INGESTION                  β”‚
                          β”‚                                         β”‚
  PDF ──► pymupdf4llm ──► β”‚  Chunking (3 strategies)                β”‚
           (pdfplumber    β”‚  β”œβ”€ fixed_size       sliding window     β”‚
            fallback)     β”‚  β”œβ”€ recursive_char   LangChain splits   β”‚
                          β”‚  └─ semantic          embedding sim     β”‚
                          β”‚           β”‚                             β”‚
                          β”‚           β–Ό                             β”‚
                          β”‚  Contextual Enrichment                  β”‚
                          β”‚  "filename | Page N | Section\ntext"    β”‚
                          β”‚           β”‚                             β”‚
                          β”‚           β–Ό                             β”‚
                          β”‚  all-MiniLM-L6-v2 (384-dim, L2-norm)    β”‚
                          β”‚           β”‚                             β”‚
                          β”‚     β”Œβ”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”                      β”‚
                          β”‚     β–Ό            β–Ό                      β”‚
                          β”‚  FAISS        BM25Okapi                 β”‚
                          β”‚  IndexFlatIP  rank-bm25                 β”‚
                          β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

                          β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                          β”‚               QUERY                     β”‚
                          β”‚                                         β”‚
  Question ──────────────►│  Adaptive Expansion (Groq)              β”‚
                          β”‚  only if top_score < 0.45               β”‚
                          β”‚           β”‚                             β”‚
                          β”‚           β–Ό                             β”‚
                          β”‚  Hybrid Search                          β”‚
                          β”‚  score = 0.7Β·semantic + 0.3Β·bm25        β”‚
                          β”‚  (top 20 candidates)                    β”‚
                          β”‚           β”‚                             β”‚
                          β”‚           β–Ό                             β”‚
                          β”‚  Cross-Encoder Re-ranking               β”‚
                          β”‚  ms-marco-MiniLM-L-6-v2                 β”‚
                          β”‚  (20 β†’ top k)                           β”‚
                          β”‚           β”‚                             β”‚
                          β”‚           β–Ό                             β”‚
                          β”‚  Confidence Gating                      β”‚
                          β”‚  cosine sim < 0.3 β†’ skip LLM            β”‚
                          β”‚           β”‚                             β”‚
                          β”‚           β–Ό                             β”‚
                          β”‚  Groq (Llama 3.3 70B) streaming         β”‚
                          β”‚           β”‚                             β”‚
                          β”‚           β–Ό                             β”‚
                          β”‚  Cited answer  [Source N] notation      β”‚
                          β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

                          β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                          β”‚             EVALUATION                  β”‚
                          β”‚                                         β”‚
                          β”‚  Retrieval: Acc@1/3/5, MRR              β”‚
                          β”‚  End-to-end: faithfulness + relevance   β”‚
                          β”‚             (LLM-as-judge, temp=0.0)    β”‚
                          β”‚  Config sweep: chunking strategy grid   β”‚
                          β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

Key Features

  • Cross-encoder re-ranking β€” bi-encoder retrieval (FAISS) produces 20 candidates; a cross-encoder (ms-marco-MiniLM-L-6-v2) scores each (query, chunk) pair jointly and re-orders them. Cross-encoders are too slow for full-corpus search but dramatically improve precision over the shortlist.

  • Hybrid search β€” FAISS inner-product search and BM25 run in parallel. Both score sets are normalized to [0, 1] then fused (0.7Β·semantic + 0.3Β·keyword), catching exact-match terms that dense embeddings tend to smooth over.

  • Adaptive query expansion β€” the pipeline runs an initial search first. Only if the top cosine similarity falls below 0.45 does it call the LLM to generate 3 rephrased variants, search each, and merge results. High-confidence queries pay zero expansion cost.

  • Contextual chunk enrichment β€” each chunk is embedded as "<filename> | Page N | <section>\n<text>" rather than raw text, following Anthropic's contextual retrieval pattern. The embedding captures document location and section context, not just lexical content.

  • Confidence-based response gating β€” max_score is taken from cosine similarity before re-ranking (cross-encoder scores on a βˆ’10 to +10 scale would break the threshold). If max_score < 0.3, the LLM is skipped entirely and the user receives an honest "insufficient context" message. Scores between 0.3–0.5 append a low-confidence warning.

  • Streaming responses β€” Gradio UI streams tokens via a generator; the FastAPI backend streams as Server-Sent Events (text/event-stream) with {"text": chunk} events and a terminal {"done": true}.

  • Layout-aware PDF parsing β€” pymupdf4llm extracts structured markdown with headings and tables preserved. pdfplumber serves as a fallback for PDFs that pymupdf4llm cannot parse cleanly.

  • End-to-end evaluation β€” the eval harness measures retrieval hit rate (Acc@K, MRR), then pipes results through the LLM-as-judge to score faithfulness (are claims grounded in context?) and relevance (does the answer address the question?) at temperature 0 for determinism.

Tech Stack

Component Technology
Embedding model all-MiniLM-L6-v2 (sentence-transformers, 384-dim)
Vector index FAISS IndexFlatIP (exact inner-product search)
Keyword index BM25 (rank-bm25, BM25Okapi)
Re-ranker cross-encoder/ms-marco-MiniLM-L-6-v2
LLM Llama 3.3 70B via Groq API
PDF parsing pymupdf4llm + pdfplumber fallback
Text splitting LangChain RecursiveCharacterTextSplitter
Backend FastAPI + Uvicorn
Frontend Gradio (Blocks, streaming)
Containerization Docker
Cloud deployment HuggingFace Spaces

Quick Start

git clone <repo-url>
cd rag-qa
python -m venv venv && source venv/bin/activate
pip install -r requirements.txt

Create a .env file:

GROQ_API_KEY=your_key_here

Get a free key at console.groq.com β€” the free tier allows 14,400 requests/day.

Launch the Gradio UI:

python app_gradio.py
# Opens at http://127.0.0.1:7860

Or run the FastAPI server:

uvicorn main:app --reload
# API docs at http://127.0.0.1:8000/docs

API Endpoints

Method Endpoint Description
GET /health Liveness check
GET /stats Index size, chunks ingested, model info
POST /ingest Upload a PDF; returns chunk count and IDs
POST /query Ask a question; returns answer + sources + confidence
POST /query/stream Same as /query but streams as SSE
POST /debug/chunks Return raw retrieved chunks without calling the LLM

All POST query endpoints accept {"question": "...", "top_k": 5}.

The /debug/chunks endpoint is useful for diagnosing retrieval issues β€” it shows each chunk's rank, cosine similarity score, source, page, section header, and a 300-character text preview.

Evaluation

The evaluation framework has three layers:

Retrieval metrics (evaluation/eval.py) β€” given a test set of (question, expected_source, expected_page) tuples, measures whether the correct chunk appears in the top-k results:

Metric Definition
Accuracy@K Fraction of queries where correct chunk is in top K
MRR Mean Reciprocal Rank of the first correct result

End-to-end metrics (evaluation/eval_e2e.py) β€” runs the full pipeline and scores each answer:

Metric Scorer
Faithfulness LLM judge: are all claims grounded in retrieved context?
Relevance LLM judge: does the answer address the question?
Keyword coverage Fraction of expected answer keywords found in output

Chunking strategy comparison (evaluation/compare.py) β€” grid search over chunk sizes, overlap values, and re-ranking on/off. Results table (fill in with your benchmark numbers):

Strategy Chunk size Overlap Reranking Acc@5 MRR Faithfulness Relevance
recursive_char 300 0 No β€” β€” β€” β€”
recursive_char 500 50 No β€” β€” β€” β€”
recursive_char 500 50 Yes β€” β€” β€” β€”
recursive_char 1000 100 Yes β€” β€” β€” β€”
semantic 500 β€” Yes β€” β€” β€” β€”

Project Structure

rag-qa/
β”œβ”€β”€ main.py                   # FastAPI server
β”œβ”€β”€ app_gradio.py             # Gradio web UI
β”œβ”€β”€ requirements.txt
β”œβ”€β”€ Dockerfile
β”œβ”€β”€ .env.example
β”‚
β”œβ”€β”€ ingestion/
β”‚   β”œβ”€β”€ pdf_reader.py         # pymupdf4llm + pdfplumber extraction
β”‚   β”œβ”€β”€ chunker.py            # fixed_size, recursive_char, semantic strategies
β”‚   β”œβ”€β”€ embedder.py           # SentenceTransformer + contextual enrichment
β”‚   └── pipeline.py           # Orchestrates extract β†’ chunk β†’ embed β†’ index
β”‚
β”œβ”€β”€ retrieval/
β”‚   β”œβ”€β”€ index.py              # FAISS IndexFlatIP wrapper
β”‚   β”œβ”€β”€ bm25_index.py         # BM25Okapi wrapper
β”‚   β”œβ”€β”€ reranker.py           # Cross-encoder re-ranking
β”‚   └── searcher.py           # Hybrid search + expansion + confidence score
β”‚
β”œβ”€β”€ generation/
β”‚   └── generator.py          # Groq client, confidence gating, streaming
β”‚
β”œβ”€β”€ evaluation/
β”‚   β”œβ”€β”€ judge.py              # LLM-as-judge (faithfulness + relevance)
β”‚   β”œβ”€β”€ eval.py               # Retrieval metrics: Acc@K, MRR
β”‚   β”œβ”€β”€ eval_e2e.py           # End-to-end pipeline evaluation
β”‚   └── compare.py            # Configuration grid comparison
β”‚
└── data/                     # Sample PDFs for testing

How It Works

PDF Extraction β€” pymupdf4llm converts each page to structured markdown, preserving headings and table layout. If it fails, pdfplumber extracts flat text. Pages are processed independently so that page_num metadata stays accurate for citations.

Chunking β€” three strategies are available. fixed_size is a naive sliding window (fast baseline). recursive_character uses LangChain's splitter with markdown-aware separators (\n## , \n\n, . ) and extracts the nearest heading above each chunk as section_header metadata. semantic embeds individual sentences and splits at cosine similarity drops below a threshold, grouping topically coherent content together.

Embedding β€” before embedding, each chunk's text is prefixed with "<filename> | Page N [| Section]\n". This means the stored vector encodes not just the chunk's content but where in the document it came from, which improves retrieval precision for location-specific queries.

Retrieval β€” a query vector is searched against FAISS (semantic) and BM25 (keyword) simultaneously. Both result sets are min-max normalized to [0, 1] and fused by weighted sum. The top 20 fused candidates are then re-scored by the cross-encoder, which reads the full (query, chunk) string pair and produces a relevance score independent of embedding geometry.

Generation β€” max_score (cosine similarity of the top-ranked chunk before re-ranking) determines whether to call the LLM. Below 0.3, the call is skipped. Between 0.3 and 0.5, the answer is appended with a confidence warning. Above 0.5, the model receives a numbered context block and is prompted to cite every claim with [Source N] notation and synthesize across sources.

Deployment

Docker:

docker build -t rag-qa .
docker run -p 7860:7860 -e GROQ_API_KEY=your_key rag-qa

HuggingFace Spaces:

Set GROQ_API_KEY in your Space's Settings β†’ Repository secrets, then push the repository. The Dockerfile is picked up automatically by Spaces.

Future Improvements

  • OCR support β€” integrate pytesseract or surya-ocr for scanned PDFs that contain no extractable text layer
  • Persistent vector store β€” replace the in-memory FAISS index with a disk-backed store (FAISS write_index with reload on startup, or a dedicated vector DB) so the index survives server restarts
  • Multi-document cross-referencing β€” surface when the same fact is corroborated across multiple uploaded documents rather than citing only one source
  • GPU-accelerated embedding β€” the sentence-transformers model runs on CPU by default; passing device="cuda" cuts batch embedding time significantly for large document sets