metadata
title: Multi Document RAG
emoji: π
colorFrom: blue
colorTo: indigo
sdk: docker
pinned: false
π Multi-Document RAG System
A production-ready Retrieval-Augmented Generation system for academic research papers with cross-document reasoning and RAGAS evaluation.
Architecture
ββββββββββββββββ ββββββββββββββββ ββββββββββββββββββββ
β Chainlit UI ββββββΆβ FastAPI ββββββΆβ Groq LLM API β
β (Port 8001) β β (Port 8000) β β llama-3.3-70b β
ββββββββββββββββ ββββββββ¬ββββββββ β deepseek-r1 β
β ββββββββββββββββββββ
β
ββββββββΌββββββββ ββββββββββββββββββββ
β Core Engine ββββββΆβ Google Gemini β
β - Ingestion β β Embedding API β
β - Retrieval β β (3072-dim) β
β - Generation β ββββββββββββββββββββ
β - Evaluation β
ββββββββ¬ββββββββ
β
ββββββββΌββββββββ
β Qdrant β
β Vector DB β
β (Port 6333) β
ββββββββββββββββ
Features
- Multi-format ingestion: PDF, DOCX, TXT with automatic metadata extraction
- Smart query routing: Routes simple vs. complex queries to appropriate models
- Self-correction: Detects low-relevance retrievals and rewrites queries automatically
- Cross-document reasoning: Compare and contrast findings across multiple papers
- RAGAS evaluation: 5 metrics β faithfulness, answer relevancy, context precision, context recall, answer correctness
- Structured API: Full REST API with async endpoints throughout
Prerequisites
- Python 3.11+
- Docker (for Qdrant vector database)
- API Keys:
Quick Start
1. Clone and configure
git clone <repo-url>
cd multi_doc_rag
cp .env.example .env
# Edit .env and fill in your API keys
2. Start Qdrant (Docker)
docker-compose up -d
This starts Qdrant on port 6333 (~150MB image).
3. Install dependencies
# Using UV (recommended)
uv venv
uv pip install -r requirements.txt
# Or using pip
python -m venv .venv
.venv\Scripts\activate # Windows
pip install -r requirements.txt
4. Start FastAPI backend
uvicorn app.main:app --reload --port 8000
API docs available at: http://localhost:8000/docs
5. Start Chainlit UI (optional)
chainlit run chainlit_app/app.py --port 8001
API Usage
Ingest a document
curl -X POST http://localhost:8000/api/v1/ingest \
-F "file=@paper.pdf" \
-F "source=arxiv"
Batch ingest
curl -X POST http://localhost:8000/api/v1/ingest/batch \
-F "files=@paper1.pdf" \
-F "files=@paper2.pdf" \
-F "source=upload"
Query documents
curl -X POST http://localhost:8000/api/v1/query \
-H "Content-Type: application/json" \
-d '{
"question": "What dataset was used in the study?",
"mode": "standard",
"top_k": 5
}'
Compare documents
curl -X POST http://localhost:8000/api/v1/query/compare \
-H "Content-Type: application/json" \
-d '{
"question": "How do the methodologies differ?",
"doc_ids": ["doc-id-1", "doc-id-2"],
"aspect": "methodology"
}'
List documents
curl http://localhost:8000/api/v1/documents
Delete a document
curl -X DELETE http://localhost:8000/api/v1/documents/{doc_id}
Run RAGAS evaluation
curl -X POST http://localhost:8000/api/v1/evaluate \
-H "Content-Type: application/json" \
-d '{"sample_size": 10}'
Check evaluation report
curl http://localhost:8000/api/v1/evaluate/{eval_id}
Health check
curl http://localhost:8000/api/v1/health
Running Tests
pytest tests/ -v
Tests use mocked external services (Groq, Gemini, Qdrant) so no API keys are needed.
RAGAS Evaluation
The system includes 20 sample questions across 5 categories:
| Category | Count | Example |
|---|---|---|
| Factual | 4 | "What dataset was used?" |
| Comparative | 4 | "How do methods differ?" |
| Summarization | 4 | "Summarize key findings" |
| Multi-hop | 4 | "What evidence supports the claim?" |
| Out-of-scope | 4 | "What is NVIDIA's stock price?" |
Running evaluation
- Ingest at least one document first
- Use the
/api/v1/evaluateendpoint or the/evalcommand in Chainlit - Reports are saved to
evaluation/reports/
Architecture Decisions
| Component | Choice | Rationale |
|---|---|---|
| LLM | Groq API | Fast inference, no GPU needed locally |
| Embeddings | Gemini API | High-quality 3072-dim vectors, no local models |
| Vector DB | Qdrant | Lightweight (~150MB), rich filtering, easy setup |
| Framework | FastAPI | Async-native, auto-generated docs, Pydantic validation |
| UI | Chainlit | Minimal setup for prototyping, chat-native interface |
| Parsing | pypdf + python-docx | Lightweight, no heavy dependencies |
| Orchestration | LangChain | Prompt templates and chain composition |
Troubleshooting
Qdrant connection refused
# Ensure Qdrant is running
docker-compose up -d
docker ps # Check container status
Groq rate limits
The system uses exponential backoff (3 retries). If you hit persistent rate limits:
- Reduce
TOP_K_RETRIEVALin.env - Increase delay between requests
- Check your Groq API usage dashboard
Empty search results
- Ensure documents are ingested (
GET /api/v1/documents) - Check Qdrant collection exists
- Verify embedding dimension matches (should be 3072)
Docker image too large
The Docker image should stay under 500MB. If it grows:
- Ensure no
torch,transformers, orsentence-transformersin dependencies - Check
requirements.txtfor unexpected heavy packages
Project Structure
multi_doc_rag/
βββ app/ # FastAPI application
β βββ main.py # Entry point
β βββ config.py # Settings
β βββ api/routes/ # API endpoints
β βββ core/
β β βββ ingestion/ # Document loading, chunking, tagging
β β βββ retrieval/ # Embedding, search, routing, self-correction
β β βββ generation/ # LLM client, prompts, RAG chains
β β βββ evaluation/ # RAGAS evaluation pipeline
β βββ models/ # Pydantic schemas & enums
β βββ utils/ # Logging & helpers
βββ chainlit_app/ # Chainlit UI
βββ evaluation/ # Test questions & reports
βββ tests/ # Test suite
βββ docker-compose.yml # Qdrant only
βββ Dockerfile # App container (slim)
βββ requirements.txt # Dependencies (no torch!)
License
MIT