--- title: Multi Document RAG emoji: πŸ“š colorFrom: blue colorTo: indigo sdk: docker pinned: false --- # πŸ“š Multi-Document RAG System A production-ready **Retrieval-Augmented Generation** system for academic research papers with cross-document reasoning and RAGAS evaluation. ## Architecture ``` β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚ Chainlit UI │────▢│ FastAPI │────▢│ Groq LLM API β”‚ β”‚ (Port 8001) β”‚ β”‚ (Port 8000) β”‚ β”‚ llama-3.3-70b β”‚ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚ deepseek-r1 β”‚ β”‚ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚ β”Œβ”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚ Core Engine │────▢│ Google Gemini β”‚ β”‚ - Ingestion β”‚ β”‚ Embedding API β”‚ β”‚ - Retrieval β”‚ β”‚ (3072-dim) β”‚ β”‚ - Generation β”‚ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚ - Evaluation β”‚ β””β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚ β”Œβ”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β” β”‚ Qdrant β”‚ β”‚ Vector DB β”‚ β”‚ (Port 6333) β”‚ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ ``` ## Features - **Multi-format ingestion**: PDF, DOCX, TXT with automatic metadata extraction - **Smart query routing**: Routes simple vs. complex queries to appropriate models - **Self-correction**: Detects low-relevance retrievals and rewrites queries automatically - **Cross-document reasoning**: Compare and contrast findings across multiple papers - **RAGAS evaluation**: 5 metrics β€” faithfulness, answer relevancy, context precision, context recall, answer correctness - **Structured API**: Full REST API with async endpoints throughout ## Prerequisites - **Python** 3.11+ - **Docker** (for Qdrant vector database) - **API Keys**: - [Groq API Key](https://console.groq.com/) - [Google AI API Key](https://aistudio.google.com/apikey) ## Quick Start ### 1. Clone and configure ```bash git clone cd multi_doc_rag cp .env.example .env # Edit .env and fill in your API keys ``` ### 2. Start Qdrant (Docker) ```bash docker-compose up -d ``` This starts Qdrant on port 6333 (~150MB image). ### 3. Install dependencies ```bash # Using UV (recommended) uv venv uv pip install -r requirements.txt # Or using pip python -m venv .venv .venv\Scripts\activate # Windows pip install -r requirements.txt ``` ### 4. Start FastAPI backend ```bash uvicorn app.main:app --reload --port 8000 ``` API docs available at: http://localhost:8000/docs ### 5. Start Chainlit UI (optional) ```bash chainlit run chainlit_app/app.py --port 8001 ``` ## API Usage ### Ingest a document ```bash curl -X POST http://localhost:8000/api/v1/ingest \ -F "file=@paper.pdf" \ -F "source=arxiv" ``` ### Batch ingest ```bash curl -X POST http://localhost:8000/api/v1/ingest/batch \ -F "files=@paper1.pdf" \ -F "files=@paper2.pdf" \ -F "source=upload" ``` ### Query documents ```bash curl -X POST http://localhost:8000/api/v1/query \ -H "Content-Type: application/json" \ -d '{ "question": "What dataset was used in the study?", "mode": "standard", "top_k": 5 }' ``` ### Compare documents ```bash curl -X POST http://localhost:8000/api/v1/query/compare \ -H "Content-Type: application/json" \ -d '{ "question": "How do the methodologies differ?", "doc_ids": ["doc-id-1", "doc-id-2"], "aspect": "methodology" }' ``` ### List documents ```bash curl http://localhost:8000/api/v1/documents ``` ### Delete a document ```bash curl -X DELETE http://localhost:8000/api/v1/documents/{doc_id} ``` ### Run RAGAS evaluation ```bash curl -X POST http://localhost:8000/api/v1/evaluate \ -H "Content-Type: application/json" \ -d '{"sample_size": 10}' ``` ### Check evaluation report ```bash curl http://localhost:8000/api/v1/evaluate/{eval_id} ``` ### Health check ```bash curl http://localhost:8000/api/v1/health ``` ## Running Tests ```bash pytest tests/ -v ``` Tests use mocked external services (Groq, Gemini, Qdrant) so no API keys are needed. ## RAGAS Evaluation The system includes 20 sample questions across 5 categories: | Category | Count | Example | |----------|-------|---------| | Factual | 4 | "What dataset was used?" | | Comparative | 4 | "How do methods differ?" | | Summarization | 4 | "Summarize key findings" | | Multi-hop | 4 | "What evidence supports the claim?" | | Out-of-scope | 4 | "What is NVIDIA's stock price?" | ### Running evaluation 1. Ingest at least one document first 2. Use the `/api/v1/evaluate` endpoint or the `/eval` command in Chainlit 3. Reports are saved to `evaluation/reports/` ## Architecture Decisions | Component | Choice | Rationale | |-----------|--------|-----------| | LLM | Groq API | Fast inference, no GPU needed locally | | Embeddings | Gemini API | High-quality 3072-dim vectors, no local models | | Vector DB | Qdrant | Lightweight (~150MB), rich filtering, easy setup | | Framework | FastAPI | Async-native, auto-generated docs, Pydantic validation | | UI | Chainlit | Minimal setup for prototyping, chat-native interface | | Parsing | pypdf + python-docx | Lightweight, no heavy dependencies | | Orchestration | LangChain | Prompt templates and chain composition | ## Troubleshooting ### Qdrant connection refused ```bash # Ensure Qdrant is running docker-compose up -d docker ps # Check container status ``` ### Groq rate limits The system uses exponential backoff (3 retries). If you hit persistent rate limits: - Reduce `TOP_K_RETRIEVAL` in `.env` - Increase delay between requests - Check your Groq API usage dashboard ### Empty search results - Ensure documents are ingested (`GET /api/v1/documents`) - Check Qdrant collection exists - Verify embedding dimension matches (should be 3072) ### Docker image too large The Docker image should stay under 500MB. If it grows: - Ensure no `torch`, `transformers`, or `sentence-transformers` in dependencies - Check `requirements.txt` for unexpected heavy packages ## Project Structure ``` multi_doc_rag/ β”œβ”€β”€ app/ # FastAPI application β”‚ β”œβ”€β”€ main.py # Entry point β”‚ β”œβ”€β”€ config.py # Settings β”‚ β”œβ”€β”€ api/routes/ # API endpoints β”‚ β”œβ”€β”€ core/ β”‚ β”‚ β”œβ”€β”€ ingestion/ # Document loading, chunking, tagging β”‚ β”‚ β”œβ”€β”€ retrieval/ # Embedding, search, routing, self-correction β”‚ β”‚ β”œβ”€β”€ generation/ # LLM client, prompts, RAG chains β”‚ β”‚ └── evaluation/ # RAGAS evaluation pipeline β”‚ β”œβ”€β”€ models/ # Pydantic schemas & enums β”‚ └── utils/ # Logging & helpers β”œβ”€β”€ chainlit_app/ # Chainlit UI β”œβ”€β”€ evaluation/ # Test questions & reports β”œβ”€β”€ tests/ # Test suite β”œβ”€β”€ docker-compose.yml # Qdrant only β”œβ”€β”€ Dockerfile # App container (slim) └── requirements.txt # Dependencies (no torch!) ``` ## License MIT