dharandhamo's picture
chore: add huggingface spaces configuration metadata
e409f74
|
Raw
History Blame Contribute Delete
7.67 kB
metadata
title: Multi Document RAG
emoji: πŸ“š
colorFrom: blue
colorTo: indigo
sdk: docker
pinned: false

πŸ“š Multi-Document RAG System

A production-ready Retrieval-Augmented Generation system for academic research papers with cross-document reasoning and RAGAS evaluation.

Architecture

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”     β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”     β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚  Chainlit UI │────▢│  FastAPI      │────▢│  Groq LLM API    β”‚
β”‚  (Port 8001) β”‚     β”‚  (Port 8000)  β”‚     β”‚  llama-3.3-70b   β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜     β””β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”˜     β”‚  deepseek-r1     β”‚
                            β”‚              β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                            β”‚
                     β”Œβ”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”     β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                     β”‚  Core Engine  │────▢│  Google Gemini    β”‚
                     β”‚  - Ingestion  β”‚     β”‚  Embedding API    β”‚
                     β”‚  - Retrieval  β”‚     β”‚  (3072-dim)       β”‚
                     β”‚  - Generation β”‚     β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                     β”‚  - Evaluation β”‚
                     β””β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”˜
                            β”‚
                     β”Œβ”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”
                     β”‚  Qdrant      β”‚
                     β”‚  Vector DB   β”‚
                     β”‚  (Port 6333) β”‚
                     β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

Features

  • Multi-format ingestion: PDF, DOCX, TXT with automatic metadata extraction
  • Smart query routing: Routes simple vs. complex queries to appropriate models
  • Self-correction: Detects low-relevance retrievals and rewrites queries automatically
  • Cross-document reasoning: Compare and contrast findings across multiple papers
  • RAGAS evaluation: 5 metrics β€” faithfulness, answer relevancy, context precision, context recall, answer correctness
  • Structured API: Full REST API with async endpoints throughout

Prerequisites

Quick Start

1. Clone and configure

git clone <repo-url>
cd multi_doc_rag
cp .env.example .env
# Edit .env and fill in your API keys

2. Start Qdrant (Docker)

docker-compose up -d

This starts Qdrant on port 6333 (~150MB image).

3. Install dependencies

# Using UV (recommended)
uv venv
uv pip install -r requirements.txt

# Or using pip
python -m venv .venv
.venv\Scripts\activate  # Windows
pip install -r requirements.txt

4. Start FastAPI backend

uvicorn app.main:app --reload --port 8000

API docs available at: http://localhost:8000/docs

5. Start Chainlit UI (optional)

chainlit run chainlit_app/app.py --port 8001

API Usage

Ingest a document

curl -X POST http://localhost:8000/api/v1/ingest \
  -F "file=@paper.pdf" \
  -F "source=arxiv"

Batch ingest

curl -X POST http://localhost:8000/api/v1/ingest/batch \
  -F "files=@paper1.pdf" \
  -F "files=@paper2.pdf" \
  -F "source=upload"

Query documents

curl -X POST http://localhost:8000/api/v1/query \
  -H "Content-Type: application/json" \
  -d '{
    "question": "What dataset was used in the study?",
    "mode": "standard",
    "top_k": 5
  }'

Compare documents

curl -X POST http://localhost:8000/api/v1/query/compare \
  -H "Content-Type: application/json" \
  -d '{
    "question": "How do the methodologies differ?",
    "doc_ids": ["doc-id-1", "doc-id-2"],
    "aspect": "methodology"
  }'

List documents

curl http://localhost:8000/api/v1/documents

Delete a document

curl -X DELETE http://localhost:8000/api/v1/documents/{doc_id}

Run RAGAS evaluation

curl -X POST http://localhost:8000/api/v1/evaluate \
  -H "Content-Type: application/json" \
  -d '{"sample_size": 10}'

Check evaluation report

curl http://localhost:8000/api/v1/evaluate/{eval_id}

Health check

curl http://localhost:8000/api/v1/health

Running Tests

pytest tests/ -v

Tests use mocked external services (Groq, Gemini, Qdrant) so no API keys are needed.

RAGAS Evaluation

The system includes 20 sample questions across 5 categories:

Category Count Example
Factual 4 "What dataset was used?"
Comparative 4 "How do methods differ?"
Summarization 4 "Summarize key findings"
Multi-hop 4 "What evidence supports the claim?"
Out-of-scope 4 "What is NVIDIA's stock price?"

Running evaluation

  1. Ingest at least one document first
  2. Use the /api/v1/evaluate endpoint or the /eval command in Chainlit
  3. Reports are saved to evaluation/reports/

Architecture Decisions

Component Choice Rationale
LLM Groq API Fast inference, no GPU needed locally
Embeddings Gemini API High-quality 3072-dim vectors, no local models
Vector DB Qdrant Lightweight (~150MB), rich filtering, easy setup
Framework FastAPI Async-native, auto-generated docs, Pydantic validation
UI Chainlit Minimal setup for prototyping, chat-native interface
Parsing pypdf + python-docx Lightweight, no heavy dependencies
Orchestration LangChain Prompt templates and chain composition

Troubleshooting

Qdrant connection refused

# Ensure Qdrant is running
docker-compose up -d
docker ps  # Check container status

Groq rate limits

The system uses exponential backoff (3 retries). If you hit persistent rate limits:

  • Reduce TOP_K_RETRIEVAL in .env
  • Increase delay between requests
  • Check your Groq API usage dashboard

Empty search results

  • Ensure documents are ingested (GET /api/v1/documents)
  • Check Qdrant collection exists
  • Verify embedding dimension matches (should be 3072)

Docker image too large

The Docker image should stay under 500MB. If it grows:

  • Ensure no torch, transformers, or sentence-transformers in dependencies
  • Check requirements.txt for unexpected heavy packages

Project Structure

multi_doc_rag/
β”œβ”€β”€ app/                    # FastAPI application
β”‚   β”œβ”€β”€ main.py             # Entry point
β”‚   β”œβ”€β”€ config.py           # Settings
β”‚   β”œβ”€β”€ api/routes/         # API endpoints
β”‚   β”œβ”€β”€ core/
β”‚   β”‚   β”œβ”€β”€ ingestion/      # Document loading, chunking, tagging
β”‚   β”‚   β”œβ”€β”€ retrieval/      # Embedding, search, routing, self-correction
β”‚   β”‚   β”œβ”€β”€ generation/     # LLM client, prompts, RAG chains
β”‚   β”‚   └── evaluation/     # RAGAS evaluation pipeline
β”‚   β”œβ”€β”€ models/             # Pydantic schemas & enums
β”‚   └── utils/              # Logging & helpers
β”œβ”€β”€ chainlit_app/           # Chainlit UI
β”œβ”€β”€ evaluation/             # Test questions & reports
β”œβ”€β”€ tests/                  # Test suite
β”œβ”€β”€ docker-compose.yml      # Qdrant only
β”œβ”€β”€ Dockerfile              # App container (slim)
└── requirements.txt        # Dependencies (no torch!)

License

MIT