Spaces:
Runtime error
Runtime error
| # RAG System Architecture | |
| ## Design Goals | |
| This system implements production-grade Retrieval-Augmented Generation with four core principles: | |
| 1. **No hallucinations by default** β the system abstains rather than generating from insufficient context (Sufficient Context module, CRAG evaluation) | |
| 2. **Model-agnostic** β switching between Ollama, Claude, and OpenAI requires only an environment variable change; no code modifications | |
| 3. **Observable** β every retrieval decision, latency, cache hit, and model call is traced via Langfuse and exposed via Prometheus metrics | |
| 4. **Extensible** β each technique is a composable module; adding a new retrieval strategy does not require modifying the generation pipeline | |
| --- | |
| ## Component Map | |
| ### Entry Points | |
| **`main.py`** β Typer CLI. All system capabilities are available via command line: | |
| - `ingest` β load documents into a knowledge base | |
| - `query` β standard RAG query | |
| - `chat` β multi-turn conversation with history compression | |
| - `adaptive` β auto-selects retrieval strategy (no-retrieval/single/iterative) | |
| - `cot` β Chain-of-Thought RAG with step-by-step reasoning trace | |
| - `agent` β agentic loop with tool access | |
| - `raptor-ingest` β recursive tree ingestion for long documents | |
| - `graph stats/entity/communities/global-query` β knowledge graph operations | |
| - `lightrag query/stats` β LightRAG dual-level retrieval | |
| - `eval` β RAGAS evaluation harness | |
| - `benchmark` β comparative benchmark across retrieval configurations | |
| **`api.py`** β FastAPI server with 35+ REST endpoints. Full OpenAPI schema at `/docs`. Implements streaming SSE, async request handling, and Prometheus metrics instrumentation. | |
| **`demo.py`** β Streamlit UI with: | |
| - Source card display (chunk text, similarity score, source file) | |
| - Sufficiency bar (confidence gauge before each answer) | |
| - CoT reasoning trace (collapsible step-by-step) | |
| - Agent tool call log | |
| - Side-by-side mode comparison (Naive vs Hybrid vs Full Stack) | |
| ### Core Modules | |
| **`core/retrieval.py`** | |
| Implements the full retrieval pipeline: | |
| 1. Query embedding (sentence-transformers, MPS/CUDA/CPU) | |
| 2. Dense retrieval from ChromaDB (cosine similarity) | |
| 3. Sparse retrieval via BM25 (rank-bm25) | |
| 4. RRF fusion of dense and sparse rankings | |
| 5. HyDE query expansion (optional) | |
| 6. Multi-query expansion (optional) | |
| 7. Cross-encoder reranking (ms-marco cross-encoder) | |
| 8. MMR diversity selection | |
| 9. CRAG relevance evaluation | |
| **`core/ingestion.py`** | |
| Document ingestion pipeline: | |
| 1. File loading (PDF via pdfplumber, DOCX via python-docx, TXT/MD direct, URLs via httpx) | |
| 2. Chunking: recursive character splitting, semantic sentence-boundary splitting, or hierarchical section-aware splitting | |
| 3. Contextual Retrieval: Haiku generates per-chunk context summaries (optional) | |
| 4. Quality scoring: language detection, PII scanning, section detection | |
| 5. Deduplication via SHA-256 content hash | |
| 6. ChromaDB upsert with metadata (source, chunk index, embedding model, timestamp) | |
| **`core/generation.py`** | |
| Generation pipeline with backend abstraction: | |
| - `OllamaBackend`: local inference via Ollama REST API | |
| - `ClaudeBackend`: Anthropic API with streaming support | |
| - `OpenAIBackend`: OpenAI API with streaming support | |
| - Semantic cache check before any LLM call | |
| - Sufficient Context evaluation before generation | |
| - Response parsing and source citation injection | |
| **`core/sufficient_context.py`** | |
| Implements the Google ICLR 2025 approach: | |
| - **Density scoring**: token overlap between query terms and retrieved context | |
| - **Coverage scoring**: are all sub-questions in the query addressed by at least one chunk? | |
| - **Optional LLM self-rating**: ask the LLM to rate its own confidence on a 0β1 scale | |
| - Ensemble score: 0.5 Γ density + 0.3 Γ coverage + 0.2 Γ self_rating | |
| - If score < threshold (default 0.45): retrieve more chunks or abstain | |
| **`core/contextual_retrieval.py`** | |
| Anthropic Nov 2024 contextual retrieval: | |
| - Input: full document text + chunk text | |
| - Output: contextualized chunk = "[Context: {haiku_summary}]\n\n{chunk_text}" | |
| - Haiku prompt: "Please give a short succinct context to situate this chunk within the overall document." | |
| - Result: 49% reduction in retrieval failures on Anthropic's benchmarks | |
| **`core/graph_rag.py`** | |
| Microsoft GraphRAG implementation: | |
| - Entity and relation extraction via LLM (Claude/Ollama) | |
| - Knowledge graph stored as NetworkX DiGraph, serialized to JSON | |
| - Community detection via Louvain algorithm | |
| - Community summarization via LLM | |
| - At query time: entity matching + community-level context retrieval | |
| **`core/light_rag.py`** | |
| HKUDS LightRAG implementation: | |
| - Dual-level graph: low (entity-level) + high (community-level) | |
| - Auto-routing: classify query as low-level or high-level before retrieval | |
| - Incremental updates: new documents extend the graph without full re-ingestion | |
| **`core/cot_rag.py`** | |
| Chain-of-Thought RAG (EMNLP 2025): | |
| - Decompose question into reasoning steps via LLM | |
| - For each step: retrieve step-specific context, generate intermediate answer | |
| - Synthesize final answer from all intermediate results | |
| - Returns full reasoning trace for transparency | |
| **`core/adaptive_rag.py`** | |
| Adaptive + Self-RAG: | |
| - Query classifier routes to NO_RETRIEVAL / SINGLE_STEP / ITERATIVE | |
| - NO_RETRIEVAL: answer directly from parametric knowledge | |
| - SINGLE_STEP: standard retrieval + generation | |
| - ITERATIVE: retrieve β generate partial answer β rewrite query β retrieve again (up to N iterations) | |
| **`core/raptor.py`** | |
| RAPTOR recursive tree ingestion: | |
| - Chunk document into leaf nodes | |
| - Cluster leaves by semantic similarity (GMM clustering) | |
| - Summarize each cluster via LLM β intermediate nodes | |
| - Recurse until single root node | |
| - Store all levels in ChromaDB; query retrieves from all levels simultaneously | |
| **`core/agent.py`** | |
| Agentic RAG with Claude `tool_use` API: | |
| - `search_docs(query, collection, top_k)`: vector store search | |
| - `search_web(query)`: DuckDuckGo/Tavily web search | |
| - `query_sql(question)`: text-to-SQL against configured database | |
| - `calculate(expression)`: arithmetic and unit conversion | |
| - Max iterations configurable (default 8) | |
| - Full tool call log returned with response | |
| **`core/security.py`** | |
| - PII redaction: regex patterns for SSN, credit cards, emails, phone numbers | |
| - Optional Presidio ML-based PII detection | |
| - Prompt injection detection: 10 regex patterns on retrieved chunks before generation | |
| - Audit logging: JSONL with hashed queries, PII flags, session IDs, timestamps | |
| **`core/feedback.py`** | |
| - SQLite feedback store: (session_id, question, answer, rating, timestamp) | |
| - Contrastive pair mining: thumbs-up answers become positive examples, thumbs-down become negatives | |
| - Exports training data in sentence-transformers MNR format | |
| - Triggers fine-tuning pipeline via `finetune` CLI command | |
| --- | |
| ## Data Flow | |
| ### Ingestion | |
| ``` | |
| Input (file/URL) | |
| β load_document() [pdfplumber / docx / httpx] | |
| β chunk() [recursive | semantic | hierarchical] | |
| β [optional] contextualize_chunks() [Haiku API per chunk] | |
| β [optional] analyze_document() [quality score, PII, language] | |
| β deduplicate() [SHA-256 hash check] | |
| β embed() [sentence-transformers, MPS] | |
| β ChromaDB upsert [with metadata] | |
| β [optional] extract_triples() [knowledge graph update] | |
| ``` | |
| ### Query | |
| ``` | |
| Input (question) | |
| β [optional] route() [which collection?] | |
| β semantic_cache_check() [cosine similarity β₯ 0.95 β return cached] | |
| β retrieve() [dense + BM25 + RRF] | |
| β [optional] hyde_expand() [hypothetical document embedding] | |
| β [optional] multi_query_expand() [N query variants] | |
| β rerank() [cross-encoder] | |
| β mmr_diversity() [reduce redundancy] | |
| β sufficient_context_score() [density + coverage] | |
| β [if insufficient] retrieve_more() or abstain() | |
| β [optional] crag_evaluate() [chunk relevance scoring] | |
| β [optional] web_search_fallback() [DuckDuckGo/Tavily] | |
| β generate() [Claude / Ollama / OpenAI] | |
| β extract_sources() [citation injection] | |
| β cache_store() [for future cache hits] | |
| β return QueryResponse [answer + sources + latency + tokens] | |
| ``` | |
| --- | |
| ## Storage | |
| | Store | Technology | Purpose | | |
| |-------|-----------|---------| | |
| | Vector store | ChromaDB (embedded) | Chunk embeddings + metadata | | |
| | Knowledge graph | NetworkX + JSON | Entity-relation graph | | |
| | Feedback | SQLite | User ratings + training pairs | | |
| | Semantic cache | In-memory dict | Query-response cache | | |
| | Audit log | JSONL file | Security audit trail | | |
| ChromaDB runs embedded (no separate process). For production scale, the ChromaDB client can be pointed at a remote ChromaDB server by changing `CHROMA_PERSIST_DIR` to a server URL. | |
| --- | |
| ## Deployment | |
| ### Local | |
| ```bash | |
| pip install -r requirements.txt | |
| cp .env.example .env | |
| # Edit .env | |
| python3 main.py serve | |
| ``` | |
| ### Docker | |
| ```bash | |
| docker compose up | |
| # API: http://localhost:8000 | |
| # Prometheus: http://localhost:9090 | |
| ``` | |
| The Docker image pins Python 3.11, installs dependencies from requirements.txt, and runs uvicorn with 1 worker (increase `API_WORKERS` for multi-core production deployments). | |
| ### Hugging Face Spaces | |
| The `hf_space/` directory contains a self-contained Streamlit app with auto-ingestion of sample documents on first boot. Deploy by copying `hf_space/` to a Streamlit Space and setting `ANTHROPIC_API_KEY` in Secrets. | |
| --- | |
| ## Performance Characteristics | |
| | Operation | Latency (p50) | Notes | | |
| |-----------|--------------|-------| | |
| | Dense retrieval only | ~50ms | ChromaDB HNSW index | | |
| | + BM25 + RRF | ~70ms | Rank fusion is CPU-bound | | |
| | + Cross-encoder reranking | ~250ms | MiniLM-L6 cross-encoder | | |
| | + HyDE | ~600ms | One LLM call for hypothetical doc | | |
| | + Multi-query | ~800ms | Three LLM calls for variants | | |
| | + Contextual Retrieval | +2s/chunk at ingest | One-time cost, not per query | | |
| | Cache hit | 0ms | Cosine lookup | | |
| | Full stack (Claude backend) | ~2β4s | End-to-end with generation | | |
| | CoT-RAG (4 steps) | ~8β15s | Four retrieval + generation passes | | |
| | Agentic (3 tool calls) | ~10β20s | Depends on tools invoked | | |
| Semantic cache eliminates LLM latency for repeated or near-duplicate queries. For a corpus with predictable query patterns, 30β60% cache hit rates are achievable. | |