rag-system / docs /system_architecture.md
joshsears's picture
Polish: BGE-large embeddings, contextual retrieval, 142 tests passing, lint clean
21ca2ea
|
Raw
History Blame Contribute Delete
10.2 kB

RAG System Architecture

Design Goals

This system implements production-grade Retrieval-Augmented Generation with four core principles:

  1. No hallucinations by default β€” the system abstains rather than generating from insufficient context (Sufficient Context module, CRAG evaluation)
  2. Model-agnostic β€” switching between Ollama, Claude, and OpenAI requires only an environment variable change; no code modifications
  3. Observable β€” every retrieval decision, latency, cache hit, and model call is traced via Langfuse and exposed via Prometheus metrics
  4. Extensible β€” each technique is a composable module; adding a new retrieval strategy does not require modifying the generation pipeline

Component Map

Entry Points

main.py β€” Typer CLI. All system capabilities are available via command line:

  • ingest β€” load documents into a knowledge base
  • query β€” standard RAG query
  • chat β€” multi-turn conversation with history compression
  • adaptive β€” auto-selects retrieval strategy (no-retrieval/single/iterative)
  • cot β€” Chain-of-Thought RAG with step-by-step reasoning trace
  • agent β€” agentic loop with tool access
  • raptor-ingest β€” recursive tree ingestion for long documents
  • graph stats/entity/communities/global-query β€” knowledge graph operations
  • lightrag query/stats β€” LightRAG dual-level retrieval
  • eval β€” RAGAS evaluation harness
  • benchmark β€” comparative benchmark across retrieval configurations

api.py β€” FastAPI server with 35+ REST endpoints. Full OpenAPI schema at /docs. Implements streaming SSE, async request handling, and Prometheus metrics instrumentation.

demo.py β€” Streamlit UI with:

  • Source card display (chunk text, similarity score, source file)
  • Sufficiency bar (confidence gauge before each answer)
  • CoT reasoning trace (collapsible step-by-step)
  • Agent tool call log
  • Side-by-side mode comparison (Naive vs Hybrid vs Full Stack)

Core Modules

core/retrieval.py Implements the full retrieval pipeline:

  1. Query embedding (sentence-transformers, MPS/CUDA/CPU)
  2. Dense retrieval from ChromaDB (cosine similarity)
  3. Sparse retrieval via BM25 (rank-bm25)
  4. RRF fusion of dense and sparse rankings
  5. HyDE query expansion (optional)
  6. Multi-query expansion (optional)
  7. Cross-encoder reranking (ms-marco cross-encoder)
  8. MMR diversity selection
  9. CRAG relevance evaluation

core/ingestion.py Document ingestion pipeline:

  1. File loading (PDF via pdfplumber, DOCX via python-docx, TXT/MD direct, URLs via httpx)
  2. Chunking: recursive character splitting, semantic sentence-boundary splitting, or hierarchical section-aware splitting
  3. Contextual Retrieval: Haiku generates per-chunk context summaries (optional)
  4. Quality scoring: language detection, PII scanning, section detection
  5. Deduplication via SHA-256 content hash
  6. ChromaDB upsert with metadata (source, chunk index, embedding model, timestamp)

core/generation.py Generation pipeline with backend abstraction:

  • OllamaBackend: local inference via Ollama REST API
  • ClaudeBackend: Anthropic API with streaming support
  • OpenAIBackend: OpenAI API with streaming support
  • Semantic cache check before any LLM call
  • Sufficient Context evaluation before generation
  • Response parsing and source citation injection

core/sufficient_context.py Implements the Google ICLR 2025 approach:

  • Density scoring: token overlap between query terms and retrieved context
  • Coverage scoring: are all sub-questions in the query addressed by at least one chunk?
  • Optional LLM self-rating: ask the LLM to rate its own confidence on a 0–1 scale
  • Ensemble score: 0.5 Γ— density + 0.3 Γ— coverage + 0.2 Γ— self_rating
  • If score < threshold (default 0.45): retrieve more chunks or abstain

core/contextual_retrieval.py Anthropic Nov 2024 contextual retrieval:

  • Input: full document text + chunk text
  • Output: contextualized chunk = "[Context: {haiku_summary}]\n\n{chunk_text}"
  • Haiku prompt: "Please give a short succinct context to situate this chunk within the overall document."
  • Result: 49% reduction in retrieval failures on Anthropic's benchmarks

core/graph_rag.py Microsoft GraphRAG implementation:

  • Entity and relation extraction via LLM (Claude/Ollama)
  • Knowledge graph stored as NetworkX DiGraph, serialized to JSON
  • Community detection via Louvain algorithm
  • Community summarization via LLM
  • At query time: entity matching + community-level context retrieval

core/light_rag.py HKUDS LightRAG implementation:

  • Dual-level graph: low (entity-level) + high (community-level)
  • Auto-routing: classify query as low-level or high-level before retrieval
  • Incremental updates: new documents extend the graph without full re-ingestion

core/cot_rag.py Chain-of-Thought RAG (EMNLP 2025):

  • Decompose question into reasoning steps via LLM
  • For each step: retrieve step-specific context, generate intermediate answer
  • Synthesize final answer from all intermediate results
  • Returns full reasoning trace for transparency

core/adaptive_rag.py Adaptive + Self-RAG:

  • Query classifier routes to NO_RETRIEVAL / SINGLE_STEP / ITERATIVE
  • NO_RETRIEVAL: answer directly from parametric knowledge
  • SINGLE_STEP: standard retrieval + generation
  • ITERATIVE: retrieve β†’ generate partial answer β†’ rewrite query β†’ retrieve again (up to N iterations)

core/raptor.py RAPTOR recursive tree ingestion:

  • Chunk document into leaf nodes
  • Cluster leaves by semantic similarity (GMM clustering)
  • Summarize each cluster via LLM β†’ intermediate nodes
  • Recurse until single root node
  • Store all levels in ChromaDB; query retrieves from all levels simultaneously

core/agent.py Agentic RAG with Claude tool_use API:

  • search_docs(query, collection, top_k): vector store search
  • search_web(query): DuckDuckGo/Tavily web search
  • query_sql(question): text-to-SQL against configured database
  • calculate(expression): arithmetic and unit conversion
  • Max iterations configurable (default 8)
  • Full tool call log returned with response

core/security.py

  • PII redaction: regex patterns for SSN, credit cards, emails, phone numbers
  • Optional Presidio ML-based PII detection
  • Prompt injection detection: 10 regex patterns on retrieved chunks before generation
  • Audit logging: JSONL with hashed queries, PII flags, session IDs, timestamps

core/feedback.py

  • SQLite feedback store: (session_id, question, answer, rating, timestamp)
  • Contrastive pair mining: thumbs-up answers become positive examples, thumbs-down become negatives
  • Exports training data in sentence-transformers MNR format
  • Triggers fine-tuning pipeline via finetune CLI command

Data Flow

Ingestion

Input (file/URL)
    β†’ load_document() [pdfplumber / docx / httpx]
    β†’ chunk() [recursive | semantic | hierarchical]
    β†’ [optional] contextualize_chunks() [Haiku API per chunk]
    β†’ [optional] analyze_document() [quality score, PII, language]
    β†’ deduplicate() [SHA-256 hash check]
    β†’ embed() [sentence-transformers, MPS]
    β†’ ChromaDB upsert [with metadata]
    β†’ [optional] extract_triples() [knowledge graph update]

Query

Input (question)
    β†’ [optional] route() [which collection?]
    β†’ semantic_cache_check() [cosine similarity β‰₯ 0.95 β†’ return cached]
    β†’ retrieve() [dense + BM25 + RRF]
    β†’ [optional] hyde_expand() [hypothetical document embedding]
    β†’ [optional] multi_query_expand() [N query variants]
    β†’ rerank() [cross-encoder]
    β†’ mmr_diversity() [reduce redundancy]
    β†’ sufficient_context_score() [density + coverage]
    β†’ [if insufficient] retrieve_more() or abstain()
    β†’ [optional] crag_evaluate() [chunk relevance scoring]
    β†’ [optional] web_search_fallback() [DuckDuckGo/Tavily]
    β†’ generate() [Claude / Ollama / OpenAI]
    β†’ extract_sources() [citation injection]
    β†’ cache_store() [for future cache hits]
    β†’ return QueryResponse [answer + sources + latency + tokens]

Storage

Store Technology Purpose
Vector store ChromaDB (embedded) Chunk embeddings + metadata
Knowledge graph NetworkX + JSON Entity-relation graph
Feedback SQLite User ratings + training pairs
Semantic cache In-memory dict Query-response cache
Audit log JSONL file Security audit trail

ChromaDB runs embedded (no separate process). For production scale, the ChromaDB client can be pointed at a remote ChromaDB server by changing CHROMA_PERSIST_DIR to a server URL.


Deployment

Local

pip install -r requirements.txt
cp .env.example .env
# Edit .env
python3 main.py serve

Docker

docker compose up
# API: http://localhost:8000
# Prometheus: http://localhost:9090

The Docker image pins Python 3.11, installs dependencies from requirements.txt, and runs uvicorn with 1 worker (increase API_WORKERS for multi-core production deployments).

Hugging Face Spaces

The hf_space/ directory contains a self-contained Streamlit app with auto-ingestion of sample documents on first boot. Deploy by copying hf_space/ to a Streamlit Space and setting ANTHROPIC_API_KEY in Secrets.


Performance Characteristics

Operation Latency (p50) Notes
Dense retrieval only ~50ms ChromaDB HNSW index
+ BM25 + RRF ~70ms Rank fusion is CPU-bound
+ Cross-encoder reranking ~250ms MiniLM-L6 cross-encoder
+ HyDE ~600ms One LLM call for hypothetical doc
+ Multi-query ~800ms Three LLM calls for variants
+ Contextual Retrieval +2s/chunk at ingest One-time cost, not per query
Cache hit 0ms Cosine lookup
Full stack (Claude backend) ~2–4s End-to-end with generation
CoT-RAG (4 steps) ~8–15s Four retrieval + generation passes
Agentic (3 tool calls) ~10–20s Depends on tools invoked

Semantic cache eliminates LLM latency for repeated or near-duplicate queries. For a corpus with predictable query patterns, 30–60% cache hit rates are achievable.