Spaces:
Sleeping
title: RAG Visualizer
emoji: ๐ฌ
colorFrom: blue
colorTo: red
sdk: docker
app_port: 7860
pinned: false
startup_duration_timeout: 1h
๐ฌ RAG Visualizer
An X-Ray machine for Retrieval-Augmented Generation pipelines.
RAG Visualizer is an interactive, local-first tool that lets you see what happens inside a RAG pipeline โ from how your text gets chunked, to how those chunks land in vector space, to which chunks get retrieved for a given query. No cloud APIs, no black boxes. Everything runs on your machine with local Ollama models.
โจ Features
๐งช Phase 1 โ Chunking Lab
Visualize and compare 5 chunking strategies side-by-side:
| Strategy | Description |
|---|---|
| Fixed Size | Cuts text every N tokens with configurable overlap |
| Sentence | Splits on sentence boundaries using NLTK tokenizer |
| Recursive | Applies a hierarchy of separators (\n\n โ \n โ . โ ) |
| Parent-Child | Two-level nested chunking โ large parent windows with smaller child chunks inside |
| Semantic | Detects topic shifts using embedding similarity + adaptive thresholding |
- Document X-Ray Viewer โ Original text with color-coded chunk boundaries and overlap regions.
- Chunk Inspector โ Stats panel showing total chunks, average token count, and per-chunk metadata.
- File Uploader โ Attach and parse custom text or markdown documents directly in the configuration panel.
๐ Phase 2 โ Embedding Lab
- Generate embeddings using 3 local Ollama embedding models (Nomic Embed Text, Embedding Gemma, Qwen3 Embedding).
- UMAP dimensionality reduction projects high-dimensional embeddings down to 2D.
- Interactive Canvas with pan, zoom, hover tooltips, and click-to-select. Drag-panning is isolated from clicks to ensure smooth navigation without losing focus.
- Parent-child connection lines visualized in vector space.
๐ Phase 3 โ Advanced Retrieval & Reranking
- ChromaDB persistent vector store โ chunks are indexed on every run.
- Flexible Retrieval Modes โ Switch dynamically between Dense (vector similarity), Sparse (BM25 lexical search), or Hybrid (RRF fusion) search paths.
- Sonar Query Simulator โ Type a natural language query and watch the sonar ping animate across the canvas in real time.
- Sonar Probe โ Click anywhere on the 2D canvas to retrieve the nearest chunks in that region.
- Document X-Ray Highlighting โ Retrieved chunks glow dynamically in the document viewer with rank-based styling (gold for Rank 1, dashed for Rank 2, dotted for Rank 3).
- Metadata Level Filtering โ Filter your context pool on the fly (retrieve Only Parents, Only Children, or All Levels).
- Cross-Encoder Reranking โ Run a local FlashRank (
ms-marco-MiniLM-L-12-v2) engine to rerank search results. - Rank Shift Badges โ Visual indicators showing exactly how much chunks moved after reranking (
โฒ +3,โผ -1, orโข Unchanged). - Normalized Match Strength โ Converts raw vector distances into intuitive similarity percentages (e.g.
Match: 87.7%). - Reranking Lineage โ Displays the pre-reranked retrieval score for comparison (e.g.
Match: 95.0% (was Match: 87.7%)).
โ๏ธ Phase 3.2 โ LLM-as-a-Judge (The Grand Arena)
- Side-by-Side Comparison โ Compare retrieval results from two different models/strategies in a split-screen arena.
- AI Referee โ Call upon a local Ollama model to evaluate, rank, and score retrieved contexts.
- Multi-Dimensional Scorecard โ Referee grades chunks on Relevance, Completeness, Factual Plausibility, and Clarity.
- Pydantic Validator Guardrails โ Validates the referee's output to catch and override arithmetic lies and position bias.
๐ Phase 4 โ Adaptive Thresholding (Gradient Fix)
- Semantic chunking uses an adaptive gradient derivative / peak detection algorithm instead of a static threshold split.
- Computes the dynamic threshold based on document-wide mean and standard deviation of inter-sentence embedding distances.
- Uses local maxima peak detection to prevent fragmenting paragraphs, ensuring splits only happen at true topic shift peaks.
๐๏ธ Architecture
flowchart LR
%% Ingestion Flow
DOC["๐ Input Document"] --> CHUNK["โ๏ธ Chunking Engine<br/>5 Chunking Strategies"]
CHUNK --> SPLIT["๐ LangChain + NLTK<br/>Text Splitters"]
SPLIT --> EMBED["๐ง Embedding Engine<br/>Generate Semantic Vectors"]
EMBED --> OLLAMA["๐ค Ollama API<br/>Embedding Model"]
EMBED --> UMAP["๐ UMAP<br/>2D Vector Projection"]
UMAP --> DB["๐๏ธ ChromaDB<br/>Vector Storage"]
%% Retrieval Flow
USER["๐ค User Query"] --> QEMBED["๐ง Query Embedding"]
QEMBED --> OLLAMA
QEMBED --> SEARCH["๐ Similarity Search"]
SEARCH --> DB
DB --> RESULTS["๐ Relevant Chunks"]
%% Visualization
RESULTS --> XRAY["๐ฌ Document X-Ray Viewer"]
UMAP --> VIS["๐ Vector Space Renderer"]
XRAY --> UI["๐ Interactive Frontend"]
VIS --> UI
Data Flow
- User pastes text โ selects strategy + embedding model โ clicks Run Chunking
- Backend splits text into chunks โ generates embeddings via Ollama โ reduces to 2D via UMAP โ stores in ChromaDB
- Frontend renders the chunk boundaries in the X-Ray viewer and plots particles on the 2D canvas
- User queries โ backend embeds the query โ retrieves top-K from ChromaDB โ projects query point into 2D
- Frontend draws sonar lines from query to retrieved chunks, highlights them in the document viewer
๐ Folder Structure
RAG-Visualizer/
โโโ backend/
โ โโโ __init__.py
โ โโโ main.py # FastAPI app, CORS, static file serving
โ โโโ constants.py # LLM prompt templates
โ โโโ engines/
โ โ โโโ __init__.py
โ โ โโโ chunking.py # 5 chunking strategies + ChunkingEngine
โ โ โโโ embedding.py # Ollama embedding adapter (httpx)
โ โ โโโ llm_client.py # Ollama LLM generation client
โ โ โโโ reducer.py # UMAP 2D dimensionality reducer
โ โโโ models/
โ โ โโโ __init__.py
โ โ โโโ schemas.py # Pydantic models (request/response schemas)
โ โโโ routers/
โ โ โโโ __init__.py
โ โ โโโ chunk_router.py # POST /api/chunk โ chunking + embedding + UMAP
โ โ โโโ retrieval_router.py # POST /api/retrieve โ query + ChromaDB retrieval
โ โโโ storage/
โ โโโ vector_store.py # ChromaDB persistent client wrapper
โโโ frontend/
โ โโโ index.html # Single-page app (3-column layout)
โ โโโ app.js # All frontend logic, canvas rendering, API calls
โ โโโ styles.css # Superman theme design system
โโโ store/ # ChromaDB persistent data (gitignored)
โโโ .gitignore
โโโ .python-version # Python 3.11
โโโ dev.bat # Dev server launcher
โโโ pyproject.toml # Project metadata & dependencies
โโโ uv.lock # Locked dependency versions
โโโ README.md
๐ Getting Started
Prerequisites
| Tool | Version | Purpose |
|---|---|---|
| Python | โฅ 3.11 | Runtime |
| uv | Latest | Fast Python package manager |
| Ollama | Latest | Local LLM & embedding model server |
1. Install uv
# Windows (PowerShell)
powershell -ExecutionPolicy ByPass -c "irm https://astral.sh/uv/install.ps1 | iex"
# macOS / Linux
curl -LsSf https://astral.sh/uv/install.sh | sh
2. Clone the Repository
git clone https://github.com/<your-username>/RAG-Visualizer.git
cd RAG-Visualizer
3. Install Dependencies
uv sync
This reads pyproject.toml and uv.lock, creates a .venv, and installs all dependencies in seconds.
4. Pull Ollama Models
Make sure Ollama is running, then pull the required models:
# Embedding models (at least one required)
ollama pull nomic-embed-text
ollama pull qwen3-embedding:0.6b
# LLM model (for future features)
ollama pull gemma4:e2b
5. Run the Dev Server
# Using the dev script (Windows)
.\dev.bat
# Or directly with uv
uv run uvicorn backend.main:app --reload --port 8080
Open http://localhost:8080 in your browser.
๐ฎ Usage Guide
Chunking Lab
- Paste your text into the input area on the left panel
- Select a chunking strategy โ click one of the 5 strategy cards
- Tune parameters โ adjust chunk size, overlap, or semantic threshold with the sliders
- Choose an embedding model from the dropdown
- Click โก Run Chunking
- Explore:
- Document Viewer tab โ see color-coded chunk boundaries in your text
- Vector Space 2D tab โ see chunks plotted as interactive particles
- Chunk Inspector (right panel) โ browse individual chunks with metadata
Sonar Query Simulator
- Switch to the Vector Space 2D tab
- Type a query in the Sonar Query Simulator bar (e.g.,
"linear regression") - Click ๐ Query โ watch the sonar ping animate across the canvas
- Retrieved chunks appear as ranked cards with distance scores
- The Document Viewer automatically highlights retrieved chunks with rank-based glow effects
โ๏ธ API Reference
POST /api/chunk
Chunks input text, generates embeddings, reduces to 2D, and stores in ChromaDB.
Request Body:
{
"text": "Your input text...",
"runs": [
{
"strategy": "fixed_size",
"config": {
"chunk_size": 500,
"chunk_overlap": 20,
"tokenizer": "cl100k_base"
}
}
],
"embedding_model": "nomic-embed-text",
"n_neighbors": 15,
"min_dist": 0.1
}
Response: ChunkResponse with chunks, stats, 2D coordinates, and embeddings.
POST /api/retrieve
Embeds a query and retrieves the top-K most similar chunks from ChromaDB (with optional reranking, HyDE expansion, and metadata filtering).
Request Body:
{
"search_text": "What is gradient descent?",
"embedding_model": "nomic-embed-text",
"strategy": "fixed_size",
"top_k": 3,
"retrieval_mode": "dense",
"use_hyde": false,
"use_reranking": true,
"metadata": { "level": 1 }
}
Response: QueryResponse with query coordinates, retrieved chunks (with original ranks and original scores populated if reranked), and hypothetical answer text if HyDE is used.
POST /api/compare
Compares retrieval results from two different configurations side-by-side.
Request Body:
{
"search_text": "query",
"top_k": 3,
"model_a": "nomic-embed-text",
"strategy_a": "fixed_size",
"model_b": "EmbeddingGemma",
"strategy_b": "semantic",
"retrieval_mode": "dense",
"use_hyde": false,
"use_reranking": true,
"metadata": null
}
Response: CompareResponse containing results from both configuration A and configuration B.
POST /api/judge
Submits retrieval results to a local LLM judge for evaluation and scoring.
Request Body:
{
"search_query": "query",
"chunk_a": "text of chunk a",
"chunk_b": "text of chunk b"
}
Response: JudgeResponse with winner declaration, confidence, scorecards, strengths, and weaknesses.
๐ ๏ธ Tech Stack
| Layer | Technology | Role |
|---|---|---|
| Frontend | Vanilla HTML / CSS / JS | Single-page app, Canvas 2D rendering |
| Backend | FastAPI (Python 3.11) | REST API, async request handling |
| Chunking | LangChain Text Splitters, NLTK | 5 chunking strategy implementations |
| Tokenization | tiktoken (cl100k_base) |
Token counting (OpenAI-compatible) |
| Embeddings | Ollama (local models) | nomic-embed-text, EmbeddingGemma, qwen3-embedding |
| Dimensionality Reduction | UMAP (umap-learn) |
High-dim โ 2D projection for visualization |
| Vector Database | ChromaDB (persistent) | Cosine similarity search with HNSW index |
| Package Manager | uv | Dependency management & virtual environments |

