File size: 13,927 Bytes
be4fdde
 
 
 
 
 
 
 
66480d3
be4fdde
 
5fbafd7
 
 
 
 
 
137b48f
 
5fbafd7
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
bd53034
 
 
5fbafd7
 
 
bd53034
 
 
 
5fbafd7
bd53034
5fbafd7
bd53034
 
 
 
 
66480d3
bd53034
 
 
 
5fbafd7
137b48f
 
bd53034
 
 
 
137b48f
 
5fbafd7
137b48f
 
bd53034
 
 
5fbafd7
 
 
 
 
 
ce33ef6
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
5fbafd7
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
bd53034
5fbafd7
 
 
 
 
 
 
 
bd53034
 
 
 
 
5fbafd7
 
 
bd53034
5fbafd7
137b48f
 
 
 
 
 
 
 
 
 
 
 
 
bd53034
 
 
 
 
137b48f
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
5fbafd7
137b48f
5fbafd7
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
ce33ef6
5fbafd7
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
---
title: RAG Visualizer
emoji: ๐Ÿ”ฌ
colorFrom: blue
colorTo: red
sdk: docker
app_port: 7860
pinned: false
startup_duration_timeout: 1h
---

# ๐Ÿ”ฌ RAG Visualizer

**An X-Ray machine for Retrieval-Augmented Generation pipelines.**

RAG Visualizer is an interactive, local-first tool that lets you **see** what happens inside a RAG pipeline โ€” from how your text gets chunked, to how those chunks land in vector space, to which chunks get retrieved for a given query. No cloud APIs, no black boxes. Everything runs on your machine with local Ollama models.

![RAG Visualizer Demo](assets/x_ray-visualizer.gif)

---

## โœจ Features

### ๐Ÿงช Phase 1 โ€” Chunking Lab

Visualize and compare **5 chunking strategies** side-by-side:

| Strategy         | Description                                                                       |
| ---------------- | --------------------------------------------------------------------------------- |
| **Fixed Size**   | Cuts text every N tokens with configurable overlap                                |
| **Sentence**     | Splits on sentence boundaries using NLTK tokenizer                                |
| **Recursive**    | Applies a hierarchy of separators (`\n\n` โ†’ `\n` โ†’ `. ` โ†’ ` `)                    |
| **Parent-Child** | Two-level nested chunking โ€” large parent windows with smaller child chunks inside |
| **Semantic**     | Detects topic shifts using embedding similarity + adaptive thresholding           |

- **Document X-Ray Viewer** โ€” Original text with color-coded chunk boundaries and overlap regions.
- **Chunk Inspector** โ€” Stats panel showing total chunks, average token count, and per-chunk metadata.
- **File Uploader** โ€” Attach and parse custom text or markdown documents directly in the configuration panel.

### ๐ŸŒŒ Phase 2 โ€” Embedding Lab

- Generate embeddings using **3 local Ollama embedding models** (Nomic Embed Text, Embedding Gemma, Qwen3 Embedding).
- **UMAP dimensionality reduction** projects high-dimensional embeddings down to 2D.
- **Interactive Canvas** with pan, zoom, hover tooltips, and click-to-select. Drag-panning is isolated from clicks to ensure smooth navigation without losing focus.
- Parent-child connection lines visualized in vector space.

### ๐Ÿ” Phase 3 โ€” Advanced Retrieval & Reranking

- **ChromaDB** persistent vector store โ€” chunks are indexed on every run.
- **Flexible Retrieval Modes** โ€” Switch dynamically between **Dense** (vector similarity), **Sparse** (BM25 lexical search), or **Hybrid** (RRF fusion) search paths.
- **Sonar Query Simulator** โ€” Type a natural language query and watch the sonar ping animate across the canvas in real time.
- **Sonar Probe** โ€” Click anywhere on the 2D canvas to retrieve the nearest chunks in that region.
- **Document X-Ray Highlighting** โ€” Retrieved chunks glow dynamically in the document viewer with rank-based styling (gold for Rank 1, dashed for Rank 2, dotted for Rank 3).
- **Metadata Level Filtering** โ€” Filter your context pool on the fly (retrieve _Only Parents_, _Only Children_, or _All Levels_).
- **Cross-Encoder Reranking** โ€” Run a local FlashRank (`ms-marco-MiniLM-L-12-v2`) engine to rerank search results.
- **Rank Shift Badges** โ€” Visual indicators showing exactly how much chunks moved after reranking (`โ–ฒ +3`, `โ–ผ -1`, or `โ€ข Unchanged`).
- **Normalized Match Strength** โ€” Converts raw vector distances into intuitive similarity percentages (e.g. `Match: 87.7%`).
- **Reranking Lineage** โ€” Displays the pre-reranked retrieval score for comparison (e.g. `Match: 95.0% (was Match: 87.7%)`).

### โš”๏ธ Phase 3.2 โ€” LLM-as-a-Judge (The Grand Arena)

- **Side-by-Side Comparison** โ€” Compare retrieval results from two different models/strategies in a split-screen arena.
- **AI Referee** โ€” Call upon a local Ollama model to evaluate, rank, and score retrieved contexts.
- **Multi-Dimensional Scorecard** โ€” Referee grades chunks on Relevance, Completeness, Factual Plausibility, and Clarity.
- **Pydantic Validator Guardrails** โ€” Validates the referee's output to catch and override arithmetic lies and position bias.

![Grand Arena Comparison](assets/arena_comparison.gif)

### ๐Ÿ“ Phase 4 โ€” Adaptive Thresholding (Gradient Fix)

- Semantic chunking uses an **adaptive gradient derivative / peak detection** algorithm instead of a static threshold split.
- Computes the dynamic threshold based on document-wide mean and standard deviation of inter-sentence embedding distances.
- Uses local maxima peak detection to prevent fragmenting paragraphs, ensuring splits only happen at true topic shift peaks.

---

## ๐Ÿ—๏ธ Architecture

```mermaid
flowchart LR

    %% Ingestion Flow
    DOC["๐Ÿ“„ Input Document"] --> CHUNK["โœ‚๏ธ Chunking Engine<br/>5 Chunking Strategies"]

    CHUNK --> SPLIT["๐Ÿ“ LangChain + NLTK<br/>Text Splitters"]

    SPLIT --> EMBED["๐Ÿง  Embedding Engine<br/>Generate Semantic Vectors"]

    EMBED --> OLLAMA["๐Ÿค– Ollama API<br/>Embedding Model"]

    EMBED --> UMAP["๐Ÿ“‰ UMAP<br/>2D Vector Projection"]

    UMAP --> DB["๐Ÿ—„๏ธ ChromaDB<br/>Vector Storage"]

    %% Retrieval Flow
    USER["๐Ÿ‘ค User Query"] --> QEMBED["๐Ÿง  Query Embedding"]

    QEMBED --> OLLAMA

    QEMBED --> SEARCH["๐Ÿ” Similarity Search"]

    SEARCH --> DB

    DB --> RESULTS["๐Ÿ“š Relevant Chunks"]

    %% Visualization
    RESULTS --> XRAY["๐Ÿ”ฌ Document X-Ray Viewer"]

    UMAP --> VIS["๐Ÿ“Š Vector Space Renderer"]

    XRAY --> UI["๐ŸŒ Interactive Frontend"]
    VIS --> UI
```

### Data Flow

1. **User pastes text** โ†’ selects strategy + embedding model โ†’ clicks **Run Chunking**
2. **Backend** splits text into chunks โ†’ generates embeddings via Ollama โ†’ reduces to 2D via UMAP โ†’ stores in ChromaDB
3. **Frontend** renders the chunk boundaries in the X-Ray viewer and plots particles on the 2D canvas
4. **User queries** โ†’ backend embeds the query โ†’ retrieves top-K from ChromaDB โ†’ projects query point into 2D
5. **Frontend** draws sonar lines from query to retrieved chunks, highlights them in the document viewer

---

## ๐Ÿ“ Folder Structure

```
RAG-Visualizer/
โ”œโ”€โ”€ backend/
โ”‚   โ”œโ”€โ”€ __init__.py
โ”‚   โ”œโ”€โ”€ main.py                    # FastAPI app, CORS, static file serving
โ”‚   โ”œโ”€โ”€ constants.py               # LLM prompt templates
โ”‚   โ”œโ”€โ”€ engines/
โ”‚   โ”‚   โ”œโ”€โ”€ __init__.py
โ”‚   โ”‚   โ”œโ”€โ”€ chunking.py            # 5 chunking strategies + ChunkingEngine
โ”‚   โ”‚   โ”œโ”€โ”€ embedding.py           # Ollama embedding adapter (httpx)
โ”‚   โ”‚   โ”œโ”€โ”€ llm_client.py          # Ollama LLM generation client
โ”‚   โ”‚   โ””โ”€โ”€ reducer.py             # UMAP 2D dimensionality reducer
โ”‚   โ”œโ”€โ”€ models/
โ”‚   โ”‚   โ”œโ”€โ”€ __init__.py
โ”‚   โ”‚   โ””โ”€โ”€ schemas.py             # Pydantic models (request/response schemas)
โ”‚   โ”œโ”€โ”€ routers/
โ”‚   โ”‚   โ”œโ”€โ”€ __init__.py
โ”‚   โ”‚   โ”œโ”€โ”€ chunk_router.py        # POST /api/chunk โ€” chunking + embedding + UMAP
โ”‚   โ”‚   โ””โ”€โ”€ retrieval_router.py    # POST /api/retrieve โ€” query + ChromaDB retrieval
โ”‚   โ””โ”€โ”€ storage/
โ”‚       โ””โ”€โ”€ vector_store.py        # ChromaDB persistent client wrapper
โ”œโ”€โ”€ frontend/
โ”‚   โ”œโ”€โ”€ index.html                 # Single-page app (3-column layout)
โ”‚   โ”œโ”€โ”€ app.js                     # All frontend logic, canvas rendering, API calls
โ”‚   โ””โ”€โ”€ styles.css                 # Superman theme design system
โ”œโ”€โ”€ store/                         # ChromaDB persistent data (gitignored)
โ”œโ”€โ”€ .gitignore
โ”œโ”€โ”€ .python-version                # Python 3.11
โ”œโ”€โ”€ dev.bat                        # Dev server launcher
โ”œโ”€โ”€ pyproject.toml                 # Project metadata & dependencies
โ”œโ”€โ”€ uv.lock                        # Locked dependency versions
โ””โ”€โ”€ README.md
```

---

## ๐Ÿš€ Getting Started

### Prerequisites

| Tool                                 | Version | Purpose                            |
| ------------------------------------ | ------- | ---------------------------------- |
| **Python**                           | โ‰ฅ 3.11  | Runtime                            |
| **[uv](https://docs.astral.sh/uv/)** | Latest  | Fast Python package manager        |
| **[Ollama](https://ollama.com/)**    | Latest  | Local LLM & embedding model server |

### 1. Install uv

```bash
# Windows (PowerShell)
powershell -ExecutionPolicy ByPass -c "irm https://astral.sh/uv/install.ps1 | iex"

# macOS / Linux
curl -LsSf https://astral.sh/uv/install.sh | sh
```

### 2. Clone the Repository

```bash
git clone https://github.com/<your-username>/RAG-Visualizer.git
cd RAG-Visualizer
```

### 3. Install Dependencies

```bash
uv sync
```

This reads `pyproject.toml` and `uv.lock`, creates a `.venv`, and installs all dependencies in seconds.

### 4. Pull Ollama Models

Make sure Ollama is running, then pull the required models:

```bash
# Embedding models (at least one required)
ollama pull nomic-embed-text
ollama pull qwen3-embedding:0.6b

# LLM model (for future features)
ollama pull gemma4:e2b
```

### 5. Run the Dev Server

```bash
# Using the dev script (Windows)
.\dev.bat

# Or directly with uv
uv run uvicorn backend.main:app --reload --port 8080
```

Open **http://localhost:8080** in your browser.

---

## ๐ŸŽฎ Usage Guide

### Chunking Lab

1. **Paste your text** into the input area on the left panel
2. **Select a chunking strategy** โ€” click one of the 5 strategy cards
3. **Tune parameters** โ€” adjust chunk size, overlap, or semantic threshold with the sliders
4. **Choose an embedding model** from the dropdown
5. Click **โšก Run Chunking**
6. Explore:
   - **Document Viewer tab** โ€” see color-coded chunk boundaries in your text
   - **Vector Space 2D tab** โ€” see chunks plotted as interactive particles
   - **Chunk Inspector** (right panel) โ€” browse individual chunks with metadata

### Sonar Query Simulator

1. Switch to the **Vector Space 2D** tab
2. Type a query in the **Sonar Query Simulator** bar (e.g., `"linear regression"`)
3. Click **๐Ÿ” Query** โ€” watch the sonar ping animate across the canvas
4. Retrieved chunks appear as ranked cards with distance scores
5. The **Document Viewer** automatically highlights retrieved chunks with rank-based glow effects

---

## โš™๏ธ API Reference

### `POST /api/chunk`

Chunks input text, generates embeddings, reduces to 2D, and stores in ChromaDB.

**Request Body:**

```json
{
  "text": "Your input text...",
  "runs": [
    {
      "strategy": "fixed_size",
      "config": {
        "chunk_size": 500,
        "chunk_overlap": 20,
        "tokenizer": "cl100k_base"
      }
    }
  ],
  "embedding_model": "nomic-embed-text",
  "n_neighbors": 15,
  "min_dist": 0.1
}
```

**Response:** `ChunkResponse` with chunks, stats, 2D coordinates, and embeddings.

### `POST /api/retrieve`

Embeds a query and retrieves the top-K most similar chunks from ChromaDB (with optional reranking, HyDE expansion, and metadata filtering).

**Request Body:**

```json
{
  "search_text": "What is gradient descent?",
  "embedding_model": "nomic-embed-text",
  "strategy": "fixed_size",
  "top_k": 3,
  "retrieval_mode": "dense",
  "use_hyde": false,
  "use_reranking": true,
  "metadata": { "level": 1 }
}
```

**Response:** `QueryResponse` with query coordinates, retrieved chunks (with original ranks and original scores populated if reranked), and hypothetical answer text if HyDE is used.

### `POST /api/compare`

Compares retrieval results from two different configurations side-by-side.

**Request Body:**

```json
{
  "search_text": "query",
  "top_k": 3,
  "model_a": "nomic-embed-text",
  "strategy_a": "fixed_size",
  "model_b": "EmbeddingGemma",
  "strategy_b": "semantic",
  "retrieval_mode": "dense",
  "use_hyde": false,
  "use_reranking": true,
  "metadata": null
}
```

**Response:** `CompareResponse` containing results from both configuration A and configuration B.

### `POST /api/judge`

Submits retrieval results to a local LLM judge for evaluation and scoring.

**Request Body:**

```json
{
  "search_query": "query",
  "chunk_a": "text of chunk a",
  "chunk_b": "text of chunk b"
}
```

**Response:** `JudgeResponse` with winner declaration, confidence, scorecards, strengths, and weaknesses.

---

## ๐Ÿ› ๏ธ Tech Stack

| Layer                        | Technology                     | Role                                                    |
| ---------------------------- | ------------------------------ | ------------------------------------------------------- |
| **Frontend**                 | Vanilla HTML / CSS / JS        | Single-page app, Canvas 2D rendering                    |
| **Backend**                  | FastAPI (Python 3.11)          | REST API, async request handling                        |
| **Chunking**                 | LangChain Text Splitters, NLTK | 5 chunking strategy implementations                     |
| **Tokenization**             | tiktoken (`cl100k_base`)       | Token counting (OpenAI-compatible)                      |
| **Embeddings**               | Ollama (local models)          | `nomic-embed-text`, `EmbeddingGemma`, `qwen3-embedding` |
| **Dimensionality Reduction** | UMAP (`umap-learn`)            | High-dim โ†’ 2D projection for visualization              |
| **Vector Database**          | ChromaDB (persistent)          | Cosine similarity search with HNSW index                |
| **Package Manager**          | uv                             | Dependency management & virtual environments            |

## Technologies

- [Ollama](https://ollama.com/) โ€” Local LLM inference
- [ChromaDB](https://www.trychroma.com/) โ€” Open-source vector database
- [LangChain](https://www.langchain.com/) โ€” Text splitting utilities
- [UMAP](https://umap-learn.readthedocs.io/) โ€” Dimensionality reduction
- [FastAPI](https://fastapi.tiangolo.com/) โ€” Modern Python web framework