vector-db / README.md
mittalsahab's picture
Add Dockerfile and HF Spaces config
322102c
|
Raw
History Blame Contribute Delete
6.65 kB
metadata
title: Vector DB Explorer
emoji: πŸ”
colorFrom: indigo
colorTo: purple
sdk: docker
app_port: 7860
pinned: false

Vector DB Explorer

A from-scratch vector database with a web UI β€” built to make embeddings, cosine similarity, and semantic search tangible and interactive.


What it does

You give it text. It converts the text into a list of numbers (an embedding) that captures the meaning of the text, not just its keywords. When you search, your query goes through the same process and the system finds stored documents whose number-lists point in the same "direction" β€” that direction is meaning.

This is how production AI search works (Pinecone, Weaviate, pgvector). This project does the same thing from scratch with plain Python and a JSON file.


Project structure

vector-db/
β”œβ”€β”€ vector_db.py          # The database: embed β†’ store β†’ search
β”œβ”€β”€ app.py                # FastAPI web server (3 API routes)
β”œβ”€β”€ demo.py               # CLI demo: populates the DB and prints results
β”œβ”€β”€ static/
β”‚   └── index.html        # Single-page UI (HTML + CSS + JS, no build step)
β”œβ”€β”€ data.json             # Auto-generated: stores all documents + embeddings
β”œβ”€β”€ requirements.txt      # Python dependencies
└── VECTOR_DB_EXPLAINED.md  # Deep-dive explanation of every concept

How to run

Prerequisites

  • Python 3.10 or newer
  • pip

1. Clone / download the project

cd path/to/vector-db

2. Create a virtual environment

# Windows
python -m venv venv
.\venv\Scripts\activate

# macOS / Linux
python3 -m venv venv
source venv/bin/activate

3. Install dependencies

pip install -r requirements.txt

The first install downloads PyTorch + sentence-transformers (~1 GB total).
Subsequent installs are instant (everything is cached).

4. Seed the database (optional but recommended)

This populates data.json with 16 sample documents across tech, science, food, history, and sports.

python demo.py

You can skip this step and add documents through the UI instead.

5. Start the web server

python app.py

Open http://127.0.0.1:8000 in your browser.

The first startup downloads the all-MiniLM-L6-v2 embedding model (~80 MB) from
Hugging Face. This only happens once β€” it is cached locally after the first run.


Using the UI

Search

Type anything in the search bar and press Enter or click Search.
The query does not need to share any words with the stored documents β€” it searches by meaning.

Use the example chips below the search bar to try pre-built queries instantly.

Reading the results:

Element What it means
Score % (large number) Cosine similarity Γ— 100. Higher = more similar.
Filled bar Visual representation of the score. Color: green β‰₯ 50%, amber β‰₯ 30%, red below.
"Strong / Partial / Weak match" Human-readable label for the score range.
cosine = 0.XXXX The raw cosine similarity value (0 to 1).
Category / Topic badges Metadata stored alongside the document.

Clicking any document card in the Corpus section fires a search for that document automatically.

Add a document

Fill in the text area, set a category and topic, then click + Add Document.
The server embeds the text and saves it to data.json. The corpus updates immediately.

Filter the corpus

Use the filter input in the Corpus section to narrow documents by text, category, or topic.


How it works (code level)

Embedding (vector_db.py β†’ VectorDB.add)

"Machine learning models learn patterns from data"
          |
    SentenceTransformer("all-MiniLM-L6-v2")
          |
    [0.12, -0.44, 0.87, ..., 0.03]   ← 384 floats

The model was pre-trained on hundreds of millions of sentence pairs. It maps similar sentences to nearby regions of a 384-dimensional space.

Storage (vector_db.py β†’ VectorDB.save)

Records are saved to data.json as plain JSON:

{
  "id": 1,
  "text": "Machine learning models...",
  "metadata": { "category": "tech", "topic": "AI/ML" },
  "embedding": [0.048, -0.047, 0.070, ...]
}

No external database is needed. The downside is that large corpora (10k+ documents) become slow to load and search β€” that is why production systems use binary formats and vector indexes.

Search (vector_db.py β†’ VectorDB.search)

Query: "how do computers learn from data?"
    |
embed query β†’ [0.11, -0.39, ...]
    |
compare against every stored embedding using cosine similarity
    |
sort by score, return top-k

Cosine similarity formula:

similarity(A, B) = (A Β· B) / (|A| Γ— |B|)

A score of 1.0 means identical direction (same meaning). A score of 0.0 means orthogonal (unrelated). The search is a linear scan β€” every stored vector is compared β€” which is exact but O(n). Production vector databases replace this with an HNSW index (O(log n)).

Web server (app.py)

Route Method Description
/ GET Serves the single-page UI (static/index.html)
/api/documents GET Returns all documents without their embedding arrays
/api/search POST Embeds the query and returns ranked results
/api/add POST Embeds a new document and appends it to data.json

Embedding is CPU-bound. The server runs it in a thread pool (run_in_threadpool) so the async event loop stays responsive during encoding.


Running the CLI demo instead of the UI

python demo.py

This rebuilds data.json from scratch, then runs five example queries and prints ranked results to the terminal. Useful for quick testing without a browser.


Dependency overview

Package Role
sentence-transformers Loads all-MiniLM-L6-v2 and converts text β†’ embeddings
torch Neural network runtime (pulled in by sentence-transformers)
numpy Efficient array operations used by the model
fastapi Web framework for the API routes
uvicorn ASGI server that runs the FastAPI app
pydantic Request / response schema validation (used by FastAPI)

Going further

What to try How
Different embedding model Change model_name in VectorDB.__init__ to e.g. "all-mpnet-base-v2" (768 dims, higher quality)
More documents Add via the UI or extend the DOCUMENTS list in demo.py
Explore data.json Open it in any editor to see the raw 384-float vectors
See the concept deep-dive Read VECTOR_DB_EXPLAINED.md