NexusMemory / README.md
bharatverse11's picture
Nexus Memory backend on HF Spaces (Docker + Ollama)
044b758
|
Raw
History Blame Contribute Delete
32.7 kB
metadata
title: Nexus Memory Backend
emoji: 🧠
colorFrom: indigo
colorTo: purple
sdk: docker
app_port: 7860
pinned: false

Nexus Memory

A local AI assistant that remembers you, answers questions from your documents, and runs entirely on your machine β€” no cloud, no subscriptions, no data leaving your device.

Python FastAPI React LangChain Ollama ChromaDB License

Features β€’ Architecture β€’ Setup β€’ Usage β€’ API β€’ Tech Stack


What Is Nexus Memory?

Nexus Memory is a production-inspired local AI assistant built to demonstrate practical AI engineering skills. Unlike generic chatbots, Nexus Memory:

  • Remembers facts about you across conversations using semantic memory extraction
  • Reads your documents and answers questions using RAG (Retrieval-Augmented Generation)
  • Runs 100% locally via Ollama β€” your data never leaves your machine
  • Streams responses in real-time just like ChatGPT
  • Persists everything in SQLite + ChromaDB for long-term use

Built in one day as a focused demonstration of modern AI engineering β€” clean architecture, practical tooling, and shipping ability.


Features

Conversational Chat

  • ChatGPT-style streaming interface with real-time token rendering
  • Markdown rendering with code blocks, tables, and formatting
  • Multi-turn conversation with full history persistence
  • Multiple chat sessions with sidebar navigation
  • Copy message button on every bubble
  • Typing indicator during response generation

Persistent Memory System

  • Automatic memory extraction β€” AI identifies and stores important facts from your conversation
  • Semantic memory retrieval β€” finds relevant memories using vector similarity search
  • Memory panel β€” view, add, and delete memories with category badges
  • Memory badges on responses showing how many memories were used
  • Categories: technical, professional, preference, goal, general
  • Backed by SQLite (structured) + ChromaDB (semantic search)

PDF / Document Chat (RAG)

  • Upload PDF, TXT, and Markdown files
  • Automatic text extraction and intelligent chunking
  • Semantic embeddings via nomic-embed-text
  • MMR (Maximum Marginal Relevance) retrieval for diverse, relevant results
  • Source-aware answers with document chunk badges
  • Multi-document support β€” query across multiple files simultaneously

Local LLM Integration

  • Powered by Ollama β€” runs llama3, mistral, and any installed model
  • Live model selector dropdown β€” switch models mid-conversation
  • No API keys, no rate limits, no cloud dependency
  • Streaming responses via Server-Sent Events (SSE)

Modern Dark UI

  • Clean dark theme with indigo/purple accent colors
  • Smooth Framer Motion animations throughout
  • Responsive sidebar with collapsible panel
  • Drag-and-drop file upload with progress tracking
  • Memory side panel with live count badge

    

    

    

    

    

Architecture

System Overview

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚                        NEXUS MEMORY                             β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚                                                                 β”‚
β”‚   β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”‚
β”‚   β”‚                   FRONTEND (React + Vite)                β”‚  β”‚
β”‚   β”‚                                                          β”‚  β”‚
β”‚   β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”‚  β”‚
β”‚   β”‚  β”‚ Sidebar  β”‚  β”‚ ChatWindow β”‚  β”‚   MemoryPanel        β”‚  β”‚  β”‚
β”‚   β”‚  β”‚ Sessions β”‚  β”‚ Messages   β”‚  β”‚   Brain Icon Panel   β”‚  β”‚  β”‚
β”‚   β”‚  β”‚ Memory   β”‚  β”‚ Streaming  β”‚  β”‚   Add/Delete Memory  β”‚  β”‚  β”‚
β”‚   β”‚  β”‚ Tab      β”‚  β”‚ Markdown   β”‚  β”‚                      β”‚  β”‚  β”‚
β”‚   β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β”‚  β”‚
β”‚   β”‚                                                          β”‚  β”‚
β”‚   β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”    β”‚  β”‚
β”‚   β”‚  β”‚UploadSection β”‚  β”‚ModelSelector β”‚  β”‚ MessageBubbleβ”‚    β”‚  β”‚
β”‚   β”‚  β”‚ Drag & Drop  β”‚  β”‚ Live Models  β”‚  β”‚ Badges+Copy  β”‚    β”‚  β”‚
β”‚   β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜    β”‚  β”‚
β”‚   β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β”‚
β”‚                             β”‚ HTTP / SSE                        β”‚
β”‚                             β–Ό                                   β”‚
β”‚   β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”‚
β”‚   β”‚                 BACKEND (FastAPI + Python)               β”‚  β”‚
β”‚   β”‚                                                          β”‚  β”‚
β”‚   β”‚  POST /chat          GET /chat/sessions                  β”‚  β”‚
β”‚   β”‚  POST /upload        GET /memory/{session_id}            β”‚  β”‚
β”‚   β”‚  POST /memory        DELETE /memory/{id}                 β”‚  β”‚
β”‚   β”‚  GET  /health        DELETE /documents/{id}              β”‚  β”‚
β”‚   β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β”‚
β”‚              β”‚                    β”‚                             β”‚
β”‚              β–Ό                    β–Ό                             β”‚
β”‚   β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”    β”‚
β”‚   β”‚  LANGCHAIN       β”‚  β”‚         STORAGE LAYER            β”‚    β”‚
β”‚   β”‚  PIPELINE        β”‚  β”‚                                  β”‚    β”‚
β”‚   β”‚                  β”‚  β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”‚    β”‚
β”‚   β”‚  PromptTemplate  β”‚  β”‚  β”‚   SQLite    β”‚ β”‚  ChromaDB  β”‚  β”‚    β”‚
β”‚   β”‚  Memory Retrievalβ”‚  β”‚  β”‚             β”‚ β”‚            β”‚  β”‚    β”‚
β”‚   β”‚  Doc Retrieval   β”‚  β”‚  β”‚ Sessions    β”‚ β”‚ Memory     β”‚  β”‚    β”‚
β”‚   β”‚  Context Build   β”‚  β”‚  β”‚ Messages    β”‚ β”‚ Vectors    β”‚  β”‚    β”‚
β”‚   β”‚  Response Stream β”‚  β”‚  β”‚ Memories    β”‚ β”‚            β”‚  β”‚    β”‚
β”‚   β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β”‚  β”‚ Documents   β”‚ β”‚ Doc Chunks β”‚  β”‚    β”‚
β”‚            β”‚            β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β”‚    β”‚
β”‚            β–Ό            β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜    β”‚
β”‚   β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”                                          β”‚
β”‚   β”‚  OLLAMA (Local)  β”‚                                          β”‚
β”‚   β”‚                  β”‚                                          β”‚
β”‚   β”‚  llama3          β”‚                                          β”‚
β”‚   β”‚  nomic-embed-textβ”‚                                          β”‚
β”‚   β”‚  mistral / phi3  β”‚                                          β”‚
β”‚   β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜                                          β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

Chat Request Flow

User Types Message
       β”‚
       β–Ό
useChat Hook (React)
  └── fetch POST /chat  ─────────────────────────────────────────┐
                                                                 β”‚
                                                         FastAPI /chat
                                                                 β”‚
                              β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
                              β”‚                                  β”‚
                              β–Ό                                  β”‚
                   Retrieve Memories                             β”‚
                   (ChromaDB semantic search)                    β”‚
                   filter by session_id                          β”‚
                   threshold > 0.3 similarity                    β”‚
                              β”‚                                  β”‚
                              β–Ό                                  β”‚
                   Retrieve Document Chunks                      β”‚
                   (ChromaDB MMR search)                         β”‚
                   across all session documents                  β”‚
                              β”‚                                  β”‚
                              β–Ό                                  β”‚
                   Build LangChain Prompt                        β”‚
                   β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”               β”‚
                   β”‚ SYSTEM: memories + doc ctx  β”‚               β”‚
                   β”‚ HISTORY: last 6 messages    β”‚               β”‚
                   β”‚ HUMAN: user message         β”‚               β”‚
                   β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜               β”‚
                              β”‚                                  β”‚
                              β–Ό                                  β”‚
                   OllamaLLM.astream()                           β”‚
                   yields tokens ──► SSE Stream ──► Browser      β”‚
                              β”‚                                  β”‚
                              β–Ό                                  β”‚
                   Store Assistant Message (SQLite)              β”‚
                              β”‚                                  β”‚
                              β–Ό                                  β”‚
                   Extract Memories (LLM pipeline)               β”‚
                   Store Facts (SQLite + ChromaDB)               β”‚
                              β”‚                                  β”‚
                              β–Ό                                  β”‚
                   Browser renders tokens in real-time           β”‚
                   Shows 🧠 memory badge + πŸ“„ docs badge        

RAG Document Pipeline

User Uploads PDF/TXT/MD
         β”‚
         β–Ό
  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
  β”‚   File Saved    β”‚
  β”‚  to uploads/    β”‚
  β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”˜
           β”‚
           β–Ό
  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
  β”‚ Text Extraction β”‚
  β”‚  PyPDFLoader    β”‚
  β”‚  TextLoader     β”‚
  β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”˜
           β”‚
           β–Ό
  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
  β”‚         Intelligent Chunking            β”‚
  β”‚   RecursiveCharacterTextSplitter        β”‚
  β”‚   chunk_size=1000, overlap=200          β”‚
  β”‚   separators: ["\n\n","\n",". "," ",""] β”‚
  β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
           β”‚
           β–Ό
  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
  β”‚         Vector Embeddings               β”‚
  β”‚   nomic-embed-text via Ollama           β”‚
  β”‚   768-dimensional vectors               β”‚
  β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
           β”‚
           β–Ό
  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
  β”‚         ChromaDB Storage                β”‚
  β”‚   Unique collection per document        β”‚
  β”‚   Persistent on disk                    β”‚
  β”‚   Metadata: source, page, chunk_index   β”‚
  β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
           β”‚
           β–Ό
  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
  β”‚      Query Time (MMR Retrieval)         β”‚
  β”‚   fetch_k = 12 candidates               β”‚
  β”‚   return k = 4 diverse results          β”‚
  β”‚   lambda_mult = 0.7 (relevance bias)    β”‚
  β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
           β”‚
           β–Ό
  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
  β”‚      LLM Context Injection              β”‚
  β”‚   Formatted with source citations       β”‚
  β”‚   Injected into system prompt           β”‚
  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

Memory System Pipeline

Conversation Happens
        β”‚
        β–Ό
LLM Extracts Facts
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚  Prompt: "Extract important personal     β”‚
β”‚  facts from this conversation..."        β”‚
β”‚  Returns: JSON array of facts            β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                   β”‚
        β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
        β”‚                     β”‚
        β–Ό                     β–Ό
  SQLite Storage        ChromaDB Storage
  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”      β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
  β”‚ memory_id    β”‚      β”‚ Vector embedding β”‚
  β”‚ session_id   β”‚      β”‚ of fact text     β”‚
  β”‚ fact (text)  β”‚      β”‚                  β”‚
  β”‚ category     β”‚      β”‚ Metadata:        β”‚
  β”‚ confidence   β”‚      β”‚  session_id      β”‚
  β”‚ created_at   β”‚      β”‚  category        β”‚
  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜      β”‚  memory_id       β”‚
                        β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                β”‚
                                β–Ό
                   Query Time: Semantic Search
                   β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                   β”‚ similarity_search with   β”‚
                   β”‚ session_id filter        β”‚
                   β”‚ relevance threshold 0.3  β”‚
                   β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                β”‚
                                β–Ό
                   Inject into System Prompt
                   "Relevant memories about user:
                    - User prefers Python
                    - User works at a startup
                    - User loves FastAPI"

Tech Stack

Layer Technology Purpose
Frontend React 18 + Vite UI framework + build tool
Styling TailwindCSS v3 Utility-first dark theme
Animations Framer Motion Smooth transitions
Icons Lucide React Consistent icon set
Markdown React Markdown + remark-gfm AI response rendering
HTTP Client Axios + Fetch API calls + SSE streaming
Backend FastAPI High-performance async API
LLM Runtime Ollama Local model inference
AI Framework LangChain Prompt chains + retrieval
LLM llama3 / mistral Language generation
Embeddings nomic-embed-text Semantic vector creation
Vector DB ChromaDB Semantic similarity search
Relational DB SQLite Session + memory persistence
PDF Parsing PyPDF Document text extraction
Text Splitting RecursiveCharacterTextSplitter Intelligent chunking

Requirements

System Requirements

Requirement Minimum Recommended
OS macOS 12+ / Ubuntu 20+ macOS 14+
RAM 8 GB 16 GB
Storage 10 GB free 20 GB free
Python 3.11+ 3.11.8
Node.js 18+ 20+

Software Prerequisites

# 1. Python 3.11+
python3 --version   # must be 3.11.x or higher

# 2. Node.js 18+
node --version      # must be v18.x or higher
npm --version       # must be 9.x or higher

# 3. Ollama
# Download from https://ollama.com
ollama --version    # must be installed

Python Dependencies

fastapi==0.111.0
uvicorn[standard]==0.30.1
python-multipart==0.0.9
langchain==0.2.16
langchain-community==0.2.16
langchain-chroma==0.1.4
langchain-ollama==0.1.3
chromadb==0.5.3
sentence-transformers==3.0.1
pypdf==4.2.0
sqlalchemy==2.0.31
pydantic==2.8.2
pydantic-settings==2.3.4
python-dotenv==1.0.1
aiofiles==23.2.1
httpx==0.27.0

Frontend Dependencies

react: ^18.3.1
react-dom: ^18.3.1
axios: ^1.7.2
react-markdown: ^9.0.1
framer-motion: ^11.2.12
react-dropzone: ^14.2.3
lucide-react: ^0.395.0
date-fns: ^3.6.0
remark-gfm: ^4.0.0
uuid: ^10.0.0

Setup

Step 1 β€” Clone the Repository

git clone https://github.com/yourusername/nexus-memory.git
cd nexus-memory

Step 2 β€” Install and Start Ollama

# Download Ollama from https://ollama.com
# After installation, pull required models:

ollama pull llama3
ollama pull nomic-embed-text

# Verify models are available
ollama list

Expected output:

NAME                       ID              SIZE    MODIFIED
llama3:latest              365c0bd3c000    4.7 GB  X days ago
nomic-embed-text:latest    0a109f422b47    274 MB  X days ago

Step 3 β€” Backend Setup

# Navigate to project root
cd nexus-memory

# Create virtual environment at ROOT level (important!)
python3 -m venv venv

# Activate virtual environment
source venv/bin/activate          # macOS / Linux
# venv\Scripts\activate           # Windows

# Install all dependencies
pip install --upgrade pip
pip install -r requirements.txt

# Pre-compile packages to prevent watchfiles issues
python -m compileall venv/lib/python3.11/site-packages/ -q

# Start the backend
cd backend
python run.py

Backend runs at: http://localhost:8000
API docs at: http://localhost:8000/docs


Step 4 β€” Frontend Setup

# Open a NEW terminal tab
cd nexus-memory/frontend

# Install dependencies
npm install

# Start development server
npm run dev

Frontend runs at: http://localhost:5173


Step 5 β€” Verify Everything is Running

# Check system health
curl -s http://localhost:8000/health | python -m json.tool

Expected response:

{
  "api": "healthy",
  "ollama": {
    "status": "healthy",
    "models": ["nomic-embed-text:latest", "llama3:latest"],
    "current_model": "llama3"
  },
  "model": "llama3"
}

Open http://localhost:5173 β€” you should see the dark Nexus Memory interface.


Project Structure

nexus-memory/
β”‚
β”œβ”€β”€ venv/                          # Python virtual environment (root level)
β”‚
β”œβ”€β”€ backend/
β”‚   β”œβ”€β”€ main.py                    # FastAPI app + lifespan + CORS
β”‚   β”œβ”€β”€ config.py                  # Settings + environment config
β”‚   β”œβ”€β”€ database.py                # SQLAlchemy models + session
β”‚   β”œβ”€β”€ run.py                     # Uvicorn server with watch config
β”‚   β”‚
β”‚   β”œβ”€β”€ routes/
β”‚   β”‚   β”œβ”€β”€ chat.py                # POST /chat (streaming SSE)
β”‚   β”‚   β”œβ”€β”€ upload.py              # POST /upload (document processing)
β”‚   β”‚   └── memory.py              # GET/POST/DELETE /memory
β”‚   β”‚
β”‚   β”œβ”€β”€ rag/
β”‚   β”‚   β”œβ”€β”€ loader.py              # PDF + TXT document loading
β”‚   β”‚   β”œβ”€β”€ chunker.py             # RecursiveCharacterTextSplitter
β”‚   β”‚   β”œβ”€β”€ embeddings.py          # ChromaDB store + load + delete
β”‚   β”‚   └── retriever.py           # MMR semantic retrieval
β”‚   β”‚
β”‚   β”œβ”€β”€ memory/
β”‚   β”‚   β”œβ”€β”€ store.py               # SQLite + ChromaDB memory storage
β”‚   β”‚   └── extractor.py           # LLM-based fact extraction
β”‚   β”‚
β”‚   β”œβ”€β”€ services/
β”‚   β”‚   └── llm.py                 # Ollama LLM + embeddings factory
β”‚   β”‚
β”‚   β”œβ”€β”€ uploads/                   # Uploaded document storage
β”‚   β”œβ”€β”€ chroma_db/                 # ChromaDB vector persistence
β”‚   └── nexus_memory.db            # SQLite database
β”‚
β”œβ”€β”€ frontend/
β”‚   β”œβ”€β”€ index.html
β”‚   β”œβ”€β”€ vite.config.js
β”‚   β”œβ”€β”€ tailwind.config.js
β”‚   β”œβ”€β”€ postcss.config.js
β”‚   β”œβ”€β”€ package.json
β”‚   β”‚
β”‚   └── src/
β”‚       β”œβ”€β”€ main.jsx               # React entry point
β”‚       β”œβ”€β”€ App.jsx                # Root layout + session state
β”‚       β”œβ”€β”€ index.css              # Tailwind + custom styles
β”‚       β”‚
β”‚       β”œβ”€β”€ components/
β”‚       β”‚   β”œβ”€β”€ ChatWindow.jsx     # Main chat interface
β”‚       β”‚   β”œβ”€β”€ MessageBubble.jsx  # User + AI message rendering
β”‚       β”‚   β”œβ”€β”€ Sidebar.jsx        # Sessions + Memory tabs
β”‚       β”‚   β”œβ”€β”€ UploadSection.jsx  # Drag-drop document upload
β”‚       β”‚   β”œβ”€β”€ MemoryPanel.jsx    # Memory viewer + editor
β”‚       β”‚   β”œβ”€β”€ TypingIndicator.jsx # Animated typing dots
β”‚       β”‚   └── ModelSelector.jsx  # Ollama model picker (portal)
β”‚       β”‚
β”‚       β”œβ”€β”€ hooks/
β”‚       β”‚   └── useChat.js         # SSE streaming hook
β”‚       β”‚
β”‚       └── services/
β”‚           └── api.js             # All API call functions
β”‚
β”œβ”€β”€ screenshots/                   # App screenshots
β”œβ”€β”€ requirements.txt               # Python dependencies
└── README.md                      # This file

Usage Guide

Basic Chat

Open http://localhost:5173

Click New Chat or type directly in the input box

Press Enter to send (Shift+Enter for new line)

Watch the AI stream its response in real-time


Persistent Memory

The system automatically extracts and stores facts from your conversations:

You: "My name is Alex and I prefer Python over JavaScript"

β†’ Memory extracted: "User's name is Alex"
β†’ Memory extracted: "User prefers Python over JavaScript"
β†’ Stored in: SQLite + ChromaDB

Later...

You: "What language should I use for my next project?"

β†’ AI retrieves memories β†’ recommends Python with context
β†’ Shows: 🧠 2 memories badge

Manually Add Memories

Click the 🧠 Brain icon in the top bar

Click Add Memory Manually

Enter fact + select category β†’ Save


PDF / Document Chat

Click the πŸ“Ž Paperclip icon in the input bar

Drag & drop or click to select PDF, TXT, or MD file

Wait for "X chunks ready" confirmation

Ask questions about your document:

You: "What is the quarterly revenue?"

β†’ AI retrieves relevant chunks β†’ answers from document
β†’ Shows: πŸ“„ 3 chunks badge

Combined Memory + Documents

You: "Given my Python expertise, what do you think
      about the tech stack in the uploaded document?"

β†’ AI uses: memories (your Python background)
         + documents (the uploaded file context)

β†’ Shows BOTH: 🧠 memories badge + πŸ“„ chunks badge

Model Switching

Click the model selector (top right, shows current model)

Dropdown shows all installed Ollama models

Click any model to switch for the next message


API Reference

Chat

Method Endpoint Description
POST /chat Send message, returns SSE stream
GET /chat/sessions List all chat sessions
GET /chat/sessions/{id}/messages Get session message history
DELETE /chat/sessions/{id} Delete session + messages

POST /chat Request Body

{
  "message": "What is my favorite language?",
  "session_id": "uuid-here",
  "model": "llama3",
  "stream": true
}

SSE Stream Events

data: {"type": "metadata", "session_id": "...", "memories_used": 2, "docs_retrieved": 3}
data: {"type": "token", "content": "Based"}
data: {"type": "token", "content": " on"}
data: {"type": "done"}

Documents

Method Endpoint Description
POST /upload Upload + process document
GET /documents/{session_id} List session documents
DELETE /documents/{id} Delete document + vectors

POST /upload (multipart form)

file: <PDF/TXT/MD file>
session_id: "uuid-here"

Memory

Method Endpoint Description
GET /memory/{session_id} Get all session memories
GET /memory Get all memories (global)
POST /memory Add memory manually
DELETE /memory/{id} Delete specific memory

POST /memory Request Body

{
  "session_id": "uuid-here",
  "fact": "User prefers dark mode editors",
  "category": "preference"
}

System

Method Endpoint Description
GET /health System health + Ollama status
GET /docs Interactive Swagger UI

Quick Commands

# ── Start All Services ──────────────────────────────────────────

# Terminal 1: Ollama (if not running as background service)
ollama serve

# Terminal 2: Backend
cd nexus-memory/backend
source ../venv/bin/activate
python run.py

# Terminal 3: Frontend
cd nexus-memory/frontend
npm run dev


# ── Useful Debug Commands ────────────────────────────────────────

# Check health
curl -s http://localhost:8000/health | python -m json.tool

# Test chat (non-streaming)
curl -X POST http://localhost:8000/chat \
  -H "Content-Type: application/json" \
  -d '{"message": "Hello!", "session_id": "test", "stream": false}'

# Add a memory manually
curl -X POST http://localhost:8000/memory \
  -H "Content-Type: application/json" \
  -d '{"session_id": "test", "fact": "I love Python", "category": "technical"}'

# View memories
curl http://localhost:8000/memory/test | python -m json.tool

# Check SQLite database
cd nexus-memory/backend
sqlite3 nexus_memory.db ".tables"
sqlite3 nexus_memory.db "SELECT fact, category FROM memory_entries;"

# Check ChromaDB collections
cd nexus-memory
source venv/bin/activate
python3 -c "
import chromadb
client = chromadb.PersistentClient(path='backend/chroma_db')
for col in client.list_collections():
    print(f'{col.name}: {col.count()} vectors')
"


# ── Reset / Clean Data ───────────────────────────────────────────

# Delete all chat data (keeps code)
rm backend/nexus_memory.db
rm -rf backend/chroma_db/
rm -rf backend/uploads/*

# Full dependency reinstall
cd nexus-memory
rm -rf venv
python3 -m venv venv
source venv/bin/activate
pip install -r requirements.txt

Key Learnings:

  • RAG quality depends heavily on chunk size and overlap tuning
  • Local LLM latency requires careful UX design (streaming is essential)
  • Memory retrieval needs relevance thresholds to avoid noise
  • Pydantic v1/v2 compatibility requires careful version pinning
  • ChromaDB stacking contexts cause subtle UI z-index bugs

License

MIT License

Copyright (c) 2026 Nexus Memory

Permission is hereby granted, free of charge, to any person obtaining a copy
of this software and associated documentation files (the "Software"), to deal
in the Software without restriction, including without limitation the rights
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
copies of the Software, and to permit persons to whom the Software is
furnished to do so, subject to the following conditions:

The above copyright notice and this permission notice shall be included in all
copies or substantial portions of the Software.

Built with ❀️ using local AI β€” no cloud required

⬆ Back to Top