ProfillyBot / README.md
MinhDS's picture
Update README.md
492077d verified
|
Raw
History Blame Contribute Delete
14.1 kB

A newer version of the Gradio SDK is available: 6.22.0

Upgrade
metadata
title: ProfillyBot
emoji: πŸ€–
colorFrom: gray
colorTo: indigo
sdk: gradio
sdk_version: 6.0.2
python_version: '3.11'
app_file: gradio_app.py
pinned: true
license: mit
short_description: RAG chatbot for professional profile Q&A
thumbnail: >-
  https://cdn-uploads.huggingface.co/production/uploads/656f08be4893a2a26d8b17c9/1jY92qjVxCyZ0_BVHVAOF.png

πŸ€– ProfillyBot

ProfillyBot Banner

Python 3.10 Python 3.11 Python 3.12 Docker Lint Code Quality Security Scan Codecov

ProfillyBot is a RAG (Retrieval Augmented Generation) chatbot that answers questions about your professional profile using your resume, project reports, and other documents.

✨ Features

  • πŸ“„ Multi-format Support: Process PDF, Word, HTML, and text documents
  • πŸ” Hybrid Search: Combines BM25 (keyword) + Vector (semantic) search with Reciprocal Rank Fusion
  • 🧠 Extensible Retrieval: Pluggable strategy system - easily add new retrieval methods (PageIndex, GraphRAG, etc.)
  • πŸ“Œ Main Document Support: Guaranteed context - critical information always available, auto-format detection
  • πŸ¦™ Ollama Integration: Run small language models locally
  • 🎨 Clean UI: Streamlit-based interface
  • βš™οΈ Highly Configurable: YAML-based settings for easy customization
  • πŸš€ HuggingFace Spaces Ready: Deploy with one click
  • ✨ Smart Response Enhancement: Automatically removes negative language and adds professional, recruiter-friendly tone

πŸ—οΈ Architecture

System Architecture Diagram

Overall system design showing the RAG pipeline and component interactions

Detailed architecture can be found in ARCHITECTURE.md.

πŸ› οΈ Tech Stack

  • Python 3.10+ with UV package manager
  • LangChain for RAG pipeline
  • ChromaDB for vector storage
  • rank-bm25 for lexical search
  • Ollama for local LLM serving
  • Streamlit for web interface
  • sentence-transformers for embeddings
  • tiktoken for token counting
  • Ruff for linting/formatting

πŸ“¦ Installation

Prerequisites

  1. Python 3.10+
  2. UV Package Manager: Install via:
    curl -LsSf https://astral.sh/uv/install.sh | sh
    
  3. Ollama: Install from ollama.ai

Setup

  1. Clone the repository:

    git clone <your-repo-url>
    cd profillybot
    
  2. Install dependencies with UV:

    uv venv
    source .venv/bin/activate  # On Windows: .venv\Scripts\activate
    uv pip install -r requirements.txt
    
  3. Set up environment variables:

    cp .env.template .env
    # Edit .env with your settings
    
  4. Configure the chatbot:

    • Edit config.yaml to set your name, title, and preferences
    • Adjust model settings, chunking parameters, and system prompt
  5. Add your documents:

    # Place your PDFs, Word docs, HTML files in:
    mkdir -p data/documents
    # Copy your resume, project reports, LinkedIn profile, etc.
    
  6. Create your Main Profile Document (πŸ“Œ Recommended):

    The Main Document feature ensures critical information is always included in responses, never missed by vector search.

    # Create main profile document (any format: .md, .txt, .pdf, .docx, .html)
    nano data/documents/main_profile.md
    

    Include essential information:

    • Full name, title, contact info
    • Current role and key responsibilities
    • Core skills and expertise
    • Major projects with metrics
    • Education and certifications

    Example structure:

    # Your Name
    **Title**: Your Professional Title
    **Email**: your.email@example.com
    
    ## Professional Summary
    Brief summary of your experience...
    
    ## Core Skills
    - Skill 1, Skill 2, Skill 3
    
    ## Current Role
    ### Company - Role (Dates)
    - Achievement 1 with metrics
    - Achievement 2 with metrics
    
    ## Education
    - Degree, University, Year
    

    Why use this?

    • βœ… Guaranteed Context: Critical info never missed by vector similarity search
    • βœ… Priority Positioning: Appears BEFORE retrieved chunks (higher LLM attention)
    • βœ… Auto-Format Detection: Supports MD, TXT, PDF, DOCX, HTML automatically
    • βœ… Smart Token Management: LLM-based summarization if content exceeds 10k tokens

    The feature is enabled by default in config.yaml. See docs/MAIN_DOCUMENT_GUIDE.md for advanced configuration.

  7. Pull Ollama model:

    ollama pull llama3.2:3b
    # Or other small models: phi3:mini, gemma2:2b
    

πŸš€ Usage

Local Development

  1. Build the retrieval index (default: hybrid BM25 + Vector):

    # Build with default strategy from config (bm25_vector)
    python -m src.build_vectorstore
    
    # Or specify a strategy explicitly
    python -m src.build_vectorstore --strategy bm25_vector  # Hybrid (recommended)
    python -m src.build_vectorstore --strategy vector       # Vector-only
    python -m src.build_vectorstore --strategy bm25         # BM25-only
    
  2. Run the Streamlit app:

    streamlit run app.py
    
  3. Open browser: Navigate to http://localhost:8501

Screenshot Example:

ProfillyBot UI Example

Streamlit chatbot interface in action

🐳 Docker Compose

Run the complete stack with Docker Compose (includes Ollama):

# Copy environment template
cp .env.template .env

# Start all services (app + Ollama)
docker compose up -d

# View logs
docker compose logs -f

# Stop services
docker compose down

Access the app at http://localhost:7860

πŸ“‹ Makefile Commands

Use the Makefile for common development tasks:

make help          # Show all available commands
Category Command Description
Setup make setup Initial project setup (copy env, install deps)
make install Install dependencies
make install-dev Install with dev dependencies
Development make run Run Streamlit app locally
make index-build Build vector and BM25 indexes
make index-rebuild Rebuild indexes from scratch
Quality make test Run tests
make test-cov Run tests with coverage
make lint Run linter (ruff)
make format Format code
make check Run all checks
Docker make docker-up Start all services
make docker-down Stop all services
make docker-rebuild Rebuild and restart
make docker-logs Show logs (follow mode)
make docker-clean Stop and remove volumes
Ollama make ollama-pull Pull configured model
make ollama-list List available models
Cleanup make clean Remove build artifacts
make clean-all Clean everything + indexes

🌐 Deployment to Hugging Face Spaces (ZeroGPU)

ProfillyBot on Spaces uses Gradio + ZeroGPU with Qwen/Qwen3.5-4B (gradio_app.py). Streamlit (app.py) remains available for local/Ollama/Groq.

Live Space

Deploy / update

  1. Create a Gradio Space with ZeroGPU hardware (Pro account required to create ZeroGPU Spaces).
  2. Ensure this repo has:
    • data/documents/MY_CV_2_0.pdf (or your profile docs)
    • Pre-built chroma_db/ + bm25_index/ (recommended) β€” or let the app build on first start
  3. Push to the Space:
    git remote add hf https://huggingface.co/spaces/MinhDS/ProfillyBot
    git push hf main
    
  4. Confirm Space Settings: SDK Gradio, hardware ZeroGPU, app file gradio_app.py.

Important Notes for HF Spaces

  • ZeroGPU is Gradio-only β€” do not use the Docker/Ollama Dockerfile on ZeroGPU
  • SLM: Qwen/Qwen3.5-4B via the transformers provider
  • Embeddings: all-MiniLM-L6-v2 on CPU (does not consume GPU quota)
  • Vector DB: Pre-build indexes locally when possible for faster cold start
  • Secrets: optional HF_TOKEN only if you switch to a gated model

βš™οΈ Configuration

config.yaml - Main Settings

profile:
  name: "Your Name"  # ← Change this!
  title: "Your Title"

llm:
  model: "llama3.2:3b"  # Choose your model
  temperature: 0.7

# Retrieval Strategy Configuration
retrieval:
  strategy: "bm25_vector"  # Options: vector, bm25, bm25_vector
  final_k: 4               # Number of documents to return

  vector:
    search_type: "similarity"
    k: 10                  # Docs to retrieve before fusion

  bm25:
    k: 10                  # Docs to retrieve before fusion
    tokenizer: "simple"

  fusion:
    algorithm: "rrf"       # Reciprocal Rank Fusion
    weights:
      vector: 0.7          # 70% weight on semantic search
      bm25: 0.3            # 30% weight on keyword search

document_processing:
  chunk_size: 1000
  chunk_overlap: 200

.env - Environment Variables

OLLAMA_BASE_URL=http://localhost:11434
CHROMA_PERSIST_DIR=./chroma_db

πŸ“š Project Structure

profillybot/
β”œβ”€β”€ app.py                          # Streamlit app entry point
β”œβ”€β”€ pyproject.toml                  # UV/pip dependencies & config
β”œβ”€β”€ config.yaml                     # RAG & LLM settings
β”œβ”€β”€ .env.template                   # Environment variables template
β”œβ”€β”€ Makefile                        # Development commands
β”œβ”€β”€ docker-compose.yml              # Docker Compose configuration
β”œβ”€β”€ Dockerfile                      # Docker image (standalone with Ollama)
β”œβ”€β”€ Dockerfile.app                  # Docker image (for Compose)
β”œβ”€β”€ README.md
β”œβ”€β”€ data/
β”‚   └── documents/                  # Your profile documents
β”œβ”€β”€ src/
β”‚   β”œβ”€β”€ __init__.py
β”‚   β”œβ”€β”€ document_processor.py       # Load & chunk documents
β”‚   β”œβ”€β”€ vectorstore.py              # ChromaDB operations
β”‚   β”œβ”€β”€ llm_handler.py              # Ollama/LLM interface
β”‚   β”œβ”€β”€ rag_pipeline.py             # RAG chain logic
β”‚   β”œβ”€β”€ main_document_loader.py     # Main document management
β”‚   β”œβ”€β”€ response_enhancer.py        # Response post-processing
β”‚   β”œβ”€β”€ config_loader.py            # Load config.yaml & .env
β”‚   β”œβ”€β”€ build_vectorstore.py        # CLI to build indexes
β”‚   └── retrieval/                  # Extensible retrieval system
β”‚       β”œβ”€β”€ __init__.py
β”‚       β”œβ”€β”€ base.py                 # BaseRetrieverStrategy ABC
β”‚       β”œβ”€β”€ factory.py              # RetrieverFactory
β”‚       β”œβ”€β”€ fusion.py               # RRF and fusion algorithms
β”‚       β”œβ”€β”€ stores/
β”‚       β”‚   └── bm25_store.py       # BM25 index storage
β”‚       └── strategies/
β”‚           β”œβ”€β”€ vector.py           # Vector-only strategy
β”‚           β”œβ”€β”€ bm25.py             # BM25-only strategy
β”‚           └── bm25_vector.py      # Hybrid BM25+Vector strategy
β”œβ”€β”€ chroma_db/                      # Vector database (auto-generated)
β”œβ”€β”€ bm25_index/                     # BM25 index (auto-generated)
└── tests/                          # Unit tests

🎯 Recommended Models for HF Spaces

Model Size Speed Quality HF Spaces Tier
llama3.2:3b 3B Fast Good Free βœ…
phi3:mini 3.8B Fast Good Free βœ…
gemma2:2b 2B Very Fast Decent Free βœ…
llama3.1:8b 8B Medium Great Upgraded GPU

πŸ” LLM Tracing with LangSmith

Monitor prompts, responses, and latency using LangSmith.

Setup

  1. Get API Key: Sign up at smith.langchain.com (free tier: 5,000 traces/month)

  2. Configure .env:

    LANGCHAIN_TRACING_V2=true
    LANGCHAIN_API_KEY=your_api_key_here
    LANGCHAIN_PROJECT=profillybot
    
  3. Restart the app:

    streamlit run app.py
    

What You Can See

  • Full prompts sent to LLM (system prompt + context + question)
  • Complete LLM responses
  • Latency breakdown per step
  • Token usage (when available)

πŸ”§ Troubleshooting

Ollama Connection Issues

# Check Ollama is running
ollama list

# Start Ollama service
ollama serve

ChromaDB / Index Errors

# Rebuild all indexes from scratch
rm -rf chroma_db/ bm25_index/
python -m src.build_vectorstore --force-rebuild

HuggingFace Spaces Issues

  • Check logs in the Space's "Logs" tab
  • Ensure requirements.txt is generated: uv pip compile pyproject.toml -o requirements.txt
  • Verify GPU/CPU settings match your model size

πŸ“„ License

MIT License - see LICENSE file


Note: Remember to update config.yaml with your personal information before deploying!