rag-python-rag / README.md
viktor-hirenko
fix: add HF Space YAML metadata to README
cc50299
|
Raw
History Blame Contribute Delete
6.2 kB

A newer version of the Gradio SDK is available: 6.26.0

Upgrade
metadata
title: RAG Python System
emoji: πŸ€–
colorFrom: blue
colorTo: green
sdk: gradio
sdk_version: 4.44.1
app_file: app.py
pinned: false

RAG System - Intelligent Document Q&A

A complete RAG (Retrieval-Augmented Generation) system for intelligent document search and question answering using Hugging Face Inference API.

🌟 Features

  • πŸ“„ Document Conversion: Automatic conversion of PDF, DOCX, TXT files to markdown
  • 🧩 Smart Splitting: Text chunking with context preservation (LangChain)
  • πŸ” Vector Search: Fast semantic search across documents (ChromaDB)
  • πŸ€– Local LLM: Answer generation using Ollama (no cloud data transfer)
  • 🌐 Web Interface: User-friendly Gradio interface with streaming responses
  • 🌍 Multilingual: Support for English, Russian, and Ukrainian languages

πŸ—οΈ Architecture

Documents β†’ Conversion (PyMuPDF) β†’ Splitting (LangChain)
                                           ↓
User Question β†’ Search (ChromaDB) β†’ Context + Question
                                           ↓
                                 LLM (Ollama llama3.2)
                                           ↓
                                 Answer + Sources

πŸ“‹ Requirements

  • Python 3.9+
  • Ollama (for local LLM model execution)
  • 4+ GB RAM (for embedding model and LLM)

πŸš€ Installation

1. Install Ollama

# macOS / Linux
curl -fsSL https://ollama.com/install.sh | sh

# After installation, download the model
ollama pull llama3.2

2. Clone and Setup Project

# Navigate to project directory
cd /Users/v.hirenko/Desktop/DevHubVault/my-ai-projects/rag-python-rag

# Create virtual environment
python3 -m venv venv

# Activate virtual environment
source venv/bin/activate  # Linux/macOS
# or
venv\Scripts\activate  # Windows

# Install dependencies
pip install -r requirements.txt

πŸ“š Project Structure

rag-python-rag/
β”œβ”€β”€ config.py                # System configuration
β”œβ”€β”€ document_converter.py    # Document conversion
β”œβ”€β”€ text_splitter.py        # Text chunking
β”œβ”€β”€ vector_store.py         # Vector storage
β”œβ”€β”€ llm_handler.py          # LLM request handling
β”œβ”€β”€ main.py                 # Main application
β”œβ”€β”€ requirements.txt        # Dependencies
β”œβ”€β”€ documents/              # Source documents
β”œβ”€β”€ processed_docs/         # Converted documents
└── chroma_db/             # Vector database

🎯 Usage

Quick Start

# Activate virtual environment
source venv/bin/activate

# Run the application
python main.py

The application will automatically:

  1. Download test document (Think Python PDF)
  2. Convert it to markdown
  3. Split into chunks
  4. Create vector database
  5. Launch web interface at http://localhost:7860

Adding Your Own Documents

  1. Place documents (PDF, DOCX, TXT) in the documents/ folder
  2. Restart the application or run indexing:
python -c "
from main import RAGSystem
rag = RAGSystem()
rag.setup_pipeline(force_rebuild=True)
"

Using Python API

from vector_store import retrieve_context
from llm_handler import generate_answer, format_response

# Ask a question
question = "How do loops work in Python?"

# Get context from documents
context, sources = retrieve_context(question, n_results=5)

# Generate answer
answer = generate_answer(question, context)

# Format result
response = format_response(question, answer, sources)
print(response)

Streaming Answer Generation

from vector_store import retrieve_context
from llm_handler import stream_llm_answer

question = "What are Python functions?"
context, sources = retrieve_context(question)

# Stream output
for token in stream_llm_answer(question, context):
    print(token, end='', flush=True)

βš™οΈ Configuration

Main settings are in config.py:

# Embedding model
EMBEDDING_MODEL = "all-MiniLM-L6-v2"

# LLM model
OLLAMA_MODEL = "llama3.2"

# Text splitting parameters
TEXT_SPLITTER_CONFIG = {
    "chunk_size": 1000,
    "chunk_overlap": 200,
}

# Number of search results
DEFAULT_N_RESULTS = 5

πŸ§ͺ Testing Components

Document Conversion

python document_converter.py

Text Chunking

python text_splitter.py

Vector Store

python vector_store.py

LLM Handler

python llm_handler.py

πŸ”§ Troubleshooting

Issue: Model not found

# Check available models
ollama list

# Download required model
ollama pull llama3.2

Issue: Out of memory

  • Reduce chunk_size in config.py
  • Reduce DEFAULT_N_RESULTS
  • Use a lighter model (e.g., llama3.2:1b)

Issue: Slow generation

  • Use a faster model
  • Reduce number of search results
  • Consider using GPU version of Ollama

πŸ“Š Performance

On Think Python document (300+ pages):

  • Conversion: ~5 seconds
  • Indexing: ~30 seconds (847 chunks)
  • Search: < 1 second
  • Answer generation: 5-15 seconds (depends on length)

πŸ›£οΈ Roadmap

  • Support more formats (Excel, PowerPoint)
  • Embedding caching
  • REST API endpoints
  • Multimodal documents (images)
  • Chat history and dialogue context
  • Deploy to Hugging Face Spaces

πŸ“– Sources and Inspiration

Project based on article: How I Built a RAG System in One Evening

Technologies Used:

πŸ“ License

This project is created for educational purposes. Use freely!

🀝 Contributing

If you want to improve the project:

  1. Fork the repository
  2. Create a feature branch
  3. Commit your changes
  4. Push to the branch
  5. Create a Pull Request

πŸ“§ Contact

If you have questions or suggestions, create an Issue in the repository.


Made with ❀️ for learning RAG systems and local LLMs