Spaces:
Sleeping
Sleeping
A newer version of the Gradio SDK is available: 6.26.0
metadata
title: RAG Python System
emoji: π€
colorFrom: blue
colorTo: green
sdk: gradio
sdk_version: 4.44.1
app_file: app.py
pinned: false
RAG System - Intelligent Document Q&A
A complete RAG (Retrieval-Augmented Generation) system for intelligent document search and question answering using Hugging Face Inference API.
π Features
- π Document Conversion: Automatic conversion of PDF, DOCX, TXT files to markdown
- π§© Smart Splitting: Text chunking with context preservation (LangChain)
- π Vector Search: Fast semantic search across documents (ChromaDB)
- π€ Local LLM: Answer generation using Ollama (no cloud data transfer)
- π Web Interface: User-friendly Gradio interface with streaming responses
- π Multilingual: Support for English, Russian, and Ukrainian languages
ποΈ Architecture
Documents β Conversion (PyMuPDF) β Splitting (LangChain)
β
User Question β Search (ChromaDB) β Context + Question
β
LLM (Ollama llama3.2)
β
Answer + Sources
π Requirements
- Python 3.9+
- Ollama (for local LLM model execution)
- 4+ GB RAM (for embedding model and LLM)
π Installation
1. Install Ollama
# macOS / Linux
curl -fsSL https://ollama.com/install.sh | sh
# After installation, download the model
ollama pull llama3.2
2. Clone and Setup Project
# Navigate to project directory
cd /Users/v.hirenko/Desktop/DevHubVault/my-ai-projects/rag-python-rag
# Create virtual environment
python3 -m venv venv
# Activate virtual environment
source venv/bin/activate # Linux/macOS
# or
venv\Scripts\activate # Windows
# Install dependencies
pip install -r requirements.txt
π Project Structure
rag-python-rag/
βββ config.py # System configuration
βββ document_converter.py # Document conversion
βββ text_splitter.py # Text chunking
βββ vector_store.py # Vector storage
βββ llm_handler.py # LLM request handling
βββ main.py # Main application
βββ requirements.txt # Dependencies
βββ documents/ # Source documents
βββ processed_docs/ # Converted documents
βββ chroma_db/ # Vector database
π― Usage
Quick Start
# Activate virtual environment
source venv/bin/activate
# Run the application
python main.py
The application will automatically:
- Download test document (Think Python PDF)
- Convert it to markdown
- Split into chunks
- Create vector database
- Launch web interface at http://localhost:7860
Adding Your Own Documents
- Place documents (PDF, DOCX, TXT) in the
documents/folder - Restart the application or run indexing:
python -c "
from main import RAGSystem
rag = RAGSystem()
rag.setup_pipeline(force_rebuild=True)
"
Using Python API
from vector_store import retrieve_context
from llm_handler import generate_answer, format_response
# Ask a question
question = "How do loops work in Python?"
# Get context from documents
context, sources = retrieve_context(question, n_results=5)
# Generate answer
answer = generate_answer(question, context)
# Format result
response = format_response(question, answer, sources)
print(response)
Streaming Answer Generation
from vector_store import retrieve_context
from llm_handler import stream_llm_answer
question = "What are Python functions?"
context, sources = retrieve_context(question)
# Stream output
for token in stream_llm_answer(question, context):
print(token, end='', flush=True)
βοΈ Configuration
Main settings are in config.py:
# Embedding model
EMBEDDING_MODEL = "all-MiniLM-L6-v2"
# LLM model
OLLAMA_MODEL = "llama3.2"
# Text splitting parameters
TEXT_SPLITTER_CONFIG = {
"chunk_size": 1000,
"chunk_overlap": 200,
}
# Number of search results
DEFAULT_N_RESULTS = 5
π§ͺ Testing Components
Document Conversion
python document_converter.py
Text Chunking
python text_splitter.py
Vector Store
python vector_store.py
LLM Handler
python llm_handler.py
π§ Troubleshooting
Issue: Model not found
# Check available models
ollama list
# Download required model
ollama pull llama3.2
Issue: Out of memory
- Reduce
chunk_sizeinconfig.py - Reduce
DEFAULT_N_RESULTS - Use a lighter model (e.g.,
llama3.2:1b)
Issue: Slow generation
- Use a faster model
- Reduce number of search results
- Consider using GPU version of Ollama
π Performance
On Think Python document (300+ pages):
- Conversion: ~5 seconds
- Indexing: ~30 seconds (847 chunks)
- Search: < 1 second
- Answer generation: 5-15 seconds (depends on length)
π£οΈ Roadmap
- Support more formats (Excel, PowerPoint)
- Embedding caching
- REST API endpoints
- Multimodal documents (images)
- Chat history and dialogue context
- Deploy to Hugging Face Spaces
π Sources and Inspiration
Project based on article: How I Built a RAG System in One Evening
Technologies Used:
- PyMuPDF - PDF conversion
- LangChain - text splitting
- ChromaDB - vector database
- Sentence Transformers - embeddings
- Ollama - local LLM models
- Gradio - web interface
π License
This project is created for educational purposes. Use freely!
π€ Contributing
If you want to improve the project:
- Fork the repository
- Create a feature branch
- Commit your changes
- Push to the branch
- Create a Pull Request
π§ Contact
If you have questions or suggestions, create an Issue in the repository.
Made with β€οΈ for learning RAG systems and local LLMs