Spaces:
Sleeping
Sleeping
| title: RAG Python System | |
| emoji: 🤖 | |
| colorFrom: blue | |
| colorTo: green | |
| sdk: gradio | |
| sdk_version: 4.44.1 | |
| app_file: app.py | |
| pinned: false | |
| license: mit | |
| # RAG Python System 🤖 | |
| A local-first Retrieval-Augmented Generation (RAG) system that enables intelligent question-answering over your documents using Hugging Face Inference API, ChromaDB, and Gradio. | |
| ## Features | |
| - 📄 **Document Processing**: Converts PDF, DOCX, and TXT files to searchable format | |
| - 🔍 **Semantic Search**: Uses sentence transformers for accurate context retrieval | |
| - 🤖 **AI-Powered Answers**: Leverages Llama 3.2 3B via Hugging Face Inference API | |
| - 💬 **Interactive UI**: Clean Gradio interface for easy interaction | |
| - 🎯 **Source Citations**: Provides references to source documents | |
| ## How It Works | |
| 1. **Document Ingestion**: Upload or use pre-loaded documents (currently includes "Think Python" guide) | |
| 2. **Semantic Chunking**: Documents are split into meaningful chunks | |
| 3. **Vector Embeddings**: Text chunks are converted to embeddings using `all-MiniLM-L6-v2` | |
| 4. **Context Retrieval**: Relevant chunks are retrieved based on your question | |
| 5. **Answer Generation**: Llama 3.2 generates answers using the retrieved context | |
| ## Setup Instructions | |
| ### For Hugging Face Spaces Deployment | |
| 1. **Fork or Duplicate this Space** | |
| - Click the three dots menu → "Duplicate Space" | |
| 2. **Get Your Hugging Face Token** | |
| - Go to [https://huggingface.co/settings/tokens](https://huggingface.co/settings/tokens) | |
| - Create a new token with `read` permissions | |
| - Copy the token (starts with `hf_`) | |
| 3. **Add Token to Space Secrets** | |
| - Go to your Space Settings | |
| - Navigate to "Repository secrets" | |
| - Add a new secret: | |
| - **Name**: `HF_TOKEN` | |
| - **Value**: Your token from step 2 | |
| - Click "Add" | |
| 4. **Restart the Space** | |
| - The Space will automatically rebuild and start | |
| ### For Local Development | |
| 1. **Clone the Repository** | |
| ```bash | |
| git clone https://github.com/monsara/rag-python-rag.git | |
| cd rag-python-rag | |
| ``` | |
| 2. **Create Virtual Environment** | |
| ```bash | |
| python -m venv venv | |
| source venv/bin/activate # On Windows: venv\Scripts\activate | |
| ``` | |
| 3. **Install Dependencies** | |
| ```bash | |
| pip install -r requirements.txt | |
| ``` | |
| 4. **Set Environment Variable** | |
| ```bash | |
| export HF_TOKEN=hf_your_token_here | |
| ``` | |
| 5. **Run the Application** | |
| ```bash | |
| python app.py | |
| ``` | |
| 6. **Access the Interface** | |
| - Open your browser to `http://localhost:7860` | |
| ## Usage Examples | |
| Try asking questions like: | |
| - "How do if-else statements work in Python?" | |
| - "What are the different types of loops in Python?" | |
| - "How do you handle errors in Python?" | |
| - "Explain Python functions with examples" | |
| - "What is object-oriented programming in Python?" | |
| ## Architecture | |
| ``` | |
| User Query → Gradio UI → Vector Store (ChromaDB) → Context Retrieval | |
| ↓ | |
| HF Inference API | |
| ↓ | |
| Llama 3.2 3B | |
| ↓ | |
| Formatted Response | |
| ``` | |
| ## Tech Stack | |
| - **Frontend**: Gradio 4.44.1 | |
| - **LLM**: Llama 3.2 3B Instruct (via Hugging Face Inference API) | |
| - **Embeddings**: all-MiniLM-L6-v2 (Sentence Transformers) | |
| - **Vector DB**: ChromaDB | |
| - **Document Processing**: PyMuPDF, python-docx | |
| - **Text Splitting**: LangChain Text Splitters | |
| ## Rate Limits | |
| **Free Tier (Hugging Face Inference API):** | |
| - ~1000 requests/hour | |
| - 1024 max tokens per response | |
| - Shared infrastructure | |
| **For Production:** | |
| Consider upgrading to [Hugging Face Pro](https://huggingface.co/pricing) ($9/month) for: | |
| - Higher rate limits | |
| - Faster inference | |
| - Priority support | |
| ## Limitations | |
| - Currently uses a single pre-loaded document ("Think Python") | |
| - Free tier has rate limits | |
| - Response quality depends on context relevance | |
| - Max 1024 tokens per response | |
| ## Roadmap | |
| - [ ] File upload functionality | |
| - [ ] Multiple document support | |
| - [ ] Conversation history | |
| - [ ] Custom model selection | |
| - [ ] Advanced filtering options | |
| - [ ] Export conversation feature | |
| ## Contributing | |
| Contributions are welcome! Please feel free to submit a Pull Request. | |
| ## License | |
| MIT License - see LICENSE file for details | |
| ## Links | |
| - **GitHub Repository**: [monsara/rag-python-rag](https://github.com/monsara/rag-python-rag) | |
| - **Documentation**: See [README.md](README.md) for detailed technical documentation | |
| - **Hugging Face**: [Get your API token](https://huggingface.co/settings/tokens) | |
| ## Support | |
| If you encounter issues: | |
| 1. Check that `HF_TOKEN` is set correctly in Space secrets | |
| 2. Verify your token has `read` permissions | |
| 3. Check the Space logs for error messages | |
| 4. Open an issue on GitHub | |
| --- | |
| Built with ❤️ using Hugging Face, Gradio, and ChromaDB | |