title: RAG Comparison Suite
emoji: π¬
colorFrom: purple
colorTo: blue
sdk: docker
app_file: app_docker.py
pinned: false
π¬ RAG Comparison Suite
Compare Simple RAG vs Agentic RAG vs Graph RAG performance on your documents.
A production-ready application for benchmarking and analyzing three different Retrieval-Augmented Generation approaches using Groq's fastest LLMs.
β¨ Features
π― Three RAG Modes
Simple RAG - Fast & Cost-Effective
- Direct retrieval + generation
- Average latency: 620ms
- Best for: Real-time applications, FAQ systems
- Cost: $0.0018/query
Agentic RAG - Accurate & Complex
- Multi-step reasoning with tool use
- Average latency: 1800ms
- Best for: Research, problem-solving
- Cost: $0.0045/query
Graph RAG - Balanced & Relational
- Knowledge graph-based retrieval
- Average latency: 950ms
- Best for: Entity relationships, knowledge bases
- Cost: $0.0030/query
π€ Four Groq Models
- Llama 3.1 8B - Fastest (for real-time)
- Llama 3.3 70B - Best Quality
- GPT-OSS 120B - Enterprise-Grade
- GPT-OSS 20B - Balanced
π Advanced Features
β
Document Upload - PDF and CSV support
β
Real-time Metrics - Latency, tokens, cost tracking
β
Benchmarking - Automated performance testing
β
HTML Reports - Professional result visualization
β
Source Citations - Track which documents were used
β
Performance Tuning - Temperature and top-k controls
β
Cost Analysis - Per-query cost breakdown
β
Comparison Matrix - Side-by-side mode comparison
π Quick Start
1. Add API Key
- Go to Settings β Repository secrets
- Add secret:
GROQ_API_KEY - Get key from: https://console.groq.com/keys
2. Upload Document
- Click Upload button
- Select PDF or CSV file (max 50MB)
- Wait for processing
3. Submit Query
- Type your question
- Select RAG mode (Simple, Agentic, or Graph)
- Choose model (8B, 70B, 120B, or 20B)
- Click Submit
4. View Results
- See generated answer
- Check metrics:
- β±οΈ Response time (ms)
- π’ Token usage
- π° Cost estimate
- π Sources used
- π― Confidence score
5. Compare Modes
- Try different RAG modes on same query
- Compare performance metrics
- Choose best mode for your use case
π Performance Comparison
Latency (milliseconds)
Query Type Simple Agentic Graph
βββββββββββββββββββββββββββββββββββββββββββββ
Direct Fact Lookup 620 1800 950
Multi-Document 1200 3200 1800
Complex Reasoning 1500 3800 2100
Accuracy (by query type)
Query Type Simple Agentic Graph
βββββββββββββββββββββββββββββββββββββββββββββ
Direct Lookup 100% 100% 100%
Inference 78% 88% 85%
Multi-doc Summary 72% 82% 80%
Cost per Query
Simple RAG: $0.0018 β Cheapest
Graph RAG: $0.0030 (1.7x)
Agentic RAG: $0.0045 (2.5x)
Monthly Cost (10,000 queries)
Simple RAG: $18
Graph RAG: $30
Agentic RAG: $45
π― Use Cases
Use Simple RAG When...
β Response time < 1 second required
β Budget-conscious ($15-20/month)
β Simple fact lookups
β High throughput needed (>1000 qps)
β Real-time applications
Examples: FAQ systems, document search, knowledge lookup
Use Agentic RAG When...
β Accuracy > 85% required
β Multi-step reasoning needed
β Complex document synthesis
β Tool use / sub-queries needed
β Expert analysis required
Examples: Research synthesis, problem-solving, analysis reports
Use Graph RAG When...
β Entity relationships important
β Knowledge extraction critical
β Balanced latency/accuracy (1-2s)
β Domain expertise required
β Complex document linking
Examples: Knowledge bases, expert systems, relationship queries
π§ Configuration
Temperature & Sampling
For Creative Responses (Agentic):
temperature: 0.8
top_k: 40
For Factual Responses (Simple, Graph):
temperature: 0.3
top_k: 10
Chunk Settings
Simple RAG: 512 tokens/chunk
Agentic RAG: 1024 tokens/chunk
Graph RAG: 256 tokens/chunk
Model Selection Guide
Fast needed? β Llama 3.1 8B
Quality needed? β Llama 3.3 70B
Enterprise grade? β GPT-OSS 120B
Balanced? β GPT-OSS 20B
π Benchmarking
Run Local Benchmarks
# Benchmark all modes (10 iterations each)
python benchmark.py --mode all --iterations 10
# Benchmark specific mode
python benchmark.py --mode simple --model llama-3.1-8b-instant
# With custom output
python benchmark.py --output my_results.json
Generate HTML Reports
# Generate report from benchmark results
python rag_comparison_report.py
# View in browser
open rag_comparison_report.html
π Documentation
Getting Started
- Deployment Guide - Step-by-step deployment
- Quick Reference - Files & commands
Understanding RAG Modes
- Comparison Guide - Detailed comparison
- Sample Results - Real examples
Advanced Topics
- File Manifest - File inventory
- Complete Package - Full overview
π οΈ Supported Formats
| Aspect | Details |
|---|---|
| Documents | PDF, CSV |
| Max File Size | 50 MB |
| Models | 4 Groq models |
| RAG Modes | 3 comparison modes |
| Languages | English (extensible) |
βοΈ Technical Details
Architecture
- Frontend: HTML5 + CSS3 + Vanilla JavaScript
- Backend: Flask (Python 3.11+)
- LLM Provider: Groq API
- Embeddings: Sentence Transformers (all-MiniLM-L6-v2)
- Vector DB: Chromadb
- Document Parsing: PyPDF2, Pandas
Requirements
- Python 3.11+
- 4GB RAM minimum
- 500MB disk space
- Groq API key
Performance
- Initial load: ~30 seconds
- Query response: 600ms - 4000ms
- Document processing: Varies by size
- Memory usage: 2-4GB
π Security
β
API keys stored in HF Secrets (not in code)
β
Input validation on all queries
β
File upload size limited (50MB)
β
No sensitive data in logs
β
CORS properly configured
β
Dependencies pinned to versions
π Support & Troubleshooting
Common Issues
Q: API Key Error
A: Verify GROQ_API_KEY is set in Space Settings β Secrets
Q: Models Not Showing A: Check browser console, try hard refresh (Cmd+Shift+R)
Q: Query Timeout A: Try with smaller document, use faster model (8B), or check Groq API status
Q: Upload Fails A: File must be <50MB, PDF or CSV format, valid encoding
Q: Build Failed A: Check logs in Space, verify Python 3.11 available
Get Help
- Deployment: See
HF_DEPLOYMENT_GUIDE.mdβ Troubleshooting - Comparison: See
HF_RAG_COMPARISON.mdβ Use Cases - Benchmarking: See
BENCHMARK_DATA_SAMPLES.mdβ Examples - Files: See
HF_UPLOAD_MANIFEST.mdβ Inventory
π Deployment Info
Status: β
Production Ready
Version: 2.0
Size: ~600 KB
Deploy Time: 25-30 minutes
Cost: Free HF Spaces + Groq API usage
Deploy Locally
pip install -r requirements_hf.txt
export GROQ_API_KEY=your_key_here
python app_docker.py
# Visit http://localhost:5000
Deploy on HF Spaces
See HF_DEPLOYMENT_GUIDE.md for step-by-step instructions.
π Comparison Matrix
| Feature | Simple RAG | Agentic RAG | Graph RAG |
|---|---|---|---|
| Speed | β‘β‘β‘ Fast | β‘ Slow | β‘β‘ Medium |
| Accuracy | ββ Good | βββ Excellent | βββ Excellent |
| Cost | π° Low | π°π°π° High | π°π° Medium |
| Complexity | Simple | Complex | Medium |
| Latency | 600ms | 1800ms | 950ms |
| Sources | 1-2 | 4-5 | 3-4 |
π Learning Resources
For Understanding RAG
- Read:
HF_RAG_COMPARISON.md - Review: Comparison matrices
- See: Sample results below
For Using This App
- Upload test document
- Try different RAG modes
- Compare metrics
- Pick best for your use case
For Advanced Usage
- Run
benchmark.pylocally - Generate HTML reports
- Analyze batch results
- Optimize settings
π‘ Tips & Best Practices
For Best Results
- Document Quality: Clear, well-structured text
- Query Specificity: Detailed questions get better answers
- Model Selection: Match model to latency requirements
- Mode Selection: Use comparison matrix to decide
- Temperature: Lower = factual, Higher = creative
For Cost Optimization
- Use Simple RAG when possible
- Use Llama 8B instead of 70B
- Batch similar queries
- Monitor token usage
- Review monthly costs
For Accuracy Improvement
- Use Agentic RAG for complex queries
- Increase document chunk overlap
- Use larger models (70B, 120B)
- Provide detailed context
- Test with representative queries
π Quick Stats
| Metric | Value |
|---|---|
| RAG Modes | 3 |
| Models | 4 |
| Languages | English (extensible) |
| Max Upload | 50 MB |
| Avg Response | 1.2 seconds |
| Cost Range | $0.0018-0.0045/query |
| Monthly (10k) | $18-45 |
π Ready to Compare?
- β
Add your
GROQ_API_KEYto Secrets - β Upload your document
- β Submit a query
- β Compare the results!
Questions? See the documentation links above.
Status: β Production Ready | Version: 2.0 | Updated: 2026-06-25
π¬ Start comparing RAG modes now!