Spaces:
Running on Zero
Running on Zero
File size: 8,507 Bytes
6a46e44 dd3eb25 6a46e44 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 | ---
title: Hybrid Vision RAG PDF Processor
emoji: π
colorFrom: indigo
colorTo: blue
sdk: gradio
sdk_version: "5.20.0"
app_file: main.py
pinned: false
license: mit
---
# Hybrid Vision RAG PDF Processor
A demo-ready Python project that extracts PDF content, analyzes charts/graphs via a Vision LLM, and answers questions with a Hybrid RAG system (Vector + Knowledge Graph + BM25).
## Tech stack
- **Docling** β PDF content extraction (text, pages, figures)
- **OpenRouter** β Vision LLM for chart/graph analysis
- **Groq** β LLM powering the RAG pipeline
- **LlamaIndex** β Vector + Knowledge Graph + BM25 retrieval with RRF fusion
- **Gradio** β demo UI
## What's new
- β
**Clean layered architecture** β `config`, `models`, `services`, `retrieval`, `session`, `web`
- β
**Own session architecture** β `Session` entity + pluggable `SessionStore` + `SessionManager` with janitor
- β
**Pydantic configuration** β all settings loaded from `.env` with validation
- β
**Structured logging** β no more raw `print()` statements
- β
**Polished Gradio demo UI** β live progress, status badge, commands, and session controls
- β
**Pre-flight API validation** β OpenRouter + Groq are live-checked before the pipeline starts; failures pop a warning and block processing (see the **API Status** panel)
- β
**Unit tests** for session lifecycle and BM25 retrieval
- β
**Unified CLI** β `extract`, `preprocess`, `app`
## Quick start
### 1. Install dependencies
```bash
uv sync
# or
pip install -e .
```
### 2. Configure API keys
Copy `.env.example` to `.env` and add your keys:
```bash
cp .env.example .env
```
```env
OPENROUTER_API_KEY=sk-or-v1-...
GROQ_API_KEY=gsk_...
```
### 3. Launch the demo
```bash
app
# or
python -m docling_pdf_processor
```
Open `http://localhost:7860` in your browser.
---
## Docker
Run the demo in a container (secrets are injected at runtime β never baked into the image):
```bash
# 1. Put real keys in .env (the image does NOT see your host's OS env vars)
# OPENROUTER_API_KEY=sk-or-v1-...
# GROQ_API_KEY=gsk_...
# 2. Build & run
docker compose up --build
```
Then open `http://localhost:8010`.
Without compose:
```bash
docker build -t docling-pdf-processor .
docker run -p 8010:8010 --env-file .env docling-pdf-processor
```
The compose file mounts a named volume at `/app/.cache/huggingface` so Docling's model weights download once and persist across restarts.
> β οΈ **Key source caveat for containers:** on your host, `GROQ_API_KEY` may live in an OS environment variable that overrides the `.env` placeholder (the API Status panel shows `[source: OS env $GROQ_API_KEY]`). A container does **not** inherit host OS env vars β only what you pass via `--env-file` / `env_file`. So you must put the **real** Groq key in `.env` before `docker compose up`, or the pre-flight validation will block processing with `401 Invalid API key`.
---
## CLI usage
```bash
# Extract images + markdown from a PDF
extract --pdf path/to/doc.pdf --output ./extracted_images
# Preprocess extracted images with a Vision LLM
preprocess --quarters ./extracted_images --output graphs_description.pkl
# Launch the Gradio demo
app
```
---
## Project architecture
```
docling_pdf_processor/
βββ config.py # Pydantic Settings from .env
βββ models.py # Shared dataclasses
βββ exceptions.py # Domain errors
βββ logging_config.py # Structured logging setup
βββ cli.py # Unified CLI
βββ pipeline.py # Session-aware PipelineOrchestrator
βββ services/
β βββ extractor.py # Docling PDF extraction
β βββ vision.py # OpenRouter Vision-LLM
β βββ rag.py # Hybrid RAG builder + query
βββ retrieval/
β βββ bm25.py # BM25 keyword retriever
β βββ hybrid_rrf.py # RRF fusion retriever
βββ session/
β βββ models.py # Session + SessionStatus
β βββ store.py # SessionStore protocol + implementations
β βββ manager.py # Lifecycle + janitor
βββ web/
βββ gradio_app.py # Demo UI
βββ components.py # UI helpers
```
### Data flow
```
User uploads PDF
β
βΌ
βββββββββββββββββββββββ
β SessionManager β Creates UUID workspace in temp dir
β (own session arch) β
ββββββββββ¬βββββββββββββ
β
βΌ
βββββββββββββββββββββββ
β PipelineOrchestratorβ Streams progress via Queue
β β’ ExtractorService β Docling β text, pages, figures
β β’ VisionService β OpenRouter β graph descriptions
β β’ RAGService β Vector + KG + BM25 indexes
ββββββββββ¬βββββββββββββ
β
βΌ
βββββββββββββββββββββββ
β Gradio Chat UI β Session-scoped memory
β !reasoning β Inspect retriever contributions
β !compare β Side-by-side retriever comparison
βββββββββββββββββββββββ
```
---
## Session architecture
Every upload creates a `Session` with:
- A unique UUID
- An isolated workspace under `tempdir/docling_sessions/{uuid}/`
- A `SessionStatus` lifecycle: `created β extracting β analyzing β indexing β ready`
- A heartbeat timestamp for stale-session cleanup
The `SessionStore` protocol has two built-in implementations:
- `InMemorySessionStore` β default for single-process demos
- `FileSystemSessionStore` β persists session metadata to disk
A background janitor cleans sessions idle longer than `SESSION_MAX_AGE_SECONDS` (default 30 min).
---
## Chat commands
- `!reasoning <query>` β show how Vector / KG / BM25 contribute to the final answer
- `!compare <query>` β compare Vector, KG, BM25, and Hybrid answers side-by-side
---
## API validation
Before the pipeline runs, the app live-checks the two external APIs it depends on:
| API | Check | Endpoint |
|-----|-------|----------|
| OpenRouter (Vision LLM) | key + credits | `GET https://openrouter.ai/api/v1/key` |
| Groq (RAG LLM) | key + model **actually generate** (1-token chat completion) | `POST https://api.groq.com/openai/v1/chat/completions` |
- The **API Status** panel in the *Process PDF* tab shows each API as β
Valid / β Invalid, with a **masked key prefix and the key's source** (`UI field` / `OS env $VAR` / `.env`) so it's clear which credential is being tested.
- Click **Validate APIs** to re-run the checks on demand; editing any key/model field live-refreshes the panel.
- When you click **Process PDF**, validation runs first. If any API fails (bad key β 401, wrong model β 404, no credits β 402), a popup warning shows which API and why, and the pipeline **does not start**.
> **Note on key sources:** pydantic-settings loads OS environment variables *with higher priority than* `.env`. So if a key is set in your shell/system environment, it overrides a placeholder in `.env`. The source tag in the panel makes this visible β e.g. `OS env $GROQ_API_KEY` instead of `.env`.
Validation calls are cheap and run in parallel, typically returning in ~1s.
---
## Production notes
- **Concurrency** β Gradio runs with `default_concurrency_limit=1` so only one pipeline processes at a time per process instance.
- **Multi-tenant scaling** β For true multi-user production, run each session in a separate worker or process.
- **Vector store** β Uses LlamaIndex's in-memory `SimpleVectorStore` (the previous DeepLake store is abandoned and breaks on NumPy 2 / Python 3.13). Indexes are rebuilt from the extracted files each session, which is the existing behavior, so nothing is lost β but for very large corpora consider swapping in a persisted store (FAISS, Chroma, pgvector).
- **API keys** β Never commit `.env`. It is already ignored by `.gitignore`.
- **Large PDFs** β PDFs > 100 MB trigger a warning but are not blocked.
---
## Development
### Run tests
```bash
pytest
```
### Run a specific stage from the CLI
```bash
# Extraction only
extract --pdf sample.pdf --output ./out
# Vision analysis only
preprocess --quarters ./out --output ./out/graphs.pkl --api-key $OPENROUTER_API_KEY
```
---
## License
MIT
|