docling-inference / setup.md
mc0117's picture
xlsx working, room for improvments
132dd49
|
Raw
History Blame Contribute Delete
9.22 kB
# Upload Service
Document upload API: accept a **PDF, DOCX, or XLSX**, process it with DOCLING, embed chunks with OpenAI, and upsert to Pinecone. Status updates are pushed over WebSocket and available via a job status endpoint.
> **Hugging Face:** Space config lives in `README.md` (YAML). This file is the full human-readable setup and API guide.
## Requirements
- Python 3.10+
- **Deps:** `requirements.txt` includes **`requirements_local.txt`** (CPU PyTorch). For GPU servers or Hugging Face GPU Spaces use **`requirements_prod.txt`**.
- **PDF**, **DOCX**, or **XLSX** input files. Some DOCX files with embedded drawings need **LibreOffice** installed (`libreoffice` on PATH). Excel does not require LibreOffice.
- PDFs with an embedded text layer (for fast path; optional OCR for scanned PDFs)
## Setup
1. Create and activate a virtual environment (e.g. `.upload_venv`).
2. Install dependencies:
```bash
pip install -r requirements.txt # same as requirements_local.txt (CPU)
# or for GPU / prod (CUDA):
# pip install -r requirements_prod.txt
```
3. Create a Pinecone index with **dimension 1024** (or set `EMBEDDING_DIMENSIONS` to match your index).
4. Set environment variables (e.g. in `.env` in the project root):
- `PINECONE_API_KEY`
- `PINECONE_INDEX_NAME`
- `OPENAI_API_KEY`
- For **S3 URLs** (e.g. `s3://bucket/key`): `AWS_ACCESS_KEY_ID`, `AWS_SECRET_ACCESS_KEY`, and optionally `AWS_REGION`
## Environment variables
| Variable | Required | Default | Description |
|----------|----------|---------|-------------|
| `PINECONE_API_KEY` | Yes | — | Pinecone API key |
| `PINECONE_INDEX_NAME` | Yes | — | Pinecone index name (dimension must match `EMBEDDING_DIMENSIONS`) |
| `OPENAI_API_KEY` | Yes | — | OpenAI API key (for embeddings) |
| `EMBEDDING_MODEL` | No | `text-embedding-3-small` | OpenAI embedding model |
| `EMBEDDING_DIMENSIONS` | No | `1024` | Vector dimension (must match Pinecone index) |
| `UPLOAD_TEMP_DIR` | No | `/tmp/upload-service` | Directory for temporary uploaded files |
| `PINECONE_NAMESPACE` | No | `""` | Pinecone namespace (optional) |
| `SYMBIOS_BACKEND_BASE_URL` | No | `""` | Base URL for callbacks; service POSTs to `{base}/update` (progress) and `{base}/finish` (completion) |
| `SYMBIOS_BACKEND_AUTH_SECRET` | No | `""` | If set, POST /api/vectorize requires `Authorization: <secret>` or `Bearer <secret>` |
| `DOCLING_DEVICE` | No | `auto` | DOCLING accelerator: `auto` (use GPU if available), `cpu`, `cuda`, or `mps` (macOS) |
| `AWS_ACCESS_KEY_ID` | For S3 | — | Required when vectorize URL is `s3://...` |
| `AWS_SECRET_ACCESS_KEY` | For S3 | — | Required when vectorize URL is `s3://...` |
| `AWS_REGION` | No | — | AWS region for S3 (optional; boto3 may infer) |
| `DEBUG_CHUNKS_DIR` | No | `""` | If set, dumps a markdown file per document with all extracted chunks to this directory (for debugging bad chunks) |
## GPU (optional)
DOCLING can run on CPU or GPU. By default `DOCLING_DEVICE=auto` uses GPU when available.
- **CPU only:** Use the default PyTorch from pip, or install CPU wheels:
`pip install torch --index-url https://download.pytorch.org/whl/cpu`
- **GPU (CUDA):** Install PyTorch with CUDA, then install the rest:
`pip install torch --index-url https://download.pytorch.org/whl/cu121`
then `pip install -r requirements_prod.txt`
Set `DOCLING_DEVICE=cuda` or leave `auto`.
At startup and when processing, the service logs which device DOCLING is using (e.g. `DOCLING accelerator: CUDA`).
## Run
From the project root:
```bash
./run.sh
```
Or:
```bash
.upload_venv/bin/uvicorn app.main:app --host 0.0.0.0 --port 8000
```
Or (e.g. Hugging Face Space, port 7860):
```bash
python serve.py
```
Server: `http://localhost:8000` (local) or the Space URL (HF).
## API
### POST /upload
Upload a document for processing. Returns immediately with a doc_id; processing runs in the background.
- **Request:** `multipart/form-data` with a file field named `file` (`.pdf`, `.docx`, `.xlsx`).
- **Response:** `202 Accepted` with JSON:
```json
{
"doc_id": "uuid",
"message": "Processing started",
"status_ws_url": "/ws?doc_id=uuid"
}
```
**Example:** see [Test curls](#test-curls) below.
### GET /job/{doc_id}
Poll job status and result. Use the same doc_id you got from POST /upload or POST /api/vectorize.
- **Response:** JSON with `status` (`processing` | `uploaded` | `error`), optional `chunk_count`, and optional `error` message.
**Example:** see [Test curls](#test-curls) below.
### WebSocket /ws?doc_id={doc_id}
Connect with a doc_id to receive live status updates (e.g. `processing_embeddings`, `uploaded` with `chunk_count`, or `error` with `message`).
## Pipeline
1. The file is saved temporarily and a doc_id is created.
2. **DOCLING** extracts text and structure (chunks with source, covered_pages, label). OCR and table-structure models are disabled by default for speed; see `app/services/docling_processor.py`.
3. **OpenAI** embeddings (configurable model and dimensions, with retry/backoff on rate limits).
4. **Pinecone** upsert: each chunk is stored as a vector with metadata `source`, `covered_pages` (list), `label`, and `text` (chunk content, truncated if needed for metadata size limits). Metadata also includes `doc_id` and (when provided) `project_id`.
5. Status is updated and broadcast to WebSocket clients; temp file is removed.
## Logging
At startup the service logs the DOCLING accelerator in use (e.g. `DOCLING accelerator: CUDA` or `CPU`). During each job it logs the current stage: download, running DOCLING (parsing document), DOCLING done, running embeddings, embeddings done, running Pinecone upsert, Pinecone upsert done, and job complete or error. All at INFO level. Logging is configured in `app/main.py`.
## Debugging chunks
If chunks look wrong (missing text, bad splits, garbled content), enable chunk debug dumps:
```bash
DEBUG_CHUNKS_DIR=/tmp/chunk-debug uvicorn app.main:app --host 0.0.0.0 --port 8000
```
Or in `.env`:
```
DEBUG_CHUNKS_DIR=/tmp/chunk-debug
```
After each document is processed, a file like `chunk-debug/<doc_id>.<format>.chunks.md` is written (one file per job; format is sniffed from file contents). Leave `DEBUG_CHUNKS_DIR` unset in production — no files are written.
## Test curls
Quick copy-paste examples (server at `http://localhost:8000`). Use the returned or chosen doc_id to poll and for WebSocket.
```bash
# Upload a document (returns doc_id)
curl -X POST -F "file=@/path/to/document.pdf" http://localhost:8000/upload
curl -X POST -F "file=@/path/to/document.docx;type=application/vnd.openxmlformats-officedocument.wordprocessingml.document;filename=test.docx" http://localhost:8000/upload
# Poll job status (replace DOC_ID with the returned doc_id)
curl http://localhost:8000/job/DOC_ID
# Vectorize by URL (you choose doc_id and project_id; optional backend for callbacks)
curl -X POST http://localhost:8000/api/vectorize \
-H "Content-Type: application/json" \
-d '{"url":"https://example.com/sample.pdf","doc_id":"test-doc-1","project_id":"proj-1"}'
# Poll by the same doc_id
curl http://localhost:8000/job/test-doc-1
```
WebSocket for live updates: `ws://localhost:8000/ws?doc_id=DOC_ID`
## Testing
### 1. File upload (POST /upload)
Start the server, then use the [test curls](#test-curls) above:
```bash
curl -X POST -F "file=@/path/to/document.pdf" http://localhost:8000/upload
# Use the returned doc_id:
curl http://localhost:8000/job/DOC_ID
```
### 2. Vectorize by URL (POST /api/vectorize)
Requires `doc_id`, `project_id`, and a document `url` (`.pdf`, `.docx`, or `.xlsx`). Optionally set `SYMBIOS_BACKEND_BASE_URL` so the service can POST progress to `/update` and completion to `/finish`.
**Without a real backend** – run a mock that prints every POST:
```bash
# Terminal 1: mock backend (receives /update and /finish)
.upload_venv/bin/python scripts/mock_backend.py
# Terminal 2: set base URL and start upload service
export SYMBIOS_BACKEND_BASE_URL=http://localhost:9999
.upload_venv/bin/uvicorn app.main:app --host 0.0.0.0 --port 8000
# Terminal 3: trigger a job (use a public PDF URL)
curl -X POST http://localhost:8000/api/vectorize \
-H "Content-Type: application/json" \
-d '{"url":"https://example.com/sample.pdf","doc_id":"test-doc-1","project_id":"proj-1"}'
```
Use a real PDF URL (e.g. `https://arxiv.org/pdf/2408.09869`) to run the full pipeline; the mock backend will show progress and finish callbacks.
In Terminal 1 you’ll see POSTs to `/update` (progress: Downloading…, Parsing document…, etc.) and finally POST to `/finish` with `status: complete` or `failure`. Use the same doc_id to poll or connect via WebSocket:
```bash
curl http://localhost:8000/job/test-doc-1
# or connect to ws://localhost:8000/ws?doc_id=test-doc-1
```
**With backend auth:** set `SYMBIOS_BACKEND_AUTH_SECRET` and send `Authorization: <secret>` or `Authorization: Bearer <secret>` on POST /api/vectorize.
### 3. Optional: task_auth_token
If your backend expects a per-task token, include it in the body; the service sends it as the `Authorization` header when calling `/finish` and `/update`. It can be omitted for local testing.