# Upload Service Document upload API: accept a **PDF, DOCX, or XLSX**, process it with DOCLING, embed chunks with OpenAI, and upsert to Pinecone. Status updates are pushed over WebSocket and available via a job status endpoint. > **Hugging Face:** Space config lives in `README.md` (YAML). This file is the full human-readable setup and API guide. ## Requirements - Python 3.10+ - **Deps:** `requirements.txt` includes **`requirements_local.txt`** (CPU PyTorch). For GPU servers or Hugging Face GPU Spaces use **`requirements_prod.txt`**. - **PDF**, **DOCX**, or **XLSX** input files. Some DOCX files with embedded drawings need **LibreOffice** installed (`libreoffice` on PATH). Excel does not require LibreOffice. - PDFs with an embedded text layer (for fast path; optional OCR for scanned PDFs) ## Setup 1. Create and activate a virtual environment (e.g. `.upload_venv`). 2. Install dependencies: ```bash pip install -r requirements.txt # same as requirements_local.txt (CPU) # or for GPU / prod (CUDA): # pip install -r requirements_prod.txt ``` 3. Create a Pinecone index with **dimension 1024** (or set `EMBEDDING_DIMENSIONS` to match your index). 4. Set environment variables (e.g. in `.env` in the project root): - `PINECONE_API_KEY` - `PINECONE_INDEX_NAME` - `OPENAI_API_KEY` - For **S3 URLs** (e.g. `s3://bucket/key`): `AWS_ACCESS_KEY_ID`, `AWS_SECRET_ACCESS_KEY`, and optionally `AWS_REGION` ## Environment variables | Variable | Required | Default | Description | |----------|----------|---------|-------------| | `PINECONE_API_KEY` | Yes | — | Pinecone API key | | `PINECONE_INDEX_NAME` | Yes | — | Pinecone index name (dimension must match `EMBEDDING_DIMENSIONS`) | | `OPENAI_API_KEY` | Yes | — | OpenAI API key (for embeddings) | | `EMBEDDING_MODEL` | No | `text-embedding-3-small` | OpenAI embedding model | | `EMBEDDING_DIMENSIONS` | No | `1024` | Vector dimension (must match Pinecone index) | | `UPLOAD_TEMP_DIR` | No | `/tmp/upload-service` | Directory for temporary uploaded files | | `PINECONE_NAMESPACE` | No | `""` | Pinecone namespace (optional) | | `SYMBIOS_BACKEND_BASE_URL` | No | `""` | Base URL for callbacks; service POSTs to `{base}/update` (progress) and `{base}/finish` (completion) | | `SYMBIOS_BACKEND_AUTH_SECRET` | No | `""` | If set, POST /api/vectorize requires `Authorization: ` or `Bearer ` | | `DOCLING_DEVICE` | No | `auto` | DOCLING accelerator: `auto` (use GPU if available), `cpu`, `cuda`, or `mps` (macOS) | | `AWS_ACCESS_KEY_ID` | For S3 | — | Required when vectorize URL is `s3://...` | | `AWS_SECRET_ACCESS_KEY` | For S3 | — | Required when vectorize URL is `s3://...` | | `AWS_REGION` | No | — | AWS region for S3 (optional; boto3 may infer) | | `DEBUG_CHUNKS_DIR` | No | `""` | If set, dumps a markdown file per document with all extracted chunks to this directory (for debugging bad chunks) | ## GPU (optional) DOCLING can run on CPU or GPU. By default `DOCLING_DEVICE=auto` uses GPU when available. - **CPU only:** Use the default PyTorch from pip, or install CPU wheels: `pip install torch --index-url https://download.pytorch.org/whl/cpu` - **GPU (CUDA):** Install PyTorch with CUDA, then install the rest: `pip install torch --index-url https://download.pytorch.org/whl/cu121` then `pip install -r requirements_prod.txt` Set `DOCLING_DEVICE=cuda` or leave `auto`. At startup and when processing, the service logs which device DOCLING is using (e.g. `DOCLING accelerator: CUDA`). ## Run From the project root: ```bash ./run.sh ``` Or: ```bash .upload_venv/bin/uvicorn app.main:app --host 0.0.0.0 --port 8000 ``` Or (e.g. Hugging Face Space, port 7860): ```bash python serve.py ``` Server: `http://localhost:8000` (local) or the Space URL (HF). ## API ### POST /upload Upload a document for processing. Returns immediately with a doc_id; processing runs in the background. - **Request:** `multipart/form-data` with a file field named `file` (`.pdf`, `.docx`, `.xlsx`). - **Response:** `202 Accepted` with JSON: ```json { "doc_id": "uuid", "message": "Processing started", "status_ws_url": "/ws?doc_id=uuid" } ``` **Example:** see [Test curls](#test-curls) below. ### GET /job/{doc_id} Poll job status and result. Use the same doc_id you got from POST /upload or POST /api/vectorize. - **Response:** JSON with `status` (`processing` | `uploaded` | `error`), optional `chunk_count`, and optional `error` message. **Example:** see [Test curls](#test-curls) below. ### WebSocket /ws?doc_id={doc_id} Connect with a doc_id to receive live status updates (e.g. `processing_embeddings`, `uploaded` with `chunk_count`, or `error` with `message`). ## Pipeline 1. The file is saved temporarily and a doc_id is created. 2. **DOCLING** extracts text and structure (chunks with source, covered_pages, label). OCR and table-structure models are disabled by default for speed; see `app/services/docling_processor.py`. 3. **OpenAI** embeddings (configurable model and dimensions, with retry/backoff on rate limits). 4. **Pinecone** upsert: each chunk is stored as a vector with metadata `source`, `covered_pages` (list), `label`, and `text` (chunk content, truncated if needed for metadata size limits). Metadata also includes `doc_id` and (when provided) `project_id`. 5. Status is updated and broadcast to WebSocket clients; temp file is removed. ## Logging At startup the service logs the DOCLING accelerator in use (e.g. `DOCLING accelerator: CUDA` or `CPU`). During each job it logs the current stage: download, running DOCLING (parsing document), DOCLING done, running embeddings, embeddings done, running Pinecone upsert, Pinecone upsert done, and job complete or error. All at INFO level. Logging is configured in `app/main.py`. ## Debugging chunks If chunks look wrong (missing text, bad splits, garbled content), enable chunk debug dumps: ```bash DEBUG_CHUNKS_DIR=/tmp/chunk-debug uvicorn app.main:app --host 0.0.0.0 --port 8000 ``` Or in `.env`: ``` DEBUG_CHUNKS_DIR=/tmp/chunk-debug ``` After each document is processed, a file like `chunk-debug/..chunks.md` is written (one file per job; format is sniffed from file contents). Leave `DEBUG_CHUNKS_DIR` unset in production — no files are written. ## Test curls Quick copy-paste examples (server at `http://localhost:8000`). Use the returned or chosen doc_id to poll and for WebSocket. ```bash # Upload a document (returns doc_id) curl -X POST -F "file=@/path/to/document.pdf" http://localhost:8000/upload curl -X POST -F "file=@/path/to/document.docx;type=application/vnd.openxmlformats-officedocument.wordprocessingml.document;filename=test.docx" http://localhost:8000/upload # Poll job status (replace DOC_ID with the returned doc_id) curl http://localhost:8000/job/DOC_ID # Vectorize by URL (you choose doc_id and project_id; optional backend for callbacks) curl -X POST http://localhost:8000/api/vectorize \ -H "Content-Type: application/json" \ -d '{"url":"https://example.com/sample.pdf","doc_id":"test-doc-1","project_id":"proj-1"}' # Poll by the same doc_id curl http://localhost:8000/job/test-doc-1 ``` WebSocket for live updates: `ws://localhost:8000/ws?doc_id=DOC_ID` ## Testing ### 1. File upload (POST /upload) Start the server, then use the [test curls](#test-curls) above: ```bash curl -X POST -F "file=@/path/to/document.pdf" http://localhost:8000/upload # Use the returned doc_id: curl http://localhost:8000/job/DOC_ID ``` ### 2. Vectorize by URL (POST /api/vectorize) Requires `doc_id`, `project_id`, and a document `url` (`.pdf`, `.docx`, or `.xlsx`). Optionally set `SYMBIOS_BACKEND_BASE_URL` so the service can POST progress to `/update` and completion to `/finish`. **Without a real backend** – run a mock that prints every POST: ```bash # Terminal 1: mock backend (receives /update and /finish) .upload_venv/bin/python scripts/mock_backend.py # Terminal 2: set base URL and start upload service export SYMBIOS_BACKEND_BASE_URL=http://localhost:9999 .upload_venv/bin/uvicorn app.main:app --host 0.0.0.0 --port 8000 # Terminal 3: trigger a job (use a public PDF URL) curl -X POST http://localhost:8000/api/vectorize \ -H "Content-Type: application/json" \ -d '{"url":"https://example.com/sample.pdf","doc_id":"test-doc-1","project_id":"proj-1"}' ``` Use a real PDF URL (e.g. `https://arxiv.org/pdf/2408.09869`) to run the full pipeline; the mock backend will show progress and finish callbacks. In Terminal 1 you’ll see POSTs to `/update` (progress: Downloading…, Parsing document…, etc.) and finally POST to `/finish` with `status: complete` or `failure`. Use the same doc_id to poll or connect via WebSocket: ```bash curl http://localhost:8000/job/test-doc-1 # or connect to ws://localhost:8000/ws?doc_id=test-doc-1 ``` **With backend auth:** set `SYMBIOS_BACKEND_AUTH_SECRET` and send `Authorization: ` or `Authorization: Bearer ` on POST /api/vectorize. ### 3. Optional: task_auth_token If your backend expects a per-task token, include it in the body; the service sends it as the `Authorization` header when calling `/finish` and `/update`. It can be omitted for local testing.