Spaces:
Sleeping
Sleeping
| title: Document Analysis API | |
| emoji: 📄 | |
| colorFrom: blue | |
| colorTo: green | |
| sdk: docker | |
| app_port: 7860 | |
| pinned: false | |
| # AI-Powered Document Analysis API | |
| ### GUVI Hackathon — Track 2 | |
| An intelligent document processing REST API that extracts, analyses, and summarises content from **PDF**, **DOCX**, and **image** files using a hybrid AI pipeline. | |
| --- | |
| ## Description | |
| The API accepts a base64-encoded document and returns a structured JSON response containing: | |
| - **Summary** — a concise AI-generated description of the document | |
| - **Entities** — extracted names, dates, organisations, and monetary amounts | |
| - **Sentiment** — overall document tone (Positive / Neutral / Negative) | |
| **Strategy:** Gemini 2.5 Flash (free Google AI API) is the primary analysis engine for highest accuracy. If Gemini is unavailable, the system automatically falls back to a 100% offline pipeline (spaCy NER + DistilBERT sentiment + sumy TextRank summarisation). | |
| --- | |
| ## Tech Stack | |
| | Layer | Tool | | |
| |-------|------| | |
| | Backend | FastAPI + Uvicorn | | |
| | PDF extraction | pdfplumber (layout-preserving) | | |
| | DOCX extraction | python-docx (paragraphs + tables) | | |
| | OCR | pytesseract + Tesseract-OCR + Pillow | | |
| | AI — Primary | Gemini 2.5 Flash (Google AI free tier) | | |
| | AI — Summary fallback | sumy TextRank | | |
| | AI — Entity fallback | spaCy `en_core_web_sm` + regex | | |
| | AI — Sentiment fallback | DistilBERT SST-2 (HuggingFace Transformers) | | |
| | Auth | FastAPI Header dependency | | |
| --- | |
| ## Setup Instructions | |
| ### 1. Clone the repository | |
| ```bash | |
| git clone https://github.com/YOUR_USERNAME/YOUR_REPO.git | |
| cd YOUR_REPO | |
| ``` | |
| ### 2. Install system dependencies | |
| ```bash | |
| # Ubuntu / Debian | |
| sudo apt-get install -y tesseract-ocr poppler-utils | |
| # macOS | |
| brew install tesseract poppler | |
| # Windows: download installers from | |
| # https://github.com/UB-Mannheim/tesseract/wiki | |
| # https://github.com/oschwartz10612/poppler-windows/releases | |
| ``` | |
| ### 3. Install Python dependencies | |
| ```bash | |
| pip install -r requirements.txt | |
| python -m spacy download en_core_web_sm | |
| ``` | |
| ### 4. Configure environment variables | |
| ```bash | |
| cp .env.example .env | |
| # Edit .env and fill in: | |
| # API_KEY=<your chosen API key> | |
| # GEMINI_API_KEY=<from https://aistudio.google.com/apikey> | |
| ``` | |
| ### 5. Run the API | |
| ```bash | |
| cd src | |
| uvicorn main:app --host 0.0.0.0 --port 8000 | |
| ``` | |
| --- | |
| ## API Reference | |
| ### Authentication | |
| All requests must include the `x-api-key` header: | |
| ``` | |
| x-api-key: YOUR_API_KEY | |
| ``` | |
| Returns `401 Unauthorized` if the header is missing or invalid. | |
| ### Endpoint | |
| **POST** `/api/document-analyze` | |
| **Request Body (JSON):** | |
| ```json | |
| { | |
| "fileName": "sample.pdf", | |
| "fileType": "pdf", | |
| "fileBase64": "<base64-encoded file content>" | |
| } | |
| ``` | |
| Supported `fileType` values: `pdf`, `docx`, `image` | |
| **Success Response:** | |
| ```json | |
| { | |
| "status": "success", | |
| "fileName": "sample.pdf", | |
| "summary": "This document is an invoice issued by ABC Pvt Ltd to Ravi Kumar on 10 March 2026 for an amount of ₹10,000.", | |
| "entities": { | |
| "names": ["Ravi Kumar"], | |
| "dates": ["10 March 2026"], | |
| "organizations": ["ABC Pvt Ltd"], | |
| "amounts": ["₹10,000"] | |
| }, | |
| "sentiment": "Neutral" | |
| } | |
| ``` | |
| ### Example cURL | |
| ```bash | |
| curl -X POST https://your-domain.com/api/document-analyze \ | |
| -H "Content-Type: application/json" \ | |
| -H "x-api-key: YOUR_API_KEY" \ | |
| -d '{ | |
| "fileName": "invoice.pdf", | |
| "fileType": "pdf", | |
| "fileBase64": "'"$(base64 -w 0 invoice.pdf)"'" | |
| }' | |
| ``` | |
| --- | |
| ## Approach | |
| ### Text Extraction | |
| - **PDF**: pdfplumber extracts text with layout preservation. Pages returning no text are re-processed with Tesseract OCR (scanned PDFs). | |
| - **DOCX**: python-docx iterates paragraphs and table cells to capture all content. | |
| - **Image**: PIL preprocessing (grayscale → sharpen → autocontrast) maximises OCR accuracy before passing to Tesseract. | |
| ### Summary Generation | |
| Gemini 2.5 Flash is prompted to produce a 1–2 sentence summary capturing the document's purpose, key actors, dates, and amounts. Fallback uses sumy's TextRank algorithm on the first 4000 characters. | |
| ### Entity Extraction | |
| Gemini identifies and returns all four entity types (names, dates, organisations, amounts) in a structured JSON prompt. The offline fallback combines spaCy's NER (`PERSON`, `ORG` labels) with hand-crafted regex patterns for dates (multiple format support) and monetary amounts (₹, Rs., INR, $, USD, €, £). | |
| ### Sentiment Analysis | |
| Gemini classifies sentiment as Positive / Neutral / Negative based on overall document tone. The fallback uses DistilBERT SST-2 — scores below 0.65 are mapped to Neutral to avoid overconfident labelling on factual documents. | |
| --- | |
| ## Deployment (Render.com) | |
| 1. Create a new **Web Service** on [render.com](https://render.com) | |
| 2. Set **Build Command**: `pip install -r requirements.txt && python -m spacy download en_core_web_sm` | |
| 3. Set **Start Command**: `cd src && uvicorn main:app --host 0.0.0.0 --port $PORT` | |
| 4. Add environment variables: `API_KEY`, `GEMINI_API_KEY` | |
| 5. Add **Tesseract**: use a `render.yaml` or custom Dockerfile with `apt-get install tesseract-ocr poppler-utils` | |