--- title: Document Analysis API emoji: 📄 colorFrom: blue colorTo: green sdk: docker app_port: 7860 pinned: false --- # AI-Powered Document Analysis API ### GUVI Hackathon — Track 2 An intelligent document processing REST API that extracts, analyses, and summarises content from **PDF**, **DOCX**, and **image** files using a hybrid AI pipeline. --- ## Description The API accepts a base64-encoded document and returns a structured JSON response containing: - **Summary** — a concise AI-generated description of the document - **Entities** — extracted names, dates, organisations, and monetary amounts - **Sentiment** — overall document tone (Positive / Neutral / Negative) **Strategy:** Gemini 2.5 Flash (free Google AI API) is the primary analysis engine for highest accuracy. If Gemini is unavailable, the system automatically falls back to a 100% offline pipeline (spaCy NER + DistilBERT sentiment + sumy TextRank summarisation). --- ## Tech Stack | Layer | Tool | |-------|------| | Backend | FastAPI + Uvicorn | | PDF extraction | pdfplumber (layout-preserving) | | DOCX extraction | python-docx (paragraphs + tables) | | OCR | pytesseract + Tesseract-OCR + Pillow | | AI — Primary | Gemini 2.5 Flash (Google AI free tier) | | AI — Summary fallback | sumy TextRank | | AI — Entity fallback | spaCy `en_core_web_sm` + regex | | AI — Sentiment fallback | DistilBERT SST-2 (HuggingFace Transformers) | | Auth | FastAPI Header dependency | --- ## Setup Instructions ### 1. Clone the repository ```bash git clone https://github.com/YOUR_USERNAME/YOUR_REPO.git cd YOUR_REPO ``` ### 2. Install system dependencies ```bash # Ubuntu / Debian sudo apt-get install -y tesseract-ocr poppler-utils # macOS brew install tesseract poppler # Windows: download installers from # https://github.com/UB-Mannheim/tesseract/wiki # https://github.com/oschwartz10612/poppler-windows/releases ``` ### 3. Install Python dependencies ```bash pip install -r requirements.txt python -m spacy download en_core_web_sm ``` ### 4. Configure environment variables ```bash cp .env.example .env # Edit .env and fill in: # API_KEY= # GEMINI_API_KEY= ``` ### 5. Run the API ```bash cd src uvicorn main:app --host 0.0.0.0 --port 8000 ``` --- ## API Reference ### Authentication All requests must include the `x-api-key` header: ``` x-api-key: YOUR_API_KEY ``` Returns `401 Unauthorized` if the header is missing or invalid. ### Endpoint **POST** `/api/document-analyze` **Request Body (JSON):** ```json { "fileName": "sample.pdf", "fileType": "pdf", "fileBase64": "" } ``` Supported `fileType` values: `pdf`, `docx`, `image` **Success Response:** ```json { "status": "success", "fileName": "sample.pdf", "summary": "This document is an invoice issued by ABC Pvt Ltd to Ravi Kumar on 10 March 2026 for an amount of ₹10,000.", "entities": { "names": ["Ravi Kumar"], "dates": ["10 March 2026"], "organizations": ["ABC Pvt Ltd"], "amounts": ["₹10,000"] }, "sentiment": "Neutral" } ``` ### Example cURL ```bash curl -X POST https://your-domain.com/api/document-analyze \ -H "Content-Type: application/json" \ -H "x-api-key: YOUR_API_KEY" \ -d '{ "fileName": "invoice.pdf", "fileType": "pdf", "fileBase64": "'"$(base64 -w 0 invoice.pdf)"'" }' ``` --- ## Approach ### Text Extraction - **PDF**: pdfplumber extracts text with layout preservation. Pages returning no text are re-processed with Tesseract OCR (scanned PDFs). - **DOCX**: python-docx iterates paragraphs and table cells to capture all content. - **Image**: PIL preprocessing (grayscale → sharpen → autocontrast) maximises OCR accuracy before passing to Tesseract. ### Summary Generation Gemini 2.5 Flash is prompted to produce a 1–2 sentence summary capturing the document's purpose, key actors, dates, and amounts. Fallback uses sumy's TextRank algorithm on the first 4000 characters. ### Entity Extraction Gemini identifies and returns all four entity types (names, dates, organisations, amounts) in a structured JSON prompt. The offline fallback combines spaCy's NER (`PERSON`, `ORG` labels) with hand-crafted regex patterns for dates (multiple format support) and monetary amounts (₹, Rs., INR, $, USD, €, £). ### Sentiment Analysis Gemini classifies sentiment as Positive / Neutral / Negative based on overall document tone. The fallback uses DistilBERT SST-2 — scores below 0.65 are mapped to Neutral to avoid overconfident labelling on factual documents. --- ## Deployment (Render.com) 1. Create a new **Web Service** on [render.com](https://render.com) 2. Set **Build Command**: `pip install -r requirements.txt && python -m spacy download en_core_web_sm` 3. Set **Start Command**: `cd src && uvicorn main:app --host 0.0.0.0 --port $PORT` 4. Add environment variables: `API_KEY`, `GEMINI_API_KEY` 5. Add **Tesseract**: use a `render.yaml` or custom Dockerfile with `apt-get install tesseract-ocr poppler-utils`