kimnamjoon0007
Deploy Document Analysis API
fa15fa1
|
Raw
History Blame Contribute Delete
5.07 kB
---
title: Document Analysis API
emoji: πŸ“„
colorFrom: blue
colorTo: green
sdk: docker
app_port: 7860
pinned: false
---
# AI-Powered Document Analysis API
### GUVI Hackathon β€” Track 2
An intelligent document processing REST API that extracts, analyses, and summarises content from **PDF**, **DOCX**, and **image** files using a hybrid AI pipeline.
---
## Description
The API accepts a base64-encoded document and returns a structured JSON response containing:
- **Summary** β€” a concise AI-generated description of the document
- **Entities** β€” extracted names, dates, organisations, and monetary amounts
- **Sentiment** β€” overall document tone (Positive / Neutral / Negative)
**Strategy:** Gemini 2.5 Flash (free Google AI API) is the primary analysis engine for highest accuracy. If Gemini is unavailable, the system automatically falls back to a 100% offline pipeline (spaCy NER + DistilBERT sentiment + sumy TextRank summarisation).
---
## Tech Stack
| Layer | Tool |
|-------|------|
| Backend | FastAPI + Uvicorn |
| PDF extraction | pdfplumber (layout-preserving) |
| DOCX extraction | python-docx (paragraphs + tables) |
| OCR | pytesseract + Tesseract-OCR + Pillow |
| AI β€” Primary | Gemini 2.5 Flash (Google AI free tier) |
| AI β€” Summary fallback | sumy TextRank |
| AI β€” Entity fallback | spaCy `en_core_web_sm` + regex |
| AI β€” Sentiment fallback | DistilBERT SST-2 (HuggingFace Transformers) |
| Auth | FastAPI Header dependency |
---
## Setup Instructions
### 1. Clone the repository
```bash
git clone https://github.com/YOUR_USERNAME/YOUR_REPO.git
cd YOUR_REPO
```
### 2. Install system dependencies
```bash
# Ubuntu / Debian
sudo apt-get install -y tesseract-ocr poppler-utils
# macOS
brew install tesseract poppler
# Windows: download installers from
# https://github.com/UB-Mannheim/tesseract/wiki
# https://github.com/oschwartz10612/poppler-windows/releases
```
### 3. Install Python dependencies
```bash
pip install -r requirements.txt
python -m spacy download en_core_web_sm
```
### 4. Configure environment variables
```bash
cp .env.example .env
# Edit .env and fill in:
# API_KEY=<your chosen API key>
# GEMINI_API_KEY=<from https://aistudio.google.com/apikey>
```
### 5. Run the API
```bash
cd src
uvicorn main:app --host 0.0.0.0 --port 8000
```
---
## API Reference
### Authentication
All requests must include the `x-api-key` header:
```
x-api-key: YOUR_API_KEY
```
Returns `401 Unauthorized` if the header is missing or invalid.
### Endpoint
**POST** `/api/document-analyze`
**Request Body (JSON):**
```json
{
"fileName": "sample.pdf",
"fileType": "pdf",
"fileBase64": "<base64-encoded file content>"
}
```
Supported `fileType` values: `pdf`, `docx`, `image`
**Success Response:**
```json
{
"status": "success",
"fileName": "sample.pdf",
"summary": "This document is an invoice issued by ABC Pvt Ltd to Ravi Kumar on 10 March 2026 for an amount of β‚Ή10,000.",
"entities": {
"names": ["Ravi Kumar"],
"dates": ["10 March 2026"],
"organizations": ["ABC Pvt Ltd"],
"amounts": ["β‚Ή10,000"]
},
"sentiment": "Neutral"
}
```
### Example cURL
```bash
curl -X POST https://your-domain.com/api/document-analyze \
-H "Content-Type: application/json" \
-H "x-api-key: YOUR_API_KEY" \
-d '{
"fileName": "invoice.pdf",
"fileType": "pdf",
"fileBase64": "'"$(base64 -w 0 invoice.pdf)"'"
}'
```
---
## Approach
### Text Extraction
- **PDF**: pdfplumber extracts text with layout preservation. Pages returning no text are re-processed with Tesseract OCR (scanned PDFs).
- **DOCX**: python-docx iterates paragraphs and table cells to capture all content.
- **Image**: PIL preprocessing (grayscale β†’ sharpen β†’ autocontrast) maximises OCR accuracy before passing to Tesseract.
### Summary Generation
Gemini 2.5 Flash is prompted to produce a 1–2 sentence summary capturing the document's purpose, key actors, dates, and amounts. Fallback uses sumy's TextRank algorithm on the first 4000 characters.
### Entity Extraction
Gemini identifies and returns all four entity types (names, dates, organisations, amounts) in a structured JSON prompt. The offline fallback combines spaCy's NER (`PERSON`, `ORG` labels) with hand-crafted regex patterns for dates (multiple format support) and monetary amounts (β‚Ή, Rs., INR, $, USD, €, Β£).
### Sentiment Analysis
Gemini classifies sentiment as Positive / Neutral / Negative based on overall document tone. The fallback uses DistilBERT SST-2 β€” scores below 0.65 are mapped to Neutral to avoid overconfident labelling on factual documents.
---
## Deployment (Render.com)
1. Create a new **Web Service** on [render.com](https://render.com)
2. Set **Build Command**: `pip install -r requirements.txt && python -m spacy download en_core_web_sm`
3. Set **Start Command**: `cd src && uvicorn main:app --host 0.0.0.0 --port $PORT`
4. Add environment variables: `API_KEY`, `GEMINI_API_KEY`
5. Add **Tesseract**: use a `render.yaml` or custom Dockerfile with `apt-get install tesseract-ocr poppler-utils`