title: Document Analysis API
emoji: π
colorFrom: blue
colorTo: green
sdk: docker
app_port: 7860
pinned: false
AI-Powered Document Analysis API
GUVI Hackathon β Track 2
An intelligent document processing REST API that extracts, analyses, and summarises content from PDF, DOCX, and image files using a hybrid AI pipeline.
Description
The API accepts a base64-encoded document and returns a structured JSON response containing:
- Summary β a concise AI-generated description of the document
- Entities β extracted names, dates, organisations, and monetary amounts
- Sentiment β overall document tone (Positive / Neutral / Negative)
Strategy: Gemini 2.5 Flash (free Google AI API) is the primary analysis engine for highest accuracy. If Gemini is unavailable, the system automatically falls back to a 100% offline pipeline (spaCy NER + DistilBERT sentiment + sumy TextRank summarisation).
Tech Stack
| Layer | Tool |
|---|---|
| Backend | FastAPI + Uvicorn |
| PDF extraction | pdfplumber (layout-preserving) |
| DOCX extraction | python-docx (paragraphs + tables) |
| OCR | pytesseract + Tesseract-OCR + Pillow |
| AI β Primary | Gemini 2.5 Flash (Google AI free tier) |
| AI β Summary fallback | sumy TextRank |
| AI β Entity fallback | spaCy en_core_web_sm + regex |
| AI β Sentiment fallback | DistilBERT SST-2 (HuggingFace Transformers) |
| Auth | FastAPI Header dependency |
Setup Instructions
1. Clone the repository
git clone https://github.com/YOUR_USERNAME/YOUR_REPO.git
cd YOUR_REPO
2. Install system dependencies
# Ubuntu / Debian
sudo apt-get install -y tesseract-ocr poppler-utils
# macOS
brew install tesseract poppler
# Windows: download installers from
# https://github.com/UB-Mannheim/tesseract/wiki
# https://github.com/oschwartz10612/poppler-windows/releases
3. Install Python dependencies
pip install -r requirements.txt
python -m spacy download en_core_web_sm
4. Configure environment variables
cp .env.example .env
# Edit .env and fill in:
# API_KEY=<your chosen API key>
# GEMINI_API_KEY=<from https://aistudio.google.com/apikey>
5. Run the API
cd src
uvicorn main:app --host 0.0.0.0 --port 8000
API Reference
Authentication
All requests must include the x-api-key header:
x-api-key: YOUR_API_KEY
Returns 401 Unauthorized if the header is missing or invalid.
Endpoint
POST /api/document-analyze
Request Body (JSON):
{
"fileName": "sample.pdf",
"fileType": "pdf",
"fileBase64": "<base64-encoded file content>"
}
Supported fileType values: pdf, docx, image
Success Response:
{
"status": "success",
"fileName": "sample.pdf",
"summary": "This document is an invoice issued by ABC Pvt Ltd to Ravi Kumar on 10 March 2026 for an amount of βΉ10,000.",
"entities": {
"names": ["Ravi Kumar"],
"dates": ["10 March 2026"],
"organizations": ["ABC Pvt Ltd"],
"amounts": ["βΉ10,000"]
},
"sentiment": "Neutral"
}
Example cURL
curl -X POST https://your-domain.com/api/document-analyze \
-H "Content-Type: application/json" \
-H "x-api-key: YOUR_API_KEY" \
-d '{
"fileName": "invoice.pdf",
"fileType": "pdf",
"fileBase64": "'"$(base64 -w 0 invoice.pdf)"'"
}'
Approach
Text Extraction
- PDF: pdfplumber extracts text with layout preservation. Pages returning no text are re-processed with Tesseract OCR (scanned PDFs).
- DOCX: python-docx iterates paragraphs and table cells to capture all content.
- Image: PIL preprocessing (grayscale β sharpen β autocontrast) maximises OCR accuracy before passing to Tesseract.
Summary Generation
Gemini 2.5 Flash is prompted to produce a 1β2 sentence summary capturing the document's purpose, key actors, dates, and amounts. Fallback uses sumy's TextRank algorithm on the first 4000 characters.
Entity Extraction
Gemini identifies and returns all four entity types (names, dates, organisations, amounts) in a structured JSON prompt. The offline fallback combines spaCy's NER (PERSON, ORG labels) with hand-crafted regex patterns for dates (multiple format support) and monetary amounts (βΉ, Rs., INR, $, USD, β¬, Β£).
Sentiment Analysis
Gemini classifies sentiment as Positive / Neutral / Negative based on overall document tone. The fallback uses DistilBERT SST-2 β scores below 0.65 are mapped to Neutral to avoid overconfident labelling on factual documents.
Deployment (Render.com)
- Create a new Web Service on render.com
- Set Build Command:
pip install -r requirements.txt && python -m spacy download en_core_web_sm - Set Start Command:
cd src && uvicorn main:app --host 0.0.0.0 --port $PORT - Add environment variables:
API_KEY,GEMINI_API_KEY - Add Tesseract: use a
render.yamlor custom Dockerfile withapt-get install tesseract-ocr poppler-utils