kimnamjoon0007
Deploy Document Analysis API
fa15fa1
|
Raw
History Blame Contribute Delete
5.07 kB
metadata
title: Document Analysis API
emoji: πŸ“„
colorFrom: blue
colorTo: green
sdk: docker
app_port: 7860
pinned: false

AI-Powered Document Analysis API

GUVI Hackathon β€” Track 2

An intelligent document processing REST API that extracts, analyses, and summarises content from PDF, DOCX, and image files using a hybrid AI pipeline.


Description

The API accepts a base64-encoded document and returns a structured JSON response containing:

  • Summary β€” a concise AI-generated description of the document
  • Entities β€” extracted names, dates, organisations, and monetary amounts
  • Sentiment β€” overall document tone (Positive / Neutral / Negative)

Strategy: Gemini 2.5 Flash (free Google AI API) is the primary analysis engine for highest accuracy. If Gemini is unavailable, the system automatically falls back to a 100% offline pipeline (spaCy NER + DistilBERT sentiment + sumy TextRank summarisation).


Tech Stack

Layer Tool
Backend FastAPI + Uvicorn
PDF extraction pdfplumber (layout-preserving)
DOCX extraction python-docx (paragraphs + tables)
OCR pytesseract + Tesseract-OCR + Pillow
AI β€” Primary Gemini 2.5 Flash (Google AI free tier)
AI β€” Summary fallback sumy TextRank
AI β€” Entity fallback spaCy en_core_web_sm + regex
AI β€” Sentiment fallback DistilBERT SST-2 (HuggingFace Transformers)
Auth FastAPI Header dependency

Setup Instructions

1. Clone the repository

git clone https://github.com/YOUR_USERNAME/YOUR_REPO.git
cd YOUR_REPO

2. Install system dependencies

# Ubuntu / Debian
sudo apt-get install -y tesseract-ocr poppler-utils

# macOS
brew install tesseract poppler

# Windows: download installers from
#   https://github.com/UB-Mannheim/tesseract/wiki
#   https://github.com/oschwartz10612/poppler-windows/releases

3. Install Python dependencies

pip install -r requirements.txt
python -m spacy download en_core_web_sm

4. Configure environment variables

cp .env.example .env
# Edit .env and fill in:
#   API_KEY=<your chosen API key>
#   GEMINI_API_KEY=<from https://aistudio.google.com/apikey>

5. Run the API

cd src
uvicorn main:app --host 0.0.0.0 --port 8000

API Reference

Authentication

All requests must include the x-api-key header:

x-api-key: YOUR_API_KEY

Returns 401 Unauthorized if the header is missing or invalid.

Endpoint

POST /api/document-analyze

Request Body (JSON):

{
  "fileName": "sample.pdf",
  "fileType": "pdf",
  "fileBase64": "<base64-encoded file content>"
}

Supported fileType values: pdf, docx, image

Success Response:

{
  "status": "success",
  "fileName": "sample.pdf",
  "summary": "This document is an invoice issued by ABC Pvt Ltd to Ravi Kumar on 10 March 2026 for an amount of β‚Ή10,000.",
  "entities": {
    "names": ["Ravi Kumar"],
    "dates": ["10 March 2026"],
    "organizations": ["ABC Pvt Ltd"],
    "amounts": ["β‚Ή10,000"]
  },
  "sentiment": "Neutral"
}

Example cURL

curl -X POST https://your-domain.com/api/document-analyze \
  -H "Content-Type: application/json" \
  -H "x-api-key: YOUR_API_KEY" \
  -d '{
    "fileName": "invoice.pdf",
    "fileType": "pdf",
    "fileBase64": "'"$(base64 -w 0 invoice.pdf)"'"
  }'

Approach

Text Extraction

  • PDF: pdfplumber extracts text with layout preservation. Pages returning no text are re-processed with Tesseract OCR (scanned PDFs).
  • DOCX: python-docx iterates paragraphs and table cells to capture all content.
  • Image: PIL preprocessing (grayscale β†’ sharpen β†’ autocontrast) maximises OCR accuracy before passing to Tesseract.

Summary Generation

Gemini 2.5 Flash is prompted to produce a 1–2 sentence summary capturing the document's purpose, key actors, dates, and amounts. Fallback uses sumy's TextRank algorithm on the first 4000 characters.

Entity Extraction

Gemini identifies and returns all four entity types (names, dates, organisations, amounts) in a structured JSON prompt. The offline fallback combines spaCy's NER (PERSON, ORG labels) with hand-crafted regex patterns for dates (multiple format support) and monetary amounts (β‚Ή, Rs., INR, $, USD, €, Β£).

Sentiment Analysis

Gemini classifies sentiment as Positive / Neutral / Negative based on overall document tone. The fallback uses DistilBERT SST-2 β€” scores below 0.65 are mapped to Neutral to avoid overconfident labelling on factual documents.


Deployment (Render.com)

  1. Create a new Web Service on render.com
  2. Set Build Command: pip install -r requirements.txt && python -m spacy download en_core_web_sm
  3. Set Start Command: cd src && uvicorn main:app --host 0.0.0.0 --port $PORT
  4. Add environment variables: API_KEY, GEMINI_API_KEY
  5. Add Tesseract: use a render.yaml or custom Dockerfile with apt-get install tesseract-ocr poppler-utils