IDP-Machine-learning / deployment_guide.md
mrrobot2610's picture
Initial commit: IDP (Intelligent Document Processing) System
1a7ee60
|
Raw History Blame Contribute Delete
6.31 kB

Hugging Face Spaces Deployment Guide

Complete guide to deploy the IDP API on Hugging Face Spaces.

Prerequisites

  • Hugging Face account
  • Trained model weights (classifier and NER)
  • Git installed locally

Step 1: Create a New Space

  1. Go to Hugging Face Spaces

  2. Click "Create new Space"

  3. Configure:

    • Space name: idp-api (or your preferred name)
    • License: Apache 2.0 (or your choice)
    • Select SDK: Docker
    • Hardware: CPU Basic (free) or upgrade to GPU if needed
  4. Click "Create Space"

Step 2: Prepare Your Files

Create a Dockerfile in your project root:

FROM python:3.10-slim

# Install system dependencies
RUN apt-get update && apt-get install -y \
    poppler-utils \
    libgomp1 \
    libglib2.0-0 \
    libsm6 \
    libxext6 \
    libxrender-dev \
    libgl1-mesa-glx \
    && rm -rf /var/lib/apt/lists/*

# Set working directory
WORKDIR /app

# Copy requirements
COPY requirements.txt .

# Install Python dependencies
RUN pip install --no-cache-dir -r requirements.txt

# Copy application code
COPY *.py ./
COPY models/ ./models/

# Expose port 7860 (HF Spaces default)
EXPOSE 7860

# Run the API server
CMD ["uvicorn", "api_server:app", "--host", "0.0.0.0", "--port", "7860"]

Step 3: Train Your Models

Before deployment, train your models locally:

# Train classifier
python train_classifier.py

# Train NER model
python train_ner.py

This will create:

  • models/classifier/best_classifier.pt
  • models/ner/best_ner.pt

Step 4: (Optional) Optimize Models for Faster Inference

Convert to ONNX and quantize:

# Convert classifier to ONNX
python model_optimizer.py convert_classifier \
    models/classifier/best_classifier.pt \
    models/classifier/classifier.onnx

# Quantize classifier
python model_optimizer.py quantize \
    models/classifier/classifier.onnx \
    models/classifier/classifier_quantized.onnx

# Convert NER to ONNX
python model_optimizer.py convert_ner \
    models/ner/best_ner.pt \
    models/ner/ner.onnx

# Quantize NER
python model_optimizer.py quantize \
    models/ner/ner.onnx \
    models/ner/ner_quantized.onnx

If using ONNX models, update api_server.py to use ONNX inference sessions.

Step 5: Push to Hugging Face Space

Clone your Space repository:

git clone https://huggingface.co/spaces/YOUR_USERNAME/idp-api
cd idp-api

Copy your files:

# Copy Python files
cp /path/to/your/project/*.py .

# Copy models
cp -r /path/to/your/project/models .

# Copy config files
cp /path/to/your/project/requirements.txt .
cp /path/to/your/project/Dockerfile .

Create a README.md for your Space:

---
title: IDP API
emoji: 📄
colorFrom: blue
colorTo: green
sdk: docker
pinned: false
---

# Intelligent Document Processing API

API for extracting structured data from invoices, receipts, and forms.

## Features

- Lightweight OCR with PaddleOCR
- Document classification (Invoice, Receipt, Form)
- Entity extraction (dates, amounts, names, etc.)
- Confidence scoring
- CORS-enabled for web integration

## API Endpoints

### POST /process

Upload a document (PDF or image) and get structured data.

### GET /health

Health check endpoint.

See full documentation at [YOUR_REPO_URL]

Commit and push:

git add .
git commit -m "Initial deployment"
git push

Step 6: Monitor Deployment

  1. Go to your Space URL: https://huggingface.co/spaces/YOUR_USERNAME/idp-api
  2. Watch the build logs in the "Logs" tab
  3. Wait for the build to complete (may take 5-10 minutes for first build)
  4. Once running, the Space will show "Running" status

Step 7: Test Your API

Test the health endpoint:

curl https://YOUR_USERNAME-idp-api.hf.space/health

Test document processing:

curl -X POST \
  https://YOUR_USERNAME-idp-api.hf.space/process \
  -F "file=@sample_invoice.pdf"

Step 8: Configure for Production

Upgrade Hardware (Optional)

If CPU performance is insufficient:

  1. Go to Space settings
  2. Under "Hardware", upgrade to:
    • CPU Upgrade: 2 vCPUs, 16GB RAM
    • GPU T4 Small**: 15GB VRAM (if using GPU)

Set Secrets (Optional)

If you need API keys or secrets:

  1. Go to Space settings
  2. Add secrets under "Repository secrets"
  3. Access in code via os.environ.get('SECRET_NAME')

Enable Persistent Storage (Optional)

For caching or logging:

  1. Go to Space settings
  2. Enable "Persistent Storage"
  3. Data will persist in /data directory

Troubleshooting

Build Fails

Issue: Docker build fails with package errors

Solution:

  • Check requirements.txt for version conflicts
  • Review build logs for specific errors
  • Ensure Dockerfile has all system dependencies

Out of Memory

Issue: API crashes with OOM errors

Solution:

  • Upgrade to higher memory tier
  • Reduce batch size
  • Use ONNX quantized models
  • Disable GPU if not needed

Slow Inference

Issue: Processing takes >5 seconds per document

Solution:

  • Use ONNX models with INT8 quantization
  • Upgrade to GPU hardware
  • Reduce OCR DPI (lower quality but faster)
  • Cache models in memory (done by default)

Cost Optimization

Free Tier:

  • CPU Basic: Free forever
  • Suitable for demos and low-traffic apps
  • May sleep after inactivity

Paid Tier (if needed):

  • CPU Upgrade: ~$0.50/hour
  • GPU T4: ~$0.60/hour
  • Only charged when running

Tips:

  • Use CPU for development
  • Upgrade to GPU only for production with high traffic
  • Use ONNX + quantization to maximize CPU performance

Next Steps

  • Set up monitoring with Hugging Face Analytics
  • Add authentication if needed
  • Create a custom domain (paid feature)
  • Integrate with your Next.js frontend (see nextjs_integration_guide.md)

Support

For issues: