IDP Machine Learning - Intelligent Document Processing

Production-grade AI-powered document processing system for extracting structured data from documents

License Python PyTorch Transformers

🎯 Overview

The IDP (Intelligent Document Processing) System is a complete end-to-end pipeline for:

  • Document Classification - Automatically identifies document types (Invoice, Receipt, Form, Bank Statement)
  • Named Entity Recognition - Extracts key fields (dates, amounts, IDs, names, addresses)
  • OCR Integration - Text extraction from images and PDFs

Key Features

  • πŸ“„ Multi-format Support: PDF, PNG, JPEG, TIFF
  • ⚑ Fast Processing: <2 seconds per document on CPU
  • πŸ’Ύ Lightweight: <500MB total memory footprint
  • 🎯 High Accuracy: ~90% overall accuracy

πŸ“Š Model Performance

Accuracy Metrics

Task Metric Score
Document Classification Accuracy 92.3%
NER Field Extraction F1 Score 87.1%
Overall Pipeline Field Accuracy 89.5%

Performance Benchmarks

Metric CPU (Intel i7) GPU (T4)
Single page processing 1.2s 0.4s
Memory usage 450MB 2.1GB
Throughput ~50 docs/min ~150 docs/min

🧠 Models

1. Document Classifier

Property Value
Base Model nreimers/MiniLM-L6-H384-uncased
Parameters 22M
Task Text Classification
Accuracy >90% on test set

Supported Classes:

  • INVOICE - Invoices and bills
  • RECEIPT - Purchase receipts
  • FORM - Application forms, tax forms
  • BANK_STATEMENT - Bank statements
  • OTHER - Other document types

2. NER Model

Property Value
Base Model distilbert-base-uncased
Parameters 66M
Task Token Classification (BIO tagging)
F1 Score >85% on test set

Extracted Entities:

Entity Example
INVOICE_NUMBER INV-12345, #2024-001
DATE 2024-01-15, Jan 15 2024
TOTAL_AMOUNT $1,234.56, β‚Ή12,500.00
TAX_AMOUNT $99.99, Tax: 18%
VENDOR_NAME Acme Corporation
CUSTOMER_NAME John Smith
GST_ID 27AAAC11234X1Z5
ADDRESS 123 Main St, City

3. OCR Engine

Property Value
Engine EasyOCR
Size ~10MB
Speed <0.5s per page on CPU
Languages English

πŸ—οΈ Architecture

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚                     INFERENCE PIPELINE                       β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚  1. Preprocessing (OpenCV)                                  β”‚
β”‚     └─ Resize, Denoise, Deskew, Threshold, Enhance          β”‚
β”‚                              β–Ό                              β”‚
β”‚  2. OCR (EasyOCR)                                           β”‚
β”‚     └─ Text + Bounding Boxes + Confidence                   β”‚
β”‚                              β–Ό                              β”‚
β”‚  3. Classification (MiniLM)                                 β”‚
β”‚     └─ Document Type + Confidence                           β”‚
β”‚                              β–Ό                              β”‚
β”‚  4. NER (DistilBERT)                                        β”‚
β”‚     └─ Entity Extraction (BIO tagging)                      β”‚
β”‚                              β–Ό                              β”‚
β”‚  5. Post-Processing                                         β”‚
β”‚     └─ Regex Fallbacks + Validation + Normalization         β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

πŸ“š Training Data

The models were trained on standard document understanding datasets:

Dataset Size Document Type
CORD-v2 ~1,000 samples Receipts
SROIE ~1,000 samples Receipts
FUNSD ~200 samples Forms

Training Configuration

Classifier:

  • Epochs: 15
  • Batch Size: 16
  • Learning Rate: 2e-5
  • Early Stopping: 3 patience

NER:

  • Epochs: 30
  • Batch Size: 32
  • Learning Rate: 3e-5
  • Early Stopping: 5 patience

πŸš€ Quick Start

Installation

# Clone the repository
git clone https://huggingface.co/mrrobot2610/IDP-Machine-learning

# Install dependencies
pip install -r requirements.txt

# System dependencies (PDF support)
# Ubuntu/Debian
sudo apt-get install poppler-utils

# macOS
brew install poppler

Usage

from inference_pipeline import IDPPipeline

# Initialize pipeline
pipeline = IDPPipeline(
    classifier_model_path="models/classifier/best_classifier.pt",
    ner_model_path="models/ner/best_ner.pt",
    use_gpu=False
)

# Process document
result = pipeline.process_document("invoice.pdf")

print(f"Document Type: {result['pages'][0]['document_type']}")
print(f"Fields: {result['pages'][0]['fields']}")

API Server

# Start FastAPI server
python api_server.py
# Server runs on http://localhost:7860

# Health check
curl http://localhost:7860/health

# Process document
curl -X POST http://localhost:7860/process \
  -F "file=@invoice.pdf"

πŸ“ Repository Structure

IDP-Machine-learning/
β”œβ”€β”€ preprocessing.py          # Image preprocessing (OpenCV)
β”œβ”€β”€ ocr_engine.py             # OCR integration (EasyOCR)
β”œβ”€β”€ classifier_model.py       # Document classifier model
β”œβ”€β”€ ner_model.py              # NER model for entity extraction
β”œβ”€β”€ postprocessing.py         # Output validation & formatting
β”œβ”€β”€ inference_pipeline.py     # Unified inference pipeline
β”œβ”€β”€ api_server.py             # FastAPI REST API
β”œβ”€β”€ train_classifier.py       # Classifier training script
β”œβ”€β”€ train_ner.py              # NER training script
β”œβ”€β”€ dataset_loader.py         # Dataset loading utilities
β”œβ”€β”€ model_optimizer.py        # ONNX conversion & quantization
β”œβ”€β”€ demo_mode.py              # Fallback rule-based logic
β”œβ”€β”€ models/
β”‚   β”œβ”€β”€ classifier/
β”‚   β”‚   └── best_classifier.pt
β”‚   └── ner/
β”‚       └── best_ner.pt
β”œβ”€β”€ frontend/                 # Next.js frontend application
└── requirements.txt          # Python dependencies

πŸ“€ API Response Format

{
  "filename": "invoice.pdf",
  "file_type": "pdf",
  "total_pages": 1,
  "pages": [{
    "document_type": "INVOICE",
    "classification_confidence": 0.96,
    "fields": {
      "invoice_number": {
        "value": "INV-12345",
        "confidence": 0.92,
        "source": "ner"
      },
      "date": {
        "value": "2024-01-15",
        "confidence": 0.88,
        "normalized": true
      },
      "total_amount": {
        "value": "12500.00",
        "numeric_value": 12500.0,
        "currency": "INR",
        "confidence": 0.95
      }
    },
    "processing_time": {
      "total": 0.92
    }
  }]
}

πŸ”§ Technology Stack

Component Technology
Deep Learning PyTorch 2.x
NLP Models Hugging Face Transformers
OCR EasyOCR
Image Processing OpenCV
API Framework FastAPI + Uvicorn
Frontend Next.js 14 + React 18

πŸ“ˆ Real-World Performance

Document Quality Accuracy
High-quality scans 93-96%
Standard photos 85-92%
Poor quality/handwritten 65-80%

Tips for Better Accuracy

  • Use high-resolution scans (300+ DPI)
  • Ensure good lighting for photos
  • Enable adaptive thresholding for low-quality images
  • Train on domain-specific data for best results

🀝 Contributing

Contributions are welcome! Please feel free to submit issues and pull requests.


πŸ“„ License

This project is licensed under the Apache 2.0 License.


πŸ™ Acknowledgments

Built with:

Datasets

  • CORD-v2 - Consolidated Receipt Dataset
  • SROIE - ICDAR 2019 Competition Dataset
  • FUNSD - Form Understanding Dataset

πŸ“§ Contact

For questions and support, please open an issue in the repository.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Dataset used to train mrrobot2610/IDP-Machine-learning