File size: 9,651 Bytes
a983eb6
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
---
license: apache-2.0
language:
- en
tags:
- document-processing
- ocr
- ner
- text-classification
- information-extraction
- invoice
- receipt
- form
- pytorch
- transformers
datasets:
- naver-clova-ix/cord-v2
- SROIE
- FUNSD
pipeline_tag: document-question-answering
metrics:
- accuracy
- f1
library_name: transformers
---

# IDP Machine Learning - Intelligent Document Processing

<div align="center">

**Production-grade AI-powered document processing system for extracting structured data from documents**

[![License](https://img.shields.io/badge/License-Apache_2.0-blue.svg)](https://opensource.org/licenses/Apache-2.0)
[![Python](https://img.shields.io/badge/Python-3.10+-green.svg)](https://python.org)
[![PyTorch](https://img.shields.io/badge/PyTorch-2.x-orange.svg)](https://pytorch.org)
[![Transformers](https://img.shields.io/badge/Transformers-4.x-yellow.svg)](https://huggingface.co/transformers)

</div>

## 🎯 Overview

The IDP (Intelligent Document Processing) System is a complete end-to-end pipeline for:
- **Document Classification** - Automatically identifies document types (Invoice, Receipt, Form, Bank Statement)
- **Named Entity Recognition** - Extracts key fields (dates, amounts, IDs, names, addresses)
- **OCR Integration** - Text extraction from images and PDFs

### Key Features
- πŸ“„ **Multi-format Support**: PDF, PNG, JPEG, TIFF
- ⚑ **Fast Processing**: <2 seconds per document on CPU
- πŸ’Ύ **Lightweight**: <500MB total memory footprint
- 🎯 **High Accuracy**: ~90% overall accuracy

---

## πŸ“Š Model Performance

### Accuracy Metrics

| Task | Metric | Score |
|------|--------|-------|
| **Document Classification** | Accuracy | **92.3%** |
| **NER Field Extraction** | F1 Score | **87.1%** |
| **Overall Pipeline** | Field Accuracy | **89.5%** |

### Performance Benchmarks

| Metric | CPU (Intel i7) | GPU (T4) |
|--------|----------------|----------|
| Single page processing | 1.2s | 0.4s |
| Memory usage | 450MB | 2.1GB |
| Throughput | ~50 docs/min | ~150 docs/min |

---

## 🧠 Models

### 1. Document Classifier

| Property | Value |
|----------|-------|
| **Base Model** | `nreimers/MiniLM-L6-H384-uncased` |
| **Parameters** | 22M |
| **Task** | Text Classification |
| **Accuracy** | >90% on test set |

**Supported Classes:**
- `INVOICE` - Invoices and bills
- `RECEIPT` - Purchase receipts
- `FORM` - Application forms, tax forms
- `BANK_STATEMENT` - Bank statements
- `OTHER` - Other document types

### 2. NER Model

| Property | Value |
|----------|-------|
| **Base Model** | `distilbert-base-uncased` |
| **Parameters** | 66M |
| **Task** | Token Classification (BIO tagging) |
| **F1 Score** | >85% on test set |

**Extracted Entities:**
| Entity | Example |
|--------|---------|
| `INVOICE_NUMBER` | INV-12345, #2024-001 |
| `DATE` | 2024-01-15, Jan 15 2024 |
| `TOTAL_AMOUNT` | $1,234.56, β‚Ή12,500.00 |
| `TAX_AMOUNT` | $99.99, Tax: 18% |
| `VENDOR_NAME` | Acme Corporation |
| `CUSTOMER_NAME` | John Smith |
| `GST_ID` | 27AAAC11234X1Z5 |
| `ADDRESS` | 123 Main St, City |

### 3. OCR Engine

| Property | Value |
|----------|-------|
| **Engine** | EasyOCR |
| **Size** | ~10MB |
| **Speed** | <0.5s per page on CPU |
| **Languages** | English |

---

## πŸ—οΈ Architecture

```
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚                     INFERENCE PIPELINE                       β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚  1. Preprocessing (OpenCV)                                  β”‚
β”‚     └─ Resize, Denoise, Deskew, Threshold, Enhance          β”‚
β”‚                              β–Ό                              β”‚
β”‚  2. OCR (EasyOCR)                                           β”‚
β”‚     └─ Text + Bounding Boxes + Confidence                   β”‚
β”‚                              β–Ό                              β”‚
β”‚  3. Classification (MiniLM)                                 β”‚
β”‚     └─ Document Type + Confidence                           β”‚
β”‚                              β–Ό                              β”‚
β”‚  4. NER (DistilBERT)                                        β”‚
β”‚     └─ Entity Extraction (BIO tagging)                      β”‚
β”‚                              β–Ό                              β”‚
β”‚  5. Post-Processing                                         β”‚
β”‚     └─ Regex Fallbacks + Validation + Normalization         β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
```

---

## πŸ“š Training Data

The models were trained on standard document understanding datasets:

| Dataset | Size | Document Type |
|---------|------|---------------|
| **CORD-v2** | ~1,000 samples | Receipts |
| **SROIE** | ~1,000 samples | Receipts |
| **FUNSD** | ~200 samples | Forms |

### Training Configuration

**Classifier:**
- Epochs: 15
- Batch Size: 16
- Learning Rate: 2e-5
- Early Stopping: 3 patience

**NER:**
- Epochs: 30
- Batch Size: 32
- Learning Rate: 3e-5
- Early Stopping: 5 patience

---

## πŸš€ Quick Start

### Installation

```bash
# Clone the repository
git clone https://huggingface.co/mrrobot2610/IDP-Machine-learning

# Install dependencies
pip install -r requirements.txt

# System dependencies (PDF support)
# Ubuntu/Debian
sudo apt-get install poppler-utils

# macOS
brew install poppler
```

### Usage

```python
from inference_pipeline import IDPPipeline

# Initialize pipeline
pipeline = IDPPipeline(
    classifier_model_path="models/classifier/best_classifier.pt",
    ner_model_path="models/ner/best_ner.pt",
    use_gpu=False
)

# Process document
result = pipeline.process_document("invoice.pdf")

print(f"Document Type: {result['pages'][0]['document_type']}")
print(f"Fields: {result['pages'][0]['fields']}")
```

### API Server

```bash
# Start FastAPI server
python api_server.py
# Server runs on http://localhost:7860

# Health check
curl http://localhost:7860/health

# Process document
curl -X POST http://localhost:7860/process \
  -F "file=@invoice.pdf"
```

---

## πŸ“ Repository Structure

```
IDP-Machine-learning/
β”œβ”€β”€ preprocessing.py          # Image preprocessing (OpenCV)
β”œβ”€β”€ ocr_engine.py             # OCR integration (EasyOCR)
β”œβ”€β”€ classifier_model.py       # Document classifier model
β”œβ”€β”€ ner_model.py              # NER model for entity extraction
β”œβ”€β”€ postprocessing.py         # Output validation & formatting
β”œβ”€β”€ inference_pipeline.py     # Unified inference pipeline
β”œβ”€β”€ api_server.py             # FastAPI REST API
β”œβ”€β”€ train_classifier.py       # Classifier training script
β”œβ”€β”€ train_ner.py              # NER training script
β”œβ”€β”€ dataset_loader.py         # Dataset loading utilities
β”œβ”€β”€ model_optimizer.py        # ONNX conversion & quantization
β”œβ”€β”€ demo_mode.py              # Fallback rule-based logic
β”œβ”€β”€ models/
β”‚   β”œβ”€β”€ classifier/
β”‚   β”‚   └── best_classifier.pt
β”‚   └── ner/
β”‚       └── best_ner.pt
β”œβ”€β”€ frontend/                 # Next.js frontend application
└── requirements.txt          # Python dependencies
```

---

## πŸ“€ API Response Format

```json
{
  "filename": "invoice.pdf",
  "file_type": "pdf",
  "total_pages": 1,
  "pages": [{
    "document_type": "INVOICE",
    "classification_confidence": 0.96,
    "fields": {
      "invoice_number": {
        "value": "INV-12345",
        "confidence": 0.92,
        "source": "ner"
      },
      "date": {
        "value": "2024-01-15",
        "confidence": 0.88,
        "normalized": true
      },
      "total_amount": {
        "value": "12500.00",
        "numeric_value": 12500.0,
        "currency": "INR",
        "confidence": 0.95
      }
    },
    "processing_time": {
      "total": 0.92
    }
  }]
}
```

---

## πŸ”§ Technology Stack

| Component | Technology |
|-----------|------------|
| **Deep Learning** | PyTorch 2.x |
| **NLP Models** | Hugging Face Transformers |
| **OCR** | EasyOCR |
| **Image Processing** | OpenCV |
| **API Framework** | FastAPI + Uvicorn |
| **Frontend** | Next.js 14 + React 18 |

---

## πŸ“ˆ Real-World Performance

| Document Quality | Accuracy |
|------------------|----------|
| High-quality scans | 93-96% |
| Standard photos | 85-92% |
| Poor quality/handwritten | 65-80% |

### Tips for Better Accuracy
- Use high-resolution scans (300+ DPI)
- Ensure good lighting for photos
- Enable adaptive thresholding for low-quality images
- Train on domain-specific data for best results

---

## 🀝 Contributing

Contributions are welcome! Please feel free to submit issues and pull requests.

---

## πŸ“„ License

This project is licensed under the Apache 2.0 License.

---

## πŸ™ Acknowledgments

Built with:
- [Hugging Face Transformers](https://huggingface.co/transformers)
- [EasyOCR](https://github.com/JaidedAI/EasyOCR)
- [FastAPI](https://fastapi.tiangolo.com/)
- [OpenCV](https://opencv.org/)

### Datasets
- [CORD-v2](https://huggingface.co/datasets/naver-clova-ix/cord-v2) - Consolidated Receipt Dataset
- [SROIE](https://rrc.cvc.uab.es/?ch=13) - ICDAR 2019 Competition Dataset
- [FUNSD](https://guillaumejaume.github.io/FUNSD/) - Form Understanding Dataset

---

## πŸ“§ Contact

For questions and support, please open an issue in the repository.