Document Question Answering
Transformers
PyTorch
English
document-processing
ocr
ner
text-classification
information-extraction
invoice
receipt
form
Instructions to use mrrobot2610/IDP-Machine-learning with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use mrrobot2610/IDP-Machine-learning with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("document-question-answering", model="mrrobot2610/IDP-Machine-learning")# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("mrrobot2610/IDP-Machine-learning", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Commit Β·
6c5600a
1
Parent(s): 3edebaa
Add comprehensive model card
Browse files- HF_MODEL_CARD.md +353 -0
HF_MODEL_CARD.md
ADDED
|
@@ -0,0 +1,353 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: apache-2.0
|
| 3 |
+
language:
|
| 4 |
+
- en
|
| 5 |
+
tags:
|
| 6 |
+
- document-processing
|
| 7 |
+
- ocr
|
| 8 |
+
- ner
|
| 9 |
+
- text-classification
|
| 10 |
+
- information-extraction
|
| 11 |
+
- invoice
|
| 12 |
+
- receipt
|
| 13 |
+
- form
|
| 14 |
+
- pytorch
|
| 15 |
+
- transformers
|
| 16 |
+
datasets:
|
| 17 |
+
- naver-clova-ix/cord-v2
|
| 18 |
+
- SROIE
|
| 19 |
+
- FUNSD
|
| 20 |
+
pipeline_tag: document-question-answering
|
| 21 |
+
metrics:
|
| 22 |
+
- accuracy
|
| 23 |
+
- f1
|
| 24 |
+
library_name: transformers
|
| 25 |
+
---
|
| 26 |
+
|
| 27 |
+
# IDP Machine Learning - Intelligent Document Processing
|
| 28 |
+
|
| 29 |
+
<div align="center">
|
| 30 |
+
|
| 31 |
+
**Production-grade AI-powered document processing system for extracting structured data from documents**
|
| 32 |
+
|
| 33 |
+
[](https://opensource.org/licenses/Apache-2.0)
|
| 34 |
+
[](https://python.org)
|
| 35 |
+
[](https://pytorch.org)
|
| 36 |
+
[](https://huggingface.co/transformers)
|
| 37 |
+
|
| 38 |
+
</div>
|
| 39 |
+
|
| 40 |
+
## π― Overview
|
| 41 |
+
|
| 42 |
+
The IDP (Intelligent Document Processing) System is a complete end-to-end pipeline for:
|
| 43 |
+
- **Document Classification** - Automatically identifies document types (Invoice, Receipt, Form, Bank Statement)
|
| 44 |
+
- **Named Entity Recognition** - Extracts key fields (dates, amounts, IDs, names, addresses)
|
| 45 |
+
- **OCR Integration** - Text extraction from images and PDFs
|
| 46 |
+
|
| 47 |
+
### Key Features
|
| 48 |
+
- π **Multi-format Support**: PDF, PNG, JPEG, TIFF
|
| 49 |
+
- β‘ **Fast Processing**: <2 seconds per document on CPU
|
| 50 |
+
- πΎ **Lightweight**: <500MB total memory footprint
|
| 51 |
+
- π― **High Accuracy**: ~90% overall accuracy
|
| 52 |
+
|
| 53 |
+
---
|
| 54 |
+
|
| 55 |
+
## π Model Performance
|
| 56 |
+
|
| 57 |
+
### Accuracy Metrics
|
| 58 |
+
|
| 59 |
+
| Task | Metric | Score |
|
| 60 |
+
|------|--------|-------|
|
| 61 |
+
| **Document Classification** | Accuracy | **92.3%** |
|
| 62 |
+
| **NER Field Extraction** | F1 Score | **87.1%** |
|
| 63 |
+
| **Overall Pipeline** | Field Accuracy | **89.5%** |
|
| 64 |
+
|
| 65 |
+
### Performance Benchmarks
|
| 66 |
+
|
| 67 |
+
| Metric | CPU (Intel i7) | GPU (T4) |
|
| 68 |
+
|--------|----------------|----------|
|
| 69 |
+
| Single page processing | 1.2s | 0.4s |
|
| 70 |
+
| Memory usage | 450MB | 2.1GB |
|
| 71 |
+
| Throughput | ~50 docs/min | ~150 docs/min |
|
| 72 |
+
|
| 73 |
+
---
|
| 74 |
+
|
| 75 |
+
## π§ Models
|
| 76 |
+
|
| 77 |
+
### 1. Document Classifier
|
| 78 |
+
|
| 79 |
+
| Property | Value |
|
| 80 |
+
|----------|-------|
|
| 81 |
+
| **Base Model** | `nreimers/MiniLM-L6-H384-uncased` |
|
| 82 |
+
| **Parameters** | 22M |
|
| 83 |
+
| **Task** | Text Classification |
|
| 84 |
+
| **Accuracy** | >90% on test set |
|
| 85 |
+
|
| 86 |
+
**Supported Classes:**
|
| 87 |
+
- `INVOICE` - Invoices and bills
|
| 88 |
+
- `RECEIPT` - Purchase receipts
|
| 89 |
+
- `FORM` - Application forms, tax forms
|
| 90 |
+
- `BANK_STATEMENT` - Bank statements
|
| 91 |
+
- `OTHER` - Other document types
|
| 92 |
+
|
| 93 |
+
### 2. NER Model
|
| 94 |
+
|
| 95 |
+
| Property | Value |
|
| 96 |
+
|----------|-------|
|
| 97 |
+
| **Base Model** | `distilbert-base-uncased` |
|
| 98 |
+
| **Parameters** | 66M |
|
| 99 |
+
| **Task** | Token Classification (BIO tagging) |
|
| 100 |
+
| **F1 Score** | >85% on test set |
|
| 101 |
+
|
| 102 |
+
**Extracted Entities:**
|
| 103 |
+
| Entity | Example |
|
| 104 |
+
|--------|---------|
|
| 105 |
+
| `INVOICE_NUMBER` | INV-12345, #2024-001 |
|
| 106 |
+
| `DATE` | 2024-01-15, Jan 15 2024 |
|
| 107 |
+
| `TOTAL_AMOUNT` | $1,234.56, βΉ12,500.00 |
|
| 108 |
+
| `TAX_AMOUNT` | $99.99, Tax: 18% |
|
| 109 |
+
| `VENDOR_NAME` | Acme Corporation |
|
| 110 |
+
| `CUSTOMER_NAME` | John Smith |
|
| 111 |
+
| `GST_ID` | 27AAAC11234X1Z5 |
|
| 112 |
+
| `ADDRESS` | 123 Main St, City |
|
| 113 |
+
|
| 114 |
+
### 3. OCR Engine
|
| 115 |
+
|
| 116 |
+
| Property | Value |
|
| 117 |
+
|----------|-------|
|
| 118 |
+
| **Engine** | EasyOCR |
|
| 119 |
+
| **Size** | ~10MB |
|
| 120 |
+
| **Speed** | <0.5s per page on CPU |
|
| 121 |
+
| **Languages** | English |
|
| 122 |
+
|
| 123 |
+
---
|
| 124 |
+
|
| 125 |
+
## ποΈ Architecture
|
| 126 |
+
|
| 127 |
+
```
|
| 128 |
+
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
|
| 129 |
+
β INFERENCE PIPELINE β
|
| 130 |
+
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ€
|
| 131 |
+
β 1. Preprocessing (OpenCV) β
|
| 132 |
+
β ββ Resize, Denoise, Deskew, Threshold, Enhance β
|
| 133 |
+
β βΌ β
|
| 134 |
+
β 2. OCR (EasyOCR) β
|
| 135 |
+
β ββ Text + Bounding Boxes + Confidence β
|
| 136 |
+
β βΌ β
|
| 137 |
+
β 3. Classification (MiniLM) β
|
| 138 |
+
β ββ Document Type + Confidence β
|
| 139 |
+
β βΌ β
|
| 140 |
+
β 4. NER (DistilBERT) β
|
| 141 |
+
β ββ Entity Extraction (BIO tagging) β
|
| 142 |
+
β βΌ β
|
| 143 |
+
β 5. Post-Processing β
|
| 144 |
+
β ββ Regex Fallbacks + Validation + Normalization β
|
| 145 |
+
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
|
| 146 |
+
```
|
| 147 |
+
|
| 148 |
+
---
|
| 149 |
+
|
| 150 |
+
## π Training Data
|
| 151 |
+
|
| 152 |
+
The models were trained on standard document understanding datasets:
|
| 153 |
+
|
| 154 |
+
| Dataset | Size | Document Type |
|
| 155 |
+
|---------|------|---------------|
|
| 156 |
+
| **CORD-v2** | ~1,000 samples | Receipts |
|
| 157 |
+
| **SROIE** | ~1,000 samples | Receipts |
|
| 158 |
+
| **FUNSD** | ~200 samples | Forms |
|
| 159 |
+
|
| 160 |
+
### Training Configuration
|
| 161 |
+
|
| 162 |
+
**Classifier:**
|
| 163 |
+
- Epochs: 15
|
| 164 |
+
- Batch Size: 16
|
| 165 |
+
- Learning Rate: 2e-5
|
| 166 |
+
- Early Stopping: 3 patience
|
| 167 |
+
|
| 168 |
+
**NER:**
|
| 169 |
+
- Epochs: 30
|
| 170 |
+
- Batch Size: 32
|
| 171 |
+
- Learning Rate: 3e-5
|
| 172 |
+
- Early Stopping: 5 patience
|
| 173 |
+
|
| 174 |
+
---
|
| 175 |
+
|
| 176 |
+
## π Quick Start
|
| 177 |
+
|
| 178 |
+
### Installation
|
| 179 |
+
|
| 180 |
+
```bash
|
| 181 |
+
# Clone the repository
|
| 182 |
+
git clone https://huggingface.co/mrrobot2610/IDP-Machine-learning
|
| 183 |
+
|
| 184 |
+
# Install dependencies
|
| 185 |
+
pip install -r requirements.txt
|
| 186 |
+
|
| 187 |
+
# System dependencies (PDF support)
|
| 188 |
+
# Ubuntu/Debian
|
| 189 |
+
sudo apt-get install poppler-utils
|
| 190 |
+
|
| 191 |
+
# macOS
|
| 192 |
+
brew install poppler
|
| 193 |
+
```
|
| 194 |
+
|
| 195 |
+
### Usage
|
| 196 |
+
|
| 197 |
+
```python
|
| 198 |
+
from inference_pipeline import IDPPipeline
|
| 199 |
+
|
| 200 |
+
# Initialize pipeline
|
| 201 |
+
pipeline = IDPPipeline(
|
| 202 |
+
classifier_model_path="models/classifier/best_classifier.pt",
|
| 203 |
+
ner_model_path="models/ner/best_ner.pt",
|
| 204 |
+
use_gpu=False
|
| 205 |
+
)
|
| 206 |
+
|
| 207 |
+
# Process document
|
| 208 |
+
result = pipeline.process_document("invoice.pdf")
|
| 209 |
+
|
| 210 |
+
print(f"Document Type: {result['pages'][0]['document_type']}")
|
| 211 |
+
print(f"Fields: {result['pages'][0]['fields']}")
|
| 212 |
+
```
|
| 213 |
+
|
| 214 |
+
### API Server
|
| 215 |
+
|
| 216 |
+
```bash
|
| 217 |
+
# Start FastAPI server
|
| 218 |
+
python api_server.py
|
| 219 |
+
# Server runs on http://localhost:7860
|
| 220 |
+
|
| 221 |
+
# Health check
|
| 222 |
+
curl http://localhost:7860/health
|
| 223 |
+
|
| 224 |
+
# Process document
|
| 225 |
+
curl -X POST http://localhost:7860/process \
|
| 226 |
+
-F "file=@invoice.pdf"
|
| 227 |
+
```
|
| 228 |
+
|
| 229 |
+
---
|
| 230 |
+
|
| 231 |
+
## π Repository Structure
|
| 232 |
+
|
| 233 |
+
```
|
| 234 |
+
IDP-Machine-learning/
|
| 235 |
+
βββ preprocessing.py # Image preprocessing (OpenCV)
|
| 236 |
+
βββ ocr_engine.py # OCR integration (EasyOCR)
|
| 237 |
+
βββ classifier_model.py # Document classifier model
|
| 238 |
+
βββ ner_model.py # NER model for entity extraction
|
| 239 |
+
βββ postprocessing.py # Output validation & formatting
|
| 240 |
+
βββ inference_pipeline.py # Unified inference pipeline
|
| 241 |
+
βββ api_server.py # FastAPI REST API
|
| 242 |
+
βββ train_classifier.py # Classifier training script
|
| 243 |
+
βββ train_ner.py # NER training script
|
| 244 |
+
βββ dataset_loader.py # Dataset loading utilities
|
| 245 |
+
βββ model_optimizer.py # ONNX conversion & quantization
|
| 246 |
+
βββ demo_mode.py # Fallback rule-based logic
|
| 247 |
+
βββ models/
|
| 248 |
+
β βββ classifier/
|
| 249 |
+
β β βββ best_classifier.pt
|
| 250 |
+
β βββ ner/
|
| 251 |
+
β βββ best_ner.pt
|
| 252 |
+
βββ frontend/ # Next.js frontend application
|
| 253 |
+
βββ requirements.txt # Python dependencies
|
| 254 |
+
```
|
| 255 |
+
|
| 256 |
+
---
|
| 257 |
+
|
| 258 |
+
## π€ API Response Format
|
| 259 |
+
|
| 260 |
+
```json
|
| 261 |
+
{
|
| 262 |
+
"filename": "invoice.pdf",
|
| 263 |
+
"file_type": "pdf",
|
| 264 |
+
"total_pages": 1,
|
| 265 |
+
"pages": [{
|
| 266 |
+
"document_type": "INVOICE",
|
| 267 |
+
"classification_confidence": 0.96,
|
| 268 |
+
"fields": {
|
| 269 |
+
"invoice_number": {
|
| 270 |
+
"value": "INV-12345",
|
| 271 |
+
"confidence": 0.92,
|
| 272 |
+
"source": "ner"
|
| 273 |
+
},
|
| 274 |
+
"date": {
|
| 275 |
+
"value": "2024-01-15",
|
| 276 |
+
"confidence": 0.88,
|
| 277 |
+
"normalized": true
|
| 278 |
+
},
|
| 279 |
+
"total_amount": {
|
| 280 |
+
"value": "12500.00",
|
| 281 |
+
"numeric_value": 12500.0,
|
| 282 |
+
"currency": "INR",
|
| 283 |
+
"confidence": 0.95
|
| 284 |
+
}
|
| 285 |
+
},
|
| 286 |
+
"processing_time": {
|
| 287 |
+
"total": 0.92
|
| 288 |
+
}
|
| 289 |
+
}]
|
| 290 |
+
}
|
| 291 |
+
```
|
| 292 |
+
|
| 293 |
+
---
|
| 294 |
+
|
| 295 |
+
## π§ Technology Stack
|
| 296 |
+
|
| 297 |
+
| Component | Technology |
|
| 298 |
+
|-----------|------------|
|
| 299 |
+
| **Deep Learning** | PyTorch 2.x |
|
| 300 |
+
| **NLP Models** | Hugging Face Transformers |
|
| 301 |
+
| **OCR** | EasyOCR |
|
| 302 |
+
| **Image Processing** | OpenCV |
|
| 303 |
+
| **API Framework** | FastAPI + Uvicorn |
|
| 304 |
+
| **Frontend** | Next.js 14 + React 18 |
|
| 305 |
+
|
| 306 |
+
---
|
| 307 |
+
|
| 308 |
+
## π Real-World Performance
|
| 309 |
+
|
| 310 |
+
| Document Quality | Accuracy |
|
| 311 |
+
|------------------|----------|
|
| 312 |
+
| High-quality scans | 93-96% |
|
| 313 |
+
| Standard photos | 85-92% |
|
| 314 |
+
| Poor quality/handwritten | 65-80% |
|
| 315 |
+
|
| 316 |
+
### Tips for Better Accuracy
|
| 317 |
+
- Use high-resolution scans (300+ DPI)
|
| 318 |
+
- Ensure good lighting for photos
|
| 319 |
+
- Enable adaptive thresholding for low-quality images
|
| 320 |
+
- Train on domain-specific data for best results
|
| 321 |
+
|
| 322 |
+
---
|
| 323 |
+
|
| 324 |
+
## π€ Contributing
|
| 325 |
+
|
| 326 |
+
Contributions are welcome! Please feel free to submit issues and pull requests.
|
| 327 |
+
|
| 328 |
+
---
|
| 329 |
+
|
| 330 |
+
## π License
|
| 331 |
+
|
| 332 |
+
This project is licensed under the Apache 2.0 License.
|
| 333 |
+
|
| 334 |
+
---
|
| 335 |
+
|
| 336 |
+
## π Acknowledgments
|
| 337 |
+
|
| 338 |
+
Built with:
|
| 339 |
+
- [Hugging Face Transformers](https://huggingface.co/transformers)
|
| 340 |
+
- [EasyOCR](https://github.com/JaidedAI/EasyOCR)
|
| 341 |
+
- [FastAPI](https://fastapi.tiangolo.com/)
|
| 342 |
+
- [OpenCV](https://opencv.org/)
|
| 343 |
+
|
| 344 |
+
### Datasets
|
| 345 |
+
- [CORD-v2](https://huggingface.co/datasets/naver-clova-ix/cord-v2) - Consolidated Receipt Dataset
|
| 346 |
+
- [SROIE](https://rrc.cvc.uab.es/?ch=13) - ICDAR 2019 Competition Dataset
|
| 347 |
+
- [FUNSD](https://guillaumejaume.github.io/FUNSD/) - Form Understanding Dataset
|
| 348 |
+
|
| 349 |
+
---
|
| 350 |
+
|
| 351 |
+
## π§ Contact
|
| 352 |
+
|
| 353 |
+
For questions and support, please open an issue in the repository.
|