naver-clova-ix/cord-v2
Viewer β’ Updated β’ 1k β’ 8.82k β’ 130
How to use mrrobot2610/IDP-Machine-learning with Transformers:
# Use a pipeline as a high-level helper
from transformers import pipeline
pipe = pipeline("document-question-answering", model="mrrobot2610/IDP-Machine-learning") # pip install -U transformers accelerate
# Load model directly
from transformers import AutoModel
model = AutoModel.from_pretrained("mrrobot2610/IDP-Machine-learning", device_map="auto")# pip install -U transformers accelerate
# Load model directly
from transformers import AutoModel
model = AutoModel.from_pretrained("mrrobot2610/IDP-Machine-learning", device_map="auto")Production-grade AI-powered document processing system for extracting structured data from documents
The IDP (Intelligent Document Processing) System is a complete end-to-end pipeline for:
| Task | Metric | Score |
|---|---|---|
| Document Classification | Accuracy | 92.3% |
| NER Field Extraction | F1 Score | 87.1% |
| Overall Pipeline | Field Accuracy | 89.5% |
| Metric | CPU (Intel i7) | GPU (T4) |
|---|---|---|
| Single page processing | 1.2s | 0.4s |
| Memory usage | 450MB | 2.1GB |
| Throughput | ~50 docs/min | ~150 docs/min |
| Property | Value |
|---|---|
| Base Model | nreimers/MiniLM-L6-H384-uncased |
| Parameters | 22M |
| Task | Text Classification |
| Accuracy | >90% on test set |
Supported Classes:
INVOICE - Invoices and billsRECEIPT - Purchase receiptsFORM - Application forms, tax formsBANK_STATEMENT - Bank statementsOTHER - Other document types| Property | Value |
|---|---|
| Base Model | distilbert-base-uncased |
| Parameters | 66M |
| Task | Token Classification (BIO tagging) |
| F1 Score | >85% on test set |
Extracted Entities:
| Entity | Example |
|---|---|
INVOICE_NUMBER |
INV-12345, #2024-001 |
DATE |
2024-01-15, Jan 15 2024 |
TOTAL_AMOUNT |
$1,234.56, βΉ12,500.00 |
TAX_AMOUNT |
$99.99, Tax: 18% |
VENDOR_NAME |
Acme Corporation |
CUSTOMER_NAME |
John Smith |
GST_ID |
27AAAC11234X1Z5 |
ADDRESS |
123 Main St, City |
| Property | Value |
|---|---|
| Engine | EasyOCR |
| Size | ~10MB |
| Speed | <0.5s per page on CPU |
| Languages | English |
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β INFERENCE PIPELINE β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ€
β 1. Preprocessing (OpenCV) β
β ββ Resize, Denoise, Deskew, Threshold, Enhance β
β βΌ β
β 2. OCR (EasyOCR) β
β ββ Text + Bounding Boxes + Confidence β
β βΌ β
β 3. Classification (MiniLM) β
β ββ Document Type + Confidence β
β βΌ β
β 4. NER (DistilBERT) β
β ββ Entity Extraction (BIO tagging) β
β βΌ β
β 5. Post-Processing β
β ββ Regex Fallbacks + Validation + Normalization β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
The models were trained on standard document understanding datasets:
| Dataset | Size | Document Type |
|---|---|---|
| CORD-v2 | ~1,000 samples | Receipts |
| SROIE | ~1,000 samples | Receipts |
| FUNSD | ~200 samples | Forms |
Classifier:
NER:
# Clone the repository
git clone https://huggingface.co/mrrobot2610/IDP-Machine-learning
# Install dependencies
pip install -r requirements.txt
# System dependencies (PDF support)
# Ubuntu/Debian
sudo apt-get install poppler-utils
# macOS
brew install poppler
from inference_pipeline import IDPPipeline
# Initialize pipeline
pipeline = IDPPipeline(
classifier_model_path="models/classifier/best_classifier.pt",
ner_model_path="models/ner/best_ner.pt",
use_gpu=False
)
# Process document
result = pipeline.process_document("invoice.pdf")
print(f"Document Type: {result['pages'][0]['document_type']}")
print(f"Fields: {result['pages'][0]['fields']}")
# Start FastAPI server
python api_server.py
# Server runs on http://localhost:7860
# Health check
curl http://localhost:7860/health
# Process document
curl -X POST http://localhost:7860/process \
-F "file=@invoice.pdf"
IDP-Machine-learning/
βββ preprocessing.py # Image preprocessing (OpenCV)
βββ ocr_engine.py # OCR integration (EasyOCR)
βββ classifier_model.py # Document classifier model
βββ ner_model.py # NER model for entity extraction
βββ postprocessing.py # Output validation & formatting
βββ inference_pipeline.py # Unified inference pipeline
βββ api_server.py # FastAPI REST API
βββ train_classifier.py # Classifier training script
βββ train_ner.py # NER training script
βββ dataset_loader.py # Dataset loading utilities
βββ model_optimizer.py # ONNX conversion & quantization
βββ demo_mode.py # Fallback rule-based logic
βββ models/
β βββ classifier/
β β βββ best_classifier.pt
β βββ ner/
β βββ best_ner.pt
βββ frontend/ # Next.js frontend application
βββ requirements.txt # Python dependencies
{
"filename": "invoice.pdf",
"file_type": "pdf",
"total_pages": 1,
"pages": [{
"document_type": "INVOICE",
"classification_confidence": 0.96,
"fields": {
"invoice_number": {
"value": "INV-12345",
"confidence": 0.92,
"source": "ner"
},
"date": {
"value": "2024-01-15",
"confidence": 0.88,
"normalized": true
},
"total_amount": {
"value": "12500.00",
"numeric_value": 12500.0,
"currency": "INR",
"confidence": 0.95
}
},
"processing_time": {
"total": 0.92
}
}]
}
| Component | Technology |
|---|---|
| Deep Learning | PyTorch 2.x |
| NLP Models | Hugging Face Transformers |
| OCR | EasyOCR |
| Image Processing | OpenCV |
| API Framework | FastAPI + Uvicorn |
| Frontend | Next.js 14 + React 18 |
| Document Quality | Accuracy |
|---|---|
| High-quality scans | 93-96% |
| Standard photos | 85-92% |
| Poor quality/handwritten | 65-80% |
Contributions are welcome! Please feel free to submit issues and pull requests.
This project is licensed under the Apache 2.0 License.
Built with:
For questions and support, please open an issue in the repository.
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("document-question-answering", model="mrrobot2610/IDP-Machine-learning")