Jasirdeen's picture
Upload README.md with huggingface_hub
30cc0fb verified
|
Raw
History Blame Contribute Delete
5.63 kB
---
library_name: pytorch
tags:
- image-classification
- computer-vision
- document-forgery-detection
- identity-document-analysis
license: mit
---
# EfficientNet-B3 Image Forgery Detection
This repository contains a fine-tuned EfficientNet-B3 model for binary classification of synthetic identity-document images as genuine or forged.
The checkpoint is provided as a PyTorch `state_dict`:
```text
efficientnet_b3_finetuned.pt
```
## Model description
- Architecture: EfficientNet-B3
- Framework: PyTorch and `timm`
- Task: binary image classification
- Input: one identity-document image
- Output: genuine or forged probability
- Input resolution: 300 × 300 pixels
Labels:
```text
0 = genuine / bona fide
1 = forged / tampered
```
## Dataset
The model was trained and evaluated on the template/composite portion of the Synthetic dataset of ID and Travel Documents (SIDTD), generated using MIDV-2020 document templates.
Dataset composition used in this project:
```text
Genuine images: 1,000
Forged images: 1,222
Total images: 2,222
```
The forged samples include the following SIDTD annotation types:
- `Crop_and_Replace`
- `Inpaint_and_Rewrite`
The dataset itself is not included in this repository. Please follow the SIDTD dataset terms and citation requirements before downloading or redistributing it.
## Preprocessing
Each image is:
1. converted to RGB;
2. resized to 300 × 300 pixels;
3. converted to a tensor;
4. normalized using ImageNet mean and standard deviation:
```text
mean = [0.485, 0.456, 0.406]
std = [0.229, 0.224, 0.225]
```
Training used mild brightness, contrast, and saturation variation. Horizontal flipping was not used because mirrored identity documents are not realistic examples.
## Training
The model was initialized from a pretrained EfficientNet-B3. Phase 1 trained the binary classifier head. The final checkpoint was then fine-tuned by unfreezing the final EfficientNet feature block and classifier head.
Fine-tuning used:
- optimizer: AdamW;
- learning rate: `1e-4`;
- weight decay: `1e-4`;
- epochs: 20;
- binary cross-entropy equivalent: two-class cross-entropy loss;
- deterministic stratified hold-out split: 80% train, 10% validation, 10% test.
## Evaluation
On the fixed 223-image test split at a decision threshold of `0.5`:
```text
Accuracy: 98.21%
Precision: 97.60%
Recall: 99.19%
F1: 98.39%
ROC-AUC: 99.98%
```
Confusion matrix, with rows representing actual labels and columns representing predictions:
```text
[[97, 3],
[ 1, 122]]
```
Subtype recall on the forged test samples:
```text
Crop-and-replace: 94.74%
Inpainting: 100.00%
```
These results are specific to the project’s synthetic dataset and split. They should not be interpreted as real-world identity-verification performance.
## Usage
Install the required packages:
```bash
pip install torch torchvision timm pillow
```
Load the checkpoint:
```python
from pathlib import Path
import timm
import torch
from PIL import Image
from torchvision import transforms
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
model = timm.create_model(
"efficientnet_b3",
pretrained=False,
num_classes=2,
)
model.load_state_dict(
torch.load(
"efficientnet_b3_finetuned.pt",
map_location="cpu",
weights_only=True,
)
)
model.to(device)
model.eval()
preprocess = transforms.Compose([
transforms.Resize((300, 300)),
transforms.ToTensor(),
transforms.Normalize(
mean=(0.485, 0.456, 0.406),
std=(0.229, 0.224, 0.225),
),
])
image = Image.open("document.jpg").convert("RGB")
input_tensor = preprocess(image).unsqueeze(0).to(device)
with torch.no_grad():
probabilities = torch.softmax(model(input_tensor), dim=1)[0]
genuine_probability = float(probabilities[0])
forged_probability = float(probabilities[1])
label = "forged" if forged_probability >= 0.5 else "genuine"
print({
"label": label,
"genuine_probability": genuine_probability,
"forged_probability": forged_probability,
})
```
## Explainability
The project uses Grad-CAM with the final spatial convolutional layer of EfficientNet-B3. The resulting heatmap indicates image regions that contributed to the selected prediction.
Grad-CAM is an explanation aid and should not be interpreted as a guaranteed localization of the tampered region.
## Intended use
This model is intended for:
- educational computer-vision projects;
- research prototypes;
- synthetic identity-document forgery experiments;
- demonstrations of explainable image classification;
- offline error analysis.
## Limitations
- The training data is synthetic.
- The model has not been validated on real identity documents.
- Performance may depend on document templates, image quality, and dataset artifacts.
- The binary output is not equivalent to full identity verification.
- A positive or negative result should not be used as the sole basis for a KYC or access-control decision.
- The reported test split is an image-level random split and may not measure generalization to unseen document families as strictly as a grouped or external split.
## Citation and attribution
This model uses the SIDTD dataset and MIDV-2020-derived document templates. Please cite and follow the terms of the original SIDTD and MIDV-2020 resources when using this model or reproducing the experiments.
SIDTD project: https://github.com/Oriolrt/SIDTD_Dataset
## License
The model repository is released under the MIT license. Dataset licensing, attribution, and redistribution terms remain subject to the original SIDTD and MIDV-2020 sources.