alphabet-classifier / README.md
hermitkk's picture
Update model card: document both alphanumeric and alphabet models with real metrics
7731d62
|
Raw
History Blame Contribute Delete
7.25 kB
---
license: mit
tags:
- image-classification
- onnx
- pytorch
- letters
- digits
- alphanumeric
- emnist
datasets:
- emnist
metrics:
- accuracy
---
# Handwritten Character Classifier (EMNIST, MobileNetV2)
Two MobileNetV2-based ONNX models for classifying grayscale images of handwritten characters. Both models predict **case-insensitively** — they are trained on both upper and lowercase handwriting but always output a single canonical label per letter.
| Model | Classes | Val Accuracy | ONNX |
|---|---|---|---|
| **Alphanumeric** | Digits 0–9 + Letters A–Z + blank (37 total) | **91.42%** | `outputs/exports/alphanumeric_model.onnx` |
| **Alphabet** | Letters A–Z + blank (27 total) | **96.18%** | `outputs/exports/alphabet_model.onnx` |
Use the **alphabet model** when your input is guaranteed to be a letter (higher accuracy, no digit/letter confusion). Use the **alphanumeric model** when the input may be a digit or a letter.
---
## Model Details
### Shared architecture
| Property | Value |
|---|---|
| Backbone | MobileNetV2 (pretrained ImageNet) |
| Head | Linear(1280→256) → BatchNorm → ReLU → Dropout(0.3) → Linear(256→*N*) |
| Parameters | ~2.56M |
| Input | 96 × 96 grayscale (expanded to 3-channel internally) |
| Format | ONNX (opset 17) |
### Alphanumeric model — 37 classes
| Property | Value |
|---|---|
| Classes | `0``9` (indices 0–9), `A``Z` (indices 10–35), `blank` (index 36) |
| Dataset | EMNIST byclass — both upper and lowercase handwriting, labels folded to uppercase |
| Train samples | 711,932 |
| Val samples | 118,323 |
| Val accuracy | **91.42%** |
| Macro avg F1 | 0.917 |
### Alphabet model — 27 classes
| Property | Value |
|---|---|
| Classes | `A``Z` (indices 0–25), `blank` (index 26) |
| Dataset | EMNIST letters — upper and lowercase handwriting already merged at source |
| Train samples | 138,800 |
| Val samples | 22,800 |
| Val accuracy | **96.18%** |
| Macro avg F1 | 0.960 |
---
## Training Configuration
| Setting | Value |
|---|---|
| Optimizer | AdamW — backbone lr × 0.1, head lr 5e-4, weight_decay 1e-4 |
| Scheduler | LinearLR warmup (5 epochs) → CosineAnnealingLR |
| Loss | CrossEntropyLoss with inverse-frequency class weights + label smoothing 0.10 |
| Batch size | 256 |
| Max epochs | 50 |
| Early stopping | patience 10 |
| Augmentation | RandomAffine, RandomPerspective, ColorJitter, GaussianBlur, RandomErasing |
| Mixed precision | AMP (CUDA only) |
| Blank class | Synthetic white images; hard-floored weight ≥ 3.0 |
---
## Per-class Performance
### Alphanumeric model
#### Digits
| Class | Precision | Recall | F1 |
|---|---|---|---|
| 0 | 0.741 | 0.686 | 0.712 |
| 1 | 0.748 | 0.693 | 0.719 |
| 2 | 0.984 | 0.949 | 0.967 |
| 3 | 0.997 | 0.995 | 0.996 |
| 4 | 0.989 | 0.974 | 0.981 |
| 5 | 0.985 | 0.928 | 0.956 |
| 6 | 0.987 | 0.976 | 0.981 |
| 7 | 0.994 | 0.997 | 0.995 |
| 8 | 0.994 | 0.988 | 0.991 |
| 9 | 0.953 | 0.962 | 0.958 |
#### Letters
| Class | Precision | Recall | F1 |
|---|---|---|---|
| A | 0.977 | 0.969 | 0.973 |
| B | 0.914 | 0.971 | 0.942 |
| C | 0.966 | 0.985 | 0.975 |
| D | 0.961 | 0.976 | 0.969 |
| E | 0.992 | 0.989 | 0.990 |
| F | 0.985 | 0.983 | 0.984 |
| G | 0.818 | 0.821 | 0.820 |
| H | 0.967 | 0.982 | 0.974 |
| I | 0.545 | 0.684 | **0.607** |
| J | 0.930 | 0.950 | 0.940 |
| K | 0.987 | 0.993 | 0.990 |
| L | 0.573 | 0.537 | **0.554** |
| M | 0.992 | 0.998 | 0.995 |
| N | 0.985 | 0.984 | 0.984 |
| O | 0.648 | 0.696 | **0.671** |
| P | 0.986 | 0.994 | 0.990 |
| Q | 0.783 | 0.761 | 0.772 |
| R | 0.986 | 0.982 | 0.984 |
| S | 0.913 | 0.978 | 0.944 |
| T | 0.989 | 0.989 | 0.989 |
| U | 0.969 | 0.955 | 0.961 |
| V | 0.912 | 0.956 | 0.934 |
| W | 0.988 | 0.997 | 0.992 |
| X | 0.976 | 0.990 | 0.983 |
| Y | 0.896 | 0.946 | 0.920 |
| Z | 0.762 | 0.921 | 0.834 |
#### Blank
| Class | Precision | Recall | F1 |
|---|---|---|---|
| blank | 1.000 | 1.000 | **1.000** |
Hardest cases are **I** (F1=0.607) and **L** (F1=0.554), both confused with digit `1`; and **O** (F1=0.671), confused with digit `0`. These are inherent digit–letter ambiguities in alphanumeric OCR.
---
### Alphabet model
| Class | Precision | Recall | F1 |
|---|---|---|---|
| A | 0.959 | 0.976 | 0.968 |
| B | 0.996 | 0.986 | 0.991 |
| C | 0.984 | 0.979 | 0.981 |
| D | 0.976 | 0.974 | 0.975 |
| E | 0.984 | 0.986 | 0.985 |
| F | 0.994 | 0.976 | 0.985 |
| G | 0.927 | 0.874 | 0.900 |
| H | 0.976 | 0.978 | 0.977 |
| I | 0.760 | 0.761 | **0.761** |
| J | 0.974 | 0.965 | 0.969 |
| K | 0.996 | 0.995 | 0.996 |
| L | 0.765 | 0.770 | **0.768** |
| M | 0.986 | 0.999 | 0.993 |
| N | 0.980 | 0.981 | 0.981 |
| O | 0.975 | 0.980 | 0.978 |
| P | 0.991 | 0.991 | 0.991 |
| Q | 0.890 | 0.928 | 0.908 |
| R | 0.979 | 0.978 | 0.978 |
| S | 0.992 | 0.989 | 0.991 |
| T | 0.975 | 0.988 | 0.981 |
| U | 0.961 | 0.944 | 0.952 |
| V | 0.945 | 0.958 | 0.951 |
| W | 0.997 | 0.990 | 0.993 |
| X | 0.990 | 0.991 | 0.991 |
| Y | 0.965 | 0.978 | 0.971 |
| Z | 0.995 | 0.998 | 0.996 |
| blank | 1.000 | 1.000 | **1.000** |
Hardest cases are **I** (F1=0.761) and **L** (F1=0.768), which are visually similar across handwriting styles. All other letters achieve F1 ≥ 0.90, and blank is perfect.
---
## Usage
### Alphanumeric model
```python
from huggingface_hub import hf_hub_download
import onnxruntime as ort
import numpy as np
from PIL import Image
CHAR_CLASSES = list("0123456789ABCDEFGHIJKLMNOPQRSTUVWXYZ") + ["blank"]
path = hf_hub_download(
repo_id="hermitkk/alphabet-classifier",
filename="outputs/exports/alphanumeric_model.onnx",
)
session = ort.InferenceSession(path)
# Preprocess a 96x96 grayscale crop
img = Image.open("character.png").convert("L").resize((96, 96))
x = (np.array(img, dtype=np.float32) / 255.0 - 0.5) / 0.5
x = x[np.newaxis, np.newaxis, :, :] # (1, 1, 96, 96)
logits = session.run(None, {"input": x})[0]
pred = int(np.argmax(logits))
print(CHAR_CLASSES[pred]) # e.g. "A", "3", "blank"
```
### Alphabet model
```python
from huggingface_hub import hf_hub_download
import onnxruntime as ort
import numpy as np
from PIL import Image
ALPHA_CLASSES = list("ABCDEFGHIJKLMNOPQRSTUVWXYZ") + ["blank"]
path = hf_hub_download(
repo_id="hermitkk/alphabet-classifier",
filename="outputs/exports/alphabet_model.onnx",
)
session = ort.InferenceSession(path)
# Preprocess a 96x96 grayscale crop
img = Image.open("letter.png").convert("L").resize((96, 96))
x = (np.array(img, dtype=np.float32) / 255.0 - 0.5) / 0.5
x = x[np.newaxis, np.newaxis, :, :] # (1, 1, 96, 96)
logits = session.run(None, {"input": x})[0]
pred = int(np.argmax(logits))
print(ALPHA_CLASSES[pred]) # e.g. "A", "blank"
```
> **Note:** both models accept 96 × 96 single-channel float32 input, normalized to mean 0.5 / std 0.5. White pixels (blank paper) map to +1.0 and dark ink maps toward −1.0.
---
## Reproduce Training
```bash
git clone https://huggingface.co/hermitkk/alphabet-classifier
cd alphabet-classifier
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
# Train alphanumeric model
python main.py train --config config/config.yaml
# Train alphabet model
python main.py train --config config/config_alphabet.yaml
```
---
## License
MIT