| --- |
| license: mit |
| tags: |
| - image-classification |
| - onnx |
| - pytorch |
| - letters |
| - digits |
| - alphanumeric |
| - emnist |
| datasets: |
| - emnist |
| metrics: |
| - accuracy |
| --- |
| |
| # Handwritten Character Classifier (EMNIST, MobileNetV2) |
|
|
| Two MobileNetV2-based ONNX models for classifying grayscale images of handwritten characters. Both models predict **case-insensitively** — they are trained on both upper and lowercase handwriting but always output a single canonical label per letter. |
|
|
| | Model | Classes | Val Accuracy | ONNX | |
| |---|---|---|---| |
| | **Alphanumeric** | Digits 0–9 + Letters A–Z + blank (37 total) | **91.42%** | `outputs/exports/alphanumeric_model.onnx` | |
| | **Alphabet** | Letters A–Z + blank (27 total) | **96.18%** | `outputs/exports/alphabet_model.onnx` | |
|
|
| Use the **alphabet model** when your input is guaranteed to be a letter (higher accuracy, no digit/letter confusion). Use the **alphanumeric model** when the input may be a digit or a letter. |
|
|
| --- |
|
|
| ## Model Details |
|
|
| ### Shared architecture |
|
|
| | Property | Value | |
| |---|---| |
| | Backbone | MobileNetV2 (pretrained ImageNet) | |
| | Head | Linear(1280→256) → BatchNorm → ReLU → Dropout(0.3) → Linear(256→*N*) | |
| | Parameters | ~2.56M | |
| | Input | 96 × 96 grayscale (expanded to 3-channel internally) | |
| | Format | ONNX (opset 17) | |
|
|
| ### Alphanumeric model — 37 classes |
|
|
| | Property | Value | |
| |---|---| |
| | Classes | `0`–`9` (indices 0–9), `A`–`Z` (indices 10–35), `blank` (index 36) | |
| | Dataset | EMNIST byclass — both upper and lowercase handwriting, labels folded to uppercase | |
| | Train samples | 711,932 | |
| | Val samples | 118,323 | |
| | Val accuracy | **91.42%** | |
| | Macro avg F1 | 0.917 | |
|
|
| ### Alphabet model — 27 classes |
|
|
| | Property | Value | |
| |---|---| |
| | Classes | `A`–`Z` (indices 0–25), `blank` (index 26) | |
| | Dataset | EMNIST letters — upper and lowercase handwriting already merged at source | |
| | Train samples | 138,800 | |
| | Val samples | 22,800 | |
| | Val accuracy | **96.18%** | |
| | Macro avg F1 | 0.960 | |
|
|
| --- |
|
|
| ## Training Configuration |
|
|
| | Setting | Value | |
| |---|---| |
| | Optimizer | AdamW — backbone lr × 0.1, head lr 5e-4, weight_decay 1e-4 | |
| | Scheduler | LinearLR warmup (5 epochs) → CosineAnnealingLR | |
| | Loss | CrossEntropyLoss with inverse-frequency class weights + label smoothing 0.10 | |
| | Batch size | 256 | |
| | Max epochs | 50 | |
| | Early stopping | patience 10 | |
| | Augmentation | RandomAffine, RandomPerspective, ColorJitter, GaussianBlur, RandomErasing | |
| | Mixed precision | AMP (CUDA only) | |
| | Blank class | Synthetic white images; hard-floored weight ≥ 3.0 | |
| |
| --- |
| |
| ## Per-class Performance |
| |
| ### Alphanumeric model |
| |
| #### Digits |
| |
| | Class | Precision | Recall | F1 | |
| |---|---|---|---| |
| | 0 | 0.741 | 0.686 | 0.712 | |
| | 1 | 0.748 | 0.693 | 0.719 | |
| | 2 | 0.984 | 0.949 | 0.967 | |
| | 3 | 0.997 | 0.995 | 0.996 | |
| | 4 | 0.989 | 0.974 | 0.981 | |
| | 5 | 0.985 | 0.928 | 0.956 | |
| | 6 | 0.987 | 0.976 | 0.981 | |
| | 7 | 0.994 | 0.997 | 0.995 | |
| | 8 | 0.994 | 0.988 | 0.991 | |
| | 9 | 0.953 | 0.962 | 0.958 | |
| |
| #### Letters |
| |
| | Class | Precision | Recall | F1 | |
| |---|---|---|---| |
| | A | 0.977 | 0.969 | 0.973 | |
| | B | 0.914 | 0.971 | 0.942 | |
| | C | 0.966 | 0.985 | 0.975 | |
| | D | 0.961 | 0.976 | 0.969 | |
| | E | 0.992 | 0.989 | 0.990 | |
| | F | 0.985 | 0.983 | 0.984 | |
| | G | 0.818 | 0.821 | 0.820 | |
| | H | 0.967 | 0.982 | 0.974 | |
| | I | 0.545 | 0.684 | **0.607** | |
| | J | 0.930 | 0.950 | 0.940 | |
| | K | 0.987 | 0.993 | 0.990 | |
| | L | 0.573 | 0.537 | **0.554** | |
| | M | 0.992 | 0.998 | 0.995 | |
| | N | 0.985 | 0.984 | 0.984 | |
| | O | 0.648 | 0.696 | **0.671** | |
| | P | 0.986 | 0.994 | 0.990 | |
| | Q | 0.783 | 0.761 | 0.772 | |
| | R | 0.986 | 0.982 | 0.984 | |
| | S | 0.913 | 0.978 | 0.944 | |
| | T | 0.989 | 0.989 | 0.989 | |
| | U | 0.969 | 0.955 | 0.961 | |
| | V | 0.912 | 0.956 | 0.934 | |
| | W | 0.988 | 0.997 | 0.992 | |
| | X | 0.976 | 0.990 | 0.983 | |
| | Y | 0.896 | 0.946 | 0.920 | |
| | Z | 0.762 | 0.921 | 0.834 | |
| |
| #### Blank |
| |
| | Class | Precision | Recall | F1 | |
| |---|---|---|---| |
| | blank | 1.000 | 1.000 | **1.000** | |
| |
| Hardest cases are **I** (F1=0.607) and **L** (F1=0.554), both confused with digit `1`; and **O** (F1=0.671), confused with digit `0`. These are inherent digit–letter ambiguities in alphanumeric OCR. |
| |
| --- |
| |
| ### Alphabet model |
| |
| | Class | Precision | Recall | F1 | |
| |---|---|---|---| |
| | A | 0.959 | 0.976 | 0.968 | |
| | B | 0.996 | 0.986 | 0.991 | |
| | C | 0.984 | 0.979 | 0.981 | |
| | D | 0.976 | 0.974 | 0.975 | |
| | E | 0.984 | 0.986 | 0.985 | |
| | F | 0.994 | 0.976 | 0.985 | |
| | G | 0.927 | 0.874 | 0.900 | |
| | H | 0.976 | 0.978 | 0.977 | |
| | I | 0.760 | 0.761 | **0.761** | |
| | J | 0.974 | 0.965 | 0.969 | |
| | K | 0.996 | 0.995 | 0.996 | |
| | L | 0.765 | 0.770 | **0.768** | |
| | M | 0.986 | 0.999 | 0.993 | |
| | N | 0.980 | 0.981 | 0.981 | |
| | O | 0.975 | 0.980 | 0.978 | |
| | P | 0.991 | 0.991 | 0.991 | |
| | Q | 0.890 | 0.928 | 0.908 | |
| | R | 0.979 | 0.978 | 0.978 | |
| | S | 0.992 | 0.989 | 0.991 | |
| | T | 0.975 | 0.988 | 0.981 | |
| | U | 0.961 | 0.944 | 0.952 | |
| | V | 0.945 | 0.958 | 0.951 | |
| | W | 0.997 | 0.990 | 0.993 | |
| | X | 0.990 | 0.991 | 0.991 | |
| | Y | 0.965 | 0.978 | 0.971 | |
| | Z | 0.995 | 0.998 | 0.996 | |
| | blank | 1.000 | 1.000 | **1.000** | |
| |
| Hardest cases are **I** (F1=0.761) and **L** (F1=0.768), which are visually similar across handwriting styles. All other letters achieve F1 ≥ 0.90, and blank is perfect. |
| |
| --- |
| |
| ## Usage |
| |
| ### Alphanumeric model |
| |
| ```python |
| from huggingface_hub import hf_hub_download |
| import onnxruntime as ort |
| import numpy as np |
| from PIL import Image |
|
|
| CHAR_CLASSES = list("0123456789ABCDEFGHIJKLMNOPQRSTUVWXYZ") + ["blank"] |
| |
| path = hf_hub_download( |
| repo_id="hermitkk/alphabet-classifier", |
| filename="outputs/exports/alphanumeric_model.onnx", |
| ) |
| session = ort.InferenceSession(path) |
| |
| # Preprocess a 96x96 grayscale crop |
| img = Image.open("character.png").convert("L").resize((96, 96)) |
| x = (np.array(img, dtype=np.float32) / 255.0 - 0.5) / 0.5 |
| x = x[np.newaxis, np.newaxis, :, :] # (1, 1, 96, 96) |
|
|
| logits = session.run(None, {"input": x})[0] |
| pred = int(np.argmax(logits)) |
| print(CHAR_CLASSES[pred]) # e.g. "A", "3", "blank" |
| ``` |
| |
| ### Alphabet model |
| |
| ```python |
| from huggingface_hub import hf_hub_download |
| import onnxruntime as ort |
| import numpy as np |
| from PIL import Image |
|
|
| ALPHA_CLASSES = list("ABCDEFGHIJKLMNOPQRSTUVWXYZ") + ["blank"] |
| |
| path = hf_hub_download( |
| repo_id="hermitkk/alphabet-classifier", |
| filename="outputs/exports/alphabet_model.onnx", |
| ) |
| session = ort.InferenceSession(path) |
| |
| # Preprocess a 96x96 grayscale crop |
| img = Image.open("letter.png").convert("L").resize((96, 96)) |
| x = (np.array(img, dtype=np.float32) / 255.0 - 0.5) / 0.5 |
| x = x[np.newaxis, np.newaxis, :, :] # (1, 1, 96, 96) |
|
|
| logits = session.run(None, {"input": x})[0] |
| pred = int(np.argmax(logits)) |
| print(ALPHA_CLASSES[pred]) # e.g. "A", "blank" |
| ``` |
| |
| > **Note:** both models accept 96 × 96 single-channel float32 input, normalized to mean 0.5 / std 0.5. White pixels (blank paper) map to +1.0 and dark ink maps toward −1.0. |
| |
| --- |
| |
| ## Reproduce Training |
| |
| ```bash |
| git clone https://huggingface.co/hermitkk/alphabet-classifier |
| cd alphabet-classifier |
| python -m venv .venv && source .venv/bin/activate |
| pip install -r requirements.txt |
| |
| # Train alphanumeric model |
| python main.py train --config config/config.yaml |
| |
| # Train alphabet model |
| python main.py train --config config/config_alphabet.yaml |
| ``` |
| |
| --- |
| |
| ## License |
| |
| MIT |
| |