ocr-math-captcha / README.md
arkhabbazan's picture
Publish model card with embedded 2x2 captcha gallery
38970db verified
|
Raw
History Blame Contribute Delete
2.33 kB
metadata
library_name: transformers
pipeline_tag: image-to-text
base_model: microsoft/trocr-base-printed
tags:
  - trocr
  - vision-encoder-decoder
  - image-to-text
  - ocr
  - captcha
  - math-captcha
  - synthetic-data

OCR Match Captcha

A TrOCR model fine-tuned to recognize short mathematical captcha expressions such as 26+7=? and 45-6=?.

Handwritten captcha Clean captcha
Distorted captcha Math captcha

Model details

  • Architecture: Vision Encoder-Decoder / TrOCR
  • Base model: microsoft/trocr-base-printed
  • Task: Image-to-text OCR
  • Target geometry: 130x30 RGB images
  • Output: Mathematical expressions without spaces

Training configuration

Parameter Value
Epochs 5
Learning rate 5e-6
Batch size 8
Gradient accumulation 2
Operator-token weight 3.0
Generation beams 4
Maximum output length 32

Usage

from PIL import Image
from transformers import TrOCRProcessor, VisionEncoderDecoderModel

repo_id = "arkhabbazan/ocr-match-captcha"
processor = TrOCRProcessor.from_pretrained(repo_id, token=True)
model = VisionEncoderDecoderModel.from_pretrained(repo_id, token=True)

image = Image.open("captcha.jpeg").convert("RGB")
pixel_values = processor(images=image, return_tensors="pt").pixel_values
generated_ids = model.generate(pixel_values, num_beams=4, max_length=32)
prediction = processor.batch_decode(
    generated_ids, skip_special_tokens=True
)[0]
print(prediction.replace(" ", ""))

Limitations

  • Training data is primarily synthetic.
  • Unseen fonts and layouts may reduce accuracy.
  • The model is intended for short expressions, not document OCR.
  • Usage must be authorized and compliant with applicable policies.