--- library_name: transformers pipeline_tag: image-to-text base_model: microsoft/trocr-base-printed tags: - trocr - vision-encoder-decoder - image-to-text - ocr - captcha - math-captcha - synthetic-data --- # OCR Match Captcha A TrOCR model fine-tuned to recognize short mathematical captcha expressions such as `26+7=?` and `45-6=?`.
Handwritten captcha Clean captcha
Distorted captcha Math captcha
## Model details - **Architecture:** Vision Encoder-Decoder / TrOCR - **Base model:** `microsoft/trocr-base-printed` - **Task:** Image-to-text OCR - **Target geometry:** 130x30 RGB images - **Output:** Mathematical expressions without spaces ## Training configuration | Parameter | Value | |---|---:| | Epochs | 5 | | Learning rate | 5e-6 | | Batch size | 8 | | Gradient accumulation | 2 | | Operator-token weight | 3.0 | | Generation beams | 4 | | Maximum output length | 32 | ## Usage ```python from PIL import Image from transformers import TrOCRProcessor, VisionEncoderDecoderModel repo_id = "arkhabbazan/ocr-match-captcha" processor = TrOCRProcessor.from_pretrained(repo_id, token=True) model = VisionEncoderDecoderModel.from_pretrained(repo_id, token=True) image = Image.open("captcha.jpeg").convert("RGB") pixel_values = processor(images=image, return_tensors="pt").pixel_values generated_ids = model.generate(pixel_values, num_beams=4, max_length=32) prediction = processor.batch_decode( generated_ids, skip_special_tokens=True )[0] print(prediction.replace(" ", "")) ``` ## Limitations - Training data is primarily synthetic. - Unseen fonts and layouts may reduce accuracy. - The model is intended for short expressions, not document OCR. - Usage must be authorized and compliant with applicable policies.