EasyOCR: Optimized for AMD ROCm

EasyOCR is a ready-to-use optical character recognition (OCR) pipeline combining a CRAFT text-detection model with a CRNN text-recognition model, supporting 80+ languages. This repository packages inference for OCR (text detection + recognition) using ONNX Runtime, exported and validated for AMD ROCm so it runs efficiently on AMD GPUs, CPUs, and NPUs.

This is based on the implementation of EasyOCR found here. This repository contains configurations and scripts optimized for AMD® ROCm™ platforms. You can use the easyocr AMD scripts to reproduce results or export with custom configurations. More details on model performance can be found here.


Task Overview

Task: Optical character recognition (text detection + recognition)

Dataset: COCO-Text v2 subset (813 real-world images / 8,377 ground-truth text regions, sampled from COCO2014 images with ≥5 legible text regions each)

Output metrics: Word Accuracy, Character Similarity, Detection Precision/Recall/F1, FPS

Backend note: Recognition always runs on CPU regardless of target device — its ~3ms latency is negligible next to detection, so the cross-device switch overhead isn't worth it — while detection runs on the target device's execution provider.


AMD ROCm Optimization

This model export has been adapted and validated for AMD Instinct™ / Radeon™ GPUs running ROCm, as well as AMD CPUs and AMD Ryzen AI NPUs. Key points:

  • Validated backends: ONNX Runtime across CPU, GPU (MIGraphX execution provider), and NPU (VitisAI execution provider, auto-quantized internally).
  • Both detection (CRAFT) and recognition (CRNN) models are exported to ONNX with static shapes for cross-device compatibility.
  • No code changes required versus the upstream EasyOCR implementation — only environment/runtime configuration differs.
Runtime Precision Backend Hardware Notes
ONNX Runtime FP32 / FP16 / BF16 / INT8 CPU Execution Provider AMD CPU Recognition always runs here
ONNX Runtime FP32 / FP16 / BF16 / INT8 MIGraphX Execution Provider AMD Instinct™ / Radeon™ GPU (ROCm) Detection only
ONNX Runtime Auto VitisAI Execution Provider AMD Ryzen AI NPU Detection only; auto-quantized internally

Getting Started

For setup instructions, evaluation scripts, and custom configuration options, see the easyocr on GitHub.


Model Details

Model Type: Optical character recognition pipeline (text detection + text recognition)

Base Model: CRAFT (detection) + CRNN (recognition), from EasyOCR

Model Stats:

  • Detection input: (1, 3, 640, 640) float32; outputs region/link scores (1, 320, 320, 2) and feature map (1, 32, 320, 320)
  • Recognition input: (1, 1, 64, 200) float32; output CTC logits (1, 49, 97)
  • Precision tested: FP32, FP16, BF16, INT8 (CPU/GPU); auto-quantized (NPU)

Accuracy Pipeline

Higher values mean predictions agree more closely with ground truth across all metrics below — 1.0 is perfect, 0.0 is no correct matches. Word Accuracy is intentionally low in absolute terms: COCO-Text's images are real-world "text in the wild" photos (street signs, product labels, graffiti) with small, rotated, and low-contrast text, unlike clean scanned documents.

Metrics Explained

Metric Description
Word Accuracy Fraction of IoU-matched detections whose recognized text exactly matches (case-insensitive) the ground-truth string. The strictest text metric — a single wrong character fails the match.
Character Similarity Mean character-level similarity ratio between recognized and ground-truth text across matched detections. More forgiving than Word Accuracy.
Detection Precision Of all text regions the model detected, what fraction matched a real ground-truth region (IoU ≥ 0.5).
Detection Recall Of all ground-truth text regions, what fraction the model detected.
Detection F1 Harmonic mean of Detection Precision and Recall.

Accuracy Results

Full Dataset Evaluation (COCO-Text v2, 813 images, 8,377 ground-truth text regions) — filled from runs/quality_<device>.json; run make metrics to refresh:

Device Precision Word Accuracy Char Similarity Det. Precision Det. Recall Det. F1
CPU FP32 0.2191 0.3642 0.6671 0.4468 0.5352
GPU FP32 0.2191 0.3642 0.6671 0.4468 0.5352
NPU FP32 0.2183 0.3669 0.6611 0.4529 0.5375

Dig Deeper

Want to explore the full evaluation scripts, config options, and other AMD-optimized model examples?

📂 View the full project on GitHub

The GitHub repository includes:

  • Setup and prerequisites for ROCm environments
  • Scripts for CPU, GPU, and NPU runners
  • IoU-matching accuracy evaluation pipeline against COCO-Text v2
  • Benchmarking and reproduction instructions
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support