OCR: Leaderboards
Where OCR models are actually scored. Benchmark datasets with results tables first, then per-language leaderboard Spaces. Ordered by likes.
Benchmark β’ Updated β’ 35.7k β’ 294Note Results table in the card. 1,400 PDFs with 7,000 unit-test style checks (old scans, math, tables, headers). Pass rate per model, most-cited English document OCR bench. Updated Feb 2026.
llamaindex/ParseBench
Benchmark β’ Updated β’ 169k β’ 23.7k β’ 130Note Leaderboard in the card. Document parsing of real business documents into Markdown, scored on text, tables, layout. Updated Apr 2026.
PaddlePaddle/Real5-OmniDocBench
Benchmark β’ Updated β’ 5.68k β’ 38Note Leaderboard in the card, not a Space. OmniDocBench v1.5 page parsing: text, tables, formulas, reading order across document types. Largest results table here, updated Sep 2026.
BHL OCR Leaderboard
π17OCR & VLM character error rates on 18thβ19th-c. book pages
Note Historical printed books (Biodiversity Heritage Library). CER on human-corrected IMPACT ground truth, 15 models, updated Sep 2026. Static Space, results from an open uv-script harness.
slOCR Benchmark
π11Slovenian-Language OCR Benchmark Leaderboard
Note Slovene. Gradio leaderboard over the slOCR benchmark, modern and historical print, CER/WER per model. Updated May 2026.
Icelandic OCR Leaderboard
π2OCR & VLM character error rates on Icelandic documents
Note Icelandic. 100 manually transcribed pages with ALTO and PAGE layout ground truth, so line detection is scored too. Updated Sep 2026.
Tibetan OCR Leaderboard
π1Explore Tibetan OCR model rankings on the BDRC benchmark
Note Tibetan. Line-to-text benchmark from OpenPecha, CER per model, scripts from woodblock to modern print. Updated Aug 2026.