OCR: Documents
Roughly ordered by recent releases, useful updates and current usage. Practical OCR and document parsing models. Reviewed September 2026.
Image-Text-to-Text • 0.9B • Updated • 134k • 450Note 0.8B model; author-reported OmniDocBench v1.6 score: 96.58. Page-to-Markdown parser with tables and formulas; use its parsing wrapper and preserve image crops for visual regions.
StarDoc-AI/NaviDC-OCR
Image-Text-to-Text • 1B • Updated • 15.8k • 28Note Aug 2026 release. Digital and camera-captured documents, including warped pages; complete parsing needs the linked project’s pipeline. Usage: https://github.com/caipeng328/NaviDC-OCR
nvidia/NVIDIA-Nemotron-Parse-2.0
Image-Text-to-Text • 0.9B • Updated • 28.1k • 110Note Aug 2026 release. Text, bounding boxes, reading order and chart-to-table parsing; use the supplied output converters and check the component licences.
tencent/HunyuanOCR
Image-Text-to-Text • 1B • Updated • 657k • 816Note Version 1.5 adds DFlash acceleration and a llama.cpp deployment path. Document parsing, text spotting and extraction; the repository root hosts v1.5 under Tencent’s Hunyuan community licence. Usage: https://github.com/Tencent-Hunyuan/HunyuanOCR
baidu/Unlimited-OCR
Image-Text-to-Text • 3B • Updated • 2.7M • 4.2kNote Jun 2026 release. Document parsing for long outputs; follow the prescribed generation settings and postprocessing. Usage: https://github.com/baidu/Unlimited-OCR
PaddlePaddle/PaddleOCR-VL-1.6
Image-Text-to-Text • 1.0B • Updated • 33.2k • 444Note May 2026 release; author-reported OmniDocBench v1.6 score: 96.33. Multilingual document parsing with text, tables, formulas and charts; whole pages need the PaddleOCR pipeline. Usage: https://github.com/PaddlePaddle/PaddleOCR
zenosai/MonkeyOCRv2-S-Parsing
Image-Text-to-Text • 0.8B • Updated • 4.34k • 4Note Jul 2026 release. The smaller MonkeyOCR v2 parsing checkpoint; use its OCR pipeline, and distinguish its results from the larger B variant. Usage: https://github.com/Yuliang-Liu/MonkeyOCRv2
opendatalab/MinerU2.5-Pro-2605-1.2B
Image-Text-to-Text • 1B • Updated • 88.9k • 71Note Structured JSON and Markdown from documents; use the MinerU pipeline for layout analysis and output conversion. Usage: https://github.com/opendatalab/MinerU
datalab-to/surya-ocr-2
Image-Text-to-Text • 0.7B • Updated • 1.27M • 103Note Document OCR, layout and tables through the Surya package; check backend requirements and the modified model licence. Usage: https://github.com/datalab-to/surya
infly/Infinity-Parser2-Pro
Image-Text-to-Text • 35B • Updated • 268k • 91Note English/Chinese documents to Markdown or structured layout; the authors report weaker support for other languages and complex orientations.
datalab-to/chandra-ocr-2
Image-Text-to-Text • 5B • Updated • 2.9M • 491Note Documents and handwriting to Markdown, HTML or JSON; its modified OpenRAIL-M licence includes commercial-use conditions. Usage: https://github.com/datalab-to/chandra
dots-studio/dots.mocr
Image-Text-to-Text • 3B • Updated • 583k • 169Note Multilingual text recognition and document parsing; choose the transcription prompt rather than a graphics-generation task. Usage: https://github.com/rednote-hilab/dots.mocr
zai-org/GLM-OCR
Image-Text-to-Text • 1B • Updated • 2.01M • • 2.02kNote Compact Chinese/English reader for text, tables and formulas; use the SDK’s layout and recognition pipeline for whole pages. Usage: https://github.com/zai-org/GLM-OCR Batch recipe for HF Jobs: https://huggingface.co/datasets/uv-scripts/ocr/blob/main/glm-ocr.py
deepseek-ai/DeepSeek-OCR-2
Image-Text-to-Text • 3B • Updated • 974k • 1.09kNote Document and free-text OCR with native and vLLM examples; prompt choice determines the output format. Usage: https://github.com/deepseek-ai/DeepSeek-OCR-2
lightonai/LightOnOCR-2-1B
Image-Text-to-Text • 1B • Updated • 255k • 800Note Compact 1B reader for scanned pages, articles and equations; follow its image preprocessing instructions. Batch recipe for HF Jobs: https://huggingface.co/datasets/uv-scripts/ocr/blob/main/lighton-ocr2.py
allenai/olmOCR-2-7B-1025-FP8
Image-Text-to-Text • 8B • Updated • 231k • 255Note The recommended FP8 inference checkpoint for olmOCR; use its toolkit for page rendering, rotation and retries. Usage: https://github.com/allenai/olmocr
ibm-granite/granite-docling-258M
Image-Text-to-Text • 0.3B • Updated • 210k • 1.26kNote Small document converter with Docling postprocessing; English is primary, with experimental Japanese, Arabic and Chinese support. Usage: https://github.com/docling-project/docling