OCR: Documents Collection Roughly ordered by recent releases, useful updates and current usage. Practical OCR and document parsing models. Reviewed September 2026. • 17 items • Updated about 1 hour ago • 1
OCR: Languages & scripts Collection OCR for particular languages and writing systems, from whole-page readers to dedicated text-line recognisers. • 10 items • Updated about 1 hour ago • 1
OCR: Handwriting & archives Collection Handwriting, historical print and manuscript recognition. Notes on languages, text-line segmentation and transcription conventions. • 10 items • Updated about 1 hour ago • 1
OCR on the Hub Collection Curated OCR models for documents, languages, handwriting and text in images. Browse four collections with short practical notes. • 4 items • Updated about 1 hour ago • 1
OCR: Text recognition & pipelines Collection Text-line and region recognisers, plus detection and recognition pipelines for documents, manga and text in photographs. • 6 items • Updated about 1 hour ago • 1
DataFlow-Harness: A Grounded Code-Agent Platform for Constructing Editable LLM Data Pipelines Paper • 2607.16617 • Published Jul 18 • 145
Kraken PP-OCRv6 text recognition models Collection Hub mirrors of Benjamin Kiessling's multilingual PP-OCRv6 line-recognition family for Kraken: tiny, small, and medium. • 3 items • Updated 4 days ago • 6
view article Article Introducing DOI: the Digital Object Identifier to Datasets and Models +2 sasha, Sylvestre, christopher, aleroy • Oct 7, 2022 • 7
Arabic HTR Collection A set of datasets, models and tools for Arabic HTR produced by Calfa. • 5 items • Updated May 22 • 2
view article Article Internet-Scale Knowledge Retrieval: A Novel Vector Search Dataset at 10B Scale Qdrant • 6 days ago • 13
MobileMoE Collection a family of on-device MoE language models with sub-billion active parameters (0.3-0.9B active and 1.3-5.3B total) that establish a new Pareto frontier • 11 items • Updated 7 days ago • 26
Encyclopaedia Britannica illustrations, 1768–1929 Collection Illustrated pages across every out-of-copyright Britannica edition: the page classifier, the 115k pages it found, and the NLS labels it learned from. • 3 items • Updated 13 days ago • 2
view article Article Training and Finetuning Multi-Vector Embedding Models with Sentence Transformers tomaarsen • 13 days ago • 120
view article Article Extremely Fast and Accurate Transcription with Granite Speech 5.0 Turbo CTC ibm-granite • 13 days ago • 33
view article Article NVIDIA Cosmos-H-Dreams: Bringing Real-Time Generative Simulation to Surgical Robotics nvidia • Jul 27 • 77
EgoSuite-Open100K Collection The largest fully-annotated open egocentric human dataset. 100,000 hours across 15,000+ tasks and scenes. • 3 items • Updated 19 days ago • 52