OCR on the Hub Collection Curated OCR models for documents, languages, handwriting and text in images. Browse five collections with short practical notes. • 5 items • Updated 7 days ago • 12
view article Article NeoMME: an efficient Multimodal-native and Multilingual Encoder Hcompany • 20 days ago • 110
view article Article Training and Finetuning Multi-Vector Embedding Models with Sentence Transformers tomaarsen • 29 days ago • 144
Unified Multimodal Autoregressive Modeling with Shared Context-Visual Tokenizer is Key to Unification Paper • 2606.18249 • Published Jun 16 • 16
view article Article DenseOn with the LateOn: Open State-of-the-Art Single and Multi-Vector Models lightonai • Apr 21 • 46
view article Article **ColBERT-Zero: To Pre-train Or Not To Pre-train ColBERT models?** lightonai • Feb 19 • 23
view article Article Party is over: regularizing ColBERT models to fix efficient ANN methods lightonai • Jun 16 • 24
PaperFit: Vision-in-the-Loop Typesetting Optimization for Scientific Documents Paper • 2605.10341 • Published May 11 • 34
Efficient Training on Multiple Consumer GPUs with RoundPipe Paper • 2604.27085 • Published Apr 29 • 46
Repetition over Diversity: High-Signal Data Filtering for Sample-Efficient German Language Modeling Paper • 2604.28075 • Published Apr 30 • 18