Files changed (1) hide show
  1. README.md +23 -4
README.md CHANGED
@@ -23,13 +23,29 @@ tags:
23
  - forms
24
  ---
25
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
26
  <div align="center">
27
  <img src="lightonocr-banner.png" alt="LightOnOCR-2-1B Banner" width="600"/>
28
  </div>
29
 
30
- # LightOnOCR-2-1B
31
 
32
- **Best OCR model .** LightOnOCR-2-1B is our flagship OCR model, refined with RLVR training for maximum accuracy. We recommend this variant for most OCR tasks.
 
33
 
34
  ## About LightOnOCR-2
35
 
@@ -45,7 +61,7 @@ LightOnOCR-2 is an efficient end-to-end 1B-parameter vision-language model for c
45
 
46
  ---
47
 
48
- 📄 **[Paper]( https://arxiv.org/pdf/2601.14251)** | 📝 **[Blog Post](https://huggingface.co/blog/lightonai/lightonocr-2)** | 🚀 **[Demo](https://huggingface.co/spaces/lightonai/LightOnOCR-2-1B-Demo)** | 📊 **[Dataset](https://huggingface.co/datasets/lightonai/LightOnOCR-mix-0126)** | 📊 **[BBox Dataset](https://huggingface.co/datasets/lightonai/LightOnOCR-bbox-mix-0126)** | 📓 **[Finetuning Notebook](https://colab.research.google.com/drive/1WjbsFJZ4vOAAlKtcCauFLn_evo5UBRNa?usp=sharing)**
49
 
50
  ---
51
 
@@ -197,6 +213,9 @@ Apache License 2.0
197
  title = {LightOnOCR: A 1B End-to-End Multilingual Vision-Language Model for State-of-the-Art OCR},
198
  author = {Said Taghadouini and Adrien Cavaill\`{e}s and Baptiste Aubertin},
199
  year = {2026},
200
- howpublished = {\url{https://arxiv.org/pdf/2601.14251}}
201
  }
 
 
 
202
  ```
 
23
  - forms
24
  ---
25
 
26
+ # LightOnOCR-2-1B
27
+ [![Speed](https://img.shields.io/badge/Speed-5.71%20pages%2Fs-brightgreen)](https://huggingface.co/lightonai/LightOnOCR-2-1B)
28
+ [![Model Size](https://img.shields.io/badge/Model%20Size-1B%20params-orange)](https://huggingface.co/lightonai/LightOnOCR-2-1B)
29
+ [![Demo](https://img.shields.io/badge/🚀%20Demo-Spaces-orange)](https://huggingface.co/spaces/lightonai/LightOnOCR-2-1B-Demo)
30
+ [![License](https://img.shields.io/badge/License-Apache%202.0-green.svg)](https://opensource.org/licenses/Apache-2.0)
31
+ [![SOTA](https://img.shields.io/badge/OlmOCR--Bench-SOTA-gold)](https://huggingface.co/lightonai/LightOnOCR-2-1B)
32
+ [![Transformers](https://img.shields.io/badge/🤗%20Transformers-supported-yellow)](https://github.com/huggingface/transformers)
33
+ [![vLLM](https://img.shields.io/badge/vLLM-supported-purple)](https://github.com/vllm-project/vllm)
34
+ [![Blog](https://img.shields.io/badge/📝%20Blog-HuggingFace-yellow)](https://huggingface.co/blog/lightonai/lightonocr-2)
35
+ [![Dataset](https://img.shields.io/badge/📊%20Dataset-LightOnOCR--mix-blue)](https://huggingface.co/datasets/lightonai/LightOnOCR-mix-0126)
36
+ [![BBox Dataset](https://img.shields.io/badge/📊%20BBox%20Dataset-LightOnOCR--bbox--mix-blue)](https://huggingface.co/datasets/lightonai/LightOnOCR-bbox-mix-0126)
37
+ [![Colab](https://img.shields.io/badge/Finetuning-Notebook-F9AB00?logo=googlecolab)](https://colab.research.google.com/drive/1WjbsFJZ4vOAAlKtcCauFLn_evo5UBRNa?usp=sharing)
38
+ [![arXiv](https://img.shields.io/badge/arXiv-2412.13663-b31b1b.svg)](https://arxiv.org/abs/2412.13663)
39
+ [![Website](https://img.shields.io/badge/LightOn-Website-blue?logo=google-chrome)](https://lighton.ai)
40
+ [![LinkedIn](https://img.shields.io/badge/LightOn-LinkedIn-0A66C2?logo=linkedin)](https://www.linkedin.com/company/lighton/)
41
+ [![X](https://img.shields.io/badge/@LightOnIO-X-black?logo=x)](https://x.com/LightOnIO)
42
  <div align="center">
43
  <img src="lightonocr-banner.png" alt="LightOnOCR-2-1B Banner" width="600"/>
44
  </div>
45
 
 
46
 
47
+ # LightOnOCR-2-1B
48
+ **Best OCR model .** LightOnOCR-2-1B is **[LightOn's](https://lighton.ai)** flagship OCR model, refined with RLVR training for maximum accuracy. We recommend this variant for most OCR tasks.
49
 
50
  ## About LightOnOCR-2
51
 
 
61
 
62
  ---
63
 
64
+ 📄 **[Paper]( https://arxiv.org/pdf/2601.14251)** | 📝 **[Blog Post](https://huggingface.co/blog/lightonai/lightonocr-2)** | 🚀 **[Demo](https://huggingface.co/spaces/lightonai/LightOnOCR-2-1B-Demo)** | 📊 **[Dataset](https://huggingface.co/datasets/lightonai/LightOnOCR-mix-0126)** | 📊 **[BBox Dataset](https://huggingface.co/datasets/lightonai/LightOnOCR-bbox-mix-0126)** | 📓 **[Finetuning Notebook](https://colab.research.google.com/drive/1WjbsFJZ4vOAAlKtcCauFLn_evo5UBRNa?usp=sharing)** | **[LightOn blog entry](https://www.lighton.ai/lighton-blogs/lighton-opens-a-new-field-for-ai-with-lightonocr-2-document-intelligence)**
65
 
66
  ---
67
 
 
213
  title = {LightOnOCR: A 1B End-to-End Multilingual Vision-Language Model for State-of-the-Art OCR},
214
  author = {Said Taghadouini and Adrien Cavaill\`{e}s and Baptiste Aubertin},
215
  year = {2026},
216
+ howpublished = {\url{https://arxiv.org/abs/2601.14251}}
217
  }
218
+
219
+ [![Downloads](https://img.shields.io/badge/dynamic/json?url=https://huggingface.co/api/models/lightonai/LightOnOCR-2-1B&query=downloads&label=Downloads&color=blue)](https://huggingface.co/lightonai/LightOnOCR-2-1B)
220
+ [![EU](https://img.shields.io/badge/🇪🇺%20Made%20in-Europe-blue)](https://huggingface.co/lightonai)
221
  ```