File size: 14,853 Bytes
7b2177e | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 | ---
language:
- en
- as
- bn
- brx
- doi
- gu
- hi
- kn
- ks
- kok
- mai
- ml
- mni
- mr
- ne
- or
- pa
- sa
- sat
- sd
- ta
- te
- ur
pipeline_tag: image-to-text
tags:
- ocr
- document-parsing
- layout-analysis
- reading-order
- indic
- vision-language-model
- qwen
- rt-detr
---
<div align="center">
[](#the-two-models)
[](#the-two-models)
[](#the-two-models)
[](#supported-languages)
[](#license)
</div>
**Document parsing for English and 22 Indian languages, printed and handwritten.** A page image
in; reading-ordered Markdown out, with math as LaTeX and tables as HTML or Markdown, plus
per-block JSON.
<div align="center">
<img src="assets/diagram.png" alt="IndicDocParser: page image to layout detection with reading order, then block-level OCR, then Markdown" width="100%">
</div>
IndicDocParser reads a document page and returns its text in reading order. It is a modular,
two-stage parser: **IndicDocLayout** detects the blocks on the page and orders them, and
**IndicBlockOCR** transcribes the textual blocks. The two stages communicate through a structured
JSON file, so either stage can be used independently or replaced with another implementation.
[`ARCHITECTURE.md`](ARCHITECTURE.md) traces one page through the whole call path, names what
each module does, and lists the invariants that break the output silently when violated.
---
<h2 id="examples" style="color:#F97316;">Examples</h2>
Detected blocks with their reading order on the left, the transcription on the right.
<div align="center">
<img src="assets/gallery-1-english-math-ramanujan.png" alt="A page from Ramanujan's notebooks: text and display equations detected in reading order, with the transcription rendering the mathematics as LaTeX" width="100%">
<p><b>Example #1. English page with dense mathematics.</b></p>
</div>
<div align="center">
<img src="assets/gallery-2-telugu-novel.png" alt="A printed Telugu novel page: paragraph blocks and a page number detected and numbered in reading order, with the Telugu transcription beside it" width="100%">
<p><b>Example #2. Printed Telugu page.</b></p>
</div>
<div align="center">
<img src="assets/cand-hindi-maths-g10-6pr6eq-7dcd5d95-p20.png" alt="A handwritten Hindi maths exercise on ruled paper: alternating Equation and Paragraph blocks detected in reading order, with the transcription rendering the algebra as LaTeX" width="100%">
<p><b>Example #3. Handwritten Hindi maths.</b></p>
</div>
---
<h2 id="model-summary" style="color:#F97316;">Model Summary</h2>
| | IndicDocLayout | IndicBlockOCR |
| --- | --- | --- |
| **Role** | Layout detection + reading order | Block-level text recognition |
| **Architecture** | PP-DocLayoutV3 / RT-DETR | Qwen3.5-0.8B |
| **Parameters** | 33 M | 0.8 B |
| **Precision** | fp32 | bf16 |
| **In this repo** | `weights/layout` (133 MB) | `weights/ocr` (1.7 GB) |
| **Output** | Layout JSON | Markdown + block JSON |
IndicBlockOCR uses the **Sarvam-30B tokenizer**, with a vocabulary designed to cover Indian
scripts. IndicDocLayout is a fine-tune of PP-DocLayoutV3/RT-DETR, trained with a 37-class
taxonomy designed for education-domain documents.
IndicDocLayout predicts a labelled bounding box for each detected layout element.
The 37 supported labels are:
> Advertisement, Answer, Author, Chapter-end-section, Chapter-title, Chart, Code, Contact-info, Dateline, Diagram, Equation, Expression, Flag, Folio, Footer, Footnote, Header, Image, Image-caption, Index, Infobox, List, MCQ, Page-number, Paragraph, Placeholder-text, Question, Reference, Section-title, Solved-example, Sub-section-title, Sub-sub-section-title, Table, Table-caption, Table-of-contents, Title, Website-link
---
<h2 id="supported-languages" style="color:#F97316;">Supported languages</h2>
**Printed** page recognition is supported across English and the 22 constitutionally recognised Indian languages: Assamese, Bengali, Bodo, Dogri,
Gujarati, Hindi, Kannada, Kashmiri, Konkani, Maithili, Malayalam, Manipuri, Marathi, Nepali,
Odia, Punjabi, Sanskrit, Santali, Sindhi, Tamil, Telugu, Urdu.
**Handwriting** recognition currently supports English and 12 Indian languages: Hindi, Bengali,
Telugu, Marathi, Tamil, Gujarati, Kannada, Malayalam, Odia, Punjabi, Assamese, and Urdu.
Handwriting quality is still a work in progress, particularly across different writing styles. We are working on improving recognition and extending support to additional languages.
---
<h2 id="usage" style="color:#F97316;">Usage</h2>
<h3 style="color:#F97316;">Installation</h3>
The repo ships an installer that reads your driver and picks matching CUDA wheels. If you work in a
virtual environment, please activate it first, as the installer installs into whichever
Python is active.
```bash
IDP=$(python -c "from huggingface_hub import snapshot_download as d; print(d('bodhan-ai/indic-doc-parser'))")
cd "$IDP" && ./install.sh
```
It will use `uv` if that is available, and `pip` otherwise. Where running a shell script is not
convenient, [TROUBLESHOOTING.md](TROUBLESHOOTING.md) lists the two commands it runs.
<h3 style="color:#F97316;">Basic inference</h3>
```python
import sys
from huggingface_hub import snapshot_download
repo = snapshot_download("bodhan-ai/indic-doc-parser")
sys.path.insert(0, repo) # the code ships in the repo
from indic_doc_parser import IndicDocParser
parser = IndicDocParser.from_pretrained(repo)
page = parser.parse("page.png")
print(page["markdown"]) # reading-ordered Markdown
```
`page` also carries the per-block detail, which you can save as follows:
```python
import json
with open("page.json", "w", encoding="utf-8") as f:
json.dump(page, f, ensure_ascii=False, indent=2)
```
<h3 style="color:#F97316;">Running one stage at a time</h3>
To run the two stages separately:
```python
from indic_doc_parser import IndicDocLayout, IndicBlockOCR
layout = IndicDocLayout(f"{repo}/weights/layout").detect("page.png")
page = IndicBlockOCR(f"{repo}/weights/ocr").run("page.png", layout)
```
`run()` takes a layout object, a dict, or the path to a layout JSON file.
---
<h2 id="output" style="color:#F97316;">Output</h2>
`parser.parse("page.png")` returns the page metadata and its blocks in reading order:
```json
{
"image": "sample1.png",
"width": 800,
"height": 1273,
"blocks": [
{"order": 0, "label": "Header", "type": "PageHeader",
"bbox_xyxy": [345.6, 51.7, 437.1, 114.7], "conf": 0.6, "text": ""},
{"order": 1, "label": "Page-number", "type": "PageNumber",
"bbox_xyxy": [367.8, 78.9, 413.2, 107.4], "conf": 0.747, "text": "229"},
{"order": 2, "label": "Paragraph", "type": "Text",
"bbox_xyxy": [77.9, 121.3, 711.8, 199.4], "conf": 0.863,
"text": "Thus we see that, if we can prove that twice the L.H.S. of (30) ..."}
]
}
```
| field | meaning |
| --- | --- |
| `order` | reading-order rank, 0-based and gap-free |
| `label` | the raw IndicDocLayout class (37-class taxonomy) |
| `type` | coarse pipeline category: `Text`, `Table`, `Equation`, `Title`, ... |
| `bbox_xyxy` | pixel box `[x0, y0, x1, y1]` |
| `conf` | detection confidence |
| `text` | transcription; `""` for blocks not sent to the recognizer |
**Note:** Figures, charts, advertisements, running headers, and footers are not sent through the recognizer
by default. They remain in the JSON with `text: ""`, so you can see what was detected and where.
Page numbers and other margin text such as folios are transcribed.
<h3 style="color:#F97316;">Schemas</h3>
Machine-readable JSON Schema for each envelope, in [`schemas/`](schemas):
| file | describes |
| --- | --- |
| `layout_output.schema.json` | The layout file: what **IndicDocLayout** writes and **IndicBlockOCR** reads. Blocks and reading order, before any text is read, so there is **no** `text` key at all. |
| `parse_output.schema.json` | The parsed page shown above. Every block now has `text`; `""` means the block was detected but deliberately not sent to the recognizer. |
A layout from your own detector must use a `label` from the 37-class taxonomy, or declare `type`
explicitly. An unrecognised label is rejected rather than silently read as prose.
<h3 style="color:#F97316;">Table format</h3>
Tables come back as HTML by default. Choose the format when you construct the parser:
```python
parser = IndicDocParser.from_pretrained(repo) # HTML (default)
parser = IndicDocParser.from_pretrained(repo, table_format="markdown") # Markdown
```
---
<h2 id="performance" style="color:#F97316;">Performance</h2>
<h3 style="color:#F97316;">OmniDocBench 1.6 (english subset)</h3>
| OmniDocBench 1.6 (english subset) | Overall↑ | TextEdit↓ | FormulaCDM↑ | TableTEDS↑ | TableTEDS-S↑ | Read OrderEdit↓ |
| --- | :---: | :---: | :---: | :---: | :---: | :---: |
| PaddleOCRVL-1.6 | 96.36 | 0.03 | 98.55 | 93.37 | 96.33 | 0.09 |
| Chandra OCR 2 | 93.11 | 0.04 | 96.93 | 86.07 | 90.34 | 0.09 |
| **IndicOCR (ours)** | **92.76** | **0.04** | **97.53** | **85.10** | **90.58** | **0.11** |
| GPT-5.6-sol | 92.46 | 0.04 | 95.42 | 85.87 | 90.98 | 0.10 |
| Gemini 3.1 Pro | 91.15 | 0.06 | 95.53 | 83.46 | 88.77 | 0.13 |
| Surya OCR 2 (model) | 91.13 | 0.04 | 95.67 | 81.61 | 86.37 | 0.10 |
| Sarvam Vision | 90.08 | 0.04 | 97.62 | 76.82 | 82.01 | 0.10 |
| Gemma 31B | 86.71 | 0.09 | 89.48 | 79.79 | 85.19 | 0.19 |
| Nemotron Parse 2 | 79.12 | 0.159 | 78.94 | 74.32 | 81.09 | 0.29 |
<h3 style="color:#F97316;">olmOCR-Bench (<a href="https://huggingface.co/datasets/sarvamai/olmOCR-Bench-English" style="color:#F97316;">english subset</a>)</h3>
| OlmoOCRBench ([english subset](https://huggingface.co/datasets/sarvamai/olmOCR-Bench-English)) | Overall↑ | arxiv_math↑ | baseline↑ | headers_footers↑ | long_tiny_text↑ | multi_column↑ | old_scans↑ | old_scans_math↑ | table_tests↑ |
| --- | :---: | :---: | :---: | :---: | :---: | :---: | :---: | :---: | :---: |
| Chandra OCR 2 | 85.9 | 86.7 | 99.8 | 91.5 | 93.7 | 84.7 | 51 | 88.2 | 92.2 |
| Sarvam Vision | 84.3 | 86.5 | 99.6 | 96.3 | 91 | 82.2 | 49.8 | 81 | 88.3 |
| Gemini 3.1 Pro | 82.6 | 90.5 | 99 | 82.9 | 88.5 | 81.6 | 47 | 84.3 | 87.3 |
| **IndicOCR (ours)** | **82.2** | **83.2** | **99.4** | **92.9** | **89.8** | **76** | **48.3** | **77.7** | **90** |
| Surya OCR 2 (model) | 81.4 | 82.5 | 99.8 | 92.9 | 79.9 | 85.1 | 42.8 | 84.3 | 84.2 |
| Gemma 31B | 80.4 | 79 | 99.4 | 92.9 | 89.8 | 80.5 | 45.8 | 73.8 | 82.2 |
| PaddleOCRVL-1.6 | 78.7 | 85.1 | 98.4 | 96.2 | 75.3 | 83.9 | 39 | 68.3 | 83 |
| GPT | 78 | 79.3 | 93.9 | 95.4 | 87.8 | 77.4 | 43.7 | 64.6 | 82.2 |
| Nemotron Parse 2 | 68.2 | 64 | 96.7 | 90 | 79.6 | 72.8 | 31.9 | 28.6 | 81.8 |
<h3 style="color:#F97316;">IndicOCR-PR: printed accuracy by language (higher is better)</h3>
Word-level accuracy, reported as 100 x (1 - WER).
| Language | Sarvam Vision | **IndicOCR (ours)** | Gemini 3.1 Pro | SuryaOCR | Gemma 31B | Chandra OCR 2 |
| --- | :---: | :---: | :---: | :---: | :---: | :---: |
| **Overall** | 86.6 | **86.2** | 80.4 | 67.9 | 66.3 | 64.2 |
| Assamese | 89.5 | 90.2 | 90.7 | 86.4 | 70.6 | 73.5 |
| Bodo | 91.0 | 86.5 | 91.0 | 55.6 | 68.1 | 46.6 |
| Bengali | 91.6 | 91.4 | 92.5 | 81.1 | 83.9 | 79.2 |
| Dogri | 85.8 | 81.7 | 83.7 | 60.5 | 64.4 | 55.8 |
| English | 96.6 | 97.0 | 97.7 | 93.8 | 97.2 | 91.3 |
| Gujarati | 91.6 | 91.7 | 92.8 | 79.6 | 81.6 | 73.0 |
| Hindi | 95.7 | 96.0 | 96.3 | 90.3 | 93.7 | 89.3 |
| Konkani | 93.6 | 93.7 | 93.5 | 90.5 | 76.9 | 85.5 |
| Kannada | 88.8 | 88.0 | 89.8 | 75.7 | 68.3 | 69.6 |
| Kashmiri | 43.3 | 52.2 | 38.1 | 23.4 | 19.9 | 17.6 |
| Malayalam | 90.6 | 89.9 | 90.6 | 76.5 | 72.0 | 68.3 |
| Manipuri | 81.9 | 83.8 | 0.8 | 0.1 | 0.1 | 0.0 |
| Marathi | 93.9 | 93.5 | 94.5 | 84.3 | 89.1 | 83.1 |
| Maithili | 86.7 | 83.0 | 86.7 | 67.6 | 76.3 | 66.1 |
| Nepali | 92.5 | 91.5 | 93.7 | 87.6 | 87.2 | 82.1 |
| Odia | 77.5 | 75.7 | 84.8 | 64.5 | 38.7 | 62.6 |
| Punjabi | 92.2 | 93.2 | 93.5 | 86.3 | 75.1 | 84.1 |
| Sanskrit | 82.0 | 76.2 | 83.7 | 57.8 | 60.8 | 55.8 |
| Sindhi | 89.2 | 87.1 | 86.3 | 80.5 | 74.5 | 71.4 |
| Santhali | 71.9 | 74.7 | 0.2 | 0.1 | 0.2 | 0.0 |
| Tamil | 94.2 | 91.3 | 94.4 | 79.9 | 83.3 | 79.0 |
| Telugu | 84.3 | 82.3 | 85.5 | 63.1 | 66.6 | 59.6 |
| Urdu | 87.1 | 85.9 | 88.0 | 76.4 | 76.6 | 74.4 |
<h3 style="color:#F97316;">IndicOCR-HW: handwriting accuracy by language (higher is better)</h3>
Word-level accuracy, reported as 100 x (1 - WER).
| Language | Gemini 3.1 Pro | **IndicOCR (ours)** | Sarvam Vision | Gemma 31B | Chandra OCR 2 | SuryaOCR |
| --- | :---: | :---: | :---: | :---: | :---: | :---: |
| **Overall** | 72.0 | **66.7** | 55.4 | 33.9 | 24.7 | 23.0 |
| Assamese | 71.6 | 66.1 | 47.8 | 24.1 | 8.9 | 17.8 |
| Bengali | 74.8 | 71.3 | 58.3 | 35.1 | 6.6 | 10.0 |
| English | 84.4 | 80.7 | 77.7 | 78.5 | 78.2 | 72.7 |
| Gujarati | 60.0 | 55.9 | 39.2 | 23.7 | 11.8 | 11.5 |
| Hindi | 83.1 | 77.6 | 72.3 | 70.7 | 54.6 | 42.7 |
| Kannada | 73.8 | 69.6 | 57.7 | 17.2 | 11.5 | 13.2 |
| Malayalam | 63.9 | 60.5 | 45.6 | 16.0 | 15.7 | 11.8 |
| Marathi | 79.0 | 70.2 | 61.8 | 56.5 | 35.4 | 28.8 |
| Odia | 66.7 | 68.2 | 40.6 | 15.5 | 19.4 | 19.9 |
| Punjabi | 70.1 | 69.4 | 54.8 | 11.7 | 11.5 | 15.7 |
| Tamil | 80.5 | 76.8 | 60.5 | 33.4 | 18.8 | 16.8 |
| Telugu | 72.0 | 53.5 | 59.1 | 32.0 | 20.8 | 14.6 |
| Urdu | 54.4 | 46.4 | 44.4 | 25.6 | 27.6 | 22.6 |
---
<h2 id="limitations" style="color:#F97316;">Limitations</h2>
Reading order remains a challenge for **complex, multi-column layouts**. Handwriting recognition
is also still being improved, particularly across different writing styles and writing
characteristics.
We are also extending handwriting support to additional Indic languages.
---
<h2 id="hardware" style="color:#F97316;">Hardware</h2>
Latency and throughput numbers to follow.
---
<h2 id="license" style="color:#F97316;">License</h2>
Released under [Bodhan Open License 1.0]().
The release incorporates components distributed under Apache 2.0, including PP-DocLayoutV3,
Qwen3.5, and the Sarvam-30B tokenizer. See the repository license and the corresponding upstream
licenses for the applicable terms and attribution requirements.
---
<h2 id="citation" style="color:#F97316;">Citation</h2>
```bibtex
@misc{indicdocparser2026,
title = {IndicDocParser: Multilingual Document Parsing for English and 22 Indian Languages},
author = {Bodhan.AI},
year = {2026},
url = {https://huggingface.co/bodhan-ai/indic-doc-parser}
}
``` |