Sansar OCR 2B

Reads printed Sanskrit book pages (Devanagari) and writes out their text, line by line. A 2.2B-parameter vision-language model: Qwen/Qwen3.5-2B fine-tuned on 19,377 Sanskrit page images with their text. From the Sansar project at Muse Mesh (Hugging Face, Sansar Lab): Sanskrit language models, their corpus, their tokenizer, and the OCR that feeds the corpus.

  • This version: v0.1.0 = training run ocrvlm_v12_clean_q35, finished 2026-10-08.
  • Accuracy: median letter error rate 1.05% on 54 held-out pages from 53 books (transformers, the files in this repository), 1.06% with vLLM. Google Cloud Vision makes 1.69% on the same pages and the archive.org OCR of these scans 9.03%. Better than Vision on 38 of 54 pages (39 of 54 with vLLM).
  • Training labels: clean e-texts aligned to scans, the text layers of born-digital textbooks and news bulletins, and synthetic pages. No output of a commercial AI service (Google Cloud Vision, Anthropic Claude, OpenAI models, Sarvam) was used as a training label; those services appear on this card only as comparison rows in the evaluation. See Training data.

Quick start

pip install "transformers>=5.18" torch torchvision pillow numpy accelerate
from huggingface_hub import hf_hub_download
import importlib.util

# ocr_page.py (in this repository) prepares pages exactly as in training and runs greedy decoding
path = hf_hub_download("MuseMesh/sansar-ocr-2b", "ocr_page.py", revision="v0.1.0")
spec = importlib.util.spec_from_file_location("ocr_page", path); ocr = importlib.util.module_from_spec(spec)
spec.loader.exec_module(ocr)

model, processor = ocr.load("MuseMesh/sansar-ocr-2b", revision="v0.1.0")      # bfloat16 on a GPU, float32 on CPU
texts = ocr.transcribe(["page_001.jpg", "page_002.jpg"], model, processor, batch=4)
print(texts[0])

Or from the command line: python ocr_page.py --out_dir out/ pages/*.jpg (one .txt per page).

What ocr_page.py does, if you want to call transformers directly:

  1. Page image. Crop near-white outer margins (a dark scanner border stays), then resize with the aspect ratio kept so that width x height <= 1,800,000 pixels and both sides are multiples of 32. The model was trained on exactly this; much larger or smaller pages are outside what it has seen. One page per request.

  2. Prompt. The chat template with one image and this text, thinking off (enable_thinking=False):

    Transcribe all the text printed on this page exactly as printed, in reading order, line by line. Output only the text.
    
  3. Decoding. Greedy (do_sample=False), up to 3,072 new tokens, stop at <|im_end|> (248046) or <|endoftext|> (248044). generation_config.json sets all of this. Batching needs padding_side="left".

import torch
from transformers import AutoModelForImageTextToText, AutoProcessor

model = AutoModelForImageTextToText.from_pretrained("MuseMesh/sansar-ocr-2b", dtype=torch.bfloat16).to("cuda").eval()
processor = AutoProcessor.from_pretrained("MuseMesh/sansar-ocr-2b")
msgs = [{"role": "user", "content": [{"type": "image"}, {"type": "text", "text": ocr.PROMPT}]}]
prompt = processor.apply_chat_template(msgs, tokenize=False, add_generation_prompt=True, enable_thinking=False)
image = ocr.prepare_page("page_001.jpg")                      # margins trimmed, <= 1.8 MP, sides multiple of 32
x = processor(text=[prompt], images=[image], return_tensors="pt").to("cuda")
out = model.generate(**x, max_new_tokens=3072, do_sample=False)
print(processor.tokenizer.decode(out[0, x["input_ids"].shape[1]:], skip_special_tokens=True))

vLLM (much faster for many pages): load the repository with LLM(model="MuseMesh/sansar-ocr-2b", limit_mm_per_prompt={"image": 1}, max_model_len=5672, dtype="bfloat16"), send {"prompt": prompt, "multi_modal_data": {"image": image}} with the same prompt and page preparation, and SamplingParams(temperature=0.0, max_tokens=3072, stop_token_ids=[248046, 248044]). The vLLM numbers on this card were measured on the training run's checkpoint with vLLM 0.30; the release files hold the same weights (see Files).

What the output looks like

Plain text, one printed line per output line, Devanagari as printed (NFC, no ZWJ/ZWNJ), । and ॥ as dandas, digits as printed (Devanagari or Latin). On book pages it tends to leave out running heads, page numbers and footnotes, because most of its training pages had those painted out (see Training data); do not rely on either behaviour. Latin-script text (English introductions, IAST) was rare in training and is not evaluated.

Evaluation

Bench. bench_v1_pages: 54 pages from 53 archive.org scanned books (mixed print and scan quality), each paired with the same passage from a clean digital edition of the text, cut to what is printed on that page and checked page by page. No page of the bench books (nor any page sharing three 16-letter windows with a bench reference) is in the training or validation data, and the bench was never used to pick a checkpoint.

Metric. Letter error rate: edit distance on Devanagari letters and signs only (consonants, vowels, vowel signs, virama, anusvara, visarga, avagraha; no digits, dandas, Vedic accent marks, punctuation or spaces) divided by the reference's letter count, with the reference located inside the output (text the reference does not cover, such as a running head the model wrote out, is not counted against it). Median over the 54 pages, so one ruined page does not decide the score; the mean is shown too. The same scorer and pages for every engine.

engine median letter CER mean bench tier "kept" (39) tier "mid" (15) hardest third of scans (18) note
Claude Opus 5.5 (Anthropic) 0.65% 1.90% 0.63% 1.08% 2.11% general LLM, page by page; comparison only
Sansar OCR 2B v0.1.0 (transformers) 1.05% 2.31% 0.78% 1.31% 3.14% this repository's files and ocr_page.py, re-run for the release on an NVIDIA GeForce RTX 3060
Sansar OCR 2B v0.1.0 (vLLM) 1.06% 2.27% 0.81% 1.40% 3.12% the training run's own evaluation, A100
Claude Sonnet 5.5 (Anthropic) 1.14% 2.17% 1.10% 1.31% 3.12% comparison only
Google Cloud Vision 1.69% 3.89% 1.50% 3.15% 3.55% DOCUMENT_TEXT_DETECTION, 2026-10; comparison only
Codex gpt-6-astra (OpenAI) 2.65% 3.92% 2.18% 3.55% 6.66% comparison only
archive.org OCR of the same scans 9.03% 9.64% the text most of these books are online with today
Qwen3.5-2B, no fine-tuning 18.54% 28.89% 17.04% 21.88% 27.19% the base model, same prompt

The Vision, Claude and Codex rows are those services' transcriptions of the same 54 page images (2026-09/10), scored by the same scorer, shown for comparison only; none of their output was used to train this model. "kept"/"mid" are the bench's own page-quality tiers; "hardest third" are the 18 pages where the archive.org OCR did worst (a proxy for scan quality).

  • The release check. The model in this repository (repackaged: see Files) was re-run on all 54 pages with transformers and ocr_page.py for this release: median 1.05%, mean 2.31% (NVIDIA GeForce RTX 3060, batch 4). The training run's own transformers cross-check of the same weights gave 1.06%; its vLLM run 1.06%. Different engines and GPUs change a few characters on some pages (bfloat16 kernels). Everything is in eval/.
  • Small bench. With 54 pages a median difference of a few tenths of a point has a wide interval; the page-by-page count (better than Vision on 39 of 54 pages, vLLM) is the sturdier evidence.
  • The worst pages are bad for every engine (scan or reference problems, not this model's).
  • Validation. On the 711-page validation set (31 books, held out by book; the same kinds of sources as the training data) the median is 0.25% and the mean 0.93%. It is easier than the bench and not comparable to it.
  • Overall CER including punctuation, digits and spacing is higher: median 3.50% on the bench (Vision 4.56%).

Speed

GPU engine pages per hour (generation only)
NVIDIA A100-SXM4-40GB (GCP spot) vLLM, batch 64 9,216 (54 test pages) / 15,104 (711 validation pages)
NVIDIA GeForce RTX 3060 transformers generate, batch 4 244 (54 test pages)

About 847-1,035 output tokens per page (vLLM). Start-up (model load, vLLM compile) is not included.

Training

base Qwen/Qwen3.5-2B (revision 15852e8), Apache-2.0: hybrid Gated-DeltaNet/attention language model with its vision encoder
method full fine-tune of every weight (vision encoder at a lower learning rate), loss on the transcription only
steps 2,000 x 16 pages = 32,000 page samples, 80.7M tokens seen
optimiser 8-bit AdamW (bitsandbytes), lr 1.5e-5 (vision encoder 5e-6), 3% warmup, bfloat16 autocast over float32 master weights, gradient checkpointing
pages the prompt above + page image (<= 1.8 MP) + label + `<
mixture sampling weights per source (below); the last 15% of samples drawn from the gold sources (twin, textbooks, news) only
checkpoint picked by loss on a source-balanced gold validation subset (276 pages, <= 120 per source), never on the test bench: step 2,000 (val loss 0.0621; 0.1577 at step 200)
hardware 1x NVIDIA A100-SXM4-40GB, Google Cloud spot (a2-highgpu-1g, us-central1-a)
time and cost 7.6 h of VM time for training and the evaluations (1 spot stop; training resumed from its last checkpoint), about $9.9 of spot time

training/summary.json holds every argument and the validation-loss curve; training/mixture.json the sample counts; training/label_stats.json the label-set counts and exactly what was excluded.

Training data

19,377 training pages from 1,057 books, split by book from 711 validation pages (31 books) and the 54 test pages. Page images and labels are not redistributed here: most of the scans and texts carry no open licence. Where each label comes from:

source train pages weight samples drawn label
archive.org scans + a clean e-text of the same work ("twin") 11,074 1.5 27,467 a clean digital edition of the work (Sanskrit Wikisource proofread pages, GRETIL, Muktabodha, sanskritdocuments.org, Ambuda, sanskritsahitya.org, snskrt) aligned to the printed page and cut at its printed line breaks; running heads, page numbers and footnotes are painted out of the image
born-digital textbooks (NCERT, Kerala SCERT) 3,367 0.5 2,808 the PDF's own text layer, repaired by rule (split words, doubled signs, reph order); ~1 split word in 100 remains
All India Radio Sanskrit news bulletins (2023-2026) 3,839 0.2 1,270 the PDF text decoded from its fonts (the fonts' own glyph tables, and a vocabulary-based decipherment of unmapped glyphs); no OCR; ~1-1.5% character errors expected
synthetic pages 1,097 0.3 455 paragraphs from curated open Sanskrit e-text corpora (GRETIL, DCS, SARIT, Muktabodha, DharmaNexus, BORI Mahābhārata, ...) typeset in 1-2 columns with real fonts and degraded (skew, blur, noise, bleed); exact labels

No commercial AI service output is training text. Labels come only from clean e-texts aligned to scans, the text layers of born-digital textbooks and All India Radio bulletins, and synthetic pages typeset from curated e-text corpora. Excluded on purpose: page transcriptions made with Google Cloud Vision, Anthropic Claude and OpenAI Codex (an earlier, internal label set had used them), and synthetic pages typesetting paragraphs from Sarvam's Vagartha dataset. Where those services were used at all, it was as a check, never as label text: the news-bulletin decoding was validated against Vision on a sample of pages, and the comparison rows in Evaluation are their transcriptions of the test pages. archive.org's own OCR is used only to locate the printed lines on twin pages. Wikisource texts are human-proofread transcriptions by Wikisource volunteers.

All labels are NFC with ZWJ/ZWNJ removed, | as ।, and a colon typed after a letter as visarga. Every training and validation page that shares three or more 16-letter windows with a test reference was dropped.

Limitations

  • Printed Devanagari Sanskrit only. Not trained on manuscripts, handwriting, palm leaves, other scripts or Hindi/Marathi pages (some Hindi in the textbooks and news, but it was not evaluated).
  • Reading order on two-column pages, glossaries, tables and picture captions can be wrong; on scans of two-page spreads it can read across. (Fewer such pages were in training than in our earlier internal model, which also learnt from Vision's text of glossary and table pages.)
  • Loops and runaways. On a hard page it can repeat a line until the 3,072-token cap. Check for repeated lines.
  • Page furniture (running heads, page numbers, footnotes) is sometimes left out and sometimes written out.
  • Not scored: digits, punctuation, spacing and Vedic accents are outside the letter metric; the overall error rate including them is 3.50% (median).
  • It can invent text on blank pages, figures or very faint scans, as any generative OCR can. It also inherits the base model's general behaviour on anything that is not a Sanskrit page.
  • Labels are imperfect: the twin labels follow another edition of the same work (a few percent of variant readings and verse numbers), the textbook layers split words, the news decoding has rare glyph errors. The model learnt some of that.

Versions

Each version is a git tag; main is the newest. Pin one with revision="v0.1.0".

version date training run bench median (transformers / vLLM) Google Vision
v0.1.0 2026-10-08 ocrvlm_v12_clean_q35 1.05% / 1.06% 1.69%

Files

file what
model.safetensors the weights, bfloat16 (the gated-DeltaNet decay and norm parameters in float32 as in the base release); the output head tied to the input embedding
config.json, generation_config.json architecture; greedy decoding with both stop tokens
tokenizer.json, tokenizer_config.json, chat_template.jinja, processor_config.json the base model's tokenizer and processor, unchanged
ocr_page.py page preparation, loading and batch transcription (the code above)
eval/bench_v1_scores.json per-page and per-group scores of every engine in the table
eval/verification.json, eval/verification_pages.json the release check: this repository's files re-run on the 54 pages
eval/transcriptions/ the training run's own outputs for the 54 pages (vLLM and transformers)
eval/*_scores.json, eval/*_timing.json the training run's evaluation logs (test, validation)
training/summary.json, training/mixture.json, training/cloud.json, training/label_stats.json arguments, validation-loss curve, sample counts, machine time, label-set counts
LICENSE, LICENSE-CODE, LICENSE-QWEN, NOTICE, CHANGELOG.md licences, base-model attribution, version history

Repackaging against the training checkpoint: the training run saved the tied output head as a second copy of the input embedding (bit-identical; dropped); config.json records bfloat16 and use_cache: true instead of the training-time float32/false; generation_config.json adds the <|im_end|> stop token and greedy decoding. No weight changed.

Licence

This release is for research and non-commercial use. The training labels come partly from sources without an open licence (textbooks, news bulletins, scanned books), used here for research.

  • Weights (model.safetensors): CC BY-NC 4.0 (LICENSE). Non-commercial use (research, teaching, non-profit work) with attribution to "Sansar, Muse Mesh Private Limited".
  • Base model: Qwen3.5-2B by Alibaba Cloud, Apache-2.0 (LICENSE-QWEN, NOTICE). The tokenizer and processor files are the base model's and stay under Apache-2.0.
  • Code (ocr_page.py): Apache-2.0 (LICENSE-CODE).
  • Commercial use: contact kushal@muse-mesh.com.

Citation

@misc{sansar_ocr_2b_2026,
  title  = {Sansar OCR 2B: printed Sanskrit page OCR},
  author = {Muse Mesh},
  year   = {2026},
  note   = {v0.1.0, fine-tuned from Qwen3.5-2B},
  url    = {https://huggingface.co/MuseMesh/sansar-ocr-2b}
}

Contact: kushal@muse-mesh.com

Downloads last month
4
Safetensors
Model size
2B params
Tensor type
F32
·
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for MuseMesh/sansar-ocr-2b

Finetuned
Qwen/Qwen3.5-2B
Finetuned
(475)
this model

Collection including MuseMesh/sansar-ocr-2b

Evaluation results

  • median letter error rate, transformers (this repository's files) on Sansar bench_v1 pages (54 archive.org book pages, 53 books)
    self-reported
    1.050
  • median letter error rate, vLLM (the training run's evaluation) on Sansar bench_v1 pages (54 archive.org book pages, 53 books)
    self-reported
    1.060
  • mean letter error rate, transformers on Sansar bench_v1 pages (54 archive.org book pages, 53 books)
    self-reported
    2.310