Instructions to use MuseMesh/sansar-ocr-2b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use MuseMesh/sansar-ocr-2b with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="MuseMesh/sansar-ocr-2b") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("MuseMesh/sansar-ocr-2b") model = AutoModelForMultimodalLM.from_pretrained("MuseMesh/sansar-ocr-2b", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use MuseMesh/sansar-ocr-2b with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "MuseMesh/sansar-ocr-2b" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "MuseMesh/sansar-ocr-2b", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/MuseMesh/sansar-ocr-2b
- SGLang
How to use MuseMesh/sansar-ocr-2b with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "MuseMesh/sansar-ocr-2b" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "MuseMesh/sansar-ocr-2b", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "MuseMesh/sansar-ocr-2b" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "MuseMesh/sansar-ocr-2b", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use MuseMesh/sansar-ocr-2b with Docker Model Runner:
docker model run hf.co/MuseMesh/sansar-ocr-2b
Sansar OCR 2B
Reads printed Sanskrit book pages (Devanagari) and writes out their text, line by line. A 2.2B-parameter vision-language model: Qwen/Qwen3.5-2B fine-tuned on 19,377 Sanskrit page images with their text. From the Sansar project at Muse Mesh (Hugging Face, Sansar Lab): Sanskrit language models, their corpus, their tokenizer, and the OCR that feeds the corpus.
- This version: v0.1.0 = training run
ocrvlm_v12_clean_q35, finished 2026-10-08. - Accuracy: median letter error rate 1.05% on 54 held-out pages from 53 books (transformers, the files in this repository), 1.06% with vLLM. Google Cloud Vision makes 1.69% on the same pages and the archive.org OCR of these scans 9.03%. Better than Vision on 38 of 54 pages (39 of 54 with vLLM).
- Training labels: clean e-texts aligned to scans, the text layers of born-digital textbooks and news bulletins, and synthetic pages. No output of a commercial AI service (Google Cloud Vision, Anthropic Claude, OpenAI models, Sarvam) was used as a training label; those services appear on this card only as comparison rows in the evaluation. See Training data.
Quick start
pip install "transformers>=5.18" torch torchvision pillow numpy accelerate
from huggingface_hub import hf_hub_download
import importlib.util
# ocr_page.py (in this repository) prepares pages exactly as in training and runs greedy decoding
path = hf_hub_download("MuseMesh/sansar-ocr-2b", "ocr_page.py", revision="v0.1.0")
spec = importlib.util.spec_from_file_location("ocr_page", path); ocr = importlib.util.module_from_spec(spec)
spec.loader.exec_module(ocr)
model, processor = ocr.load("MuseMesh/sansar-ocr-2b", revision="v0.1.0") # bfloat16 on a GPU, float32 on CPU
texts = ocr.transcribe(["page_001.jpg", "page_002.jpg"], model, processor, batch=4)
print(texts[0])
Or from the command line: python ocr_page.py --out_dir out/ pages/*.jpg (one .txt per page).
What ocr_page.py does, if you want to call transformers directly:
Page image. Crop near-white outer margins (a dark scanner border stays), then resize with the aspect ratio kept so that width x height <= 1,800,000 pixels and both sides are multiples of 32. The model was trained on exactly this; much larger or smaller pages are outside what it has seen. One page per request.
Prompt. The chat template with one image and this text, thinking off (
enable_thinking=False):Transcribe all the text printed on this page exactly as printed, in reading order, line by line. Output only the text.Decoding. Greedy (
do_sample=False), up to 3,072 new tokens, stop at<|im_end|>(248046) or<|endoftext|>(248044).generation_config.jsonsets all of this. Batching needspadding_side="left".
import torch
from transformers import AutoModelForImageTextToText, AutoProcessor
model = AutoModelForImageTextToText.from_pretrained("MuseMesh/sansar-ocr-2b", dtype=torch.bfloat16).to("cuda").eval()
processor = AutoProcessor.from_pretrained("MuseMesh/sansar-ocr-2b")
msgs = [{"role": "user", "content": [{"type": "image"}, {"type": "text", "text": ocr.PROMPT}]}]
prompt = processor.apply_chat_template(msgs, tokenize=False, add_generation_prompt=True, enable_thinking=False)
image = ocr.prepare_page("page_001.jpg") # margins trimmed, <= 1.8 MP, sides multiple of 32
x = processor(text=[prompt], images=[image], return_tensors="pt").to("cuda")
out = model.generate(**x, max_new_tokens=3072, do_sample=False)
print(processor.tokenizer.decode(out[0, x["input_ids"].shape[1]:], skip_special_tokens=True))
vLLM (much faster for many pages): load the repository with
LLM(model="MuseMesh/sansar-ocr-2b", limit_mm_per_prompt={"image": 1}, max_model_len=5672, dtype="bfloat16"), send
{"prompt": prompt, "multi_modal_data": {"image": image}} with the same prompt and page preparation, and
SamplingParams(temperature=0.0, max_tokens=3072, stop_token_ids=[248046, 248044]). The vLLM numbers on this card
were measured on the training run's checkpoint with vLLM 0.30; the release files hold the same weights (see Files).
What the output looks like
Plain text, one printed line per output line, Devanagari as printed (NFC, no ZWJ/ZWNJ), । and ॥ as dandas,
digits as printed (Devanagari or Latin). On book pages it tends to leave out running heads, page numbers and
footnotes, because most of its training pages had those painted out (see Training data); do not rely on either
behaviour. Latin-script text (English introductions, IAST) was rare in training and is not evaluated.
Evaluation
Bench. bench_v1_pages: 54 pages from 53 archive.org scanned books (mixed print and scan quality), each
paired with the same passage from a clean digital edition of the text, cut to what is printed on that page and
checked page by page. No page of the bench books (nor any page sharing three 16-letter windows with a bench
reference) is in the training or validation data, and the bench was never used to pick a checkpoint.
Metric. Letter error rate: edit distance on Devanagari letters and signs only (consonants, vowels, vowel signs, virama, anusvara, visarga, avagraha; no digits, dandas, Vedic accent marks, punctuation or spaces) divided by the reference's letter count, with the reference located inside the output (text the reference does not cover, such as a running head the model wrote out, is not counted against it). Median over the 54 pages, so one ruined page does not decide the score; the mean is shown too. The same scorer and pages for every engine.
| engine | median letter CER | mean | bench tier "kept" (39) | tier "mid" (15) | hardest third of scans (18) | note |
|---|---|---|---|---|---|---|
| Claude Opus 5.5 (Anthropic) | 0.65% | 1.90% | 0.63% | 1.08% | 2.11% | general LLM, page by page; comparison only |
| Sansar OCR 2B v0.1.0 (transformers) | 1.05% | 2.31% | 0.78% | 1.31% | 3.14% | this repository's files and ocr_page.py, re-run for the release on an NVIDIA GeForce RTX 3060 |
| Sansar OCR 2B v0.1.0 (vLLM) | 1.06% | 2.27% | 0.81% | 1.40% | 3.12% | the training run's own evaluation, A100 |
| Claude Sonnet 5.5 (Anthropic) | 1.14% | 2.17% | 1.10% | 1.31% | 3.12% | comparison only |
| Google Cloud Vision | 1.69% | 3.89% | 1.50% | 3.15% | 3.55% | DOCUMENT_TEXT_DETECTION, 2026-10; comparison only |
| Codex gpt-6-astra (OpenAI) | 2.65% | 3.92% | 2.18% | 3.55% | 6.66% | comparison only |
| archive.org OCR of the same scans | 9.03% | 9.64% | the text most of these books are online with today | |||
| Qwen3.5-2B, no fine-tuning | 18.54% | 28.89% | 17.04% | 21.88% | 27.19% | the base model, same prompt |
The Vision, Claude and Codex rows are those services' transcriptions of the same 54 page images (2026-09/10), scored by the same scorer, shown for comparison only; none of their output was used to train this model. "kept"/"mid" are the bench's own page-quality tiers; "hardest third" are the 18 pages where the archive.org OCR did worst (a proxy for scan quality).
- The release check. The model in this repository (repackaged: see Files) was re-run on all 54 pages with
transformers and
ocr_page.pyfor this release: median 1.05%, mean 2.31% (NVIDIA GeForce RTX 3060, batch 4). The training run's own transformers cross-check of the same weights gave 1.06%; its vLLM run 1.06%. Different engines and GPUs change a few characters on some pages (bfloat16 kernels). Everything is ineval/. - Small bench. With 54 pages a median difference of a few tenths of a point has a wide interval; the page-by-page count (better than Vision on 39 of 54 pages, vLLM) is the sturdier evidence.
- The worst pages are bad for every engine (scan or reference problems, not this model's).
- Validation. On the 711-page validation set (31 books, held out by book; the same kinds of sources as the training data) the median is 0.25% and the mean 0.93%. It is easier than the bench and not comparable to it.
- Overall CER including punctuation, digits and spacing is higher: median 3.50% on the bench (Vision 4.56%).
Speed
| GPU | engine | pages per hour (generation only) |
|---|---|---|
| NVIDIA A100-SXM4-40GB (GCP spot) | vLLM, batch 64 | 9,216 (54 test pages) / 15,104 (711 validation pages) |
| NVIDIA GeForce RTX 3060 | transformers generate, batch 4 |
244 (54 test pages) |
About 847-1,035 output tokens per page (vLLM). Start-up (model load, vLLM compile) is not included.
Training
| base | Qwen/Qwen3.5-2B (revision 15852e8), Apache-2.0: hybrid Gated-DeltaNet/attention language model with its vision encoder |
| method | full fine-tune of every weight (vision encoder at a lower learning rate), loss on the transcription only |
| steps | 2,000 x 16 pages = 32,000 page samples, 80.7M tokens seen |
| optimiser | 8-bit AdamW (bitsandbytes), lr 1.5e-5 (vision encoder 5e-6), 3% warmup, bfloat16 autocast over float32 master weights, gradient checkpointing |
| pages | the prompt above + page image (<= 1.8 MP) + label + `< |
| mixture | sampling weights per source (below); the last 15% of samples drawn from the gold sources (twin, textbooks, news) only |
| checkpoint | picked by loss on a source-balanced gold validation subset (276 pages, <= 120 per source), never on the test bench: step 2,000 (val loss 0.0621; 0.1577 at step 200) |
| hardware | 1x NVIDIA A100-SXM4-40GB, Google Cloud spot (a2-highgpu-1g, us-central1-a) |
| time and cost | 7.6 h of VM time for training and the evaluations (1 spot stop; training resumed from its last checkpoint), about $9.9 of spot time |
training/summary.json holds every argument and the validation-loss curve; training/mixture.json the sample counts;
training/label_stats.json the label-set counts and exactly what was excluded.
Training data
19,377 training pages from 1,057 books, split by book from 711 validation pages (31 books) and the 54 test pages. Page images and labels are not redistributed here: most of the scans and texts carry no open licence. Where each label comes from:
| source | train pages | weight | samples drawn | label |
|---|---|---|---|---|
| archive.org scans + a clean e-text of the same work ("twin") | 11,074 | 1.5 | 27,467 | a clean digital edition of the work (Sanskrit Wikisource proofread pages, GRETIL, Muktabodha, sanskritdocuments.org, Ambuda, sanskritsahitya.org, snskrt) aligned to the printed page and cut at its printed line breaks; running heads, page numbers and footnotes are painted out of the image |
| born-digital textbooks (NCERT, Kerala SCERT) | 3,367 | 0.5 | 2,808 | the PDF's own text layer, repaired by rule (split words, doubled signs, reph order); ~1 split word in 100 remains |
| All India Radio Sanskrit news bulletins (2023-2026) | 3,839 | 0.2 | 1,270 | the PDF text decoded from its fonts (the fonts' own glyph tables, and a vocabulary-based decipherment of unmapped glyphs); no OCR; ~1-1.5% character errors expected |
| synthetic pages | 1,097 | 0.3 | 455 | paragraphs from curated open Sanskrit e-text corpora (GRETIL, DCS, SARIT, Muktabodha, DharmaNexus, BORI Mahābhārata, ...) typeset in 1-2 columns with real fonts and degraded (skew, blur, noise, bleed); exact labels |
No commercial AI service output is training text. Labels come only from clean e-texts aligned to scans, the text layers of born-digital textbooks and All India Radio bulletins, and synthetic pages typeset from curated e-text corpora. Excluded on purpose: page transcriptions made with Google Cloud Vision, Anthropic Claude and OpenAI Codex (an earlier, internal label set had used them), and synthetic pages typesetting paragraphs from Sarvam's Vagartha dataset. Where those services were used at all, it was as a check, never as label text: the news-bulletin decoding was validated against Vision on a sample of pages, and the comparison rows in Evaluation are their transcriptions of the test pages. archive.org's own OCR is used only to locate the printed lines on twin pages. Wikisource texts are human-proofread transcriptions by Wikisource volunteers.
All labels are NFC with ZWJ/ZWNJ removed, | as ।, and a colon typed after a letter as visarga. Every training
and validation page that shares three or more 16-letter windows with a test reference was dropped.
Limitations
- Printed Devanagari Sanskrit only. Not trained on manuscripts, handwriting, palm leaves, other scripts or Hindi/Marathi pages (some Hindi in the textbooks and news, but it was not evaluated).
- Reading order on two-column pages, glossaries, tables and picture captions can be wrong; on scans of two-page spreads it can read across. (Fewer such pages were in training than in our earlier internal model, which also learnt from Vision's text of glossary and table pages.)
- Loops and runaways. On a hard page it can repeat a line until the 3,072-token cap. Check for repeated lines.
- Page furniture (running heads, page numbers, footnotes) is sometimes left out and sometimes written out.
- Not scored: digits, punctuation, spacing and Vedic accents are outside the letter metric; the overall error rate including them is 3.50% (median).
- It can invent text on blank pages, figures or very faint scans, as any generative OCR can. It also inherits the base model's general behaviour on anything that is not a Sanskrit page.
- Labels are imperfect: the twin labels follow another edition of the same work (a few percent of variant readings and verse numbers), the textbook layers split words, the news decoding has rare glyph errors. The model learnt some of that.
Versions
Each version is a git tag; main is the newest. Pin one with revision="v0.1.0".
| version | date | training run | bench median (transformers / vLLM) | Google Vision |
|---|---|---|---|---|
| v0.1.0 | 2026-10-08 | ocrvlm_v12_clean_q35 |
1.05% / 1.06% | 1.69% |
Files
| file | what |
|---|---|
model.safetensors |
the weights, bfloat16 (the gated-DeltaNet decay and norm parameters in float32 as in the base release); the output head tied to the input embedding |
config.json, generation_config.json |
architecture; greedy decoding with both stop tokens |
tokenizer.json, tokenizer_config.json, chat_template.jinja, processor_config.json |
the base model's tokenizer and processor, unchanged |
ocr_page.py |
page preparation, loading and batch transcription (the code above) |
eval/bench_v1_scores.json |
per-page and per-group scores of every engine in the table |
eval/verification.json, eval/verification_pages.json |
the release check: this repository's files re-run on the 54 pages |
eval/transcriptions/ |
the training run's own outputs for the 54 pages (vLLM and transformers) |
eval/*_scores.json, eval/*_timing.json |
the training run's evaluation logs (test, validation) |
training/summary.json, training/mixture.json, training/cloud.json, training/label_stats.json |
arguments, validation-loss curve, sample counts, machine time, label-set counts |
LICENSE, LICENSE-CODE, LICENSE-QWEN, NOTICE, CHANGELOG.md |
licences, base-model attribution, version history |
Repackaging against the training checkpoint: the training run saved the tied output head as a second copy of the
input embedding (bit-identical; dropped); config.json records bfloat16 and use_cache: true instead of the
training-time float32/false; generation_config.json adds the <|im_end|> stop token and greedy decoding. No
weight changed.
Licence
This release is for research and non-commercial use. The training labels come partly from sources without an open licence (textbooks, news bulletins, scanned books), used here for research.
- Weights (
model.safetensors): CC BY-NC 4.0 (LICENSE). Non-commercial use (research, teaching, non-profit work) with attribution to "Sansar, Muse Mesh Private Limited". - Base model: Qwen3.5-2B by Alibaba Cloud, Apache-2.0 (
LICENSE-QWEN,NOTICE). The tokenizer and processor files are the base model's and stay under Apache-2.0. - Code (
ocr_page.py): Apache-2.0 (LICENSE-CODE). - Commercial use: contact kushal@muse-mesh.com.
Citation
@misc{sansar_ocr_2b_2026,
title = {Sansar OCR 2B: printed Sanskrit page OCR},
author = {Muse Mesh},
year = {2026},
note = {v0.1.0, fine-tuned from Qwen3.5-2B},
url = {https://huggingface.co/MuseMesh/sansar-ocr-2b}
}
Contact: kushal@muse-mesh.com
- Downloads last month
- 4
Model tree for MuseMesh/sansar-ocr-2b
Collection including MuseMesh/sansar-ocr-2b
Evaluation results
- median letter error rate, transformers (this repository's files) on Sansar bench_v1 pages (54 archive.org book pages, 53 books)self-reported1.050
- median letter error rate, vLLM (the training run's evaluation) on Sansar bench_v1 pages (54 archive.org book pages, 53 books)self-reported1.060
- mean letter error rate, transformers on Sansar bench_v1 pages (54 archive.org book pages, 53 books)self-reported2.310