Instructions to use MuseMesh/sansar-700m with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use MuseMesh/sansar-700m with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="MuseMesh/sansar-700m", trust_remote_code=True)# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("MuseMesh/sansar-700m", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use MuseMesh/sansar-700m with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "MuseMesh/sansar-700m" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "MuseMesh/sansar-700m", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/MuseMesh/sansar-700m
- SGLang
How to use MuseMesh/sansar-700m with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "MuseMesh/sansar-700m" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "MuseMesh/sansar-700m", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "MuseMesh/sansar-700m" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "MuseMesh/sansar-700m", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use MuseMesh/sansar-700m with Docker Model Runner:
docker model run hf.co/MuseMesh/sansar-700m
Sansar 700M
A 704.1M-parameter Sanskrit language model trained from scratch, from the Sansar project at Muse Mesh (Hugging Face): Sanskrit language models, their corpus and their tokenizer. Size class: the d700m preset: 24 layers, width 1,536, 12 heads (679.6M parameters outside the embeddings; 704.1M in total with the 8k vocabulary and untied head).
- This version: v0.1.0 = training run
f10_slp1_uni8k_d700m_plus_all_v2_3x, finished 2026-10-07 (experiment F10: 700M, F9 recipe plus weight decay, on the filtered plus_all v2 slice at three passes). - Held-out score: 0.5307 bits per Devanagari byte (pooled over four held-out sets, excluding the Bhagavad-gītā; lower is better); 0.5541 on the stricter clean_v1 sets.
- Base model: it continues Devanagari Sanskrit text; it is not instruction-tuned.
- Why this checkpoint: The best model so far: -4.3% ex-Gītā against Sansar 350M v0.2.0 (F9, 0.5547) and -4.0% on clean_v1 (0.5773), better on every held-out set. Size (318M -> 704M), the data slice (filtered, plus web crawls) and weight decay changed together, so the gain is not split between them.
Quick start
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
repo = "MuseMesh/sansar-700m"
tok = AutoTokenizer.from_pretrained(repo, revision="v0.1.0", trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(repo, revision="v0.1.0", trust_remote_code=True)
inputs = tok("संस्कृतं नाम दैवी वाक्", return_tensors="pt") # Devanagari in
out = model.generate(**inputs, max_new_tokens=50, do_sample=True, temperature=0.8, top_k=100)
print(tok.decode(out[0], skip_special_tokens=True)) # Devanagari out
Needs transformers (tested with 5.18.0), torch and sentencepiece. trust_remote_code=True runs three small files from this repository: modeling_sansar.py (the network), tokenization_sansar.py and translit.py (Devanagari <-> SLP1); read them before you run them. The weights are stored in bfloat16 (transformers 5 loads them as such); pass dtype=torch.float32 for fp32, usually the faster choice on a CPU. The tokenizer adds no BOS/EOS: training records were separated by </s> only.
What this checkpoint wrote for that prompt (CPU, torch.manual_seed(0), 50 new tokens):
संस्कृतं नाम दैवी वाक् । सैव अग्रे आचार्याणामपि प्राक् आचार्यैः प्रकाशिता आसीत् । अस्यां वाङ्मयसृष्टौ अस्माकं देशस्य कृते संस्कृतजगति अस्या एवोत्कृष्ट्
and with greedy decoding:
संस्कृतं नाम दैवी वाक् । अनादिनिधना नित्या वागुत्सृष्टा स्वयंभुवा । आदौ वेदमयी दिव्या यतः सर्वाः प्रवृत्तयः ॥ इति स्मृतेः । अनादिनिधना नित्या वागुत्सृष्टा स्वयं
To score text instead of generating it, call the model with labels=input_ids; loss is the mean negative log-likelihood per token in nats.
Model details
| Architecture | decoder-only transformer, pre-LayerNorm (LayerNorm without bias), no biases anywhere |
| Layers / heads / width | 24 / 12 / 1,536 (head dim 128); MLP 4x = 6,144 |
| Positions | rotary (base 10,000; half-split pairing), no position table |
| Attention | causal SDPA; queries and keys RMS-normalised per head (QK-norm) |
| MLP activation | squared ReLU |
| Output head | separate (untied), zero-initialised |
| Vocabulary | 8,000 (SentencePiece unigram over SLP1, MuseMesh/sansar-sanskrit-tokenizer v0.1.0) |
| Context | 512 tokens (about 3.6 kB of Devanagari text at the training data's 7.08 Devanagari bytes per token) |
| Parameters, total | 704,128,512 |
| Parameters, non-embedding | 679,552,512 |
| Token embedding | 12,288,000 |
| Output head | 12,288,000 |
| Weights in this repo | bfloat16 safetensors (trained as fp32 master weights under bf16 autocast) |
Training
| Optimizer, block matrices | Muon (679,477,248 params): momentum 0.95 (Nesterov), 5 Newton-Schulz steps in bf16, peak lr 0.02, update scaled by sqrt(max(1, rows/cols)), decoupled weight decay 0.0031 (p <- p x (1 - lr x wd), with the scheduled lr) |
| Optimizer, everything else | AdamW: embedding + head (24,576,000 params) with weight decay 0.062, LayerNorm gains without; peak lr 0.001, betas (0.9, 0.95), eps 1e-8 |
| Learning-rate schedule | linear warm-up over 500 steps, then cosine decay to 0.1 x peak at step 96,822 (Muon and AdamW share it) |
| Batch | 24 sequences x 4 accumulation steps x 512 tokens = 49,152 tokens per step |
| Steps | 96,822 |
| Tokens seen | 4,758,994,944 (3.00 passes over 1,586.3M training tokens) |
| Gradient clipping | global norm 1.0 |
| Dropout | 0.0 |
| Precision | bf16 autocast, fp32 master weights, torch.compile |
| Seed | 1337 |
| Data sampling | 512-token windows at uniformly random offsets of the token stream (records separated by </s>), so passes are counted in expectation |
| Hardware | 1x NVIDIA A100 40 GB (Google Cloud a2-highgpu-1g, spot, us-central1) |
| Wall time | 40.3 h (145,253 s), 33.6k tokens/s mean, no preemptions |
| Compute cost | about $52 of spot GPU time (watcher estimate $52.6 at $1.30/h) |
| Final validation | loss 2.1945 nats/token = 0.4491 bits per Devanagari byte on the run's own validation split (0.5% of records by near-duplicate cluster; not comparable across data slices) |
Training code: scripts/train/train.py, model.py and muon.py of the Sansar repository; the exact arguments are in training/summary.json.
Training data
train_slice_plus_all_v2 (489.9M words, 14.09M records) = the Sansar 350M v0.2.0 slice (520.4M words) plus 1.9M words of Sanskrit from open web crawls (HPLT 3.0, MaLA, CC-100), passed through quality filters that dropped 32.5M words (6.2%): transliteration junk, Pali, garbled text, strict-Hindi lines and low-vocabulary crawl text, each rule checked by eye on 30 dropped samples; a 60M-parameter A/B run found the filters neutral on quality. The same pass repaired 350,187 records that typed the visarga as an ASCII colon. After held-out masking: 14,034,636 training records, 1,586.3M tokens, 11.23 GB of Devanagari text. The run saw 4,759.0M tokens = 3.0 passes.
Held-out texts excluded by dedup key and masked inside training records (28-character windows, stride 4: 153,307 records masked, 16.2M characters removed, 22,597 records dropped); every source added since Sansar 350M v0.1.0 passed an independent leak gate (0 held-out items per set). The filter pass also removed held-out passages hidden behind ASCII ":" visargas in the added sources; the older base records still carry that blind spot, so the clean_v1 columns below re-score on the held-out items with any such overlap removed. The Gītā is still partly memorised through near-copies in commentaries, so it stays out of the headline number.
The text comes from the Sansar corpus, which collects Sanskrit in Devanagari from public sources: classical e-text collections (GRETIL, SARIT, Muktabodha, the Digital Corpus of Sanskrit, DharmaNexus and others), Sanskrit Wikisource and Wikipedia, dictionaries, and the Sanskrit parts of web-crawl datasets (AI4Bharat Sangraha, IndicCorp, MADLAD-400, the sanskrit-monolingual-pretraining collection). Every record keeps its provenance and a licence tier (T0 permissive, T1 share-alike, T2 non-commercial, T3 no licence statement or all rights reserved). The training slice mixes all tiers: a large share is licensed for non-commercial use only and some sources state no licence, which is why the weights are released under CC BY-NC 4.0 (see Licence). Old archive.org OCR of printed books is left out (it measurably hurt the models).
Preparation: Unicode NFC; standalone / and // read as daṇḍa । and double daṇḍa ॥; machine reference markers (verse numbers of digital editions) stripped; transliterated to SLP1 and tokenized; one </s> after every record; a 0.5% validation split by near-duplicate cluster.
Evaluation
Metric: bits per Devanagari byte (lower is better): the model's negative log-likelihood of a text divided by the UTF-8 byte count of the same text in Devanagari, so models with different tokenizers are measured against the same denominator. Scored teacher-forced with scripts/train/eval_bpb.py: each set's records joined with </s> into one stream, 512-token windows with stride 256 (every scored token after the first window has at least 256 tokens of context), text cleaned exactly as the training text was.
Sets (E0, frozen before any model was trained and excluded from training): dcs_gold 3,000 sentences of the Digital Corpus of Sanskrit gold standard (classical), prose 2,470 prose passages, ood 2,000 web, Wikipedia and other out-of-domain texts, vedic 1,000 accented Ṛgveda pādas, gita all 700 verses of the Bhagavad-gītā. The headline is pooled excluding the Gītā (byte-weighted over the other four): the Gītā is quoted inside commentaries and epics throughout the training text, so its column measures memorisation.
| ex-Gītā (headline) | pooled, all five | dcs_gold | prose | ood | vedic | gita (memorisation) | |
|---|---|---|---|---|---|---|---|
| E0 sets (9,170 items) | 0.5307 | 0.5140 | 0.5513 | 0.5186 | 0.5286 | 0.6670 | 0.0758 |
| clean_v1 (8,270 items) | 0.5541 | 0.5312 | 0.5547 | 0.5315 | 0.5598 | 0.6697 | 0.0818 |
Verse completion (600 verses: 200 each from the Bhagavad-gītā, Mahābhārata and Rāmāyaṇa; the model gets the first half-verse and greedily writes the second, stopping at ॥ or a newline): chrF 0.142, exact match 4/600. chrF credits shared character n-grams, so it rewards plausible vocabulary even when the half-verse is not the right one.
clean_v1 (added 2026-10-04): the same five sets minus every item that could overlap the training text in a way the held-out masking could not see. Some training text types the visarga as an ASCII colon (":" for "ः"); both the held-out mask and the leak gate ignored ":", so a 16-character window of a held-out item could survive inside such a record. clean_v1 removes every item with any such window in either training slice (900 of 9,170 items; mostly a shared stock phrase, rarely a real near-copy). It keeps 8,270 items; ood loses half its bytes and reads harder, so compare clean_v1 numbers only with clean_v1 numbers.
Reference points
Same metric and sets. The external base models were scored zero-shot through the same code path (scripts/ops/extbench_bpb.py: the same items, cleaning, Devanagari byte count and </s>-joined stream as eval_bpb.py) with their own tokenizers and a 2,048-token window (stride 1,024), which gives them more context than our 512-token window; their training data may contain the public ood and prose texts. Sansar rows are the released checkpoints.
Not in the table: Krutrim-2-instruct (12B) scored about 0.51 ex-Gītā in an earlier pass (4-bit nf4 weights, verse references stripped) that is not strictly comparable and was not re-run. On that number it is ahead of this model, so this card does not claim to beat every model.
| model | parameters | ex-Gītā | clean_v1 ex-Gītā | notes |
|---|---|---|---|---|
| Sansar 20M v0.1.0 | 27.1M | 0.7177 | n/a | TOK-v3 screen, arm v0.1, 0.34B tokens |
| Sansar 60M v0.1.0 | 63.2M | 0.6434 | n/a | F0, 1.34B tokens |
| Sansar 125M v0.1.0 | 97.2M | 0.6039 | 0.6254 | F6-clean-2x, 2.66B tokens |
| Sansar 350M v0.1.0 | 318.4M | 0.5720 | 0.5937 | F7, 2.66B tokens |
| Sansar 350M v0.2.0 | 318.4M | 0.5547 | 0.5773 | F9, 5.09B tokens |
| Sansar 700M v0.1.0 (this model) | 704.1M | 0.5307 | 0.5541 | F10, 4.76B tokens |
| Gemma 3 4B pt (zero-shot, bf16, BOS in every window) | 3.9B | 0.6965 | 0.7365 | general base model, 2,048-token window |
| Qwen3-4B-Base (zero-shot, bf16) | 4.0B | 0.7071 | 0.7450 | general base model, 2,048-token window |
| Qwen3.5-4B-Base (zero-shot, bf16) | 4.2B | 0.7114 | 0.7569 | general base model, 2,048-token window |
| Llama 3.2 3B (zero-shot, bf16, BOS in every window) | 3.2B | 0.7122 | 0.7607 | general base model, 2,048-token window |
| Sarvam-1 (zero-shot, bf16, BOS in every window) | 2.5B | 0.7465 | 0.8166 | general base model, 2,048-token window |
Checked before release
modeling_sansar.pyagainst the training code (scripts/train/model.py), fp32 on CPU, 6 inputs (one held-out text per set, cropped to 512 tokens, and 512 random ids): with the same bf16 weights the logits are identical (max |difference| 0.0e+00); against the fp32 training weights the bf16 storage moves logits by at most 0.208.- Bits per byte on the first 20 records of each set with
eval_bpb.py's own scoring: 0.48407 (training checkpoint, fp32) vs 0.48405 (this repo, bf16 weights) = -0.0039%. - Every record of every set, this repo's bf16 weights through the transformers code on a GPU (
scripts/ops/extbench_bpb.py, 512:256): ex-Gītā 0.53059 vs 0.53067 from the training checkpoint (-0.016%), clean_v1 0.55404 vs 0.55415 (-0.019%); dcs_gold 0.55111 vs 0.55126. Same numbers to the fourth decimal, up to rounding. - Tokenizer: the same ids as the evaluation pipeline on 100/100 sample texts with
fence_latin=False(95/100 with the default fence; the rest contain English words); Devanagari round trip exact on 100/100. - Left-padded batches give the same logits as single sequences (max |difference| 3.5e-05).
Limitations
- Base model. It continues Sanskrit text. It has not been instruction-tuned or aligned: it does not follow instructions, answer questions or chat, and it will not refuse anything.
- It makes things up. Output is fluent-looking Sanskrit that can be ungrammatical, mix registers and invent verses, authors and works. Do not use it as a source of quotations or facts; exact verse recall is close to zero (see verse completion).
- Other languages in the data. The quality filters of this run removed most Pali and strict-Hindi text, but some Hindi and Marathi lines remain, so the model can still drift into Hindi.
- Small and short. 704M parameters and a 512-token context. This implementation has no key/value cache, so long generations are slow (each new token re-reads the last 512).
- Vedic is the weakest register: accent marks fragment the tokenization and Vedic text is a small share of the training data.
- Script. Devanagari in and out. English words inside the text are fenced by the tokenizer so they come back unchanged; the model never saw the fence marks in training, so its predictions around them are weaker (
AutoTokenizer.from_pretrained(..., fence_latin=False)reproduces the training encoding, but then Latin letters decode as Devanagari). IAST and other Indic scripts are not transliterated for you. - What the corpus says, the model says. Most of the text is religious, philosophical and classical literature, plus modern web and news Sanskrit; the model reproduces its views and its errors, including OCR errors that survived in web-crawl sources.
- The Gītā is memorised in part (see the
gitacolumn), so do not read its score as generalisation.
Versions
Each version is a git tag on this repository; main is the newest. Pin one with revision="v0.1.0". A version is one training run, named by its run id in the Sansar experiment log.
| version | date | training run | tokens seen | ex-Gītā | clean_v1 ex-Gītā |
|---|---|---|---|---|---|
| v0.1.0 | 2026-10-07 | f10_slp1_uni8k_d700m_plus_all_v2_3x (F10) |
4.76B | 0.5307 | 0.5541 |
All runs of this size (ex-Gītā on the same E0 sets):
| run | date | recipe | ex-Gītā | status |
|---|---|---|---|---|
| F10 | 2026-10-07 | d700m, plus_all v2 slice (489.9M words), 4.76B tokens, Muon wd 0.0031 + AdamW wd 0.062 | 0.5307 | v0.1.0 |
Files
| file | what |
|---|---|
model.safetensors |
the weights, bfloat16 (no optimizer state) |
config.json, generation_config.json |
architecture and default sampling settings |
configuration_sansar.py, modeling_sansar.py |
the network for transformers (auto_map, trust_remote_code) |
tokenizer.model, tokenization_sansar.py, translit.py, tokenizer_config.json, special_tokens_map.json |
the tokenizer (MuseMesh/sansar-sanskrit-tokenizer v0.1.0) with its Devanagari <-> SLP1 wrapper |
eval/bpb.json, eval/bpb_clean_v1.json, eval/verse_scores.json |
the evaluation outputs quoted above |
eval/verification.json, eval/smoke_test.json |
the release checks and the sample generations |
training/summary.json, training/data_meta.json, training/cloud.json |
every training argument, the loss curve's evaluation points, and the data preparation record |
LICENSE, LICENSE-CODE, CHANGELOG.md |
licences and version history |
Licence
This release is for research and non-commercial use. The training data includes texts licensed for non-commercial use only and texts without a licence statement, used here for research. A commercially licensed model, trained only on permissively licensed text, is planned as a separate release with its own version line.
- Weights (
model.safetensors) andtokenizer.model: CC BY-NC 4.0 (LICENSE). Non-commercial use (research, teaching, non-profit work) with attribution to "Sansar, Muse Mesh Private Limited". - Code (
configuration_sansar.py,modeling_sansar.py,tokenization_sansar.py,translit.py): Apache-2.0 (LICENSE-CODE). - Commercial use: contact kushal@muse-mesh.com.
Citation
@misc{sansar_700m_2026,
title = {Sansar 700M: a Sanskrit language model},
author = {Muse Mesh},
year = {2026},
note = {v0.1.0},
url = {https://huggingface.co/MuseMesh/sansar-700m}
}
Contact: kushal@muse-mesh.com
- Downloads last month
- -
Datasets used to train MuseMesh/sansar-700m
chronbmm/sanskrit-monolingual-pretraining
Collection including MuseMesh/sansar-700m
Evaluation results
- bits per Devanagari byte, pooled excluding the Bhagavad-gītā (E0 held-out sets) on Sansar E0 held-out Sanskrit setsself-reported0.531
- bits per Devanagari byte, pooled over all five E0 sets on Sansar E0 held-out Sanskrit setsself-reported0.514
- bits per Devanagari byte, dcs_gold on Sansar E0 held-out Sanskrit setsself-reported0.551
- bits per Devanagari byte, prose on Sansar E0 held-out Sanskrit setsself-reported0.519
- bits per Devanagari byte, ood on Sansar E0 held-out Sanskrit setsself-reported0.529
- bits per Devanagari byte, vedic on Sansar E0 held-out Sanskrit setsself-reported0.667
- bits per Devanagari byte, gita on Sansar E0 held-out Sanskrit setsself-reported0.076
- bits per Devanagari byte, pooled excluding the Bhagavad-gītā (clean_v1 sets) on Sansar E0 held-out Sanskrit setsself-reported0.554