AI & ML interests
Sanskrit and Indic language models, tokenizers, multilingual AI agents
Recent Activity
Muse Mesh
We build language technology for Sanskrit and other Indic languages: corpora, tokenizers and language models trained from scratch. Every training run we do, finished or live, is published at mume.ai/lab.
Sansar: Sanskrit-only language models
Sansar is a family of small causal language models trained only on Sanskrit text, with a Sanskrit tokenizer of its own. Each model is trained from scratch; none is a fine-tune of an English model.
| Model | Parameters | Bits per byte (held-out, ex-Gītā) ↓ |
|---|---|---|
| sansar-350m | 318M | 0.5547 |
| sansar-125m | 97M | 0.6039 |
| sansar-60m | 63M | 0.6434 |
| sansar-20m | 27M | 0.7177 |
Scores are measured on frozen held-out Sanskrit sets: DCS gold sentences, prose, Vedic and out-of-domain texts. Held-out text is masked out of training. Every model card gives the per-set numbers, the training data and a version history.
from transformers import AutoModelForCausalLM, AutoTokenizer
tok = AutoTokenizer.from_pretrained("MuseMesh/sansar-350m", trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained("MuseMesh/sansar-350m", trust_remote_code=True)
ids = tok("संस्कृतं नाम दैवी वाक्", return_tensors="pt")
print(tok.decode(model.generate(**ids, max_new_tokens=60, do_sample=True, top_k=100)[0]))
English and Math models
Small models trained from scratch with the same recipe (Muon, rotary positions, QK-norm, squared ReLU) and a shared 32k tokenizer, released with their evaluation sets and the exact list of training documents. Collection.
| Model | Parameters | Trained on | Result |
|---|---|---|---|
| mume-english-125m | 134M | 1.33B FineWeb tokens | 1.106 bits per byte on frozen FineWeb val; on the validation stream 1.132 vs 1.176 for a plain-AdamW GPT-2 124M at the same tokens (3.7% lower) |
| mume-math-125m | 134M | 1.33B OpenWebMath tokens | MATH test (clean) 62.6% fewer bits than the English model; branch sft is a GSM8K/MATH fine-tune |
- mume-tokenizer-32k: the shared 32k unigram tokenizer.
- mume-eval-suites: the frozen test sets behind every number (FineWeb val, WikiText-103, enwik8, text8, GSM8K, MATH, OpenWebMath val) with contamination-clean variants, plus training manifests.
Data and tokenizer
- Sansar Sanskrit Corpus: 252M words of openly licensed Sanskrit from 21 sources, in Devanagari. It is split into one config per licence (CC BY, CC BY-SA, CC0, ODC-By, Apache-2.0, MIT), and every record carries its source, licence and attribution. Text overlapping our evaluation sets is removed.
- sansar-sanskrit-tokenizer: an 8k unigram SentencePiece model over SLP1 transliteration with a Devanagari wrapper. 2.6 to 3.9 tokens per Sanskrit word with only 8k entries.
Licences
The corpus records keep their source licences. Model weights and the tokenizer are released under CC BY-NC 4.0 for non-commercial research use, and the modelling code under Apache-2.0. For other uses, contact us.
Links: muse-mesh.com · GitHub · kushal@muse-mesh.com