AI & ML interests

Sanskrit and Indic language models, tokenizers, multilingual AI agents

Recent Activity

ks2812  updated a Space about 11 hours ago
MuseMesh/README
ks2812  published a dataset about 11 hours ago
MuseMesh/mume-eval-suites
ks2812  published a model about 11 hours ago
MuseMesh/mume-math-125m
View all activity

Organization Card

Muse Mesh

We build language technology for Sanskrit and other Indic languages: corpora, tokenizers and language models trained from scratch. Every training run we do, finished or live, is published at mume.ai/lab.

Sansar: Sanskrit-only language models

Sansar is a family of small causal language models trained only on Sanskrit text, with a Sanskrit tokenizer of its own. Each model is trained from scratch; none is a fine-tune of an English model.

Model Parameters Bits per byte (held-out, ex-Gītā) ↓
sansar-350m 318M 0.5547
sansar-125m 97M 0.6039
sansar-60m 63M 0.6434
sansar-20m 27M 0.7177

Scores are measured on frozen held-out Sanskrit sets: DCS gold sentences, prose, Vedic and out-of-domain texts. Held-out text is masked out of training. Every model card gives the per-set numbers, the training data and a version history.

from transformers import AutoModelForCausalLM, AutoTokenizer
tok = AutoTokenizer.from_pretrained("MuseMesh/sansar-350m", trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained("MuseMesh/sansar-350m", trust_remote_code=True)
ids = tok("संस्कृतं नाम दैवी वाक्", return_tensors="pt")
print(tok.decode(model.generate(**ids, max_new_tokens=60, do_sample=True, top_k=100)[0]))

English and Math models

Small models trained from scratch with the same recipe (Muon, rotary positions, QK-norm, squared ReLU) and a shared 32k tokenizer, released with their evaluation sets and the exact list of training documents. Collection.

Model Parameters Trained on Result
mume-english-125m 134M 1.33B FineWeb tokens 1.106 bits per byte on frozen FineWeb val; on the validation stream 1.132 vs 1.176 for a plain-AdamW GPT-2 124M at the same tokens (3.7% lower)
mume-math-125m 134M 1.33B OpenWebMath tokens MATH test (clean) 62.6% fewer bits than the English model; branch sft is a GSM8K/MATH fine-tune
  • mume-tokenizer-32k: the shared 32k unigram tokenizer.
  • mume-eval-suites: the frozen test sets behind every number (FineWeb val, WikiText-103, enwik8, text8, GSM8K, MATH, OpenWebMath val) with contamination-clean variants, plus training manifests.

Data and tokenizer

  • Sansar Sanskrit Corpus: 252M words of openly licensed Sanskrit from 21 sources, in Devanagari. It is split into one config per licence (CC BY, CC BY-SA, CC0, ODC-By, Apache-2.0, MIT), and every record carries its source, licence and attribution. Text overlapping our evaluation sets is removed.
  • sansar-sanskrit-tokenizer: an 8k unigram SentencePiece model over SLP1 transliteration with a Devanagari wrapper. 2.6 to 3.9 tokens per Sanskrit word with only 8k entries.

Licences

The corpus records keep their source licences. Model weights and the tokenizer are released under CC BY-NC 4.0 for non-commercial research use, and the modelling code under Apache-2.0. For other uses, contact us.

Links: muse-mesh.com · GitHub · kushal@muse-mesh.com