--- title: README emoji: 🪷 colorFrom: yellow colorTo: red sdk: static pinned: false --- # Muse Mesh We build language technology for **Sanskrit** and other Indic languages: corpora, tokenizers and language models trained from scratch. Every training run we do, finished or live, is published at the **[Sansar Lab](https://saansar.com)**. ## Sansar: Sanskrit-only language models **Sansar** is a family of small causal language models trained only on Sanskrit text, with a Sanskrit tokenizer of its own. Each model is trained from scratch; none is a fine-tune of an English model. | Model | Parameters | Bits per byte (held-out, ex-Gītā) ↓ | |---|---|---| | [sansar-700m](https://huggingface.co/MuseMesh/sansar-700m) | 704M | **0.5307** | | [sansar-350m](https://huggingface.co/MuseMesh/sansar-350m) | 318M | 0.5547 | | [sansar-125m](https://huggingface.co/MuseMesh/sansar-125m) | 97M | 0.6039 | | [sansar-60m](https://huggingface.co/MuseMesh/sansar-60m) | 63M | 0.6434 | | [sansar-20m](https://huggingface.co/MuseMesh/sansar-20m) | 27M | 0.7177 | Scores are measured on frozen held-out Sanskrit sets: DCS gold sentences, prose, Vedic and out-of-domain texts. Held-out text is masked out of training. Every model card gives the per-set numbers, the training data and a version history. Scored the same way, general base models need more bits per byte: Gemma 3 4B 0.6965, Qwen3-4B 0.7071, Llama 3.2 3B 0.7122, Sarvam-1 0.7465. sansar-700m beats all four on every held-out set with 704M parameters (Krutrim-2 12B, scored earlier in a pass that is not strictly comparable, is not in this list). A live demo is at **[saansar.com/demo](https://saansar.com/demo)**. ```python from transformers import AutoModelForCausalLM, AutoTokenizer tok = AutoTokenizer.from_pretrained("MuseMesh/sansar-700m", trust_remote_code=True) model = AutoModelForCausalLM.from_pretrained("MuseMesh/sansar-700m", trust_remote_code=True) ids = tok("संस्कृतं नाम दैवी वाक्", return_tensors="pt") print(tok.decode(model.generate(**ids, max_new_tokens=60, do_sample=True, top_k=100)[0])) ``` ## English and Math models Small models trained from scratch with the same recipe (Muon, rotary positions, QK-norm, squared ReLU) and a shared 32k tokenizer, released with their evaluation sets and the exact list of training documents. [Collection](https://huggingface.co/collections/MuseMesh/mume-english-and-math-language-models-6ac4f187f33abd6c9a9cdde6). | Model | Parameters | Trained on | Result | |---|---|---|---| | [mume-english-125m](https://huggingface.co/MuseMesh/mume-english-125m) | 134M | 1.33B FineWeb tokens | 1.106 bits per byte on frozen FineWeb val; on the validation stream 1.132 vs 1.176 for a plain-AdamW GPT-2 124M at the same tokens (3.7% lower) | | [mume-math-125m](https://huggingface.co/MuseMesh/mume-math-125m) | 134M | 1.33B OpenWebMath tokens | MATH test (clean) 62.6% fewer bits than the English model; branch `sft` is a GSM8K/MATH fine-tune | - **[mume-tokenizer-32k](https://huggingface.co/MuseMesh/mume-tokenizer-32k)**: the shared 32k unigram tokenizer. - **[mume-eval-suites](https://huggingface.co/datasets/MuseMesh/mume-eval-suites)**: the frozen test sets behind every number (FineWeb val, WikiText-103, enwik8, text8, GSM8K, MATH, OpenWebMath val) with contamination-clean variants, plus training manifests. ## Data and tokenizer - **[Sansar Sanskrit Corpus](https://huggingface.co/datasets/MuseMesh/sansar-sanskrit-corpus)**: 252M words of openly licensed Sanskrit from 21 sources, in Devanagari. It is split into one config per licence (CC BY, CC BY-SA, CC0, ODC-By, Apache-2.0, MIT), and every record carries its source, licence and attribution. Text overlapping our evaluation sets is removed. - **[sansar-sanskrit-tokenizer](https://huggingface.co/MuseMesh/sansar-sanskrit-tokenizer)**: an 8k unigram SentencePiece model over SLP1 transliteration with a Devanagari wrapper. 2.6 to 3.9 tokens per Sanskrit word with only 8k entries. ## Licences The corpus records keep their source licences. Model weights and the tokenizer are released under **CC BY-NC 4.0** for non-commercial research use, and the modelling code under **Apache-2.0**. For other uses, contact us. **Links:** [muse-mesh.com](https://muse-mesh.com) · [GitHub](https://github.com/muse-mesh) · kushal@muse-mesh.com