Download README.md from MuseMesh/README: direct link, hf CLI and curl.
- Browser
- Download file 4.34 kB
-
https://huggingface.co/spaces/MuseMesh/README/resolve/main/README.md
- Command line
-
hf download hf://spaces/MuseMesh/README/README.md
-
curl -L -o README.md https://huggingface.co/spaces/MuseMesh/README/resolve/main/README.md
title: README
emoji: 🪷
colorFrom: yellow
colorTo: red
sdk: static
pinned: false
Muse Mesh
We build language technology for Sanskrit and other Indic languages: corpora, tokenizers and language models trained from scratch. Every training run we do, finished or live, is published at mume.ai/lab.
Sansar: Sanskrit-only language models
Sansar is a family of small causal language models trained only on Sanskrit text, with a Sanskrit tokenizer of its own. Each model is trained from scratch; none is a fine-tune of an English model.
| Model | Parameters | Bits per byte (held-out, ex-Gītā) ↓ |
|---|---|---|
| sansar-700m | 704M | 0.5307 |
| sansar-350m | 318M | 0.5547 |
| sansar-125m | 97M | 0.6039 |
| sansar-60m | 63M | 0.6434 |
| sansar-20m | 27M | 0.7177 |
Scores are measured on frozen held-out Sanskrit sets: DCS gold sentences, prose, Vedic and out-of-domain texts. Held-out text is masked out of training. Every model card gives the per-set numbers, the training data and a version history.
Scored the same way, general base models need more bits per byte: Gemma 3 4B 0.6965, Qwen3-4B 0.7071, Llama 3.2 3B 0.7122, Sarvam-1 0.7465. sansar-700m beats all four on every held-out set with 704M parameters (Krutrim-2 12B, scored earlier in a pass that is not strictly comparable, is not in this list). A live demo is at mume.ai/sansar.
from transformers import AutoModelForCausalLM, AutoTokenizer
tok = AutoTokenizer.from_pretrained("MuseMesh/sansar-700m", trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained("MuseMesh/sansar-700m", trust_remote_code=True)
ids = tok("संस्कृतं नाम दैवी वाक्", return_tensors="pt")
print(tok.decode(model.generate(**ids, max_new_tokens=60, do_sample=True, top_k=100)[0]))
English and Math models
Small models trained from scratch with the same recipe (Muon, rotary positions, QK-norm, squared ReLU) and a shared 32k tokenizer, released with their evaluation sets and the exact list of training documents. Collection.
| Model | Parameters | Trained on | Result |
|---|---|---|---|
| mume-english-125m | 134M | 1.33B FineWeb tokens | 1.106 bits per byte on frozen FineWeb val; on the validation stream 1.132 vs 1.176 for a plain-AdamW GPT-2 124M at the same tokens (3.7% lower) |
| mume-math-125m | 134M | 1.33B OpenWebMath tokens | MATH test (clean) 62.6% fewer bits than the English model; branch sft is a GSM8K/MATH fine-tune |
- mume-tokenizer-32k: the shared 32k unigram tokenizer.
- mume-eval-suites: the frozen test sets behind every number (FineWeb val, WikiText-103, enwik8, text8, GSM8K, MATH, OpenWebMath val) with contamination-clean variants, plus training manifests.
Data and tokenizer
- Sansar Sanskrit Corpus: 252M words of openly licensed Sanskrit from 21 sources, in Devanagari. It is split into one config per licence (CC BY, CC BY-SA, CC0, ODC-By, Apache-2.0, MIT), and every record carries its source, licence and attribution. Text overlapping our evaluation sets is removed.
- sansar-sanskrit-tokenizer: an 8k unigram SentencePiece model over SLP1 transliteration with a Devanagari wrapper. 2.6 to 3.9 tokens per Sanskrit word with only 8k entries.
Licences
The corpus records keep their source licences. Model weights and the tokenizer are released under CC BY-NC 4.0 for non-commercial research use, and the modelling code under Apache-2.0. For other uses, contact us.
Links: muse-mesh.com · GitHub · kushal@muse-mesh.com