--- license: apache-2.0 language: - uz - en - ru pipeline_tag: text-generation tags: - uzbek - o'zbek - chat - instruction-tuned library_name: transformers --- # MustaqiLLM MustaqiLLM is a 5.17-billion-parameter Uzbek **chat and text-classification model**. It follows Uzbek instructions reliably, writes fluent Uzbek in both Latin and Cyrillic script, and is strong on sentiment and news classification. It is **not** a knowledge model: on multiple-choice knowledge benchmarks it performs at chance. Read the [Evaluation](#evaluation) and [Limitations](#limitations) sections before using it — they are specific about what works and what does not. | | | |---|---| | Parameters | 5.17 B | | Architecture | `NeuronLMForCausalLM` (custom, ships with the repo) | | Layers / hidden | 36 / 3584 | | Attention | GQA, 28 query heads : 4 KV heads, head_dim 128, QK-norm | | Position encoding | RoPE, θ = 500000 | | Context length | 4096 tokens | | Vocabulary | 48,000 (BPE) | | Embeddings | untied | | Weights dtype | bfloat16 (embeddings and `lm_head` stored fp32) | | Languages | Uzbek (Latin + Cyrillic), English, Russian | --- ## Quick start The architecture is custom, so `trust_remote_code=True` is **required** — the modeling code ships inside this repository. ```python import torch from transformers import AutoModelForCausalLM, AutoTokenizer model_id = "NeuronUz/MustaqiLLM" tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True) model = AutoModelForCausalLM.from_pretrained( model_id, trust_remote_code=True, dtype=torch.bfloat16, # weights are bf16; do not load in fp32 device_map="cuda", ).eval() messages = [{"role": "user", "content": "O'zbekistonning poytaxti qaysi shahar?"}] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, return_tensors="pt", return_dict=True, ).to(model.device) with torch.no_grad(): out = model.generate( **inputs, max_new_tokens=256, do_sample=False, # greedy is fine for a short answer like this; # for open chat use the sampling settings below eos_token_id=5, # <|im_end|> -- also the repo default pad_token_id=3, # ) print(tokenizer.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True)) ``` ``` Oʻzbekistonning poytaxti - Toshkent. ``` ### Chat template The model uses ChatML. `tokenizer.apply_chat_template` applies it for you; the raw form is: ``` <|im_start|>system {system}<|im_end|> <|im_start|>user {user}<|im_end|> <|im_start|>assistant {assistant}<|im_end|> ``` A system turn is optional, and for general chat you should leave it out — a generic system prompt measurably increases repetition (see Generation settings). Task-specific system prompts, in Uzbek, work well. ### Generation settings These are measured, not guessed. 35 decoding configurations were swept over 120 held-out Uzbek prompts across 14 categories with 2 seeds each — 8,400 generations — scored automatically for verbatim sentence repetition and for failure to emit `<|im_end|>` within the token budget. Because those metrics see repetition but not fluency, the finalists were then compared head-to-head by an LLM judge over 1,200 pairwise judgements with randomised A/B order. **Recommended for open chat:** ```python out = model.generate( **inputs, max_new_tokens=512, do_sample=True, temperature=0.7, top_p=0.9, repetition_penalty=1.05, # not optional -- see below (1.05-1.10 all work) use_cache=True, ) ``` | setting | value | why | |---|---|---| | `repetition_penalty` | **1.05–1.10** for chat | The single most important setting. Without it the model restates whole sentences verbatim. Duplicate-sentence rate at temperature 0.7: **4.0% at `1.00`, 1.5% at `1.03`, 1.3% at `1.05`, 1.0% at `1.10`, 0.1% at `1.15`.** Do not read that as "higher is better" — see the note below the table. | | `do_sample` / `temperature` / `top_p` | `True`, `0.7`, `0.9` for chat; `False` (greedy) for classification, extraction and short answers | `generation_config.json` ships `do_sample: true` with **no** `temperature` or `top_p`, so the unconfigured default is temperature 1.0 / top_p 1.0 — pass these explicitly. Terse tasks showed a 0% repetition rate under every configuration tested, so greedy is safe there. | | `eos_token_id` | **5** (`<\|im_end\|>`) | The turn terminator, already the default in `config.json` / `generation_config.json` — you do not need to pass it. Do **not** override it with the pretraining EOS (``), which never appears in chat data: generation would then run to `max_new_tokens`. | | system prompt | **omit it** for general chat | A generic system turn measurably degrades output. Duplicate-sentence rate over a 24-prompt subset: **0.0% with no system prompt, 1.9% with a generic Uzbek one, 5.2% with a generic English one** (at temperature 0.7, `repetition_penalty` 1.05); without a repetition penalty the same comparison is 11.2% / 23.7% / 14.9%. Task-specific system prompts (a required format, a persona) are fine — it is the generic "you are a helpful assistant" turn that hurts. | | `dtype` | `torch.bfloat16` | Trained in bf16. `float16` is also safe — no overflow, and output quality is indistinguishable — so pre-Ampere GPUs are supported. `float32` doubles memory for **half** the throughput (205 vs 412 tok/s) and changes nothing. | **More penalty is not better past ~1.10.** The automatic metrics keep improving as `repetition_penalty` rises, but fluency does not. Judged head-to-head on the same prompts, `rp=1.15` — the cleanest configuration by repetition metrics — *lost* to gentler settings: 30.6% win rate against `rp=1.10` and 38.8% against `rp=1.05`. Between 1.05 and 1.10 the judge is a coin flip (52.2%), so anywhere in that band is fine. Below it there is a real floor: `rp=1.05` beats `rp=1.03` at 60.4%. Sampling with a penalty beats greedy outright (58.8%). **Greedy decoding degrades as the output gets longer**, which is why it is recommended above only for short outputs. Over the full 120-prompt sweep at a 384-token budget, greedy produced 17.1% duplicate sentences and failed to terminate on 21.7% of prompts, against 1.3% and 4.2% for `t=0.7, rp=1.05`. On chat and long-form prompts with a 768-token budget the gap widens: | configuration | never emits `<\|im_end\|>` | duplicate sentences | worst case | |---|---:|---:|---:| | greedy | 23.1% | 27.3% | one sentence repeated **9.8×** | | `t=0.7, top_p=0.9` | 15.4% | 8.0% | 2.2× | | `t=0.7, top_p=0.9, rp=1.05` | 11.5% | 3.0% | 1.7× | | **`t=0.7, top_p=0.9, rp=1.10`** | **0.0%** | **0.8%** | **1.1×** | Lowering the temperature makes this worse, not better, because sharpening the distribution locks the model into the repeat loop. Without a repetition penalty, duplicate sentences rise from 1.5% at temperature 0.9 to 8.8% at 0.5; a separate probe at temperature 0.3 reached 19.1%, the worst of any configuration tested. Determinism is genuinely in tension with quality here: greedy plus `repetition_penalty=1.10` still leaves 7.9% duplicate sentences — better than greedy alone, but far short of sampling. If you need reproducible output, sample with a fixed seed rather than decoding greedily. Two categories are much harder than the rest and need a larger `max_new_tokens`: Uzbek **Cyrillic** prompts (37.6% hit the token cap, 12.4% duplicate sentences, pooled across all configurations) and **refusals** (21.8% and 9.2%) — the model has trouble ending a turn once it starts declining a request. Everything else — translation, short answers, grammar and style rewriting, multi-turn — sat at or near 0% on both metrics under every configuration tested. Batch size changes greedy output: identical prompts decoded at batch 1 and batch 12 matched in only 24 of 32 cases, because left-padding shifts the numerics. Fix the batch size when comparing runs. Memory: the checkpoint is 11.0 GB on disk (embeddings and `lm_head` are stored fp32); loading with `dtype=torch.bfloat16` as above casts them down to ~10.3 GB of weights, so a single 16 GB GPU is enough for inference. `config.json` sets `use_cache: false`, but `generation_config.json` sets `use_cache: true`, so `generate()` uses the KV cache. Pass `use_cache=True` explicitly if you write your own decode loop. ### Classification The model is usable as a constrained label picker: put the label set in the prompt, ask for the label only, decode **greedily**, and cap `max_new_tokens`. Terse tasks showed a 0% repetition rate under every decoding configuration tested, so no repetition penalty is needed here — and greedy keeps the output reproducible. These are the exact prompts behind the news (0.6531) and sentiment (0.9259) scores in [Evaluation](#uzbek-benchmarks). Reuse them verbatim to reproduce those numbers. ```python import re import torch from transformers import AutoModelForCausalLM, AutoTokenizer model_id = "NeuronUz/MustaqiLLM" tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True) model = AutoModelForCausalLM.from_pretrained( model_id, trust_remote_code=True, dtype=torch.bfloat16, device_map="cuda", ).eval() def classify(prompt: str, text: str, max_chars: int = 4000) -> str: if len(text) > max_chars: text = text[:max_chars].rsplit(" ", 1)[0] inputs = tokenizer.apply_chat_template( [{"role": "user", "content": prompt.format(text=text)}], add_generation_prompt=True, return_tensors="pt", return_dict=True, ).to(model.device) with torch.no_grad(): out = model.generate( **inputs, max_new_tokens=12, # a label is a few tokens; do not give it room to ramble do_sample=False, # greedy -- labels must be deterministic pad_token_id=3, # ) return tokenizer.decode( out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True ).strip() ``` **News topic, 10-way.** Numbered labels: one digit is easier to emit and to parse than a multi-word category name. ```python NEWS_LABELS = [ "Siyosat", "Iqtisodiyot", "Texnologiya", "Sport", "Madaniyat", "Salomatlik", "Oila va Jamiyat", "Ta'lim", "Ekologiya", "Xorijiy Yangiliklar", ] NEWS_PROMPT = ( "Classify the given Uzbek news article into one of the following categories. " "Respond with only the category number.\n\n" + "".join(f"{i} - {name}\n" for i, name in enumerate(NEWS_LABELS)) + "\nArticle: {text}\n\nAnswer:" ) raw = classify(NEWS_PROMPT, "O'zbekiston Markaziy banki asosiy stavkani o'zgarishsiz qoldirdi.") match = re.search(r"\d+", raw) label = NEWS_LABELS[int(match.group())] if match and int(match.group()) < 10 else None print(raw, "->", label) ``` ``` 1 -> Iqtisodiyot ``` **Sentiment, binary.** ```python SENTIMENT_PROMPT = ( "Given the following Uzbek text, determine the sentiment as either " "'Positive' or 'Negative'. Respond with only one label.\n\n" "Text: {text}\n\nLabel:" ) raw = classify(SENTIMENT_PROMPT, "Mahsulot juda sifatli, yetkazib berish tez bo'ldi.") print(raw) # Positive ``` **Your own label set.** The same shape works for any closed label set — put one label per line, demand the label (or its number) and nothing else, and parse the output with a prefix match or a regex rather than an exact-string comparison, so a stray token never becomes an invalid prediction. Two practical notes: - **A task-specific system prompt is fine here** and often helps — it is the generic "you are a helpful assistant" turn that degrades output (see [Generation settings](#generation-settings)). Put the required output format in it. - **English prompt text with Uzbek labels** is what was measured. Uzbek prompt wording also works; if you change the wording, re-measure — label boundaries (especially `Siyosat` vs `Xorijiy Yangiliklar`, and `Oila va Jamiyat`, the weakest class at 0.4273) are sensitive to how the categories are described. - **Do not batch-compare greedy runs at different batch sizes.** Left-padding shifts the numerics; identical prompts matched in only 24 of 32 cases between batch 1 and batch 12. ### Serving **vLLM and SGLang cannot load this model.** They reimplement each architecture internally rather than executing a repository's Python, and `NeuronLMForCausalLM` is not in their model registries — `trust_remote_code` only covers the config and tokenizer there. Use the `transformers` backend, or convert the weights (the architecture is Qwen3-equivalent apart from *fused* `qkv_proj` / `gate_up_proj` and `out_proj` naming; splitting those tensors and renaming to the Qwen3 layout yields a checkpoint vLLM will serve). --- ## Evaluation Full public benchmark suite, greedy decoding, `transformers` backend, seed 42, complete test sets (no subsampling). Scores are accuracy unless noted. ### Uzbek benchmarks | benchmark | n | score | invalid rate | |---|---:|---:|---:| | uzlib (Uzbek linguistic MCQ) | 1,861 | 0.2875 | 0.0000 | | TUMLU-Uzbek (Uzbek MMLU) | 700 | 0.3286 | 0.0000 | | MMLU-Uz (translated MMLU) | 14,042 | 0.2584 | 0.0000 | | News topic classification (10-way, `risqaliyevds/uzbek-zero-shot-classification`) | 96,970 | **0.6531** | 0.0000 | | Sentiment (binary) | 10,000 | **0.9259** | 0.0001 | Random baselines: 0.25 for the 4-way MCQ tasks, 0.10 for news, 0.50 for sentiment. ### English | benchmark | n | score | invalid rate | |---|---:|---:|---:| | MMLU (English) | 14,042 | 0.2619 | 0.0000 | ### Translation (FLORES+) | direction | n | BLEU | COMET | length ratio | |---|---:|---:|---:|---:| | English → Uzbek | 2,009 | 5.17 | 0.7397 | 1.018 | | Uzbek → English | 2,009 | 1.83 | 0.5376 | 1.229 | ### uzlib, per split | split | n | score | |---|---:|---:| | fill_in | 52 | 0.3077 | | correct_word (orthography) | 1,501 | 0.3011 | | meaning_in_context | 72 | 0.2639 | | meaning | 236 | 0.2034 | ### News, per class | class | n | score | |---|---:|---:| | Sport | 16,113 | 0.8743 | | Texnologiya (Technology) | 5,177 | 0.7309 | | Madaniyat (Culture) | 2,405 | 0.7081 | | Siyosat (Politics) | 29,500 | 0.6794 | | Iqtisodiyot (Economy) | 10,755 | 0.6596 | | Salomatlik (Health) | 3,505 | 0.6579 | | Ta'lim (Education) | 1,987 | 0.6548 | | Ekologiya (Ecology) | 1,784 | 0.5667 | | Xorijiy Yangiliklar (World news) | 11,732 | 0.5124 | | Oila va Jamiyat (Family & Society) | 14,012 | 0.4273 | ## Limitations - **MCQ knowledge tasks are at chance.** uzlib, MMLU-Uz and MMLU-English all sit within noise of the 0.25 baseline over ~30,000 questions, with near-zero invalid rates — correct format, wrong answer. This is missing knowledge, not parsing. Do not use it for factual QA, exams, or retrieval-free knowledge tasks. TUMLU-Uzbek (0.3286) is the only MCQ result above chance, on a 700-item sample (±3.5%). - **Uzbek → English translation is weak** (BLEU 1.83, length ratio 1.229): it over-generates. English → Uzbek is usable (COMET 0.7397) but below dedicated MT systems. - **Script conversion does not work** despite being trained for it — Latin→Cyrillic requests often return the input unchanged. - **Cyrillic artifacts.** The Cyrillic data was machine-transliterated; loanwords and brand names can be mangled (`Facebook` → `Факебоок`) and stray Cyrillic characters leak into Latin words. Cyrillic chat is coherent, but its orthography is less reliable than Latin. - **Self-identification.** Identity data predates the current name, so the model calls itself "NeuronAI 5B". - **Uneven news classification:** 0.4273 on the diffuse "Oila va Jamiyat" class vs 0.8743 on Sport. - **Safety.** No safety alignment, RLHF, or red-teaming; no refusal training beyond what the instruction data incidentally contains. It can produce incorrect, biased, or unsafe content and will state false facts fluently. Evaluate before any user-facing deployment. --- ## Intended use **Suitable for:** Uzbek-language chat and assistance; text classification (sentiment, topic); Uzbek text generation and rewriting in Latin or Cyrillic; English → Uzbek translation where approximate meaning suffices; a base for further fine-tuning. **Not suitable for:** factual question answering or anything knowledge-intensive; exam-style multiple choice; Uzbek → English translation; script transliteration; any application where a confidently-stated wrong fact causes harm (medical, legal, financial advice). ## License Apache 2.0. Training data licensing follows the sources of the underlying public datasets. ## Citation ```bibtex @misc{mustaqillm, title = {MustaqiLLM: an instruction-tuned Uzbek language model}, author = {NeuronUz}, year = {2026}, url = {https://huggingface.co/NeuronUz/MustaqiLLM} } ```