NeuronAI-2B / README.md
kmamaroziqov's picture
Match NeuronAI-4B card format with 2B evidence and Apache license
f99a82c verified
|
Raw
History Blame Contribute Delete
11.1 kB
---
language:
- uz
- en
license: apache-2.0
library_name: transformers
pipeline_tag: text-generation
base_model: Qwen/Qwen3.5-2B-Base
tags:
- qwen3.5
- uzbek
- conversational
- translation
- text-generation-inference
datasets:
- HuggingFaceFW/fineweb-2
- tahrirchi/uz-books
- tahrirchi/uz-crawl
- HuggingFaceFW/fineweb-edu
- HuggingFaceTB/finemath
---
# NeuronAI-2B
**NeuronAI-2B** is an Uzbek-first, bilingual assistant model built from
Qwen3.5-2B-Base. It combines an Uzbek tokenizer retrofit, continued pretraining,
annealing, and assistant-only supervised fine-tuning. The published weights are
fully merged—no LoRA adapter is needed.
![Strict eight-task benchmark comparison](assets/overall_score.png)
> **License:** [Apache License 2.0](https://www.apache.org/licenses/LICENSE-2.0).
> Commercial and non-commercial use are permitted under the license terms. This
> differs from the NeuronAI-4B release, which is licensed for non-commercial use.
## Quick start
Install a recent Transformers build with Qwen3.5 support:
```bash
pip install -U "transformers>=5.1" accelerate torch
```
```python
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "NeuronUz/NeuronAI-2B"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
dtype=torch.bfloat16,
device_map={"": 0},
).eval()
messages = [
{"role": "system", "content": "Siz foydali va aniq AI yordamchisiz."},
{"role": "user", "content": "Alisher Navoiy haqida qisqacha aytib bering."},
]
inputs = tokenizer.apply_chat_template(
messages,
add_generation_prompt=True,
enable_thinking=False,
return_tensors="pt",
return_dict=True,
).to(model.device)
with torch.inference_mode():
output = model.generate(
**inputs,
max_new_tokens=1024,
do_sample=True,
temperature=0.7,
top_p=0.8,
top_k=20,
min_p=0.0,
repetition_penalty=1.0,
use_cache=True,
)
reply = tokenizer.decode(
output[0, inputs["input_ids"].shape[1]:],
skip_special_tokens=True,
).strip()
print(reply)
```
This is the recommended quality-oriented preset for general assistant use:
non-thinking mode with Qwen3.5's instruct sampling settings. Greedy decoding
can cause repetition and lower response quality; reserve `do_sample=False` for
deterministic evaluation or classification. The generation metadata already
registers `<|im_end|>` and `<|endoftext|>` as end-of-sequence tokens. Keep the
combined prompt and output within the validated 4,096-token serving limit.
### Serve with vLLM
```bash
pip install -U vllm
vllm serve NeuronUz/NeuronAI-2B \
--dtype bfloat16 \
--max-model-len 4096 \
--tensor-parallel-size 1 \
--generation-config vllm \
--default-chat-template-kwargs '{"enable_thinking":false}' \
--language-model-only \
--enable-prefix-caching \
--mamba-block-size 16 \
--mamba-cache-mode align
```
```bash
curl http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "NeuronUz/NeuronAI-2B",
"messages": [
{"role": "user", "content": "O‘zbekiston haqida uchta fakt ayting."}
],
"max_tokens": 1024,
"temperature": 0.7,
"top_p": 0.8,
"top_k": 20,
"min_p": 0.0,
"presence_penalty": 1.5,
"repetition_penalty": 1.0,
"chat_template_kwargs": {"enable_thinking": false}
}'
```
## Benchmarks
All five model result sets below cover the same full eight-task suite.
Classification and multiple-choice tasks use accuracy; FLORES+ translation
uses COMET. The weighted score is normalized by the 0.95 sum of the published
task weights. All eight NeuronAI-2B tasks completed and passed the
invalid-output gate.
![Per-task comparison](assets/tasks_comparison.png)
| Benchmark | Metric | Weight | **NeuronAI-2B** | Qwen3.5-2B | alloma-8B | alloma-3B | alloma-1B |
| --- | --- | ---: | ---: | ---: | ---: | ---: | ---: |
| UzLiB | accuracy | 0.20 | **49.60%** | 28.69% | 42.40% | 32.08% | 23.32% |
| TUMLU-Uzbek | accuracy | 0.20 | **32.57%** | 31.29% | 20.71% | 27.71% | 22.00% |
| FLORES+ en→uz | COMET | 0.15 | 0.8762 | 0.7010 | **0.8779** | 0.8673 | 0.7383 |
| Uzbek news | accuracy | 0.10 | **78.55%** | 36.75% | 57.77% | 13.60% | 25.41% |
| MMLU English | accuracy | 0.10 | **54.07%** | 52.39% | 53.47% | 38.73% | 21.98% |
| MMLU Uzbek | accuracy | 0.10 | **46.85%** | 37.10% | 40.04% | 32.74% | 21.11% |
| FLORES+ uz→en | COMET | 0.05 | 0.8535 | 0.8072 | **0.8713** | 0.7954 | 0.7636 |
| Uzbek sentiment | accuracy | 0.05 | **95.50%** | 76.87% | 79.94% | 38.85% | 79.54% |
| **Normalized weighted score** | | 1.00 | **0.5954** | 0.4528 | 0.5187 | 0.4147 | 0.3661 |
Alloma runs used the `APST` apostrophe preprocessing required by their model
cards; NeuronAI and stock Qwen did not. The alloma-8B column combines its full
model-card-protocol evaluation with separately archived full UzLiB,
TUMLU-Uzbek, and MMLU-Uzbek runs. Exact source files, scores, and run IDs are
included in [`benchmark_results.json`](benchmark_results.json).
### Run the benchmarks on your computer
The repository includes a portable Alloma-style benchmark runner. It covers
FLORES+ (both directions), Uzbek sentiment, Uzbek news, MMLU English, MMLU Uzbek,
and TUMLU-Uzbek.
```bash
pip install -r https://huggingface.co/NeuronUz/NeuronAI-2B/resolve/main/benchmark-requirements.txt
wget https://huggingface.co/NeuronUz/NeuronAI-2B/resolve/main/benchmark.py
python benchmark.py --limit 200 --output quick-results.json
```
The quick command uses the same seed on 200 examples per dataset. Run all public
examples and add COMET with:
```bash
pip install unbabel-comet
python benchmark.py --limit 0 --comet --output full-results.json
```
Run one task when you only need a short check:
```bash
python benchmark.py --tasks mmlu-uz --limit 200 --output mmlu-uz.json
python benchmark.py --tasks flores --limit 200 --output flores.json
```
`--limit 0` means the full dataset. Only full runs are comparable with the table
above; 200-example quick runs are sanity checks. COMET downloads the
`Unbabel/wmt22-comet-da` evaluator and needs additional disk/RAM.
## Uzbek tokenizer efficiency
The tokenizer is an in-place, primarily **Latin-script Uzbek** retrofit rather
than a vocabulary extension. The initial 20,000-document figure was measured on
training-source `uz-crawl`, so we replaced it with a larger corpus-stratified
test: 118,832 held-out-source documents plus a separate 100,000-document
training-source control. Documents were selected with deterministic SHA-256
bottom-k sampling (seed `20260825`), exact duplicates were excluded from the
selected sample, tiny texts were filtered, and raw source text was tokenized
without apostrophe normalization.
![Uzbek tokenizer fertility](assets/tokenizer_fertility.png)
| Corpus | Status | Documents | Words | NeuronAI-2B | Qwen3.5-2B | Reduction (95% CI) |
| --- | --- | ---: | ---: | ---: | ---: | ---: |
| Community OSCAR Uzbek | Held-out web source | 100,000 | 7,618,770 | **2.0304** | 3.3639 | **39.64%** (39.57–39.71%) |
| Uzbek legal corpus | Held-out legal source/domain | 18,832 | 2,534,566 | **2.3747** | 2.9705 | **20.06%** (19.55–20.57%) |
| uz-crawl | Training-source control | 100,000 | 20,825,680 | **2.3206** | 3.3224 | **30.15%** (30.02–30.30%) |
Across the two held-out sources combined, the tokenizer uses **35.19%
fewer tokens overall** and **40.90% fewer tokens on Latin-dominant text**,
matching its intended Latin-Uzbek focus.
The paired intervals use 5,000 bootstrap replicates over 1,000 deterministic
document buckets. OSCAR may still have incidental overlap with other public web
corpora and was previously checked in a post-hoc weak-token coverage analysis,
but it contributed no tokenizer-training rows. The legal corpus does not appear
in the tokenizer or training source manifests and is the cleanest
source-and-domain holdout in this test. Full results and script/length
breakdowns: [`fertility_large_20260825.json`](fertility_large_20260825.json) and
[`fertility_large_20260825.md`](fertility_large_20260825.md).
Fertility measures tokenization efficiency—not model quality or measured
decoding speed. The evaluated 2B and 4B custom tokenizer files are byte-identical,
as are their evaluated stock-base tokenizer files; SHA-256 fingerprints are
recorded in the JSON result.
## Training
| Item | Value |
| --- | --- |
| Parameters | 1,881,825,088 (1.882B) |
| Prepared train examples | 151,968 (152,152 source rows) |
| Prepared grouped dev examples | 1,535 (1,537 source rows) |
| Train/dev prompt-group overlap | 0 |
| Sequence length / packing | 2,048 / disabled |
| Training duration / seed | 1 epoch / 42 |
| Batch size | 16 micro × 2 accumulation × 1 GPU = 32 effective |
| Optimizer | Fused AdamW; betas 0.9/0.95; weight decay 0.01; gradient clipping 1.0 |
| Learning-rate schedule | Peak 1e-4; cosine decay; 142 warmup steps (2.99%) |
| LoRA | rank 64, alpha 128, dropout 0.05; 12 projection types; 67,276,800 trainable parameters |
| Loss | Fused causal-LM cross-entropy on assistant-response tokens; prompt tokens masked |
| Precision | bf16 training with TF32; merged embeddings and normalization tensors retained in fp32 |
The mixture is Uzbek-first and includes general assistant conversations,
translation, Uzbek language and literature, spelling, classification, math,
and English-retention examples. Training data is not distributed with this
model repository.
## Intended use
Good fits include Uzbek research, education, commercial and non-commercial
prototyping, translation experiments, writing assistance, retrieval-augmented
generation, and local/offline applications. Users remain responsible for
validating the model for their application and complying with the Apache 2.0
license and applicable law.
## Limitations
- This is a public-suite-selected checkpoint. The benchmark results are useful
for reproducibility and relative comparison, but they are not a locked,
independent estimate of real-world generalization.
- LoRA rank, learning rate, batch size, and dropout were not exhaustively swept;
the table reports the released run, not globally optimal hyperparameters.
- TUMLU-Uzbek is the weakest reported Uzbek task and should not be treated as
solved at 32.57% accuracy.
- The model can hallucinate, repeat biases in its data, or produce unsafe or
outdated content. It has not been comprehensively safety-evaluated.
- Do not rely on it without expert review for medical, legal, financial, public
safety, or other high-stakes decisions.
- SFT used sequences up to 2,048 tokens; serving at longer inherited context
lengths has not been validated here. The published inference examples use
4,096 tokens.
## License
NeuronAI-2B is released under the
[Apache License 2.0](https://www.apache.org/licenses/LICENSE-2.0). Commercial
and non-commercial use, modification, and distribution are permitted subject
to its terms. This summary does not replace the license text; see
[`LICENSE`](LICENSE).