NeuronAI-2B / fertility_large_20260825.md
kmamaroziqov's picture
Match NeuronAI-4B card format with 2B evidence and Apache license
f99a82c verified
|
Raw
History Blame Contribute Delete
1.53 kB
# NeuronAI-2B large tokenizer fertility evaluation
Generated: 2026-08-24T19:25:09+00:00
Fertility is the number of tokenizer tokens per whitespace-delimited word; lower is better.
The confidence interval is a paired 5,000-replicate bootstrap over 1,000 deterministic document buckets.
| Corpus | Domain status | Documents | Words | NeuronAI-2B | Qwen3.5-2B | Token reduction (95% CI) |
| --- | --- | ---: | ---: | ---: | ---: | ---: |
| OSCAR Uzbek | held-out web source | 100,000 | 7,618,770 | 2.0304 | 3.3639 | 39.64% (39.57%–39.71%) |
| Uzbek legal corpus | held-out legal source/domain | 18,832 | 2,534,566 | 2.3747 | 2.9705 | 20.06% (19.55%–20.57%) |
| uz-crawl control | training-source control | 100,000 | 20,825,680 | 2.3206 | 3.3224 | 30.15% (30.02%–30.30%) |
## Interpretation
- The 2B and 4B evaluated tokenizer pairs are byte-identical, so their token counts are exactly identical. SHA-256 fingerprints are recorded in the JSON result.
- The held-out-source corpora are the fairness readout; the training-source corpus is retained only as a control.
- Exact web overlap cannot be ruled out for OSCAR because both corpora ultimately derive from public web data.
- OSCAR was checked in a post-hoc weak-token coverage analysis but contributed no tokenizer-training rows.
- The legal corpus is absent from the tokenizer and CPT source manifests and is the cleanest source-and-domain holdout in this run.
- Results use raw source text without apostrophe normalization or other model-specific preprocessing.