Instructions to use vpermilp/spellman with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- fastText
How to use vpermilp/spellman with fastText:
from huggingface_hub import hf_hub_download import fasttext model = fasttext.load_model(hf_hub_download("vpermilp/spellman", "model.bin")) - Notebooks
- Google Colab
- Kaggle
spellman — Cyrillic-optimized language detection
This is the runtime model for spellman,
a fast Rust language detector covering 30 languages: the big world
languages plus twenty Cyrillic-script languages (Yakut, Tuvan, Udmurt,
Komi, Mari, Ossetic, Chechen, …) that most detectors don't know exist.
It is built for the text that breaks language detectors — tweets,
comments, two-word utterances, and close pairs like Russian/Ukrainian —
and runs at 3.3 µs per document on one CPU core (0.5 µs across 14
cores; Apple M4 Max).
7.9 MB, one int8 lookup table. The trained network folds algebraically into a single table, stored and scored as int8 with one scale per language column; inference is pure lookups, no matmul.
Usage
The featurizer lives in the Rust runtime, so the model is consumed through the spellman crate or CLI (not a standalone checkpoint):
echo "Съешь ещё этих мягких французских булок" | spellman detect --model hf:vpermilp/spellman # rus
printf 'Қазақша\nHello\n' | spellman detect --model hf:vpermilp/spellman --lines # kaz / eng
use spellman_detector::SingleDetector;
let mut det = SingleDetector::from_hub(1024)?;
let d = det.detect("Күн сайын аҕыйах тыл үөрэт")?; // sah
One storage format. The model ships once, at the repo root, as int8
with per-column scales — the format the runtime scores directly. It
matched f16 accuracy on every eval (within ±0.02pp) at half the size and
ran faster on every path measured, so the earlier f16/, int8-row/,
int8-col/ and fp8-* variants were removed (they remain in this repo's
history).
Languages (30)
- Cyrillic (21): Russian, Ukrainian, Belarusian, Bulgarian, Macedonian, Serbian, Kazakh, Kyrgyz, Tajik, Uzbek (Cyrillic), Tatar, Bashkir, Chuvash, Yakut, Tuvan, Mongolian, Ossetic, Chechen, Udmurt, Meadow Mari, Komi-Zyrian
- Latin (5): English, Spanish, French, Portuguese, German
- Script-routed (4, no model columns): Mandarin, Japanese, Hindi, Arabic
Accuracy (v16, 2026-10-04)
| eval | accuracy |
|---|---|
| held-out mix (718,647 rows, all registers) | 98.74% |
| held-out mix, texts of ≤20 characters (59,932) | 95.06% |
| Tatoeba (37,051, out-of-domain) | 99.08% |
| single words / pairs / triples (Tatoeba ladder) | 72.04 / 89.78 / 95.37 |
| wild Russian tweets (2,606, label-audited) | 96.55% |
| COSMUS gold-labeled wild Russian (2,713, unseen) | 99.15% |
| real short utterances (574, ≤19 chars) | 95.82% |
| literary Russian (2,000 sentences, held-out novel) | 98.40% |
| throughput | ~3.3 µs/doc bulk, ~6.7 µs single on one core; ~0.5 µs/doc on 14 cores (Apple M4 Max) |
Every accuracy row is measured with this artifact through the Rust runtime; throughput was measured on the v14 artifact, which has the same shape. All referee files are held out of training by construction since v16 (the mixer drops any row whose text appears in one).
Against v14 on identical rows: Ukrainian 96.1 → 98.6% (texts of ≤20 characters 92.4 → 97.6%), other languages mistaken for Russian on short texts 0.85 → 0.25%, short Russian comments from an unseen source 88.1 → 93.2%. Wild Russian tweets (96.78 → 96.55) and literary Russian (98.95 → 98.40) are slightly lower.
On identical rows it beats fastText lid.176 and open-LID SOTA GlotLID v3 on the held-out mix (98.74 vs 81.97 and 92.92), on wild Russian tweets (96.55 vs GlotLID's 82.73), short utterances (95.82 vs 71.25) and single words (72.04 vs 43.9), while running ~100× faster. GlotLID keeps a small lead on clean out-of-domain sentences (99.25 vs 99.08), and on the cleaned COSMUS Russian set all three are within 0.6pp (lid.176 99.37, spellman 99.15, GlotLID 98.82). Full tables and methodology: docs/benchmarks.md.
Files
model.json— the runtime contract: language order, hash id/seed, n-gram config, canonicalization/lexical flags, and θ — the confidence threshold under which a detection is flagged uncertain (0.67, calibrated so the flag actually predicts errors).model.safetensors— the folded tableP[262145×30]as int8, its per-language-column f32scales[30], and the classbias.
Features are canonicalized char 1–5-grams plus whole-word and word-bigram keys, sign-hashed into 2^18 buckets (fmix32).
Training data
Trained on a 4.2M-row, 26-language mix published with its byte-exact
recipe and per-upstream license table at
vpermilp/spellman
(dataset repo type). Sources include FineWeb-2 line-windows, Tatoeba,
per-language community and parallel corpora, a wild social-media lane,
a verified short-utterance lane, and the six open crawl datasets
(vpermilp/lid-*: news/library sites + Telegram/VK for
sah/tyv/kpv/mhr/oss/udm) built for this project. Since v16 Russian's
short texts come from open chat and forum dumps (Telegram dialogues, the
Dvach chat, forum messages) instead of a tweet corpus whose ru tag
turned out to be mostly Ukrainian. Every source passes LID hygiene; twin
languages are arbitrated by spelling and by vocabulary (rows whose
letters or words belong to a twin language of their label are dropped).
No upstream is NC-licensed (2026-08 license audit); the per-source
license table, including sources whose license is the uploader's own
label, is on the dataset card.
Version history
- v16 (2026-10-04): label-noise cleanup — orthographic and lexical twin gates, a 41k-row tweet lane that was mostly Ukrainian under a Russian label retired, short Russian from chat/forum sources instead, referees held out of training by construction, θ 0.67 by error-detection F1. Same architecture and format as v14.
- 2026-10-01: one storage format — the int8-col table at the root; the f16 / int8-row / int8-col / fp8 variant directories removed. Same v14 weights.
- v14 (2026-08-31): architecture-review retrain — 2^18 buckets, full-length training truncation, θ by error-detection F1; int8-row root artifact.
- v13c (2026-08-30): the six crawl datasets + per-language cap 32k → 120k (+3.4pp wild Russian, +4.0pp short vs v12).
- v12 (2026-08-24): commercial-clean rebuild — every NC-licensed upstream replaced or row-filtered.
The full measured experiment log (including what failed): docs/experiments.md.
- Downloads last month
- -