spellman — Cyrillic-optimized language detection

This is the runtime model for spellman, a fast Rust language detector covering 30 languages: the big world languages plus twenty Cyrillic-script languages (Yakut, Tuvan, Udmurt, Komi, Mari, Ossetic, Chechen, …) that most detectors don't know exist. It is built for the text that breaks language detectors — tweets, comments, two-word utterances, and close pairs like Russian/Ukrainian — and runs at 3.3 µs per document on one CPU core (0.5 µs across 14 cores; Apple M4 Max).

7.9 MB, one int8 lookup table. The trained network folds algebraically into a single table, stored and scored as int8 with one scale per language column; inference is pure lookups, no matmul.

Usage

The featurizer lives in the Rust runtime, so the model is consumed through the spellman crate or CLI (not a standalone checkpoint):

echo "Съешь ещё этих мягких французских булок" | spellman detect --model hf:vpermilp/spellman   # rus
printf 'Қазақша\nHello\n' | spellman detect --model hf:vpermilp/spellman --lines               # kaz / eng
use spellman_detector::SingleDetector;
let mut det = SingleDetector::from_hub(1024)?;
let d = det.detect("Күн сайын аҕыйах тыл үөрэт")?;   // sah

One storage format. The model ships once, at the repo root, as int8 with per-column scales — the format the runtime scores directly. It matched f16 accuracy on every eval (within ±0.02pp) at half the size and ran faster on every path measured, so the earlier f16/, int8-row/, int8-col/ and fp8-* variants were removed (they remain in this repo's history).

Languages (30)

  • Cyrillic (21): Russian, Ukrainian, Belarusian, Bulgarian, Macedonian, Serbian, Kazakh, Kyrgyz, Tajik, Uzbek (Cyrillic), Tatar, Bashkir, Chuvash, Yakut, Tuvan, Mongolian, Ossetic, Chechen, Udmurt, Meadow Mari, Komi-Zyrian
  • Latin (5): English, Spanish, French, Portuguese, German
  • Script-routed (4, no model columns): Mandarin, Japanese, Hindi, Arabic

Accuracy (v16, 2026-10-04)

eval accuracy
held-out mix (718,647 rows, all registers) 98.74%
held-out mix, texts of ≤20 characters (59,932) 95.06%
Tatoeba (37,051, out-of-domain) 99.08%
single words / pairs / triples (Tatoeba ladder) 72.04 / 89.78 / 95.37
wild Russian tweets (2,606, label-audited) 96.55%
COSMUS gold-labeled wild Russian (2,713, unseen) 99.15%
real short utterances (574, ≤19 chars) 95.82%
literary Russian (2,000 sentences, held-out novel) 98.40%
throughput ~3.3 µs/doc bulk, ~6.7 µs single on one core; ~0.5 µs/doc on 14 cores (Apple M4 Max)

Every accuracy row is measured with this artifact through the Rust runtime; throughput was measured on the v14 artifact, which has the same shape. All referee files are held out of training by construction since v16 (the mixer drops any row whose text appears in one).

Against v14 on identical rows: Ukrainian 96.1 → 98.6% (texts of ≤20 characters 92.4 → 97.6%), other languages mistaken for Russian on short texts 0.85 → 0.25%, short Russian comments from an unseen source 88.1 → 93.2%. Wild Russian tweets (96.78 → 96.55) and literary Russian (98.95 → 98.40) are slightly lower.

On identical rows it beats fastText lid.176 and open-LID SOTA GlotLID v3 on the held-out mix (98.74 vs 81.97 and 92.92), on wild Russian tweets (96.55 vs GlotLID's 82.73), short utterances (95.82 vs 71.25) and single words (72.04 vs 43.9), while running ~100× faster. GlotLID keeps a small lead on clean out-of-domain sentences (99.25 vs 99.08), and on the cleaned COSMUS Russian set all three are within 0.6pp (lid.176 99.37, spellman 99.15, GlotLID 98.82). Full tables and methodology: docs/benchmarks.md.

Files

  • model.json — the runtime contract: language order, hash id/seed, n-gram config, canonicalization/lexical flags, and θ — the confidence threshold under which a detection is flagged uncertain (0.67, calibrated so the flag actually predicts errors).
  • model.safetensors — the folded table P [262145×30] as int8, its per-language-column f32 scales [30], and the class bias.

Features are canonicalized char 1–5-grams plus whole-word and word-bigram keys, sign-hashed into 2^18 buckets (fmix32).

Training data

Trained on a 4.2M-row, 26-language mix published with its byte-exact recipe and per-upstream license table at vpermilp/spellman (dataset repo type). Sources include FineWeb-2 line-windows, Tatoeba, per-language community and parallel corpora, a wild social-media lane, a verified short-utterance lane, and the six open crawl datasets (vpermilp/lid-*: news/library sites + Telegram/VK for sah/tyv/kpv/mhr/oss/udm) built for this project. Since v16 Russian's short texts come from open chat and forum dumps (Telegram dialogues, the Dvach chat, forum messages) instead of a tweet corpus whose ru tag turned out to be mostly Ukrainian. Every source passes LID hygiene; twin languages are arbitrated by spelling and by vocabulary (rows whose letters or words belong to a twin language of their label are dropped). No upstream is NC-licensed (2026-08 license audit); the per-source license table, including sources whose license is the uploader's own label, is on the dataset card.

Version history

  • v16 (2026-10-04): label-noise cleanup — orthographic and lexical twin gates, a 41k-row tweet lane that was mostly Ukrainian under a Russian label retired, short Russian from chat/forum sources instead, referees held out of training by construction, θ 0.67 by error-detection F1. Same architecture and format as v14.
  • 2026-10-01: one storage format — the int8-col table at the root; the f16 / int8-row / int8-col / fp8 variant directories removed. Same v14 weights.
  • v14 (2026-08-31): architecture-review retrain — 2^18 buckets, full-length training truncation, θ by error-detection F1; int8-row root artifact.
  • v13c (2026-08-30): the six crawl datasets + per-language cap 32k → 120k (+3.4pp wild Russian, +4.0pp short vs v12).
  • v12 (2026-08-24): commercial-clean rebuild — every NC-licensed upstream replaced or row-filtered.

The full measured experiment log (including what failed): docs/experiments.md.

Downloads last month
-
Safetensors
Model size
7.86M params
Tensor type
F32
·
F16
·
I8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support