metadata
title: BuzzASR
emoji: π
colorFrom: yellow
colorTo: red
sdk: static
pinned: false
license: mit
π BuzzASR β A Swarm of 100+ Monolingual Speech Recognition Models
One specialist model per language. BuzzASR is a suite of 102 monolingual ASR models, each a fine-tune of Whisper-large-v3 specialized to a single language.
The work has been accepted at EMNLP 2026 Β· built by Lemn Lab in collaboration with EleutherAI.
The idea
Big multilingual models spread themselves thin across hundreds of languages. We asked whether a single specialist, the same size as Whisper, could beat the giant generalists. It can to a considerable extent :)
Highlights
- β Beats Whisper-large-v3 on 89 of 102 languages (FLEURS)
- π Open-source state-of-the-art (beats Whisper, Omnilingual 1B/7B, MMS, Qwen3-ASR, Cohere) on 27 languages
- π Cuts character error rate ~3x on average, and 3.45x across Whisper's 51 worst languages
- π§© Each model is the same 1.55B parameters as Whisper β about 4.5x smaller than the strongest baseline (Omni-7B)
A few standouts (combined FLEURS + Common Voice, normalized CER %)
| Language | BuzzASR CER | Whisper zero-shot | Note |
|---|---|---|---|
| Cantonese | 13.0 | 36.4 | SOTA β 2x better than the next-best system |
| Punjabi | 8.8 | 42.3 | SOTA |
| Mongolian | 5.2 | 38.2 | SOTA |
| Korean | 4.6 | 5.7 | SOTA |
| Amharic | 9.1 | 191 | >20x reduction over Whisper |
Find your language
Browse all 102 models in the Models tab above, or go to huggingface.co/BuzzASR/<language>.
Usage
from transformers import WhisperForConditionalGeneration, WhisperProcessor
model = WhisperForConditionalGeneration.from_pretrained("BuzzASR/mongolian")
proc = WhisperProcessor.from_pretrained("BuzzASR/mongolian")
# the language/task prompt is baked in β just call model.generate(input_features)
Links
- π Paper β Findings of EMNLP 2026 (https://arxiv.org/abs/2609.09554)
- π Project & docs β https://lemn-lab.github.io/buzz-asr/