Bangla Zipformer - ONNX (sherpa-onnx)
Bangla speech recognition with sherpa-onnx: no custom code, no PyTorch. ONNX export of a 65.5M-parameter Zipformer transducer, fine-tuned from an English LibriSpeech model on about 1,210 hours of Bangla speech.
Three precisions (fp32, fp16, int8), two decoding modes (beam search, greedy), CPU or GPU. PyTorch version: SayedShaun/bangla-zipformer-pt.
| Test WER / CER | 14.4% / 4.3% (fp32, beam search); int8 14.4% / 4.3% |
| Dev WER / CER | 14.2% / 4.3% |
| Model | Zipformer2 transducer, non-streaming, 65.5M parameters (average of epochs 23-25) |
| Size | int8 71 MB, fp16 129 MB, fp32 254 MB |
| Speed | CPU, int8: 62 ms per clip with beam search (46 ms greedy), 4 threads, clips of 1-15 s. GPU, fp32: 29 ms |
| Install | pip install sherpa-onnx (CPU) or the CUDA 12 wheel (GPU), see below |
| Input | mono audio of any sample rate, up to ~20 s per call (split longer audio first) |
| License | CC BY-SA 4.0 |
This is the recommended way to use the model. Install is one
pip install sherpa-onnx: no PyTorch, no icefall, no k2, no custom code, and it also runs on a GPU. For fine-tuning or inspecting the network, use the PyTorch version instead.
Quick start (CPU)
pip install sherpa-onnx numpy soundfile huggingface_hub
import sherpa_onnx
import soundfile as sf
from huggingface_hub import snapshot_download
d = snapshot_download("SayedShaun/bangla-zipformer-onnx", allow_patterns=["int8/*", "tokens.txt"]) # "fp32" | "fp16" | "int8"
asr = sherpa_onnx.OfflineRecognizer.from_transducer(
encoder=f"{d}/int8/encoder.onnx", decoder=f"{d}/int8/decoder.onnx", joiner=f"{d}/int8/joiner.onnx",
tokens=f"{d}/tokens.txt", num_threads=4, sample_rate=16000, feature_dim=80,
decoding_method="modified_beam_search", max_active_paths=4, # or decoding_method="greedy_search" (faster)
)
audio, sr = sf.read("audio.wav", dtype="float32") # any sample rate: sherpa-onnx resamples
if audio.ndim > 1:
audio = audio.mean(axis=1) # stereo -> mono
stream = asr.create_stream()
stream.accept_waveform(sr, audio)
asr.decode_stream(stream)
print(stream.result.text)
The files follow sherpa-onnx's standard transducer layout. sherpa-onnx also has C++, Java, C#, Go, Swift and Android / iOS / WebAssembly APIs; only the Python API was tested for this card.
soundfile and huggingface_hub are only used above to read audio and download the files; you can use any equivalent (for example huggingface-cli download).
Quick start (GPU)
pip install "sherpa-onnx==1.13.4+cuda12.cudnn9" -f https://k2-fsa.github.io/sherpa/onnx/cuda.html # needs CUDA 12 and cuDNN 9
Use the same code with the fp32 files and provider="cuda" added to from_transducer(...). Use fp32 on the GPU: fp16 and int8 are slower there (table below).
The CUDA 12 and cuDNN 9 libraries must be findable (for example on LD_LIBRARY_PATH).
Which version should I use?
| You want | Use | Why |
|---|---|---|
| The default on a CPU | int8 + beam search |
Smallest (71 MB) and fastest; within 0.12 WER points of fp32 (62 ms per clip) |
| Lowest latency on a CPU | int8 + greedy search |
46 ms per clip, about 0.2-0.3 WER points worse than beam |
| Best accuracy on a CPU | fp32 + beam search |
Reference accuracy; fp16 is identical but gives no speed-up on CPU |
| A GPU | fp32 + beam search, provider="cuda" |
29 ms per clip; fp16 and int8 are slower on the GPU |
Results
Full dev (5,282 utterances) and test (10,485 utterances) sets, run with sherpa-onnx 1.13.8 on CPU.
Beam search (modified_beam_search, 4 paths)
| Version | Size | Test WER | Test CER | Dev WER | Dev CER |
|---|---|---|---|---|---|
| sherpa-onnx fp32 | 254 MB | 14.40 | 4.32 | 14.18 | 4.30 |
| sherpa-onnx fp16 | 129 MB | 14.39 | 4.32 | 14.19 | 4.30 |
| sherpa-onnx int8 | 71 MB | 14.44 | 4.32 | 14.30 | 4.31 |
| PyTorch fp32 (reference) | 262 MB | 14.37 | 4.30 | 14.17 | 4.29 |
Greedy search (greedy_search)
| Version | Size | Test WER | Test CER | Dev WER | Dev CER |
|---|---|---|---|---|---|
| sherpa-onnx fp32 | 254 MB | 14.70 | 4.47 | 14.47 | 4.41 |
| sherpa-onnx fp16 | 129 MB | 14.70 | 4.47 | 14.48 | 4.41 |
| sherpa-onnx int8 | 71 MB | 14.73 | 4.49 | 14.52 | 4.43 |
| PyTorch fp32 (reference) | 262 MB | 14.67 | 4.46 | 14.42 | 4.40 |
WER counts punctuation attached to words; CER ignores punctuation. Precision matters little for accuracy: fp16 is identical to fp32 and int8 is within 0.12 WER points. The PyTorch reference rows use its own decoder, so they differ from sherpa-onnx by a few hundredths of a point.
WER with punctuation stripped
| Version | Test, beam | Dev, beam | Test, greedy | Dev, greedy |
|---|---|---|---|---|
| fp32 | 13.85 | 13.69 | 14.13 | 13.99 |
| fp16 | 13.84 | 13.70 | 14.13 | 14.00 |
| int8 | 13.89 | 13.81 | 14.17 | 14.02 |
Speed
Mean time to transcribe one clip (60 dev clips of 1-15 s: 20 each of short, medium and long). RTF = processing time / audio length (lower is faster; 0.01 means 100x faster than real time).
CPU, 4 threads (32-core machine)
| Version | Beam: ms / clip | Greedy: ms / clip | Beam: RTF | RAM |
|---|---|---|---|---|
| fp32 | 77 ms | 62 ms | 0.0129 | 1121 MB |
| fp16 | 69 ms | 63 ms | 0.0116 | 1181 MB |
| int8 | 62 ms | 46 ms | 0.0104 | 733 MB |
- int8 is the fastest and smallest (about 25% faster than fp32 in greedy mode, and a third less memory). fp16 gives no speed-up on CPU.
- Beam search costs 6-16 ms more per clip than greedy (the spread is mostly run-to-run noise); use greedy if latency matters more than the last 0.3 WER points.
GPU (RTX 5080, sherpa-onnx 1.13.4 CUDA build, 4 CPU threads for the CPU-side work)
| Version | Beam: ms / clip | Greedy: ms / clip | GPU memory |
|---|---|---|---|
| fp32 (recommended) | 29 ms | 20 ms | 2.4 GB |
| fp16 (slow on GPU) | 259 ms | 242 ms | 1.9 GB |
| int8 (partly CPU) | 56 ms | 42 ms | 1.9 GB |
- fp32 on the GPU is about 2.6x faster than on the CPU (29 ms vs 77 ms with beam search). fp16 is much slower on the GPU and int8 falls back to the CPU for some operators, so use fp32 there.
- The GPU build holds about 2.4 GB of GPU memory per process.
Method: 5 repeats per clip after warm-up, beam and greedy run back to back on the same clip, per-clip median, audio passed as an array (file reading not counted). The CPU 4-thread and GPU fp32 rows are medians of 3 runs; the single-thread rows and the GPU fp16 / int8 rows are single runs. Run-to-run variation on a shared machine is about 10-20%, so small differences are noise.
Single thread
| Version | Beam: ms / clip | Greedy: ms / clip | Beam: RTF | RAM |
|---|---|---|---|---|
| fp32 | 127 ms | 120 ms | 0.0214 | 1116 MB |
| fp16 | 129 ms | 122 ms | 0.0216 | 1181 MB |
| int8 | 80 ms | 74 ms | 0.0135 | 733 MB |
Training data
772,605 utterances, 1,209.6 hours of Bangla speech from nine corpora. Dev and test are held-out utterances from the same sources.
| Source | Hours | Share | Test utts | Test WER* |
|---|---|---|---|---|
| IndicVoices | 632.0 | 52.2% | 4,482 | 11.7% |
| OpenSLR-53 Bengali | 211.3 | 17.5% | 2,867 | 11.1% |
| Vaani | 159.1 | 13.2% | 1,638 | 24.7% |
| Kathbath | 82.3 | 6.8% | 634 | 10.0% |
| Common Voice | 73.2 | 6.1% | 539 | 14.4% |
| Shrutilipi | 27.6 | 2.3% | 219 | 20.2% |
| FLEURS | 15.3 | 1.3% | 60 | 14.6% |
| BEN10 (dialects) | 6.0 | 0.5% | 16 | 78.4% |
| OpenSLR-37 | 2.9 | 0.2% | 30 | 15.7% |
| Total | 1,209.6 | 100% | 10,485 | 14.4% |
*PyTorch model, beam search. Kathbath, OpenSLR and IndicVoices are easiest (10-12%); Vaani (25%) and Shrutilipi (20%) are harder.
BEN10 regional-dialect speech is handled poorly (~78%) - it is tiny (6 training hours, 16 test utterances), so treat that number as indicative only.
Split statistics and text cleaning
| Split | Utterances | Hours | Mean length (s) | Words | Unique words |
|---|---|---|---|---|---|
| train | 772,605 | 1,209.6 | 5.6 | 8,202,235 | 242,087 |
| dev | 5,282 | 8.1 | 5.5 | 55,168 | 12,976 |
| test | 10,485 | 16.4 | 5.6 | 112,188 | 20,269 |
- Training used utterances of 1-20 s (about 725,700 utterances, ~1,168 hours).
- 67,092 training transcripts (8.7%) contained
{...}annotation tags (glosses / transliterations). They were removed from all splits before training and scoring. - 16 kHz mono audio, 80-dim fbank features; 500-token SentencePiece tokenizer (
bpe.model). - Speaker separation between train and dev/test was not verified, so scores may be optimistic for unseen speakers.
How it was trained
- Base: encoder initialised from the English LibriSpeech Zipformer; decoder, joiner and tokenizer trained from scratch.
- Code: icefall Zipformer recipe, commit
3f848bb, with k2, lhotse and PyTorch 2.11. - Recipe: pruned RNN-T loss, ScaledAdam + Eden schedule, 25 epochs (~3 h each) on one GPU in bf16, about 400 s of audio per batch.
- Final weights: average of epochs 23-25. Decoding: greedy or beam search (beam 4).
Full settings
- Model: non-streaming Zipformer2, icefall's default (medium) size, 65,549,011 parameters. Encoder dims 192-256-384-512-384-256, layers 2-2-3-4-3-2, stateless decoder (context 2, dim 512), joiner dim 512, 500 tokens.
- Loss: pruned RNN-T, simple-loss scale 0.5, prune range 5, no CTC branch.
- Optimiser: base LR 0.003, 2000 warm-up steps. Epochs 1-5 used icefall's fine-tuning defaults (
--lr-batches 100000 --lr-epochs 100, nearly constant LR); training was resumed from epoch 5 with--lr-batches 30000 --lr-epochs 3for a decaying LR. - SpecAugment on (icefall default), no MUSAN noise, seed 42.
- Local changes to icefall are limited to the fine-tuning / decoding scripts (Bangla data module, bf16 autocast ported from icefall's
train.py, a GPU memory cap, experiment logging) and the ONNX fp16 conversion import. Model, loss and beam search are stock icefall. - ONNX export: icefall's
export-onnx.pyfrom the averaged weights; fp16 viaonnxconverter_common(keep_io_types), int8 via onnxruntime dynamic MatMul quantisation.
Limitations
- Scores are on held-out data from the training sources; expect worse results on other domains, noisy phone audio and regional dialects.
- The model was trained on utterances up to ~20 s: split longer audio before decoding (sherpa-onnx does not do it for you). Not a streaming model.
- Output follows the training text: no
{...}tags, and it may contain punctuation.
License
- Model weights, tokenizer and this card: CC BY-SA 4.0 (
LICENSE). Attribution to the data sources is required.
Training-data licenses and open points
Share-alike was chosen on purpose: part of the training data is CC BY-SA (OpenSLR), and whether weights count as a derivative of the data is legally unsettled, so this is the conservative choice.
| Source | Declared license |
|---|---|
| IndicVoices, Kathbath, Shrutilipi, Vaani | CC BY 4.0 (gated on the Hugging Face Hub) |
| FLEURS | CC BY 4.0 |
| OpenSLR-53, OpenSLR-37 | CC BY-SA 4.0 |
| Common Voice (cv-corpus-26.0) | CC0 as published by Common Voice (not re-checked for this release) |
| BEN10 | not verified (0.5% of training hours) |
Checked 2026-10-04; each source's own terms apply, including access terms of gated datasets. The English base model declares no license on its Hugging Face card (built with icefall, Apache-2.0, on LibriSpeech, CC BY 4.0) - an open point. Not legal advice: verify with the original sources before relying on these terms.
References
Model and code
- Z. Yao et al. Zipformer: A faster and better encoder for automatic speech recognition. ICLR 2024. arXiv:2310.11230
- F. Kuang et al. Pruned RNN-T for fast, memory-efficient ASR training. 2022. arXiv:2206.13236
- icefall (commit
3f848bb6d0acc970c9b294a30ca0a04a7c9c78d1), k2 - P. Zelasko et al. Lhotse: a speech data representation library for the modern deep learning ecosystem. NeurIPS 2021 DCAI workshop. arXiv:2110.12561
- V. Panayotov et al. LibriSpeech: an ASR corpus based on public domain audio books. ICASSP 2015 (the base model's training data)
Training data 6. IndicVoices: T. Javed et al. 2024. arXiv:2403.01926 7. Kathbath: T. Javed et al. IndicSUPERB. 2022. arXiv:2208.11761 8. Shrutilipi: K. S. Bhogale et al. 2022. arXiv:2208.12666 9. Vaani: S. Pulikodan et al. VAANI: Capturing the language landscape for an inclusive digital India. 2026. arXiv:2603.28714 10. FLEURS: A. Conneau et al. 2022. arXiv:2205.12446 11. Common Voice: R. Ardila et al. LREC 2020. arXiv:1912.06670 12. OpenSLR-53: O. Kjartansson et al. SLTU 2018. doi:10.21437/SLTU.2018-11 13. OpenSLR-37: K. Sodimana et al. SLTU 2018 (the citation its page requests). doi:10.21437/SLTU.2018-14 14. BEN10: no citation information found.