Bangla Zipformer - PyTorch
Bangla speech recognition: a 65.5M-parameter Zipformer transducer, fine-tuned from an English LibriSpeech model on about 1,210 hours of Bangla speech.
CPU-friendly ONNX versions (fp32 / fp16 / int8): SayedShaun/bangla-zipformer-onnx.
Just want to transcribe audio? Use the ONNX version with sherpa-onnx: one
pip install sherpa-onnx, no PyTorch, icefall or k2, essentially the same accuracy, and about 3x faster on a CPU and 5x faster on a GPU than this repo. This PyTorch repo is for research use: fine-tuning, inspecting the network, or running the original weights.
| Test WER / CER | 14.4% / 4.3% (beam search) |
| Dev WER / CER | 14.2% / 4.3% |
| Model | Zipformer2 transducer, non-streaming, 65.5M parameters |
| Weights | model.pt, 250 MB (fp32), average of epochs 23-25 |
| Input | 16 kHz mono audio, up to ~20 s per call (longer audio is split automatically) |
| License | CC BY-SA 4.0 (weights), Apache-2.0 (code) |
Quick start
The model code comes from icefall, so you need to install it and a few packages once:
git clone https://github.com/k2-fsa/icefall
git -C icefall checkout 3f848bb6d0acc970c9b294a30ca0a04a7c9c78d1 # the version this model was trained with
export ICEFALL_ROOT=$PWD/icefall
pip install torch lhotse sentencepiece soundfile soxr numpy kaldialign pypinyin tensorboard dill huggingface_hub
pip install k2 # needs a wheel matching your torch / CUDA: https://k2-fsa.github.io/k2/installation/index.html
Then, using the repo name:
import sys
from huggingface_hub import snapshot_download
sys.path.insert(0, snapshot_download("SayedShaun/bangla-zipformer-pt"))
from asr import BanglaASR
asr = BanglaASR.from_pretrained("SayedShaun/bangla-zipformer-pt") # uses the GPU if available
print(asr.transcribe("audio.wav")) # beam search (default)
print(asr.transcribe("audio.wav", decoding="greedy")) # faster
Accepts a file path (wav / flac / ogg ...) or (samples, sample_rate). Don't want to install icefall and k2? Use the ONNX version, which runs with sherpa-onnx (no PyTorch, no icefall).
Using the model in PyTorch
asr.py is a thin wrapper. To use the weights directly, set up icefall as in the quick start (ICEFALL_ROOT), then:
import os, sys, argparse
import torch, sentencepiece as spm, soundfile as sf
ICEFALL = os.environ["ICEFALL_ROOT"]
sys.path[:0] = [ICEFALL, f"{ICEFALL}/egs/librispeech/ASR/zipformer"] # the `icefall` package + the stock Zipformer recipe
from train import add_model_arguments, get_model, get_params # recipe modules (not pip packages)
from beam_search import greedy_search_batch, modified_beam_search
from lhotse import Fbank, FbankConfig
sp = spm.SentencePieceProcessor(); sp.load("bpe.model")
# 1. build the network: default (medium, non-streaming) architecture - do not change these arguments
parser = argparse.ArgumentParser(); add_model_arguments(parser)
params = get_params(); params.update(vars(parser.parse_args([])))
params.update(context_size=2, use_transducer=True, use_ctc=False, use_attention_decoder=False,
blank_id=sp.piece_to_id("<blk>"), unk_id=sp.piece_to_id("<unk>"), vocab_size=sp.get_piece_size())
model = get_model(params)
# 2. load the weights
model.load_state_dict(torch.load("model.pt", map_location="cpu")["model"], strict=True)
model.eval()
# 3. features -> encoder -> search -> text
wave, sr = sf.read("audio.wav", dtype="float32") # 16 kHz mono, up to ~20 s
feats = torch.from_numpy(Fbank(FbankConfig(num_mel_bins=80)).extract(wave, sr))[None] # (1, T, 80)
with torch.no_grad():
enc, enc_lens = model.forward_encoder(feats, torch.tensor([feats.shape[1]])) # (1, T', 512)
tokens = modified_beam_search(model=model, encoder_out=enc, encoder_out_lens=enc_lens, beam=4) # or greedy_search_batch(...)
print(sp.decode(tokens)[0])
| Checkpoint | model.pt is a dict with one key, "model": the fp32 state dict (743 tensors, 65,549,011 parameters) |
| Model class | AsrModel with encoder_embed, encoder (Zipformer2), decoder (stateless, context 2), joiner, simple_am_proj, simple_lm_proj |
| Input features | 80-dim Kaldi-style fbank (25 ms window, 10 ms shift, no dither) on float samples in [-1, 1], as computed by lhotse.Fbank |
| Encoder output | (batch, T', 512) at about 4x lower frame rate than the features (about 40 ms per frame) |
| Tokens | 500-piece SentencePiece unigram model: <blk> = 0, <sos/eos> = 1, <unk> = 2 |
| Search | greedy_search_batch and modified_beam_search from the recipe's beam_search.py (they need k2) |
| Precision | trained in bf16 AMP, weights stored in fp32; run inference in fp32 (or torch.autocast) |
Results
Full dev (5,282 utterances, 8.1 h) and test (10,485 utterances, 16.4 h) sets.
| Decoding | Test WER | Test CER | Dev WER | Dev CER |
|---|---|---|---|---|
| Beam search (beam 4, default) | 14.37 | 4.30 | 14.17 | 4.29 |
| Greedy | 14.67 | 4.46 | 14.42 | 4.40 |
WER counts punctuation attached to words; CER ignores punctuation. With punctuation stripped, WER is 13.82 (beam) / 14.11 (greedy) on test and 13.69 / 13.94 on dev. The averaging window (epochs 23-25) and beam size were chosen on dev only.
Beam search or greedy?
Beam search (beam 4) is the default. Both decoders use the same model; they differ in accuracy and speed.
| Decoding | Test WER | Test CER | Dev WER | Dev CER | Time per clip, CPU | Time per clip, GPU |
|---|---|---|---|---|---|---|
| Beam search (default) | 14.37 | 4.30 | 14.17 | 4.29 | 199 ms | 155 ms |
| Greedy | 14.67 | 4.46 | 14.42 | 4.40 | 157 ms | 41 ms |
- Use beam search when accuracy matters and latency does not: offline or batch transcription, subtitles, labelling data. It is about 0.3 WER points (about 2% relative) better, consistently on dev and test.
- Use greedy for low latency or high throughput, especially on a GPU. On the GPU greedy is about 3.8x faster than beam search; on the CPU the gap is smaller (beam search is about 27% slower). The loss is small: about 0.3 WER points.
- Why the GPU does not help beam search much: it speeds up the encoder (about 14 ms vs 108 ms on CPU for a 5 s clip), but the beam search runs frame by frame and still takes roughly 90 ms for that clip, against about 15 ms for greedy on the GPU.
- Beam size: the default is 4 (
from_pretrained(..., beam=4)). Beam 8 was no better on dev (14.21 vs 14.17 WER) and is slower. - Need it faster? The ONNX version (sherpa-onnx) runs about 3x faster than this loader on CPU (int8, beam search) and about 5x faster on a GPU (fp32), at nearly the same accuracy, with both decoders.
asr.transcribe("audio.wav") # beam search (default)
asr.transcribe("audio.wav", decoding="greedy") # greedy
Timings: median of 5 repeats per clip on 60 dev clips of 1-15 s (20 each of short, medium and long), beam and greedy run back to back on the same clip, audio passed as an array; 32-core CPU with 4 threads and one GPU (RTX 5080, peak model memory about 315 MB). On this CPU the PyTorch loader was no faster with 4 threads than with 1 (I did not investigate why).
Training data
772,605 utterances, 1,209.6 hours of Bangla speech from nine corpora. Dev and test are held-out utterances from the same sources.
| Source | Hours | Share | Test utts | Test WER* |
|---|---|---|---|---|
| IndicVoices | 632.0 | 52.2% | 4,482 | 11.7% |
| OpenSLR-53 Bengali | 211.3 | 17.5% | 2,867 | 11.1% |
| Vaani | 159.1 | 13.2% | 1,638 | 24.7% |
| Kathbath | 82.3 | 6.8% | 634 | 10.0% |
| Common Voice | 73.2 | 6.1% | 539 | 14.4% |
| Shrutilipi | 27.6 | 2.3% | 219 | 20.2% |
| FLEURS | 15.3 | 1.3% | 60 | 14.6% |
| BEN10 (dialects) | 6.0 | 0.5% | 16 | 78.4% |
| OpenSLR-37 | 2.9 | 0.2% | 30 | 15.7% |
| Total | 1,209.6 | 100% | 10,485 | 14.4% |
*PyTorch model, beam search. Kathbath, OpenSLR and IndicVoices are easiest (10-12%); Vaani (25%) and Shrutilipi (20%) are harder.
BEN10 regional-dialect speech is handled poorly (~78%) - it is tiny (6 training hours, 16 test utterances), so treat that number as indicative only.
Split statistics and text cleaning
| Split | Utterances | Hours | Mean length (s) | Words | Unique words |
|---|---|---|---|---|---|
| train | 772,605 | 1,209.6 | 5.6 | 8,202,235 | 242,087 |
| dev | 5,282 | 8.1 | 5.5 | 55,168 | 12,976 |
| test | 10,485 | 16.4 | 5.6 | 112,188 | 20,269 |
- Training used utterances of 1-20 s (about 725,700 utterances, ~1,168 hours).
- 67,092 training transcripts (8.7%) contained
{...}annotation tags (glosses / transliterations). They were removed from all splits before training and scoring. - 16 kHz mono audio, 80-dim fbank features; 500-token SentencePiece tokenizer (
bpe.model). - Speaker separation between train and dev/test was not verified, so scores may be optimistic for unseen speakers.
How it was trained
- Base: encoder initialised from the English LibriSpeech Zipformer; decoder, joiner and tokenizer trained from scratch.
- Code: icefall Zipformer recipe, commit
3f848bb, with k2, lhotse and PyTorch 2.11. - Recipe: pruned RNN-T loss, ScaledAdam + Eden schedule, 25 epochs (~3 h each) on one GPU in bf16, about 400 s of audio per batch.
- Final weights: average of epochs 23-25. Decoding: greedy or beam search (beam 4).
Full settings
- Model: non-streaming Zipformer2, icefall's default (medium) size, 65,549,011 parameters. Encoder dims 192-256-384-512-384-256, layers 2-2-3-4-3-2, stateless decoder (context 2, dim 512), joiner dim 512, 500 tokens.
- Loss: pruned RNN-T, simple-loss scale 0.5, prune range 5, no CTC branch.
- Optimiser: base LR 0.003, 2000 warm-up steps. Epochs 1-5 used icefall's fine-tuning defaults (
--lr-batches 100000 --lr-epochs 100, nearly constant LR); training was resumed from epoch 5 with--lr-batches 30000 --lr-epochs 3for a decaying LR. - SpecAugment on (icefall default), no MUSAN noise, seed 42.
- Local changes to icefall are limited to the fine-tuning / decoding scripts (Bangla data module, bf16 autocast ported from icefall's
train.py, a GPU memory cap, experiment logging) and the ONNX fp16 conversion import. Model, loss and beam search are stock icefall.
Limitations
- Scores are on held-out data from the training sources; expect worse results on other domains, noisy phone audio and regional dialects.
- Up to ~20 s per call (the loader splits longer audio automatically); 16 kHz mono; not a streaming model.
- Output follows the training text: no
{...}tags, and it may contain punctuation.
License
- Model weights, tokenizer and this card: CC BY-SA 4.0 (
LICENSE). Attribution to the data sources is required. - Loader code (
asr.py): Apache-2.0 (LICENSE-CODE-Apache-2.0).
Training-data licenses and open points
Share-alike was chosen on purpose: part of the training data is CC BY-SA (OpenSLR), and whether weights count as a derivative of the data is legally unsettled, so this is the conservative choice.
| Source | Declared license |
|---|---|
| IndicVoices, Kathbath, Shrutilipi, Vaani | CC BY 4.0 (gated on the Hugging Face Hub) |
| FLEURS | CC BY 4.0 |
| OpenSLR-53, OpenSLR-37 | CC BY-SA 4.0 |
| Common Voice (cv-corpus-26.0) | CC0 as published by Common Voice (not re-checked for this release) |
| BEN10 | not verified (0.5% of training hours) |
Checked 2026-10-04; each source's own terms apply, including access terms of gated datasets. The English base model declares no license on its Hugging Face card (built with icefall, Apache-2.0, on LibriSpeech, CC BY 4.0) - an open point. Not legal advice: verify with the original sources before relying on these terms.
References
Model and code
- Z. Yao et al. Zipformer: A faster and better encoder for automatic speech recognition. ICLR 2024. arXiv:2310.11230
- F. Kuang et al. Pruned RNN-T for fast, memory-efficient ASR training. 2022. arXiv:2206.13236
- icefall (commit
3f848bb6d0acc970c9b294a30ca0a04a7c9c78d1), k2 - P. Zelasko et al. Lhotse: a speech data representation library for the modern deep learning ecosystem. NeurIPS 2021 DCAI workshop. arXiv:2110.12561
- V. Panayotov et al. LibriSpeech: an ASR corpus based on public domain audio books. ICASSP 2015 (the base model's training data)
Training data 6. IndicVoices: T. Javed et al. 2024. arXiv:2403.01926 7. Kathbath: T. Javed et al. IndicSUPERB. 2022. arXiv:2208.11761 8. Shrutilipi: K. S. Bhogale et al. 2022. arXiv:2208.12666 9. Vaani: S. Pulikodan et al. VAANI: Capturing the language landscape for an inclusive digital India. 2026. arXiv:2603.28714 10. FLEURS: A. Conneau et al. 2022. arXiv:2205.12446 11. Common Voice: R. Ardila et al. LREC 2020. arXiv:1912.06670 12. OpenSLR-53: O. Kjartansson et al. SLTU 2018. doi:10.21437/SLTU.2018-11 13. OpenSLR-37: K. Sodimana et al. SLTU 2018 (the citation its page requests). doi:10.21437/SLTU.2018-14 14. BEN10: no citation information found.