Bangla Zipformer - PyTorch

Bangla speech recognition: a 65.5M-parameter Zipformer transducer, fine-tuned from an English LibriSpeech model on about 1,210 hours of Bangla speech. CPU-friendly ONNX versions (fp32 / fp16 / int8): SayedShaun/bangla-zipformer-onnx.

Just want to transcribe audio? Use the ONNX version with sherpa-onnx: one pip install sherpa-onnx, no PyTorch, icefall or k2, essentially the same accuracy, and about 3x faster on a CPU and 5x faster on a GPU than this repo. This PyTorch repo is for research use: fine-tuning, inspecting the network, or running the original weights.

Test WER / CER 14.4% / 4.3% (beam search)
Dev WER / CER 14.2% / 4.3%
Model Zipformer2 transducer, non-streaming, 65.5M parameters
Weights model.pt, 250 MB (fp32), average of epochs 23-25
Input 16 kHz mono audio, up to ~20 s per call (longer audio is split automatically)
License CC BY-SA 4.0 (weights), Apache-2.0 (code)

Quick start

The model code comes from icefall, so you need to install it and a few packages once:

git clone https://github.com/k2-fsa/icefall
git -C icefall checkout 3f848bb6d0acc970c9b294a30ca0a04a7c9c78d1      # the version this model was trained with
export ICEFALL_ROOT=$PWD/icefall

pip install torch lhotse sentencepiece soundfile soxr numpy kaldialign pypinyin tensorboard dill huggingface_hub
pip install k2        # needs a wheel matching your torch / CUDA: https://k2-fsa.github.io/k2/installation/index.html

Then, using the repo name:

import sys
from huggingface_hub import snapshot_download
sys.path.insert(0, snapshot_download("SayedShaun/bangla-zipformer-pt"))
from asr import BanglaASR

asr = BanglaASR.from_pretrained("SayedShaun/bangla-zipformer-pt")           # uses the GPU if available
print(asr.transcribe("audio.wav"))                                  # beam search (default)
print(asr.transcribe("audio.wav", decoding="greedy"))               # faster

Accepts a file path (wav / flac / ogg ...) or (samples, sample_rate). Don't want to install icefall and k2? Use the ONNX version, which runs with sherpa-onnx (no PyTorch, no icefall).

Using the model in PyTorch

asr.py is a thin wrapper. To use the weights directly, set up icefall as in the quick start (ICEFALL_ROOT), then:

import os, sys, argparse
import torch, sentencepiece as spm, soundfile as sf

ICEFALL = os.environ["ICEFALL_ROOT"]
sys.path[:0] = [ICEFALL, f"{ICEFALL}/egs/librispeech/ASR/zipformer"]    # the `icefall` package + the stock Zipformer recipe

from train import add_model_arguments, get_model, get_params           # recipe modules (not pip packages)
from beam_search import greedy_search_batch, modified_beam_search
from lhotse import Fbank, FbankConfig

sp = spm.SentencePieceProcessor(); sp.load("bpe.model")

# 1. build the network: default (medium, non-streaming) architecture - do not change these arguments
parser = argparse.ArgumentParser(); add_model_arguments(parser)
params = get_params(); params.update(vars(parser.parse_args([])))
params.update(context_size=2, use_transducer=True, use_ctc=False, use_attention_decoder=False,
              blank_id=sp.piece_to_id("<blk>"), unk_id=sp.piece_to_id("<unk>"), vocab_size=sp.get_piece_size())
model = get_model(params)

# 2. load the weights
model.load_state_dict(torch.load("model.pt", map_location="cpu")["model"], strict=True)
model.eval()

# 3. features -> encoder -> search -> text
wave, sr = sf.read("audio.wav", dtype="float32")                        # 16 kHz mono, up to ~20 s
feats = torch.from_numpy(Fbank(FbankConfig(num_mel_bins=80)).extract(wave, sr))[None]     # (1, T, 80)
with torch.no_grad():
    enc, enc_lens = model.forward_encoder(feats, torch.tensor([feats.shape[1]]))          # (1, T', 512)
    tokens = modified_beam_search(model=model, encoder_out=enc, encoder_out_lens=enc_lens, beam=4)   # or greedy_search_batch(...)
print(sp.decode(tokens)[0])
Checkpoint model.pt is a dict with one key, "model": the fp32 state dict (743 tensors, 65,549,011 parameters)
Model class AsrModel with encoder_embed, encoder (Zipformer2), decoder (stateless, context 2), joiner, simple_am_proj, simple_lm_proj
Input features 80-dim Kaldi-style fbank (25 ms window, 10 ms shift, no dither) on float samples in [-1, 1], as computed by lhotse.Fbank
Encoder output (batch, T', 512) at about 4x lower frame rate than the features (about 40 ms per frame)
Tokens 500-piece SentencePiece unigram model: <blk> = 0, <sos/eos> = 1, <unk> = 2
Search greedy_search_batch and modified_beam_search from the recipe's beam_search.py (they need k2)
Precision trained in bf16 AMP, weights stored in fp32; run inference in fp32 (or torch.autocast)

Results

Full dev (5,282 utterances, 8.1 h) and test (10,485 utterances, 16.4 h) sets.

Decoding Test WER Test CER Dev WER Dev CER
Beam search (beam 4, default) 14.37 4.30 14.17 4.29
Greedy 14.67 4.46 14.42 4.40

WER counts punctuation attached to words; CER ignores punctuation. With punctuation stripped, WER is 13.82 (beam) / 14.11 (greedy) on test and 13.69 / 13.94 on dev. The averaging window (epochs 23-25) and beam size were chosen on dev only.

Beam search or greedy?

Beam search (beam 4) is the default. Both decoders use the same model; they differ in accuracy and speed.

Decoding Test WER Test CER Dev WER Dev CER Time per clip, CPU Time per clip, GPU
Beam search (default) 14.37 4.30 14.17 4.29 199 ms 155 ms
Greedy 14.67 4.46 14.42 4.40 157 ms 41 ms
  • Use beam search when accuracy matters and latency does not: offline or batch transcription, subtitles, labelling data. It is about 0.3 WER points (about 2% relative) better, consistently on dev and test.
  • Use greedy for low latency or high throughput, especially on a GPU. On the GPU greedy is about 3.8x faster than beam search; on the CPU the gap is smaller (beam search is about 27% slower). The loss is small: about 0.3 WER points.
  • Why the GPU does not help beam search much: it speeds up the encoder (about 14 ms vs 108 ms on CPU for a 5 s clip), but the beam search runs frame by frame and still takes roughly 90 ms for that clip, against about 15 ms for greedy on the GPU.
  • Beam size: the default is 4 (from_pretrained(..., beam=4)). Beam 8 was no better on dev (14.21 vs 14.17 WER) and is slower.
  • Need it faster? The ONNX version (sherpa-onnx) runs about 3x faster than this loader on CPU (int8, beam search) and about 5x faster on a GPU (fp32), at nearly the same accuracy, with both decoders.
asr.transcribe("audio.wav")                      # beam search (default)
asr.transcribe("audio.wav", decoding="greedy")   # greedy

Timings: median of 5 repeats per clip on 60 dev clips of 1-15 s (20 each of short, medium and long), beam and greedy run back to back on the same clip, audio passed as an array; 32-core CPU with 4 threads and one GPU (RTX 5080, peak model memory about 315 MB). On this CPU the PyTorch loader was no faster with 4 threads than with 1 (I did not investigate why).

Training data

772,605 utterances, 1,209.6 hours of Bangla speech from nine corpora. Dev and test are held-out utterances from the same sources.

Source Hours Share Test utts Test WER*
IndicVoices 632.0 52.2% 4,482 11.7%
OpenSLR-53 Bengali 211.3 17.5% 2,867 11.1%
Vaani 159.1 13.2% 1,638 24.7%
Kathbath 82.3 6.8% 634 10.0%
Common Voice 73.2 6.1% 539 14.4%
Shrutilipi 27.6 2.3% 219 20.2%
FLEURS 15.3 1.3% 60 14.6%
BEN10 (dialects) 6.0 0.5% 16 78.4%
OpenSLR-37 2.9 0.2% 30 15.7%
Total 1,209.6 100% 10,485 14.4%

*PyTorch model, beam search. Kathbath, OpenSLR and IndicVoices are easiest (10-12%); Vaani (25%) and Shrutilipi (20%) are harder. BEN10 regional-dialect speech is handled poorly (~78%) - it is tiny (6 training hours, 16 test utterances), so treat that number as indicative only.

Split statistics and text cleaning
Split Utterances Hours Mean length (s) Words Unique words
train 772,605 1,209.6 5.6 8,202,235 242,087
dev 5,282 8.1 5.5 55,168 12,976
test 10,485 16.4 5.6 112,188 20,269
  • Training used utterances of 1-20 s (about 725,700 utterances, ~1,168 hours).
  • 67,092 training transcripts (8.7%) contained {...} annotation tags (glosses / transliterations). They were removed from all splits before training and scoring.
  • 16 kHz mono audio, 80-dim fbank features; 500-token SentencePiece tokenizer (bpe.model).
  • Speaker separation between train and dev/test was not verified, so scores may be optimistic for unseen speakers.

How it was trained

  • Base: encoder initialised from the English LibriSpeech Zipformer; decoder, joiner and tokenizer trained from scratch.
  • Code: icefall Zipformer recipe, commit 3f848bb, with k2, lhotse and PyTorch 2.11.
  • Recipe: pruned RNN-T loss, ScaledAdam + Eden schedule, 25 epochs (~3 h each) on one GPU in bf16, about 400 s of audio per batch.
  • Final weights: average of epochs 23-25. Decoding: greedy or beam search (beam 4).
Full settings
  • Model: non-streaming Zipformer2, icefall's default (medium) size, 65,549,011 parameters. Encoder dims 192-256-384-512-384-256, layers 2-2-3-4-3-2, stateless decoder (context 2, dim 512), joiner dim 512, 500 tokens.
  • Loss: pruned RNN-T, simple-loss scale 0.5, prune range 5, no CTC branch.
  • Optimiser: base LR 0.003, 2000 warm-up steps. Epochs 1-5 used icefall's fine-tuning defaults (--lr-batches 100000 --lr-epochs 100, nearly constant LR); training was resumed from epoch 5 with --lr-batches 30000 --lr-epochs 3 for a decaying LR.
  • SpecAugment on (icefall default), no MUSAN noise, seed 42.
  • Local changes to icefall are limited to the fine-tuning / decoding scripts (Bangla data module, bf16 autocast ported from icefall's train.py, a GPU memory cap, experiment logging) and the ONNX fp16 conversion import. Model, loss and beam search are stock icefall.

Limitations

  • Scores are on held-out data from the training sources; expect worse results on other domains, noisy phone audio and regional dialects.
  • Up to ~20 s per call (the loader splits longer audio automatically); 16 kHz mono; not a streaming model.
  • Output follows the training text: no {...} tags, and it may contain punctuation.

License

  • Model weights, tokenizer and this card: CC BY-SA 4.0 (LICENSE). Attribution to the data sources is required.
  • Loader code (asr.py): Apache-2.0 (LICENSE-CODE-Apache-2.0).
Training-data licenses and open points

Share-alike was chosen on purpose: part of the training data is CC BY-SA (OpenSLR), and whether weights count as a derivative of the data is legally unsettled, so this is the conservative choice.

Source Declared license
IndicVoices, Kathbath, Shrutilipi, Vaani CC BY 4.0 (gated on the Hugging Face Hub)
FLEURS CC BY 4.0
OpenSLR-53, OpenSLR-37 CC BY-SA 4.0
Common Voice (cv-corpus-26.0) CC0 as published by Common Voice (not re-checked for this release)
BEN10 not verified (0.5% of training hours)

Checked 2026-10-04; each source's own terms apply, including access terms of gated datasets. The English base model declares no license on its Hugging Face card (built with icefall, Apache-2.0, on LibriSpeech, CC BY 4.0) - an open point. Not legal advice: verify with the original sources before relying on these terms.

References

Model and code

  1. Z. Yao et al. Zipformer: A faster and better encoder for automatic speech recognition. ICLR 2024. arXiv:2310.11230
  2. F. Kuang et al. Pruned RNN-T for fast, memory-efficient ASR training. 2022. arXiv:2206.13236
  3. icefall (commit 3f848bb6d0acc970c9b294a30ca0a04a7c9c78d1), k2
  4. P. Zelasko et al. Lhotse: a speech data representation library for the modern deep learning ecosystem. NeurIPS 2021 DCAI workshop. arXiv:2110.12561
  5. V. Panayotov et al. LibriSpeech: an ASR corpus based on public domain audio books. ICASSP 2015 (the base model's training data)

Training data 6. IndicVoices: T. Javed et al. 2024. arXiv:2403.01926 7. Kathbath: T. Javed et al. IndicSUPERB. 2022. arXiv:2208.11761 8. Shrutilipi: K. S. Bhogale et al. 2022. arXiv:2208.12666 9. Vaani: S. Pulikodan et al. VAANI: Capturing the language landscape for an inclusive digital India. 2026. arXiv:2603.28714 10. FLEURS: A. Conneau et al. 2022. arXiv:2205.12446 11. Common Voice: R. Ardila et al. LREC 2020. arXiv:1912.06670 12. OpenSLR-53: O. Kjartansson et al. SLTU 2018. doi:10.21437/SLTU.2018-11 13. OpenSLR-37: K. Sodimana et al. SLTU 2018 (the citation its page requests). doi:10.21437/SLTU.2018-14 14. BEN10: no citation information found.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Papers for SayedShaun/bangla-zipformer-pt