Whisper Small — Sinhala

Fine-tuned variants of openai/whisper-small for Sinhala automatic speech recognition (ASR), plus the training data splits used to produce them.

This repo bundles nine training runs (run1 to run9) and the dataset splits they were trained and evaluated on, so the full experiment is reproducible from one place. Run names match the ErrorAnalysis/ folders and the project presentation. (The project test plan numbers runs in a different order: its "Run 6" is run5 here.)

Repo layout

models/
  run1/   full fine-tune, v1 data — merged, ready-to-use model
  run2/   LoRA (AMD hardware), v1 data — merged, ready-to-use model
  run3/   LoRA adapter only, v1 data — best checkpoint (epoch 2), run interrupted at epoch 2.5
  run4/   LoRA adapter only, v1 data — r=32, six target modules
  run5/   full fine-tune, v4 data (speaker-disjoint, spacing-normalised)
  run6/   full fine-tune, v5 data, lr 3e-5 linear, 5 epochs   <- best model
  run7/   as run6 but cosine scheduler
  run8/   as run6 but 4 epochs
  run9/   as run6 but lr 2e-5 and 4 epochs
checkpoints/
  run5/   resume notes and training log for run5
data/
  stratified/     (v1) stratified train/test/validation split
  stratified_v2/  (v2) regenerated stratified split
  stratified_v3/  speaker-disjoint split
  stratified_v4/  (v4) speaker-disjoint, spacing-normalised split

The v5 split (used by run6 to run9) is in Yohan2003/whisper-sl-data under data/stratified_v5/, together with a dedicated dataset card.

Results

Test WER / CER in percent. run1 to run4 were scored on the v1 test set (15,483 utterances), which is not speaker-disjoint, so its numbers are optimistic. run5 to run9 were scored on the speaker-disjoint test set (15,860 utterances); v5 also normalises about 8.5% of the labels, so run5 vs run6 is indicative only. Text is lower-cased with punctuation removed before scoring.

Run Type Data LR · scheduler Epochs Test WER Test CER
run1 Full v1 3e-5 · linear 4 17.36 3.49
run2 LoRA (merged) v1 5e-5 · linear 4 21.01 5.73
run3 LoRA adapter v1 3e-5 · cosine interrupted at 2.5 59.07 18.03
run4 LoRA adapter (r=32) v1 1e-4 · cosine 4 25.99 7.06
run5 Full v4 3e-5 · linear 4 19.06 4.90
run6 Full v5 3e-5 · linear 5 17.15 4.62
run7 Full v5 3e-5 · cosine 5 17.44 4.72
run8 Full v5 3e-5 · linear 4 17.90 4.76
run9 Full v5 2e-5 · linear 4 19.07 5.02

Which run should I use?

Run Use when
run6 You want the best Sinhala accuracy. Full weights, no PEFT dependency.
run1 You want the recipe that started the error analysis (v1 data).
run2 A merged LoRA model.
run3, run4 You want a small LoRA adapter to load on top of openai/whisper-small.
run5, run7, run8, run9 Comparisons: v4 data, cosine scheduler, fewer epochs, lower learning rate.

Usage

Full / merged models (run1, run2, run5 to run9)

from transformers import WhisperForConditionalGeneration, WhisperProcessor

repo = "Yohan2003/whisper-small-sinhala"
model = WhisperForConditionalGeneration.from_pretrained(repo, subfolder="models/run6")
processor = WhisperProcessor.from_pretrained(repo, subfolder="models/run6")

LoRA adapters (run3, run4)

from transformers import WhisperForConditionalGeneration, WhisperProcessor
from peft import PeftModel
from huggingface_hub import snapshot_download

base = WhisperForConditionalGeneration.from_pretrained("openai/whisper-small")
path = snapshot_download("Yohan2003/whisper-small-sinhala", allow_patterns="models/run3/*")
model = PeftModel.from_pretrained(base, f"{path}/models/run3")
processor = WhisperProcessor.from_pretrained(f"{path}/models/run3")

Training data

All runs were trained on Sinhala speech aggregated from OpenSLR-52, YouTube, BizBrains, and Linga sources (~154,828 examples total). See the data/ folder or the dataset card for split details and column schema.

Limitations

  • stratified (v1) and stratified_v2 may share speakers between train and eval splits; use stratified_v3 or later for a speaker-independent evaluation.
  • Errors that remain in the best model are mostly spelling conventions (spoken vs written forms), rare words and long utterances; see the project's error analysis.
  • The LoRA runs can repeat words in long outputs; no_repeat_ngram_size=3 is a recommended decoding setting for them.
  • English forgetting was measured only for some early runs; it was not evaluated for run5 to run9.
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Yohan2003/whisper-small-sinhala

Adapter
(299)
this model

Dataset used to train Yohan2003/whisper-small-sinhala