Whisper Small — Sinhala
Fine-tuned variants of openai/whisper-small for Sinhala automatic speech recognition (ASR), plus the training data splits used to produce them.
This repo bundles nine training runs (run1 to run9) and the dataset splits they were trained and evaluated on, so the full experiment is reproducible from one place. Run names match the ErrorAnalysis/ folders and the project presentation. (The project test plan numbers runs in a different order: its "Run 6" is run5 here.)
Repo layout
models/
run1/ full fine-tune, v1 data — merged, ready-to-use model
run2/ LoRA (AMD hardware), v1 data — merged, ready-to-use model
run3/ LoRA adapter only, v1 data — best checkpoint (epoch 2), run interrupted at epoch 2.5
run4/ LoRA adapter only, v1 data — r=32, six target modules
run5/ full fine-tune, v4 data (speaker-disjoint, spacing-normalised)
run6/ full fine-tune, v5 data, lr 3e-5 linear, 5 epochs <- best model
run7/ as run6 but cosine scheduler
run8/ as run6 but 4 epochs
run9/ as run6 but lr 2e-5 and 4 epochs
checkpoints/
run5/ resume notes and training log for run5
data/
stratified/ (v1) stratified train/test/validation split
stratified_v2/ (v2) regenerated stratified split
stratified_v3/ speaker-disjoint split
stratified_v4/ (v4) speaker-disjoint, spacing-normalised split
The v5 split (used by run6 to run9) is in Yohan2003/whisper-sl-data under data/stratified_v5/, together with a dedicated dataset card.
Results
Test WER / CER in percent. run1 to run4 were scored on the v1 test set (15,483 utterances), which is not speaker-disjoint, so its numbers are optimistic. run5 to run9 were scored on the speaker-disjoint test set (15,860 utterances); v5 also normalises about 8.5% of the labels, so run5 vs run6 is indicative only. Text is lower-cased with punctuation removed before scoring.
| Run | Type | Data | LR · scheduler | Epochs | Test WER | Test CER |
|---|---|---|---|---|---|---|
run1 |
Full | v1 | 3e-5 · linear | 4 | 17.36 | 3.49 |
run2 |
LoRA (merged) | v1 | 5e-5 · linear | 4 | 21.01 | 5.73 |
run3 |
LoRA adapter | v1 | 3e-5 · cosine | interrupted at 2.5 | 59.07 | 18.03 |
run4 |
LoRA adapter (r=32) | v1 | 1e-4 · cosine | 4 | 25.99 | 7.06 |
run5 |
Full | v4 | 3e-5 · linear | 4 | 19.06 | 4.90 |
run6 |
Full | v5 | 3e-5 · linear | 5 | 17.15 | 4.62 |
run7 |
Full | v5 | 3e-5 · cosine | 5 | 17.44 | 4.72 |
run8 |
Full | v5 | 3e-5 · linear | 4 | 17.90 | 4.76 |
run9 |
Full | v5 | 2e-5 · linear | 4 | 19.07 | 5.02 |
Which run should I use?
| Run | Use when |
|---|---|
run6 |
You want the best Sinhala accuracy. Full weights, no PEFT dependency. |
run1 |
You want the recipe that started the error analysis (v1 data). |
run2 |
A merged LoRA model. |
run3, run4 |
You want a small LoRA adapter to load on top of openai/whisper-small. |
run5, run7, run8, run9 |
Comparisons: v4 data, cosine scheduler, fewer epochs, lower learning rate. |
Usage
Full / merged models (run1, run2, run5 to run9)
from transformers import WhisperForConditionalGeneration, WhisperProcessor
repo = "Yohan2003/whisper-small-sinhala"
model = WhisperForConditionalGeneration.from_pretrained(repo, subfolder="models/run6")
processor = WhisperProcessor.from_pretrained(repo, subfolder="models/run6")
LoRA adapters (run3, run4)
from transformers import WhisperForConditionalGeneration, WhisperProcessor
from peft import PeftModel
from huggingface_hub import snapshot_download
base = WhisperForConditionalGeneration.from_pretrained("openai/whisper-small")
path = snapshot_download("Yohan2003/whisper-small-sinhala", allow_patterns="models/run3/*")
model = PeftModel.from_pretrained(base, f"{path}/models/run3")
processor = WhisperProcessor.from_pretrained(f"{path}/models/run3")
Training data
All runs were trained on Sinhala speech aggregated from OpenSLR-52, YouTube, BizBrains, and Linga sources (~154,828 examples total). See the data/ folder or the dataset card for split details and column schema.
Limitations
stratified(v1) andstratified_v2may share speakers between train and eval splits; usestratified_v3or later for a speaker-independent evaluation.- Errors that remain in the best model are mostly spelling conventions (spoken vs written forms), rare words and long utterances; see the project's error analysis.
- The LoRA runs can repeat words in long outputs;
no_repeat_ngram_size=3is a recommended decoding setting for them. - English forgetting was measured only for some early runs; it was not evaluated for
run5torun9.
Model tree for Yohan2003/whisper-small-sinhala
Base model
openai/whisper-small