Automatic Speech Recognition
Transformers
Safetensors
English
whisper
african-accented-english
clinical-asr
afrispeech
Eval Results (legacy)
Instructions to use Professor/whisper-small-afrispeech with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Professor/whisper-small-afrispeech with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("automatic-speech-recognition", model="Professor/whisper-small-afrispeech")# Load model directly from transformers import AutoProcessor, AutoModelForSpeechSeq2Seq processor = AutoProcessor.from_pretrained("Professor/whisper-small-afrispeech") model = AutoModelForSpeechSeq2Seq.from_pretrained("Professor/whisper-small-afrispeech", device_map="auto") - Notebooks
- Google Colab
- Kaggle
| language: en | |
| license: cc-by-nc-sa-4.0 | |
| library_name: transformers | |
| pipeline_tag: automatic-speech-recognition | |
| tags: | |
| - automatic-speech-recognition | |
| - whisper | |
| - african-accented-english | |
| - clinical-asr | |
| - afrispeech | |
| base_model: openai/whisper-small | |
| datasets: | |
| - intronhealth/afrispeech-200 | |
| metrics: | |
| - wer | |
| model-index: | |
| - name: whisper-small-afrispeech | |
| results: | |
| - task: | |
| type: automatic-speech-recognition | |
| name: Automatic Speech Recognition | |
| dataset: | |
| name: AfriSpeech-200 (16 kHz) | |
| type: Professor/afrispeech-200-16khz | |
| metrics: | |
| - type: wer | |
| value: 22.94 | |
| name: Test WER | |
| # whisper-small-afrispeech | |
| Fine-tune of [`openai/whisper-small`](https://huggingface.co/openai/whisper-small) on | |
| **AfriSpeech-200** for **African-accented English** speech recognition, spanning both | |
| **general** and **clinical/medical** domains. | |
| - **Base model:** openai/whisper-small (244M) | |
| - **Training data:** [Professor/afrispeech-200-16khz](https://huggingface.co/datasets/Professor/afrispeech-200-16khz) β a clean 16 kHz mono build of [AfriSpeech-200](https://huggingface.co/datasets/intronhealth/afrispeech-200) (Intron Health): ~200 h, 120 accents, 13 countries, clinical + general. | |
| - **Language:** English (African accents) | |
| ## Results (test split, 6,178 clips) | |
| ### Fine-tuned vs. zero-shot baseline | |
| Zero-shot = `openai/whisper-small` evaluated on the *same* test set with no fine-tuning. | |
| | Metric | Zero-shot (baseline) | **Fine-tuned (this model)** | Improvement | | |
| |---|---|---|---| | |
| | Overall WER | 49.15% | **22.94%** | **β26.2 pts (β53%)** | | |
| | Clinical WER | 56.28% | **29.84%** | **β26.4 pts (β47%)** | | |
| | General WER | 42.21% | **16.23%** | **β26.0 pts (β62%)** | | |
| Fine-tuning **roughly halves WER across the board** β a 26-point absolute reduction | |
| overall β and cuts clinical error from 56% to 30%. | |
| ### WER by domain | |
| | Domain | WER | n | | |
| |---|---|---| | |
| | General | 16.23% | 2,670 | | |
| | **Clinical** | **29.84%** | 3,508 | | |
| | **Overall** | **22.94%** | 6,178 | | |
| **Clinical speech is ~1.8Γ harder than general** β medical terminology (drug names, | |
| conditions, dosages) drives most of the error, so domain matters when reporting WER. | |
| ### WER by accent (70 accents with n β₯ 25) | |
| Performance varies widely across accents (**~8% to ~62%**): | |
| | Best accents | WER | | Hardest accents | WER | | |
| |---|---|---|---|---| | |
| | okirika | 8.2% | | agatu | 61.6% | | |
| | brass | 9.4% | | mwaghavul | 58.1% | | |
| | afrikaans | 11.7% | | efik | 57.8% | | |
| | twi | 11.9% | | bini | 51.0% | | |
| | ikwere | 12.1% | | idah | 47.7% | | |
| The hardest accents are mostly smaller, under-represented Nigerian accents β an important | |
| coverage/equity consideration. Full per-accent numbers: `eval_wer_breakdown.csv`. | |
| ## Usage | |
| ```python | |
| import torch | |
| from transformers import WhisperForConditionalGeneration, WhisperProcessor | |
| model = WhisperForConditionalGeneration.from_pretrained("Professor/whisper-small-afrispeech") | |
| processor = WhisperProcessor.from_pretrained("Professor/whisper-small-afrispeech") | |
| # audio: a 16 kHz mono waveform (numpy array) | |
| inputs = processor(audio, sampling_rate=16000, return_tensors="pt") | |
| ids = model.generate(inputs.input_features, language="english", task="transcribe") | |
| print(processor.batch_decode(ids, skip_special_tokens=True)[0]) | |
| ``` | |
| ## Training | |
| - **Objective:** full fine-tune (all parameters), 3 epochs | |
| - **Precision / hardware:** bf16 on a single NVIDIA A40 (48 GB) | |
| - **Batch:** 32 Β· **LR:** 1e-5 (500 warmup steps, linear decay) Β· **best model by eval WER** | |
| - **Data handling:** on-the-fly log-Mel feature extraction; clips > 30 s filtered out | |
| (Whisper's fixed 30 s window); labels truncated to Whisper's 448-token decoder cap. | |
| - Runtime: ~83 min. | |
| ## Limitations & biases | |
| - **Clinical error rate is high (~30%)** β verify medical terms; not for unsupervised clinical use. | |
| - **Large accent disparity (~8β62% WER)** β much weaker on under-represented accents. | |
| - **β€ 30 s audio** β longer inputs are truncated to 30 s by Whisper's feature extractor. | |
| - **English only**; inherits Whisper's general biases. | |
| ## Dataset, license & attribution | |
| Trained on AfriSpeech-200, created by **Intron Health** and released under | |
| **CC-BY-NC-SA-4.0** (non-commercial, share-alike, attribution). This derivative model | |
| carries the same license. Please cite AfriSpeech-200: | |
| ```bibtex | |
| @article{olatunji2023afrispeech, | |
| title={AfriSpeech-200: Pan-African Accented Speech Dataset for Clinical and General Domain ASR}, | |
| author={Olatunji, Tobi and others}, | |
| journal={Transactions of the Association for Computational Linguistics}, | |
| year={2023} | |
| } | |
| ``` | |