M-Kiriku-ASR / README.md
khady's picture
Update README.md
86699c1 verified
|
Raw
History Blame Contribute Delete
11.2 kB
---
license: apache-2.0
language:
- wo
- ff
- srr
metrics:
- wer
- cer
base_model:
- openai/whisper-large-v3
pipeline_tag: automatic-speech-recognition
tags:
- whisper
- fine-tuned
- wolof
- pulaar
- sereer
- senegal
- african-languages
- low-resource
- multilingual
---
<div align="center">
<img src="https://huggingface.co/datasets/AIHubSN/Sunubaat_Wolof_ASR/resolve/main/IMG-20260721-WA0021.jpg" alt="AI Hub Senegal" height= "500" width="1200"/>
</div>
# M-Kiriku ASR: Multilingual Speech Recognition for Senegal
**The first open-source multilingual ASR model covering Wolof, Pulaar, and Sérère three of Senegal's most widely spoken national languages.**
M-Kiriku ASR is a fine-tuned version of `openai/whisper-large-v3` developed by **IA Hub Senegal**.
The model is built upon a subset of the **Kallama Speech Dataset**. We started from approximately **90 hours** of multilingual recordings from the Kallama project that had **not previously undergone linguistic review **.
Since these recordings originate from **radio broadcasts and conversations**, IA Hub Senegal first carried out an extensive audio preprocessing pipeline to improve data quality. This included removing music, jingles, background noise, and other non-speech segments before linguistic validation.
After audio cleaning, IA Hub Senegal performed transcription review, linguistic validation, and quality assurance, resulting in a high-quality **75.4-hour** multilingual corpus composed of:
- **34.9 hours** of **Wolof**
- **24.3 hours** of **Sérère**
- **16.2 hours** of **Pulaar**
To train M-Kiriku ASR, this curated corpus was merged with existing open-source speech datasets from the community, including **Google FLEURS**, **Alpha**, **Urban**, **Wolof Banking**, and other publicly available resources. The resulting multilingual training corpus totals **approximately 150 hours** of speech.
By combining newly curated data from the Kallama project with existing open datasets, M-Kiriku ASR provides one of the largest open multilingual ASR resources for Senegalese languages and aims to accelerate speech technology research and development across West Africa.
## Why This Matters
Wolof, Pulaar, and Sérère are low-resource languages with **zero native support** in the original Whisper vocabulary. These three languages alone cover over 80% of Senegal's population, yet remain invisible to mainstream speech technology.
M-Kiriku ASR doesn't just fine-tune weights — it expands the model's fundamental capabilities:
* **Custom Vocabulary:** Injected language-specific characters (`ñ`, `ë`, `ŋ`, `ɗ`, `ɓ`, `ƴ`) and custom language tokens (`<|wo|>`, `<|pu|>`, `<|se|>`) into the tokenizer.
* **Smart Initialization:** Used warm-starting from French `<|fr|>` embeddings to give the decoder a linguistic head-start, leveraging shared phonological features (CV syllable structure, nasal vowels).
* **150h Expert Corpus:** Trained on three curated datasets spanning agriculture, banking, urban life, and daily conversation.
* **Text Normalization:** A custom `SenegalNormalizer` standardizes orthographic variants, Unicode inconsistencies, and punctuation across all three languages.
* **Discriminative Learning Rate:** The encoder (acoustic features) learns at half the rate of the decoder (language generation), preserving Whisper's universal audio understanding while rapidly adapting text generation to new languages.
## Performance & Results
M-Kiriku ASR achieves benchmark-setting Word Error Rates across all three target languages.
| Language | WER | CER | Training Data | Eval Samples |
| :--- | :--- | :--- | :--- | :--- |
| **Wolof** | **16.34%** | 8.56% | 88.27h | 2,315 |
| **Pulaar** | **45.28%** | 24.45% | 25.21h | 490 |
| **Sérère** | **42.99%** | 19.80% | 33.87h | 557 |
| **Global** | **27.71%** | 13.97% | ~150h | 3,362 |
### Training Configuration
| Parameter | Value |
| :--- | :--- |
| Base Model | `openai/whisper-large-v3` (1.55B params, 128 mel bins) |
| Learning Rate | 2e-5 decoder / 1e-5 encoder (discriminative) |
| Schedule | Cosine with 10% warmup |
| Batch Size | 16 effective (2 × 8 gradient accumulation) |
| Epochs | 10 (39,900 steps) |
| Training Time | ~72 hours on RTX A6000 |
| Precision | FP16 with gradient checkpointing |
| Regularization | Dropout 0.1 on attention and residuals |
| Augmentation | Gaussian noise, time stretch, pitch shift (30% probability) |
| Audio Filtering | Removed samples < 0.5s or > 30s |
| Text Normalization | Custom SenegalNormalizer (Unicode, punctuation, casing) |
| Best Model Selection | Early stopping on WER global (patience 5) |
## Quick Start
### Simple Transcription
```python
import torch
import librosa
from transformers import WhisperProcessor, WhisperForConditionalGeneration
MODEL_ID = "AIHubSN/m-kiriku-asr"
device = "cuda" if torch.cuda.is_available() else "cpu"
processor = WhisperProcessor.from_pretrained(MODEL_ID)
model = WhisperForConditionalGeneration.from_pretrained(
MODEL_ID, torch_dtype=torch.float16
).to(device)
# Load audio at 16kHz
audio, sr = librosa.load("audio.wav", sr=16000)
# Prepare features
input_features = processor.feature_extractor(
audio, sampling_rate=sr, return_tensors="pt"
).input_features.to(device, dtype=torch.float16)
# Transcribe
with torch.no_grad():
generated_ids = model.generate(input_features, max_new_tokens=440)
text = processor.batch_decode(generated_ids, skip_special_tokens=True)[0]
print(text)
```
### Specifying Language
M-Kiriku ASR uses custom language tokens (`<|wo|>`, `<|pu|>`, `<|se|>`) that are not part of the standard Whisper vocabulary. To force a specific language, pass `decoder_input_ids` directly:
```python
# Build the decoder prompt for Wolof
sot_id = processor.tokenizer.convert_tokens_to_ids("<|startoftranscript|>")
wo_id = processor.tokenizer.convert_tokens_to_ids("<|wo|>") # or <|pu|> / <|se|>
task_id = processor.tokenizer.convert_tokens_to_ids("<|transcribe|>")
nots_id = processor.tokenizer.convert_tokens_to_ids("<|notimestamps|>")
decoder_input_ids = torch.tensor([[sot_id, wo_id, task_id, nots_id]]).to(device)
with torch.no_grad():
generated_ids = model.generate(
input_features,
decoder_input_ids=decoder_input_ids,
max_new_tokens=440,
)
text = processor.batch_decode(generated_ids, skip_special_tokens=True)[0]
```
> **Note:** Do not use the `language=` parameter of the standard Whisper pipeline — it does not recognize `wo`, `pu`, or `se`. Always use `decoder_input_ids` for language control.
### Transcribing Long Audio (> 30 seconds)
Whisper processes a maximum of 30 seconds per inference. For longer audio, chunk with overlap to avoid cutting words at boundaries:
```python
import math
def transcribe_long(audio_path, lang_code="wo"):
audio, sr = librosa.load(audio_path, sr=16000)
duration = len(audio) / sr
# Chunk with 5s overlap
chunk_sec, overlap = 30, 5
step = chunk_sec - overlap
# Prepare decoder prompt
sot_id = processor.tokenizer.convert_tokens_to_ids("<|startoftranscript|>")
lang_id = processor.tokenizer.convert_tokens_to_ids(f"<|{lang_code}|>")
task_id = processor.tokenizer.convert_tokens_to_ids("<|transcribe|>")
nots_id = processor.tokenizer.convert_tokens_to_ids("<|notimestamps|>")
decoder_ids = torch.tensor([[sot_id, lang_id, task_id, nots_id]]).to(device)
texts = []
for i in range(max(1, math.ceil((duration - overlap) / step))):
start = i * step
end = min(start + chunk_sec, duration)
chunk = audio[int(start * sr):int(end * sr)]
if len(chunk) < sr * 0.5:
continue
features = processor.feature_extractor(
chunk, sampling_rate=sr, return_tensors="pt"
).input_features.to(device, dtype=torch.float16)
with torch.no_grad():
ids = model.generate(features, decoder_input_ids=decoder_ids, max_new_tokens=440)
texts.append(processor.batch_decode(ids, skip_special_tokens=True)[0].strip())
return " ".join(texts)
print(transcribe_long("long_audio.wav", lang_code="se")) # Sérère
```
### Production Deployment with Faster-Whisper
For high-throughput deployment, convert to CTranslate2 with INT8 quantization (~4× faster):
```bash
pip install faster-whisper
ct2-whisper-converter --model AIHubSN/m-kiriku-asr --output_dir m-kiriku-ct2 --quantization int8
```
```python
from faster_whisper import WhisperModel
model = WhisperModel("m-kiriku-ct2", device="cuda", compute_type="int8")
segments, info = model.transcribe("audio.wav", vad_filter=True)
for segment in segments:
print(f"[{segment.start:.1f}s → {segment.end:.1f}s] {segment.text}")
```
## Architecture
M-Kiriku ASR is a standard Whisper encoder-decoder Transformer with the following modifications:
* **Encoder:** 32 Transformer layers processing 128-bin log-Mel spectrograms (1500 frames per 30s of audio). Pretrained weights from OpenAI, fine-tuned at reduced learning rate (1e-5) to preserve acoustic representations.
* **Decoder:** 32 Transformer layers with causal self-attention and cross-attention to the encoder. Fine-tuned at full learning rate (2e-5) to learn the three target languages.
* **Vocabulary:** Extended from 51,866 to 51,873 tokens (+7: `ŋ`, `ɗ`, `ɓ`, `ƴ`, `<|wo|>`, `<|pu|>`, `<|se|>`).
* **Language Token Initialization:** `<|wo|>`, `<|pu|>`, `<|se|>` embeddings warm-started from `<|fr|>` (French), leveraging shared CV syllable structure and nasal vowels between French and Senegalese languages.
## Limitations
* **Pulaar and Sérère** have higher WER due to limited training data (25h and 34h respectively). Performance will improve significantly with more data.
* **Dialectal variation** within each language (e.g., Pulaar Fuuta vs. Pulaar Jeeri) is not fully covered.
* **Standard Whisper pipeline** does not support the custom language tokens. Use `decoder_input_ids` for language specification (see Quick Start above).
## Produced By
This project is a testament to Senegalese AI excellence:
- **AI Hub Senegal** Coordination, linguistic validation, community mobilization, and project support.
- **Contact:** contact@aihubsenegal.com
## Citation & Attribution
If you use **M-Kiriku ASR** in your research, products, benchmarks, publications, or demonstrations, **please acknowledge AI Hub Senegal** by citing this repository and mentioning **AI Hub Senegal**.
We kindly ask users to include an acknowledgement such as:
> *"M-Kiriku ASR, developed by AI Hub Senegal."*
Your citation helps recognize the work of the contributors and supports the continued development of open-source AI resources for African languages.
## Citation
```bibtex
@misc{iahubsn2026m_kiriku_asr,
title = {M-Kiriku ASR: Whisper Large-v3 Fine-tuned for Wolof, Pulaar, and S\'{e}r\`{e}re},
author = {IA Hub Senegal},
year = {2026},
publisher = {Hugging Face},
howpublished = {\url{https://huggingface.co/AIHubSN/m-kiriku-asr}},
note = {Fine-tuned on ~150 hours of expert-verified multilingual Senegalese speech}
}
```
**Developed with 🤍 in Dakar for the African AI Renaissance.**