Automatic Speech Recognition
Transformers
Safetensors
Mende (Sierra Leone)
wav2vec2-bert
speech
asr
african-languages
w2v-bert
khaya
Eval Results (legacy)
Instructions to use KhayaAI/w2v-bert-men with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use KhayaAI/w2v-bert-men with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("automatic-speech-recognition", model="KhayaAI/w2v-bert-men")# Load model directly from transformers import AutoProcessor, AutoModelForCTC processor = AutoProcessor.from_pretrained("KhayaAI/w2v-bert-men") model = AutoModelForCTC.from_pretrained("KhayaAI/w2v-bert-men", device_map="auto") - Notebooks
- Google Colab
- Kaggle
| language: | |
| - men | |
| license: apache-2.0 | |
| library_name: transformers | |
| pipeline_tag: automatic-speech-recognition | |
| tags: | |
| - speech | |
| - asr | |
| - african-languages | |
| - w2v-bert | |
| - khaya | |
| - men | |
| - arxiv:2607.21540 | |
| metrics: | |
| - wer | |
| model-index: | |
| - name: w2v-bert-men | |
| results: | |
| - task: | |
| type: automatic-speech-recognition | |
| name: Automatic Speech Recognition | |
| metrics: | |
| - type: wer | |
| value: 12.9 | |
| name: WER | |
|  | |
| # w2v-bert-men — Mende ASR | |
| **DONDO** (Democratizing Oral Neural Dialect Ontology) open speech-recognition base model for **Mende** (Sierra Leone). | |
| Fine-tuned from the **w2v-BERT 2.0** self-supervised speech encoder. | |
| - **Language:** Mende (`men`) | |
| - **Task:** Automatic Speech Recognition (CTC) | |
| - **Backbone:** w2v-BERT 2.0 | |
| - **License:** Apache-2.0 (attribution only; commercial use permitted) | |
| - **Paper:** [arXiv:2607.21540](https://arxiv.org/abs/2607.21540) | |
| - **Demo:** [Try it in your browser](https://huggingface.co/spaces/Ghana-NLP/Sierra-Leone-ASR) | |
| - **Part of:** the DONDO family — <https://huggingface.co/KhayaAI> | |
| ## Intended use | |
| A **base model** for Mende ASR. Use it directly to transcribe read or relatively | |
| clean speech, or as an initialisation to **fine-tune on your own in-domain data** | |
| (conversational, broadcast, clinical, etc.) with comparatively little labelled audio. | |
| DONDO is also intended as an **open test bed**: a shared, reproducible model on | |
| which new low-resource speech techniques can be demonstrated for the benefit of | |
| all. Aligned research groups are welcome to build on it and collaborate. | |
| ## Training data | |
| Primarily read speech derived from **religious texts** paired with verified | |
| transcripts in the standard orthography. Such data is license-clear, exists for | |
| many otherwise under-resourced languages, and is orthographically consistent. Its | |
| main limitation is domain narrowness, which is why this is released as a base model. | |
| ## Evaluation | |
| | Metric | Value | | |
| |---|---| | |
| | WER (in-domain test) | **12.9%** | | |
| WER is computed on in-domain test material and should not be read as a guarantee | |
| for other domains. | |
| ## How to use | |
| ```python | |
| import torch, torchaudio | |
| from transformers import AutoProcessor, AutoModelForCTC | |
| model_id = "KhayaAI/w2v-bert-men" | |
| processor = AutoProcessor.from_pretrained(model_id) | |
| model = AutoModelForCTC.from_pretrained(model_id) | |
| speech, sr = torchaudio.load("audio.wav") | |
| if sr != 16000: | |
| speech = torchaudio.functional.resample(speech, sr, 16000) | |
| inputs = processor(speech.squeeze().numpy(), sampling_rate=16000, return_tensors="pt") | |
| with torch.no_grad(): | |
| logits = model(**inputs).logits | |
| pred_ids = torch.argmax(logits, dim=-1) | |
| print(processor.batch_decode(pred_ids)[0]) | |
| ``` | |
| ## Limitations | |
| Trained largely on read religious speech, so the model may underperform on | |
| spontaneous, code-switched or noisy audio until fine-tuned. Orthographic | |
| conventions vary across communities; evaluation for the smallest languages rests | |
| on limited test sets. | |
| ## License | |
| Released under the **Apache-2.0** license. You are free to use, modify, redistribute and build upon this model, **including for commercial purposes**. The only substantive requirement is **attribution**. | |
| ## Professional services & hosted APIs | |
| This model is free to use under Apache-2.0. If your team would like help | |
| **deploying it offline on your own infrastructure and data**, Khaya AI offers | |
| professional services (integration, fine-tuning and on-premises deployment). For | |
| ready-to-use and more advanced ASR, including **APIs** for interested parties, | |
| see Khaya Studio. | |
| - Khaya AI — <https://khaya.ai> | |
| - Khaya Studio — <https://studio.khaya.ai> | |
| - Khaya Studio ASR — <https://studio.khaya.ai/asr> | |
| ## Citation | |
| ```bibtex | |
| @article{azunre2026dondo, | |
| title = {DONDO: Open w2v-BERT Speech Recognition Base Models for African Languages}, | |
| author = {Azunre, Paul and Ibrahim, Naafi and Budu, Joel and Adu-Gyamfi, Lawrence}, | |
| year = {2026}, | |
| eprint = {2607.21540}, | |
| archivePrefix = {arXiv}, | |
| primaryClass = {cs.CL}, | |
| note = {Democratizing Oral Neural Dialect Ontology. Funded by the Huniki Federation.} | |
| } | |
| ``` | |
| ## Acknowledgements | |
| Funded by the **Huniki Federation**. We thank **Ghana-NLP** and **Algorine Research** for their support with benchmarking, testing and data, and **Hugging Face** for compute credits. We also thank the language communities and data contributors whose recordings and transcriptions made this work possible. | |