keystats's picture
Update README.md
481d7db verified
|
Raw
History Blame Contribute Delete
7.28 kB
---
license: apache-2.0
language:
- ln
tags:
- automatic-speech-recognition
- lingala
- wav2vec2-bert
- ctc
- low-resource
base_model: keystats/w2v-bert-2.0-lingala-main-best-3
datasets:
- google/WaxalNLP
- KasuleTrevor/Lingala_100hrs
pipeline_tag: automatic-speech-recognition
---
# w2v-bert-2.0-lingala
A Lingala automatic speech recognition (ASR) model, further
fine-tuned from
[keystats/w2v-bert-2.0-lingala-main-best-3](https://huggingface.co/keystats/w2v-bert-2.0-lingala-main-best-3),
which was itself continued from
[keystats/w2v-bert-2.0-lingala-main-best](https://huggingface.co/keystats/w2v-bert-2.0-lingala-main-best),
which was fine-tuned from
[facebook/w2v-bert-2.0](https://huggingface.co/facebook/w2v-bert-2.0).
This checkpoint is the latest link in that chain: starting from
`main-best-3`'s weights, training continued for **1 additional epoch**
at a learning rate of **1e-5**, everything else — data, splits,
casing, filtering — kept the same.
## Model description
`facebook/w2v-bert-2.0` — a large-scale, multilingual self-supervised
speech encoder pretrained with a BERT-style masked prediction
objective — is used as the backbone, with a character-level CTC
(Connectionist Temporal Classification) head.
**Text casing note:** training targets were kept in their raw, cased
form (same tokenizer/vocab used throughout this lineage since
`main-best`).
## Training data
Same pool and split usage as `main-best` and `main-best-3` — WAXAL
`validation` held out, `test` untouched:
| Source | Role |
|---|---|
| [google/WaxalNLP](https://huggingface.co/datasets/google/WaxalNLP) (`lin_asr` config) | `train` split pooled into training, `validation` split held out untouched as the fixed evaluation benchmark |
| [KasuleTrevor/Lingala_100hrs](https://huggingface.co/datasets/KasuleTrevor/Lingala_100hrs) | All splits pooled into training |
**WAXAL's `validation` split is the only data used for evaluation,
and it was never included in training, across this entire lineage.**
## Training procedure
- **Base checkpoint:** `keystats/w2v-bert-2.0-lingala-main-best-3`
(continued fine-tuning)
- **Architecture:** `Wav2Vec2BertForCTC`, `add_adapter=True`
- **Processor:** same `Wav2Vec2BertProcessor` / tokenizer used since
`main-best` (character-level, raw/cased text, `|` word delimiter,
`[PAD]` as CTC blank)
- **Sample rate:** 16 kHz mono
- **Epochs:** 1 additional epoch on top of `main-best-3`
- **Learning rate:** 1e-5, cosine schedule, 10% warmup
- **Effective batch size:** 32 (per-device batch size 4 × gradient
accumulation 8)
- **Precision:** fp16, gradient checkpointing enabled
- **Regularization:** attention/hidden/feature-projection dropout 0.05
- **Data filtering:** same CTC-feasibility filtering used throughout
this lineage
- **Seed:** 42 (deterministic — same seed for Python/NumPy/PyTorch/CUDA)
## Evaluation results
Evaluated on the WAXAL Lingala `validation` split, greedy decoding
vs. greedy + KenLM (`keystats/waxal-kenlm-models-best`). Adding the
KLM gives a consistent, meaningful WER/CER improvement over greedy
decoding alone — pair the two for the best results.
## How to use
### Option 1 — model alone (greedy decoding)
```python
import torch
import librosa
from transformers import Wav2Vec2BertForCTC, Wav2Vec2BertProcessor
MODEL_ID = "keystats/w2v-bert-2.0-lingala"
DEVICE = "cuda" if torch.cuda.is_available() else "cpu"
processor = Wav2Vec2BertProcessor.from_pretrained(MODEL_ID)
model = Wav2Vec2BertForCTC.from_pretrained(MODEL_ID).to(DEVICE).eval()
audio_array, sr = librosa.load("path/to/audio.wav", sr=16000, mono=True)
inputs = processor(audio_array, sampling_rate=16000, return_tensors="pt")
with torch.no_grad():
logits = model(input_features=inputs.input_features.to(DEVICE)).logits
predicted_ids = torch.argmax(logits, dim=-1)
transcription = processor.batch_decode(predicted_ids)[0]
print(transcription) # cased, punctuated Lingala text
```
### Option 2 — model + KLM (recommended, higher accuracy)
```python
# pip install pyctcdecode
# pip install https://github.com/kpu/kenlm/archive/master.zip
import torch
import librosa
from huggingface_hub import hf_hub_download
from transformers import Wav2Vec2BertForCTC, Wav2Vec2BertProcessor
from pyctcdecode import build_ctcdecoder
MODEL_ID = "keystats/w2v-bert-2.0-lingala"
KLM_REPO_ID = "keystats/waxal-kenlm-models-best"
DEVICE = "cuda" if torch.cuda.is_available() else "cpu"
processor = Wav2Vec2BertProcessor.from_pretrained(MODEL_ID)
model = Wav2Vec2BertForCTC.from_pretrained(MODEL_ID).to(DEVICE).eval()
klm_path = hf_hub_download(repo_id=KLM_REPO_ID, repo_type="dataset",
filename="lingala/lingala_5gram_correct-best.arpa")
def build_vocab_list(tokenizer, vocab_size):
vocab_dict = tokenizer.get_vocab()
vocab_list = [None] * vocab_size
for tok, idx in sorted(vocab_dict.items(), key=lambda kv: kv[1]):
if idx < vocab_size:
vocab_list[idx] = tok
pad_id = tokenizer.pad_token_id
if pad_id is not None and pad_id < len(vocab_list):
vocab_list[pad_id] = ""
word_delim = getattr(tokenizer, "word_delimiter_token", None)
if word_delim:
delim_id = vocab_dict.get(word_delim)
if delim_id is not None:
vocab_list[delim_id] = " "
return vocab_list
vocab_list = build_vocab_list(processor.tokenizer, model.config.vocab_size)
decoder = build_ctcdecoder(
vocab_list,
kenlm_model_path=klm_path,
alpha=0.5,
beta=0.7,
)
audio_array, sr = librosa.load("path/to/audio.wav", sr=16000, mono=True)
inputs = processor(audio_array, sampling_rate=16000, return_tensors="pt")
with torch.no_grad():
logits = model(input_features=inputs.input_features.to(DEVICE)).logits
transcription = decoder.decode(logits.cpu().numpy()[0], beam_width=100)
print(transcription)
```
## Intended uses & limitations
- Intended for transcribing spoken Lingala audio into cased,
punctuated text.
- As a CTC-based model, it assumes single-speaker, forward-only
audio and has no mechanism for overlapping speech from multiple
speakers.
- This is the fourth checkpoint in a continued fine-tuning lineage
(`main-best` → `main-best-3` → this model). If earlier checkpoints
had already converged, this additional low-LR epoch may yield only
marginal gains — compare validation metrics against `main-best-3`
before choosing between them.
## Citation
```bibtex
@misc{keystats_wav2vec2bert_lingala_final,
title={w2v-bert-2.0-lingala: A Lingala ASR model, continued fine-tune of w2v-bert-2.0-lingala-main-best-3},
author={keystats},
year={2026},
howpublished={\url{https://huggingface.co/keystats/w2v-bert-2.0-lingala}}
}
@misc{waxal,
title={WAXAL: A Multilingual African Speech Dataset},
author={Google},
howpublished={\url{https://huggingface.co/datasets/google/WaxalNLP}}
}
@misc{kasule_lingala_100hrs,
title={Lingala\_100hrs},
author={KasuleTrevor},
howpublished={\url{https://huggingface.co/datasets/KasuleTrevor/Lingala_100hrs}}
}
@inproceedings{w2vbert2,
title={Seamless: Multilingual Expressive and Streaming Speech Translation},
author={Seamless Communication and others},
year={2023},
howpublished={\url{https://huggingface.co/facebook/w2v-bert-2.0}}
}
```