Text-to-Speech
Transformers
Plateau Malagasy
pl-bert
plbert
albert
malagasy
african-languages
low-resource
masked-language-modeling
phoneme
Instructions to use mimba/plbert-plt with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use mimba/plbert-plt with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-to-speech", model="mimba/plbert-plt")# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("mimba/plbert-plt", device_map="auto") - Notebooks
- Google Colab
- Kaggle
File size: 6,553 Bytes
64042a3 4a9920f 64042a3 4a9920f | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 | ---
language:
- plt
license: cc-by-nc-sa-4.0
pretty_name: Mimba PL-BERT PLT (Plateau Malagasy Phonetic-Level BERT)
tags:
- text-to-speech
- pl-bert
- plbert
- albert
- malagasy
- plt
- african-languages
- low-resource
- masked-language-modeling
- phoneme
library_name: transformers
---
# Mimba PL-BERT PLT — Phonetic-Level BERT for Plateau Malagasy
A **phonetic-level pre-trained language model (PL-BERT)** for **Plateau Malagasy (PLT)**,
trained from scratch to provide text/phoneme-context-aware embeddings for
**StyleTTS2** synthesis. Adapted from the original
[PL-BERT](https://github.com/yl4579/PL-BERT) architecture (Li et al., used in
StyleTTS2) and trained on a large phonemized Malagasy corpus with the same
55-symbol phoneme vocabulary used across every Mimba PLT model (StyleTTS2,
NeuTTS-Nano, Supertonic 3).
> ⚠️ **Not a standalone TTS model.** PL-BERT is a text/phoneme encoder only —
> it produces contextual embeddings consumed by a downstream acoustic model
> (StyleTTS2 Stage 1/Stage 2). It cannot synthesize audio by itself.
## Summary
| | |
|---|---|
| Language | Plateau Malagasy (`plt`) |
| Architecture | ALBERT (`transformers.AlbertModel` + 2 prediction heads) |
| Phoneme vocabulary | 55 symbols (`phoneme_symbols.pkl`, shared with StyleTTS2/NeuTTS-Nano) |
| Hidden size | 768 |
| Attention heads | 12 |
| Hidden layers | 12 |
| Intermediate size | 2048 |
| Max position embeddings | 512 |
| Dropout | 0.1 |
| Training objective | Masked language modeling, dual head (phoneme-level + word-level) |
| Training steps | 1,000,000 |
| Final masked-phoneme accuracy | **65.52%** (measured on ~200K masked positions) |
| Checkpoint format | `step_{N}.t7` (`{'net': state_dict, 'optimizer': ..., 'step': N}`) |
### Loss curve
<!DOCTYPE html>
<html>
<head>
<style>
.conteneur-images {
display: flex; /* Active le mode horizontal */
gap: 10px; /* Espace de 10px entre les images */
}
.conteneur-images img {
width: 25%; /* 100% / 4 images = 25% */
height: auto; /* Maintient les proportions */
flex-shrink: 1; /* Permet de rétrécir si besoin */
}
</style>
</head>
<body>
<p class="conteneur-images">
<img src="assets/0.png" width="900" alt="Loss curve PL BERT">
<img src="assets/1.png" width="900" alt="Loss curve PL BERT">
<img src="assets/2.png" width="900" alt="Loss curve PL BERT">
</p>
</body>
</html>
## Training details
The model is trained with a dual masked-language-modeling objective — one
head predicts the masked **phoneme token** (55-way classification), the other
predicts the masked **word form** (large open vocabulary of Malagasy word
forms, zipfian-distributed due to the language's agglutinative morphology).
Only the phoneme-level task is used downstream by StyleTTS2, but the joint
objective helps the encoder learn richer contextual representations.
Training ran for 1,000,000 steps with a cosine learning-rate decay
(`1e-4 → 1e-6`) applied over the final ~270K steps. Accuracy on masked
phoneme positions (measured periodically on held-out batches, not just
training loss) tracked as follows:
| Step | Masked-phoneme accuracy |
|---|---|
| 136,000 | 58.55% |
| 727,514 | 61–63% |
| 900,000 | 63.92% |
| 1,000,000 | **65.52%** |
Accuracy plateaued in the final third of training despite the LR decay
reaching down to `1e-6` — this is treated as the effective ceiling for this
model size/corpus, not a sign that more steps would help. For reference,
comparable phoneme/sup-phoneme masked-LM setups in other languages (e.g.
Mixed-Phoneme BERT, PnG-BERT) report converged accuracies around 70–75%;
this PLT model sits somewhat below that range, likely due to corpus size and
language-specific factors rather than an implementation issue.
## Usage
```python
import torch, yaml
from transformers import AlbertConfig, AlbertModel
from huggingface_hub import hf_hub_download
from collections import OrderedDict
REPO = "mimba/plbert-plt"
class CustomAlbert(AlbertModel):
def forward(self, *args, **kwargs):
return super().forward(*args, **kwargs).last_hidden_state
def load_plbert(repo_id=REPO, step=1_000_000):
config_path = hf_hub_download(repo_id, "config.yml")
plbert_config = yaml.safe_load(open(config_path))
config = AlbertConfig(**plbert_config["model_params"])
bert = CustomAlbert(config)
ckpt_path = hf_hub_download(repo_id, f"step_{step}.t7")
checkpoint = torch.load(ckpt_path, map_location="cpu")
state_dict = checkpoint["net"]
new_state_dict = OrderedDict()
for k, v in state_dict.items():
name = k[7:] if k.startswith("module.") else k
if name.startswith("encoder."):
new_state_dict[name[8:]] = v
new_state_dict.pop("embeddings.position_ids", None)
bert.load_state_dict(new_state_dict, strict=False)
return bert
model = load_plbert()
model.eval()
```
**As a StyleTTS2 `PLBERT_dir`**: download `config.yml` + `step_1000000.t7`
into `Utils/PLBERT/` of the StyleTTS2 repo — `util.py`'s `load_plbert()`
(shown above) is what `train_first.py`/`train_second.py` call automatically.
## Relation to other Mimba datasets/models
```
mimba/text2text (source text corpus)
-> mimba/plt-tts-dataset (audio + text, 4 speakers)
-> phonemized PLT corpus (IPA phonemization, mode PHRASE)
-> mimba/plbert-plt <- this model
-> mimba/styletts2-plt-corpus (StyleTTS2-ready corpus)
-> mimba/styletts2-plt-stage1 / stage2 (StyleTTS2 checkpoints)
```
## Limitations
- Masked-phoneme accuracy (65.52%) is below reference points from other
languages' phoneme-level BERT models (~70–75%); treat this as this
model's practical ceiling rather than an intermediate result.
- The word-level prediction head operates over a very large, zipfian
vocabulary (agglutinative morphology) and is noisy on rare word forms —
this does not affect StyleTTS2 usage, which only consumes phoneme-level
embeddings.
- Trained on synthetic/derived text sources (see `mimba/text2text` and
`mimba/plt-tts-dataset` cards for provenance); verify licensing
independently before commercial use.
## Citation
```bibtex
@misc{mimba2026plbertplt,
title = {Mimba PL-BERT PLT: A Phonetic-Level BERT for Plateau Malagasy},
author = {Mimba Ngouana Fofou},
year = {2026},
}
```
### Contact
For questions or contributions, open a discussion in the "Community" tab of
this repository.
##### *Contact: [@Mimba](baounabaouna@gmail.com)*
|