---
language:
- plt
license: cc-by-nc-sa-4.0
pretty_name: Mimba PL-BERT PLT (Plateau Malagasy Phonetic-Level BERT)
tags:
- text-to-speech
- pl-bert
- plbert
- albert
- malagasy
- plt
- african-languages
- low-resource
- masked-language-modeling
- phoneme
library_name: transformers
---
# Mimba PL-BERT PLT — Phonetic-Level BERT for Plateau Malagasy
A **phonetic-level pre-trained language model (PL-BERT)** for **Plateau Malagasy (PLT)**,
trained from scratch to provide text/phoneme-context-aware embeddings for
**StyleTTS2** synthesis. Adapted from the original
[PL-BERT](https://github.com/yl4579/PL-BERT) architecture (Li et al., used in
StyleTTS2) and trained on a large phonemized Malagasy corpus with the same
55-symbol phoneme vocabulary used across every Mimba PLT model (StyleTTS2,
NeuTTS-Nano, Supertonic 3).
> ⚠️ **Not a standalone TTS model.** PL-BERT is a text/phoneme encoder only —
> it produces contextual embeddings consumed by a downstream acoustic model
> (StyleTTS2 Stage 1/Stage 2). It cannot synthesize audio by itself.
## Summary
| | |
|---|---|
| Language | Plateau Malagasy (`plt`) |
| Architecture | ALBERT (`transformers.AlbertModel` + 2 prediction heads) |
| Phoneme vocabulary | 55 symbols (`phoneme_symbols.pkl`, shared with StyleTTS2/NeuTTS-Nano) |
| Hidden size | 768 |
| Attention heads | 12 |
| Hidden layers | 12 |
| Intermediate size | 2048 |
| Max position embeddings | 512 |
| Dropout | 0.1 |
| Training objective | Masked language modeling, dual head (phoneme-level + word-level) |
| Training steps | 1,000,000 |
| Final masked-phoneme accuracy | **65.52%** (measured on ~200K masked positions) |
| Checkpoint format | `step_{N}.t7` (`{'net': state_dict, 'optimizer': ..., 'step': N}`) |
### Loss curve
## Training details
The model is trained with a dual masked-language-modeling objective — one
head predicts the masked **phoneme token** (55-way classification), the other
predicts the masked **word form** (large open vocabulary of Malagasy word
forms, zipfian-distributed due to the language's agglutinative morphology).
Only the phoneme-level task is used downstream by StyleTTS2, but the joint
objective helps the encoder learn richer contextual representations.
Training ran for 1,000,000 steps with a cosine learning-rate decay
(`1e-4 → 1e-6`) applied over the final ~270K steps. Accuracy on masked
phoneme positions (measured periodically on held-out batches, not just
training loss) tracked as follows:
| Step | Masked-phoneme accuracy |
|---|---|
| 136,000 | 58.55% |
| 727,514 | 61–63% |
| 900,000 | 63.92% |
| 1,000,000 | **65.52%** |
Accuracy plateaued in the final third of training despite the LR decay
reaching down to `1e-6` — this is treated as the effective ceiling for this
model size/corpus, not a sign that more steps would help. For reference,
comparable phoneme/sup-phoneme masked-LM setups in other languages (e.g.
Mixed-Phoneme BERT, PnG-BERT) report converged accuracies around 70–75%;
this PLT model sits somewhat below that range, likely due to corpus size and
language-specific factors rather than an implementation issue.
## Usage
```python
import torch, yaml
from transformers import AlbertConfig, AlbertModel
from huggingface_hub import hf_hub_download
from collections import OrderedDict
REPO = "mimba/plbert-plt"
class CustomAlbert(AlbertModel):
def forward(self, *args, **kwargs):
return super().forward(*args, **kwargs).last_hidden_state
def load_plbert(repo_id=REPO, step=1_000_000):
config_path = hf_hub_download(repo_id, "config.yml")
plbert_config = yaml.safe_load(open(config_path))
config = AlbertConfig(**plbert_config["model_params"])
bert = CustomAlbert(config)
ckpt_path = hf_hub_download(repo_id, f"step_{step}.t7")
checkpoint = torch.load(ckpt_path, map_location="cpu")
state_dict = checkpoint["net"]
new_state_dict = OrderedDict()
for k, v in state_dict.items():
name = k[7:] if k.startswith("module.") else k
if name.startswith("encoder."):
new_state_dict[name[8:]] = v
new_state_dict.pop("embeddings.position_ids", None)
bert.load_state_dict(new_state_dict, strict=False)
return bert
model = load_plbert()
model.eval()
```
**As a StyleTTS2 `PLBERT_dir`**: download `config.yml` + `step_1000000.t7`
into `Utils/PLBERT/` of the StyleTTS2 repo — `util.py`'s `load_plbert()`
(shown above) is what `train_first.py`/`train_second.py` call automatically.
## Relation to other Mimba datasets/models
```
mimba/text2text (source text corpus)
-> mimba/plt-tts-dataset (audio + text, 4 speakers)
-> phonemized PLT corpus (IPA phonemization, mode PHRASE)
-> mimba/plbert-plt <- this model
-> mimba/styletts2-plt-corpus (StyleTTS2-ready corpus)
-> mimba/styletts2-plt-stage1 / stage2 (StyleTTS2 checkpoints)
```
## Limitations
- Masked-phoneme accuracy (65.52%) is below reference points from other
languages' phoneme-level BERT models (~70–75%); treat this as this
model's practical ceiling rather than an intermediate result.
- The word-level prediction head operates over a very large, zipfian
vocabulary (agglutinative morphology) and is noisy on rare word forms —
this does not affect StyleTTS2 usage, which only consumes phoneme-level
embeddings.
- Trained on synthetic/derived text sources (see `mimba/text2text` and
`mimba/plt-tts-dataset` cards for provenance); verify licensing
independently before commercial use.
## Citation
```bibtex
@misc{mimba2026plbertplt,
title = {Mimba PL-BERT PLT: A Phonetic-Level BERT for Plateau Malagasy},
author = {Mimba Ngouana Fofou},
year = {2026},
}
```
### Contact
For questions or contributions, open a discussion in the "Community" tab of
this repository.
##### *Contact: [@Mimba](baounabaouna@gmail.com)*