Mimba PL-BERT PLT β€” Phonetic-Level BERT for Plateau Malagasy

A phonetic-level pre-trained language model (PL-BERT) for Plateau Malagasy (PLT), trained from scratch to provide text/phoneme-context-aware embeddings for StyleTTS2 synthesis. Adapted from the original PL-BERT architecture (Li et al., used in StyleTTS2) and trained on a large phonemized Malagasy corpus with the same 55-symbol phoneme vocabulary used across every Mimba PLT model (StyleTTS2, NeuTTS-Nano, Supertonic 3).

⚠️ Not a standalone TTS model. PL-BERT is a text/phoneme encoder only β€” it produces contextual embeddings consumed by a downstream acoustic model (StyleTTS2 Stage 1/Stage 2). It cannot synthesize audio by itself.

Summary

Language Plateau Malagasy (plt)
Architecture ALBERT (transformers.AlbertModel + 2 prediction heads)
Phoneme vocabulary 55 symbols (phoneme_symbols.pkl, shared with StyleTTS2/NeuTTS-Nano)
Hidden size 768
Attention heads 12
Hidden layers 12
Intermediate size 2048
Max position embeddings 512
Dropout 0.1
Training objective Masked language modeling, dual head (phoneme-level + word-level)
Training steps 1,000,000
Final masked-phoneme accuracy 65.52% (measured on ~200K masked positions)
Checkpoint format step_{N}.t7 ({'net': state_dict, 'optimizer': ..., 'step': N})

Loss curve

Loss curve PL BERT Loss curve PL BERT Loss curve PL BERT

Training details

The model is trained with a dual masked-language-modeling objective β€” one head predicts the masked phoneme token (55-way classification), the other predicts the masked word form (large open vocabulary of Malagasy word forms, zipfian-distributed due to the language's agglutinative morphology). Only the phoneme-level task is used downstream by StyleTTS2, but the joint objective helps the encoder learn richer contextual representations.

Training ran for 1,000,000 steps with a cosine learning-rate decay (1e-4 β†’ 1e-6) applied over the final ~270K steps. Accuracy on masked phoneme positions (measured periodically on held-out batches, not just training loss) tracked as follows:

Step Masked-phoneme accuracy
136,000 58.55%
727,514 61–63%
900,000 63.92%
1,000,000 65.52%

Accuracy plateaued in the final third of training despite the LR decay reaching down to 1e-6 β€” this is treated as the effective ceiling for this model size/corpus, not a sign that more steps would help. For reference, comparable phoneme/sup-phoneme masked-LM setups in other languages (e.g. Mixed-Phoneme BERT, PnG-BERT) report converged accuracies around 70–75%; this PLT model sits somewhat below that range, likely due to corpus size and language-specific factors rather than an implementation issue.

Usage

import torch, yaml
from transformers import AlbertConfig, AlbertModel
from huggingface_hub import hf_hub_download
from collections import OrderedDict

REPO = "mimba/plbert-plt"

class CustomAlbert(AlbertModel):
    def forward(self, *args, **kwargs):
        return super().forward(*args, **kwargs).last_hidden_state

def load_plbert(repo_id=REPO, step=1_000_000):
    config_path = hf_hub_download(repo_id, "config.yml")
    plbert_config = yaml.safe_load(open(config_path))
    config = AlbertConfig(**plbert_config["model_params"])
    bert = CustomAlbert(config)

    ckpt_path = hf_hub_download(repo_id, f"step_{step}.t7")
    checkpoint = torch.load(ckpt_path, map_location="cpu")
    state_dict = checkpoint["net"]

    new_state_dict = OrderedDict()
    for k, v in state_dict.items():
        name = k[7:] if k.startswith("module.") else k
        if name.startswith("encoder."):
            new_state_dict[name[8:]] = v
    new_state_dict.pop("embeddings.position_ids", None)
    bert.load_state_dict(new_state_dict, strict=False)
    return bert

model = load_plbert()
model.eval()

As a StyleTTS2 PLBERT_dir: download config.yml + step_1000000.t7 into Utils/PLBERT/ of the StyleTTS2 repo β€” util.py's load_plbert() (shown above) is what train_first.py/train_second.py call automatically.

Relation to other Mimba datasets/models

mimba/text2text                     (source text corpus)
   -> mimba/plt-tts-dataset          (audio + text, 4 speakers)
   -> phonemized PLT corpus          (IPA phonemization, mode PHRASE)
   -> mimba/plbert-plt                <- this model
   -> mimba/styletts2-plt-corpus     (StyleTTS2-ready corpus)
   -> mimba/styletts2-plt-stage1 / stage2   (StyleTTS2 checkpoints)

Limitations

  • Masked-phoneme accuracy (65.52%) is below reference points from other languages' phoneme-level BERT models (~70–75%); treat this as this model's practical ceiling rather than an intermediate result.
  • The word-level prediction head operates over a very large, zipfian vocabulary (agglutinative morphology) and is noisy on rare word forms β€” this does not affect StyleTTS2 usage, which only consumes phoneme-level embeddings.
  • Trained on synthetic/derived text sources (see mimba/text2text and mimba/plt-tts-dataset cards for provenance); verify licensing independently before commercial use.

Citation

@misc{mimba2026plbertplt,
  title  = {Mimba PL-BERT PLT: A Phonetic-Level BERT for Plateau Malagasy},
  author = {Mimba Ngouana Fofou},
  year   = {2026},
}

Contact

For questions or contributions, open a discussion in the "Community" tab of this repository.

Contact: @Mimba
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support