PL-ModernBERT-gl / README.md
cmagui's picture
Update model, config.yml, and README
50c5641 verified
|
Raw
History Blame Contribute Delete
6.39 kB
---
language:
- gl
license: apache-2.0
tags:
- phoneme-level
- modernbert
- galician
- text-to-speech
- tts
- masked-lm
- proxecto-nos
metrics:
- precision
---
# PL-ModernBERT-gl: Phoneme-level ModernBERT for Galician
## Overview
- [Model Description](#model-description)
- [Intended Uses and Limitations](#intended-uses-and-limitations)
- [Training Details](#training-details)
- [Evaluation](#evaluation)
- [Citation](#citation)
- [Additional Information](#additional-information)
- [Funding and Acknowledgements](#funding-and-acknowledgements)
---
## Model Description
**PL-ModernBERT-gl** is a phoneme-level masked language model trained on Galician text. It is based on the state-of-the-art **ModernBERT** architecture, adapting the Phoneme-Level BERT (PL-BERT) framework to learn contextualized phoneme representations via masked language modeling and alignment objectives.
This model is designed to support phoneme-based text-to-speech (TTS) systems, including but not limited to *StyleTTS2*. Thanks to its Galician-specific phoneme vocabulary and contextual embedding capabilities, it can serve as a high-precision phoneme encoder for any TTS architecture requiring phoneme-level features.
### Key Improvements & Features:
* **Native Galician Pipeline:** Unlike the original PL-BERT architecture which relied on English phonemizers, this model integrates the open-source linguistic tool **Cotovía** for native Galician grapheme-to-phoneme transcription and text normalization.
* **1:1 Alignment System:** Implements a strict sequential alignment between graphemes and phonemes, successfully handling Galician digraphs (e.g., `ll`, `rr`, `ch`, `nh`, `qu`), silent characters (e.g., silent `h`), and morphosyntactic contractions (e.g., `para` + `a``pra`).
* **Dual-Head Architecture:** The core encoder branches into two parallel prediction layers during training:
* **MLM Head (Masked Language Modeling):** Predicts the identity of masked phonemes.
* **P2G Head (Phoneme-to-Grapheme):** Predicts the corresponding grapheme for aligned feature extraction.
---
## Intended Uses and Limitations
### Intended uses
* Integration into phoneme-based Galician TTS pipelines (such as *StyleTTS2*).
* Contextualized phoneme embedding extraction for downstream speech and linguistic tasks in the Galician language.
### Limitations
* Not designed for general-purpose word-level NLP tasks (e.g., standard text classification, Named Entity Recognition, or sentiment analysis).
* Strictly supports Galician phoneme tokens mapped through the provided tokenizer utilities.
---
## Training Details
### Training Data
The model was pre-trained on a curated Galician corpus of **11.6 million words** compiled from diverse domains under open licenses:
| Corpus Source | Domain | Word Count | License |
| :--- | :--- | :---: | :--- |
| *Enciclopedia Galega Universal* | Encyclopedic | 4,751,813 | CC-BY-4.0 |
| *URCO Editora* | Literature / Novels | 1,937,517 | CC-BY-4.0 |
| *Wikipedia* | Encyclopedic | 1,794,937 | CC-BY-SA-4.0 |
| *Mallando no Android* | Technology Blog | 1,553,327 | CC-BY-4.0 |
| *Revista Pincha* | Culture Blog | 669,020 | CC-BY-4.0 |
| *A Nosa Terra* | Journalistic | 652,372 | CC-BY-4.0 |
| *Servizo Publicacións USC* | Academic / Scientific | 280,858 | CC-BY-4.0 |
| **Total** | | **11,639,844** | |
#### Prosodic Enrichment
To capture expressive prosody and speech modulation, the training corpus was enriched with two dedicated subsets of **66,338 interrogative sentences** and **11,546 exclamative sentences**.
### Dynamic Masking Strategy
To force the model to prioritize prosodic clues over highly frequent phonemes, an inverse-frequency dynamic masking algorithm was implemented:
* **Common Phonemes:** 15% masking probability.
* **Standard Punctuation Marks:** 40% masking probability.
* **Critical Prosodic Marks (`!`, `?`):** 80% masking probability.
### Training Configuration
| Parameter | Value |
| :--- | :--- |
| **Model Type / Core Architecture** | ModernBERT (12 layers, 12 attention heads) |
| **Vocabulary Size** | 69 |
| **Hidden Size** | 768 |
| **Intermediate Size (FFN)** | 2048 |
| **Dropout** | 0.1 |
| **Batch Size** | 192 |
| **Total Steps** | 1,000,000 |
| **Precision** | Mixed Precision (`fp16`) |
| **Initial Learning Rate** | 1e-4 |
| **Scheduler Type** | `onecycle` (Cos annealing strategy, Warmup ratio: 0.1) |
| **Max Sequence Length** | 512 |
| **Base Word Mask Probability** | 0.15 |
| **Base Phoneme Mask Probability** | 0.1 |
| **Replacement Probability** | 0.2 |
## Evaluation
The model was evaluated using an independent test partition from the core phonemized Galician dataset. Performance was tracked via Masked Language Modeling accuracy (predicting masked phonemes) and Phoneme-to-Grapheme symbolic alignment metrics (predicting corresponding characters):
| Metric | Value |
|-------------------|------|
| **MLM Precision** | 89.54% |
| **P2G Precision** | 97.88% |
## Citation
If this model contributes to your research, please cite it as follows:
```bibtex
@misc{proxectonos/PL-ModernBERT-gl,
author = {{Proxecto Nós}},
title = {{PL-ModernBERT-gl: Phoneme-level ModernBERT for Galician}},
year = {2026},
publisher = {Hugging Face},
howpublished = {\url{[https://huggingface.co/proxectonos/PL-ModernBERT-gl](https://huggingface.co/proxectonos/PL-ModernBERT-gl)}},
}
```
## Additional Information
### Licensing
This model is licensed under the **Apache License 2.0**.
### Authors and Credits
* **Project Oversight:** [Proxecto Nós](https://nos.gal/gl/proxecto-nos)
* **Technical Development:** [Gradiant](https://www.gradiant.org/) (Centro Tecnolóxico de Telecomunicacións de Galicia)
## Funding and Acknowledgements
This work is funded by the Ministerio para la Transformación Digital y de la Función Pública - Funded by EU – NextGenerationEU within the framework of the project Desarrollo de Modelos ALIA.
We would like to express our gratitude to the engineering and research teams at **Gradiant** for the technical development of this model, as well as to the **Aholab Signal Processing Laboratory (HiTZ)** and the **Language Technologies Laboratory (BSC)** for their technical support and collaboration.