| ---
|
| language:
|
| - gl
|
| license: apache-2.0
|
| tags:
|
| - phoneme-level
|
| - modernbert
|
| - galician
|
| - text-to-speech
|
| - tts
|
| - masked-lm
|
| - proxecto-nos
|
| metrics:
|
| - precision
|
| ---
|
|
|
| # PL-ModernBERT-gl: Phoneme-level ModernBERT for Galician
|
|
|
| ## Overview
|
| - [Model Description](#model-description)
|
| - [Intended Uses and Limitations](#intended-uses-and-limitations)
|
| - [Training Details](#training-details)
|
| - [Evaluation](#evaluation)
|
| - [Citation](#citation)
|
| - [Additional Information](#additional-information)
|
| - [Funding and Acknowledgements](#funding-and-acknowledgements)
|
| ---
|
|
|
| ## Model Description
|
|
|
| **PL-ModernBERT-gl** is a phoneme-level masked language model trained on Galician text. It is based on the state-of-the-art **ModernBERT** architecture, adapting the Phoneme-Level BERT (PL-BERT) framework to learn contextualized phoneme representations via masked language modeling and alignment objectives.
|
|
|
| This model is designed to support phoneme-based text-to-speech (TTS) systems, including but not limited to *StyleTTS2*. Thanks to its Galician-specific phoneme vocabulary and contextual embedding capabilities, it can serve as a high-precision phoneme encoder for any TTS architecture requiring phoneme-level features.
|
|
|
| ### Key Improvements & Features:
|
| * **Native Galician Pipeline:** Unlike the original PL-BERT architecture which relied on English phonemizers, this model integrates the open-source linguistic tool **Cotovía** for native Galician grapheme-to-phoneme transcription and text normalization.
|
| * **1:1 Alignment System:** Implements a strict sequential alignment between graphemes and phonemes, successfully handling Galician digraphs (e.g., `ll`, `rr`, `ch`, `nh`, `qu`), silent characters (e.g., silent `h`), and morphosyntactic contractions (e.g., `para` + `a` → `pra`).
|
| * **Dual-Head Architecture:** The core encoder branches into two parallel prediction layers during training:
|
| * **MLM Head (Masked Language Modeling):** Predicts the identity of masked phonemes.
|
| * **P2G Head (Phoneme-to-Grapheme):** Predicts the corresponding grapheme for aligned feature extraction.
|
|
|
| ---
|
|
|
| ## Intended Uses and Limitations
|
|
|
| ### Intended uses
|
| * Integration into phoneme-based Galician TTS pipelines (such as *StyleTTS2*).
|
| * Contextualized phoneme embedding extraction for downstream speech and linguistic tasks in the Galician language.
|
|
|
| ### Limitations
|
| * Not designed for general-purpose word-level NLP tasks (e.g., standard text classification, Named Entity Recognition, or sentiment analysis).
|
| * Strictly supports Galician phoneme tokens mapped through the provided tokenizer utilities.
|
|
|
| ---
|
|
|
| ## Training Details
|
|
|
| ### Training Data
|
| The model was pre-trained on a curated Galician corpus of **11.6 million words** compiled from diverse domains under open licenses:
|
|
|
| | Corpus Source | Domain | Word Count | License |
|
| | :--- | :--- | :---: | :--- |
|
| | *Enciclopedia Galega Universal* | Encyclopedic | 4,751,813 | CC-BY-4.0 |
|
| | *URCO Editora* | Literature / Novels | 1,937,517 | CC-BY-4.0 |
|
| | *Wikipedia* | Encyclopedic | 1,794,937 | CC-BY-SA-4.0 |
|
| | *Mallando no Android* | Technology Blog | 1,553,327 | CC-BY-4.0 |
|
| | *Revista Pincha* | Culture Blog | 669,020 | CC-BY-4.0 |
|
| | *A Nosa Terra* | Journalistic | 652,372 | CC-BY-4.0 |
|
| | *Servizo Publicacións USC* | Academic / Scientific | 280,858 | CC-BY-4.0 |
|
| | **Total** | | **11,639,844** | |
|
|
|
| #### Prosodic Enrichment
|
| To capture expressive prosody and speech modulation, the training corpus was enriched with two dedicated subsets of **66,338 interrogative sentences** and **11,546 exclamative sentences**.
|
|
|
| ### Dynamic Masking Strategy
|
| To force the model to prioritize prosodic clues over highly frequent phonemes, an inverse-frequency dynamic masking algorithm was implemented:
|
| * **Common Phonemes:** 15% masking probability.
|
| * **Standard Punctuation Marks:** 40% masking probability.
|
| * **Critical Prosodic Marks (`!`, `?`):** 80% masking probability.
|
|
|
| ### Training Configuration
|
|
|
| | Parameter | Value |
|
| | :--- | :--- |
|
| | **Model Type / Core Architecture** | ModernBERT (12 layers, 12 attention heads) |
|
| | **Vocabulary Size** | 69 |
|
| | **Hidden Size** | 768 |
|
| | **Intermediate Size (FFN)** | 2048 |
|
| | **Dropout** | 0.1 |
|
| | **Batch Size** | 192 |
|
| | **Total Steps** | 1,000,000 |
|
| | **Precision** | Mixed Precision (`fp16`) |
|
| | **Initial Learning Rate** | 1e-4 |
|
| | **Scheduler Type** | `onecycle` (Cos annealing strategy, Warmup ratio: 0.1) |
|
| | **Max Sequence Length** | 512 |
|
| | **Base Word Mask Probability** | 0.15 |
|
| | **Base Phoneme Mask Probability** | 0.1 |
|
| | **Replacement Probability** | 0.2 |
|
|
|
| ## Evaluation
|
| The model was evaluated using an independent test partition from the core phonemized Galician dataset. Performance was tracked via Masked Language Modeling accuracy (predicting masked phonemes) and Phoneme-to-Grapheme symbolic alignment metrics (predicting corresponding characters):
|
|
|
| | Metric | Value |
|
| |-------------------|------|
|
| | **MLM Precision** | 89.54% |
|
| | **P2G Precision** | 97.88% |
|
|
|
| ## Citation
|
| If this model contributes to your research, please cite it as follows:
|
|
|
| ```bibtex
|
|
|
| @misc{proxectonos/PL-ModernBERT-gl,
|
| author = {{Proxecto Nós}},
|
| title = {{PL-ModernBERT-gl: Phoneme-level ModernBERT for Galician}},
|
| year = {2026},
|
| publisher = {Hugging Face},
|
| howpublished = {\url{[https://huggingface.co/proxectonos/PL-ModernBERT-gl](https://huggingface.co/proxectonos/PL-ModernBERT-gl)}},
|
| }
|
|
|
| ```
|
|
|
| ## Additional Information
|
|
|
| ### Licensing
|
| This model is licensed under the **Apache License 2.0**.
|
|
|
| ### Authors and Credits
|
| * **Project Oversight:** [Proxecto Nós](https://nos.gal/gl/proxecto-nos)
|
| * **Technical Development:** [Gradiant](https://www.gradiant.org/) (Centro Tecnolóxico de Telecomunicacións de Galicia)
|
|
|
| ## Funding and Acknowledgements
|
| This work is funded by the Ministerio para la Transformación Digital y de la Función Pública - Funded by EU – NextGenerationEU within the framework of the project Desarrollo de Modelos ALIA.
|
|
|
| We would like to express our gratitude to the engineering and research teams at **Gradiant** for the technical development of this model, as well as to the **Aholab Signal Processing Laboratory (HiTZ)** and the **Language Technologies Laboratory (BSC)** for their technical support and collaboration. |