maubert-phone / README.md
angelo-ortiz's picture
Upload folder using huggingface_hub
2cb426c verified
|
Raw
History Blame Contribute Delete
2.79 kB
metadata
library_name: maubert
license: mit
tags:
  - acoustic-unit-discovery
  - hubert
  - maubert
  - speech
language_bcp47:
  - ab
  - am
  - ba
  - be
  - bg
  - bn
  - ca
  - ckb
  - cs
  - cv
  - dv
  - el
  - eu
  - gn
  - ha
  - hi
  - hsb
  - hu
  - hy-AM
  - id
  - it
  - ja
  - ka
  - kk
  - kmr
  - ko
  - ky
  - lij
  - lt
  - ml
  - mn
  - mr
  - mt
  - myv
  - nan-tw
  - nl
  - pa-IN
  - pl
  - pt
  - ro
  - ru
  - rw
  - sk
  - sl
  - sq
  - sr
  - sv-SE
  - tk
  - tt
  - ug
  - ur
  - uz
  - vi
  - yo
  - yue

MauBERT-phone

MauBERT-phone is a speech model adapting a pre-trained HuBERT base model to predict phones on a subset of 788 hours and 55 languages of Common Voice 16.1 leveraging the phone annotations from VoxCommunis. It was introduced by the MauBERT paper at ACL 2026.

You can load it with maubert:

from maubert import MauBERTPhone

model = MauBERTPhone.from_pretrained("coml/maubert-phone")

Check out the maubert repository if you are interested in further retraining or fine-tuning, and the corresponding paper for our results.

Files:

  • model.safetensors and config.json: HuggingFace Transformers checkpoint and config.

The manifests, phone alignments, language table and phone inventories used to train this model are released separately in coml/vox-communis-artifacts. The inventories are required to reduce the phone head when extracting features from the phone projection layer:

from huggingface_hub import snapshot_download

data_dir = snapshot_download("coml/vox-communis-artifacts", repo_type="dataset")
# then pass f"{data_dir}/canonical_inventories.jsonl" as `inventory_file`

Citing

@inproceedings{ortiztandazo-etal-2026-maubert,
    title = "{M}au{BERT}: Universal Phonetic Inductive Biases for Few-Shot Acoustic Units Discovery",
    author = "Ortiz Tandazo, Angelo  and
      Khentout, Manel  and
      Benchekroun, Youssef  and
      Hueber, Thomas  and
      Dupoux, Emmanuel",
    editor = "Liakata, Maria  and
      Moreira, Viviane P.  and
      Zhang, Jiajun  and
      Jurgens, David",
    booktitle = "Proceedings of the 64th Annual Meeting of the {A}ssociation for {C}omputational {L}inguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2026",
    address = "San Diego, California, United States",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2026.acl-long.24/",
    doi = "10.18653/v1/2026.acl-long.24",
    pages = "568--585",
    ISBN = "979-8-89176-390-6",
}