DiscoPhon finetuned baselines

Finetuned checkpoints of the four DiscoPhon baselines. Each pretrained model is finetuned on each of the 12 benchmark languages, with 10 minutes, 1 hour, or 10 hours of training data: 4 × 12 × 3 = 144 checkpoints.

Key Pretrained model
spidr-mmsulab coml/spidr-mmsulab
spidr-vp20 coml/spidr-vp20
hubert-mmsulab-it2 coml/hubert-base-mmsulab
hubert-vp20-it2 coml/hubert-base-vp20

Languages: cmn, deu, eng, eus, fra, jpn, swa, tam, tha, tur, ukr, wol.

Layout

{key}/ft-{language}-{duration}/

with duration in 10min, 1h, 10h. For example, spidr-vp20/ft-deu-1h/ is SpidR VP-20 finetuned on 1 hour of German. Each directory holds the checkpoint selected on validation (best.pt) and the validation scores. SpidR directories also hold the model config.json, the same as the pretrained model's, which spidr reads from the directory of the checkpoint. For HuBERT the K-means used to compute the finetuning targets is in the repository of the pretrained model.

The units and scores of these models on the benchmark are in the artifacts dataset, under the same names (e.g. spidr-vp20/ft-deu-1h/).

Usage

Download one checkpoint:

from huggingface_hub import hf_hub_download

path = hf_hub_download("coml/discophon-finetuned-baselines", "spidr-vp20/ft-deu-1h/best.pt")
hf_hub_download("coml/discophon-finetuned-baselines", "spidr-vp20/ft-deu-1h/config.json")  # SpidR only

or all the checkpoints of one model:

hf download coml/discophon-finetuned-baselines --include "spidr-vp20/*" --local-dir discophon-finetuned-baselines

Load it with spidr or minimal_hubert, or extract units and features directly with discophon.baselines:

from minimal_hubert import HuBERT
from spidr.models import build_model

spidr = build_model(model_type="spidr", checkpoint=path)  # SpidR
hubert = HuBERT.from_pretrained(path)  # HuBERT

See the baselines guide for the finetuning recipe.

Citation

@inproceedings{poli2026discophon,
  title     = {{DiscoPhon: Benchmarking the Unsupervised Discovery of Phoneme Inventories With Discrete Speech Units}},
  author    = {Maxime Poli and Manel Khentout and Angelo {Ortiz Tandazo} and Ewan Dunbar and Emmanuel Chemla and Emmanuel Dupoux},
  year      = {2026},
  booktitle = {{Interspeech 2026}},
  pages     = {6664--6669},
  doi       = {10.21437/Interspeech.2026-2791},
  issn      = {2958-1796},
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including coml/discophon-finetuned-baselines