How to use from the
Use from the
Transformers library
# Use a pipeline as a high-level helper
from transformers import pipeline

pipe = pipeline("feature-extraction", model="mlr2000/vocoder-large-speaker-encoder", trust_remote_code=True)
# Load model directly
from transformers import AutoModel
model = AutoModel.from_pretrained("mlr2000/vocoder-large-speaker-encoder", trust_remote_code=True, device_map="auto")
Quick Links

VocBulwark speaker encoder

Standalone speaker encoder: turns a raw waveform into a fixed 768-dimensional speaker embedding. It is the conditioning front-end for the companion VocBulwark vocoder โ€” compute an embedding once from a reference clip of a speaker, then pass it to the vocoder to synthesize in that voice. The embedding is also usable on its own for speaker verification / similarity.

Wav2Vec2-based encoder with attentive statistics pooling, trained with a GE2E objective. Self-contained: loads with trust_remote_code=True, no training repo required.

Companion Models

This speaker encoder is part of a set of 6 repositories:

Repo Role
mlr2000/vocoder-large Large vocoder (generates the watermarked audio)
mlr2000/vocoder-large-watermark-detector Watermark detector for the large model
mlr2000/vocoder-large-speaker-encoder Speaker encoder (this repo)
mlr2000/vocoder-small Small vocoder
mlr2000/vocoder-small-watermark-detector Watermark detector for the small model
mlr2000/vocoder-small-speaker-encoder Speaker encoder for the small model

Usage

import torchaudio, torchaudio.functional as AF
from transformers import AutoModel

enc = AutoModel.from_pretrained("mlr2000/vocoder-large-speaker-encoder", trust_remote_code=True).eval()

wav, sr = torchaudio.load("reference.wav")             # [C, T]
wav = wav.mean(0, keepdim=True)                        # mono [1, T]
if sr != enc.config.raw_sample_rate:                   # encoder expects 22.05 kHz
    wav = AF.resample(wav, sr, enc.config.raw_sample_rate)

emb = enc.embed(wav)          # [1, 768]  โ€” feed as speaker_embedding to the vocoder

See example_roundtrip.ipynb in this repo for the full pipeline (reference clip โ†’ embedding โ†’ vocode โ†’ detect watermark).

Notes

  • Input: mono waveform at 22.05 kHz (config.raw_sample_rate). Resample first if your audio differs.
  • Output: [B, 768] L2-comparable speaker embeddings.
  • Pair with the VocBulwark vocoder trained jointly with this encoder โ€” an embedding from a different encoder will not condition the vocoder correctly.
  • This encoder is paired with mlr2000/vocoder-large. For the small vocoder (mlr2000/vocoder-small) use the companion small speaker encoder (mlr2000/vocoder-small-speaker-encoder).
  • Not intended for speaker identification or surveillance of individuals without consent.

Citation

If you use this model, please cite:

@misc{muletta2026,
  title  = {Training a Discriminator-Free Foundation Vocoder 
             with Integrated Audio Watermarking},
  author = {Muletta, Romolo and Deriu, Jan},
  year   = {2026},
  note   = {VT2 Project Report, ZHAW School of Engineering}
}

License

cc-by-4.0. Trained on MLS (CC-BY-4.0) and Common Voice (CC0); builds on BigVGAN (MIT) and wav2vec 2.0 (Apache-2.0). Please retain attribution when redistributing or building on this model.

Downloads last month
6
Safetensors
Model size
18.7M params
Tensor type
F32
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support