Text-to-Speech
Transformers
Safetensors
hifi_gan
feature-extraction
audio
vocoder
hifi-gan
bigvgan
neural-vocoder
speaker-conditioning
audio-watermarking
custom_code
Instructions to use mlr2000/vocoder-small with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use mlr2000/vocoder-small with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-to-speech", model="mlr2000/vocoder-small", trust_remote_code=True)# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("mlr2000/vocoder-small", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
| library_name: transformers | |
| tags: | |
| - audio | |
| - text-to-speech | |
| - vocoder | |
| - hifi-gan | |
| - bigvgan | |
| - neural-vocoder | |
| - speaker-conditioning | |
| - audio-watermarking | |
| pipeline_tag: text-to-speech | |
| license: cc-by-4.0 | |
| # VocBulwark HiFi-GAN β watermarking neural vocoder (inference export) | |
| Speaker-conditioned **BigVGAN / HiFi-GAN** neural vocoder. It turns an input mel-spectrogram (*what* is said) into a 24 kHz waveform, conditioned on a **precomputed 768-d speaker embedding** (*whose* voice). Every clip it generates carries a **fixed 32-bit provenance watermark** for provenance β see [Watermark](#watermark). | |
| This is the **lean, inference-only** vocoder: the frozen perceptual-loss base models (Whisper / WavLM / Wav2Vec2), the training discriminators, **and the speaker encoder** have all been stripped β you pass the speaker embedding in. Use the companion **speaker-encoder** repo to turn a reference clip into that embedding. The modeling code is bundled, so it loads with `trust_remote_code=True` **without** the training repo. | |
| For higher watermark capacity (50-bit) and a larger generator (1536 initial channels vs 512), see the large model variant `mlr2000/vocoder-large`. | |
| ## Companion Models | |
| This model is part of a set of 6 repositories: | |
| | Repo | Role | | |
| |------|------| | |
| | `mlr2000/vocoder-large` | Large vocoder | | |
| | `mlr2000/vocoder-large-watermark-detector` | Watermark detector for the large model | | |
| | `mlr2000/vocoder-large-speaker-encoder` | Speaker encoder for the large model | | |
| | `mlr2000/vocoder-small` | Small vocoder (this repo) | | |
| | `mlr2000/vocoder-small-watermark-detector` | Watermark detector for the small model | | |
| | `mlr2000/vocoder-small-speaker-encoder` | Speaker encoder for the small model | | |
| ## Model summary | |
| | | | | |
| |---|---| | |
| | Architecture | `HiFiGANArchitecture` (BigVGAN generator, snakebeta activation) | | |
| | Inputs | log-mel spectrogram (**96 mel channels**) + speaker embedding (**768-d**) | | |
| | Output | mono waveform, **24 kHz** | | |
| | Speaker conditioning | precomputed embedding (from the companion speaker encoder) | | |
| | Generator | initial channels 512, upsample rates [4, 4, 2, 2, 2, 2] | | |
| | Watermark | 32-bit fixed VocBulwark signature, always embedded | | |
| | Framework | π€ Transformers, PyTorch, safetensors | | |
| ## Usage | |
| ```python | |
| import torch | |
| from transformers import AutoModel | |
| model = AutoModel.from_pretrained("mlr2000/vocoder-small", trust_remote_code=True).eval() | |
| mel = torch.randn(1, model.config.hifigan_in_channels, 200) # [B, mel, T] | |
| emb = torch.randn(1, model.config.speaker_embedding_size) # [B, 768] from the speaker encoder | |
| with torch.no_grad(): | |
| audio = model(mel_spectrogram=mel, speaker_embedding=emb).audio | |
| # audio: [B, 1, samples] @ model.config.target_sample_rate | |
| ``` | |
| See **`example_roundtrip.ipynb`** in this repo for the full pipeline (reference | |
| clip β speaker encoder β embedding β vocode β verify watermark). | |
| ## Training | |
| | | | | |
| |---|---| | |
| | Training data | Multilingual LibriSpeech (8 languages, ~22,200h) and Common Voice (14 languages, ~3,000h) | | |
| | Training steps | 1,000,000 | | |
| | Hardware | 2 Γ NVIDIA H100 PCIe GPUs | | |
| | Training objective | Discriminator-free: mel spectrogram + WavLM + wav2vec 2.0 + Whisper encoder losses | | |
| | Effective batch size | 32 | | |
| | Learning rate | 1e-3 | | |
| ## Watermark | |
| Every clip this model generates carries a **fixed 32-bit provenance watermark** (`config.fixed_watermark`). It is embedded automatically inside `forward` and **cannot be disabled or changed** through this interface β there is deliberately no watermark argument to override. To check whether a given audio came from this model, use the companion **detector** repo, which extracts the bits and compares them to the same fixed code. | |
| To verify whether a given audio clip was generated by this model, use the companion **detector** repo (`mlr2000/vocoder-small-watermark-detector`), which extracts the embedded bits and compares them to the known fixed code. | |
| ## Notes | |
| - Inputs: log-mel spectrogram (`config.hifigan_in_channels` channels) and a `[B, config.speaker_embedding_size]` speaker embedding. | |
| - Output: mono waveform at `config.target_sample_rate`. | |
| - Use the speaker embedding from the speaker encoder this vocoder was **trained with**. A mismatched encoder will not condition it correctly. | |
| - Not intended for voice cloning of real individuals without consent, or any deceptive or impersonation use. | |
| ## Citation | |
| If you use this model, please cite: | |
| ```bibtex | |
| @misc{muletta2026, | |
| title = {Training a Discriminator-Free Foundation Vocoder | |
| with Integrated Audio Watermarking}, | |
| author = {Muletta, Romolo and Deriu, Jan}, | |
| year = {2026}, | |
| note = {VT2 Project Report, ZHAW School of Engineering} | |
| } | |
| ``` | |
| ## License | |
| `cc-by-4.0`. Trained on MLS (CC-BY-4.0) and Common Voice (CC0); builds on | |
| BigVGAN (MIT) and wav2vec 2.0 (Apache-2.0). Please retain attribution when | |
| redistributing or building on this model. |