Vocetta2 500K

A complete English text-to-speech system in 499,528 parameters — three tiny networks and a dictionary-based grapheme-to-phoneme front end, all running on CPU in real time.

Params 499,528 (duration 5,344 + acoustic 265,642 + decoder 228,542)
Audio 24 kHz mono
Speed ~125x real time on CPU (RTF 0.008)
Intelligibility WER 0.117 on a 24-sentence diverse held-out set (9/24 exact)
Naturalness SCOREQ 2.70, DNSMOS-OVRL 3.20, DNSMOS-SIG 3.46
License MIT (runtime and weights); bundled G2P data is Apache-2.0

Listen to samples/ first — those eight files were rendered by this exact checkpoint.

How it works

Text becomes audio through four stages, all trained by distillation from a larger teacher TTS:

text ──► G2P ──► phoneme ids
                  │
      duration.pt │  ids ──────────────► frame count per phoneme
                  │
      acoustic.pt │  ids + durations ───► mel spectrogram [100 bands, T frames]
                  │
      decoder.pt  │  mel + noise ───────► complex spectrum ──► iSTFT ──► audio

G2P (microtts/g2p/). A dictionary-first English front end: words are looked up in bundled pronunciation dictionaries (gold and silver), and words the dictionaries miss go through a small bundled neural fallback model. NumPy only, no espeak, no torch, no network access. The output is a string of IPA phonemes, mapped to a frozen 62-symbol vocabulary (<bos> and <eos> bracket each utterance).

Duration student (duration.pt, 5,344 params). A 3-layer 1D convolutional network over the phoneme sequence. It predicts how many mel frames each phoneme occupies, using learned position, sequence-length and duration features, with residual blocks around each conv pair. The output is exponentiated log-duration, rounded and clamped to at least one frame per phoneme.

Acoustic student (acoustic.pt, 265,642 params). Embeds the phoneme ids, refines them with token-context convolutions, expands them to the frame grid by repeating each phoneme its predicted number of frames, then runs a second stack of convolutions over frames and projects to 100 mel bands. Trained against teacher mels with an L1 + spectral-convergence objective plus anti-smoothing (normalized-L1, temporal-delta, channel-statistics) terms and a hinge PatchGAN critic.

Decoder (decoder.pt, 228,542 params). Mel to waveform. A ConvNeXt1D stack (depthwise conv, LayerNorm, two pointwise layers with a GELU between them, residual) maps the mel to a complex spectrum of 513 bins, which the iSTFT turns into audio. The magnitude head is exponential with bin 0 and the Nyquist bin zeroed, and a DC-blocking filter removes the remaining offset. The decoder is noise-fed: a 4-channel noise input is projected and added to the mel embedding. At inference, zero noise is the best choice.

A 100-value tilt.npy is shipped alongside the weights and applied to the mel before decoding; it is a measured per-band correction between the acoustic's mel and the teacher's mel distribution, and the published numbers include it.

The budget choice

The 500K budget is split mostly in favour of the decoder, because decoder width dominated every capacity comparison made while building this family.

component params share
duration 5,344 1%
acoustic 265,642 53%
decoder 228,542 46%

Acoustic size was swept through one fixed reference renderer and peaks at h64: h48 reaches 2.77, h56 2.96, h64 3.34, h72 3.15 SCOREQ of its own rendered mel. Both directions lose, so the freed parameters would not buy a better model at this budget. The duration student is deliberately tiny; a larger one was measured to trade SCOREQ for word accuracy rather than improve both.

For comparison within the same family: a 276,765-parameter model reaches SCOREQ 2.04 with WER 0.101, and a 147,532-parameter model reaches 1.48 with WER 0.105. This model reaches 2.70 — naturalness scales with the budget much more steeply than word accuracy does.

The training recipe

Every student is trained by distillation: a larger teacher TTS renders a text corpus once, and the students learn to reproduce the teacher's intermediate representations.

Stage 0 — build the pack

Pick a teacher TTS and a text corpus (thousands of sentences of varied, spoken-style text). For each line store: phoneme ids, teacher audio, per-phoneme durations, and the mel spectrogram of the teacher audio (100 bands, n_fft 1024, hop 256). One .npz per line (train/build_pack.py). Watch the duration units when the teacher's frame rate differs from the mel hop — the script shows the conversion. ~10,300 lines were used here.

Stage 1 — duration student

Train ids → frame counts against the teacher's durations. Loss: smooth-L1 on log-duration plus a term on the total length (weight 0.35). Learning rate 2e-3 with AdamW, ~4k steps at batch 32.

Stage 2 — acoustic student

Train ids + durations → mel against the teacher's mel. Loss: L1 plus normalized L1, temporal-delta L1 (weight 0.10 — this anti-smoothing term is load-bearing), channel-statistics L1, and a hinge-loss PatchGAN critic on the mel from an early step with weight 0.1. The learning rate matters most: 2e-3, constant. ~12k steps at batch 8.

Stage 3 — build the deployment pack

Convert the teacher pack into the decoder's training distribution (train/build_deploy_pack.py): the acoustic student's own mels become the inputs, and full vocoder renders of those exact mels become the audio targets. The decoder's objective is then "reproduce the vocoder's output on the mels you will actually receive", and the usual acoustic→decoder distribution gap does not exist to begin with. This is the single most important departure from the classic recipe and is what the released decoder was trained on.

Stage 4 — decoder

  1. Initialize, do not train from scratch. Slice the first N channels of a pretrained neural vocoder's decoder (train/init_decoder.py; dim 67, pw 197, 4 blocks for this model). Training a decoder this small from scratch does not reach intelligibility.
  2. Train on the deployment pack (train_decoder.py --mix-prob 0.0 --noise-zeros) with waveform L1 + multi-resolution spectral loss + a hinge PatchGAN discriminator + a cosine-gram temporal-structure loss (weight 0.4 — this supervises temporal texture and is what keeps the output from sounding mushy), plus high-band-excess and quiet-ceiling terms. Schedule length is the biggest single lever here: this decoder ran 115k steps at constant lr 2e-4, because the quality curve was still rising at 40k. If you shorten it, expect WER to reach its ceiling early while SCOREQ stays below the mel's own ceiling.
  3. Compute the tilt file (train/compute_tilt.py) and apply it at inference.

There is no joint fine-tune stage in this release. Co-training the acoustic and decoder together was measured twice on this model (a thin loss stack, then a full one against achievable targets); both lost about 0.35 SCOREQ versus the frozen-acoustic decoder. Gradients flowing back through the vocoder push the acoustic into a region that renders worse downstream. The released recipe therefore trains the decoder only.

Two traps measured during development, worth stating plainly:

  • Use a constant learning rate. Decaying the rate (e.g. ×0.9998 per step) silently starves these small models: the loss looks fine while the output stays muddy. Removing the decay was the largest single fix in this family's history.
  • Do not over-trust one metric. A decoder tuned only against SCOREQ once learned to emit muffled, unintelligible audio that scored 3.07 SCOREQ (WER 1.0). Every candidate here was checked with WER, DNSMOS and listening.

Use it from scratch

Download this repository, then install the runtime dependencies and run the script below. Two commands do the download — pick whichever tool you have:

# either the Hugging Face CLI (install it if you don't have it)
pip install -U huggingface_hub
hf download VantoraLabs/Vocetta2-500k --local-dir vocetta2-500k

# or git-lfs (install git-lfs first: https://git-lfs.com)
git lfs install && git clone https://huggingface.co/VantoraLabs/Vocetta2-500k

cd vocetta2-500k
pip install -r requirements.txt          # numpy, torch, soundfile

Then, from inside that folder:

from microtts import MicroTTS

tts = MicroTTS.load(".")                  # reads duration.pt / acoustic.pt / decoder.pt / tilt.npy
wav = tts.synthesize("Hello world.")      # float32 numpy array, 24 kHz
tts.save("out.wav", wav)                  # saves RMS-normalized to -26 dBFS

MicroTTS.load accepts device="cpu" (default) or "cuda". The full pipeline loads in well under a second after the first import and needs no network access. If you already have phoneme ids, call tts.synthesize_ids(ids) directly and skip the G2P.

Two practical notes:

  • Loudness. synthesize returns the raw output; its loudness is not normalized. MicroTTS.normalize(wav) applies RMS normalization (0.05, about -26 dBFS), and MicroTTS.save does it for you.
  • Noise and tilt. synthesize(..., noise_scale=0.0) is the default and gives the best measured quality; nonzero noise costs quality on every metric. The shipped tilt.npy is applied by default (use_tilt=True) and is what the published numbers include; use_tilt=False disables it.

Python 3.10+, CPU is enough.

Benchmarks

Measured on this exact checkpoint through the shipped runtime. All word-error numbers use Whisper small as the judge; naturalness scores use the standard open models (SCOREQ, DNSMOS), each evaluated on the raw rendered audio.

Intelligibility (word error rate, lower is better)

set WER sentences exactly right
24-sentence diverse held-out set 0.117 9 / 24

Naturalness / audio quality (higher is better, 24-sentence diverse set)

setting SCOREQ DNSMOS-OVRL DNSMOS-SIG
noise_scale = 0.0, tilt on (recommended) 2.70 3.20 3.46
same mels, tilt off 2.66 3.20 3.45

DNSMOS-SIG catches metallic distortion; 3.46 says the output is not buzzy.

Speed (24 diverse sentences, warm-up excluded, zero noise, CPU)

device RTF real-time factor
CPU 0.008 ~125x faster than real time

RTF includes the duration, acoustic, decoder forwards and the tilt correction; the G2P adds ~3 ms per sentence on top.

For reference, a GPU is not faster for this model: measured on an RTX 3080, the same pipeline runs at RTF 0.014 (~71x) versus ~126x on a desktop CPU, because a 500K-parameter network is small enough that kernel-launch overhead dominates. On a single 9.5 s utterance the gap narrows (CPU 173x vs GPU 109x) but CPU still leads. Ship it on CPU.

Reproducing the scores. Word error: transcribe the rendered wavs with openai/whisper-small and compute WER against the input text (scripts in benchmark/). Naturalness: pip install scoreq speechmos, then score each wav with the library's own defaults.

What is in this folder

duration.pt, acoustic.pt, decoder.pt   the weights (2.0 MB total, fp32)
tilt.npy                               100-value mel correction applied at inference
model.safetensors                      same weights, prefixed keys (dur./ac./dec.),
                                       fp32 (auto-detected by HF Hub so the params
                                       count shows on the repo card)
microtts/                              runtime package (frontend, models, g2p)
  g2p/g2p_data/                        dictionaries + fallback model, Apache-2.0
samples/                               eight rendered examples
train/                                 the training scripts (see the recipe above)
benchmark/                             RTF + WER scripts, eval sentences, results
README.md, LICENSE, requirements.txt

Limits, stated plainly

  • English only. The G2P dictionaries are US English.
  • One voice. This is a single-voice model; there is no speaker conditioning.
  • Utterances cap at 207 phoneme tokens and 2400 mel frames (25 s).
  • Free-form conversational text is harder than templated text; word errors are common there (WER 0.117 on the diverse set above).
  • Naturalness and word accuracy trade against each other at this size. This checkpoint is the naturalness-optimised variant: it outscores a 276,765-parameter sibling on SCOREQ (2.70 vs 2.04) while scoring slightly worse on WER (0.117 vs 0.101). A same-size variant with the retrained duration student measures SCOREQ 2.69, WER 0.106 — pick per use case.
  • No text normalization beyond the front end's number handling. Unusual punctuation or markup should be stripped before synthesis.

Credits

Runtime, weights and training scripts: MIT (this release). The bundled grapheme-to-phoneme dictionaries and the fallback model are Apache-2.0; see microtts/g2p/g2p_data/NOTICE.md.

Downloads last month
29
Safetensors
Model size
500k params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support