Vocetta2 1M

A complete English text-to-speech system in 1,031,754 parameters that scores SCOREQ 4.04 on the standard 24-sentence diverse held-out set, runs at ~90x faster than real time on a CPU (RTF 0.011), and needs nothing but Python and three pip packages to use.

Params 1,031,754 (duration 5,344 + acoustic 438,516 + decoder 587,894)
Audio 24 kHz mono
Speed ~90x real time on CPU (RTF 0.011); no GPU required
Intelligibility WER 0.083 on a 24-sentence diverse held-out set (12/24 exact)
Naturalness SCOREQ 4.04, DNSMOS-OVRL 3.44, DNSMOS-SIG 3.64, UTMOS 3.97
License MIT (runtime and weights); bundled G2P data is Apache-2.0

Listen to samples/ first. Those eight files were rendered by this exact checkpoint.

This is the largest release in the Vocetta2 family: 147K, 276K, 500K, and now 1M. Each step up doubles the budget and distills from a larger teacher TTS through the same three-stage design. This release is the one that crossed SCOREQ 4.0, which this family's earlier trajectory (1.48 at 147K, 2.04 at 276K, 2.70 at 500K) had suggested would land near 3.0 to 3.3 at this size. The section "How it reached 4.0" explains what actually made that possible, and most of it was procedure, not size.

How it works

Text becomes audio through four stages:

text ──► G2P ──► phoneme ids
                  │
      duration.pt │  ids ──────────────► frame count per phoneme
                  │
      acoustic.pt │  ids + durations ───► mel spectrogram [100 bands, T frames]
                  │
      decoder.pt  │  mel + noise ───────► complex spectrum ──► iSTFT ──► audio

G2P (microtts/g2p/). A dictionary-first English front end: words are looked up in bundled pronunciation dictionaries (gold and silver), and words the dictionaries miss go through a small bundled neural fallback model. NumPy only, no espeak, no torch, no network access. The output is a string of IPA phonemes, mapped to a frozen 62-symbol vocabulary (<bos> and <eos> bracket each utterance).

Duration student (duration.pt, 5,344 params). A 3-layer 1D convolutional network over the phoneme sequence. It predicts how many mel frames each phoneme occupies, using learned position, sequence-length and duration features. The shipped checkpoint is the third stage of a short pacing chain: a base run followed by two fine-tunes at lower learning rates. Pacing is a real quality lever on the full system: at the champion stage the two fine-tunes moved the system from 3.886 to 3.940 (+0.02, then +0.03) with nothing else changed.

Acoustic student (acoustic.pt, 438,516 params). Embed phoneme ids, refine them with token-context convolutions, expand them to the frame grid by repeating each phoneme its predicted number of frames, then run a second stack of convolutions over frames and project to 100 mel bands. Width 72 and depth 5 for this model. Trained against teacher mels with an L1 + spectral-convergence objective plus anti-smoothing (normalized-L1, temporal-delta, channel-statistics) terms and a hinge PatchGAN critic.

Decoder (decoder.pt, 587,894 params). Mel to waveform. A ConvNeXt1D stack (depthwise conv, LayerNorm, two pointwise layers with a GELU between them, residual) maps the mel to a complex spectrum of 513 bins, which the iSTFT turns into audio. The magnitude head is exponential with bin 0 and the Nyquist bin zeroed, and a DC-blocking filter removes the remaining offset. The decoder is noise-fed: a 4-channel noise input is projected and added to the mel embedding. At inference, zero noise is the best choice, measured on every metric.

A 100-value tilt.npy ships alongside the weights as a per-band mel correction. For this checkpoint the shipped vector is all zeros: zero tilt measured best for this acoustic, so applying it is a no-op. The file and the tool that computes it (train/compute_tilt.py) are kept so the convention stays explicit and a retrained acoustic has a defined home for its own vector.

The budget choice

component params share
duration 5,344 0.5%
acoustic 438,516 42.5%
decoder 587,894 57%

Under a 1M total, the split went mostly to the decoder, which matches the family's pattern: decoder capacity dominated every comparison made while building these models. The acoustic side earned its share too. Two measured levers shaped it:

Depth 5 on the acoustic was the single biggest architecture find in this project: about +0.42 gate SCOREQ over depth 3 at matched width. The value came from the reference recipe this family distills from; earlier work on this line had inherited depth 3 and never tested it. Depth 7 lost, and raising the token-stage depth also lost, so 5 is a measured optimum, not a default.

Ablations settled the shapes. The decoder depth was saturated at 4 blocks (deeper cost 100K parameters for a WER tie and a worse metallic profile), h72 was chosen over h80/h96 by full-system comparison, and the deployment-pack recipe (below) made the decoder's job well defined.

Family line: 147,532 params reaches SCOREQ 1.48, 276,765 reaches 2.04, 499,528 reaches 2.70, and this release reaches 4.04. The largest single jump in that line (500K to 1M: +1.34) came from a stack of measured changes rather than from size alone: depth-5 acoustics (about +0.42 gate), full schedules on rented compute after years of undersized local ones, deployment-pack decoder training, duration pacing, and checkpoint averaging. Size alone would not have gotten there.

How it reached 4.0

The line's first full system scored 3.49, and it climbed to 4.04 in two very different phases. The first phase was ordinary training work: longer decoder schedules (115k to 410k steps, +0.26), an acoustic polish and a spectral-convergence adjustment (system 3.75 to 3.89), and duration pacing (3.89 to 3.94). The second phase, the final night, is where the last 0.09 came from, and it used two ideas that were not training at all.

The gate: measure the ceiling, not the system

Rendering the acoustic student's own mels through the full teacher vocoder and scoring that audio (the "gate") measures how much quality the acoustic is leaving on the table, independent of the small decoder. For this line the system score tracks roughly 0.92 x gate, and that ratio stayed remarkably stable across configs. Gate evaluations take minutes at the 24-sentence set while full system evaluations take longer and depend on the decoder, so the gate became the fast signal for where to spend effort.

Gate history on this model:

stage gate SCOREQ
base acoustic, depth 5 4.083
+ polish fine-tune 4.123
+ sc-weight bite 4.239
+ averaging with one polish sibling 4.333
+ five-way blend with three fresh twins 4.372

Averaging, the free lever

The biggest discovery of the final session: same-recipe checkpoints, diverged by nothing but training noise and a short continuation, average into a model that beats every member. Every averaging move paid, and every training move past that point regressed. The final night's ladder, in which the only trained artifacts were six short acoustic twins (about 10 GPU-minutes each):

3.9462  champion decoder + base acoustic
3.9528  + 12k-step decoder continuation on a rebuilt pack
3.9709  + decoder checkpoint soup (average of two continuations)
3.9931  + acoustic soup (gate raised to 4.333)
3.9954  + cross-regime decoder soup
4.0286  + three fresh acoustic twins in the gate blend (gate 4.372)
4.0377  + four-member decoder soup
4.0384  + soup of the two best soups (the shipped decoder)

Twelve distinct configurations in this budget score 4.0 or better; the shipped one is the best measured. The shipped weights are weight averages all the way down (see train/README.md for the full composition), which is why the train/average_checkpoints.py script is part of this release.

What did not work (all measured, do not retry)

  • Retraining the decoder on the new acoustic (capture re-adaptation): 3.9714 against 3.9931 for the un-adapted soup. The decoder transfers to an improved acoustic without retraining.
  • Larger training pools: 10k rows scored 3.9328, 20k scored 3.9528, and 41.8k scored 3.9223 at matched settings. 20k is a genuine sweet spot.
  • Training longer: past about 600k total decoder steps the score regresses.
  • Every single-knob ablation: adversarial weight, quiet-ceiling weight, high-band weight, hint weight, crop length, discriminator learning rate, betas, and start-step variants were all run as isolated arms, and all landed at or below the incumbent. The recipe values were already at their optimum.

Honest notes on the numbers

  • The campaign's SCOREQ for this configuration is 4.0384. The released files re-measure at 4.0375 with a locally installed scorer build. The renders are bit-identical; the difference is scorer version noise. Both are above 4.0.
  • SCOREQ, DNSMOS and UTMOS are automatic predictors. They were used alongside WER and listening at every decision point, and a checkpoint once gamed SCOREQ while being unintelligible (WER 1.0), which is why no single metric was trusted.
  • 4.0 was reached at this scale. The margin is thin (about 0.04), and the next ceiling question (what would a 2M or 4M version of this recipe reach) is untested.

Use it from scratch

Download this repository, install the runtime dependencies, and run the script below. Two commands do the download; pick whichever tool you have:

# either the Hugging Face CLI (install it if you don't have it)
pip install -U huggingface_hub
hf download VantoraLabs/Vocetta2-1m --local-dir vocetta2-1m

# or git-lfs (install git-lfs first: https://git-lfs.com)
git lfs install && git clone https://huggingface.co/VantoraLabs/Vocetta2-1m

cd vocetta2-1m
pip install -r requirements.txt          # numpy, torch, soundfile

Then, from inside that folder:

from microtts import MicroTTS

tts = MicroTTS.load(".")                  # reads duration.pt / acoustic.pt / decoder.pt / tilt.npy
wav = tts.synthesize("Hello world.")      # float32 numpy array, 24 kHz
tts.save("out.wav", wav)                  # saves RMS-normalized to -26 dBFS

MicroTTS.load accepts device="cpu" (default) or "cuda". The full pipeline loads in well under a second and needs no network access. If you already have phoneme ids, call tts.synthesize_ids(ids) directly and skip the G2P.

Two practical notes:

  • Loudness. synthesize returns the raw output; its loudness is not normalized. MicroTTS.normalize(wav) applies RMS normalization (0.05, about -26 dBFS), and MicroTTS.save does it for you.
  • Noise. synthesize(..., noise_scale=0.0) is the default and gives the best measured quality; nonzero noise costs quality on every metric.

Python 3.10+, CPU is enough. Everything in benchmark/ runs on the same machine with the dependencies from requirements.txt (the WER script also needs transformers and jiwer, installed on first use in the same environment).

Benchmarks

Measured on this exact checkpoint through the shipped runtime. Word errors use Whisper small as the judge; naturalness scores use the standard open scorers (SCOREQ, DNSMOS, UTMOS) on the raw rendered audio, normalized to RMS 0.05.

Intelligibility (word error rate, lower is better)

set WER sentences exactly right
24-sentence diverse held-out set 0.083 12 / 24

Naturalness (higher is better, 24-sentence diverse set)

setting SCOREQ DNSMOS-OVRL DNSMOS-SIG UTMOS
noise_scale = 0.0, shipped tilt (recommended) 4.04 3.44 3.64 3.97

DNSMOS-SIG catches metallic distortion; 3.64 says the output is not buzzy. UTMOS 3.97 is the same automatic MOS estimator several other released TTS systems report, which makes it the most portable single number here.

Speed (24 diverse sentences, warm-up excluded, CPU)

device RTF real-time factor
CPU 0.0098 to 0.0121 across runs 83x to 102x (~90x typical)
GPU (GTX 750, legacy card) 0.012 83x

RTF includes the duration, acoustic and decoder forwards; the G2P adds a few milliseconds per sentence on top. A GPU does not meaningfully help this model: at this size the per-operation launch overhead dominates the real arithmetic, which is why both rows land within noise of each other. Ship it on CPU.

Reproducing the scores. benchmark/rtf.py and benchmark/wer.py are included and reproduce the tables above (wer.py needs transformers and jiwer). Naturalness: pip install scoreq speechmos and score each rendered wav with the libraries' own defaults, plus UTMOS from torch.hub ("tarepan/SpeechMOS:v1.2.0") in a separate interpreter (the two packages disagree about the module name speechmos, so they cannot share a process). The 24 sentences are in benchmark/eval_sentences.txt.

How it compares

Parameters are inference-time and exclude the grapheme-to-phoneme front end, matching the convention of the tables below. This model is 1.03M in that count.

Against the small-TTS field, on the same metric suite. The following rows come from the published comparison table of the sanoTTS release (huggingface.co/ampixa/sanoTTS), which scores a diverse 24-sentence set with SCOREQ / UTMOS / DNSMOS-SIG. This model was evaluated with the same scorer suite on the same 24 sentences, so the columns are directly comparable even though each system renders with its own runtime:

System Params SCOREQ UTMOS DNS-SIG
sanoTTS (amy) 1.46 M 4.13 4.10 3.61
Vocetta2 1M (this model) 1.03 M 4.04 3.97 3.64
TinyTTS 1.62 M 3.94 3.65 3.62
Inflect Nano 4.63 M 3.81 3.65 3.58
Kitten TTS nano 15 M 3.02 3.58 3.43
sanoTTS (heart-nano) 0.29 M 2.29 2.45 3.35
Piper (large teacher) ~15 M 4.71 4.47 3.65
Kokoro (large teacher) 82 M 4.89 4.52 3.69

Among sub-2M systems in that table this model is second on SCOREQ and UTMOS (behind a system 1.4x its size) and first on DNSMOS-SIG. The gap to the large teachers (Piper at about 15x the size, Kokoro at about 80x) remains wide and this release does not claim to close it.

Against a mid-size classic architecture. SupraTTS-0.1-Beta (huggingface.co/SupraLabs/SupraTTS-0.1-Beta) is a Glow-TTS + HiFi-GAN system, about 44M parameters across its acoustic model and vocoder, trained on LJSpeech. On its own evaluation protocol (10 out-of-domain plus 10 held-out LJSpeech sentences, 5 seeds, UTMOS as the automatic MOS estimate) it reports UTMOS 3.59 +/- 0.07 and WER 3.64%. This model's UTMOS on its own set is 3.97 at about 1/43rd of the parameters. The protocols differ (different sets, different seeds, different language material), so read that as a directional comparison, not a controlled one. The parameter ratio is the robust part.

Within the family. Each row measured on its own release, same scorer suite:

model params SCOREQ WER
Vocetta2-147k 147,532 1.48 0.105
Vocetta2-276K 276,765 2.04 0.101
Vocetta2-500k 499,528 2.70 0.117
Vocetta2 1M (this) 1,031,754 4.04 0.083

The family's word accuracy was flat between 147K and 500K (0.101 to 0.117) while naturalness climbed; this release is the first to improve both at once.

What is in this folder

duration.pt, acoustic.pt, decoder.pt   the weights (~4.0 MB total, fp32)
tilt.npy                               per-band mel correction (all zeros for this model)
model.safetensors                      same weights, prefixed keys (dur./ac./dec.),
                                       fp32 (auto-detected by HF Hub so the params
                                       count shows on the repo card)
config.json                            model metadata + architecture sizes; also the
                                       file that makes download counting work
microtts/                              runtime package (frontend, models, g2p)
  g2p/g2p_data/                        dictionaries + fallback model, Apache-2.0
samples/                               eight rendered examples
train/                                 the training scripts + the full recipe table
benchmark/                             RTF + WER scripts, eval sentences, results
README.md, LICENSE, requirements.txt

Reproduction on your own machine

Two levels, both spelled out in train/README.md:

  • Inference: a laptop CPU, Python 3.10+, three pip packages. About 90x real time. This is what the "Use it from scratch" section above is.
  • Full retrain: the train/ scripts are teacher-agnostic (you supply a text corpus and a teacher TTS that reports phonemes, audio and durations) and reproduce every stage: pack build, duration chain, acoustic chain, deployment pack, decoder, and the checkpoint averaging that produced the final weights. Total GPU time was roughly a GPU-day on an RTX 3090-class card, decoder stages dominating. The recipe table lists the exact commands, steps, batch sizes and learning rates per stage.

Limits, stated plainly

  • English only. The G2P dictionaries are US English.
  • One voice. This is a single-voice model; there is no speaker conditioning.
  • Utterances cap at 207 phoneme tokens and 2400 mel frames (25 s).
  • Free-form conversational text is harder than templated text. On the diverse set above, WER is 0.083 and 12 of 24 sentences come out perfect; expect a few word errors on unusual names and rare words.
  • No text normalization beyond the front end's number handling. Unusual punctuation or markup should be stripped before synthesis.
  • The bundled G2P fallback (used only for words missing from the dictionaries) is shared family infrastructure and is not counted in the 1,031,754, same convention as the comparison tables above.
  • The 4.0 margin is thin and the number comes from automatic predictors. Listen before trusting any single figure, including these.

Credits

Runtime, weights and training scripts: MIT (this release). The bundled grapheme-to-phoneme dictionaries and the fallback model are Apache-2.0; see microtts/g2p/g2p_data/NOTICE.md.

Downloads last month
-
Safetensors
Model size
1.03M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support