Training recipe: Matcha-TTS-PL
The complete, flattened procedure that produces the released model: a Polish Matcha-TTS acoustic model (base training, then fine-tuning on 8 readers) plus a HiFi-GAN vocoder fine-tuned to it. Everything below uses public data and public code and runs on one consumer or data-centre GPU in about 6 GPU-hours (RTX 4090: base 3.5 h, target 1 h, vocoder 1.7 h). Intermediate experiments are not part of this document.
0. Ingredients
| item | source | licence |
|---|---|---|
| Matcha-TTS code | github.com/shivammehta25/Matcha-TTS | MIT |
| Warm-start checkpoint | Matcha-TTS release matcha_vctk.ckpt (trained on VCTK) |
MIT weights; VCTK corpus CC BY 4.0 |
| Vocoder | HiFi-GAN universal v1 g_02500000 + discriminator do_02500000 (Matcha-TTS release / HF mirror AlexAlexBabarika/hifigan-universal-v1) |
MIT |
| HiFi-GAN training code | github.com/jik876/hifi-gan | MIT |
| Wolne Lektury audiobooks | wolnelektury.pl (repack datadriven-company/WolneLektury-TTS-Polish on Hugging Face) |
CC BY-SA 3.0 PL (attribution: author, title, reader, director) |
AZON spontaneous speech (pwr-azon_spont) |
Politechnika Wrocławska | CC BY-SA 4.0 |
| Phonemizer | espeak-ng 1.52 via phonemizer |
GPL-3.0 (runtime dependency only) |
| Speech recogniser for data filtering | faster-whisper large-v3-turbo |
MIT |
Tools referenced below live in the scripts/ directory of the playground repository
(machinekind/tts-pl-playground); matcha_patch/ holds the Polish
cleaner and configs. Apply python scripts/patch_matcha.py <repo> inside the Matcha-TTS clone once (idempotent).
1. Environment
git clone https://github.com/shivammehta25/Matcha-TTS.git
python3.12 -m venv .venv && . .venv/bin/activate
pip install torch torchaudio # CUDA build for training, CPU build is enough for inference
pip install Cython numpy && (cd Matcha-TTS && pip install -e . --no-deps --no-build-isolation)
pip install -r matcha_patch/requirements_min.txt librosa soundfile phonemizer faster-whisper huggingface_hub tensorboard
(cd Matcha-TTS && python ../scripts/patch_matcha.py ..)
export PHONEMIZER_ESPEAK_LIBRARY=/path/to/libespeak-ng.so # macOS: /opt/homebrew/lib/libespeak-ng.dylib
What the patch changes in Matcha-TTS:
polish_cleaners: espeak-ngplphonemization with punctuation preserved, text normalisation (numbers, abbreviations).- Symbol table: adds U+0303 (combining tilde, nasal vowels) →
n_vocab = 179. - Style token: optional 4th filelist column
path|speaker|text|style;MatchaTTS(n_styles=K)adds a zero-initialised embedding tableE_k ∈ R^{K×64}summed with the speaker embedding before the encoder and the decoder. - Monotonic alignment search through pinned memory;
torchaudio.load→soundfile; DataLoaderspawn; numpy 2 fixes.
2. Data
2.1 Ingest
scripts/ingest_wolnelektury.py: download audiobooks, segment on silences to 1–15 s, align text, writewavs/*.wav(22.05 kHz mono, peak −0.45 dBFS),piper_metadata.csv(file|reader|text) andsources.json(book → author, title, reader, director, licence URL) for attribution.scripts/ingest_azon.py: unpack, resample to 22.05 kHz, keep speaker ids.
2.2 Per-clip statistics and text/audio agreement
python scripts/clip_stats.py data/wl --out data_v2/clip_stats_wl.csv --workers 10
python scripts/clip_stats.py data/azon --out data_v2/clip_stats_azon.csv
python scripts/wl_book_meta.py data/wl/wl_api_cache.json --out data_v2/wl_books.json # genre per book (WL API)
python scripts/whisper_check.py data/wl --out data_v2/whisper_wl.csv --model large-v3-turbo --device cuda
Per clip: duration, RMS, silence fraction, F0 median and spread (pyin), UTMOS (tarepan/SpeechMOS), DNSMOS,
characters per second, and the character error rate between the label and a Whisper transcript. Clips with CER > 0.2
are text/audio mismatches (mis-segmented or mislabelled) and are dropped: about a quarter of the speaking-rate
outliers turned out to be mislabelled this way.
2.3 Build the training sets
python scripts/build_mix_v3.py --out data_v2 --q-oversample 4
Filters: 1–15 s, UTMOS ≥ 3.0 (AZON ≥ 2.8), DNSMOS ≥ 3.3, Whisper CER ≤ 0.2, Wolne Lektury prose only (verse, drama
and fables excluded by genre). Reader selection is by consistency, not hours: for each reader
score = z(UTMOS std) + z(DNSMOS std) + z(spread of per-clip F0 median) − z(UTMOS mean), lower is better; the 15 most
consistent readers form the base set, the 8 best are the fine-tune targets (ids 0–7). Speaker-balanced sampling
∝ √hours (1–4×), per-speaker validation split of 2 %.
Question labels. Audiobook readers read most questions with a falling contour, so a raw "?" would teach "slightly less fall". For clips ending in "?", the final-0.9 s pitch slope decides: rising → keep "?" and oversample 4×; falling but starting with a wh-word (co, kto, gdzie, kiedy, dlaczego, jak, ile, który…) → keep "?" (falling is correct Polish); falling otherwise → relabel "?" as "." so the label matches the audio.
Style tokens. 18 designed tokens = pitch range {flat, mid, wide} × speaking rate {slow, normal, fast} × question
{no, yes}, assigned per clip from the reader-relative F0 spread, rate and the question flag (style_map.json; the
neutral token is the one the map names neutral). Written as the 4th filelist column.
Outputs: data_v2/mix_v3 (base: 15 WL readers + 5 AZON speakers, 9.5k clips, ≈ 27 h) and data_v2/mix_v3_target
(8 readers, 5.2k clips, ≈ 14 h). The filelists are reproducible from the public sources with the commands above.
3. Acoustic model (Matcha-TTS)
Common settings: batch 64, bf16, Adam, out_size per Matcha default, checkpoints every N steps plus final.ckpt,
validation every 5 epochs, mel statistics computed per dataset (matcha.utils.generate_data_statistics).
0.25 s/step on an H100, 0.31 s/step on an RTX 4090.
| stage | data | init | steps | lr |
|---|---|---|---|---|
| base | mix_v3 (20 speakers, 18 style tokens) |
matcha_vctk.ckpt, weights only; speaker, style and symbol tables re-initialised |
40 000 | 1e-4 |
| target | mix_v3_target (8 readers) |
base final.ckpt |
12 000 | 5e-5 |
MIX=mix_v3 RUN_NAME=matcha_v3_base STYLE=1 MAX_STEPS=40000 CKPT_EVERY_STEPS=4000 \
WARM_CKPT=checkpoints/matcha_vctk.ckpt scripts/train_matcha.sh
MIX=mix_v3_target RUN_NAME=matcha_v3_target STYLE=1 MAX_STEPS=12000 CKPT_EVERY_STEPS=2000 \
WARM_CKPT=runs/matcha_v3_base/checkpoints/final.ckpt EXTRA="model.optimizer.lr=5e-5" scripts/train_matcha.sh
Warm start is weights-only with a shape-aware partial copy (scripts/train_matcha_warm.py): tables that changed size
(speaker, style, symbol embeddings) are copied row by row where shapes overlap, the rest is re-initialised.
Hydra rejects = inside override values, so checkpoint files named step_step=N.ckpt must be copied to a plain name
before being passed as +warm_ckpt=.
The base checkpoint (15 readers + AZON, 40k steps) is not published; to add a new voice, fine-tune the target checkpoint on one to two hours of clean recordings of one speaker with the
target command above, pointing MIX at your own filelist.
4. Vocoder (HiFi-GAN fine-tuned to the acoustic model)
Matcha's predicted mels are smoother than real ones (lower harmonic contrast, ≈ 7 dB less energy above 5 kHz). A vocoder trained only on real mels renders that blur as a phasey second layer and electronic "breaths" in pauses. Fine-tuning the universal HiFi-GAN on the target model's own mels removes it (the standard "fine-tuning" setting of the HiFi-GAN paper).
# teacher-forced mels for every target clip: MAS alignment on the real mel -> decoder (10 ODE steps, T 0.5), exact GT length
python scripts/gen_mels_tf.py --ckpt runs/matcha_v3_target/checkpoints/final.ckpt \
--filelist data_v2/mix_v3_target/matcha_train.txt --out vocoder_ft --max-clips 6000 --val 60 --device cuda
# fine-tune generator + discriminators from g_02500000 / do_02500000
FT_LR=2e-5 FT_DISC_LR_SCALE=0.5 FT_WARMUP=2000 FT_BATCH=16 FT_CKPT_EVERY=5000 \
scripts/finetune_hifigan.sh vocoder_ft runs/hifigan_ft 93 # 93 epochs × 323 steps ≈ 30k steps, 0.2 s/step
scripts/patch_hifigan_train.py adapts the official trainer: it restores the learning rate after loading the optimizer
state from do_02500000 (which otherwise silently overrides it), trains the generator alone on the mel loss for the
first 2 000 steps before enabling the adversarial and feature-matching losses, and fixes torch ≥ 2 / librosa ≥ 0.10
incompatibilities. Generator lr 2e-5, discriminators 1e-5, batch 16, segment 8192, 30k steps. The released vocoder is
the 30k-step generator; checkpoints from 15k on sound alike.
5. Evaluation
scripts/synth_samples.py (10 conversational sentences, 4–10 ODE steps, temperature 0.5, length scale 0.9) →
scripts/eval_synth.py: Whisper large-v3 WER/CER, UTMOS, F0 spread in semitones, speaking rate, silence ratio.
Note that UTMOS does not register the vocoder artefacts described in §4 (it even scores the fine-tuned vocoder slightly
lower); the vocoder decision was made by listening, with scripts/diag_vocoder_copy.py (copy-synthesis of real clips)
and scripts/diag_phone_artifacts.py (per-phoneme roughness) as diagnostics.
6. Inference
- Recommended settings: 4 ODE steps, temperature 0.5, length scale 0.9–1.0, the fine-tuned vocoder, a 10 kHz low-pass
on the output (removes a HiFi-GAN upsampling tone at the Nyquist frequency;
effects.pyin the playground). - Voice blending: average speaker-embedding rows with weights (
--voice "0.5*6+0.5*3"). The ONNX exports carry extra rows with blends already baked in (scripts/bake_voice.py), so a blend is just a speaker id. - Style tokens: an extra embedding added to the speaker embedding; ids and meanings in
style_map.json. - ONNX (acoustic model + vocoder in one graph, inputs
x,x_lengths,scales=[temperature, length_scale],spks):scripts/export_matcha_onnx.py <ckpt> out.onnx --n-timesteps 4 --vocoder-name hifigan_univ_v1 --vocoder-checkpoint-path <fine-tuned g_*>. - Measured on an NVIDIA GB10 (PyTorch bf16 +
torch.compile, batch 1, 4 steps): ≈ 23 ms to first audio, RTF ≈ 0.006. Apple M-series CPU, ONNX Runtime: ≈ 0.4 s for a 4 s sentence.
7. Licences of the result
Weights trained on CC BY-SA material are released under CC BY-SA 4.0 (compatible with CC BY-SA 3.0 PL).
ATTRIBUTION.md lists every Wolne Lektury book (author, title, reader, director, URL) and the AZON corpus; the warm
start (VCTK) and HiFi-GAN are MIT. Readers' voices are personal attributes not covered by the Creative Commons licence;
the release recommends blended voices under neutral names and disclosure of synthetic speech to listeners.