# Training recipe: Matcha-TTS-PL The complete, flattened procedure that produces the released model: a Polish Matcha-TTS acoustic model (base training, then fine-tuning on 8 readers) plus a HiFi-GAN vocoder fine-tuned to it. Everything below uses public data and public code and runs on one consumer or data-centre GPU in about 6 GPU-hours (RTX 4090: base 3.5 h, target 1 h, vocoder 1.7 h). Intermediate experiments are not part of this document. ## 0. Ingredients | item | source | licence | |---|---|---| | Matcha-TTS code | github.com/shivammehta25/Matcha-TTS | MIT | | Warm-start checkpoint | Matcha-TTS release `matcha_vctk.ckpt` (trained on VCTK) | MIT weights; VCTK corpus CC BY 4.0 | | Vocoder | HiFi-GAN universal v1 `g_02500000` + discriminator `do_02500000` (Matcha-TTS release / HF mirror `AlexAlexBabarika/hifigan-universal-v1`) | MIT | | HiFi-GAN training code | github.com/jik876/hifi-gan | MIT | | Wolne Lektury audiobooks | wolnelektury.pl (repack `datadriven-company/WolneLektury-TTS-Polish` on Hugging Face) | CC BY-SA 3.0 PL (attribution: author, title, reader, director) | | AZON spontaneous speech (`pwr-azon_spont`) | Politechnika Wrocławska | CC BY-SA 4.0 | | Phonemizer | espeak-ng 1.52 via `phonemizer` | GPL-3.0 (runtime dependency only) | | Speech recogniser for data filtering | faster-whisper `large-v3-turbo` | MIT | Tools referenced below live in the `scripts/` directory of the playground repository ([machinekind/tts-pl-playground](https://github.com/machinekind/tts-pl-playground)); `matcha_patch/` holds the Polish cleaner and configs. Apply `python scripts/patch_matcha.py ` inside the Matcha-TTS clone once (idempotent). ## 1. Environment ```bash git clone https://github.com/shivammehta25/Matcha-TTS.git python3.12 -m venv .venv && . .venv/bin/activate pip install torch torchaudio # CUDA build for training, CPU build is enough for inference pip install Cython numpy && (cd Matcha-TTS && pip install -e . --no-deps --no-build-isolation) pip install -r matcha_patch/requirements_min.txt librosa soundfile phonemizer faster-whisper huggingface_hub tensorboard (cd Matcha-TTS && python ../scripts/patch_matcha.py ..) export PHONEMIZER_ESPEAK_LIBRARY=/path/to/libespeak-ng.so # macOS: /opt/homebrew/lib/libespeak-ng.dylib ``` What the patch changes in Matcha-TTS: - `polish_cleaners`: espeak-ng `pl` phonemization with punctuation preserved, text normalisation (numbers, abbreviations). - Symbol table: adds U+0303 (combining tilde, nasal vowels) → `n_vocab = 179`. - **Style token**: optional 4th filelist column `path|speaker|text|style`; `MatchaTTS(n_styles=K)` adds a zero-initialised embedding table `E_k ∈ R^{K×64}` summed with the speaker embedding before the encoder and the decoder. - Monotonic alignment search through pinned memory; `torchaudio.load` → `soundfile`; DataLoader `spawn`; numpy 2 fixes. ## 2. Data ### 2.1 Ingest - `scripts/ingest_wolnelektury.py`: download audiobooks, segment on silences to 1–15 s, align text, write `wavs/*.wav` (22.05 kHz mono, peak −0.45 dBFS), `piper_metadata.csv` (`file|reader|text`) and `sources.json` (book → author, title, reader, director, licence URL) for attribution. - `scripts/ingest_azon.py`: unpack, resample to 22.05 kHz, keep speaker ids. ### 2.2 Per-clip statistics and text/audio agreement ```bash python scripts/clip_stats.py data/wl --out data_v2/clip_stats_wl.csv --workers 10 python scripts/clip_stats.py data/azon --out data_v2/clip_stats_azon.csv python scripts/wl_book_meta.py data/wl/wl_api_cache.json --out data_v2/wl_books.json # genre per book (WL API) python scripts/whisper_check.py data/wl --out data_v2/whisper_wl.csv --model large-v3-turbo --device cuda ``` Per clip: duration, RMS, silence fraction, F0 median and spread (pyin), UTMOS (`tarepan/SpeechMOS`), DNSMOS, characters per second, and the character error rate between the label and a Whisper transcript. Clips with CER > 0.2 are text/audio mismatches (mis-segmented or mislabelled) and are dropped: about a quarter of the speaking-rate outliers turned out to be mislabelled this way. ### 2.3 Build the training sets ```bash python scripts/build_mix_v3.py --out data_v2 --q-oversample 4 ``` Filters: 1–15 s, UTMOS ≥ 3.0 (AZON ≥ 2.8), DNSMOS ≥ 3.3, Whisper CER ≤ 0.2, Wolne Lektury **prose only** (verse, drama and fables excluded by genre). Reader selection is by **consistency**, not hours: for each reader `score = z(UTMOS std) + z(DNSMOS std) + z(spread of per-clip F0 median) − z(UTMOS mean)`, lower is better; the 15 most consistent readers form the base set, the 8 best are the fine-tune targets (ids 0–7). Speaker-balanced sampling ∝ √hours (1–4×), per-speaker validation split of 2 %. **Question labels.** Audiobook readers read most questions with a falling contour, so a raw "?" would teach "slightly less fall". For clips ending in "?", the final-0.9 s pitch slope decides: rising → keep "?" and oversample 4×; falling but starting with a wh-word (co, kto, gdzie, kiedy, dlaczego, jak, ile, który…) → keep "?" (falling is correct Polish); falling otherwise → relabel "?" as "." so the label matches the audio. **Style tokens.** 18 designed tokens = pitch range {flat, mid, wide} × speaking rate {slow, normal, fast} × question {no, yes}, assigned per clip from the reader-relative F0 spread, rate and the question flag (`style_map.json`; the neutral token is the one the map names `neutral`). Written as the 4th filelist column. Outputs: `data_v2/mix_v3` (base: 15 WL readers + 5 AZON speakers, 9.5k clips, ≈ 27 h) and `data_v2/mix_v3_target` (8 readers, 5.2k clips, ≈ 14 h). The filelists are reproducible from the public sources with the commands above. ## 3. Acoustic model (Matcha-TTS) Common settings: batch 64, bf16, Adam, `out_size` per Matcha default, checkpoints every N steps plus `final.ckpt`, validation every 5 epochs, mel statistics computed per dataset (`matcha.utils.generate_data_statistics`). 0.25 s/step on an H100, 0.31 s/step on an RTX 4090. | stage | data | init | steps | lr | |---|---|---|---|---| | base | `mix_v3` (20 speakers, 18 style tokens) | `matcha_vctk.ckpt`, weights only; speaker, style and symbol tables re-initialised | 40 000 | 1e-4 | | target | `mix_v3_target` (8 readers) | base `final.ckpt` | 12 000 | 5e-5 | ```bash MIX=mix_v3 RUN_NAME=matcha_v3_base STYLE=1 MAX_STEPS=40000 CKPT_EVERY_STEPS=4000 \ WARM_CKPT=checkpoints/matcha_vctk.ckpt scripts/train_matcha.sh MIX=mix_v3_target RUN_NAME=matcha_v3_target STYLE=1 MAX_STEPS=12000 CKPT_EVERY_STEPS=2000 \ WARM_CKPT=runs/matcha_v3_base/checkpoints/final.ckpt EXTRA="model.optimizer.lr=5e-5" scripts/train_matcha.sh ``` Warm start is weights-only with a shape-aware partial copy (`scripts/train_matcha_warm.py`): tables that changed size (speaker, style, symbol embeddings) are copied row by row where shapes overlap, the rest is re-initialised. Hydra rejects `=` inside override values, so checkpoint files named `step_step=N.ckpt` must be copied to a plain name before being passed as `+warm_ckpt=`. The base checkpoint (15 readers + AZON, 40k steps) is not published; to add a new voice, fine-tune the target checkpoint on one to two hours of clean recordings of one speaker with the target command above, pointing `MIX` at your own filelist. ## 4. Vocoder (HiFi-GAN fine-tuned to the acoustic model) Matcha's predicted mels are smoother than real ones (lower harmonic contrast, ≈ 7 dB less energy above 5 kHz). A vocoder trained only on real mels renders that blur as a phasey second layer and electronic "breaths" in pauses. Fine-tuning the universal HiFi-GAN on the target model's own mels removes it (the standard "fine-tuning" setting of the HiFi-GAN paper). ```bash # teacher-forced mels for every target clip: MAS alignment on the real mel -> decoder (10 ODE steps, T 0.5), exact GT length python scripts/gen_mels_tf.py --ckpt runs/matcha_v3_target/checkpoints/final.ckpt \ --filelist data_v2/mix_v3_target/matcha_train.txt --out vocoder_ft --max-clips 6000 --val 60 --device cuda # fine-tune generator + discriminators from g_02500000 / do_02500000 FT_LR=2e-5 FT_DISC_LR_SCALE=0.5 FT_WARMUP=2000 FT_BATCH=16 FT_CKPT_EVERY=5000 \ scripts/finetune_hifigan.sh vocoder_ft runs/hifigan_ft 93 # 93 epochs × 323 steps ≈ 30k steps, 0.2 s/step ``` `scripts/patch_hifigan_train.py` adapts the official trainer: it restores the learning rate after loading the optimizer state from `do_02500000` (which otherwise silently overrides it), trains the generator alone on the mel loss for the first 2 000 steps before enabling the adversarial and feature-matching losses, and fixes torch ≥ 2 / librosa ≥ 0.10 incompatibilities. Generator lr 2e-5, discriminators 1e-5, batch 16, segment 8192, 30k steps. The released vocoder is the 30k-step generator; checkpoints from 15k on sound alike. ## 5. Evaluation `scripts/synth_samples.py` (10 conversational sentences, 4–10 ODE steps, temperature 0.5, length scale 0.9) → `scripts/eval_synth.py`: Whisper large-v3 WER/CER, UTMOS, F0 spread in semitones, speaking rate, silence ratio. Note that UTMOS does not register the vocoder artefacts described in §4 (it even scores the fine-tuned vocoder slightly lower); the vocoder decision was made by listening, with `scripts/diag_vocoder_copy.py` (copy-synthesis of real clips) and `scripts/diag_phone_artifacts.py` (per-phoneme roughness) as diagnostics. ## 6. Inference - Recommended settings: 4 ODE steps, temperature 0.5, length scale 0.9–1.0, the fine-tuned vocoder, a 10 kHz low-pass on the output (removes a HiFi-GAN upsampling tone at the Nyquist frequency; `effects.py` in the playground). - Voice blending: average speaker-embedding rows with weights (`--voice "0.5*6+0.5*3"`). The ONNX exports carry extra rows with blends already baked in (`scripts/bake_voice.py`), so a blend is just a speaker id. - Style tokens: an extra embedding added to the speaker embedding; ids and meanings in `style_map.json`. - ONNX (acoustic model + vocoder in one graph, inputs `x`, `x_lengths`, `scales=[temperature, length_scale]`, `spks`): `scripts/export_matcha_onnx.py out.onnx --n-timesteps 4 --vocoder-name hifigan_univ_v1 --vocoder-checkpoint-path `. - Measured on an NVIDIA GB10 (PyTorch bf16 + `torch.compile`, batch 1, 4 steps): ≈ 23 ms to first audio, RTF ≈ 0.006. Apple M-series CPU, ONNX Runtime: ≈ 0.4 s for a 4 s sentence. ## 7. Licences of the result Weights trained on CC BY-SA material are released under **CC BY-SA 4.0** (compatible with CC BY-SA 3.0 PL). `ATTRIBUTION.md` lists every Wolne Lektury book (author, title, reader, director, URL) and the AZON corpus; the warm start (VCTK) and HiFi-GAN are MIT. Readers' voices are personal attributes not covered by the Creative Commons licence; the release recommends blended voices under neutral names and disclosure of synthetic speech to listeners.