Matcha-TTS-PL / RECIPE.md
mcPear's picture
RECIPE: new voice needs one to two hours of clean recordings
627141b verified
|
Raw
History Blame Contribute Delete
11 kB

Training recipe: Matcha-TTS-PL

The complete, flattened procedure that produces the released model: a Polish Matcha-TTS acoustic model (base training, then fine-tuning on 8 readers) plus a HiFi-GAN vocoder fine-tuned to it. Everything below uses public data and public code and runs on one consumer or data-centre GPU in about 6 GPU-hours (RTX 4090: base 3.5 h, target 1 h, vocoder 1.7 h). Intermediate experiments are not part of this document.

0. Ingredients

item source licence
Matcha-TTS code github.com/shivammehta25/Matcha-TTS MIT
Warm-start checkpoint Matcha-TTS release matcha_vctk.ckpt (trained on VCTK) MIT weights; VCTK corpus CC BY 4.0
Vocoder HiFi-GAN universal v1 g_02500000 + discriminator do_02500000 (Matcha-TTS release / HF mirror AlexAlexBabarika/hifigan-universal-v1) MIT
HiFi-GAN training code github.com/jik876/hifi-gan MIT
Wolne Lektury audiobooks wolnelektury.pl (repack datadriven-company/WolneLektury-TTS-Polish on Hugging Face) CC BY-SA 3.0 PL (attribution: author, title, reader, director)
AZON spontaneous speech (pwr-azon_spont) Politechnika Wrocławska CC BY-SA 4.0
Phonemizer espeak-ng 1.52 via phonemizer GPL-3.0 (runtime dependency only)
Speech recogniser for data filtering faster-whisper large-v3-turbo MIT

Tools referenced below live in the scripts/ directory of the playground repository (machinekind/tts-pl-playground); matcha_patch/ holds the Polish cleaner and configs. Apply python scripts/patch_matcha.py <repo> inside the Matcha-TTS clone once (idempotent).

1. Environment

git clone https://github.com/shivammehta25/Matcha-TTS.git
python3.12 -m venv .venv && . .venv/bin/activate
pip install torch torchaudio            # CUDA build for training, CPU build is enough for inference
pip install Cython numpy && (cd Matcha-TTS && pip install -e . --no-deps --no-build-isolation)
pip install -r matcha_patch/requirements_min.txt librosa soundfile phonemizer faster-whisper huggingface_hub tensorboard
(cd Matcha-TTS && python ../scripts/patch_matcha.py ..)
export PHONEMIZER_ESPEAK_LIBRARY=/path/to/libespeak-ng.so   # macOS: /opt/homebrew/lib/libespeak-ng.dylib

What the patch changes in Matcha-TTS:

  • polish_cleaners: espeak-ng pl phonemization with punctuation preserved, text normalisation (numbers, abbreviations).
  • Symbol table: adds U+0303 (combining tilde, nasal vowels) → n_vocab = 179.
  • Style token: optional 4th filelist column path|speaker|text|style; MatchaTTS(n_styles=K) adds a zero-initialised embedding table E_k ∈ R^{K×64} summed with the speaker embedding before the encoder and the decoder.
  • Monotonic alignment search through pinned memory; torchaudio.load → soundfile; DataLoader spawn; numpy 2 fixes.

2. Data

2.1 Ingest

  • scripts/ingest_wolnelektury.py: download audiobooks, segment on silences to 1–15 s, align text, write wavs/*.wav (22.05 kHz mono, peak −0.45 dBFS), piper_metadata.csv (file|reader|text) and sources.json (book → author, title, reader, director, licence URL) for attribution.
  • scripts/ingest_azon.py: unpack, resample to 22.05 kHz, keep speaker ids.

2.2 Per-clip statistics and text/audio agreement

python scripts/clip_stats.py data/wl   --out data_v2/clip_stats_wl.csv   --workers 10
python scripts/clip_stats.py data/azon --out data_v2/clip_stats_azon.csv
python scripts/wl_book_meta.py data/wl/wl_api_cache.json --out data_v2/wl_books.json     # genre per book (WL API)
python scripts/whisper_check.py data/wl --out data_v2/whisper_wl.csv --model large-v3-turbo --device cuda

Per clip: duration, RMS, silence fraction, F0 median and spread (pyin), UTMOS (tarepan/SpeechMOS), DNSMOS, characters per second, and the character error rate between the label and a Whisper transcript. Clips with CER > 0.2 are text/audio mismatches (mis-segmented or mislabelled) and are dropped: about a quarter of the speaking-rate outliers turned out to be mislabelled this way.

2.3 Build the training sets

python scripts/build_mix_v3.py --out data_v2 --q-oversample 4

Filters: 1–15 s, UTMOS ≥ 3.0 (AZON ≥ 2.8), DNSMOS ≥ 3.3, Whisper CER ≤ 0.2, Wolne Lektury prose only (verse, drama and fables excluded by genre). Reader selection is by consistency, not hours: for each reader score = z(UTMOS std) + z(DNSMOS std) + z(spread of per-clip F0 median) − z(UTMOS mean), lower is better; the 15 most consistent readers form the base set, the 8 best are the fine-tune targets (ids 0–7). Speaker-balanced sampling ∝ √hours (1–4×), per-speaker validation split of 2 %.

Question labels. Audiobook readers read most questions with a falling contour, so a raw "?" would teach "slightly less fall". For clips ending in "?", the final-0.9 s pitch slope decides: rising → keep "?" and oversample 4×; falling but starting with a wh-word (co, kto, gdzie, kiedy, dlaczego, jak, ile, który…) → keep "?" (falling is correct Polish); falling otherwise → relabel "?" as "." so the label matches the audio.

Style tokens. 18 designed tokens = pitch range {flat, mid, wide} × speaking rate {slow, normal, fast} × question {no, yes}, assigned per clip from the reader-relative F0 spread, rate and the question flag (style_map.json; the neutral token is the one the map names neutral). Written as the 4th filelist column.

Outputs: data_v2/mix_v3 (base: 15 WL readers + 5 AZON speakers, 9.5k clips, ≈ 27 h) and data_v2/mix_v3_target (8 readers, 5.2k clips, ≈ 14 h). The filelists are reproducible from the public sources with the commands above.

3. Acoustic model (Matcha-TTS)

Common settings: batch 64, bf16, Adam, out_size per Matcha default, checkpoints every N steps plus final.ckpt, validation every 5 epochs, mel statistics computed per dataset (matcha.utils.generate_data_statistics). 0.25 s/step on an H100, 0.31 s/step on an RTX 4090.

stage data init steps lr
base mix_v3 (20 speakers, 18 style tokens) matcha_vctk.ckpt, weights only; speaker, style and symbol tables re-initialised 40 000 1e-4
target mix_v3_target (8 readers) base final.ckpt 12 000 5e-5
MIX=mix_v3        RUN_NAME=matcha_v3_base   STYLE=1 MAX_STEPS=40000 CKPT_EVERY_STEPS=4000 \
  WARM_CKPT=checkpoints/matcha_vctk.ckpt scripts/train_matcha.sh
MIX=mix_v3_target RUN_NAME=matcha_v3_target STYLE=1 MAX_STEPS=12000 CKPT_EVERY_STEPS=2000 \
  WARM_CKPT=runs/matcha_v3_base/checkpoints/final.ckpt EXTRA="model.optimizer.lr=5e-5" scripts/train_matcha.sh

Warm start is weights-only with a shape-aware partial copy (scripts/train_matcha_warm.py): tables that changed size (speaker, style, symbol embeddings) are copied row by row where shapes overlap, the rest is re-initialised. Hydra rejects = inside override values, so checkpoint files named step_step=N.ckpt must be copied to a plain name before being passed as +warm_ckpt=.

The base checkpoint (15 readers + AZON, 40k steps) is not published; to add a new voice, fine-tune the target checkpoint on one to two hours of clean recordings of one speaker with the target command above, pointing MIX at your own filelist.

4. Vocoder (HiFi-GAN fine-tuned to the acoustic model)

Matcha's predicted mels are smoother than real ones (lower harmonic contrast, ≈ 7 dB less energy above 5 kHz). A vocoder trained only on real mels renders that blur as a phasey second layer and electronic "breaths" in pauses. Fine-tuning the universal HiFi-GAN on the target model's own mels removes it (the standard "fine-tuning" setting of the HiFi-GAN paper).

# teacher-forced mels for every target clip: MAS alignment on the real mel -> decoder (10 ODE steps, T 0.5), exact GT length
python scripts/gen_mels_tf.py --ckpt runs/matcha_v3_target/checkpoints/final.ckpt \
    --filelist data_v2/mix_v3_target/matcha_train.txt --out vocoder_ft --max-clips 6000 --val 60 --device cuda
# fine-tune generator + discriminators from g_02500000 / do_02500000
FT_LR=2e-5 FT_DISC_LR_SCALE=0.5 FT_WARMUP=2000 FT_BATCH=16 FT_CKPT_EVERY=5000 \
  scripts/finetune_hifigan.sh vocoder_ft runs/hifigan_ft 93       # 93 epochs × 323 steps ≈ 30k steps, 0.2 s/step

scripts/patch_hifigan_train.py adapts the official trainer: it restores the learning rate after loading the optimizer state from do_02500000 (which otherwise silently overrides it), trains the generator alone on the mel loss for the first 2 000 steps before enabling the adversarial and feature-matching losses, and fixes torch ≥ 2 / librosa ≥ 0.10 incompatibilities. Generator lr 2e-5, discriminators 1e-5, batch 16, segment 8192, 30k steps. The released vocoder is the 30k-step generator; checkpoints from 15k on sound alike.

5. Evaluation

scripts/synth_samples.py (10 conversational sentences, 4–10 ODE steps, temperature 0.5, length scale 0.9) → scripts/eval_synth.py: Whisper large-v3 WER/CER, UTMOS, F0 spread in semitones, speaking rate, silence ratio. Note that UTMOS does not register the vocoder artefacts described in §4 (it even scores the fine-tuned vocoder slightly lower); the vocoder decision was made by listening, with scripts/diag_vocoder_copy.py (copy-synthesis of real clips) and scripts/diag_phone_artifacts.py (per-phoneme roughness) as diagnostics.

6. Inference

  • Recommended settings: 4 ODE steps, temperature 0.5, length scale 0.9–1.0, the fine-tuned vocoder, a 10 kHz low-pass on the output (removes a HiFi-GAN upsampling tone at the Nyquist frequency; effects.py in the playground).
  • Voice blending: average speaker-embedding rows with weights (--voice "0.5*6+0.5*3"). The ONNX exports carry extra rows with blends already baked in (scripts/bake_voice.py), so a blend is just a speaker id.
  • Style tokens: an extra embedding added to the speaker embedding; ids and meanings in style_map.json.
  • ONNX (acoustic model + vocoder in one graph, inputs x, x_lengths, scales=[temperature, length_scale], spks): scripts/export_matcha_onnx.py <ckpt> out.onnx --n-timesteps 4 --vocoder-name hifigan_univ_v1 --vocoder-checkpoint-path <fine-tuned g_*>.
  • Measured on an NVIDIA GB10 (PyTorch bf16 + torch.compile, batch 1, 4 steps): ≈ 23 ms to first audio, RTF ≈ 0.006. Apple M-series CPU, ONNX Runtime: ≈ 0.4 s for a 4 s sentence.

7. Licences of the result

Weights trained on CC BY-SA material are released under CC BY-SA 4.0 (compatible with CC BY-SA 3.0 PL). ATTRIBUTION.md lists every Wolne Lektury book (author, title, reader, director, URL) and the AZON corpus; the warm start (VCTK) and HiFi-GAN are MIT. Readers' voices are personal attributes not covered by the Creative Commons licence; the release recommends blended voices under neutral names and disclosure of synthetic speech to listeners.