Archive: screening models (80 M, 88.8 h, 30k steps) used to choose the recipe. The main model: VoiceHub/mytts-en-v1 · the data: VoiceHub/mytts-en-v1-data

mytts-en: zero-shot English text-to-speech

Read this first: every checkpoint here is a small screening model, not the planned model. The 14 runs published so far are Tier-1 screening runs: ~80 M parameters, 30k training steps (about 1-1.5 GPU-hours each), trained on only 88.8 hours of speech (31,987 clips, 3,587 voices), 0.8 % of the provided ~11,150-hour SynDataLab-EN/echo-clones-4m-en. Screening runs exist only to choose settings (each run changes one technique and is compared with two baseline seeds), so their quality is far below what the main model is planned to reach.

The main model (~206 M parameters, the full ~11,150 h, 200k steps) has its own page: VoiceHub/mytts-en-v1 (weights, listening samples and scores; the data: VoiceHub/mytts-en-v1-data).

mytts-en clones a voice from a 3-10 s prompt and reads any English text in it, at 48 kHz. It is a flow-matching diffusion transformer that generates Semantic-DACVAE codec latents (25 frames/s x 128 channels) conditioned on the text and on the prompt's own latents, which it continues in context. Code: kadirnar/dacvae-next (the mytts package). A hobby project. Trained only on the provided data: SynDataLab-EN/echo-clones-4m-en, synthetic EchoTTS speech (no external corpus).

Listen: every checkpoint below reads the same 5 prompts and texts on the samples page. Start with its Start here block.

Architecture

  • DiT backbone over the latent sequence (prompt frames + target frames), trained as a rectified flow (flow matching): the prompt's latents stay clean and the model infills the target, so voice cloning is in-context, with no speaker encoder.
  • Cross-attention text path: character tokens (a normalized English text front end: numbers, dates, money, fillers) are encoded once and read by every block through cross-attention, with LARoPE (length-aware rotary positions that align text and speech positions by their relative progress).
  • Auxiliary heads during training: a text CTC head in the middle of the stack and a REPA head that aligns a deep block with self-supervised speech features (w2v-BERT 2.0, or HuBERT-large in the EN-25 arms); both are dropped at inference.
  • Codec: Aratako/Semantic-DACVAE-Japanese (48 kHz, 25 Hz, 128 channels); it encodes the prompt and decodes the generated latents.
  • Sampling (the EN-E1 protocol): 32 Euler steps with a sway time grid, classifier-free guidance (joint, w 4), initial noise scale 0.9, output normalized to at most -16 LUFS (peak-limited); the target length comes from the prompt's speaking rate (band rule).

Status

The screening checkpoints (Tier-1) train the small preset (~80 M parameters) for 30k steps on the en_t1 subset: 88.8 hours of speech (31,987 clips, 3,587 voices), 0.8 % of the provided ~11,150-hour SynDataLab-EN/echo-clones-4m-en. Each run changes one technique against two baseline seeds: they show how each technique sounds and scores, not the final quality. Their training logs, raw exports and evaluation files are in VoiceHub/mytts-en-ablations.

The main model (~206 M parameters, the full ~11,150 h, 200k steps) has its own page: VoiceHub/mytts-en-v1 (weights, listening samples and scores; the data: VoiceHub/mytts-en-v1-data).

Checkpoints

14 checkpoints of 14 runs: the main model first (once it trains), then the finished screening runs (the highest echo-dev UTMOS first), then any screening run still training; only runs trained exclusively on the provided data are listed. WER = corpus WER (Whisper-large-v3, en-v2 normalization), SIM-o = speaker similarity to the prompt (WavLM-large

  • ECAPA-TDNN), UTMOS = UTMOS22 strong; GT = the ground truth's UTMOS on that set (echo-dev 4.20, seed-dev 3.52: the ceiling, EN-22 validation); - = not evaluated at that step (the Tier-1 protocol scores the final export). EN-E1 = the run's screening verdict against both baseline seeds. Samples = the checkpoint's section on the samples page. Shown: of the main model every checkpoint from step 20k on; of each screening run only its final checkpoint (the one its EN-E1 evaluation scores). Checkpoints before step 20k are not published: they have not learned to align text and speech yet (on the five listening items, a mean WER of 82-111 % and UTMOS 1.3-1.7 at 5k-10k steps, against 5.5 % and 2.8 at 30k). A screening run's later steps before its final one have learned to align too; they are left out only so that each run is heard at the checkpoint it is judged by.

Screening runs

14 runs: Tier-1, ~80 M parameters, 30k steps, 88.8 hours of speech (31,987 clips, 3,587 voices); used only to choose settings.

Run Step What the run tests echo-dev WER echo-dev SIM-o echo-dev UTMOS (GT 4.20) seed-dev WER seed-dev SIM-o seed-dev UTMOS (GT 3.52) EN-E1 verdict Files Samples
en55-t06 30000 EN-55 (#100) noise schedule: flow t_mean -0.6 / t_std 0.9 instead of -1.2 / 0.8 3.34 % 0.676 3.09 8.01 % 0.448 2.96 worse on echo_dev, seed_dev safetensors listen
en55-pc06-t06 30000 EN-55 (#100) noise schedule t_mean -0.6 / t_std 0.9 with the EN-29 pc06 prompting (fixed picker, p_cross 0.6) 2.14 % 0.687 3.09 6.56 % 0.449 2.99 worse on seed_dev safetensors listen
en55-pc06-t08 30000 EN-55 (#100) noise schedule t_mean -0.8 / t_std 0.8 with the EN-29 pc06 prompting (fixed picker, p_cross 0.6); the Tier-1 arm of the main-run (M4) recipe 1.68 % 0.681 3.04 5.30 % 0.440 2.94 no significant gain safetensors listen
en55-pc06-t10 30000 EN-55 (#100) noise schedule t_mean -1.0 / t_std 0.8 with the EN-29 pc06 prompting (fixed picker, p_cross 0.6) 2.07 % 0.663 2.95 5.21 % 0.413 2.79 no significant gain safetensors listen
en55-t08 30000 EN-55 (#100) noise schedule: flow t_mean -0.8 / t_std 0.8 (WavTTS) instead of -1.2 / 0.8 3.38 % 0.662 2.93 8.01 % 0.435 2.78 worse on echo_dev, seed_dev safetensors listen
en29-fixed-pc06-tf04 30000 EN-29 (#53) prompting: fixed picker, p_cross 0.6, prompt-text drop 0.4 1.24 % 0.649 2.76 5.97 % 0.407 2.58 adopted: adopt safetensors listen
en29-fixed 30000 EN-29 (#53) prompting: the fixed cross-prompt picker (p_cross 0.35, prompt-text drop 0.2) 1.73 % 0.639 2.76 5.21 % 0.392 2.65 no significant gain safetensors listen
en29-fixed-pc06 30000 EN-29 (#53) prompting: fixed picker, p_cross 0.6 1.06 % 0.652 2.71 6.30 % 0.424 2.63 adopted: adopt safetensors listen
en55-repafade 30000 EN-55 (#100) REPA at full weight to 18k steps, then a linear fade to 0 by 21k 3.15 % 0.628 2.70 7.28 % 0.398 2.51 worse on echo_dev, seed_dev safetensors listen
en-t1-base-s2 30000 Tier-1 baseline, seed 2 (the same recipe as en-t1-base): the seed-to-seed spread every arm is judged against 2.51 % 0.630 2.70 5.84 % 0.392 2.62 baseline (seed 2); vs the other seed: no significant gain safetensors listen
en-t1-base 30000 Tier-1 baseline, seed 1234: small preset (~80 M), en_t1 catalog (88.8 h of echo-clones-4m-en), 30k steps, w2v-BERT 2.0 L16 REPA targets 1.96 % 0.631 2.68 5.96 % 0.393 2.58 baseline (seed 1234) safetensors listen
en25-hubl24 30000 EN-25 (#49) REPA teacher: HuBERT-large-ll60k layer 24 targets (bf16) instead of w2v-BERT 2.0 L16 1.68 % 0.617 2.59 5.37 % 0.381 2.45 no significant gain safetensors listen
en25-hubl18 30000 EN-25 (#49) REPA teacher: HuBERT-large-ll60k layer 18 targets 2.03 % 0.607 2.59 5.49 % 0.369 2.54 no significant gain safetensors listen
en25-hubl24fp32 30000 EN-25 (#49) REPA teacher: HuBERT-large-ll60k layer 24 targets extracted in fp32 1.75 % 0.614 2.55 5.82 % 0.394 2.34 no significant gain safetensors listen

Usage

Install the code (kadirnar/dacvae-next, branch roadmap/en-echo: pip install -e .), download one checkpoint folder and load it; the folder is a pickle-free release (safetensors + JSON), so nothing is unpickled.

import soundfile as sf
from huggingface_hub import snapshot_download

from mytts.flow.sampler import SamplerConfig
from mytts.infer import Synthesizer

ckpt = "checkpoints/en-t1-base/step_0030000"  # any folder of the table above
local = snapshot_download("VoiceHub/mytts-en", allow_patterns=[f"{ckpt}/*"])
syn = Synthesizer.from_export(f"{local}/{ckpt}", device="cuda")  # prints: watermark: AudioSeal ... on every output
sc = SamplerConfig(steps=32, sway=-1.0, cfg_mode="joint", cfg_w=4.0, noise_scale=0.9)  # the EN-E1 protocol
wav, sr = syn.synthesize("Hi! This is my cloned voice reading a brand new sentence.", lang="en", sc=sc,
                         prompt_audio="prompt.wav", prompt_text="The exact transcript of prompt.wav.", out_lufs=-16.0, seed=0)
sf.write("output.wav", wav, sr)  # 48 kHz, AudioSeal-watermarked

Single files: hf_hub_download("VoiceHub/mytts-en", f"{ckpt}/config.json") (and model.safetensors, extras.safetensors when present) into one folder works the same. From the command line: python scripts/synthesize.py --model <folder> --text ... --prompt-audio prompt.wav --prompt-text ... --out out.wav. Prompts of 3-10 s with an exact transcript work best. A text longer than 250 characters is split into chunks at sentence ends (a single longer sentence at a comma or space), which are generated one by one and cross-faded; the long sample on the samples page is two sentences, split at the sentence end.

Watermark

The codec decodes without a watermark, so every checkpoint records the AudioSeal watermark (meta.watermark in config.json; facebook/audioseal, 16-bit message 0100110101010100): Synthesizer.from_export and scripts/synthesize.py add it to every output as the last step, so generated speech is marked in a machine-readable way and detectable as AI-generated. Check a file with AudioSealWatermarker().detect(wav, sr) (mytts.watermark) or python scripts/watermark_check.py detect <files>. The mark is added by the inference code, not the weights: code that decodes the latents itself produces unmarked audio.

Training data

  • SynDataLab-EN/echo-clones-4m-en: about 4 M synthetic EchoTTS voice clones of 4,000 reference voices, ~11,150 h; the screening runs use en_t1 (88.8 h). 236 voices are held out for evaluation and never trained on.

Every checkpoint here was trained exclusively on these provided datasets (the allowlist configs/data/en_train_sources.txt; scripts/hub_showcase.py publishes a checkpoint only when its training catalog, and the checkpoint it was initialized from, records nothing else). No external corpus is used for training.

Licence of the weights: apache-2.0 (owner decision 2026-09-29, hobby project).

Evaluation protocol (EN-E1)

  • echo-dev: sentences voiced by held-out echo voices x prompt draws d0 + d1 (in-domain; SIM there is in-distribution).
  • seed-dev: the frozen dev half of Seed-TTS test-en (Common Voice prompts, real voices) x sampling seeds 0 + 1; an evaluation benchmark only (never trained on, not on the samples page).
  • Judge: Whisper-large-v3 (greedy, repetition-loop guard) with the frozen en-v2 normalization; corpus WER (errors / reference words); SIM-o = WavLM-large + ECAPA-TDNN cosine to the prompt; UTMOS22 strong. Every number is over 2 draws.
  • UTMOS ceiling: the ground truth scores 4.20 on the echo voices (their codec resynthesis about the same) and 3.52 on Seed test-en (EN-22 validation); a checkpoint's UTMOS is shown next to it (GT).
  • Verdict rule: a technique is adopted only with a paired corpus-WER gain (95 % CI excluding 0) against BOTH baseline seeds on one set, significantly worse than neither seed on the other, SIM-o drop <= 0.01 and UTMOS drop within the seeds' spread.
  • Validation gates passed (EN-22, the judge on ground-truth audio): Seed test-en GT WER 2.14 % (seed-tts-eval protocol; PASS), LS-PC GT WER 2.47 % (F5 protocol; PASS).

Limitations

  • Trained on synthetic voices (EchoTTS clones): speaker similarity on real voices (seed-dev SIM-o 0.39) is well below the in-domain echo voices (0.63) and below the ground-truth ceiling (0.73); real-voice cloning is the main gap.
  • Screening checkpoints (~80 M parameters, 30k steps, 88.8 h): expect mispronunciations, especially on rare words, long numbers and texts far from conversational English. English only.
  • UTMOS is an English MOS predictor; read small differences with care.

Responsible use

Clone a voice only with the consent of its owner, never to impersonate a real person or to deceive, and disclose synthetic speech. Every output of the inference code carries the AudioSeal watermark; do not remove it.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train VoiceHub/mytts-en