Instructions to use VoiceHub/mytts-en with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- AudioSeal
How to use VoiceHub/mytts-en with AudioSeal:
# Watermark Generator from audioseal import AudioSeal model = AudioSeal.load_generator("VoiceHub/mytts-en") # pass a tensor (tensor_wav) of shape (batch, channels, samples) and a sample rate wav, sr = tensor_wav, 16000 watermark = model.get_watermark(wav, sr) watermarked_audio = wav + watermark# Watermark Detector from audioseal import AudioSeal detector = AudioSeal.load_detector("VoiceHub/mytts-en") result, message = detector.detect_watermark(watermarked_audio, sr) - Notebooks
- Google Colab
- Kaggle
Archive: screening models (80 M, 88.8 h, 30k steps) used to choose the recipe. The main model: VoiceHub/mytts-en-v1 · the data: VoiceHub/mytts-en-v1-data
mytts-en: zero-shot English text-to-speech
Read this first: every checkpoint here is a small screening model, not the planned model. The 14 runs published so far are Tier-1 screening runs: ~80 M parameters, 30k training steps (about 1-1.5 GPU-hours each), trained on only 88.8 hours of speech (31,987 clips, 3,587 voices), 0.8 % of the provided ~11,150-hour SynDataLab-EN/echo-clones-4m-en. Screening runs exist only to choose settings (each run changes one technique and is compared with two baseline seeds), so their quality is far below what the main model is planned to reach.
The main model (~206 M parameters, the full ~11,150 h, 200k steps) has its own page: VoiceHub/mytts-en-v1 (weights, listening samples and scores; the data: VoiceHub/mytts-en-v1-data).
mytts-en clones a voice from a 3-10 s prompt and reads any English text in it, at 48 kHz. It is a flow-matching diffusion
transformer that generates Semantic-DACVAE codec latents (25 frames/s x 128 channels) conditioned on the text and on the
prompt's own latents, which it continues in context. Code: kadirnar/dacvae-next (the mytts package).
A hobby project. Trained only on the provided data: SynDataLab-EN/echo-clones-4m-en, synthetic EchoTTS speech (no external corpus).
Listen: every checkpoint below reads the same 5 prompts and texts on the samples page. Start with its Start here block.
Architecture
- DiT backbone over the latent sequence (prompt frames + target frames), trained as a rectified flow (flow matching): the prompt's latents stay clean and the model infills the target, so voice cloning is in-context, with no speaker encoder.
- Cross-attention text path: character tokens (a normalized English text front end: numbers, dates, money, fillers) are encoded once and read by every block through cross-attention, with LARoPE (length-aware rotary positions that align text and speech positions by their relative progress).
- Auxiliary heads during training: a text CTC head in the middle of the stack and a REPA head that aligns a deep block with self-supervised speech features (w2v-BERT 2.0, or HuBERT-large in the EN-25 arms); both are dropped at inference.
- Codec: Aratako/Semantic-DACVAE-Japanese (48 kHz, 25 Hz, 128 channels); it encodes the prompt and decodes the generated latents.
- Sampling (the EN-E1 protocol): 32 Euler steps with a sway time grid, classifier-free guidance (joint, w 4), initial noise scale 0.9, output normalized to at most -16 LUFS (peak-limited); the target length comes from the prompt's speaking rate (band rule).
Status
The screening checkpoints (Tier-1) train the small preset (~80 M parameters) for 30k steps on the en_t1
subset: 88.8 hours of speech (31,987 clips, 3,587 voices), 0.8 % of the provided ~11,150-hour SynDataLab-EN/echo-clones-4m-en. Each run changes one
technique against two baseline seeds: they show how each technique sounds and scores, not the final quality. Their
training logs, raw exports and evaluation files are in VoiceHub/mytts-en-ablations.
The main model (~206 M parameters, the full ~11,150 h, 200k steps) has its own page: VoiceHub/mytts-en-v1 (weights, listening samples and scores; the data: VoiceHub/mytts-en-v1-data).
Checkpoints
14 checkpoints of 14 runs: the main model first (once it trains), then the finished screening runs (the highest echo-dev UTMOS first), then any screening run still training; only runs trained exclusively on the provided data are listed. WER = corpus WER (Whisper-large-v3, en-v2 normalization), SIM-o = speaker similarity to the prompt (WavLM-large
- ECAPA-TDNN), UTMOS = UTMOS22 strong; GT = the ground truth's UTMOS on that set (echo-dev 4.20, seed-dev
3.52: the ceiling, EN-22 validation);
-= not evaluated at that step (the Tier-1 protocol scores the final export). EN-E1 = the run's screening verdict against both baseline seeds. Samples = the checkpoint's section on the samples page. Shown: of the main model every checkpoint from step 20k on; of each screening run only its final checkpoint (the one its EN-E1 evaluation scores). Checkpoints before step 20k are not published: they have not learned to align text and speech yet (on the five listening items, a mean WER of 82-111 % and UTMOS 1.3-1.7 at 5k-10k steps, against 5.5 % and 2.8 at 30k). A screening run's later steps before its final one have learned to align too; they are left out only so that each run is heard at the checkpoint it is judged by.
Screening runs
14 runs: Tier-1, ~80 M parameters, 30k steps, 88.8 hours of speech (31,987 clips, 3,587 voices); used only to choose settings.
| Run | Step | What the run tests | echo-dev WER | echo-dev SIM-o | echo-dev UTMOS (GT 4.20) | seed-dev WER | seed-dev SIM-o | seed-dev UTMOS (GT 3.52) | EN-E1 verdict | Files | Samples |
|---|---|---|---|---|---|---|---|---|---|---|---|
| en55-t06 | 30000 | EN-55 (#100) noise schedule: flow t_mean -0.6 / t_std 0.9 instead of -1.2 / 0.8 | 3.34 % | 0.676 | 3.09 | 8.01 % | 0.448 | 2.96 | worse on echo_dev, seed_dev | safetensors | listen |
| en55-pc06-t06 | 30000 | EN-55 (#100) noise schedule t_mean -0.6 / t_std 0.9 with the EN-29 pc06 prompting (fixed picker, p_cross 0.6) | 2.14 % | 0.687 | 3.09 | 6.56 % | 0.449 | 2.99 | worse on seed_dev | safetensors | listen |
| en55-pc06-t08 | 30000 | EN-55 (#100) noise schedule t_mean -0.8 / t_std 0.8 with the EN-29 pc06 prompting (fixed picker, p_cross 0.6); the Tier-1 arm of the main-run (M4) recipe | 1.68 % | 0.681 | 3.04 | 5.30 % | 0.440 | 2.94 | no significant gain | safetensors | listen |
| en55-pc06-t10 | 30000 | EN-55 (#100) noise schedule t_mean -1.0 / t_std 0.8 with the EN-29 pc06 prompting (fixed picker, p_cross 0.6) | 2.07 % | 0.663 | 2.95 | 5.21 % | 0.413 | 2.79 | no significant gain | safetensors | listen |
| en55-t08 | 30000 | EN-55 (#100) noise schedule: flow t_mean -0.8 / t_std 0.8 (WavTTS) instead of -1.2 / 0.8 | 3.38 % | 0.662 | 2.93 | 8.01 % | 0.435 | 2.78 | worse on echo_dev, seed_dev | safetensors | listen |
| en29-fixed-pc06-tf04 | 30000 | EN-29 (#53) prompting: fixed picker, p_cross 0.6, prompt-text drop 0.4 | 1.24 % | 0.649 | 2.76 | 5.97 % | 0.407 | 2.58 | adopted: adopt | safetensors | listen |
| en29-fixed | 30000 | EN-29 (#53) prompting: the fixed cross-prompt picker (p_cross 0.35, prompt-text drop 0.2) | 1.73 % | 0.639 | 2.76 | 5.21 % | 0.392 | 2.65 | no significant gain | safetensors | listen |
| en29-fixed-pc06 | 30000 | EN-29 (#53) prompting: fixed picker, p_cross 0.6 | 1.06 % | 0.652 | 2.71 | 6.30 % | 0.424 | 2.63 | adopted: adopt | safetensors | listen |
| en55-repafade | 30000 | EN-55 (#100) REPA at full weight to 18k steps, then a linear fade to 0 by 21k | 3.15 % | 0.628 | 2.70 | 7.28 % | 0.398 | 2.51 | worse on echo_dev, seed_dev | safetensors | listen |
| en-t1-base-s2 | 30000 | Tier-1 baseline, seed 2 (the same recipe as en-t1-base): the seed-to-seed spread every arm is judged against | 2.51 % | 0.630 | 2.70 | 5.84 % | 0.392 | 2.62 | baseline (seed 2); vs the other seed: no significant gain | safetensors | listen |
| en-t1-base | 30000 | Tier-1 baseline, seed 1234: small preset (~80 M), en_t1 catalog (88.8 h of echo-clones-4m-en), 30k steps, w2v-BERT 2.0 L16 REPA targets | 1.96 % | 0.631 | 2.68 | 5.96 % | 0.393 | 2.58 | baseline (seed 1234) | safetensors | listen |
| en25-hubl24 | 30000 | EN-25 (#49) REPA teacher: HuBERT-large-ll60k layer 24 targets (bf16) instead of w2v-BERT 2.0 L16 | 1.68 % | 0.617 | 2.59 | 5.37 % | 0.381 | 2.45 | no significant gain | safetensors | listen |
| en25-hubl18 | 30000 | EN-25 (#49) REPA teacher: HuBERT-large-ll60k layer 18 targets | 2.03 % | 0.607 | 2.59 | 5.49 % | 0.369 | 2.54 | no significant gain | safetensors | listen |
| en25-hubl24fp32 | 30000 | EN-25 (#49) REPA teacher: HuBERT-large-ll60k layer 24 targets extracted in fp32 | 1.75 % | 0.614 | 2.55 | 5.82 % | 0.394 | 2.34 | no significant gain | safetensors | listen |
Usage
Install the code (kadirnar/dacvae-next, branch roadmap/en-echo: pip install -e .), download one checkpoint
folder and load it; the folder is a pickle-free release (safetensors + JSON), so nothing is unpickled.
import soundfile as sf
from huggingface_hub import snapshot_download
from mytts.flow.sampler import SamplerConfig
from mytts.infer import Synthesizer
ckpt = "checkpoints/en-t1-base/step_0030000" # any folder of the table above
local = snapshot_download("VoiceHub/mytts-en", allow_patterns=[f"{ckpt}/*"])
syn = Synthesizer.from_export(f"{local}/{ckpt}", device="cuda") # prints: watermark: AudioSeal ... on every output
sc = SamplerConfig(steps=32, sway=-1.0, cfg_mode="joint", cfg_w=4.0, noise_scale=0.9) # the EN-E1 protocol
wav, sr = syn.synthesize("Hi! This is my cloned voice reading a brand new sentence.", lang="en", sc=sc,
prompt_audio="prompt.wav", prompt_text="The exact transcript of prompt.wav.", out_lufs=-16.0, seed=0)
sf.write("output.wav", wav, sr) # 48 kHz, AudioSeal-watermarked
Single files: hf_hub_download("VoiceHub/mytts-en", f"{ckpt}/config.json") (and model.safetensors, extras.safetensors when
present) into one folder works the same. From the command line: python scripts/synthesize.py --model <folder> --text ... --prompt-audio prompt.wav --prompt-text ... --out out.wav. Prompts of 3-10 s with an exact transcript work best. A text longer than 250
characters is split into chunks at sentence ends (a single longer sentence at a comma or space), which are generated one by
one and cross-faded; the long sample on the samples page is two sentences, split at the sentence end.
Watermark
The codec decodes without a watermark, so every checkpoint records the AudioSeal watermark (meta.watermark in
config.json; facebook/audioseal, 16-bit message 0100110101010100):
Synthesizer.from_export and scripts/synthesize.py add it to every output as the last step, so generated speech is marked
in a machine-readable way and detectable as AI-generated. Check a file with AudioSealWatermarker().detect(wav, sr)
(mytts.watermark) or python scripts/watermark_check.py detect <files>. The mark is added by the inference code, not the
weights: code that decodes the latents itself produces unmarked audio.
Training data
- SynDataLab-EN/echo-clones-4m-en: about 4 M synthetic EchoTTS voice clones of 4,000
reference voices, ~11,150 h; the screening runs use
en_t1(88.8 h). 236 voices are held out for evaluation and never trained on.
Every checkpoint here was trained exclusively on these provided datasets (the allowlist configs/data/en_train_sources.txt;
scripts/hub_showcase.py publishes a checkpoint only when its training catalog, and the checkpoint it was initialized from,
records nothing else). No external corpus is used for training.
Licence of the weights: apache-2.0 (owner decision 2026-09-29, hobby project).
Evaluation protocol (EN-E1)
- echo-dev: sentences voiced by held-out echo voices x prompt draws d0 + d1 (in-domain; SIM there is in-distribution).
- seed-dev: the frozen dev half of Seed-TTS test-en (Common Voice prompts, real voices) x sampling seeds 0 + 1; an evaluation benchmark only (never trained on, not on the samples page).
- Judge: Whisper-large-v3 (greedy, repetition-loop guard) with the frozen en-v2 normalization; corpus WER (errors / reference words); SIM-o = WavLM-large + ECAPA-TDNN cosine to the prompt; UTMOS22 strong. Every number is over 2 draws.
- UTMOS ceiling: the ground truth scores 4.20 on the echo voices (their codec resynthesis about the same) and
3.52 on Seed test-en (EN-22 validation); a checkpoint's UTMOS is shown next to it (
GT). - Verdict rule: a technique is adopted only with a paired corpus-WER gain (95 % CI excluding 0) against BOTH baseline seeds on one set, significantly worse than neither seed on the other, SIM-o drop <= 0.01 and UTMOS drop within the seeds' spread.
- Validation gates passed (EN-22, the judge on ground-truth audio): Seed test-en GT WER 2.14 % (seed-tts-eval protocol; PASS), LS-PC GT WER 2.47 % (F5 protocol; PASS).
Limitations
- Trained on synthetic voices (EchoTTS clones): speaker similarity on real voices (seed-dev SIM-o
0.39) is well below the in-domain echo voices (0.63) and below the ground-truth ceiling (0.73); real-voice cloning is the main gap. - Screening checkpoints (~80 M parameters, 30k steps, 88.8 h): expect mispronunciations, especially on rare words, long numbers and texts far from conversational English. English only.
- UTMOS is an English MOS predictor; read small differences with care.
Responsible use
Clone a voice only with the consent of its owner, never to impersonate a real person or to deceive, and disclose synthetic speech. Every output of the inference code carries the AudioSeal watermark; do not remove it.