DACFlow-EN-10k: English text-to-speech that copies a voice

DACFlow-EN-10k reads English text aloud in a voice you choose. Give it a short recording of a voice (3-10 seconds, with the exact words spoken in it) and any English text, and it speaks that text in that voice. This is zero-shot voice cloning: the voice does not have to be in the training data. The output is 48 kHz audio. The model has 206 M parameters and is trained from scratch only on the provided data: 9,543 hours of synthetic English speech (3,400,911 clips, 3,587 voices) from SynDataLab-EN/echo-clones-4m-en (the curated part of its ~11,150 hours), nothing else. A hobby project; the weights are apache-2.0. The name: DAC for the Semantic-DACVAE audio codec whose latents it generates, Flow for flow matching, EN for English, 10k for its training data, about ten thousand hours of speech. The code: kadirnar/dacvae-next, Python package mytts.

Status: training, step 100k of 200k. A new checkpoint (a copy of the model saved at that point of its training) appears here every 10k steps (from step 20k on), with its samples and scores: this page updates itself. Early checkpoints sound rougher than later ones.

On this page: Listen · Results · How to use · Training · Data · Licence

Listen

The newest checkpoint, step 100k, reads five texts. Each text is spoken in a different voice, copied from the short voice prompt next to it. The model never heard these voices in training (held-out voices of the provided data).

What the model readsVoice prompt (the input)Model output, step 100k
Short sentence
I left my umbrella at the office again, so I'm definitely getting soaked on the way home.

A higher voice (~202 Hz). It says: “I keep telling myself I just need to power through, like, push a little harder, and then it'll be fine. But it's not getting fine. It's getting— it's getting worse.”
Question
Have you ever noticed that the quietest person in the room usually has the most interesting story to tell?

A lower voice (~149 Hz). It says: “Wait, that mole— has it always looked like that? No, stop.”
Numbers, dates, abbreviations
Dr. Patel moved my appointment to Tuesday, March 3rd, at 4:15 p.m., and the co-pay went up from $20 to $35.

A higher voice (~181 Hz). It says: “this woman on the train was literally eating a whole bag of chips, like, crunching so loud, and I'm sitting there like, okay, do I move? whatever.”
Conversation (~10 s)
So I finally tried that new ramen place downtown, and honestly, it was worth the wait. The broth was amazing, but next time I'm definitely skipping the extra spicy option.

A lower voice (~115 Hz). It says: “I'm sorry, but the subway doors closed on my bag. Again.”
Long passage (~20 s)
When the storm passed, the whole neighborhood came outside to look at the damage, and although a few fences had fallen and the old oak tree had lost its biggest branch, everyone was relieved that nobody was hurt. They spent the rest of the afternoon clearing the street and sharing whatever food they had left.

The deepest voice (~108 Hz). It says: “Honestly I think we need to just call IT and have them... you know, actually fix it this time. I'm too old for this.”

Older checkpoints (8): the same five texts and voices. Click one to open it.

Step 90k: echo-dev WER 0.74 % · UTMOS 3.78
What the model readsModel output, step 90k
Short sentence
Question
Numbers, dates, abbreviations
Conversation (~10 s)
Long passage (~20 s)
Step 80k: echo-dev WER 0.78 % · UTMOS 3.73
What the model readsModel output, step 80k
Short sentence
Question
Numbers, dates, abbreviations
Conversation (~10 s)
Long passage (~20 s)
Step 70k: echo-dev WER 0.67 % · UTMOS 3.69
What the model readsModel output, step 70k
Short sentence
Question
Numbers, dates, abbreviations
Conversation (~10 s)
Long passage (~20 s)
Step 60k: echo-dev WER 0.69 % · UTMOS 3.62
What the model readsModel output, step 60k
Short sentence
Question
Numbers, dates, abbreviations
Conversation (~10 s)
Long passage (~20 s)
Step 50k: echo-dev WER 0.74 % · UTMOS 3.56
What the model readsModel output, step 50k
Short sentence
Question
Numbers, dates, abbreviations
Conversation (~10 s)
Long passage (~20 s)
Step 40k: echo-dev WER 0.74 % · UTMOS 3.50
What the model readsModel output, step 40k
Short sentence
Question
Numbers, dates, abbreviations
Conversation (~10 s)
Long passage (~20 s)
Step 30k: echo-dev WER 1.06 % · UTMOS 3.43
What the model readsModel output, step 30k
Short sentence
Question
Numbers, dates, abbreviations
Conversation (~10 s)
Long passage (~20 s)
Step 20k: echo-dev WER 1.20 % · UTMOS 3.33
What the model readsModel output, step 20k
Short sentence
Question
Numbers, dates, abbreviations
Conversation (~10 s)
Long passage (~20 s)

How the samples are made: one take per text, no cherry-picking, with the settings of How to use (32 Euler steps + sway, joint CFG w 4.0, initial noise 0.9, at most -16 LUFS (peak-limited), duration by the band rule). Every output carries an inaudible AudioSeal watermark, checked before upload; the prompts are the original recordings.

Results

How well each checkpoint does on voices and sentences it never saw in training (lower WER is better, higher SIM-o and UTMOS are better):

Checkpoint echo-dev WER ↓ echo-dev SIM-o ↑ echo-dev UTMOS ↑ (real speech: 4.20) seed-dev WER ↓ seed-dev SIM-o ↑ seed-dev UTMOS ↑ (real speech: 3.52) Samples Weights
step 100k (newest) 0.69 % 0.787 3.85 1.54 % 0.499 3.78 listen files
step 90k 0.74 % 0.782 3.78 - - - listen files
step 80k 0.78 % 0.776 3.73 1.58 % 0.507 3.68 listen files
step 70k 0.67 % 0.766 3.69 - - - listen files
step 60k 0.69 % 0.759 3.62 1.61 % 0.537 3.56 listen files
step 50k 0.74 % 0.754 3.56 - - - listen files
step 40k 0.74 % 0.748 3.50 1.85 % 0.536 3.48 listen files
step 30k 1.06 % 0.741 3.43 - - - listen files
step 20k 1.20 % 0.722 3.33 2.47 % 0.506 3.27 listen files
  • WER (word error rate): Whisper-large-v3 writes down what it hears and this is compared with the text; 2 % means about one word in fifty is wrong.
  • SIM-o (speaker similarity, 0 to 1): how close the output's voice is to the prompt's (WavLM-large + ECAPA-TDNN).
  • UTMOS (1 to 5): a predicted listener rating of naturalness (UTMOS22). The column header gives the score of real recordings of the same set: the practical ceiling.
  • echo-dev: held-out voices of the training data's kind (synthetic EchoTTS voices, none of them trained on). seed-dev: real people's voices (the dev half of Seed-TTS test-en, Common Voice recordings). seed-dev is harder: the model learned only from synthetic voices.
  • -: not measured at that step (echo-dev is scored every 10k steps, seed-dev every 20k). Every number averages two takes per sentence.

How to use

1. Install (Python 3.10 or newer; a CUDA GPU is recommended):

git clone -b roadmap/en-echo https://github.com/kadirnar/dacvae-next
cd dacvae-next
pip install -e ".[codec]"

2. Copy a voice in Python. You need a short recording of the voice (3-10 s, clean, one speaker) and the exact words spoken in it:

import soundfile as sf
from huggingface_hub import snapshot_download

from mytts.flow.sampler import SamplerConfig
from mytts.infer import Synthesizer

ckpt = "checkpoints/step_0100000"  # the newest checkpoint; any row of the Results table works
local = snapshot_download("VoiceHub/DACFlow-EN-10k", allow_patterns=[f"{ckpt}/*"])  # downloads only this checkpoint
tts = Synthesizer.from_export(f"{local}/{ckpt}", device="cuda")
sc = SamplerConfig(steps=32, cfg_mode="joint", cfg_w=4.0, noise_scale=0.9)  # the settings of the samples above
wav, sr = tts.synthesize(
    "Any English text you like.",
    lang="en",
    sc=sc,
    prompt_audio="my_voice.wav",  # the voice to copy
    prompt_text="The exact words spoken in my_voice.wav.",
    out_lufs=-16.0,
    seed=0,
)
sf.write("output.wav", wav, sr)  # 48 kHz, with the AudioSeal watermark

Or from the command line (these are its default settings):

hf download VoiceHub/DACFlow-EN-10k --include "checkpoints/step_0100000/*" --local-dir DACFlow-EN-10k
python scripts/synthesize.py --model DACFlow-EN-10k/checkpoints/step_0100000 \
    --text "Any English text you like." --prompt-audio my_voice.wav \
    --prompt-text "The exact words spoken in my_voice.wav." --out output.wav

Tips: a clean prompt with an exact transcript works best. A long text is split at sentence ends and read piece by piece. python scripts/watermark_check.py detect output.wav checks the watermark.

Training

  • Model: the base preset, 206 M parameters: a diffusion transformer trained with flow matching. It generates the latents of the Semantic-DACVAE audio codec (48 kHz, 25 frames per second) from the text, continuing the voice prompt in context (no separate speaker encoder).
  • Data: the curated training catalog echo_en_v1 of the provided data (EN-10, #34 curation; see Data).
  • Recipe: 200k steps on one RTX 5090, about 27 minutes of speech per step (40,000 latent frames); learning rate 2.5e-4 after 5k warm-up steps, held, then lowered over the last 20 % of the steps (from step 160k). The settings were chosen with small screening runs: the flow-matching noise schedule t_mean -0.8 / t_std 0.8 (EN-55, #100) and cross-utterance voice prompts with p_cross 0.6 (EN-29, #53).
  • Code and history: the recipe configs/train/en_full.yaml (PR #154, schedule PR #156); the run EN-34, #58; the data processing EN-12, #36; the whole English plan EN-01, #78.

Data

  • The training data, tokenized: VoiceHub/DACFlow-EN-10k-data. Exactly what the model trains on: every clip of the training catalog echo_en_v1 as codec tokens (Semantic-DACVAE latents, 25 frames per second) with its transcript, and the catalog itself (the exact list of training clips). No audio: the codec turns the tokens back into sound. With it (and the speech-feature targets in the backup below) the training can be repeated without encoding the 1.5 TB of audio again.
  • The full working backup: VoiceHub/DACFlow-EN-10k-backup: everything else the training and its evaluation use (speech-feature (REPA) targets, annotations, evaluation sets, experiments), kept to resume them; not needed to use the model.
  • The source audio: SynDataLab-EN/echo-clones-4m-en: about 4 million synthetic English clips (EchoTTS voice clones of 4,000 reference voices, ~11,150 hours; apache-2.0). 236 of its voices are held out for testing and never trained on.
  • Nothing else is used for training: a checkpoint is published here only when its training catalog (and any checkpoint it started from) records no other source (the allowlist configs/data/en_train_sources.txt, checked by scripts/hub_showcase.py before every upload).

Licence

apache-2.0 (owner decision 2026-09-29, hobby project) for the weights and the samples. The voice prompts are clips of held-out voices of SynDataLab-EN/echo-clones-4m-en (apache-2.0 on its card).

Limitations and responsible use

  • The model learned only from synthetic voices: copying real people's voices is its weakest point (compare the seed-dev and echo-dev SIM-o above). English only; rare words and long numbers can be mispronounced.
  • Copy a voice only with its owner's consent, never to impersonate anyone or to deceive, and say that the speech is synthetic.
  • Every output of this code carries an inaudible AudioSeal watermark (16-bit message 0100110101010100; python scripts/watermark_check.py detect <files> finds it). The mark is added by the inference code, not by the weights: do not remove it.
Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Datasets used to train VoiceHub/DACFlow-EN-10k