Text-to-Speech
AudioSeal
Safetensors
English
zero-shot
voice-cloning
english
flow-matching
diffusion-transformer
dacvae
Instructions to use VoiceHub/DACFlow-EN-10k with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- AudioSeal
How to use VoiceHub/DACFlow-EN-10k with AudioSeal:
# Watermark Generator from audioseal import AudioSeal model = AudioSeal.load_generator("VoiceHub/DACFlow-EN-10k") # pass a tensor (tensor_wav) of shape (batch, channels, samples) and a sample rate wav, sr = tensor_wav, 16000 watermark = model.get_watermark(wav, sr) watermarked_audio = wav + watermark# Watermark Detector from audioseal import AudioSeal detector = AudioSeal.load_detector("VoiceHub/DACFlow-EN-10k") result, message = detector.detect_watermark(watermarked_audio, sr) - Notebooks
- Google Colab
- Kaggle
showcase: step 100k: 5 listening samples + the prompt-choice comparison (3 items with their old prompt) (AudioSeal-watermarked), delete 85 files the manifest no longer lists (an older listening set)
Browse filesThis view is limited to 50 files because it contains too many changes. See raw diff
- .gitattributes +6 -0
- README.md +79 -271
- prompts/02-question.wav +2 -2
- prompts/03-numbers.wav +1 -1
- prompts/05-long.wav +2 -2
- samples/{step_0030000 → prompt_choice/old_output}/02-question.wav +1 -1
- samples/{step_0030000 → prompt_choice/old_output}/03-numbers.wav +1 -1
- samples/{step_0020000 → prompt_choice/old_output}/05-long.wav +1 -1
- samples/{step_0020000 → prompt_choice/old_prompt}/02-question.wav +2 -2
- samples/{step_0040000 → prompt_choice/old_prompt}/03-numbers.wav +2 -2
- samples/{step_0020000/01-short.wav → prompt_choice/old_prompt/05-long.wav} +2 -2
- samples/step_0020000/03-numbers.wav +0 -3
- samples/step_0020000/04-conversational.wav +0 -3
- samples/step_0030000/01-short.wav +0 -3
- samples/step_0030000/04-conversational.wav +0 -3
- samples/step_0030000/05-long.wav +0 -3
- samples/step_0040000/01-short.wav +0 -3
- samples/step_0040000/02-question.wav +0 -3
- samples/step_0040000/04-conversational.wav +0 -3
- samples/step_0040000/05-long.wav +0 -3
- samples/step_0050000/01-short.wav +0 -3
- samples/step_0050000/02-question.wav +0 -3
- samples/step_0050000/03-numbers.wav +0 -3
- samples/step_0050000/04-conversational.wav +0 -3
- samples/step_0050000/05-long.wav +0 -3
- samples/step_0060000/01-short.wav +0 -3
- samples/step_0060000/02-question.wav +0 -3
- samples/step_0060000/03-numbers.wav +0 -3
- samples/step_0060000/04-conversational.wav +0 -3
- samples/step_0060000/05-long.wav +0 -3
- samples/step_0070000/01-short.wav +0 -3
- samples/step_0070000/02-question.wav +0 -3
- samples/step_0070000/03-numbers.wav +0 -3
- samples/step_0070000/04-conversational.wav +0 -3
- samples/step_0070000/05-long.wav +0 -3
- samples/step_0080000/01-short.wav +0 -3
- samples/step_0080000/02-question.wav +0 -3
- samples/step_0080000/03-numbers.wav +0 -3
- samples/step_0080000/04-conversational.wav +0 -3
- samples/step_0080000/05-long.wav +0 -3
- samples/step_0090000/01-short.wav +0 -3
- samples/step_0090000/02-question.wav +0 -3
- samples/step_0090000/03-numbers.wav +0 -3
- samples/step_0090000/04-conversational.wav +0 -3
- samples/step_0090000/05-long.wav +0 -3
- samples/step_0100000/01-short.wav +1 -1
- samples/step_0100000/02-question.wav +2 -2
- samples/step_0100000/03-numbers.wav +2 -2
- samples/step_0100000/04-conversational.wav +1 -1
- samples/step_0100000/05-long.wav +2 -2
.gitattributes
CHANGED
|
@@ -128,3 +128,9 @@ samples/step_0190000/02-question.wav filter=lfs diff=lfs merge=lfs -text
|
|
| 128 |
samples/step_0190000/03-numbers.wav filter=lfs diff=lfs merge=lfs -text
|
| 129 |
samples/step_0190000/04-conversational.wav filter=lfs diff=lfs merge=lfs -text
|
| 130 |
samples/step_0190000/05-long.wav filter=lfs diff=lfs merge=lfs -text
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 128 |
samples/step_0190000/03-numbers.wav filter=lfs diff=lfs merge=lfs -text
|
| 129 |
samples/step_0190000/04-conversational.wav filter=lfs diff=lfs merge=lfs -text
|
| 130 |
samples/step_0190000/05-long.wav filter=lfs diff=lfs merge=lfs -text
|
| 131 |
+
samples/prompt_choice/old_output/02-question.wav filter=lfs diff=lfs merge=lfs -text
|
| 132 |
+
samples/prompt_choice/old_output/03-numbers.wav filter=lfs diff=lfs merge=lfs -text
|
| 133 |
+
samples/prompt_choice/old_output/05-long.wav filter=lfs diff=lfs merge=lfs -text
|
| 134 |
+
samples/prompt_choice/old_prompt/02-question.wav filter=lfs diff=lfs merge=lfs -text
|
| 135 |
+
samples/prompt_choice/old_prompt/03-numbers.wav filter=lfs diff=lfs merge=lfs -text
|
| 136 |
+
samples/prompt_choice/old_prompt/05-long.wav filter=lfs diff=lfs merge=lfs -text
|
README.md
CHANGED
|
@@ -31,295 +31,93 @@ The name: **DAC** for the Semantic-DACVAE audio codec whose latents it generates
|
|
| 31 |
English, **10k** for its training data, about ten thousand hours of speech.
|
| 32 |
The code: [kadirnar/dacvae-next](https://github.com/kadirnar/dacvae-next/tree/roadmap/en-echo), Python package `mytts`.
|
| 33 |
|
| 34 |
-
> **Status: training
|
| 35 |
-
>
|
| 36 |
-
>
|
|
|
|
|
|
|
|
|
|
| 37 |
|
| 38 |
**On this page:** [Listen](#listen) · [Results](#results) · [How to use](#how-to-use) · [Training](#training) · [Data](#data) · [Licence](#licence)
|
| 39 |
|
| 40 |
## Listen
|
| 41 |
|
| 42 |
-
The
|
| 43 |
-
|
| 44 |
|
| 45 |
<table>
|
| 46 |
-
<tr><th>What the model reads</th><th>Voice prompt (the input)</th><th>Model output
|
| 47 |
-
<tr><td><b>Short sentence</b><br>I left my umbrella at the office again, so I'm definitely getting soaked on the way home.</td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/prompts/01-short.wav"></audio><br><sub>A higher voice (~202 Hz). It says: “I keep telling myself I just need to power through, like, push a little harder, and then it'll be fine. But it's not getting fine. It's getting— it's getting worse.”</sub></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/
|
| 48 |
-
<tr><td><b>Question</b><br>Have you ever noticed that the quietest person in the room usually has the most interesting story to tell?</td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/prompts/02-question.wav"></audio><br><sub>A lower voice (~
|
| 49 |
-
<tr><td><b>Numbers, dates, abbreviations</b><br>Dr. Patel moved my appointment to Tuesday, March 3rd, at 4:15 p.m., and the co-pay went up from $20 to $35.</td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/prompts/03-numbers.wav"></audio><br><sub>A higher voice (~
|
| 50 |
-
<tr><td><b>Conversation (~10 s)</b><br>So I finally tried that new ramen place downtown, and honestly, it was worth the wait. The broth was amazing, but next time I'm definitely skipping the extra spicy option.</td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/prompts/04-conversational.wav"></audio><br><sub>A lower voice (~115 Hz). It says: “I'm sorry, but the subway doors closed on my bag. Again.”</sub></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/
|
| 51 |
-
<tr><td><b>Long passage (~20 s)</b><br>When the storm passed, the whole neighborhood came outside to look at the damage, and although a few fences had fallen and the old oak tree had lost its biggest branch, everyone was relieved that nobody was hurt. They spent the rest of the afternoon clearing the street and sharing whatever food they had left.</td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/prompts/05-long.wav"></audio><br><sub>The deepest voice (~
|
| 52 |
</table>
|
| 53 |
|
| 54 |
-
|
| 55 |
-
|
| 56 |
-
|
| 57 |
-
<summary><b>Step 180k</b>: echo-dev WER 1.04 % · UTMOS 3.98</summary>
|
| 58 |
-
|
| 59 |
-
<table>
|
| 60 |
-
<tr><th>What the model reads</th><th>Model output, step 180k</th></tr>
|
| 61 |
-
<tr><td><b>Short sentence</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0180000/01-short.wav"></audio></td></tr>
|
| 62 |
-
<tr><td><b>Question</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0180000/02-question.wav"></audio></td></tr>
|
| 63 |
-
<tr><td><b>Numbers, dates, abbreviations</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0180000/03-numbers.wav"></audio></td></tr>
|
| 64 |
-
<tr><td><b>Conversation (~10 s)</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0180000/04-conversational.wav"></audio></td></tr>
|
| 65 |
-
<tr><td><b>Long passage (~20 s)</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0180000/05-long.wav"></audio></td></tr>
|
| 66 |
-
</table>
|
| 67 |
-
|
| 68 |
-
</details>
|
| 69 |
-
|
| 70 |
-
<details>
|
| 71 |
-
<summary><b>Step 170k</b>: echo-dev WER 1.01 % · UTMOS 3.97</summary>
|
| 72 |
-
|
| 73 |
-
<table>
|
| 74 |
-
<tr><th>What the model reads</th><th>Model output, step 170k</th></tr>
|
| 75 |
-
<tr><td><b>Short sentence</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0170000/01-short.wav"></audio></td></tr>
|
| 76 |
-
<tr><td><b>Question</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0170000/02-question.wav"></audio></td></tr>
|
| 77 |
-
<tr><td><b>Numbers, dates, abbreviations</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0170000/03-numbers.wav"></audio></td></tr>
|
| 78 |
-
<tr><td><b>Conversation (~10 s)</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0170000/04-conversational.wav"></audio></td></tr>
|
| 79 |
-
<tr><td><b>Long passage (~20 s)</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0170000/05-long.wav"></audio></td></tr>
|
| 80 |
-
</table>
|
| 81 |
-
|
| 82 |
-
</details>
|
| 83 |
-
|
| 84 |
-
<details>
|
| 85 |
-
<summary><b>Step 160k</b>: echo-dev WER 0.81 % · UTMOS 3.98</summary>
|
| 86 |
-
|
| 87 |
-
<table>
|
| 88 |
-
<tr><th>What the model reads</th><th>Model output, step 160k</th></tr>
|
| 89 |
-
<tr><td><b>Short sentence</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0160000/01-short.wav"></audio></td></tr>
|
| 90 |
-
<tr><td><b>Question</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0160000/02-question.wav"></audio></td></tr>
|
| 91 |
-
<tr><td><b>Numbers, dates, abbreviations</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0160000/03-numbers.wav"></audio></td></tr>
|
| 92 |
-
<tr><td><b>Conversation (~10 s)</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0160000/04-conversational.wav"></audio></td></tr>
|
| 93 |
-
<tr><td><b>Long passage (~20 s)</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0160000/05-long.wav"></audio></td></tr>
|
| 94 |
-
</table>
|
| 95 |
-
|
| 96 |
-
</details>
|
| 97 |
-
|
| 98 |
-
<details>
|
| 99 |
-
<summary><b>Step 150k</b>: echo-dev WER 0.53 % · UTMOS 3.93</summary>
|
| 100 |
-
|
| 101 |
-
<table>
|
| 102 |
-
<tr><th>What the model reads</th><th>Model output, step 150k</th></tr>
|
| 103 |
-
<tr><td><b>Short sentence</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0150000/01-short.wav"></audio></td></tr>
|
| 104 |
-
<tr><td><b>Question</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0150000/02-question.wav"></audio></td></tr>
|
| 105 |
-
<tr><td><b>Numbers, dates, abbreviations</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0150000/03-numbers.wav"></audio></td></tr>
|
| 106 |
-
<tr><td><b>Conversation (~10 s)</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0150000/04-conversational.wav"></audio></td></tr>
|
| 107 |
-
<tr><td><b>Long passage (~20 s)</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0150000/05-long.wav"></audio></td></tr>
|
| 108 |
-
</table>
|
| 109 |
-
|
| 110 |
-
</details>
|
| 111 |
-
|
| 112 |
-
<details>
|
| 113 |
-
<summary><b>Step 140k</b>: echo-dev WER 0.58 % · UTMOS 3.92</summary>
|
| 114 |
-
|
| 115 |
-
<table>
|
| 116 |
-
<tr><th>What the model reads</th><th>Model output, step 140k</th></tr>
|
| 117 |
-
<tr><td><b>Short sentence</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0140000/01-short.wav"></audio></td></tr>
|
| 118 |
-
<tr><td><b>Question</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0140000/02-question.wav"></audio></td></tr>
|
| 119 |
-
<tr><td><b>Numbers, dates, abbreviations</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0140000/03-numbers.wav"></audio></td></tr>
|
| 120 |
-
<tr><td><b>Conversation (~10 s)</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0140000/04-conversational.wav"></audio></td></tr>
|
| 121 |
-
<tr><td><b>Long passage (~20 s)</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0140000/05-long.wav"></audio></td></tr>
|
| 122 |
-
</table>
|
| 123 |
-
|
| 124 |
-
</details>
|
| 125 |
-
|
| 126 |
-
<details>
|
| 127 |
-
<summary><b>Step 130k</b>: echo-dev WER 0.76 % · UTMOS 3.93</summary>
|
| 128 |
-
|
| 129 |
-
<table>
|
| 130 |
-
<tr><th>What the model reads</th><th>Model output, step 130k</th></tr>
|
| 131 |
-
<tr><td><b>Short sentence</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0130000/01-short.wav"></audio></td></tr>
|
| 132 |
-
<tr><td><b>Question</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0130000/02-question.wav"></audio></td></tr>
|
| 133 |
-
<tr><td><b>Numbers, dates, abbreviations</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0130000/03-numbers.wav"></audio></td></tr>
|
| 134 |
-
<tr><td><b>Conversation (~10 s)</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0130000/04-conversational.wav"></audio></td></tr>
|
| 135 |
-
<tr><td><b>Long passage (~20 s)</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0130000/05-long.wav"></audio></td></tr>
|
| 136 |
-
</table>
|
| 137 |
-
|
| 138 |
-
</details>
|
| 139 |
-
|
| 140 |
-
<details>
|
| 141 |
-
<summary><b>Step 120k</b>: echo-dev WER 0.53 % · UTMOS 3.89</summary>
|
| 142 |
-
|
| 143 |
-
<table>
|
| 144 |
-
<tr><th>What the model reads</th><th>Model output, step 120k</th></tr>
|
| 145 |
-
<tr><td><b>Short sentence</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0120000/01-short.wav"></audio></td></tr>
|
| 146 |
-
<tr><td><b>Question</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0120000/02-question.wav"></audio></td></tr>
|
| 147 |
-
<tr><td><b>Numbers, dates, abbreviations</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0120000/03-numbers.wav"></audio></td></tr>
|
| 148 |
-
<tr><td><b>Conversation (~10 s)</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0120000/04-conversational.wav"></audio></td></tr>
|
| 149 |
-
<tr><td><b>Long passage (~20 s)</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0120000/05-long.wav"></audio></td></tr>
|
| 150 |
-
</table>
|
| 151 |
-
|
| 152 |
-
</details>
|
| 153 |
-
|
| 154 |
-
<details>
|
| 155 |
-
<summary><b>Step 110k</b>: echo-dev WER 1.10 % · UTMOS 3.86</summary>
|
| 156 |
-
|
| 157 |
-
<table>
|
| 158 |
-
<tr><th>What the model reads</th><th>Model output, step 110k</th></tr>
|
| 159 |
-
<tr><td><b>Short sentence</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0110000/01-short.wav"></audio></td></tr>
|
| 160 |
-
<tr><td><b>Question</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0110000/02-question.wav"></audio></td></tr>
|
| 161 |
-
<tr><td><b>Numbers, dates, abbreviations</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0110000/03-numbers.wav"></audio></td></tr>
|
| 162 |
-
<tr><td><b>Conversation (~10 s)</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0110000/04-conversational.wav"></audio></td></tr>
|
| 163 |
-
<tr><td><b>Long passage (~20 s)</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0110000/05-long.wav"></audio></td></tr>
|
| 164 |
-
</table>
|
| 165 |
-
|
| 166 |
-
</details>
|
| 167 |
-
|
| 168 |
-
<details>
|
| 169 |
-
<summary><b>Step 100k</b>: echo-dev WER 0.69 % · UTMOS 3.85</summary>
|
| 170 |
-
|
| 171 |
-
<table>
|
| 172 |
-
<tr><th>What the model reads</th><th>Model output, step 100k</th></tr>
|
| 173 |
-
<tr><td><b>Short sentence</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0100000/01-short.wav"></audio></td></tr>
|
| 174 |
-
<tr><td><b>Question</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0100000/02-question.wav"></audio></td></tr>
|
| 175 |
-
<tr><td><b>Numbers, dates, abbreviations</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0100000/03-numbers.wav"></audio></td></tr>
|
| 176 |
-
<tr><td><b>Conversation (~10 s)</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0100000/04-conversational.wav"></audio></td></tr>
|
| 177 |
-
<tr><td><b>Long passage (~20 s)</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0100000/05-long.wav"></audio></td></tr>
|
| 178 |
-
</table>
|
| 179 |
-
|
| 180 |
-
</details>
|
| 181 |
-
|
| 182 |
-
<details>
|
| 183 |
-
<summary><b>Step 90k</b>: echo-dev WER 0.74 % · UTMOS 3.78</summary>
|
| 184 |
-
|
| 185 |
-
<table>
|
| 186 |
-
<tr><th>What the model reads</th><th>Model output, step 90k</th></tr>
|
| 187 |
-
<tr><td><b>Short sentence</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0090000/01-short.wav"></audio></td></tr>
|
| 188 |
-
<tr><td><b>Question</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0090000/02-question.wav"></audio></td></tr>
|
| 189 |
-
<tr><td><b>Numbers, dates, abbreviations</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0090000/03-numbers.wav"></audio></td></tr>
|
| 190 |
-
<tr><td><b>Conversation (~10 s)</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0090000/04-conversational.wav"></audio></td></tr>
|
| 191 |
-
<tr><td><b>Long passage (~20 s)</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0090000/05-long.wav"></audio></td></tr>
|
| 192 |
-
</table>
|
| 193 |
-
|
| 194 |
-
</details>
|
| 195 |
-
|
| 196 |
-
<details>
|
| 197 |
-
<summary><b>Step 80k</b>: echo-dev WER 0.78 % · UTMOS 3.73</summary>
|
| 198 |
-
|
| 199 |
-
<table>
|
| 200 |
-
<tr><th>What the model reads</th><th>Model output, step 80k</th></tr>
|
| 201 |
-
<tr><td><b>Short sentence</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0080000/01-short.wav"></audio></td></tr>
|
| 202 |
-
<tr><td><b>Question</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0080000/02-question.wav"></audio></td></tr>
|
| 203 |
-
<tr><td><b>Numbers, dates, abbreviations</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0080000/03-numbers.wav"></audio></td></tr>
|
| 204 |
-
<tr><td><b>Conversation (~10 s)</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0080000/04-conversational.wav"></audio></td></tr>
|
| 205 |
-
<tr><td><b>Long passage (~20 s)</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0080000/05-long.wav"></audio></td></tr>
|
| 206 |
-
</table>
|
| 207 |
-
|
| 208 |
-
</details>
|
| 209 |
-
|
| 210 |
-
<details>
|
| 211 |
-
<summary><b>Step 70k</b>: echo-dev WER 0.67 % · UTMOS 3.69</summary>
|
| 212 |
-
|
| 213 |
-
<table>
|
| 214 |
-
<tr><th>What the model reads</th><th>Model output, step 70k</th></tr>
|
| 215 |
-
<tr><td><b>Short sentence</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0070000/01-short.wav"></audio></td></tr>
|
| 216 |
-
<tr><td><b>Question</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0070000/02-question.wav"></audio></td></tr>
|
| 217 |
-
<tr><td><b>Numbers, dates, abbreviations</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0070000/03-numbers.wav"></audio></td></tr>
|
| 218 |
-
<tr><td><b>Conversation (~10 s)</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0070000/04-conversational.wav"></audio></td></tr>
|
| 219 |
-
<tr><td><b>Long passage (~20 s)</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0070000/05-long.wav"></audio></td></tr>
|
| 220 |
-
</table>
|
| 221 |
-
|
| 222 |
-
</details>
|
| 223 |
-
|
| 224 |
-
<details>
|
| 225 |
-
<summary><b>Step 60k</b>: echo-dev WER 0.69 % · UTMOS 3.62</summary>
|
| 226 |
-
|
| 227 |
-
<table>
|
| 228 |
-
<tr><th>What the model reads</th><th>Model output, step 60k</th></tr>
|
| 229 |
-
<tr><td><b>Short sentence</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0060000/01-short.wav"></audio></td></tr>
|
| 230 |
-
<tr><td><b>Question</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0060000/02-question.wav"></audio></td></tr>
|
| 231 |
-
<tr><td><b>Numbers, dates, abbreviations</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0060000/03-numbers.wav"></audio></td></tr>
|
| 232 |
-
<tr><td><b>Conversation (~10 s)</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0060000/04-conversational.wav"></audio></td></tr>
|
| 233 |
-
<tr><td><b>Long passage (~20 s)</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0060000/05-long.wav"></audio></td></tr>
|
| 234 |
-
</table>
|
| 235 |
-
|
| 236 |
-
</details>
|
| 237 |
-
|
| 238 |
-
<details>
|
| 239 |
-
<summary><b>Step 50k</b>: echo-dev WER 0.74 % · UTMOS 3.56</summary>
|
| 240 |
-
|
| 241 |
-
<table>
|
| 242 |
-
<tr><th>What the model reads</th><th>Model output, step 50k</th></tr>
|
| 243 |
-
<tr><td><b>Short sentence</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0050000/01-short.wav"></audio></td></tr>
|
| 244 |
-
<tr><td><b>Question</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0050000/02-question.wav"></audio></td></tr>
|
| 245 |
-
<tr><td><b>Numbers, dates, abbreviations</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0050000/03-numbers.wav"></audio></td></tr>
|
| 246 |
-
<tr><td><b>Conversation (~10 s)</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0050000/04-conversational.wav"></audio></td></tr>
|
| 247 |
-
<tr><td><b>Long passage (~20 s)</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0050000/05-long.wav"></audio></td></tr>
|
| 248 |
-
</table>
|
| 249 |
|
| 250 |
-
|
| 251 |
|
| 252 |
-
|
| 253 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 254 |
|
| 255 |
<table>
|
| 256 |
-
<tr><th>What the model reads</th><th>
|
| 257 |
-
<tr><td><b>
|
| 258 |
-
<tr><td><b>
|
| 259 |
-
<tr><td><b>
|
| 260 |
-
<tr><td><b>Conversation (~10 s)</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0040000/04-conversational.wav"></audio></td></tr>
|
| 261 |
-
<tr><td><b>Long passage (~20 s)</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0040000/05-long.wav"></audio></td></tr>
|
| 262 |
</table>
|
| 263 |
|
| 264 |
-
|
| 265 |
-
|
| 266 |
-
|
| 267 |
-
<summary><b>Step 30k</b>: echo-dev WER 1.06 % · UTMOS 3.43</summary>
|
| 268 |
|
| 269 |
-
|
| 270 |
-
<tr><th>What the model reads</th><th>Model output, step 30k</th></tr>
|
| 271 |
-
<tr><td><b>Short sentence</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0030000/01-short.wav"></audio></td></tr>
|
| 272 |
-
<tr><td><b>Question</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0030000/02-question.wav"></audio></td></tr>
|
| 273 |
-
<tr><td><b>Numbers, dates, abbreviations</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0030000/03-numbers.wav"></audio></td></tr>
|
| 274 |
-
<tr><td><b>Conversation (~10 s)</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0030000/04-conversational.wav"></audio></td></tr>
|
| 275 |
-
<tr><td><b>Long passage (~20 s)</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0030000/05-long.wav"></audio></td></tr>
|
| 276 |
-
</table>
|
| 277 |
-
|
| 278 |
-
</details>
|
| 279 |
|
| 280 |
-
|
| 281 |
-
<summary><b>Step 20k</b>: echo-dev WER 1.20 % · UTMOS 3.33</summary>
|
| 282 |
|
| 283 |
-
|
| 284 |
-
|
| 285 |
-
<tr><td><b>Short sentence</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0020000/01-short.wav"></audio></td></tr>
|
| 286 |
-
<tr><td><b>Question</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0020000/02-question.wav"></audio></td></tr>
|
| 287 |
-
<tr><td><b>Numbers, dates, abbreviations</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0020000/03-numbers.wav"></audio></td></tr>
|
| 288 |
-
<tr><td><b>Conversation (~10 s)</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0020000/04-conversational.wav"></audio></td></tr>
|
| 289 |
-
<tr><td><b>Long passage (~20 s)</b></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0020000/05-long.wav"></audio></td></tr>
|
| 290 |
-
</table>
|
| 291 |
|
| 292 |
-
|
|
|
|
|
|
|
|
|
|
| 293 |
|
| 294 |
-
|
| 295 |
-
|
| 296 |
-
are the original recordings.
|
| 297 |
|
| 298 |
-
##
|
| 299 |
|
| 300 |
-
|
| 301 |
-
UTMOS** are better):
|
| 302 |
|
| 303 |
| Checkpoint | echo-dev WER ↓ | echo-dev SIM-o ↑ | echo-dev UTMOS ↑ (real speech: 4.20) | seed-dev WER ↓ | seed-dev SIM-o ↑ | seed-dev UTMOS ↑ (real speech: 3.52) | Samples | Weights |
|
| 304 |
|---|---:|---:|---:|---:|---:|---:|---|---|
|
| 305 |
-
|
|
| 306 |
-
| step 180k | 1.04 % | 0.792 | 3.98 | 1.65 % | 0.261 | 3.29 |
|
| 307 |
-
| step 170k | 1.01 % | 0.796 | 3.97 | - | - | - |
|
| 308 |
-
| step 160k | 0.81 % | 0.796 | 3.98 | 1.58 % | 0.349 | 3.66 |
|
| 309 |
-
| step 150k | 0.53 % | 0.794 | 3.93 | - | - | - |
|
| 310 |
-
| step 140k | 0.58 % | 0.793 | 3.92 | 1.60 % | 0.404 | 3.81 |
|
| 311 |
-
| step 130k | 0.76 % | 0.792 | 3.93 | - | - | - |
|
| 312 |
-
| step 120k | 0.53 % | 0.790 | 3.89 | 1.45 % | 0.451 | 3.86 |
|
| 313 |
-
| step 110k | 1.10 % | 0.789 | 3.86 | - | - | - |
|
| 314 |
-
| step 100k | 0.69 % | 0.787 | 3.85 | 1.54 % | 0.499 | 3.78 | [listen](#listen) | [files](https://huggingface.co/VoiceHub/DACFlow-EN-10k/tree/main/checkpoints/step_0100000) |
|
| 315 |
-
| step 90k | 0.74 % | 0.782 | 3.78 | - | - | - |
|
| 316 |
-
| step 80k | 0.78 % | 0.776 | 3.73 | 1.58 % | 0.507 | 3.68 |
|
| 317 |
-
| step 70k | 0.67 % | 0.766 | 3.69 | - | - | - |
|
| 318 |
-
| step 60k | 0.69 % | 0.759 | 3.62 | 1.61 % | 0.537 | 3.56 |
|
| 319 |
-
| step 50k | 0.74 % | 0.754 | 3.56 | - | - | - |
|
| 320 |
-
| step 40k | 0.74 % | 0.748 | 3.50 | 1.85 % | 0.536 | 3.48 |
|
| 321 |
-
| step 30k | 1.06 % | 0.741 | 3.43 | - | - | - |
|
| 322 |
-
| step 20k | 1.20 % | 0.722 | 3.33 | 2.47 % | 0.506 | 3.27 |
|
| 323 |
|
| 324 |
- **WER** (word error rate): Whisper-large-v3 writes down what it hears and this is compared with the text; 2 % means about
|
| 325 |
one word in fifty is wrong.
|
|
@@ -329,8 +127,10 @@ UTMOS** are better):
|
|
| 329 |
- **echo-dev**: held-out voices of the training data's kind (synthetic EchoTTS voices, none of them trained on). **seed-dev**:
|
| 330 |
real people's voices (the dev half of Seed-TTS test-en, Common Voice recordings). seed-dev is harder: the model learned
|
| 331 |
only from synthetic voices.
|
|
|
|
|
|
|
| 332 |
- `-`: not measured at that step (echo-dev is scored every 10k steps, seed-dev every 20k). Every number averages two takes
|
| 333 |
-
per sentence.
|
| 334 |
|
| 335 |
## How to use
|
| 336 |
|
|
@@ -339,6 +139,7 @@ UTMOS** are better):
|
|
| 339 |
```bash
|
| 340 |
git clone -b roadmap/en-echo https://github.com/kadirnar/dacvae-next
|
| 341 |
cd dacvae-next
|
|
|
|
| 342 |
pip install -e ".[codec]"
|
| 343 |
```
|
| 344 |
|
|
@@ -352,9 +153,9 @@ from huggingface_hub import snapshot_download
|
|
| 352 |
from mytts.flow.sampler import SamplerConfig
|
| 353 |
from mytts.infer import Synthesizer
|
| 354 |
|
| 355 |
-
ckpt = "checkpoints/
|
| 356 |
local = snapshot_download("VoiceHub/DACFlow-EN-10k", allow_patterns=[f"{ckpt}/*"]) # downloads only this checkpoint
|
| 357 |
-
tts = Synthesizer.from_export(f"{local}/{ckpt}", device="cuda")
|
| 358 |
sc = SamplerConfig(steps=32, cfg_mode="joint", cfg_w=4.0, noise_scale=0.9) # the settings of the samples above
|
| 359 |
wav, sr = tts.synthesize(
|
| 360 |
"Any English text you like.",
|
|
@@ -362,21 +163,27 @@ wav, sr = tts.synthesize(
|
|
| 362 |
sc=sc,
|
| 363 |
prompt_audio="my_voice.wav", # the voice to copy
|
| 364 |
prompt_text="The exact words spoken in my_voice.wav.",
|
|
|
|
| 365 |
out_lufs=-16.0,
|
| 366 |
seed=0,
|
| 367 |
)
|
| 368 |
sf.write("output.wav", wav, sr) # 48 kHz, with the AudioSeal watermark
|
| 369 |
```
|
| 370 |
|
| 371 |
-
**Or from the command line**
|
| 372 |
|
| 373 |
```bash
|
| 374 |
-
hf download VoiceHub/DACFlow-EN-10k --include "checkpoints/
|
| 375 |
-
python scripts/synthesize.py --model DACFlow-EN-10k/checkpoints/
|
| 376 |
--text "Any English text you like." --prompt-audio my_voice.wav \
|
| 377 |
--prompt-text "The exact words spoken in my_voice.wav." --out output.wav
|
| 378 |
```
|
| 379 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 380 |
Tips: a clean prompt with an exact transcript works best. A long text is split at sentence ends and read piece by piece.
|
| 381 |
`python scripts/watermark_check.py detect output.wav` checks the watermark.
|
| 382 |
|
|
@@ -386,7 +193,8 @@ Tips: a clean prompt with an exact transcript works best. A long text is split a
|
|
| 386 |
generates the latents of the [Semantic-DACVAE](https://huggingface.co/Aratako/Semantic-DACVAE-Japanese) audio codec (48 kHz,
|
| 387 |
25 frames per second) from the text, continuing the voice prompt in context (no separate speaker encoder).
|
| 388 |
- **Data**: the curated training catalog `echo_en_v1` of the provided data ([EN-10, #34](https://github.com/kadirnar/dacvae-next/issues/34) curation; see [Data](#data)).
|
| 389 |
-
- **Recipe**: 200k steps on one RTX 5090, about 27 minutes of speech per step (40,000 latent frames); learning rate 2.5e-4 after 5k warm-up steps, held, then lowered over the last 20 % of the steps (from step 160k).
|
|
|
|
| 390 |
The settings were chosen with small screening runs: the flow-matching noise schedule t_mean -0.8 / t_std 0.8 ([EN-55, #100](https://github.com/kadirnar/dacvae-next/issues/100))
|
| 391 |
and cross-utterance voice prompts with p_cross 0.6 ([EN-29, #53](https://github.com/kadirnar/dacvae-next/issues/53)).
|
| 392 |
- **Code and history**: the recipe [configs/train/en_full.yaml](https://github.com/kadirnar/dacvae-next/blob/roadmap/en-echo/configs/train/en_full.yaml)
|
|
|
|
| 31 |
English, **10k** for its training data, about ten thousand hours of speech.
|
| 32 |
The code: [kadirnar/dacvae-next](https://github.com/kadirnar/dacvae-next/tree/roadmap/en-echo), Python package `mytts`.
|
| 33 |
|
| 34 |
+
> **Status: training is finished. The model is step 100k** (the EMA weights saved at that step). The run was planned for 200k steps
|
| 35 |
+
> and stopped at step 194,271 on 2026-10-04; this page no longer changes by itself. Why step 100k: later checkpoints sound a little
|
| 36 |
+
> better on held-out voices of the training data's kind, but copy real people's voices much worse. Speaker similarity on real voices
|
| 37 |
+
> (seed-dev SIM-o) falls from 0.50 at step 100k to 0.22 at step 190k, where 51 % of the outputs score below 0.2 (5.7 % at step
|
| 38 |
+
> 100k), with the same settings. Step 100k, with each prompt's own bandwidth as the condition (auto), is the best balance. Details:
|
| 39 |
+
> [#177](https://github.com/kadirnar/dacvae-next/issues/177).
|
| 40 |
|
| 41 |
**On this page:** [Listen](#listen) · [Results](#results) · [How to use](#how-to-use) · [Training](#training) · [Data](#data) · [Licence](#licence)
|
| 42 |
|
| 43 |
## Listen
|
| 44 |
|
| 45 |
+
The model, **step 100k**, reads five texts. Each text is spoken in a different voice, copied from the short voice prompt next
|
| 46 |
+
to it. The model never heard these voices in training (held-out voices of the provided data).
|
| 47 |
|
| 48 |
<table>
|
| 49 |
+
<tr><th>What the model reads</th><th>Voice prompt (the input)</th><th>Model output (step 100k)</th></tr>
|
| 50 |
+
<tr><td><b>Short sentence</b><br>I left my umbrella at the office again, so I'm definitely getting soaked on the way home.</td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/prompts/01-short.wav"></audio><br><sub>A higher voice (~202 Hz). It says: “I keep telling myself I just need to power through, like, push a little harder, and then it'll be fine. But it's not getting fine. It's getting— it's getting worse.”</sub></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0100000/01-short.wav"></audio></td></tr>
|
| 51 |
+
<tr><td><b>Question</b><br>Have you ever noticed that the quietest person in the room usually has the most interesting story to tell?</td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/prompts/02-question.wav"></audio><br><sub>A lower voice (~119 Hz). It says: “Look, I'm not saying it's easy, but you always pull it together. Just take it step by step, you know?”</sub></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0100000/02-question.wav"></audio></td></tr>
|
| 52 |
+
<tr><td><b>Numbers, dates, abbreviations</b><br>Dr. Patel moved my appointment to Tuesday, March 3rd, at 4:15 p.m., and the co-pay went up from $20 to $35.</td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/prompts/03-numbers.wav"></audio><br><sub>A higher voice (~201 Hz). It says: “Would you just— I know you mean well, but every time you mention it I feel like an idiot. It's probably nothing.”</sub></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0100000/03-numbers.wav"></audio></td></tr>
|
| 53 |
+
<tr><td><b>Conversation (~10 s)</b><br>So I finally tried that new ramen place downtown, and honestly, it was worth the wait. The broth was amazing, but next time I'm definitely skipping the extra spicy option.</td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/prompts/04-conversational.wav"></audio><br><sub>A lower voice (~115 Hz). It says: “I'm sorry, but the subway doors closed on my bag. Again.”</sub></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0100000/04-conversational.wav"></audio></td></tr>
|
| 54 |
+
<tr><td><b>Long passage (~20 s)</b><br>When the storm passed, the whole neighborhood came outside to look at the damage, and although a few fences had fallen and the old oak tree had lost its biggest branch, everyone was relieved that nobody was hurt. They spent the rest of the afternoon clearing the street and sharing whatever food they had left.</td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/prompts/05-long.wav"></audio><br><sub>The deepest voice (~102 Hz). It says: “I mean, come on, it's not like I woke up and decided to have the worst day ever. Things just went wrong one after another, like dominoes!”</sub></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0100000/05-long.wav"></audio></td></tr>
|
| 55 |
</table>
|
| 56 |
|
| 57 |
+
How the samples are made: one take per text, no cherry-picking, with the settings of [How to use](#how-to-use) (32 Euler steps + sway, joint CFG w 4.0, initial noise 0.9, at most -16 LUFS (peak-limited), duration by the band rule, each prompt's own bandwidth as the condition (auto, clamped to 9,798-16,839 Hz)).
|
| 58 |
+
Every output carries an inaudible [AudioSeal](https://huggingface.co/facebook/audioseal) watermark, checked before upload; the prompts
|
| 59 |
+
are the original recordings.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 60 |
|
| 61 |
+
### How the sample voices were chosen
|
| 62 |
|
| 63 |
+
The voice prompt decides most of how an output sounds. In a test of this model on 45 held-out voices, 10 texts each, the choice of
|
| 64 |
+
prompt explained about 63 % of the differences in predicted naturalness (UTMOS), the text about 4 %. The model also copies the
|
| 65 |
+
recording: a prompt with a narrow frequency band (a dull, muffled recording) or with background noise gives an output that sounds
|
| 66 |
+
the same. So three of the five prompts above were replaced by clips that were checked first: held-out voices of the provided data
|
| 67 |
+
that pass the same prompt rules, recorded with a wide band and no background noise (measured on the prompts). Below, the model
|
| 68 |
+
(step 100k) reads the same text with the old and the new prompt (same settings and seed), with each output's band and predicted
|
| 69 |
+
naturalness. Details: [#177](https://github.com/kadirnar/dacvae-next/issues/177).
|
| 70 |
|
| 71 |
<table>
|
| 72 |
+
<tr><th>What the model reads</th><th>Old voice prompt</th><th>Output with it</th><th>New voice prompt</th><th>Output with it</th></tr>
|
| 73 |
+
<tr><td><b>Question</b><br><sub>Replaced: a dull recording (band 8.9 kHz)</sub></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/prompt_choice/old_prompt/02-question.wav"></audio><br><sub><code>spk_0342</code> · band 8.9 kHz</sub></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/prompt_choice/old_output/02-question.wav"></audio><br><sub>band 8.2 kHz · UTMOS 3.84</sub></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/prompts/02-question.wav"></audio><br><sub><code>spk_0327</code> · band 17.0 kHz</sub></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0100000/02-question.wav"></audio><br><sub>band 15.0 kHz · UTMOS 4.25</sub></td></tr>
|
| 74 |
+
<tr><td><b>Numbers, dates, abbreviations</b><br><sub>Replaced: background noise (signal-to-noise 38 dB)</sub></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/prompt_choice/old_prompt/03-numbers.wav"></audio><br><sub><code>spk_2669</code> · band 17.7 kHz</sub></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/prompt_choice/old_output/03-numbers.wav"></audio><br><sub>band 17.5 kHz · UTMOS 3.70</sub></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/prompts/03-numbers.wav"></audio><br><sub><code>spk_0409</code> · band 15.5 kHz</sub></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0100000/03-numbers.wav"></audio><br><sub>band 15.3 kHz · UTMOS 4.38</sub></td></tr>
|
| 75 |
+
<tr><td><b>Long passage (~20 s)</b><br><sub>Replaced: a dull recording (band 10.9 kHz)</sub></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/prompt_choice/old_prompt/05-long.wav"></audio><br><sub><code>spk_1437</code> · band 10.9 kHz</sub></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/prompt_choice/old_output/05-long.wav"></audio><br><sub>band 8.3 kHz · UTMOS 4.15</sub></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/prompts/05-long.wav"></audio><br><sub><code>spk_1964</code> · band 15.7 kHz</sub></td><td><audio controls src="https://huggingface.co/VoiceHub/DACFlow-EN-10k/resolve/main/samples/step_0100000/05-long.wav"></audio><br><sub>band 14.5 kHz · UTMOS 4.38</sub></td></tr>
|
|
|
|
|
|
|
| 76 |
</table>
|
| 77 |
|
| 78 |
+
Band: the highest frequency with real content in the recording (the estimator of the training data). UTMOS: predicted
|
| 79 |
+
naturalness, 1 to 5. Both measured on the take of the same text in the sweep of [#177](https://github.com/kadirnar/dacvae-next/issues/177) (same model, settings
|
| 80 |
+
and seed); the players hold this page's own takes, which can differ slightly.
|
|
|
|
| 81 |
|
| 82 |
+
## Results
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 83 |
|
| 84 |
+
### The model: step 100k
|
|
|
|
| 85 |
|
| 86 |
+
How well the model does on voices and sentences it never saw in training (**lower WER** is better, **higher SIM-o and UTMOS** are
|
| 87 |
+
better). The settings are those of [How to use](#how-to-use): each prompt's own bandwidth as the condition (auto, clamped to 9,798-16,839 Hz).
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 88 |
|
| 89 |
+
| Test set | WER ↓ | SIM-o ↑ | UTMOS ↑ |
|
| 90 |
+
|---|---:|---:|---:|
|
| 91 |
+
| **seed-dev**: real people's voices (545 sentences, two takes each) | 1.42 % | 0.526 | 3.80 (real speech: 3.52) |
|
| 92 |
+
| **echo-dev v2**: held-out voices of the training data's kind (198 voices, 1,000 sentences, one take each) | 0.53 % | 0.816 | 3.99 |
|
| 93 |
|
| 94 |
+
2.4 % of the seed-dev outputs score SIM-o below 0.2 (the voice is not copied).
|
| 95 |
+
For comparison, step 190k with the default bandwidth condition (auto was not measured there): seed-dev WER 2.06 % · SIM-o 0.223 · UTMOS 3.00, 51 % below 0.2; echo-dev v2 WER 0.46 % · SIM-o 0.820 · UTMOS 4.14.
|
|
|
|
| 96 |
|
| 97 |
+
### All checkpoints of the run
|
| 98 |
|
| 99 |
+
Every checkpoint of the training run, scored with the default bandwidth condition; the model is step 100k:
|
|
|
|
| 100 |
|
| 101 |
| Checkpoint | echo-dev WER ↓ | echo-dev SIM-o ↑ | echo-dev UTMOS ↑ (real speech: 4.20) | seed-dev WER ↓ | seed-dev SIM-o ↑ | seed-dev UTMOS ↑ (real speech: 3.52) | Samples | Weights |
|
| 102 |
|---|---:|---:|---:|---:|---:|---:|---|---|
|
| 103 |
+
| step 190k | 0.64 % | 0.790 | 3.98 | - | - | - | pending | [files](https://huggingface.co/VoiceHub/DACFlow-EN-10k/tree/main/checkpoints/step_0190000) |
|
| 104 |
+
| step 180k | 1.04 % | 0.792 | 3.98 | 1.65 % | 0.261 | 3.29 | pending | [files](https://huggingface.co/VoiceHub/DACFlow-EN-10k/tree/main/checkpoints/step_0180000) |
|
| 105 |
+
| step 170k | 1.01 % | 0.796 | 3.97 | - | - | - | pending | [files](https://huggingface.co/VoiceHub/DACFlow-EN-10k/tree/main/checkpoints/step_0170000) |
|
| 106 |
+
| step 160k | 0.81 % | 0.796 | 3.98 | 1.58 % | 0.349 | 3.66 | pending | [files](https://huggingface.co/VoiceHub/DACFlow-EN-10k/tree/main/checkpoints/step_0160000) |
|
| 107 |
+
| step 150k | 0.53 % | 0.794 | 3.93 | - | - | - | pending | [files](https://huggingface.co/VoiceHub/DACFlow-EN-10k/tree/main/checkpoints/step_0150000) |
|
| 108 |
+
| step 140k | 0.58 % | 0.793 | 3.92 | 1.60 % | 0.404 | 3.81 | pending | [files](https://huggingface.co/VoiceHub/DACFlow-EN-10k/tree/main/checkpoints/step_0140000) |
|
| 109 |
+
| step 130k | 0.76 % | 0.792 | 3.93 | - | - | - | pending | [files](https://huggingface.co/VoiceHub/DACFlow-EN-10k/tree/main/checkpoints/step_0130000) |
|
| 110 |
+
| step 120k | 0.53 % | 0.790 | 3.89 | 1.45 % | 0.451 | 3.86 | pending | [files](https://huggingface.co/VoiceHub/DACFlow-EN-10k/tree/main/checkpoints/step_0120000) |
|
| 111 |
+
| step 110k | 1.10 % | 0.789 | 3.86 | - | - | - | pending | [files](https://huggingface.co/VoiceHub/DACFlow-EN-10k/tree/main/checkpoints/step_0110000) |
|
| 112 |
+
| **step 100k** (the model) | 0.69 % | 0.787 | 3.85 | 1.54 % | 0.499 | 3.78 | [listen](#listen) | [files](https://huggingface.co/VoiceHub/DACFlow-EN-10k/tree/main/checkpoints/step_0100000) |
|
| 113 |
+
| step 90k | 0.74 % | 0.782 | 3.78 | - | - | - | pending | [files](https://huggingface.co/VoiceHub/DACFlow-EN-10k/tree/main/checkpoints/step_0090000) |
|
| 114 |
+
| step 80k | 0.78 % | 0.776 | 3.73 | 1.58 % | 0.507 | 3.68 | pending | [files](https://huggingface.co/VoiceHub/DACFlow-EN-10k/tree/main/checkpoints/step_0080000) |
|
| 115 |
+
| step 70k | 0.67 % | 0.766 | 3.69 | - | - | - | pending | [files](https://huggingface.co/VoiceHub/DACFlow-EN-10k/tree/main/checkpoints/step_0070000) |
|
| 116 |
+
| step 60k | 0.69 % | 0.759 | 3.62 | 1.61 % | 0.537 | 3.56 | pending | [files](https://huggingface.co/VoiceHub/DACFlow-EN-10k/tree/main/checkpoints/step_0060000) |
|
| 117 |
+
| step 50k | 0.74 % | 0.754 | 3.56 | - | - | - | pending | [files](https://huggingface.co/VoiceHub/DACFlow-EN-10k/tree/main/checkpoints/step_0050000) |
|
| 118 |
+
| step 40k | 0.74 % | 0.748 | 3.50 | 1.85 % | 0.536 | 3.48 | pending | [files](https://huggingface.co/VoiceHub/DACFlow-EN-10k/tree/main/checkpoints/step_0040000) |
|
| 119 |
+
| step 30k | 1.06 % | 0.741 | 3.43 | - | - | - | pending | [files](https://huggingface.co/VoiceHub/DACFlow-EN-10k/tree/main/checkpoints/step_0030000) |
|
| 120 |
+
| step 20k | 1.20 % | 0.722 | 3.33 | 2.47 % | 0.506 | 3.27 | pending | [files](https://huggingface.co/VoiceHub/DACFlow-EN-10k/tree/main/checkpoints/step_0020000) |
|
| 121 |
|
| 122 |
- **WER** (word error rate): Whisper-large-v3 writes down what it hears and this is compared with the text; 2 % means about
|
| 123 |
one word in fifty is wrong.
|
|
|
|
| 127 |
- **echo-dev**: held-out voices of the training data's kind (synthetic EchoTTS voices, none of them trained on). **seed-dev**:
|
| 128 |
real people's voices (the dev half of Seed-TTS test-en, Common Voice recordings). seed-dev is harder: the model learned
|
| 129 |
only from synthetic voices.
|
| 130 |
+
- **echo-dev v2**: a larger set of the same kind (198 held-out voices), the echo set the model was chosen on; the table of all
|
| 131 |
+
checkpoints uses the first echo-dev set.
|
| 132 |
- `-`: not measured at that step (echo-dev is scored every 10k steps, seed-dev every 20k). Every number averages two takes
|
| 133 |
+
per sentence (echo-dev v2: one take per sentence).
|
| 134 |
|
| 135 |
## How to use
|
| 136 |
|
|
|
|
| 139 |
```bash
|
| 140 |
git clone -b roadmap/en-echo https://github.com/kadirnar/dacvae-next
|
| 141 |
cd dacvae-next
|
| 142 |
+
git checkout 029dec6 # the code that made the samples above
|
| 143 |
pip install -e ".[codec]"
|
| 144 |
```
|
| 145 |
|
|
|
|
| 153 |
from mytts.flow.sampler import SamplerConfig
|
| 154 |
from mytts.infer import Synthesizer
|
| 155 |
|
| 156 |
+
ckpt = "checkpoints/step_0100000" # the model (step 100k)
|
| 157 |
local = snapshot_download("VoiceHub/DACFlow-EN-10k", allow_patterns=[f"{ckpt}/*"]) # downloads only this checkpoint
|
| 158 |
+
tts = Synthesizer.from_export(f"{local}/{ckpt}", device="cuda", bandwidth_range=(9797.6, 16839.0))
|
| 159 |
sc = SamplerConfig(steps=32, cfg_mode="joint", cfg_w=4.0, noise_scale=0.9) # the settings of the samples above
|
| 160 |
wav, sr = tts.synthesize(
|
| 161 |
"Any English text you like.",
|
|
|
|
| 163 |
sc=sc,
|
| 164 |
prompt_audio="my_voice.wav", # the voice to copy
|
| 165 |
prompt_text="The exact words spoken in my_voice.wav.",
|
| 166 |
+
bandwidth_hz="auto", # condition on the prompt's own bandwidth (clamped to the range above)
|
| 167 |
out_lufs=-16.0,
|
| 168 |
seed=0,
|
| 169 |
)
|
| 170 |
sf.write("output.wav", wav, sr) # 48 kHz, with the AudioSeal watermark
|
| 171 |
```
|
| 172 |
|
| 173 |
+
**Or from the command line**:
|
| 174 |
|
| 175 |
```bash
|
| 176 |
+
hf download VoiceHub/DACFlow-EN-10k --include "checkpoints/step_0100000/*" --local-dir DACFlow-EN-10k
|
| 177 |
+
python scripts/synthesize.py --model DACFlow-EN-10k/checkpoints/step_0100000 --bandwidth auto --bandwidth-range 9797.6 16839.0 \
|
| 178 |
--text "Any English text you like." --prompt-audio my_voice.wav \
|
| 179 |
--prompt-text "The exact words spoken in my_voice.wav." --out output.wav
|
| 180 |
```
|
| 181 |
|
| 182 |
+
`auto` measures the voice prompt's frequency band and asks the model for an output of the same band (a dull prompt
|
| 183 |
+
gives a dull output; see [How the sample voices were chosen](#how-the-sample-voices-were-chosen)).
|
| 184 |
+
The range is the 5th to 95th percentile of the training data's band; it is passed explicitly because a released
|
| 185 |
+
`config.json` records only the top of it. `auto` needs the code of [PR #180](https://github.com/kadirnar/dacvae-next/pull/180) or later (the commit above has it).
|
| 186 |
+
|
| 187 |
Tips: a clean prompt with an exact transcript works best. A long text is split at sentence ends and read piece by piece.
|
| 188 |
`python scripts/watermark_check.py detect output.wav` checks the watermark.
|
| 189 |
|
|
|
|
| 193 |
generates the latents of the [Semantic-DACVAE](https://huggingface.co/Aratako/Semantic-DACVAE-Japanese) audio codec (48 kHz,
|
| 194 |
25 frames per second) from the text, continuing the voice prompt in context (no separate speaker encoder).
|
| 195 |
- **Data**: the curated training catalog `echo_en_v1` of the provided data ([EN-10, #34](https://github.com/kadirnar/dacvae-next/issues/34) curation; see [Data](#data)).
|
| 196 |
+
- **Recipe**: planned 200k steps on one RTX 5090, about 27 minutes of speech per step (40,000 latent frames); learning rate 2.5e-4 after 5k warm-up steps, held, then lowered over the last 20 % of the steps (from step 160k).
|
| 197 |
+
Training was stopped at step 194,271 on 2026-10-04. The model is the EMA export of step 100k, before the learning rate was lowered.
|
| 198 |
The settings were chosen with small screening runs: the flow-matching noise schedule t_mean -0.8 / t_std 0.8 ([EN-55, #100](https://github.com/kadirnar/dacvae-next/issues/100))
|
| 199 |
and cross-utterance voice prompts with p_cross 0.6 ([EN-29, #53](https://github.com/kadirnar/dacvae-next/issues/53)).
|
| 200 |
- **Code and history**: the recipe [configs/train/en_full.yaml](https://github.com/kadirnar/dacvae-next/blob/roadmap/en-echo/configs/train/en_full.yaml)
|
prompts/02-question.wav
CHANGED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
-
size
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:ec5010cbb72864e62cd95a87840c6f74a750f4ac135ddcda43ce34417b63c346
|
| 3 |
+
size 503852
|
prompts/03-numbers.wav
CHANGED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
size 634924
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:4f1d94263fbe54bdcd239706933b97365dac4d34ea62f75e319da06821cdeff7
|
| 3 |
size 634924
|
prompts/05-long.wav
CHANGED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
-
size
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:9f6966023f8946ea8d9d915dbd3a8ac00c242217b131dcc979eed1063df2b9e0
|
| 3 |
+
size 708652
|
samples/{step_0030000 → prompt_choice/old_output}/02-question.wav
RENAMED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
size 675884
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:0d251f50641106b27a64844cdaae6ddf68539441021cea535a1419ff1595be34
|
| 3 |
size 675884
|
samples/{step_0030000 → prompt_choice/old_output}/03-numbers.wav
RENAMED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
size 802604
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:45490e02ea584a23db51da972998d91353bd6204fc8ebab5adb609e28f6befc1
|
| 3 |
size 802604
|
samples/{step_0020000 → prompt_choice/old_output}/05-long.wav
RENAMED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
size 1699244
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:20652f389b6b9a5b7424ee8dd42765e4aa33eaa6d0176812698226990fdba49e
|
| 3 |
size 1699244
|
samples/{step_0020000 → prompt_choice/old_prompt}/02-question.wav
RENAMED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
-
size
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:337e2af17f4096f6afa7beb9a151859157c3ee8ec3cbfeca9e63aa2a61529cab
|
| 3 |
+
size 380972
|
samples/{step_0040000 → prompt_choice/old_prompt}/03-numbers.wav
RENAMED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
-
size
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:c7949d82713134b3b4480cfe85cd599e5ddfc73e860a96de1b09bee04b9102fd
|
| 3 |
+
size 634924
|
samples/{step_0020000/01-short.wav → prompt_choice/old_prompt/05-long.wav}
RENAMED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
-
size
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:b0047e2d525a230a7d05e70ac838100b92c425687e7cc0a3c5ab0e065c04af95
|
| 3 |
+
size 622636
|
samples/step_0020000/03-numbers.wav
DELETED
|
@@ -1,3 +0,0 @@
|
|
| 1 |
-
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:c125ca86b2c4234331edff603762bff5c2560c0279daa8ed0b3864acb1ba234e
|
| 3 |
-
size 802604
|
|
|
|
|
|
|
|
|
|
|
|
samples/step_0020000/04-conversational.wav
DELETED
|
@@ -1,3 +0,0 @@
|
|
| 1 |
-
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:e4bd481b97e04d9663bac0b8378c1aabf60a4876f2aebe2d23e30327049af680
|
| 3 |
-
size 1129004
|
|
|
|
|
|
|
|
|
|
|
|
samples/step_0030000/01-short.wav
DELETED
|
@@ -1,3 +0,0 @@
|
|
| 1 |
-
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:ef8a4af218c3446f88cc66c5a1685af66d3287a1dd4023d5a286e79cf8f5bfa0
|
| 3 |
-
size 491564
|
|
|
|
|
|
|
|
|
|
|
|
samples/step_0030000/04-conversational.wav
DELETED
|
@@ -1,3 +0,0 @@
|
|
| 1 |
-
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:268121f68f855288f80d35f0dac735393ac2b7aeab9fe3c3239b1fd6e7031711
|
| 3 |
-
size 1129004
|
|
|
|
|
|
|
|
|
|
|
|
samples/step_0030000/05-long.wav
DELETED
|
@@ -1,3 +0,0 @@
|
|
| 1 |
-
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:7d7d7dec305fcd3195d9f0d77c5aea0c033080778de8ef4edbe7e63880f26e84
|
| 3 |
-
size 1699244
|
|
|
|
|
|
|
|
|
|
|
|
samples/step_0040000/01-short.wav
DELETED
|
@@ -1,3 +0,0 @@
|
|
| 1 |
-
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:49a9ad994685c295924654ab4bb913396c512f83fe8c61d611af31d41a109e2a
|
| 3 |
-
size 491564
|
|
|
|
|
|
|
|
|
|
|
|
samples/step_0040000/02-question.wav
DELETED
|
@@ -1,3 +0,0 @@
|
|
| 1 |
-
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:191d8b96c3c567b1fe15eaa95505db8e15d06cf3f554063fbeecd855b006e53e
|
| 3 |
-
size 675884
|
|
|
|
|
|
|
|
|
|
|
|
samples/step_0040000/04-conversational.wav
DELETED
|
@@ -1,3 +0,0 @@
|
|
| 1 |
-
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:7ec715e56ca0035b5a992d05e424eee61738b3f4565d84e05cde53e3b94f1070
|
| 3 |
-
size 1129004
|
|
|
|
|
|
|
|
|
|
|
|
samples/step_0040000/05-long.wav
DELETED
|
@@ -1,3 +0,0 @@
|
|
| 1 |
-
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:5525bb18667843cb8fec70df78f4f5478b79faef1a753e7ec7a02a11122a6691
|
| 3 |
-
size 1699244
|
|
|
|
|
|
|
|
|
|
|
|
samples/step_0050000/01-short.wav
DELETED
|
@@ -1,3 +0,0 @@
|
|
| 1 |
-
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:90977dc5598fed3c02d897870e0661c045b25375ea5594e11991564d0c6112e2
|
| 3 |
-
size 491564
|
|
|
|
|
|
|
|
|
|
|
|
samples/step_0050000/02-question.wav
DELETED
|
@@ -1,3 +0,0 @@
|
|
| 1 |
-
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:c11a29b4306ff85b6a9f1834389201a3a8eaa6dc4084dcac4a2130bb28c89eb9
|
| 3 |
-
size 675884
|
|
|
|
|
|
|
|
|
|
|
|
samples/step_0050000/03-numbers.wav
DELETED
|
@@ -1,3 +0,0 @@
|
|
| 1 |
-
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:cd6993d65d2ef0ee2e695bed119d56fb0a3ec261dde9b126244814d0a0cb49c1
|
| 3 |
-
size 802604
|
|
|
|
|
|
|
|
|
|
|
|
samples/step_0050000/04-conversational.wav
DELETED
|
@@ -1,3 +0,0 @@
|
|
| 1 |
-
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:9f8c8b9578823d51830602bb7f3dafa2ac0639a4c647d43871d5ef7b6ed8c433
|
| 3 |
-
size 1129004
|
|
|
|
|
|
|
|
|
|
|
|
samples/step_0050000/05-long.wav
DELETED
|
@@ -1,3 +0,0 @@
|
|
| 1 |
-
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:3253fea011132cb064196f174d8856a924ecae7f3dedce0c75ab91192f0f7de5
|
| 3 |
-
size 1699244
|
|
|
|
|
|
|
|
|
|
|
|
samples/step_0060000/01-short.wav
DELETED
|
@@ -1,3 +0,0 @@
|
|
| 1 |
-
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:aa41e9039c23632e6e23dc63defe11549bf664c606ad8dd02a17de95f44a5954
|
| 3 |
-
size 491564
|
|
|
|
|
|
|
|
|
|
|
|
samples/step_0060000/02-question.wav
DELETED
|
@@ -1,3 +0,0 @@
|
|
| 1 |
-
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:0302c8685d9e47b71c2fc891e7ae18475ce287747186d1f061a82c6b626d4ad6
|
| 3 |
-
size 675884
|
|
|
|
|
|
|
|
|
|
|
|
samples/step_0060000/03-numbers.wav
DELETED
|
@@ -1,3 +0,0 @@
|
|
| 1 |
-
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:21cd0b3472609ece712c22c0c12d72134387b17ac90ef5d9084f4b8725fd25e8
|
| 3 |
-
size 802604
|
|
|
|
|
|
|
|
|
|
|
|
samples/step_0060000/04-conversational.wav
DELETED
|
@@ -1,3 +0,0 @@
|
|
| 1 |
-
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:1c935be979bca4df237b0dc95c0f2188150c5e3b04324c4e702a02dbb1977d6b
|
| 3 |
-
size 1129004
|
|
|
|
|
|
|
|
|
|
|
|
samples/step_0060000/05-long.wav
DELETED
|
@@ -1,3 +0,0 @@
|
|
| 1 |
-
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:4e2199443c386340cefcb37edc19c71a9ee2e5c10ffc97b4d5b6d0d994511165
|
| 3 |
-
size 1699244
|
|
|
|
|
|
|
|
|
|
|
|
samples/step_0070000/01-short.wav
DELETED
|
@@ -1,3 +0,0 @@
|
|
| 1 |
-
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:abec54dbe34fda663904f8c0266a7cb02519b557a6b64cd106ee7ad0f95916b1
|
| 3 |
-
size 491564
|
|
|
|
|
|
|
|
|
|
|
|
samples/step_0070000/02-question.wav
DELETED
|
@@ -1,3 +0,0 @@
|
|
| 1 |
-
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:55da8e9b91b52837b0e508a0aef7e7de32ed47add1cf5faf6e6c63a1434932da
|
| 3 |
-
size 675884
|
|
|
|
|
|
|
|
|
|
|
|
samples/step_0070000/03-numbers.wav
DELETED
|
@@ -1,3 +0,0 @@
|
|
| 1 |
-
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:4c68b434f4d130f86610bfb7ed090a91595a7fe16072f7c6703ba7c032efd346
|
| 3 |
-
size 802604
|
|
|
|
|
|
|
|
|
|
|
|
samples/step_0070000/04-conversational.wav
DELETED
|
@@ -1,3 +0,0 @@
|
|
| 1 |
-
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:6437479879b21c243c6360a6270f7eb662d593315c8ebe79e07ad31190889435
|
| 3 |
-
size 1129004
|
|
|
|
|
|
|
|
|
|
|
|
samples/step_0070000/05-long.wav
DELETED
|
@@ -1,3 +0,0 @@
|
|
| 1 |
-
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:be36c40ab725b3638cfdd159f68c283e0a38e19636335255ee5e3a8c1e819693
|
| 3 |
-
size 1699244
|
|
|
|
|
|
|
|
|
|
|
|
samples/step_0080000/01-short.wav
DELETED
|
@@ -1,3 +0,0 @@
|
|
| 1 |
-
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:220a06fd91bc5b539bfab4cc4cf6d7fa8bb207b1bd4a8136b0e9fa0b820ef99e
|
| 3 |
-
size 491564
|
|
|
|
|
|
|
|
|
|
|
|
samples/step_0080000/02-question.wav
DELETED
|
@@ -1,3 +0,0 @@
|
|
| 1 |
-
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:fdf40c4150f1063d6b141721f574208fcb3b4b4a218044880e31cd1b9801b666
|
| 3 |
-
size 675884
|
|
|
|
|
|
|
|
|
|
|
|
samples/step_0080000/03-numbers.wav
DELETED
|
@@ -1,3 +0,0 @@
|
|
| 1 |
-
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:f508b564c626e54bf17a85b9d1cde757a7865b63bb2c6438006cbf314422d9ce
|
| 3 |
-
size 802604
|
|
|
|
|
|
|
|
|
|
|
|
samples/step_0080000/04-conversational.wav
DELETED
|
@@ -1,3 +0,0 @@
|
|
| 1 |
-
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:297a1d51596503186a06986a690e9b899e360df647bcfd970b839a4636fcfcf4
|
| 3 |
-
size 1129004
|
|
|
|
|
|
|
|
|
|
|
|
samples/step_0080000/05-long.wav
DELETED
|
@@ -1,3 +0,0 @@
|
|
| 1 |
-
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:7204173c814b5924fdd90afa379c8a42e31b3c54b4068f9a0063ea3393ea7539
|
| 3 |
-
size 1699244
|
|
|
|
|
|
|
|
|
|
|
|
samples/step_0090000/01-short.wav
DELETED
|
@@ -1,3 +0,0 @@
|
|
| 1 |
-
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:912169451f2095e83992c48b83187a6e6eaf20224fe36d55866ae4b339eaec16
|
| 3 |
-
size 491564
|
|
|
|
|
|
|
|
|
|
|
|
samples/step_0090000/02-question.wav
DELETED
|
@@ -1,3 +0,0 @@
|
|
| 1 |
-
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:331f3aa5925718dfd401626289b343560d11f02f84fb231c9cc884b783308162
|
| 3 |
-
size 675884
|
|
|
|
|
|
|
|
|
|
|
|
samples/step_0090000/03-numbers.wav
DELETED
|
@@ -1,3 +0,0 @@
|
|
| 1 |
-
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:eec195dfc8db1a8fa38e2c5e2e46f971aafdfe66eed096dce2f12d1bd7e7f87e
|
| 3 |
-
size 802604
|
|
|
|
|
|
|
|
|
|
|
|
samples/step_0090000/04-conversational.wav
DELETED
|
@@ -1,3 +0,0 @@
|
|
| 1 |
-
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:9fad8387f13bb3e23258cb435f0aa0f08d8303929b683a559c83b88915fef1cb
|
| 3 |
-
size 1129004
|
|
|
|
|
|
|
|
|
|
|
|
samples/step_0090000/05-long.wav
DELETED
|
@@ -1,3 +0,0 @@
|
|
| 1 |
-
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:b372727c0000be7091c5c9d29ed10ca65c116edf98b81fc4099df47ca103d9ff
|
| 3 |
-
size 1699244
|
|
|
|
|
|
|
|
|
|
|
|
samples/step_0100000/01-short.wav
CHANGED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
size 491564
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:be50f96dc34d32d24739068756b6f52e88e3b0234396ee501502ba6732c548f6
|
| 3 |
size 491564
|
samples/step_0100000/02-question.wav
CHANGED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
-
size
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:1a64e3bbded645444bab1abce831686febe10ad2efc5a45d72227d2d2821fcb8
|
| 3 |
+
size 583724
|
samples/step_0100000/03-numbers.wav
CHANGED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
-
size
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:81225f9e0b8b42522ba01e69e5a7207495916da59b101e694a437734ef9a3440
|
| 3 |
+
size 841004
|
samples/step_0100000/04-conversational.wav
CHANGED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
size 1129004
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:9cdf15771b9bc38db03573d062ef083d2c6164f8189977ed34cca7c93e365c0a
|
| 3 |
size 1129004
|
samples/step_0100000/05-long.wav
CHANGED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
-
size
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:252c18556b414370fdd7330dc92e07094b33533494e4cc87a3a6bf1a111d8524
|
| 3 |
+
size 1691564
|