diff --git a/.gitattributes b/.gitattributes index f126de2a3ac3dd1266ab8f947c51988ee5671144..0fac566fb3b788a2d95442d0413c869ddb6f3ccf 100644 --- a/.gitattributes +++ b/.gitattributes @@ -128,3 +128,9 @@ samples/step_0190000/02-question.wav filter=lfs diff=lfs merge=lfs -text samples/step_0190000/03-numbers.wav filter=lfs diff=lfs merge=lfs -text samples/step_0190000/04-conversational.wav filter=lfs diff=lfs merge=lfs -text samples/step_0190000/05-long.wav filter=lfs diff=lfs merge=lfs -text +samples/prompt_choice/old_output/02-question.wav filter=lfs diff=lfs merge=lfs -text +samples/prompt_choice/old_output/03-numbers.wav filter=lfs diff=lfs merge=lfs -text +samples/prompt_choice/old_output/05-long.wav filter=lfs diff=lfs merge=lfs -text +samples/prompt_choice/old_prompt/02-question.wav filter=lfs diff=lfs merge=lfs -text +samples/prompt_choice/old_prompt/03-numbers.wav filter=lfs diff=lfs merge=lfs -text +samples/prompt_choice/old_prompt/05-long.wav filter=lfs diff=lfs merge=lfs -text diff --git a/README.md b/README.md index fb0e448af671db324008199de787d82c8de7af45..46d8bc31915cadc4b98a7eedd4a3e04ba9fc5f0f 100644 --- a/README.md +++ b/README.md @@ -31,295 +31,93 @@ The name: **DAC** for the Semantic-DACVAE audio codec whose latents it generates English, **10k** for its training data, about ten thousand hours of speech. The code: [kadirnar/dacvae-next](https://github.com/kadirnar/dacvae-next/tree/roadmap/en-echo), Python package `mytts`. -> **Status: training, step 190k of 200k.** A new checkpoint (a copy of the model saved at that point of its training) -> appears here every 10k steps (from step 20k on), with its samples and scores: this page updates itself. Early checkpoints sound rougher than -> later ones. +> **Status: training is finished. The model is step 100k** (the EMA weights saved at that step). The run was planned for 200k steps +> and stopped at step 194,271 on 2026-10-04; this page no longer changes by itself. Why step 100k: later checkpoints sound a little +> better on held-out voices of the training data's kind, but copy real people's voices much worse. Speaker similarity on real voices +> (seed-dev SIM-o) falls from 0.50 at step 100k to 0.22 at step 190k, where 51 % of the outputs score below 0.2 (5.7 % at step +> 100k), with the same settings. Step 100k, with each prompt's own bandwidth as the condition (auto), is the best balance. Details: +> [#177](https://github.com/kadirnar/dacvae-next/issues/177). **On this page:** [Listen](#listen) · [Results](#results) · [How to use](#how-to-use) · [Training](#training) · [Data](#data) · [Licence](#licence) ## Listen -The newest checkpoint, **step 190k**, reads five texts. Each text is spoken in a different voice, copied from -the short voice prompt next to it. The model never heard these voices in training (held-out voices of the provided data). +The model, **step 100k**, reads five texts. Each text is spoken in a different voice, copied from the short voice prompt next +to it. The model never heard these voices in training (held-out voices of the provided data).
| What the model reads | Voice prompt (the input) | Model output, step 190k |
|---|---|---|
| Short sentence I left my umbrella at the office again, so I'm definitely getting soaked on the way home. | A higher voice (~202 Hz). It says: “I keep telling myself I just need to power through, like, push a little harder, and then it'll be fine. But it's not getting fine. It's getting— it's getting worse.” | |
| Question Have you ever noticed that the quietest person in the room usually has the most interesting story to tell? | A lower voice (~149 Hz). It says: “Wait, that mole— has it always looked like that? No, stop.” | |
| Numbers, dates, abbreviations Dr. Patel moved my appointment to Tuesday, March 3rd, at 4:15 p.m., and the co-pay went up from $20 to $35. | A higher voice (~181 Hz). It says: “this woman on the train was literally eating a whole bag of chips, like, crunching so loud, and I'm sitting there like, okay, do I move? whatever.” | |
| Conversation (~10 s) So I finally tried that new ramen place downtown, and honestly, it was worth the wait. The broth was amazing, but next time I'm definitely skipping the extra spicy option. | A lower voice (~115 Hz). It says: “I'm sorry, but the subway doors closed on my bag. Again.” | |
| Long passage (~20 s) When the storm passed, the whole neighborhood came outside to look at the damage, and although a few fences had fallen and the old oak tree had lost its biggest branch, everyone was relieved that nobody was hurt. They spent the rest of the afternoon clearing the street and sharing whatever food they had left. | The deepest voice (~108 Hz). It says: “Honestly I think we need to just call IT and have them... you know, actually fix it this time. I'm too old for this.” | |
| What the model reads | Voice prompt (the input) | Model output (step 100k) |
| Short sentence I left my umbrella at the office again, so I'm definitely getting soaked on the way home. | A higher voice (~202 Hz). It says: “I keep telling myself I just need to power through, like, push a little harder, and then it'll be fine. But it's not getting fine. It's getting— it's getting worse.” | |
| Question Have you ever noticed that the quietest person in the room usually has the most interesting story to tell? | A lower voice (~119 Hz). It says: “Look, I'm not saying it's easy, but you always pull it together. Just take it step by step, you know?” | |
| Numbers, dates, abbreviations Dr. Patel moved my appointment to Tuesday, March 3rd, at 4:15 p.m., and the co-pay went up from $20 to $35. | A higher voice (~201 Hz). It says: “Would you just— I know you mean well, but every time you mention it I feel like an idiot. It's probably nothing.” | |
| Conversation (~10 s) So I finally tried that new ramen place downtown, and honestly, it was worth the wait. The broth was amazing, but next time I'm definitely skipping the extra spicy option. | A lower voice (~115 Hz). It says: “I'm sorry, but the subway doors closed on my bag. Again.” | |
| Long passage (~20 s) When the storm passed, the whole neighborhood came outside to look at the damage, and although a few fences had fallen and the old oak tree had lost its biggest branch, everyone was relieved that nobody was hurt. They spent the rest of the afternoon clearing the street and sharing whatever food they had left. | The deepest voice (~102 Hz). It says: “I mean, come on, it's not like I woke up and decided to have the worst day ever. Things just went wrong one after another, like dominoes!” |
| What the model reads | Model output, step 180k |
|---|---|
| Short sentence | |
| Question | |
| Numbers, dates, abbreviations | |
| Conversation (~10 s) | |
| Long passage (~20 s) |
| What the model reads | Model output, step 170k |
|---|---|
| Short sentence | |
| Question | |
| Numbers, dates, abbreviations | |
| Conversation (~10 s) | |
| Long passage (~20 s) |
| What the model reads | Model output, step 160k |
|---|---|
| Short sentence | |
| Question | |
| Numbers, dates, abbreviations | |
| Conversation (~10 s) | |
| Long passage (~20 s) |
| What the model reads | Model output, step 150k |
|---|---|
| Short sentence | |
| Question | |
| Numbers, dates, abbreviations | |
| Conversation (~10 s) | |
| Long passage (~20 s) |
| What the model reads | Model output, step 140k |
|---|---|
| Short sentence | |
| Question | |
| Numbers, dates, abbreviations | |
| Conversation (~10 s) | |
| Long passage (~20 s) |
| What the model reads | Model output, step 130k |
|---|---|
| Short sentence | |
| Question | |
| Numbers, dates, abbreviations | |
| Conversation (~10 s) | |
| Long passage (~20 s) |
| What the model reads | Model output, step 120k |
|---|---|
| Short sentence | |
| Question | |
| Numbers, dates, abbreviations | |
| Conversation (~10 s) | |
| Long passage (~20 s) |
| What the model reads | Model output, step 110k |
|---|---|
| Short sentence | |
| Question | |
| Numbers, dates, abbreviations | |
| Conversation (~10 s) | |
| Long passage (~20 s) |
| What the model reads | Model output, step 100k |
|---|---|
| Short sentence | |
| Question | |
| Numbers, dates, abbreviations | |
| Conversation (~10 s) | |
| Long passage (~20 s) |
| What the model reads | Model output, step 90k |
|---|---|
| Short sentence | |
| Question | |
| Numbers, dates, abbreviations | |
| Conversation (~10 s) | |
| Long passage (~20 s) |
| What the model reads | Model output, step 80k |
|---|---|
| Short sentence | |
| Question | |
| Numbers, dates, abbreviations | |
| Conversation (~10 s) | |
| Long passage (~20 s) |
| What the model reads | Model output, step 70k |
|---|---|
| Short sentence | |
| Question | |
| Numbers, dates, abbreviations | |
| Conversation (~10 s) | |
| Long passage (~20 s) |
| What the model reads | Model output, step 60k |
|---|---|
| Short sentence | |
| Question | |
| Numbers, dates, abbreviations | |
| Conversation (~10 s) | |
| Long passage (~20 s) |
| What the model reads | Model output, step 50k |
|---|---|
| Short sentence | |
| Question | |
| Numbers, dates, abbreviations | |
| Conversation (~10 s) | |
| Long passage (~20 s) |
| What the model reads | Model output, step 40k | |||
|---|---|---|---|---|
| Short sentence | ||||
| Question | ||||
| Numbers, dates, abbreviations | ||||
| Conversation (~10 s) | ||||
| Long passage (~20 s) | ||||
| What the model reads | Old voice prompt | Output with it | New voice prompt | Output with it |
| Question Replaced: a dull recording (band 8.9 kHz) | spk_0342 · band 8.9 kHz | band 8.2 kHz · UTMOS 3.84 | spk_0327 · band 17.0 kHz | band 15.0 kHz · UTMOS 4.25 |
| Numbers, dates, abbreviations Replaced: background noise (signal-to-noise 38 dB) | spk_2669 · band 17.7 kHz | band 17.5 kHz · UTMOS 3.70 | spk_0409 · band 15.5 kHz | band 15.3 kHz · UTMOS 4.38 |
| Long passage (~20 s) Replaced: a dull recording (band 10.9 kHz) | spk_1437 · band 10.9 kHz | band 8.3 kHz · UTMOS 4.15 | spk_1964 · band 15.7 kHz | band 14.5 kHz · UTMOS 4.38 |
| What the model reads | Model output, step 30k |
|---|---|
| Short sentence | |
| Question | |
| Numbers, dates, abbreviations | |
| Conversation (~10 s) | |
| Long passage (~20 s) |
| What the model reads | Model output, step 20k |
|---|---|
| Short sentence | |
| Question | |
| Numbers, dates, abbreviations | |
| Conversation (~10 s) | |
| Long passage (~20 s) |