Drop datasets: field — ZeroBench-TTS is eval data, not training data

#2
Files changed (1) hide show
  1. README.md +9 -4
README.md CHANGED
@@ -15,8 +15,12 @@ tags:
15
  - voice-cloning
16
  - vietnamese-tts
17
  - tieng-viet
18
- datasets:
19
- - zeroweight-ai/ZeroBench-TTS
 
 
 
 
20
  metrics:
21
  - wer
22
  model-index:
@@ -133,12 +137,13 @@ for why both matter.
133
 
134
  | | **ZeroTTS** | XTTS-v2-vietnamse | viXTTS |
135
  |---|:-:|:-:|:-:|
136
- | **WER** ↓ | **1.03%** | 16.42% | 18.40% |
 
137
  | **Naturalness** (UTMOS) ↑ | **2.91** | 2.43 | 2.35 |
138
  | **Voice similarity** (SSIM) ↑ | 0.936 | **0.940** | 0.935 |
139
  | **Dead air** (excess silence) ↓ | **0.029 s** | 0.532 s | 0.233 s |
140
 
141
- **16× fewer word errors**, ~0.5 MOS more natural, an order of magnitude less
142
  dead air. Median WER is **0.00%** on all four subsets.
143
 
144
  WER by subset:
 
15
  - voice-cloning
16
  - vietnamese-tts
17
  - tieng-viet
18
+ # NOTE deliberately NO `datasets:` field. It is the only thing that populates
19
+ # the Hub's cross-link, but the Hub renders it as "Models trained or fine-tuned
20
+ # on <dataset>" — which for our own held-out benchmark reads as train/test
21
+ # contamination. ZeroBench-TTS is EVALUATION data; every voice in it is held
22
+ # out of training. The `model-index` block below states that correctly, and the
23
+ # body links the benchmark in prose.
24
  metrics:
25
  - wer
26
  model-index:
 
137
 
138
  | | **ZeroTTS** | XTTS-v2-vietnamse | viXTTS |
139
  |---|:-:|:-:|:-:|
140
+ | **WER — raw text** ↓ | **1.03%** | 16.42% | 18.40% |
141
+ | **WER — pre-normalized text** ↓ | **0.56%** | 7.27% | 8.61% |
142
  | **Naturalness** (UTMOS) ↑ | **2.91** | 2.43 | 2.35 |
143
  | **Voice similarity** (SSIM) ↑ | 0.936 | **0.940** | 0.935 |
144
  | **Dead air** (excess silence) ↓ | **0.029 s** | 0.532 s | 0.233 s |
145
 
146
+ **16× fewer word errors** on the real task, **13× fewer** even after handing every model a perfect text frontend; ~0.5 MOS more natural, an order of magnitude less
147
  dead air. Median WER is **0.00%** on all four subsets.
148
 
149
  WER by subset: