File size: 12,488 Bytes
9528ca8
0478022
9528ca8
 
 
 
 
 
0478022
9598ed6
 
 
0478022
9528ca8
0478022
9f6e0ef
 
 
 
0478022
9528ca8
0478022
9528ca8
0478022
9528ca8
0478022
9598ed6
 
 
 
 
 
 
0478022
9528ca8
 
 
0478022
9528ca8
0478022
9528ca8
 
0478022
9528ca8
 
 
 
 
 
 
 
 
 
0478022
9528ca8
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
9f6e0ef
9528ca8
9f6e0ef
 
 
 
 
9528ca8
9f6e0ef
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
9528ca8
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
# Inflect v2 evaluation and benchmark report

This page is the detailed evidence record for
[Inflect-Micro-v2](https://huggingface.co/owensong/Inflect-Micro-v2) and
[Inflect-Nano-v2](https://huggingface.co/owensong/Inflect-Nano-v2). The model
cards show the headline results; this document defines the datasets, systems,
normalization, uncertainty, exclusions, runtime boundaries, and raw artifacts
behind those results.

Read each metric separately: WER measures intelligibility, UTMOS22 predicts
naturalness, listening tests record preference, and runtime measures deployment
cost. None of them is an overall quality score.

## Evidence at a glance

| Release | Complete parameters | FP32 weights | UTMOS22 ↑ | Two-ASR semantic WER ↓ | Community preference ↑ | 4-thread CPU ↑ |
| --- | ---: | ---: | ---: | ---: | ---: | ---: |
| **Inflect-Micro-v2** | **9,356,513** | **37.53 MB** | **4.395** | **3.99%** | **66.2%** | **6.28×** |
| **Inflect-Nano-v2** | **3,966,721** | **15.97 MB** | **4.386** | **4.21%** | **63.9%** | **10.72×** |

![Community blind listening](../assets/evidence/human-preference.svg)

![Predicted quality versus footprint](../assets/evidence/quality-vs-footprint.svg)

![Two-ASR semantic WER](../assets/evidence/asr-consensus.svg)

## Compared systems

The benchmark includes KittenTTS Nano, Piper Low, and Supertonic 3 because they
are established compact or local TTS systems. In this package-level comparison,
each has a larger deployable weight footprint than both Inflect releases. That
provides useful size context, but no single result establishes overall
superiority.

Voice variants are separate rows when voice identity can affect quality or
intelligibility. Kitten and Piper may be shown as equal-weight two-voice means
in family-level figures. Supertonic 3-step and 8-step are never pooled.

## Complete automated result matrix

All systems received matched text within each benchmark. Lower WER is better;
higher UTMOS22 is better.

| System / voice | Complete weights ↓ | UTMOS22 ↑ | Qwen3-ASR WER ↓ | Nemotron 3.5 WER ↓ | Whisper large-v3 WER ↓ | Two-ASR consensus ↓ |
| --- | ---: | ---: | ---: | ---: | ---: | ---: |
| Inflect-Micro-v2 | 37.53 MB | **4.395** | 2.52% | 5.45% | 2.73% | **3.99%** |
| Inflect-Nano-v2 | 15.97 MB | **4.386** | 2.78% | 5.63% | 2.65% | **4.21%** |
| KittenTTS Nano · Bruno | 56.77 MB | **4.217** | 2.15% | 3.96% | 2.17% | **3.06%** |
| KittenTTS Nano · Hugo | 56.77 MB | **4.191** | 2.39% | 3.80% | 2.11% | **3.09%** |
| Piper Low · Danny | 63.10 MB | **4.195** | 2.62% | 5.60% | 2.55% | **4.11%** |
| Piper Low · Ryan | 63.10 MB | **4.289** | 2.81% | 5.51% | 2.87% | **4.16%** |
| Supertonic 3 · M2 · 8-step | 398.08 MB | **4.295** | 2.05% | 3.56% | 8.08% | **2.81%** |
| Supertonic 3 · M2 · 3-step | 398.08 MB | **2.471** | 3.03% | 6.04% | 3.22% | **4.54%** |

The two-ASR consensus is the unweighted mean of Qwen3-ASR and Nemotron 3.5
corpus-level semantic WER. Whisper remains visible but is excluded from the
headline consensus for every system because its Supertonic 8-step result
contains insertion-heavy outliers that conflict with the other recognizers and
human listening. This rule is applied globally, not selectively to Inflect.

![Three-ASR intelligibility audit](../assets/evidence/modern400-three-asr.svg)

## Uncertainty for the Inflect releases

Intervals use 10,000 matched-prompt bootstrap samples.

| Release | UTMOS22 mean (95% CI) | Qwen3-ASR WER (95% CI) | Nemotron WER (95% CI) | Whisper WER (95% CI) |
| --- | ---: | ---: | ---: | ---: |
| Inflect-Micro-v2 | 4.395 (4.381–4.408) | 2.52% (1.88%–3.24%) | 5.45% (4.50%–6.50%) | 2.73% (2.05%–3.49%) |
| Inflect-Nano-v2 | 4.386 (4.372–4.399) | 2.78% (2.13%–3.52%) | 5.63% (4.72%–6.62%) | 2.65% (2.03%–3.35%) |

## Modern400 intelligibility protocol

- **400 identical unseen English prompts per system**
- **3,200 generated clips** across eight system/voice configurations
- 200 manually constructed modern and stress prompts
- 200 deterministic Google FLEURS `en_us` test prompts
- exact normalized-text exclusion against 87,362 training transcripts
- deterministic frontend and inference settings within each system
- explicit resampling to every recognizer's required sample rate
- corpus-level WER plus matched-prompt bootstrap intervals

The three independent recognizers are:

| Evaluator | Model |
| --- | --- |
| Qwen3-ASR | `Qwen/Qwen3-ASR-1.7B` |
| Nemotron 3.5 ASR | `nvidia/nemotron-3.5-asr-streaming-0.6b` |
| Whisper | `openai/whisper-large-v3` |

### Semantic normalization

Raw orthographic WER is not sufficient for TTS. A recognizer may transcribe
“twelve point six” as `12.6` or “July twenty-first” as `July 21st` without the
audio being wrong. The semantic normalizer canonicalizes equivalent number,
currency, date, time, unit, abbreviation, and punctuation forms before scoring.
The release retains both raw and semantic hypotheses so normalization cannot
silently hide errors.

### Where Inflect errors remain

The following diagnostic uses the same two-ASR consensus as the headline and
does not introduce a new score.

| Prompt category | Micro WER ↓ | Nano WER ↓ |
| --- | ---: | ---: |
| Names Places | 16.67% | 17.54% |
| Short | 4.29% | 9.05% |
| Rare Words | 7.89% | 7.89% |
| Numbers | 6.25% | 4.69% |
| Homographs | 3.57% | 5.71% |
| External Heldout | 5.39% | 5.48% |
| Technical | 4.10% | 3.12% |
| Narrative | 0.57% | 1.56% |
| Punctuation | 0.00% | 1.54% |
| Long Range | 0.83% | 1.24% |
| Conversational | 0.49% | 1.13% |
| Descriptive | 0.44% | 0.88% |
| Long | 0.81% | 0.58% |
| Emotional | 0.15% | 0.30% |
| Question | 0.00% | 0.00% |

![Prompt-category semantic WER](../assets/evidence/category-semantic-wer.svg)

## UTMOS22 predicted-naturalness protocol

- **500 identical unseen prompts per system**
- `tarepan/SpeechMOS` v1.2.0
- one score per generated clip
- mean, median, standard deviation, percentiles, extrema, and 95% bootstrap CI
- Kitten and Piper family figures use equal-weight means across their two voices

UTMOS22 is a learned non-intrusive quality predictor. It is useful for
large-scale screening but is **not human MOS**, can be sensitive to voice and
recording characteristics, and must be interpreted beside intelligibility and
listening evidence.

## Community blind listening

The final community study stored anonymous pairwise choices rather than user
identity. Model names were hidden during voting, left/right order was
randomized, and a tie counts as half a win. The dataset contains **147 pairwise

decisions**; it does not claim a laboratory-controlled listener panel.

| System / voice | Wins | Losses | Ties | Appearances | Preference rate ↑ |
| --- | ---: | ---: | ---: | ---: | ---: |
| KittenTTS Nano · Hugo | 22 | 8 | 3 | 33 | **71.2%** |
| Inflect-Micro-v2 | 21 | 10 | 3 | 34 | **66.2%** |
| Inflect-Nano-v2 | 22 | 12 | 2 | 36 | **63.9%** |
| Supertonic 3 · James · 8-step | 16 | 12 | 3 | 31 | **56.5%** |
| KittenTTS Nano · Bruno | 16 | 13 | 7 | 36 | **54.2%** |
| Piper Low · Danny | 10 | 8 | 6 | 24 | **54.2%** |
| Piper Low · Ryan | 14 | 13 | 6 | 33 | **51.5%** |
| Inflect-Nano-v1 | 7 | 22 | 3 | 32 | **26.6%** |
| Supertonic 3 · James · 3-step | 1 | 29 | 2 | 32 | **6.2%** |

Preference rate is `(wins + 0.5 × ties) / appearances`. Unequal appearances
mean this table is descriptive; it should not be treated as a perfectly
balanced tournament or formal MOS.

## Runtime and throughput

### Managed CPU reference

The published CPU reference uses a Hugging Face **CPU Upgrade** instance
(8 vCPU, 32 GB RAM) with four framework threads and end-to-end
text-to-waveform timing. Inflect received 100 fixed Modern400 prompts across
three complete passes. The first pass populated runtime caches; the published
steady-state result pools passes two and three.

| Release | Steady-state RTF ↓ | Audio / wall time ↑ |
| --- | ---: | ---: |
| **Inflect-Micro-v2** | **0.1593** | **6.28×** |
| **Inflect-Nano-v2** | **0.0933** | **10.72×** |

The earlier Ryzen 9 3900X CPU comparison is superseded because background host
saturation and inconsistent framework thread policies made it unsuitable as
release evidence.

### Directional compact-system context

The same managed CPU and four-thread policy were used for a shorter comparator
confirmation: the identical 50-prompt prefix repeated twice. KittenTTS and
Piper are equal-work pooled across two voices.

| System | Audio / wall time ↑ |
| --- | ---: |
| Piper Low | 31.37× |
| KittenTTS Nano | 13.33× |
| **Inflect-Nano-v2** | **10.72×** |
| Supertonic 3 · 3-step | 10.15× |
| **Inflect-Micro-v2** | **6.28×** |
| Supertonic 3 · 8-step | 4.37× |

This is deployment context, not a perfectly matched speed leaderboard:
Inflect's figures come from the larger steady-state run, comparator rows use the
shorter confirmation pass, and several comparator packages use optimized ONNX
while the public Inflect package uses PyTorch.

The public runtime also collapses training-time weight normalization once,
immediately after loading the checkpoint. A fixed-seed regression test produced
bit-for-bit identical waveforms before and after this reparameterization. The
downloaded checkpoints retain **9,356,513** and **3,966,721** complete
parameters; the optimized live inference graphs contain **9,344,753** and
**3,958,801** parameters respectively. This is a deployment optimization, not
a smaller checkpoint or a changed acoustic model.

### Retained CUDA engineering snapshot

The package retains a 50-prompt CUDA snapshot for the two Inflect releases. It
is useful for regression checking but is not a cross-system leaderboard because
the archived report does not identify the GPU model.

| Release | Prompts | RTF ↓ | Audio / wall time ↑ | Median latency | p95 latency |
| --- | ---: | ---: | ---: | ---: | ---: |
| Inflect-Micro-v2 | 50 | 0.0174 | 57.49× | 78.6 ms | 82.7 ms |
| Inflect-Nano-v2 | 50 | 0.0171 | 58.41× | 80.2 ms | 85.3 ms |

## Complete model footprint

![Complete model weight footprint](../assets/evidence/model-footprint.svg)

“Complete” means the downloadable weight file required for text-to-waveform
generation, including Inflect's integrated waveform decoder. Voice variants
that share identical weights are counted once in family-level footprint
figures. Package size is not the same as peak runtime memory.

## Frozen artifacts and reproducibility

| Artifact | Purpose |
| --- | --- |
| [`evaluation/final/modern400/prompts.jsonl`](../evaluation/final/modern400/prompts.jsonl) | Frozen Modern400 prompt manifest |
| [`evaluation/final/modern400/summary.json`](../evaluation/final/modern400/summary.json) | Three-ASR summaries, intervals, and category results |
| [`evaluation/final/modern400/exclusion_report.json`](../evaluation/final/modern400/exclusion_report.json) | Exact-text training exclusion audit |
| [`evaluation/final/modern400/`](../evaluation/final/modern400/) | Compressed row-level hypotheses and scoring reports |
| [`evaluation/final/utmos22/report.json`](../evaluation/final/utmos22/report.json) | UTMOS22 clip scores and distributions |
| [`evaluation/final/runtime_snapshot/`](../evaluation/final/runtime_snapshot/) | Retained engineering snapshots |
| [`release_manifest.json`](../release_manifest.json) | File sizes and SHA-256 integrity hashes |

- Modern400 prompt SHA-256: `b7504ce2dce44a2da82770a6a5dfd2a034fe17e2113980f8a69663ade417a34c`
- UTMOS22 prompt SHA-256: `5b7aa33852e6c738eb0e1dbf9fcfcc90772b97ad219167e44aaf3917d063b056`
- Modern400 bootstrap samples: `10,000`
- UTMOS22 bootstrap samples: `10,000`

## Interpretation limits

- ASR WER measures recognizer agreement with intended text, not naturalness.
- UTMOS22 predicts perceived quality; it is not a substitute for human MOS.
- Community preference is valuable human evidence but is not a controlled lab study.
- One fixed voice per release limits conclusions about speaker diversity.
- Runtime depends on hardware, framework, provider, compilation, thread policy, and host load.
- Results apply to these frozen checkpoints, frontends, prompts, and inference settings.
- No benchmark result should be presented as proof that one model is best for every deployment.