File size: 4,527 Bytes
c64650d | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 | # Official conversion validation
This report records the July 21, 2026 comparison between GGUF files published
by `cstr` and files generated directly from pinned
`hexgrad/Kokoro-82M` revision
`f3ff3571791e39611d31c381e3a41a3af07b4987`.
## Sizes
| Artifact | cstr | Generated here | Difference |
|---|---:|---:|---:|
| F16 model | 163,728,096 | 163,728,096 | 0 bytes |
| Q8_0 model | 141,322,336 | 141,322,752 | +416 bytes |
| Shared voice pack | 522,560 | 522,560 | 0 bytes |
| All 54 generated voices | n/a | 28,218,272 | n/a |
The 416-byte Q8_0 difference is container metadata/alignment; both files have
459 tensors with the same names and the same type counts: 111 Q8_0, 142 F16,
and 206 F32.
Most generated voices are 522,560 bytes; `jf_gongitsune` is 522,592 bytes
because its longer metadata name changes GGUF alignment. The cstr voice
repository has three voices that also occur in the official
Hexgrad v1.0 voice set: `af_heart`, `ef_dora`, and `ff_siwis`. All three
published `voice.pack` tensors are element-for-element identical to our F32
conversions (max error and RMSE both zero). Its four `df_*`/`dm_*` voices are
for the separate German model and have no counterpart in the official v1.0
voice set.
## Tensor errors
Every statistic below is computed after dequantization over all 81,731,256
model elements.
### cstr versus generated here
| Type | Max absolute | RMSE | MAE | Relative L2 | Cosine |
|---|---:|---:|---:|---:|---:|
| F16 | 2.89440155e-4 | 3.33034834e-7 | 3.16075518e-9 | 5.34848643e-6 | 0.999999999986 |
| Q8_0 | 3.76281738e-2 | 1.06066509e-4 | 8.01437548e-6 | 1.70338414e-3 | 0.999998549352 |
For F16, 376 of 459 tensors are exact and 83 differ. The largest difference is
in `pred.N.2.conv2.weight`. The difference is intentional: this repository
uses PyTorch's actual weight-normalization formula,
`v * (g / clamp_min(norm, eps))`. The inherited converter used
`v * clamp_min(g / norm, eps)`, which changes negative learned `weight_g`
values. The corrected formula was checked directly against
`torch._weight_norm` (maximum float32 discrepancy 1.49e-8).
Q8_0 files are independently quantized, so direct Q8-to-Q8 element differences
include different rounding decisions. Comparing each Q8 model with its own F16
baseline is more meaningful:
| Pair | Max absolute | RMSE | MAE | Relative L2 |
|---|---:|---:|---:|---:|
| cstr Q8_0 versus cstr F16 | 1.91650391e-2 | 3.19161067e-4 | 1.16402241e-4 | 5.12567592e-3 |
| Our Q8_0 versus our F16 | 1.91955566e-2 | 3.19376048e-4 | 1.16557019e-4 | 5.12912849e-3 |
The quantization-error distributions are therefore nearly identical; our
global RMSE is 0.067% higher in this comparison.
## Fixed-seed audio output
The four model files synthesized the same sentence with `af_heart`, CPU, four
threads, length scale 1.0, and `KOKORO_SEED=42`. All outputs had exactly 61,800
PCM16 samples (2.575 seconds at 24 kHz).
| Reference -> candidate | RMS delta | Relative RMS | Correlation | SNR dB | Spectral convergence | Log-spectral distance dB |
|---|---:|---:|---:|---:|---:|---:|
| cstr F16 -> our F16 | 6.2137e-4 | 0.01267 | 0.999920 | 37.94 | 0.00549 | 0.529 |
| cstr F16 -> cstr Q8_0 | 0.03647 | 0.74362 | 0.72357 | 2.57 | 0.10702 | 3.315 |
| our F16 -> our Q8_0 | 0.04121 | 0.84010 | 0.64731 | 1.51 | 0.11562 | 3.236 |
| cstr Q8_0 -> our Q8_0 | 0.03570 | 0.72735 | 0.73548 | 2.77 | 0.10940 | 2.809 |
PCM errors are phase-sensitive: small model changes propagate through the
source generator and can produce large sample-aligned differences despite much
smaller magnitude-spectrum differences. These statistics validate
deterministic execution and expose regressions, but do not by themselves prove
perceptual equivalence. A release-quality evaluation should add several texts
per language plus STOI/PESQ or human listening and ASR transcript checks.
## Reproducing comparisons
Tensor comparison:
```bash
python3 tools/compare_gguf.py reference.gguf candidate.gguf --top 20
```
Fixed-seed model output comparison:
```bash
python3 tools/compare_model_outputs.py \
--cli build/kokoro-cli \
--voice models/generated/voices/kokoro-voice-af_heart.gguf \
--model reference=models/kokoro-82m-f16.gguf \
--model candidate=models/generated/kokoro-82m-f16.gguf \
--reference reference \
--text "The documented Q Eight model works." \
--seed 42 --output-dir validation/model-output
```
The output tool preserves every WAV and stderr log and writes `report.json`
with reference-relative and all pairwise waveform/spectral metrics.
|