kokoro.cpp-GGUF / README.md
simonfxr's picture
Publish Kokoro-82M F16/Q8_0 and all 54 official voices
c64650d verified
|
Raw
History Blame Contribute Delete
7.08 kB
---
language:
- en
- es
- fr
- hi
- it
- ja
- pt
- zh
license: apache-2.0
library_name: kokoro.cpp
pipeline_tag: text-to-speech
base_model: hexgrad/Kokoro-82M
base_model_relation: quantized
inference: false
tags:
- audio
- text-to-speech
- tts
- kokoro
- gguf
- ggml
- cpp
- cpu
- gpu
- multilingual
---
# Kokoro-82M for `kokoro.cpp` — F16, Q8_0, and all 54 voices
Reproducible GGUF conversions of the official
[Hexgrad Kokoro-82M v1.0](https://huggingface.co/hexgrad/Kokoro-82M) checkpoint for
[**`kokoro.cpp`**](https://github.com/simonfxr/kokoro.cpp), a standalone C/C++ runtime powered by GGML.
> **What is included:** the reference F16 model, a validated Q8_0 model, and
> every one of the 54 official v1.0 voice packs across all supported languages.
> No unofficial or third-party voices are mixed into this repository.
## Files
| File | Precision | Size | Notes |
|---|---:|---:|---|
| `kokoro-82m-f16.gguf` | F16/F32 | 156.14 MiB | Reference-quality conversion |
| `kokoro-82m-q8_0.gguf` | Q8_0/F16/F32 | 134.78 MiB | 111 matrix tensors quantized; sensitive/unsupported tensors retained |
| `voices/kokoro-voice-*.gguf` | F32 | 26.91 MiB total | All 54 official voices; about 510 KiB each |
| `manifest.json` | — | — | SHA256, size, source revision, and checkpoint provenance |
| `VALIDATION.md` | — | — | Tensor and fixed-seed audio comparison results |
Voice packs remain F32 because they are small and are indexed style embeddings,
not matrix-multiplication weights. A model file and one voice file are both
required for synthesis.
## Quick start with `kokoro.cpp`
```bash
# Download one model and one voice.
hf download simonfxr/kokoro.cpp-GGUF \
kokoro-82m-q8_0.gguf \
voices/kokoro-voice-af_heart.gguf \
--local-dir models
# Build the runtime. See the GitHub README for platform/backend options.
git clone --recursive https://github.com/simonfxr/kokoro.cpp.git
cd kokoro.cpp
cmake -S . -B build -DCMAKE_BUILD_TYPE=Release
cmake --build build -j
# Synthesize 24 kHz mono audio.
./build/kokoro-cli \
-m ../models/kokoro-82m-q8_0.gguf \
-v ../models/voices/kokoro-voice-af_heart.gguf \
-l en-us \
-o output.wav \
"Hello from Kokoro dot C P P."
```
GPU execution is available when `kokoro.cpp` is built with a supported Vulkan,
Metal, or CUDA backend. GGUF is the container format; llama.cpp and whisper.cpp
do not implement this Kokoro architecture.
## Languages and voices
Kokoro reports eight languages. Its pipeline distinguishes American and British
English, resulting in nine language variants. `kokoro.cpp` takes the espeak-ng
voice shown in the third column, not Kokoro's one-letter pipeline code.
| Language | Kokoro code | `kokoro.cpp -l` | Voices |
|---|:---:|:---:|---|
| American English | `a` | `en-us` | 20 |
| British English | `b` | `en-gb` | 8 |
| Spanish | `e` | `es` | 3 |
| French | `f` | `fr-fr` | 1 |
| Hindi | `h` | `hi` | 4 |
| Italian | `i` | `it` | 2 |
| Japanese | `j` | `ja` | 5 |
| Brazilian Portuguese | `p` | `pt-br` | 3 |
| Mandarin Chinese | `z` | `cmn` | 8 |
See [`voices/README.md`](voices/README.md) for every voice name and ready-to-copy
commands. The machine-readable mapping is in
[`voices/languages.json`](voices/languages.json).
## Conversion and provenance
- Base model: [`hexgrad/Kokoro-82M`](https://huggingface.co/hexgrad/Kokoro-82M)
- Pinned source revision: `f3ff3571791e39611d31c381e3a41a3af07b4987`
- Official checkpoint SHA256: `496dba118d1a58f5f3db2efc88dbdc216e0483fc89fe6e47ee1f2c53f18ad1e4`
- Runtime and converters: [`simonfxr/kokoro.cpp`](https://github.com/simonfxr/kokoro.cpp)
- Output checksums: [`manifest.json`](manifest.json)
The conversion pipeline fuses PyTorch weight normalization, preserves Kokoro's
sparse 178-entry IPA vocabulary, maps all 459 runtime tensors, and emits GGUF
metadata expected by `kokoro.cpp`. Q8_0 quantizes only eligible two-dimensional
matrix weights whose GGML row width supports 32-element blocks. Convolutional
kernels remain F16 and rank-one tensors remain F32.
Reproduce these files from the official weights:
```bash
git clone --recursive https://github.com/simonfxr/kokoro.cpp.git
cd kokoro.cpp
python3 -m venv .venv-convert
. .venv-convert/bin/activate
pip install -r models/requirements-convert.txt
./models/convert-original.sh
```
## Validation
The generated files were checked for tensor names, shapes, dtypes, metadata,
and source hashes. Both models were loaded by `kokoro.cpp`; F16 and Q8_0 English
and Q8_0 Spanish synthesis produced valid 24 kHz WAV files.
Against the previously published cstr files:
- The F16 file has the same size and cosine similarity `0.999999999986` over all
81,731,256 dequantized elements.
- Our Q8_0-versus-F16 relative L2 error is `0.00512913`, essentially matching
cstr's `0.00512568`.
- Shared official voices `af_heart`, `ef_dora`, and `ff_siwis` are exactly equal
element-for-element.
- Fixed-seed cstr-F16 versus this F16 output has waveform correlation `0.999920`
and log-spectral distance `0.529 dB`.
Full methodology, caveats, and pairwise results are in
[`VALIDATION.md`](VALIDATION.md). PCM comparisons are phase-sensitive and do not
replace multilingual listening tests or perceptual metrics.
## Limitations
- Maximum input is 510 raw phoneme tokens because the model adds two boundary
tokens to PL-BERT's 512-position context.
- Phonemization is provided by espeak-ng in `kokoro.cpp`, while the official
Python package uses Misaki. Pronunciation can therefore differ.
- Mandarin tone numbers are not represented in Kokoro's 178-token vocabulary.
- espeak-ng Japanese kanji coverage is incomplete; pre-phonemized IPA can be
supplied when needed.
- Q8_0 is smaller and validated, but F16 remains the reference choice for
maximum fidelity. Lower-bit quantizations are intentionally not published.
- Voice availability does not imply identical quality across languages. Consult
the [official model card](https://huggingface.co/hexgrad/Kokoro-82M) and
[official voice notes](https://huggingface.co/hexgrad/Kokoro-82M/blob/main/VOICES.md).
## Intended use
This repository is for local and embedded text-to-speech inference with
`kokoro.cpp`, including research, accessibility, applications, and deployment
permitted by the upstream Apache-2.0 license. Users remain responsible for
appropriate disclosure, consent, and compliance when publishing synthetic
speech. Do not use generated audio to impersonate people or facilitate fraud.
## License and attribution
The Kokoro-82M weights and official voice packs are Apache-2.0 licensed. The
GGUF files are numerical conversions of those weights and retain that license.
`kokoro.cpp` has its own software/dependency licensing details; consult its
[LICENSE](https://github.com/simonfxr/kokoro.cpp/blob/main/LICENSE) and
[NOTICE](https://github.com/simonfxr/kokoro.cpp/blob/main/NOTICE) before distributing binaries.
Architecture: StyleTTS 2 / ISTFTNet. Original model, training information,
datasets, acknowledgements, and citations are documented by
[Hexgrad](https://huggingface.co/hexgrad/Kokoro-82M).