--- language: - en - es - fr - hi - it - ja - pt - zh license: apache-2.0 library_name: kokoro.cpp pipeline_tag: text-to-speech base_model: hexgrad/Kokoro-82M base_model_relation: quantized inference: false tags: - audio - text-to-speech - tts - kokoro - gguf - ggml - cpp - cpu - gpu - multilingual --- # Kokoro-82M for `kokoro.cpp` — F16, Q8_0, and all 54 voices Reproducible GGUF conversions of the official [Hexgrad Kokoro-82M v1.0](https://huggingface.co/hexgrad/Kokoro-82M) checkpoint for [**`kokoro.cpp`**](https://github.com/simonfxr/kokoro.cpp), a standalone C/C++ runtime powered by GGML. > **What is included:** the reference F16 model, a validated Q8_0 model, and > every one of the 54 official v1.0 voice packs across all supported languages. > No unofficial or third-party voices are mixed into this repository. ## Files | File | Precision | Size | Notes | |---|---:|---:|---| | `kokoro-82m-f16.gguf` | F16/F32 | 156.14 MiB | Reference-quality conversion | | `kokoro-82m-q8_0.gguf` | Q8_0/F16/F32 | 134.78 MiB | 111 matrix tensors quantized; sensitive/unsupported tensors retained | | `voices/kokoro-voice-*.gguf` | F32 | 26.91 MiB total | All 54 official voices; about 510 KiB each | | `manifest.json` | — | — | SHA256, size, source revision, and checkpoint provenance | | `VALIDATION.md` | — | — | Tensor and fixed-seed audio comparison results | Voice packs remain F32 because they are small and are indexed style embeddings, not matrix-multiplication weights. A model file and one voice file are both required for synthesis. ## Quick start with `kokoro.cpp` ```bash # Download one model and one voice. hf download simonfxr/kokoro.cpp-GGUF \ kokoro-82m-q8_0.gguf \ voices/kokoro-voice-af_heart.gguf \ --local-dir models # Build the runtime. See the GitHub README for platform/backend options. git clone --recursive https://github.com/simonfxr/kokoro.cpp.git cd kokoro.cpp cmake -S . -B build -DCMAKE_BUILD_TYPE=Release cmake --build build -j # Synthesize 24 kHz mono audio. ./build/kokoro-cli \ -m ../models/kokoro-82m-q8_0.gguf \ -v ../models/voices/kokoro-voice-af_heart.gguf \ -l en-us \ -o output.wav \ "Hello from Kokoro dot C P P." ``` GPU execution is available when `kokoro.cpp` is built with a supported Vulkan, Metal, or CUDA backend. GGUF is the container format; llama.cpp and whisper.cpp do not implement this Kokoro architecture. ## Languages and voices Kokoro reports eight languages. Its pipeline distinguishes American and British English, resulting in nine language variants. `kokoro.cpp` takes the espeak-ng voice shown in the third column, not Kokoro's one-letter pipeline code. | Language | Kokoro code | `kokoro.cpp -l` | Voices | |---|:---:|:---:|---| | American English | `a` | `en-us` | 20 | | British English | `b` | `en-gb` | 8 | | Spanish | `e` | `es` | 3 | | French | `f` | `fr-fr` | 1 | | Hindi | `h` | `hi` | 4 | | Italian | `i` | `it` | 2 | | Japanese | `j` | `ja` | 5 | | Brazilian Portuguese | `p` | `pt-br` | 3 | | Mandarin Chinese | `z` | `cmn` | 8 | See [`voices/README.md`](voices/README.md) for every voice name and ready-to-copy commands. The machine-readable mapping is in [`voices/languages.json`](voices/languages.json). ## Conversion and provenance - Base model: [`hexgrad/Kokoro-82M`](https://huggingface.co/hexgrad/Kokoro-82M) - Pinned source revision: `f3ff3571791e39611d31c381e3a41a3af07b4987` - Official checkpoint SHA256: `496dba118d1a58f5f3db2efc88dbdc216e0483fc89fe6e47ee1f2c53f18ad1e4` - Runtime and converters: [`simonfxr/kokoro.cpp`](https://github.com/simonfxr/kokoro.cpp) - Output checksums: [`manifest.json`](manifest.json) The conversion pipeline fuses PyTorch weight normalization, preserves Kokoro's sparse 178-entry IPA vocabulary, maps all 459 runtime tensors, and emits GGUF metadata expected by `kokoro.cpp`. Q8_0 quantizes only eligible two-dimensional matrix weights whose GGML row width supports 32-element blocks. Convolutional kernels remain F16 and rank-one tensors remain F32. Reproduce these files from the official weights: ```bash git clone --recursive https://github.com/simonfxr/kokoro.cpp.git cd kokoro.cpp python3 -m venv .venv-convert . .venv-convert/bin/activate pip install -r models/requirements-convert.txt ./models/convert-original.sh ``` ## Validation The generated files were checked for tensor names, shapes, dtypes, metadata, and source hashes. Both models were loaded by `kokoro.cpp`; F16 and Q8_0 English and Q8_0 Spanish synthesis produced valid 24 kHz WAV files. Against the previously published cstr files: - The F16 file has the same size and cosine similarity `0.999999999986` over all 81,731,256 dequantized elements. - Our Q8_0-versus-F16 relative L2 error is `0.00512913`, essentially matching cstr's `0.00512568`. - Shared official voices `af_heart`, `ef_dora`, and `ff_siwis` are exactly equal element-for-element. - Fixed-seed cstr-F16 versus this F16 output has waveform correlation `0.999920` and log-spectral distance `0.529 dB`. Full methodology, caveats, and pairwise results are in [`VALIDATION.md`](VALIDATION.md). PCM comparisons are phase-sensitive and do not replace multilingual listening tests or perceptual metrics. ## Limitations - Maximum input is 510 raw phoneme tokens because the model adds two boundary tokens to PL-BERT's 512-position context. - Phonemization is provided by espeak-ng in `kokoro.cpp`, while the official Python package uses Misaki. Pronunciation can therefore differ. - Mandarin tone numbers are not represented in Kokoro's 178-token vocabulary. - espeak-ng Japanese kanji coverage is incomplete; pre-phonemized IPA can be supplied when needed. - Q8_0 is smaller and validated, but F16 remains the reference choice for maximum fidelity. Lower-bit quantizations are intentionally not published. - Voice availability does not imply identical quality across languages. Consult the [official model card](https://huggingface.co/hexgrad/Kokoro-82M) and [official voice notes](https://huggingface.co/hexgrad/Kokoro-82M/blob/main/VOICES.md). ## Intended use This repository is for local and embedded text-to-speech inference with `kokoro.cpp`, including research, accessibility, applications, and deployment permitted by the upstream Apache-2.0 license. Users remain responsible for appropriate disclosure, consent, and compliance when publishing synthetic speech. Do not use generated audio to impersonate people or facilitate fraud. ## License and attribution The Kokoro-82M weights and official voice packs are Apache-2.0 licensed. The GGUF files are numerical conversions of those weights and retain that license. `kokoro.cpp` has its own software/dependency licensing details; consult its [LICENSE](https://github.com/simonfxr/kokoro.cpp/blob/main/LICENSE) and [NOTICE](https://github.com/simonfxr/kokoro.cpp/blob/main/NOTICE) before distributing binaries. Architecture: StyleTTS 2 / ISTFTNet. Original model, training information, datasets, acknowledgements, and citations are documented by [Hexgrad](https://huggingface.co/hexgrad/Kokoro-82M).