Kokoro / README.md
Thorsten-Voice's picture
Update README.md
734e593 verified
|
Raw
History Blame Contribute Delete
8.03 kB
---
language: de
license: apache-2.0
tags:
- text-to-speech
- tts
- german
- kokoro
- styletts2
datasets:
- thorsten-voice
---
# Thorsten-Voice/Kokoro
A German fine-tune of [Kokoro-82M](https://huggingface.co/hexgrad/Kokoro-82M) on the [Thorsten-Voice](https://www.thorsten-voice.de/) dataset — a fast, high-quality, CPU-friendly text-to-speech model that speaks with Thorsten's own voice.
Kokoro-82M is a compact (82M parameter) TTS model based on the StyleTTS2 architecture. Its small size means it runs comfortably on CPU, in real time or faster, without requiring a GPU — making it well suited for local, offline use.
## Credits
This model would not exist without:
- **[hexgrad/Kokoro-82M](https://huggingface.co/hexgrad/Kokoro-82M)** — the original English Kokoro-82M model and architecture.
- **[kikiri-tts](https://github.com/semidark/kikiri-tts)** by semidark — the German fine-tuning recipe, patched StyleTTS2/Kokoro submodules, and documentation this training was based on. Please check out their project and the German community voices (Martin, Victoria) they've already published.
Fine-tuned on the [Thorsten-Voice](https://www.thorsten-voice.de/) dataset (CC0 / public domain).
## Files in this repository
| File | Description |
|---|---|
| `config.json` | Kokoro-82M architecture config (unchanged from the base model) |
| `model.pth` | **Default checkpoint (epoch 5).** Fine-tuned weights (`bert`, `bert_encoder`, `predictor`, `text_encoder`, `decoder`), converted from the Stage 2 StyleTTS2 checkpoint |
| `voices/thorsten.pt` | Voicepack matching the default (epoch 5) checkpoint |
| `model_ep{1,2,3,4,6,7,8,9,10}.pth` | All other Stage 2 checkpoints (epochs 1–4, 6–10), same converted, ready-to-use format as `model.pth` |
| `voices/thorsten_ep{1,2,3,4,6,7,8,9,10}.pt` | Matching voicepacks for each of the above |
## Installation
This model requires the German-language forks of `misaki` and `kokoro` (the official PyPI `misaki` package does not include the `de` submodule needed for German G2P), plus the `espeak-ng` system package that `misaki` relies on for phonemization:
```bash
# System dependency (required by misaki for German G2P)
# macOS:
brew install espeak-ng
# Debian/Ubuntu:
sudo apt-get install espeak-ng
# Python dependencies
pip install huggingface_hub soundfile numpy torch
pip install "git+https://github.com/semidark/misaki.git@6d252a2e02f3b030f22f56686f1a73786c16ffc8"
pip install "git+https://github.com/semidark/kokoro.git"
```
## Usage
```python
import numpy as np
import soundfile as sf
import torch
from huggingface_hub import hf_hub_download
from kokoro import KModel, KPipeline
REPO_ID = "Thorsten-Voice/Kokoro"
device = "cuda" if torch.cuda.is_available() else "cpu"
config_path = hf_hub_download(repo_id=REPO_ID, filename="config.json")
model_path = hf_hub_download(repo_id=REPO_ID, filename="model.pth")
voice_path = hf_hub_download(repo_id=REPO_ID, filename="voices/thorsten.pt")
kmodel = KModel(repo_id="hexgrad/Kokoro-82M", config=config_path, model=model_path)
kmodel = kmodel.to(device).eval()
pipeline = KPipeline(lang_code="d", repo_id="hexgrad/Kokoro-82M", model=kmodel)
# Workaround: misaki's German G2P can emit 'ʏ' (short ü), which is not in
# Kokoro's vocabulary (only 'y' is). See "Known limitations" below.
_original_g2p = pipeline.g2p
pipeline.g2p = lambda text: (lambda ps, tok: (ps.replace("ʏ", "y"), tok))(*_original_g2p(text))
voice = torch.load(voice_path, map_location="cpu", weights_only=True)
text = "Hallo, hier spricht Thorsten."
audio_chunks = [audio for _, _, audio in pipeline(text, voice=voice, speed=1.0)]
combined = np.concatenate(audio_chunks)
sf.write("output.wav", combined, 24000)
```
A ready-to-run version of this snippet is included as [`inference.py`](./inference.py):
```bash
# Default checkpoint (epoch 5)
python inference.py "Hallo, hier spricht Thorsten." output.wav
# Any other epoch (1-10) - e.g. epoch 10, faster/tighter delivery
python inference.py "Hallo, hier spricht Thorsten." output.wav ep10
python inference.py "Hallo, hier spricht Thorsten." output.wav ep3
```
## Audio samples
Sample outputs from the default (epoch 5) checkpoint, covering German pronunciation edge cases (umlauts, ich/ach-laut, eszett, consonant clusters, numbers, prosody) and technical/loanword pronunciation overrides:
<audio controls src="https://huggingface.co/Thorsten-Voice/Kokoro/resolve/main/test_audio_epoch5/test_04.wav"></audio> Zwei weiße Zwerge zwängen sich zwischen zwei Zweige.
<audio controls src="https://huggingface.co/Thorsten-Voice/Kokoro/resolve/main/test_audio_epoch5/test_02.wav"></audio> Ich mache mich auf den Weg nach Aachen, um auch Nachts wach zu sein.
<audio controls src="https://huggingface.co/Thorsten-Voice/Kokoro/resolve/main/test_audio_epoch5/override_02.wav"></audio> Lade die API oder ein JSON herunter.
All 14 samples are available under [`test_audio_epoch5/`](./test_audio_epoch5).
## Training details
- **Base model:** Kokoro-82M (English), architecture unchanged
- **Dataset:** Thorsten-Voice (12,283 training sentences after cleaning/filtering)
- **Training recipe:** [kikiri-tts](https://github.com/semidark/kikiri-tts), a patched StyleTTS2 fine-tuning pipeline for Kokoro
- **Stage 1** (acoustic/alignment): 10 epochs
- **Stage 2** (adversarial/prosody, with GAN + WavLM discriminator): 10 epochs
### Stage 2 checkpoint comparison
Validation loss stayed essentially flat across the second half of Stage 2 training, with epoch 5 and epoch 10 tied for the lowest value. Epoch 10 has a slightly lower F0 (pitch) loss, suggesting more refined prosody after additional adversarial fine-tuning — but in informal listening comparisons, **epoch 5 was judged more natural**, with a slightly slower, less "clipped" speaking pace. The metrics alone did not predict this; it only became apparent by listening to both checkpoints on identical sentences.
All 10 Stage 2 checkpoints are included in this repository, already converted to Kokoro's inference format and ready to use via `inference.py` (see [Usage](#usage) above) — no separate conversion step needed.
| Epoch | Validation loss | Duration loss | F0 loss | `inference.py` variant |
|---|---|---|---|---|
| 1 | 0.288 | 0.455 | 2.281 | `ep1` |
| 2 | 0.285 | 0.432 | 2.131 | `ep2` |
| 3 | 0.283 | 0.427 | 2.085 | `ep3` |
| 4 | 0.272 | 0.439 | 2.015 | `ep4` |
| 5 | 0.269 | 0.420 | 1.957 | `ep5` / default |
| 6 | 0.274 | 0.427 | 1.953 | `ep6` |
| 7 | 0.271 | 0.422 | 1.883 | `ep7` |
| 8 | 0.271 | 0.420 | 1.869 | `ep8` |
| 9 | 0.271 | 0.425 | 1.846 | `ep9` |
| 10 | 0.269 | 0.416 | **1.800** | `ep10` |
Only epochs 5 and 10 were carefully compared by ear; the others are provided as-is for anyone curious to explore the full training trajectory. Feedback on the intermediate checkpoints is welcome.
## Known limitations
- The German G2P frontend (`misaki.de.DEG2P`) can emit the short-ü symbol `ʏ` (e.g. in "Brücke"), which is **not** part of Kokoro's 178-symbol vocabulary — only the long-ü symbol `y` is. Left unhandled, this silently breaks short-ü words during inference. [`inference.py`](./inference.py) includes a small workaround that patches the G2P output to replace `ʏ` with `y` (the same substitution used when preparing the training data). If you're writing your own inference code instead of using the provided script, make sure to apply this substitution yourself.
- Hyphenated German compound adjectives (e.g. *"atmosphärisch-optisches"*) can occasionally produce unnatural prosody or stumbling at the hyphen boundary. This is a known limitation of the underlying G2P frontend (`misaki`/`espeak-ng`) shared with other TTS systems (including Piper), not specific to this fine-tune.
- As with the base Kokoro model, very long sentences may be truncated; splitting text on sentence boundaries is recommended for longer inputs.
## License
Released under Apache 2.0, consistent with the base Kokoro-82M model and the CC0-licensed Thorsten-Voice dataset used for fine-tuning.