File size: 3,350 Bytes
fffdc72 dbc15ec fffdc72 dbc15ec | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 | ---
license: cc-by-nc-sa-4.0
language:
- km
base_model: fishaudio/openaudio-s1-mini
pipeline_tag: text-to-speech
tags:
- text-to-speech
- khmer
- lora
- fish-speech
- voice-cloning
---
# Khmer TTS (Panhapich)
LoRA fine-tunes of [fishaudio/openaudio-s1-mini](https://huggingface.co/fishaudio/openaudio-s1-mini)
for **Khmer** text-to-speech, trained in two stages on top of the frozen pretrained
checkpoint:
1. **Khmer base** — LoRA rank 32 fine-tune on the
[DDD-Cambodia/khmer-speech-dataset](https://huggingface.co/datasets/DDD-Cambodia/khmer-speech-dataset)
corpus (~70h, multi-speaker), teaching the model Khmer pronunciation and rhythm.
2. **Voice clone** — a second, low-rank (r=8) LoRA adapted from the Khmer base
checkpoint onto a specific speaker's recordings, so the model speaks Khmer *as that
person*.
Training code: [Pich09/voice-clone](https://github.com/Pich09/voice-clone)
(`kaggle/khmer_tts_kaggle.ipynb` runs the full pipeline end-to-end on a free Kaggle
GPU). See that repo's `ARCHITECTURE.md` for how the pipeline fits together.
## Status
🚧 Training in progress. This repo is used as a **checkpoint relay**: each Kaggle
session pushes its LoRA checkpoint here (with its training step and validation loss)
so training can resume across free, session-limited GPU runs. There is no
production-ready checkpoint here yet — check `registry.json` for the current best.
## Repo layout
```
registry.json # every checkpoint pushed so far, with step/val_loss
lock.json # advisory "who's training right now" lock
checkpoints/step_<step>_<trainer>/ # one LoRA checkpoint per push
```
`registry.json` entries look like:
```json
{"step": 8000, "val_loss": 2.13, "trainer_id": "panhapich", "shard_index": 0,
"path": "checkpoints/step_0008000_panhapich", "created": 1753000000.0}
```
The "best" checkpoint is the lowest `val_loss` (ties broken by highest step); if no
run has a parseable val_loss yet, the most-trained (highest step) one is used instead.
See `khmer_tts/collab/registry.py` in the training repo for the exact selection logic.
## Model architecture
Fish Speech is **autoregressive**, not diffusion: a text frontend turns Khmer
graphemes into token ids (no phonemizer/G2P), an LLM-style transformer predicts
discrete audio (VQ) tokens autoregressively, and a frozen DAC-style vocoder decodes
those tokens to a waveform. Only the middle (acoustic codec-LM) stage is LoRA
fine-tuned here — the text frontend and vocoder are untouched.
## Intended use & limitations
- Non-commercial Khmer TTS / voice-cloning research, inheriting `openaudio-s1-mini`'s
CC-BY-NC-SA-4.0 license.
- **Not intended for**: impersonation without consent, voice-authentication bypass,
or generating misleading audio of real individuals.
- Voice-clone-stage checkpoints require the speaker's explicit, recorded consent
before any use beyond personal experimentation.
- Pronunciation quality is bounded by the training data's coverage of Khmer dialects,
numbers/dates/currency normalization, and code-switched text — see the training
repo's `MODEL_CARD.md` for the filled-in evaluation once available.
## License
Derived from `fishaudio/openaudio-s1-mini`, licensed CC-BY-NC-SA-4.0 (non-commercial,
share-alike). These checkpoints inherit that license.
|