Gooya 0.1 — Persian TTS
Gooya 0.1 is a Persian text-to-speech model: a LoRA fine-tune of
MOSS-TTS-Nano (100M) trained on 48k Persian clips. Pronunciation comes from the
Gooya clean G2P (Reza2kn/negara-g2p-clean), so the model reads phonemes —
including the greeting-ezafe rule (سلام من → salAm man) — not raw graphemes.
- Architecture: global-local transformer (GPT-2 768 hidden, 12 layers + 1 local transformer) + 16-codebook audio tokenizer
- Base:
OpenMOSS-Team/MOSS-TTS-Nano-100M - Fine-tune: LoRA r16/α32 on 48
transformer.h.*linears (1.97% of params), bf16, 3 epochs - Data:
Reza2kn/gooya-v7-chizzled-multispeaker(raw Persian + audio), same-speaker ref pairs - Output: 48 kHz mono WAV (via MOSS-Audio-Tokenizer-Nano)
Usage
# 1) phonemize Persian -> char-spaced phonemes
python scripts/negara_tts_phonemize.py "سلام من رضا هستم"
# سلام من رضا هستم salAm man rezA hastam
# 2) synthesize (voice-clone mode)
python scripts/demo_gooya_01.py "سلام من رضا هستم" \
--ref /path/to/reference.wav --out out.wav
Everything loads from Hugging Face by repo id — no local model files needed.
Gate C (held-out, 200 rows)
6 held-out rows synthesized in two timbres:
| clip | judges |
|---|---|
gate_c/NN_manaref.wav |
pronunciation in base voice |
gate_c/NN_demoref.wav |
pronunciation + cross-voice clone |
Current verdict: pronunciation is good; manaref clips sound clearly better than demoref clones. Voice cloning works but is imperfect.
Known limitations
- Phoneme input only — feed raw Persian through the G2P first.
- Cloned timbre drifts on longer/lower-quality refs.
- Trained on 48k rows; pronunciation edge cases (rare homographs, foreign names) still benefit from the G2P rule set.
Cite / license
Experimental research artifact. No warranty. Training recipe in
finetune_config.json; full pipeline in the gooya-1 repo
(scripts/box_v7_negara.sh).