gooya-0.1 / model.md
Reza2kn's picture
Mirror Reza2kn/gooya-0.1 at revision 852a9ed389527da1607912d1620e789358b94ac9
cd2fdb7 verified
|
Raw
History Blame Contribute Delete
2 kB

Gooya 0.1 — Persian TTS

Gooya 0.1 is a Persian text-to-speech model: a LoRA fine-tune of MOSS-TTS-Nano (100M) trained on 48k Persian clips. Pronunciation comes from the Gooya clean G2P (Reza2kn/negara-g2p-clean), so the model reads phonemes — including the greeting-ezafe rule (سلام من → salAm man) — not raw graphemes.

  • Architecture: global-local transformer (GPT-2 768 hidden, 12 layers + 1 local transformer) + 16-codebook audio tokenizer
  • Base: OpenMOSS-Team/MOSS-TTS-Nano-100M
  • Fine-tune: LoRA r16/α32 on 48 transformer.h.* linears (1.97% of params), bf16, 3 epochs
  • Data: Reza2kn/gooya-v7-chizzled-multispeaker (raw Persian + audio), same-speaker ref pairs
  • Output: 48 kHz mono WAV (via MOSS-Audio-Tokenizer-Nano)

Usage

# 1) phonemize Persian -> char-spaced phonemes
python scripts/negara_tts_phonemize.py "سلام من رضا هستم"
#    سلام من رضا هستم	salAm man rezA hastam

# 2) synthesize (voice-clone mode)
python scripts/demo_gooya_01.py "سلام من رضا هستم" \
    --ref /path/to/reference.wav --out out.wav

Everything loads from Hugging Face by repo id — no local model files needed.

Gate C (held-out, 200 rows)

6 held-out rows synthesized in two timbres:

clip judges
gate_c/NN_manaref.wav pronunciation in base voice
gate_c/NN_demoref.wav pronunciation + cross-voice clone

Current verdict: pronunciation is good; manaref clips sound clearly better than demoref clones. Voice cloning works but is imperfect.

Known limitations

  • Phoneme input only — feed raw Persian through the G2P first.
  • Cloned timbre drifts on longer/lower-quality refs.
  • Trained on 48k rows; pronunciation edge cases (rare homographs, foreign names) still benefit from the G2P rule set.

Cite / license

Experimental research artifact. No warranty. Training recipe in finetune_config.json; full pipeline in the gooya-1 repo (scripts/box_v7_negara.sh).