StyleTTS2 Somali Single-Speaker Female Voice (Full Training Progression)

This repository hosts a high-quality, single-speaker Somali Text-to-Speech (TTS) model finetuned using the StyleTTS2 architecture. The model was trained on approximately 1 hour and 15 minutes (1,050 audio clips) of clean, high-quality Somali female speech, finetuned on top of pre-trained English LibriTTS weights.

This repository includes the training progression across two main stages, showcasing a dramatic improvement in native Somali phonetic accuracy as the model matured through advanced epochs.

Model Details

  • Architecture: StyleTTS2
  • Dataset: ~1.25 hours of single-speaker clean female Somali speech
  • Base Model: LibriTTS (English)
  • Sample Audio: See the output_somali.wav file in this repository to hear a sample of the synthesized voice.

Training Evolution & Phonetic Accuracy

Because this model was finetuned on an English base, it had to "unlearn" English phonetic rules and adopt native Somali phonology. You will notice a distinct evolution depending on which checkpoint you use:

Stage 1 (stage_1_epoch_2nd_00004_pruned.pth)

This checkpoint represents the early phase of training. While the vocal timbre, pacing, and basic flow are highly natural and successfully cloned, the model still suffers from English-biased phonetic bottlenecks:

  • Phonetic Softening of 'X': The Somali voiceless pharyngeal fricative X ([ħ]) is often softened to a standard English H ([h]) or occasionally mispronounced as "ks".
  • Phonetic Softening of 'C': The Somali voiced pharyngeal fricative C ([ʕ]) is often softened to a standard vowel transition or glottal stop.

Stage 2 (epoch_2nd_00004.pth through epoch_2nd_00039.pth)

During the second stage of training, these phonetic limitations have been significantly improved. While not 100% flawless in every single edge case, transitioning into deep Stage 2 allows the model to map the correct phonemes much more accurately. By utilizing these later checkpoints, the model synthesizes native Somali phonology—including the proper deep pharyngeal 'X' and 'C' sounds—with a high degree of realism and consistency.

Latest Training Metrics (Late Stage 2)

The model shows excellent convergence and stability in the later epochs. Below are the recorded metrics from the advanced training stages (Epoch 42 / Step 410) reflecting the state of the model around and beyond Epoch 39:

Metric Value Metric Value
Total Loss 0.24934 Sty Loss 0.05516
Gen Loss 5.18974 S2S Loss 0.03662
Disc Loss 3.93945 Mono Loss 0.05514
F0 Loss 1.97537 CE Loss 0.01814
Dur Loss 0.42321 Norm Loss 0.33834
LM Loss 1.12631 Diff Loss 0.36748

Usage & Checkpoint Selection

  • For production use or maximum linguistic accuracy: It is highly recommended to use the latest available Stage 2 checkpoints (e.g., epoch_2nd_00034.pth or epoch_2nd_00039.pth). These will provide the most native and realistic Somali pronunciation.
  • For testing/lightweight integration: The Stage 1 pruned checkpoint (stage_1_epoch_2nd_00004_pruned.pth) is provided as a smaller snapshot of the model's early development, though it comes with the phonetic limitations described above.
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Space using Zyroxx66/StyleTTS2-Somali 1