Instructions to use Anadilorg/Anadil_Turkmen_TTS with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- VoxCPM
How to use Anadilorg/Anadil_Turkmen_TTS with VoxCPM:
import soundfile as sf from voxcpm import VoxCPM model = VoxCPM.from_pretrained("Anadilorg/Anadil_Turkmen_TTS") wav = model.generate( text="VoxCPM is an innovative end-to-end TTS model from ModelBest, designed to generate highly expressive speech.", prompt_wav_path=None, # optional: path to a prompt speech for voice cloning prompt_text=None, # optional: reference text cfg_value=2.0, # LM guidance on LocDiT, higher for better adherence to the prompt, but maybe worse inference_timesteps=10, # LocDiT inference timesteps, higher for better result, lower for fast speed normalize=True, # enable external TN tool denoise=True, # enable external Denoise tool retry_badcase=True, # enable retrying mode for some bad cases (unstoppable) retry_badcase_max_times=3, # maximum retrying times retry_badcase_ratio_threshold=6.0, # maximum length restriction for bad case detection (simple but effective), it could be adjusted for slow pace speech ) sf.write("output.wav", wav, 16000) print("saved: output.wav") - Notebooks
- Google Colab
- Kaggle
AnadilTurkmenTTS — Türkmen Text-to-Speech Modeli
AnadilTurkmenTTS, openbmb/VoxCPM2 temel modeli üzerine LoRA (Low-Rank Adaptation) ile ince ayar yapılmış, Türkmen (ISO 639-3: tuk) için açık kaynak bir konuşma sentezi (text-to-speech) modelidir. 2.031 segmentlik tek konuşmacılı bir veri setiyle (Mozilla Common Voice 26.0 tuk katkısı) eğitilmiştir — "Anadil" ailesinin (Lazca, Zazaca, Adığece, Kurmancî, Ermenice, Ladino, Kırgızca, Uyghur, Sakha/Yakut) en yeni üyesidir.
AnadilTurkmenTTS is an open-source text-to-speech LoRA adapter for Turkmen (Turkic language, tuk), fine-tuned on top of VoxCPM2 on 2,031 single-speaker utterances from Mozilla Common Voice 26.0. Turkmen is a Turkic language spoken by ~6 million people in Turkmenistan, Iran, Afghanistan, and other regions — this model is the newest member of the "Anadil" family of minority-language TTS adapters, continuing the effort to give low-resource languages a voice in digital tools.
📋 Model Bilgileri
| Özellik | Detay |
|---|---|
| Dil | Türkmen (tuk) |
| Konuşucu | spk_tmp_001 (tek konuşmacı) |
| Temel model | openbmb/VoxCPM2 |
| Yöntem | LoRA (r=32, α=32, LM + DiT katmanları) |
| Eğitim verisi | 2.031 segment — Common Voice 26.0 tuk |
| Eğitim adımı | 4.999 (kaydedilen son checkpoint) |
| Adapter boyutu | ~72 MB (384 tensor, ~18,1M parametre, F32) |
| Sample rate | 48 kHz çıkış |
| Lisans | MIT |
🔊 Sentez Örnekleri
Aşağıdaki cümleler eğitim sırasında modele görülmüş (in-training) cümlelerdir. Otomatik kalite metrikleri (WER, UTMOS vb.) henüz hesaplanmamıştır.
| # | Türkmen Metin | Sentez |
|---|---|---|
| 1 | O wagt muňa kän ünsem bermändim. | |
| 2 | Ejesi görgüli ýol boýy ogluna guwanyp geldi. | |
| 3 | Mergen göz gytagyny murty ýaňy garalyp başlan oglanlara aýlady. | |
| 4 | Ýok hiç wagt uçarman bolmagy arzuw etmedim | |
| 5 | Men muňa ýok diýemok. |
⚡ Hızlı Başlangıç
Model, hedef konuşmacının referans ses klibi (reference wav) üzerinden ses klonlaması yapar — kendi ses kaydınızı verin.
pip install torch torchaudio soundfile safetensors voxcpm
git clone https://huggingface.co/Anadilorg/Anadil_Turkmen_TTS
cd Anadil_Turkmen_TTS
# Tek cümle sentezle (kendi referans sesinizle)
python inference.py "Salam!" --reference my_voice.wav --out hello.wav
# Demo repo'daki 5 örnek cümlenin tümünü üret (1a..5a.wav)
python inference.py --list-samples
Python API
from inference import AnadilTurkmenTTS
tts = AnadilTurkmenTTS() # openbmb/VoxCPM2 + bu repo'daki LoRA ağırlıkları
tts.synthesize(
"Salam!",
reference_wav="1.wav",
out_path="hello.wav",
cfg_value=2.0,
inference_timesteps=10,
)
Gradio arayüzü (yerel)
pip install gradio
python demo.py # http://localhost:7860
🧪 Doğrulama
python test_smoke.py --weights-only # saniyeler içinde ağırlık bütünlük kontrolü
python test_smoke.py # tam test: temel modeli indirir, Türkmen bir cümle sentezler
⚠️ Kısıtlamalar
- Küçük veri seti: 2.031 segment — ailedeki diğer modellere göre çok daha az; genel konuşurken modelin üslubu eğitim konuşucusuna yakındır.
- Tek konuşmacı: Model tek bir konuşmacı etiketi (
spk_tmp_001) üzerine eğitilmiştir; farklı sesler referans klip ile sıfır-örnekçi klonlama yoluyla yaklaşılmaya çalışılır. - Kalite metrikleri: Henüz otomatik metrikler (WER, UTMOS vb.) raporlanmamıştır; örnekler eğitim içi cümlelerle sınırlıdır.
- Sentez için ~9 GB bellek önerilir (fp32); CUDA tercih edilir, Apple Silicon daha yavaştır.
Citation
@software{anadil-turkmen-tts,
author = {Anadilorg},
title = {AnadilTurkmenTTS: An Open-Source Text-to-Speech LoRA Adapter for Turkmen},
url = {https://huggingface.co/Anadilorg/Anadil_Turkmen_TTS},
version = {0.1.0},
year = {2026}
}
Lisans
MIT — ayrıntılar için LICENSE dosyasına bakınız. Eğitim verisi Mozilla Common Voice 26.0 (tuk) katkısından gelmektedir; ilgili lisans koşullarını (CC-0) lütfen gözden geçiriniz.
- Downloads last month
- 5
Model tree for Anadilorg/Anadil_Turkmen_TTS
Base model
openbmb/VoxCPM2