Azerbaijani TTS — 9.36M, 24 kHz, offline
A text-to-speech model that speaks Azerbaijani entirely offline: no server, no API key, no network at inference. 9.36M parameters, 24 kHz mono, single speaker, 2-4x faster than real time on a laptop CPU.
A fine-tune of owensong/Inflect-Micro-v2
by Owen Song — 200,000 steps on 25.07 hours of Azerbaijani speech, 409 of its
410 tensors carried over bit-identically.
| Hear it | The playground — runs in your browser, nothing installed |
| Code | github.com/HuseynliIlqar/Inflect_Micro_v2_Azerbaijan |
| In this repo | PyTorch at the root, ONNX in onnx/, WAVs in samples/ |
Samples
| Text | Audio |
|---|---|
Salam, bu model tamamilə yerli maşında işləyir. |
|
Payız gəlmişdi və şəhərin küçələri saralmış yarpaqlarla örtülmüşdü. |
|
Sən bu kitabı oxumusan? Mənə çox maraqlı gəldi. |
|
II Dünya müharibəsi 01/09/1939 tarixində başladı və 25% artım oldu. |
The last one shows why the text layer matters: II becomes İkinci,
01/09/1939 becomes birinci sentyabr min doqquz yüz otuz doqquz, 25%
becomes iyirmi beş faiz.
Usage
The weights are useless without the Azerbaijani text layer, which lives in the GitHub repository together with the CLI, the tests and the training pipeline:
git lfs install
git clone https://github.com/HuseynliIlqar/Inflect_Micro_v2_Azerbaijan
cd Inflect_Micro_v2_Azerbaijan
pip install -r requirements.txt
python say.py "Salam, necəsiniz?"
That clone already contains these weights. To use the copy in this repository instead:
from huggingface_hub import snapshot_download
from aztts import AzTTS
tts = AzTTS(snapshot_download("ilqarrrr/Inflect_Micro_v2_Azerbaijan"))
tts.save("Salam, necəsiniz?", "out/salam.wav")
Normalisation and chunking run automatically. Calling the model/ package
directly means calling normalize_az yourself.
Without PyTorch
onnx/ holds the same model as two graphs, verified against the PyTorch module
at a waveform correlation of 0.9999999999916:
| Graph | Inputs | Outputs |
|---|---|---|
onnx/duration.onnx |
tokens, lengths, length_scale |
m_p_exp, logs_p_exp, y_mask |
onnx/decode.onnx |
m_p_exp, logs_p_exp, y_mask, zp_noise, noise_scale |
waveform |
This is what the browser playground runs, through ONNX Runtime Web. web/ in
the GitHub repository is a working implementation, including the Azerbaijani
text layer ported to JavaScript.
Model details
| Architecture | VITS (compact), 9,356,513 parameters |
| Audio | 24 kHz, mono, single speaker |
| Frontend | eSpeak NG (az) through phonemizer, with stress |
| Training | 200,000 steps = 1,343 epochs, ~3.7 days on one A40, ~$39 |
| Data | 25.07 hours, 9,674 clips |
| Speed | 2-4x real time on CPU; ~0.3 s to load |
The text layer ahead of the model does two things the checkpoint cannot: rewrites digits, Roman numerals, dates, units and abbreviations into spoken words, and cuts sentences into ~15-word chunks, because this model's intonation flattens towards the end of a long sentence.
Limitations
- One voice. No voice cloning, no multi-speaker support, no emotion control.
- At 9.36M parameters the speech is clear and intelligible but not fully natural; the vocoder sometimes leaves a metallic resonance.
- Rare names, foreign words and unusual spellings are at the mercy of the phoneme frontend.
- The training audio is most likely synthetic, so the quality ceiling is the
system that produced it. The evidence is in
docs/MODEL.mdin the repository.
Licence and commercial use
Apache-2.0 for this project's code and for these weights, inherited from the base model. Commercial use is allowed, with one condition that comes from the phonemiser rather than the model.
Phonemisation calls phonemizer and
eSpeak NG in-process, and both are
GPL-3.0-or-later. Running them changes nothing. Shipping a combined work —
a desktop app, a container image, a binary handed to customers — brings GPL-3.0
obligations for those components.
The checkpoint itself does not need eSpeak: config.json declares
accepts_prephonemized_input: true, so phonemes can come from any frontend.
THIRD_PARTY_NOTICES.md in the GitHub repository sets out both routes.
Training data. ughurabbasov/azerbaijani-tts-dataset
declares no licence file; asked in its Hugging Face community tab, the author
confirmed free use and asked to be credited. Attribution is the condition of
use — if you build on this model, credit the dataset too.
Files. The PyTorch export sits at the root (model.pth, config.json, the
frontend and the VITS runtime); checksums.sha256 verifies 24 of them.
NOTICES.md is this project's full third-party notice, while LICENSE and
THIRD_PARTY_NOTICES.md are the upstream package's own, unchanged.
Ethical use
Do not use this voice to impersonate a real person or to produce deceptive content. Disclose synthetic speech where the context could otherwise mislead.
Credits and citation
Built on two openly published works: owensong/Inflect-Micro-v2 by Owen
Song (the checkpoint) and ughurabbasov/azerbaijani-tts-dataset by
ughurabbasov (the audio). Please cite all three:
@software{azerbaijani_tts,
title = {Azerbaijani TTS: an offline 9.36M-parameter VITS model},
author = {Huseynli, Ilqar},
year = {2026},
url = {https://github.com/HuseynliIlqar/Inflect_Micro_v2_Azerbaijan},
note = {Adapted from owensong/Inflect-Micro-v2 (Apache-2.0);
trained on ughurabbasov/azerbaijani-tts-dataset,
used with the author's permission}
}
- Downloads last month
- 10
Model tree for ilqarrrr/Inflect_Micro_v2_Azerbaijan
Base model
owensong/Inflect-Micro-v2