| --- |
| license: apache-2.0 |
| language: [as, bn, brx, doi, en, gu, hi, kn, kok, mai, ml, mni, mr, ne, or, pa, sa, sat, sd, ta, te, ur] |
| library_name: onnx |
| pipeline_tag: text-to-speech |
| tags: [onnx, text-to-speech, tts, indic, parler-tts, phoonnx] |
| base_model: ai4bharat/indic-parler-tts |
| --- |
| |
| # phoonnx — Indic Parler-TTS (ONNX) |
|
|
| ONNX export of [`ai4bharat/indic-parler-tts`](https://huggingface.co/ai4bharat/indic-parler-tts) |
| for the [phoonnx](https://github.com/OpenVoiceOS/phoonnx) `indic_parler` engine. |
|
|
| Indic Parler-TTS speaks 20 Indic languages and English. You select the voice with a |
| natural-language description ("Rohit's voice is clear and expressive..."), not with a |
| speaker id or a reference clip. |
|
|
| ## Graphs |
|
|
| | File | Size | Purpose | |
| |---|---|---| |
| | `text_encoder.onnx` | 1.4 GB | Flan-T5 encoder over the voice description | |
| | `decoder_prefill.onnx` | 2.1 GB | First AR step; emits self-attention **and** cross-attention KV | |
| | `decoder_decode.onnx` | 1.3 GB | Later AR steps; consumes both caches, emits self-attention KV only | |
| | `dac_decoder.onnx` | 217 MB | DAC 44.1 kHz codec decoder (9 codebooks) | |
| | `tokenizer.json` | | prompt tokenizer (the text you want spoken) | |
| | `description_tokenizer.json` | | description tokenizer (Flan-T5 vocabulary) | |
| | `config.json` | | self-describing phoonnx config (`engine: indic_parler`) | |
|
|
| The two tokenizers are **different vocabularies**. Upstream is explicit about this: one |
| tokenizer for the prompt, one for the description. |
|
|
| ## Architecture |
|
|
| description --> Flan-T5 encoder --> encoder states --> cross-attention (all 24 layers) |
| prompt --> embed_prompts --> prepended to the decoder input embeddings |
| decoder --> 9 delayed DAC codebooks --> DAC decoder --> 44.1 kHz mono |
| |
| Cross-attention keys and values do not change while decoding, so `decoder_prefill` |
| computes them once and `decoder_decode` reads them back unchanged. |
|
|
| ## Parity |
|
|
| Every graph is float32 and was checked against the PyTorch model |
| (`parler_tts.ParlerTTSForConditionalGeneration`, eager attention) on ser9 CPU: |
|
|
| | Check | Result | |
| |---|---| |
| | Text encoder, max abs diff (hi / en / ta) | 3.9e-07 / 6.5e-06 / 3.2e-07 | |
| | Prefill KV, max abs diff (self K/V, cross K/V) | 8.1e-06 / 4.5e-06 / 2.4e-06 / 1.2e-06 | |
| | Greedy codes, 60 steps | identical, all 3 languages | |
| | Waveform correlation vs torch | 0.9999999917 / 1.0000000000 / 1.0000000000 | |
| | Waveform RMSE vs torch | 1.3e-08 / 5.0e-08 / 1.5e-08 | |
|
|
| No quantised variants are published: the quantised graphs were not verified, and |
| phoonnx does not ship unverified exports. |
|
|
| ## Notice |
|
|
| Upstream model: `ai4bharat/indic-parler-tts` by [AI4Bharat](https://ai4bharat.iitm.ac.in/), |
| built on [`huggingface/parler-tts`](https://github.com/huggingface/parler-tts) by Yoach |
| Lacombe, Vaibhav Srivastav and Sanchit Gandhi. Apache-2.0, preserved from upstream. |
| The DAC codec is Descript's, as vendored by upstream. |
|
|
| The weights are unmodified: the export wraps the PyTorch modules and traces them. No |
| surgery, no quantisation. |
|
|