--- license: apache-2.0 language: [as, bn, brx, doi, en, gu, hi, kn, kok, mai, ml, mni, mr, ne, or, pa, sa, sat, sd, ta, te, ur] library_name: onnx pipeline_tag: text-to-speech tags: [onnx, text-to-speech, tts, indic, parler-tts, phoonnx] base_model: ai4bharat/indic-parler-tts --- # phoonnx — Indic Parler-TTS (ONNX) ONNX export of [`ai4bharat/indic-parler-tts`](https://huggingface.co/ai4bharat/indic-parler-tts) for the [phoonnx](https://github.com/OpenVoiceOS/phoonnx) `indic_parler` engine. Indic Parler-TTS speaks 20 Indic languages and English. You select the voice with a natural-language description ("Rohit's voice is clear and expressive..."), not with a speaker id or a reference clip. ## Graphs | File | Size | Purpose | |---|---|---| | `text_encoder.onnx` | 1.4 GB | Flan-T5 encoder over the voice description | | `decoder_prefill.onnx` | 2.1 GB | First AR step; emits self-attention **and** cross-attention KV | | `decoder_decode.onnx` | 1.3 GB | Later AR steps; consumes both caches, emits self-attention KV only | | `dac_decoder.onnx` | 217 MB | DAC 44.1 kHz codec decoder (9 codebooks) | | `tokenizer.json` | | prompt tokenizer (the text you want spoken) | | `description_tokenizer.json` | | description tokenizer (Flan-T5 vocabulary) | | `config.json` | | self-describing phoonnx config (`engine: indic_parler`) | The two tokenizers are **different vocabularies**. Upstream is explicit about this: one tokenizer for the prompt, one for the description. ## Architecture description --> Flan-T5 encoder --> encoder states --> cross-attention (all 24 layers) prompt --> embed_prompts --> prepended to the decoder input embeddings decoder --> 9 delayed DAC codebooks --> DAC decoder --> 44.1 kHz mono Cross-attention keys and values do not change while decoding, so `decoder_prefill` computes them once and `decoder_decode` reads them back unchanged. ## Parity Every graph is float32 and was checked against the PyTorch model (`parler_tts.ParlerTTSForConditionalGeneration`, eager attention) on ser9 CPU: | Check | Result | |---|---| | Text encoder, max abs diff (hi / en / ta) | 3.9e-07 / 6.5e-06 / 3.2e-07 | | Prefill KV, max abs diff (self K/V, cross K/V) | 8.1e-06 / 4.5e-06 / 2.4e-06 / 1.2e-06 | | Greedy codes, 60 steps | identical, all 3 languages | | Waveform correlation vs torch | 0.9999999917 / 1.0000000000 / 1.0000000000 | | Waveform RMSE vs torch | 1.3e-08 / 5.0e-08 / 1.5e-08 | No quantised variants are published: the quantised graphs were not verified, and phoonnx does not ship unverified exports. ## Notice Upstream model: `ai4bharat/indic-parler-tts` by [AI4Bharat](https://ai4bharat.iitm.ac.in/), built on [`huggingface/parler-tts`](https://github.com/huggingface/parler-tts) by Yoach Lacombe, Vaibhav Srivastav and Sanchit Gandhi. Apache-2.0, preserved from upstream. The DAC codec is Descript's, as vendored by upstream. The weights are unmodified: the export wraps the PyTorch modules and traces them. No surgery, no quantisation.