Jarbas's picture
ONNX export of ai4bharat/indic-parler-tts for phoonnx
aed7256 verified
|
Raw
History Blame Contribute Delete
3.01 kB
metadata
license: apache-2.0
language:
  - as
  - bn
  - brx
  - doi
  - en
  - gu
  - hi
  - kn
  - kok
  - mai
  - ml
  - mni
  - mr
  - ne
  - or
  - pa
  - sa
  - sat
  - sd
  - ta
  - te
  - ur
library_name: onnx
pipeline_tag: text-to-speech
tags:
  - onnx
  - text-to-speech
  - tts
  - indic
  - parler-tts
  - phoonnx
base_model: ai4bharat/indic-parler-tts

phoonnx — Indic Parler-TTS (ONNX)

ONNX export of ai4bharat/indic-parler-tts for the phoonnx indic_parler engine.

Indic Parler-TTS speaks 20 Indic languages and English. You select the voice with a natural-language description ("Rohit's voice is clear and expressive..."), not with a speaker id or a reference clip.

Graphs

File Size Purpose
text_encoder.onnx 1.4 GB Flan-T5 encoder over the voice description
decoder_prefill.onnx 2.1 GB First AR step; emits self-attention and cross-attention KV
decoder_decode.onnx 1.3 GB Later AR steps; consumes both caches, emits self-attention KV only
dac_decoder.onnx 217 MB DAC 44.1 kHz codec decoder (9 codebooks)
tokenizer.json prompt tokenizer (the text you want spoken)
description_tokenizer.json description tokenizer (Flan-T5 vocabulary)
config.json self-describing phoonnx config (engine: indic_parler)

The two tokenizers are different vocabularies. Upstream is explicit about this: one tokenizer for the prompt, one for the description.

Architecture

description --> Flan-T5 encoder --> encoder states --> cross-attention (all 24 layers)
prompt      --> embed_prompts   --> prepended to the decoder input embeddings
decoder     --> 9 delayed DAC codebooks --> DAC decoder --> 44.1 kHz mono

Cross-attention keys and values do not change while decoding, so decoder_prefill computes them once and decoder_decode reads them back unchanged.

Parity

Every graph is float32 and was checked against the PyTorch model (parler_tts.ParlerTTSForConditionalGeneration, eager attention) on ser9 CPU:

Check Result
Text encoder, max abs diff (hi / en / ta) 3.9e-07 / 6.5e-06 / 3.2e-07
Prefill KV, max abs diff (self K/V, cross K/V) 8.1e-06 / 4.5e-06 / 2.4e-06 / 1.2e-06
Greedy codes, 60 steps identical, all 3 languages
Waveform correlation vs torch 0.9999999917 / 1.0000000000 / 1.0000000000
Waveform RMSE vs torch 1.3e-08 / 5.0e-08 / 1.5e-08

No quantised variants are published: the quantised graphs were not verified, and phoonnx does not ship unverified exports.

Notice

Upstream model: ai4bharat/indic-parler-tts by AI4Bharat, built on huggingface/parler-tts by Yoach Lacombe, Vaibhav Srivastav and Sanchit Gandhi. Apache-2.0, preserved from upstream. The DAC codec is Descript's, as vendored by upstream.

The weights are unmodified: the export wraps the PyTorch modules and traces them. No surgery, no quantisation.