Luigi's picture
Add PrimeTTS v2-Stream-Clean: token-level input + v2-clean audio (RIGHT=16)
9994aaf verified
|
Raw
History Blame Contribute Delete
1.44 kB

PrimeTTS v2-Stream-Clean — token-level streaming, v2-clean audio

The definitive streaming variant: token-level band-attention encoder (input streams incrementally) + v2's non-causal clean vocoder (no parasite noise). Reconciles token-level input streaming with clean audio — the causal v2-Stream sacrificed quality unnecessarily; this doesn't.

  • v2streamclean_enc.onnx — text (x,tone,lang,x_lengths,noise_scale,length_scale) → z[1,192,T]. Band encoder, token-level. Run once per phrase.
  • v2streamclean_dec.onnx — z[1,192,Tc] → wav[1,1,Tc·256]. Clean non-causal vocoder. Run per chunk, overlap-save.
  • onnx_stream.py — reference runner (uses the right params).

Streaming params: chunk = 24, left = 64, RIGHT = 16 (the clean non-causal vocoder needs 16 future frames for bit-exact chunking; the causal one used 4). 16 kHz, zh-TW + English.

from onnx_stream import StreamingTTS   # RIGHT=16 baked in
tts = StreamingTTS("v2streamclean_enc.onnx", "v2streamclean_dec.onnx")
z = tts.encode(phone_ids, tone_ids, lang_ids)      # once
for pcm in tts.stream(z): play(pcm)                # per 24-frame chunk, clean audio

Frontend (text→ids): g2pw bopomofo + g2p_en, 88 syms/6 tones/2 langs, add_blank. sherpa-onnx: OfflineTtsMbistftStreamModel(enc, dec, num_threads=2, right_lookahead=16).

License: Apache-2.0 · part of Luigi/PrimeTTS.