Piper TTS for Synaptics Torq (SL2619)

Piper (VITS) text-to-speech, voice en_US-libritts_r-medium (904 speakers, 22.05 kHz), split across CPU and NPU for the Synaptics SL2619 board.

The VITS graph is cut where the per-phoneme durations are ceiled and summed — the point at which the output length becomes exact:

text --[espeak]--> phoneme ids --> [partA]  (CPU, onnxruntime)  --> z [1,192,F], g [1,512,1]
                                              z,g --> [partB]   (NPU, bf16 vmfb) --> audio [F*256]

partA holds 85% of the nodes (small shape/attention ops) but partB — the HiFi-GAN vocoder — holds 82% of the time and is pure convolution, which is what the NPU accelerates. Because partA yields the exact frame count F, the correct static vocoder window is known before the vocoder runs.

Measured on the SL2619 board: 2.2× real time end to end, versus 1.0× for the same model run entirely on the CPU. The bf16 NPU vocoder matches the fp32 CPU vocoder at 39.6 dB SNR (correlation 0.99998).

Contents

path what it is
onnx/partA.onnx Text encoder + duration predictor. Runs on the CPU under onnxruntime.
onnx/partB_static_{1,2,4,6,8}s.bf16io.onnx The vocoder, statically shaped per window, bf16 I/O — the source the VMFBs were compiled from.
onnx/en_US-libritts_r-medium.onnx The original monolithic Piper voice, for reference.
vmfb/partB_static_{1,2,4,6,8}s.vmfb The vocoder compiled for the Torq NPU (NSS-only, bf16), one per window.
tflite/partB_static_4s.{int8,int16x8}.tflite Quantized vocoder conversions (int8 was slower than bf16 on this target; kept for reference).
voice/en_US-libritts_r-medium.onnx.json Voice config, including the phoneme→id map.
espeak/phonemizerd Small resident espeak-ng phonemizer daemon (aarch64) mirroring libpiper's phonemization.
espeak/espeak-ng-data.tar.gz espeak-ng dictionaries the daemon needs.

The vocoder ships as five separate VMFBs because the NPU model is statically shaped; each covers 1, 2, 4, 6 or 8 seconds of audio. Per sentence you pick the smallest window that fits, edge-pad the latent up to it, and trim the output back to F × 256 samples.

Usage

These files are consumed by the piper_tts demo in synaptics-torq/torq-examples, which writes a .wav and plays it on the board's speaker:

cd piper_tts
pip install -r requirements.txt
cd .. && python setup_demos.py piper_tts
cd piper_tts && python src/infer.py --text "Hello from the Synaptics board."

The demo reads each window's frame width from the VMFB signature rather than hardcoding it, so recompiling with a different set of windows needs no code change.

Licensing

The Piper voice and models are MIT. espeak/phonemizerd links espeak-ng (GPLv3) statically; its source ships with the demo at piper_tts/piper_core/phonemizerd.c in torq-examples, with the build command in its header comment.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support