| --- |
| license: mit |
| language: |
| - en |
| pipeline_tag: text-to-speech |
| tags: |
| - piper |
| - vits |
| - text-to-speech |
| - synaptics |
| - torq |
| - npu |
| --- |
| |
| # Piper TTS for Synaptics Torq (SL2619) |
|
|
| Piper (VITS) text-to-speech, voice `en_US-libritts_r-medium` (904 speakers, |
| 22.05 kHz), **split across CPU and NPU** for the Synaptics SL2619 board. |
|
|
| The VITS graph is cut where the per-phoneme durations are ceiled and summed — |
| the point at which the output length becomes exact: |
|
|
| ``` |
| text --[espeak]--> phoneme ids --> [partA] (CPU, onnxruntime) --> z [1,192,F], g [1,512,1] |
| z,g --> [partB] (NPU, bf16 vmfb) --> audio [F*256] |
| ``` |
|
|
| partA holds 85% of the *nodes* (small shape/attention ops) but partB — the |
| HiFi-GAN vocoder — holds 82% of the *time* and is pure convolution, which is |
| what the NPU accelerates. Because partA yields the exact frame count `F`, the |
| correct static vocoder window is known before the vocoder runs. |
|
|
| Measured on the SL2619 board: **2.2× real time** end to end, versus 1.0× for the |
| same model run entirely on the CPU. The bf16 NPU vocoder matches the fp32 CPU |
| vocoder at **39.6 dB SNR** (correlation 0.99998). |
|
|
| ## Contents |
|
|
| | path | what it is | |
| |---|---| |
| | `onnx/partA.onnx` | Text encoder + duration predictor. Runs on the CPU under onnxruntime. | |
| | `onnx/partB_static_{1,2,4,6,8}s.bf16io.onnx` | The vocoder, statically shaped per window, bf16 I/O — the source the VMFBs were compiled from. | |
| | `onnx/en_US-libritts_r-medium.onnx` | The original monolithic Piper voice, for reference. | |
| | `vmfb/partB_static_{1,2,4,6,8}s.vmfb` | The vocoder compiled for the Torq NPU (NSS-only, bf16), one per window. | |
| | `tflite/partB_static_4s.{int8,int16x8}.tflite` | Quantized vocoder conversions (int8 was slower than bf16 on this target; kept for reference). | |
| | `voice/en_US-libritts_r-medium.onnx.json` | Voice config, including the phoneme→id map. | |
| | `espeak/phonemizerd` | Small resident espeak-ng phonemizer daemon (aarch64) mirroring libpiper's phonemization. | |
| | `espeak/espeak-ng-data.tar.gz` | espeak-ng dictionaries the daemon needs. | |
|
|
| The vocoder ships as five separate VMFBs because the NPU model is statically |
| shaped; each covers 1, 2, 4, 6 or 8 seconds of audio. Per sentence you pick the |
| smallest window that fits, edge-pad the latent up to it, and trim the output back |
| to `F × 256` samples. |
|
|
| ## Usage |
|
|
| These files are consumed by the `piper_tts` demo in |
| [synaptics-torq/torq-examples](https://github.com/synaptics-torq/torq-examples), |
| which writes a `.wav` and plays it on the board's speaker: |
|
|
| ```sh |
| cd piper_tts |
| pip install -r requirements.txt |
| cd .. && python setup_demos.py piper_tts |
| cd piper_tts && python src/infer.py --text "Hello from the Synaptics board." |
| ``` |
|
|
| The demo reads each window's frame width from the VMFB signature rather than |
| hardcoding it, so recompiling with a different set of windows needs no code |
| change. |
|
|
| ## Licensing |
|
|
| The Piper voice and models are MIT. `espeak/phonemizerd` links espeak-ng |
| (**GPLv3**) statically; its source ships with the demo at |
| `piper_tts/piper_core/phonemizerd.c` in torq-examples, with the build command in |
| its header comment. |
|
|