Inflect-Nano-v2 / README.md
mlboydaisuke's picture
Inflect-Nano-v2 LiteRT: dynamic-length fp32/fp16 graphs, exact streaming, Pi 5 bench, say.py
a8864da verified
|
Raw
History Blame Contribute Delete
8.26 kB
---
license: apache-2.0
library_name: litert
pipeline_tag: text-to-speech
base_model: owensong/Inflect-Nano-v2
language:
- en
tags:
- litert
- tflite
- tts
- text-to-speech
- on-device
- streaming
- raspberry-pi
- vits
- inflect
---
# Inflect-Nano-v2 — LiteRT, dynamic length + exact streaming
[Inflect-Nano-v2](https://huggingface.co/owensong/Inflect-Nano-v2) (4.0M params,
VITS-family end-to-end TTS, English, fixed male voice, 24 kHz, Apache-2.0)
converted to **LiteRT CPU/XNNPACK** graphs with a **dynamic sequence length**
and **exact intra-sentence streaming** (overlap-discard chunking reproduces the
full decode at corr 1.000000). Built for small-CPU targets (Raspberry Pi
class): **RTF 0.111 measured on a Raspberry Pi 5**, time-to-first-audio
~220 ms — Piper-class speed.
![Inflect-Nano-v2 — LiteRT on-device output](samples/sample.png)
**Listen** (golden sentence, same inputs/noise):
[`samples/litert_golden.wav`](https://huggingface.co/litert-community/Inflect-Nano-v2/resolve/main/samples/litert_golden.wav)
— this port (fp32) ·
[`samples/ref_torch.wav`](https://huggingface.co/litert-community/Inflect-Nano-v2/resolve/main/samples/ref_torch.wav)
— the PyTorch reference (waveform corr 1.000000).
The upstream repo ships a PyTorch checkpoint + runtime. This port re-authors
the VITS inference graph in TF (weights loaded from `model.pth`) and converts
with the official `TFLiteConverter`, keeping both sequence axes dynamic.
## Graphs
| Graph | Inputs | Outputs | fp32 | fp16 |
| ----- | ------ | ------- | ---- | ---- |
| `inflect_text_encoder.tflite` | tokens [1,N] int32 | m_p [1,N,128], logs_p [1,N,128], logw [1,N,1] | 3.5 MB | 1.8 MB |
| `inflect_decoder.tflite` | z_p [1,T,128] | wav [1,256·T] @ 24 kHz | 12.6 MB | 6.4 MB |
Host glue: `durations = ceil(exp(logw)/speed)`; expand `m_p`/`logs_p` with
`np.repeat`; `z_p = m_p + randn·exp(logs_p)·variation` (noise generated
host-side for reproducibility); decoder → waveform. use_sdp=false in this
checkpoint, so the duration predictor is deterministic convs — no spline flows
anywhere.
## Measured on a real Raspberry Pi 5 (2026-08-06)
Pi 5 8 GB, Raspberry Pi OS 64-bit, Python 3.13.5, `ai-edge-litert` 2.1.6,
4 threads. `vcgencmd get_throttled` = 0x0 before/after each fp32 run.
| Sentence | N tokens | audio | encoder | decoder | sentence | RTF | TTFA |
| -------- | -------- | ----- | ------- | ------- | -------- | --- | ---- |
| [0] | 75 | 2.07 s | 2.7 ms | 220.4 ms | 223.1 ms | 0.108 | 218.8 ms |
| [1] | 121 | 3.21 s | 4.9 ms | 351.1 ms | 356.0 ms | 0.111 | 221.2 ms |
| [2] | 269 | 7.19 s | 22.2 ms | 788.3 ms | 810.5 ms | 0.113 | 239.0 ms |
Overall **RTF 0.111** (fp32), streaming exact (corr 1.000000) and full-waveform
ref-corr **1.000000** — Piper-class speed (Piper lessac-low baseline on this
device class: RTF 0.10, 147 ms/phrase). ⚠ fp16 is speed-identical but showed a
real quality break on one bench sentence (ref-corr 0.388) — the flow layers are
fp16-sensitive; **deploy fp32**.
### GPU (v3dv WebGPU) status — measured on the Pi 5, 2026-08-06
With Mesa built from git (v3dv Vulkan 1.3, driver 26.2.99,
`V3D_WEBGPU_OVERRIDE=1`): the static-chunk decoder
(`inflect_decoder_static228.tflite`) **compiles and runs fully accelerated**
(`is_fully_accelerated=True`) with output corr 0.9904 vs CPU (fp16-class
divergence; the flow layers are precision-sensitive — listen before adopting).
Dynamic graphs do not compile (static shapes required). CPU remains the
recommended deployment (faster on this board).
## Verification (vs. PyTorch reference, same inputs/noise)
| Check | Result |
| ----- | ------ |
| text encoder (m_p / logs_p / logw) | maxerr ≤ 2.3e-6 |
| decoder wav, golden sentence | **corr 1.000000**, maxerr 2.6e-5 |
| dynamic lengths | N = 49 / 165 / 217, T = 134 / 430 / 527 on the same graphs |
| streaming vs full decode | **corr 1.000000** (maxerr ≤ 1e-6) |
| fp16 decoder | corr 0.999879 (but see the fp16 warning above) |
## Speed (Mac M-series, 4 threads, XNNPACK)
| Sentence | N | T | audio | encoder | decoder | RTF |
| -------- | - | - | ----- | ------- | ------- | --- |
| golden | 165 | 430 | 4.59 s | 3 ms | 88 ms | **0.020** |
| short | 49 | 134 | 1.43 s | 1 ms | 28 ms | 0.020 |
| long | 217 | 527 | 5.62 s | 3 ms | 99 ms | 0.018 |
In a `python:3.12-slim` **linux/arm64** container (same aarch64
`ai-edge-litert` 2.1.6 wheel the Pi uses): waveform corr **1.000000** vs the
Mac output, streaming corr 1.000000.
## Quickstart
```bash
hf download litert-community/Inflect-Nano-v2 --local-dir inflect-litert
cd inflect-litert
```
**One-command benchmark** (no espeak needed on the device — inputs are
pre-tokenized in `bench_inputs.npz`):
```bash
pip install numpy ai-edge-litert
python bench.py --models-dir . # fp32, 4 threads
python bench.py --models-dir . --precision fp16 --write-wavs
```
Reports encoder/decoder latency, RTF, streaming time-to-first-audio, and
waveform-correlation identity checks (full decode vs the bundled reference,
streamed vs full).
**Drop-in synthesis** (Piper replacement; mirrors the upstream `inference.py`
sentence handling — pauses, edge-fade, per-sentence seed):
```bash
pip install numpy ai-edge-litert phonemizer espeakng-loader num2words Unidecode
python say.py "Hello! How can I help you today?" --models-dir . --frontend-dir frontend -o hello.wav
```
```python
from say import InflectTTS
tts = InflectTTS(models_dir=".", frontend_dir="frontend")
for sentence, pcm in tts.stream(text): # float32 @ 24 kHz per sentence
play(pcm)
```
`frontend/` contains the upstream Apache-2.0 text frontend
([owensong/Inflect-Nano-v2](https://huggingface.co/owensong/Inflect-Nano-v2));
it phonemizes with espeak-ng (GPL-3.0, in-process — same situation as Piper's
frontend).
## Streaming (exact)
The decoder (flow + HiFi-GAN generator) is **fully convolutional with no
normalization layers**, so overlap-discard chunking is *exact*: 100-frame
chunks (+64 frames context each side) reproduce the full decode at
corr 1.000000. Time-to-first-audio is encoder + one chunk (~220 ms on the
Pi 5). This is the model to use when true sub-sentence streaming matters
(compare: KittenTTS's AdaIN statistics make its chunked mode approximate).
## Files
| File | Purpose |
| ---- | ------- |
| `inflect_{text_encoder,decoder}.tflite` | fp32 graphs (recommended) |
| `inflect_{text_encoder,decoder}_fp16.tflite` | fp16-weight variants (⚠ flow layers are fp16-sensitive — deploy fp32) |
| `inflect_decoder_static228.tflite` | static 228-frame decoder chunk (GPU-delegate experiment; CPU deployment recommended) |
| `frontend/` | upstream Apache-2.0 text frontend (phonemization + cleaners) |
| `bench.py` + `bench_inputs.npz` | one-command device benchmark (numpy + ai-edge-litert only) |
| `make_bench_inputs.py` | regenerate bench inputs (needs espeak on the host) |
| `say.py` | drop-in `say(text)` / `stream(text)` synthesis module + CLI |
| `samples/` | output samples: this port vs the PyTorch reference, same inputs |
## Conversion notes
- **litert-torch dynamic export is a dead end** (0.9.2): beyond the known
dynamic-LSTM wall, even a plain conv stack exported with `torch.export.Dim`
bakes the trace length into internal RESHAPEs and fails at any other length;
`F.embedding` doesn't lower with a symbolic axis at all. The TF/Keras →
`TFLiteConverter` path handles all of it (shape-computed reshapes, fused
dynamic LSTM).
- VITS's relative-position attention (window 4) uses pad/reshape "skew" tricks;
in TF they convert fine. (Under torch.export they generate unprovable
divisibility/stride guards.)
- `tf_keras` (Keras 2) is required: Keras 3 models leave READ_VARIABLE resource
ops in the converted graph.
## Text frontend / licensing
English phonemization via the upstream frontend = **espeak-ng (GPL-3.0)** +
num2words. Run espeak as a separate process, or swap a DeepPhonemizer-based
neural G2P for a GPL-free stack (reference implementation:
[litert-community/Kokoro-G2P-en-US](https://huggingface.co/litert-community/Kokoro-G2P-en-US);
note this model uses its own symbol table, so the G2P output must be remapped).
Model: Apache-2.0 (BigVGAN/VITS third-party notices in the upstream repo).