---
license: apache-2.0
language:
- en
tags:
- text-to-speech
- vits
- ax650
- npu
- onnx
pipeline_tag: text-to-speech
---
# Inflect-Nano-v2 for AX650 NPU
High-quality English TTS (VITS architecture) optimized for **AX650 NPU3**, running entirely on-chip with zero PyTorch dependency at inference time.
- **Encoder + Decoder**: AX650 NPU (AXMODEL, U16 precision)
- **Duration Predictor**: ONNX Runtime (onnxruntime)
- **Embedding**: NumPy lookup
- **Real-time Factor**: ~0.04x (24× faster than real-time)
## 🎧 Audio Samples
*All samples generated on AX650 NPU, 24kHz.*
**"Hello and welcome to Inflect Nano V2 text to speech."** (3.3s, RTF 0.06x)
**"The weather today is sunny with a gentle breeze."** (2.4s, RTF 0.07x)
**"Once upon a time, there lived a curious robot who loved to sing."** (3.7s, RTF 0.05x)
**"Neural voice synthesis is the next frontier of artificial intelligence."** (3.9s, RTF 0.05x)
**"This audio was generated entirely on the AX650 NPU chip."** (4.4s, RTF 0.04x)
## 🚀 Quick Start
```bash
bash setup.sh
bash run.sh "Hello world." output.wav
```
## 📡 API
```bash
# CLI
python -m inflect_tts_sdk --text "Hello" --output out.wav
# OpenAI-compatible Server
python -m inflect_tts_sdk.server --port 8000
```
## 📊 Performance
| Component | Time | Device |
|-----------|------|--------|
| Encoder | 2.4 ms | NPU |
| Decoder | 89.6 ms | NPU |
| DP + Align | 14.7 ms | CPU |
| **Total RTF** | **~0.04x** | 24× real-time |
*Measured on AX650 NPU3, 105 tokens → 2.7s audio.*
## 🧩 Model
| Component | Size | Input | Output |
|-----------|------|-------|--------|
| Encoder | 1.0 MB | x_emb [1,72,200] | m_p, logs_p, x, x_mask |
| Decoder | 5.6 MB | z_p [1,128,500] | waveform [1,1,128000] |
- Quantization: Encoder Conv U16 / MatMul INT8, Decoder Conv+ConvTranspose U16
- Parameters: 4.0M
- Sample rate: 24 kHz
- Source: [owensong/Inflect-Nano-v2](https://hf-mirror.com/owensong/Inflect-Nano-v2)
- Compiled: Pulsar2 6.0-lite, AX650 NPU3
## 📁 File Structure
- `models/` — AXMODEL, dp.onnx, emb_weight.npy
- `python/inflect_tts_sdk/` — Python SDK (CLI + OpenAI Server)
- `cpp/inflect_tts_sdk/` — C++ SDK (CLI + Server)
- `samples/` — audio samples
## License & Attribution(许可与署名)
本仓库是 [owensong/Inflect-Nano-v2](https://huggingface.co/owensong/Inflect-Nano-v2)
(作者 Owen Song,Apache-2.0)的 NPU 转换与再分发。原始 Inflect 代码与权重以
Apache License 2.0 发布;本仓库的转换产物(AXMODEL、SDK)同样以 Apache-2.0 许可提供。
- 许可全文:`LICENSE`(Apache License 2.0)
- 署名与修改说明:`NOTICE.md`
- 第三方组件声明:`THIRD_PARTY_NOTICES.md`、`third_party/`