File size: 3,065 Bytes
4e75a38
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
---
license: cc-by-4.0
language:
- en
library_name: mnn
base_model:
- kyutai/pocket-tts
- IgnitiveLabs/PocketTTS-ONNX
pipeline_tag: text-to-speech
tags:
- tts
- voice-cloning
- mnn
- onnx
- int8
- pocket-tts
---

# Pocket TTS MNN (English)

MNN conversion of [IgnitiveLabs/PocketTTS-ONNX](https://huggingface.co/IgnitiveLabs/PocketTTS-ONNX), itself an ONNX export of Kyutai's [Pocket TTS](https://huggingface.co/kyutai/pocket-tts) (`english_2026-04`).

Weights use **INT8 weight quantization** via `MNNConvert --weightQuantBits 8` (text conditioner kept FP32). Unsupported `ONNX::IsNaN` ops were rewritten to `Not(Equal(x, x))` before conversion.

## Files

```
mnn/
  flow_lm_main_w8.mnn          # ~73 MB
  flow_lm_flow_w8.mnn          # ~9.5 MB
  mimi_decoder_w8.mnn          # ~10 MB
  mimi_encoder_w8.mnn          # ~9.5 MB (converted; runtime uses ONNX fallback)
  text_conditioner_fp32.mnn    # ~16 MB (preferred)
  text_conditioner_w8.mnn      # ~16 MB
flow_lm_main_int8.onnx         # empty-seq priming fallback
mimi_encoder.onnx              # voice encode fallback
bundle.json
tokenizer.model
bos_before_voice.npy
pocket_tts_mnn.py
pocket_tts_onnx.py
reference_sample.wav
```

## Hybrid runtime

MNN runs the generation hot path (frame loop, flow net, decoder, text conditioner).

ONNX is still required for:

1. **Empty-sequence priming** (`seq_len=0`) — not supported by MNN on `flow_lm_main`
2. **`mimi_encoder`** — MNN fails streaming Concat shape inference

## Quick start

```bash
pip install MNN onnxruntime numpy soundfile sentencepiece scipy huggingface_hub safetensors

huggingface-cli download developerabu/pocket-tts-mnn --local-dir pocket-tts-mnn
cd pocket-tts-mnn
python generate_mnn.py "Hello, this is a test." reference_sample.wav out.wav
```

Or in Python:

```python
from pocket_tts_mnn import PocketTTSMnn

tts = PocketTTSMnn(models_dir=".")
audio = tts.generate("Hello, this is a test.", voice="reference_sample.wav")
tts.save_audio(audio, "out.wav")
```

## Benchmark (Apple M1, CPU, 5 runs)

| Backend | Medium RTFx | Stream TTFB |
|---------|-------------|-------------|
| ONNX INT8 | ~4.1× | ~100 ms |
| MNN INT8 (this repo) | ~2.4× | ~123 ms |

MNN is mainly targeted at mobile; desktop Python numbers above are not the expected mobile ranking.

## Attribution

- Original model: [Kyutai Labs](https://kyutai.org/) — [kyutai/pocket-tts](https://huggingface.co/kyutai/pocket-tts) (CC-BY-4.0)
- ONNX export used as source: [IgnitiveLabs/PocketTTS-ONNX](https://huggingface.co/IgnitiveLabs/PocketTTS-ONNX)
- ONNX inference reference: [KevinAHM/pocket-tts-onnx](https://huggingface.co/KevinAHM/pocket-tts-onnx)
- MNN conversion / packaging: developerabu

## License

Original model weights: **CC-BY-4.0** (provide attribution to Kyutai Labs).  
Conversion scripts and packaging in this repo: MIT.

## Acceptable use

Same restrictions as Kyutai / IgnitiveLabs releases: no unauthorized voice cloning, no deceptive synthetic audio presented as real, no harmful/unlawful content, and disclose AI-generated speech to end users.