| --- |
| license: other |
| language: |
| - en |
| tags: |
| - gguf |
| - tts |
| - neutts |
| - text-to-speech |
| base_model: neuphonic/neutts-nano |
| --- |
| |
| # NeuTTS Nano GGUF |
|
|
| GGUF conversions of [Neuphonic's NeuTTS Nano](https://huggingface.co/neuphonic/neutts-nano) β a compact (~228 M params, ~117 M active) on-device TTS LM with instant voice cloning. |
|
|
| The stack is split so each runtime owns the part it's good at: |
|
|
| - **LLM-part** β text-and-reference-codes to speech token IDs, stock `llama` arch, runs in **llama.cpp**. |
| - **Codec-part** β speech token IDs to/from 24 kHz PCM via [NeuCodec](https://huggingface.co/neuphonic/neucodec), runs in [codec.cpp](https://github.com/mybigday/codec.cpp). |
|
|
| ## Files |
|
|
| ### LLM-part (llama arch, vocab 194256, RoPE ΞΈ=500000 + linear scaling x32) |
|
|
| `neutts-nano-<quant>.gguf` |
|
|
| | File | Size | |
| | ----------------- | ------- | |
| | `neutts-nano-f32.gguf` | 925 MB | |
| | `neutts-nano-f16.gguf` | 467 MB | |
| | `neutts-nano-bf16.gguf` | 467 MB | |
| | `neutts-nano-q8_0.gguf` | 253 MB | |
| | `neutts-nano-q6_k.gguf` | 245 MB | |
| | `neutts-nano-q5_k_m.gguf` | 217 MB | |
| | `neutts-nano-q5_k_s.gguf` | 214 MB | |
| | `neutts-nano-q5_1.gguf` | 216 MB | |
| | `neutts-nano-q5_0.gguf` | 209 MB | |
| | `neutts-nano-q4_k_m.gguf` | 210 MB | |
| | `neutts-nano-q4_k_s.gguf` | 205 MB | |
| | `neutts-nano-q4_1.gguf` | 202 MB | |
| | `neutts-nano-q4_0.gguf` | 194 MB | |
| | `neutts-nano-q3_k_l.gguf` | 200 MB | |
| | `neutts-nano-q3_k_m.gguf` | 196 MB | |
| | `neutts-nano-q3_k_s.gguf` | 190 MB | |
| | `neutts-nano-q2_k.gguf` | 190 MB | |
|
|
| ### Codec-part (NeuCodec, `neucodec` arch, 65536 codebook, 24 kHz, 80 Hz token rate) |
|
|
| `codec[-<quant>].gguf` |
|
|
| | File | Size | |
| | ------------------- | ------- | |
| | `codec-f32.gguf` | 928 MB | |
| | `codec-f16.gguf` | 465 MB | |
| | `codec-q8_0.gguf` | 326 MB | |
| | `codec-q5_k_m.gguf` | 270 MB | |
| | `codec-q4_k_m.gguf` | 252 MB | |
|
|
| ## Prompt format |
|
|
| The reference flow (from `neuphonic/neutts`): |
|
|
| ``` |
| user: Convert the text to speech:<|TEXT_PROMPT_START|>{ref_phones} {input_phones}<|TEXT_PROMPT_END|> |
| assistant:<|SPEECH_GENERATION_START|>{ref_codes_as_speech_tokens} |
| ``` |
|
|
| - `{ref_phones}` / `{input_phones}` β IPA-phonemized text (espeak-ng / phonemizer). |
| - `{ref_codes_as_speech_tokens}` β concatenation of `<|speech_N|>` strings for each NeuCodec token of a reference clip. |
| - Generation continues with more `<|speech_N|>` tokens until `<|SPEECH_GENERATION_END|>`. |
|
|
| Speech-token vocab range: `<|speech_0|>` = id 128262 β¦ `<|speech_65535|>` = id 193797. |
| Stop token: `<|SPEECH_GENERATION_END|>` = id 128261. |
|
|
| ## Notes |
|
|
| - Reference encoding (audio β speech tokens) needs the codec encode path. |
| - Phonemization isn't part of the GGUF β apply espeak-ng (or equivalent) on the host before tokenizing prompts; raw English text will be off-distribution. |
| - Source weights: `neuphonic/neutts-nano` (LLM) and `neuphonic/neucodec` (codec). |
|
|