Instructions to use Masterx/chatterbox-nano-ONNX with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Chatterbox
How to use Masterx/chatterbox-nano-ONNX with Chatterbox:
# pip install chatterbox-tts import torchaudio as ta from chatterbox.tts import ChatterboxTTS model = ChatterboxTTS.from_pretrained(device="cuda") text = "Ezreal and Jinx teamed up with Ahri, Yasuo, and Teemo to take down the enemy's Nexus in an epic late-game pentakill." wav = model.generate(text) ta.save("test-1.wav", wav, model.sr) # If you want to synthesize with a different voice, specify the audio prompt AUDIO_PROMPT_PATH="YOUR_FILE.wav" wav = model.generate(text, audio_prompt_path=AUDIO_PROMPT_PATH) ta.save("test-2.wav", wav, model.sr) - Notebooks
- Google Colab
- Kaggle
Chatterbox Nano โ ONNX
ONNX export of Resemble AI's Chatterbox Nano
(t3_nano_v1.safetensors: a GPT-2-small T3 backbone, ~110M) with the distilled one-step S3Gen
(meanflow) decoder. Upstream: https://github.com/resemble-ai/chatterbox.
The graphs follow exactly the file naming, graph layout and input/output contract of Resemble AI's
official chatterbox-turbo-ONNX (Nano is
Turbo's architecture at GPT-2-small scale and shares its tokenizer, s3gen_meanflow decoder and
paralinguistic tags [laugh], [chuckle], [cough], โฆ), so a Turbo runtime runs Nano by swapping
the directory. The conditional_decoder_* graphs are Turbo's official decoder graphs, unchanged:
Nano's s3gen_meanflow.safetensors is byte-identical (same sha256) to Turbo's.
Files
| File | Bytes |
|---|---|
LICENSE |
1,068 |
NOTICE.md |
396 |
config.json |
1,234 |
default_voice.wav |
714,320 |
generation_config.json |
55 |
onnx/conditional_decoder_q4.onnx |
2,179,022 |
onnx/conditional_decoder_q4.onnx_data |
246,397,384 |
onnx/conditional_decoder_q4f16.onnx |
2,394,210 |
onnx/conditional_decoder_q4f16.onnx_data |
162,996,136 |
onnx/embed_tokens_q4.onnx |
3,197 |
onnx/embed_tokens_q4.onnx_data |
27,964,788 |
onnx/language_model_q4.onnx |
108,402 |
onnx/language_model_q4.onnx_data |
83,330,000 |
onnx/language_model_q4f16.onnx |
109,622 |
onnx/language_model_q4f16.onnx_data |
64,861,690 |
onnx/speech_encoder_q4.onnx |
1,105,930 |
onnx/speech_encoder_q4.onnx_data |
130,857,924 |
preprocessor_config.json |
130 |
tokenizer.json |
3,562,272 |
tokenizer_config.json |
414 |
Graph inputs / outputs
| Graph | Inputs | Outputs |
|---|---|---|
conditional_decoder_q4.onnx |
speech_tokens int64[batch_size, num_speech_tokens]speaker_embeddings float32[batch_size, 192]speaker_features float32[batch_size, feature_dim, 80] |
waveform float32[batch_size, num_samples] |
conditional_decoder_q4f16.onnx |
speech_tokens int64[batch_size, num_speech_tokens]speaker_embeddings float32[batch_size, 192]speaker_features float32[batch_size, feature_dim, 80] |
waveform float32[batch_size, num_samples] |
embed_tokens_q4.onnx |
input_ids int64[batch_size, sequence_length] |
inputs_embeds float32[batch_size, sequence_length, 768] |
language_model_q4.onnx |
inputs_embeds float32[batch_size, sequence_length, 768]attention_mask int64[batch_size, total_sequence_length]position_ids int64[batch_size, sequence_length]past_key_values.{0..11}.{key,value} float32[batch_size, 12, past_sequence_length, 64] |
logits float32[batch_size, sequence_length, 6563]present.{0..11}.{key,value} |
language_model_q4f16.onnx |
inputs_embeds float32[batch_size, sequence_length, 768]attention_mask int64[batch_size, total_sequence_length]position_ids int64[batch_size, sequence_length]past_key_values.{0..11}.{key,value} float16[batch_size, 12, past_sequence_length, 64] |
logits float32[batch_size, sequence_length, 6563]present.{0..11}.{key,value} |
speech_encoder_q4.onnx |
audio_values float32[batch_size, num_samples] |
audio_features float32[batch_size, sequence_length, 768]audio_tokens int64[batch_size, audio_sequence_length]speaker_embeddings float32[1, 192]speaker_features float32[batch_size, feature_dim, 80] |
How it was exported
- speech_encoder / embed_tokens:
torch.onnx.export(opset 20, TorchScript exporter) of the onnx-community Chatterbox conversion recipe'sPrepareConditionalsModel, adapted to Turbo/Nano conditioning exactly as upstreamtts_turbo.prepare_conditionalsdoes it: every 16 kHz view (x-vector, flow prompt tokens, 375 T3 prompt tokens) is derived from the 10 s-cropped 24 kHz reference, no perceiver, no positional embedding, and CAMPPlus's original BatchNorm dense layer written as explicit arithmetic (bit-exact to PyTorch).embed_tokensmaps the two trailing50256ids to the start-of-speech token, as in Turbo's graph. Thenonnxslim. - language_model: GPT-2 small +
speech_head, built node-for-node in the layout of Turbo's officiallanguage_model*.onnx(GroupQueryAttention with KV cache,LayerNormalization, tanh-GELU,wpeGather onposition_ids; opset 21 +com.microsoft), weights copied fromt3_nano_v1.safetensors. - Quantization (Turbo's official style):
MatMulNBits4-bit, block 32, asymmetric (zero points), noaccuracy_level(int8 activation quantization breaks GPT-2's outlier dimensions); embedding tables asGatherBlockQuantized4-bit.q4f16additionally stores weights and KV cache in fp16. - conditional_decoder_q4 / _q4f16: Turbo's official graphs, copied unchanged (identical
s3gen_meanflow.safetensors).
Validation
Parity vs PyTorch (reference: default_voice.wav, 7.4 s)
| Graph | Output | max abs diff | cosine |
|---|---|---|---|
| speech_encoder (fp32, before quantization) | audio_features | 0 | 1.0000000 |
| audio_tokens | 186/186 identical | ||
speaker_embeddings vs upstream s3gen.embed_ref |
6.4e-6 | 1.0000000 | |
| speaker_features | 8.2e-2 | 0.9999999 | |
| speech_encoder_q4 | audio_features | 3.25 | 0.9642 |
| audio_tokens | 96.8 % identical | ||
| speaker_embeddings | 7.8e-2 | 0.99975 | |
| embed_tokens (fp32) | inputs_embeds | 0 | 1.0000000 |
| embed_tokens_q4 | inputs_embeds | 0.115 | 0.9964 |
| language_model (fp32, before quantization) | logits, 41 teacher-forced steps | 1.7e-4 | 1.0000000 (top-1 100 %) |
| language_model_q4 | logits, prefill | 7.0 | 0.952 |
| language_model_q4f16 | logits, prefill | 7.3 | 0.951 |
4-bit logit drift is of the same kind as Turbo's official q4 graphs; it does not show up in intelligibility or speaker similarity (below).
End-to-end (10 English sentences, default_voice.wav reference)
Decoding: greedy argmax with repetition penalty 1.2 (how WinSTT runs it).
WER from NVIDIA Parakeet TDT 0.6B v3 (ONNX) transcripts, normalized (lowercase, no punctuation).
Speaker similarity: cosine between Chatterbox VoiceEncoder embeddings of the output and the
reference.
| Graph set | WER | Speaker similarity |
|---|---|---|
| fp32 speech_encoder / embed / LM + decoder_q4 | 1.68 % | 0.928 |
| q4 (speech_encoder_q4, embed_tokens_q4, language_model_q4, conditional_decoder_q4) | 1.68 % | 0.927 |
| q4f16 (โฆ language_model_q4f16, conditional_decoder_q4f16) | 1.68 % | 0.926 |
Usage
Identical to chatterbox-turbo-ONNX:
run speech_encoder once on the 24 kHz reference, prepend its audio_features to embed_tokens
of the tokenized text (+ two 50256 start-of-speech ids), decode speech tokens with the KV-cached
language_model (upstream samples with temperature 0.8, top-k 1000, top-p 0.95 and
repetition penalty 1.2; greedy + penalty also works), prepend
audio_tokens, and run conditional_decoder with speaker_embeddings / speaker_features to get
24 kHz audio. English only.
Watermark note
Resemble AI's Python package embeds an imperceptible Perth watermark into every generated waveform. These ONNX graphs do not contain the watermarker: their output is the raw vocoder waveform. If you deploy this model, Resemble AI asks that you keep watermarking generated audio (e.g. run Perth on the output) and use it responsibly.
License and credit
MIT, inherited from Resemble AI's Chatterbox (see LICENSE, NOTICE.md). All credit for the model
goes to Resemble AI; this is only a format conversion.
- Downloads last month
- 3
Model tree for Masterx/chatterbox-nano-ONNX
Base model
ResembleAI/chatterbox-nano