Chatterbox Nano โ€” ONNX

ONNX export of Resemble AI's Chatterbox Nano (t3_nano_v1.safetensors: a GPT-2-small T3 backbone, ~110M) with the distilled one-step S3Gen (meanflow) decoder. Upstream: https://github.com/resemble-ai/chatterbox.

The graphs follow exactly the file naming, graph layout and input/output contract of Resemble AI's official chatterbox-turbo-ONNX (Nano is Turbo's architecture at GPT-2-small scale and shares its tokenizer, s3gen_meanflow decoder and paralinguistic tags [laugh], [chuckle], [cough], โ€ฆ), so a Turbo runtime runs Nano by swapping the directory. The conditional_decoder_* graphs are Turbo's official decoder graphs, unchanged: Nano's s3gen_meanflow.safetensors is byte-identical (same sha256) to Turbo's.

Files

File Bytes
LICENSE 1,068
NOTICE.md 396
config.json 1,234
default_voice.wav 714,320
generation_config.json 55
onnx/conditional_decoder_q4.onnx 2,179,022
onnx/conditional_decoder_q4.onnx_data 246,397,384
onnx/conditional_decoder_q4f16.onnx 2,394,210
onnx/conditional_decoder_q4f16.onnx_data 162,996,136
onnx/embed_tokens_q4.onnx 3,197
onnx/embed_tokens_q4.onnx_data 27,964,788
onnx/language_model_q4.onnx 108,402
onnx/language_model_q4.onnx_data 83,330,000
onnx/language_model_q4f16.onnx 109,622
onnx/language_model_q4f16.onnx_data 64,861,690
onnx/speech_encoder_q4.onnx 1,105,930
onnx/speech_encoder_q4.onnx_data 130,857,924
preprocessor_config.json 130
tokenizer.json 3,562,272
tokenizer_config.json 414

Graph inputs / outputs

Graph Inputs Outputs
conditional_decoder_q4.onnx speech_tokens int64[batch_size, num_speech_tokens]
speaker_embeddings float32[batch_size, 192]
speaker_features float32[batch_size, feature_dim, 80]
waveform float32[batch_size, num_samples]
conditional_decoder_q4f16.onnx speech_tokens int64[batch_size, num_speech_tokens]
speaker_embeddings float32[batch_size, 192]
speaker_features float32[batch_size, feature_dim, 80]
waveform float32[batch_size, num_samples]
embed_tokens_q4.onnx input_ids int64[batch_size, sequence_length] inputs_embeds float32[batch_size, sequence_length, 768]
language_model_q4.onnx inputs_embeds float32[batch_size, sequence_length, 768]
attention_mask int64[batch_size, total_sequence_length]
position_ids int64[batch_size, sequence_length]
past_key_values.{0..11}.{key,value} float32[batch_size, 12, past_sequence_length, 64]
logits float32[batch_size, sequence_length, 6563]
present.{0..11}.{key,value}
language_model_q4f16.onnx inputs_embeds float32[batch_size, sequence_length, 768]
attention_mask int64[batch_size, total_sequence_length]
position_ids int64[batch_size, sequence_length]
past_key_values.{0..11}.{key,value} float16[batch_size, 12, past_sequence_length, 64]
logits float32[batch_size, sequence_length, 6563]
present.{0..11}.{key,value}
speech_encoder_q4.onnx audio_values float32[batch_size, num_samples] audio_features float32[batch_size, sequence_length, 768]
audio_tokens int64[batch_size, audio_sequence_length]
speaker_embeddings float32[1, 192]
speaker_features float32[batch_size, feature_dim, 80]

How it was exported

  • speech_encoder / embed_tokens: torch.onnx.export (opset 20, TorchScript exporter) of the onnx-community Chatterbox conversion recipe's PrepareConditionalsModel, adapted to Turbo/Nano conditioning exactly as upstream tts_turbo.prepare_conditionals does it: every 16 kHz view (x-vector, flow prompt tokens, 375 T3 prompt tokens) is derived from the 10 s-cropped 24 kHz reference, no perceiver, no positional embedding, and CAMPPlus's original BatchNorm dense layer written as explicit arithmetic (bit-exact to PyTorch). embed_tokens maps the two trailing 50256 ids to the start-of-speech token, as in Turbo's graph. Then onnxslim.
  • language_model: GPT-2 small + speech_head, built node-for-node in the layout of Turbo's official language_model*.onnx (GroupQueryAttention with KV cache, LayerNormalization, tanh-GELU, wpe Gather on position_ids; opset 21 + com.microsoft), weights copied from t3_nano_v1.safetensors.
  • Quantization (Turbo's official style): MatMulNBits 4-bit, block 32, asymmetric (zero points), no accuracy_level (int8 activation quantization breaks GPT-2's outlier dimensions); embedding tables as GatherBlockQuantized 4-bit. q4f16 additionally stores weights and KV cache in fp16.
  • conditional_decoder_q4 / _q4f16: Turbo's official graphs, copied unchanged (identical s3gen_meanflow.safetensors).

Validation

Parity vs PyTorch (reference: default_voice.wav, 7.4 s)

Graph Output max abs diff cosine
speech_encoder (fp32, before quantization) audio_features 0 1.0000000
audio_tokens 186/186 identical
speaker_embeddings vs upstream s3gen.embed_ref 6.4e-6 1.0000000
speaker_features 8.2e-2 0.9999999
speech_encoder_q4 audio_features 3.25 0.9642
audio_tokens 96.8 % identical
speaker_embeddings 7.8e-2 0.99975
embed_tokens (fp32) inputs_embeds 0 1.0000000
embed_tokens_q4 inputs_embeds 0.115 0.9964
language_model (fp32, before quantization) logits, 41 teacher-forced steps 1.7e-4 1.0000000 (top-1 100 %)
language_model_q4 logits, prefill 7.0 0.952
language_model_q4f16 logits, prefill 7.3 0.951

4-bit logit drift is of the same kind as Turbo's official q4 graphs; it does not show up in intelligibility or speaker similarity (below).

End-to-end (10 English sentences, default_voice.wav reference)

Decoding: greedy argmax with repetition penalty 1.2 (how WinSTT runs it). WER from NVIDIA Parakeet TDT 0.6B v3 (ONNX) transcripts, normalized (lowercase, no punctuation). Speaker similarity: cosine between Chatterbox VoiceEncoder embeddings of the output and the reference.

Graph set WER Speaker similarity
fp32 speech_encoder / embed / LM + decoder_q4 1.68 % 0.928
q4 (speech_encoder_q4, embed_tokens_q4, language_model_q4, conditional_decoder_q4) 1.68 % 0.927
q4f16 (โ€ฆ language_model_q4f16, conditional_decoder_q4f16) 1.68 % 0.926

Usage

Identical to chatterbox-turbo-ONNX: run speech_encoder once on the 24 kHz reference, prepend its audio_features to embed_tokens of the tokenized text (+ two 50256 start-of-speech ids), decode speech tokens with the KV-cached language_model (upstream samples with temperature 0.8, top-k 1000, top-p 0.95 and repetition penalty 1.2; greedy + penalty also works), prepend audio_tokens, and run conditional_decoder with speaker_embeddings / speaker_features to get 24 kHz audio. English only.

Watermark note

Resemble AI's Python package embeds an imperceptible Perth watermark into every generated waveform. These ONNX graphs do not contain the watermarker: their output is the raw vocoder waveform. If you deploy this model, Resemble AI asks that you keep watermarking generated audio (e.g. run Perth on the output) and use it responsibly.

License and credit

MIT, inherited from Resemble AI's Chatterbox (see LICENSE, NOTICE.md). All credit for the model goes to Resemble AI; this is only a format conversion.

Downloads last month
3
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for Masterx/chatterbox-nano-ONNX

Quantized
(12)
this model