Chatterbox Multilingual V3 โ€” ONNX

ONNX export of Resemble AI's Chatterbox Multilingual V3 (t3_mtl23ls_v3.safetensors + s3gen.pt from ResembleAI/chatterbox), a 0.5B zero-shot voice-cloning TTS for 23 languages. Upstream: https://github.com/resemble-ai/chatterbox.

The four graphs keep exactly the file names, graph layout and input/output contract of the onnx-community V2 export, so any runtime written for V2 runs V3 unchanged (it is what WinSTT ships). Differences that matter to a V2 runtime:

  • tokenizer.json now carries an NFKD normalizer โ€” V3 is trained on full-case NFKD text (text_preproc = "NFKD,fullcase"); ids are identical to upstream MTLTokenizer for every Space demo sentence of the 20 non-CJK languages checked. Do not lowercase.
  • The speech encoder crops the reference to the first 6 s inside the graph (V3's REF_COND_DURATION_S = 6; V2 used 10 s for the decoder prompt).
  • The speech encoder uses CAMPPlus's original BatchNorm dense layer (the V2 export swapped it for a LayerNorm, which changes the speaker x-vector direction).
  • Upstream V3 drops the audio of the final speech token before STOP (wav[:(n-1)*960]); do that in your runtime as well.
  • The KV-cache IO is declared [batch_size, 16, past_sequence_length, 64] exactly as in V2 (the current genai builder emits a symbolic last dim; it is pinned to 64 here).
  • Upstream frontends this export does not include: Chinese โ†’ Cangjie conversion (zh) and Russian stress marks (ru); there is no Japanese kana conversion either. Without them zh/ja text is mostly [UNK]. Korean works through the NFKD normalizer (3.9 % CER, below).
  • Hebrew is unsupported / poor. V3's upstream frontend feeds Hebrew undiacritized, and with greedy decoding the output is largely unintelligible (Whisper-small CER 52 % on 2 sentences through the Python runtime, 84 % on 1 sentence through WinSTT's Rust runtime; generation tends to stop early). WinSTT does not advertise he for this model.

Files

File Bytes
default_voice.wav 714,320
onnx/conditional_decoder.onnx 6,371,670
onnx/conditional_decoder.onnx_data 533,970,816
onnx/embed_tokens.onnx 17,616
onnx/embed_tokens.onnx_data 68,808,704
onnx/language_model.onnx 171,445
onnx/language_model.onnx_data 2,080,632,832
onnx/language_model_q4.onnx 228,633
onnx/language_model_q4.onnx_data 353,621,248
onnx/speech_encoder.onnx 1,204,600
onnx/speech_encoder.onnx_data 591,274,880
tokenizer.json 72,765

Graph inputs / outputs

Graph Inputs Outputs
conditional_decoder.onnx speech_tokens int64[batch_size, num_speech_tokens]
speaker_embeddings float32[batch_size, 192]
speaker_features float32[batch_size, feature_dim, 80]
waveform float32[batch_size, num_samples]
embed_tokens.onnx input_ids int64[batch_size, sequence_length]
position_ids int64[batch_size, sequence_length]
exaggeration float32[batch_size]
inputs_embeds float32[batch_size, sequence_length, 1024]
language_model.onnx attention_mask int64[batch_size, total_sequence_length]
inputs_embeds float32[batch_size, sequence_length, 1024]
past_key_values.{0..29}.{key,value} float32[batch_size, 16, past_sequence_length, 64]
logits float32[batch_size, sequence_length, 8194]
present.{0..29}.{key,value}
language_model_q4.onnx attention_mask int64[batch_size, total_sequence_length]
inputs_embeds float32[batch_size, sequence_length, 1024]
past_key_values.{0..29}.{key,value} float32[batch_size, 16, past_sequence_length, 64]
logits float32[batch_size, sequence_length, 8194]
present.{0..29}.{key,value}
speech_encoder.onnx audio_values float32[batch_size, num_samples] audio_features float32[batch_size, sequence_length, 1024]
audio_tokens int64[batch_size, audio_sequence_length]
speaker_embeddings float32[1, 192]
speaker_features float32[batch_size, feature_dim, 80]

How it was exported

  • language_model: Llama backbone of t3_mtl23ls_v3.safetensors (30 layers, 16 heads, head dim 64, speech head 8194) built with onnxruntime_genai.models.builder (exclude_embeds=true), the same tool and layout as the V2 export: GroupQueryAttention with rotary cos/sin caches, SimplifiedLayerNormalization, opset 21 + com.microsoft. language_model.onnx is fp32; language_model_q4.onnx is MatMulNBits 4-bit, block 32, symmetric, accuracy_level=4 (V2's settings), lm_head included.
  • speech_encoder / embed_tokens: torch.onnx.export (opset 20, TorchScript exporter) of the onnx-community conversion recipe's PrepareConditionalsModel / InputsEmbeds with the V3 T3 weights, a 6 s input crop and CAMPPlus's original BatchNorm dense layer written as explicit arithmetic; then onnxslim.
  • conditional_decoder: the recipe's ConditionalDecoder (S3Gen flow matching + HiFT vocoder) from s3gen.pt, opset 17, then onnxslim. It samples its initial noise in-graph (RandomNormalLike), as V2's does.

Validation

Parity vs PyTorch (reference: default_voice.wav)

Graph Check max abs diff cosine / agreement
speech_encoder speaker_embeddings vs upstream s3gen.embed_ref 6.0e-6 1.0000000
speaker_features (prompt mel) 8.2e-2 0.99999995
audio_features (T3 conditioning) 0.12 0.99981
audio_tokens vs upstream S3 tokenizer 149/150 identical
embed_tokens inputs_embeds 0 1.0000000
language_model (fp32) prefill logits 3.3e-5 1.0000000
107-step greedy run 107/107 tokens identical, same STOP
language_model_q4 prefill logits 2.18 0.99367
teacher-forced along the PyTorch path top-1 81 %, logit cosine mean 0.9965 (min 0.973)

conditional_decoder is not compared numerically: it draws random noise in-graph, so it is validated end to end below.

End-to-end (q4 rung: speech_encoder, embed_tokens, language_model_q4, conditional_decoder)

Greedy decoding with repetition penalty 1.2 (how WinSTT runs it), default_voice.wav reference. WER: NVIDIA Parakeet TDT 0.6B v3 transcripts, normalized (lowercase, no punctuation; digits are not expanded, so sentences with numbers score a few points of WER even when spoken correctly). CER: Whisper-small. Speaker similarity: cosine of Chatterbox VoiceEncoder embeddings.

Language Sentences WER / CER Speaker similarity
English 10 2.5 % WER 0.926
French 10 4.6 % WER 0.828
German 10 4.4 % WER 0.871
Spanish 10 3.0 % WER 0.839
Korean 3 3.9 % CER
Hebrew 2 52 % CER (unsupported)

The same graphs through WinSTT's Rust runtime (3 sentences each): en 0 %, fr 5.7 %, de 0 %, es 0 % WER; Korean 4.8 % CER.

V3 vs the V2 export (same 5 sentences per language, same reference, same runtime)

en fr de es mean WER speaker similarity
V2 (onnx-community) 3.2 % 3.3 % 2.0 % 20.8 % 7.3 % 0.860
V3 (this repo) 4.8 % 3.3 % 0 % 0 % 2.0 % 0.866

V2 arm: the onnx-community V2 language_model_q4 with V2's speech encoder / embeddings re-exported from the V2 checkpoint the way that repo built them (LayerNorm dense, no crop).

A/B of this export's choices on 6 sentences: BatchNorm dense + 6 s crop (shipped) 3.4 % WER, similarity 0.880, vs LayerNorm + no crop (V2-style) 5.4 % WER, similarity 0.873.

Usage

Identical to the onnx-community V2 export: run speech_encoder once on the 24 kHz reference; tokenize [<lang>] + text with tokenizer.json (it adds <EXAGGERATION>, <s>, </s> and two <START_SPEECH> ids); embed_tokens with position_ids and exaggeration; prepend audio_features; decode speech tokens with the KV-cached language_model_q4 (or the fp32 language_model); prepend audio_tokens; run conditional_decoder with speaker_embeddings / speaker_features; drop the last token's 960 samples. Output is 24 kHz mono.

Watermark note

Resemble AI's Python package embeds an imperceptible Perth watermark into every generated waveform. These ONNX graphs do not contain the watermarker: their output is the raw vocoder waveform. If you deploy this model, Resemble AI asks that you keep watermarking generated audio (e.g. run Perth on the output) and use it responsibly.

License and credit

MIT, inherited from Resemble AI's Chatterbox (see LICENSE, NOTICE.md). All credit for the model goes to Resemble AI; this is only a format conversion.

Downloads last month
15
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for Masterx/chatterbox-multilingual-v3-ONNX

Quantized
(37)
this model