Instructions to use Masterx/chatterbox-multilingual-v3-ONNX with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Chatterbox
How to use Masterx/chatterbox-multilingual-v3-ONNX with Chatterbox:
# pip install chatterbox-tts import torchaudio as ta from chatterbox.tts import ChatterboxTTS model = ChatterboxTTS.from_pretrained(device="cuda") text = "Ezreal and Jinx teamed up with Ahri, Yasuo, and Teemo to take down the enemy's Nexus in an epic late-game pentakill." wav = model.generate(text) ta.save("test-1.wav", wav, model.sr) # If you want to synthesize with a different voice, specify the audio prompt AUDIO_PROMPT_PATH="YOUR_FILE.wav" wav = model.generate(text, audio_prompt_path=AUDIO_PROMPT_PATH) ta.save("test-2.wav", wav, model.sr) - Notebooks
- Google Colab
- Kaggle
Chatterbox Multilingual V3 โ ONNX
ONNX export of Resemble AI's Chatterbox Multilingual V3 (t3_mtl23ls_v3.safetensors +
s3gen.pt from ResembleAI/chatterbox), a 0.5B
zero-shot voice-cloning TTS for 23 languages. Upstream: https://github.com/resemble-ai/chatterbox.
The four graphs keep exactly the file names, graph layout and input/output contract of the onnx-community V2 export, so any runtime written for V2 runs V3 unchanged (it is what WinSTT ships). Differences that matter to a V2 runtime:
tokenizer.jsonnow carries an NFKD normalizer โ V3 is trained on full-case NFKD text (text_preproc = "NFKD,fullcase"); ids are identical to upstreamMTLTokenizerfor every Space demo sentence of the 20 non-CJK languages checked. Do not lowercase.- The speech encoder crops the reference to the first 6 s inside the graph (V3's
REF_COND_DURATION_S = 6; V2 used 10 s for the decoder prompt). - The speech encoder uses CAMPPlus's original BatchNorm dense layer (the V2 export swapped it for a LayerNorm, which changes the speaker x-vector direction).
- Upstream V3 drops the audio of the final speech token before STOP (
wav[:(n-1)*960]); do that in your runtime as well. - The KV-cache IO is declared
[batch_size, 16, past_sequence_length, 64]exactly as in V2 (the current genai builder emits a symbolic last dim; it is pinned to 64 here). - Upstream frontends this export does not include: Chinese โ Cangjie conversion (
zh) and Russian stress marks (ru); there is no Japanese kana conversion either. Without themzh/jatext is mostly[UNK]. Korean works through the NFKD normalizer (3.9 % CER, below). - Hebrew is unsupported / poor. V3's upstream frontend feeds Hebrew undiacritized, and with
greedy decoding the output is largely unintelligible (Whisper-small CER 52 % on 2 sentences
through the Python runtime, 84 % on 1 sentence through WinSTT's Rust runtime; generation tends to
stop early). WinSTT does not advertise
hefor this model.
Files
| File | Bytes |
|---|---|
default_voice.wav |
714,320 |
onnx/conditional_decoder.onnx |
6,371,670 |
onnx/conditional_decoder.onnx_data |
533,970,816 |
onnx/embed_tokens.onnx |
17,616 |
onnx/embed_tokens.onnx_data |
68,808,704 |
onnx/language_model.onnx |
171,445 |
onnx/language_model.onnx_data |
2,080,632,832 |
onnx/language_model_q4.onnx |
228,633 |
onnx/language_model_q4.onnx_data |
353,621,248 |
onnx/speech_encoder.onnx |
1,204,600 |
onnx/speech_encoder.onnx_data |
591,274,880 |
tokenizer.json |
72,765 |
Graph inputs / outputs
| Graph | Inputs | Outputs |
|---|---|---|
conditional_decoder.onnx |
speech_tokens int64[batch_size, num_speech_tokens]speaker_embeddings float32[batch_size, 192]speaker_features float32[batch_size, feature_dim, 80] |
waveform float32[batch_size, num_samples] |
embed_tokens.onnx |
input_ids int64[batch_size, sequence_length]position_ids int64[batch_size, sequence_length]exaggeration float32[batch_size] |
inputs_embeds float32[batch_size, sequence_length, 1024] |
language_model.onnx |
attention_mask int64[batch_size, total_sequence_length]inputs_embeds float32[batch_size, sequence_length, 1024]past_key_values.{0..29}.{key,value} float32[batch_size, 16, past_sequence_length, 64] |
logits float32[batch_size, sequence_length, 8194]present.{0..29}.{key,value} |
language_model_q4.onnx |
attention_mask int64[batch_size, total_sequence_length]inputs_embeds float32[batch_size, sequence_length, 1024]past_key_values.{0..29}.{key,value} float32[batch_size, 16, past_sequence_length, 64] |
logits float32[batch_size, sequence_length, 8194]present.{0..29}.{key,value} |
speech_encoder.onnx |
audio_values float32[batch_size, num_samples] |
audio_features float32[batch_size, sequence_length, 1024]audio_tokens int64[batch_size, audio_sequence_length]speaker_embeddings float32[1, 192]speaker_features float32[batch_size, feature_dim, 80] |
How it was exported
- language_model: Llama backbone of
t3_mtl23ls_v3.safetensors(30 layers, 16 heads, head dim 64, speech head 8194) built withonnxruntime_genai.models.builder(exclude_embeds=true), the same tool and layout as the V2 export: GroupQueryAttention with rotary cos/sin caches, SimplifiedLayerNormalization, opset 21 +com.microsoft.language_model.onnxis fp32;language_model_q4.onnxisMatMulNBits4-bit, block 32, symmetric,accuracy_level=4(V2's settings), lm_head included. - speech_encoder / embed_tokens:
torch.onnx.export(opset 20, TorchScript exporter) of the onnx-community conversion recipe'sPrepareConditionalsModel/InputsEmbedswith the V3 T3 weights, a 6 s input crop and CAMPPlus's original BatchNorm dense layer written as explicit arithmetic; thenonnxslim. - conditional_decoder: the recipe's
ConditionalDecoder(S3Gen flow matching + HiFT vocoder) froms3gen.pt, opset 17, thenonnxslim. It samples its initial noise in-graph (RandomNormalLike), as V2's does.
Validation
Parity vs PyTorch (reference: default_voice.wav)
| Graph | Check | max abs diff | cosine / agreement |
|---|---|---|---|
| speech_encoder | speaker_embeddings vs upstream s3gen.embed_ref |
6.0e-6 | 1.0000000 |
| speaker_features (prompt mel) | 8.2e-2 | 0.99999995 | |
| audio_features (T3 conditioning) | 0.12 | 0.99981 | |
| audio_tokens vs upstream S3 tokenizer | 149/150 identical | ||
| embed_tokens | inputs_embeds | 0 | 1.0000000 |
| language_model (fp32) | prefill logits | 3.3e-5 | 1.0000000 |
| 107-step greedy run | 107/107 tokens identical, same STOP | ||
| language_model_q4 | prefill logits | 2.18 | 0.99367 |
| teacher-forced along the PyTorch path | top-1 81 %, logit cosine mean 0.9965 (min 0.973) |
conditional_decoder is not compared numerically: it draws random noise in-graph, so it is validated
end to end below.
End-to-end (q4 rung: speech_encoder, embed_tokens, language_model_q4, conditional_decoder)
Greedy decoding with repetition penalty 1.2 (how WinSTT runs it), default_voice.wav reference.
WER: NVIDIA Parakeet TDT 0.6B v3 transcripts, normalized (lowercase, no punctuation; digits are
not expanded, so sentences with numbers score a few points of WER even when spoken correctly).
CER: Whisper-small. Speaker similarity: cosine of Chatterbox VoiceEncoder embeddings.
| Language | Sentences | WER / CER | Speaker similarity |
|---|---|---|---|
| English | 10 | 2.5 % WER | 0.926 |
| French | 10 | 4.6 % WER | 0.828 |
| German | 10 | 4.4 % WER | 0.871 |
| Spanish | 10 | 3.0 % WER | 0.839 |
| Korean | 3 | 3.9 % CER | |
| Hebrew | 2 | 52 % CER (unsupported) |
The same graphs through WinSTT's Rust runtime (3 sentences each): en 0 %, fr 5.7 %, de 0 %, es 0 % WER; Korean 4.8 % CER.
V3 vs the V2 export (same 5 sentences per language, same reference, same runtime)
| en | fr | de | es | mean WER | speaker similarity | |
|---|---|---|---|---|---|---|
| V2 (onnx-community) | 3.2 % | 3.3 % | 2.0 % | 20.8 % | 7.3 % | 0.860 |
| V3 (this repo) | 4.8 % | 3.3 % | 0 % | 0 % | 2.0 % | 0.866 |
V2 arm: the onnx-community V2 language_model_q4 with V2's speech encoder / embeddings re-exported
from the V2 checkpoint the way that repo built them (LayerNorm dense, no crop).
A/B of this export's choices on 6 sentences: BatchNorm dense + 6 s crop (shipped) 3.4 % WER, similarity 0.880, vs LayerNorm + no crop (V2-style) 5.4 % WER, similarity 0.873.
Usage
Identical to the onnx-community V2 export:
run speech_encoder once on the 24 kHz reference; tokenize [<lang>] + text with tokenizer.json
(it adds <EXAGGERATION>, <s>, </s> and two <START_SPEECH> ids); embed_tokens with
position_ids and exaggeration; prepend audio_features; decode speech tokens with the
KV-cached language_model_q4 (or the fp32 language_model); prepend audio_tokens; run
conditional_decoder with speaker_embeddings / speaker_features; drop the last token's 960
samples. Output is 24 kHz mono.
Watermark note
Resemble AI's Python package embeds an imperceptible Perth watermark into every generated waveform. These ONNX graphs do not contain the watermarker: their output is the raw vocoder waveform. If you deploy this model, Resemble AI asks that you keep watermarking generated audio (e.g. run Perth on the output) and use it responsibly.
License and credit
MIT, inherited from Resemble AI's Chatterbox (see LICENSE, NOTICE.md). All credit for the model
goes to Resemble AI; this is only a format conversion.
- Downloads last month
- 15
Model tree for Masterx/chatterbox-multilingual-v3-ONNX
Base model
ResembleAI/chatterbox