Instructions to use KitsuMate/chatterbox-onnx with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Chatterbox
How to use KitsuMate/chatterbox-onnx with Chatterbox:
# pip install chatterbox-tts import torchaudio as ta from chatterbox.tts import ChatterboxTTS model = ChatterboxTTS.from_pretrained(device="cuda") text = "Ezreal and Jinx teamed up with Ahri, Yasuo, and Teemo to take down the enemy's Nexus in an epic late-game pentakill." wav = model.generate(text) ta.save("test-1.wav", wav, model.sr) # If you want to synthesize with a different voice, specify the audio prompt AUDIO_PROMPT_PATH="YOUR_FILE.wav" wav = model.generate(text, audio_prompt_path=AUDIO_PROMPT_PATH) ta.save("test-2.wav", wav, model.sr) - Notebooks
- Google Colab
- Kaggle
Chatterbox ONNX
KitsuMate layout of the English ResembleAI Chatterbox model using the pinned ONNX Community export at 3cab09af388d3f02bba43443fce88c1f4525ac43.
The repository includes the upstream FP32, FP16, Q4, and Q4F16 language-model
variants. The conditional decoder, token embedding, and speech encoder published
without a suffix are shared by the upstream precision variants. The additional
q4f16-suffixed copies provide a uniformly named four-graph profile for the
KitsuMate model catalog.
Every imported ONNX file is byte-for-byte unchanged from ONNX Community
revision 3cab09af388d3f02bba43443fce88c1f4525ac43. Each graph must remain beside
its matching .onnx_data file. The English tokenizer and default voice are
also redistributed unchanged under the upstream MIT license. This is the
English checkpoint; it intentionally excludes the unrelated Cangjie mapping.