Chatterbox Nano FP16 MLX

A native FP16 MLX conversion of ResembleAI/chatterbox-nano, prepared by Rybib for Apple Silicon and on-device experimentation.

Chatterbox Nano and its architecture were created by Resemble AI. This is an independent conversion, not an official Resemble AI release.

Package

  • Native MLX tensor names and layouts
  • FP16 floating-point weights; no 4-bit quantization
  • Nano's distilled mean-flow speech decoder
  • English tokenizer and default voice conditioning
  • Approximately 697 MB for speech generation
  • Bundled S3TokenizerV2 companion model for offline voice cloning
  • Approximately 1.16 GB total with voice cloning

The converter strictly verifies that every required Nano decoder and vocoder parameter is mapped. It also merges PyTorch weight-normalization parameters into inference weights and converts convolution/GPT-2 layouts for MLX.

Quality

Listening tests on Apple Silicon found the output perceptually equivalent to the official FP32 Nano model. This is not a formal benchmark; applications should perform their own voice-cloning and long-form tests.

Usage

This package requires a Chatterbox Nano-capable MLX Audio runtime. At the time of conversion, upstream MLX Audio did not correctly map all Nano mean-flow decoder tensors without the accompanying runtime/converter fixes.

Voice cloning additionally requires the S3TokenizerV2 speech tokenizer. It is included under S3TokenizerV2/ so applications can clone voices completely offline after one installation.

Provenance and license

The upstream Chatterbox model is distributed under the MIT License. Its copyright notice is preserved in LICENSE.

Downloads last month
56
Safetensors
Model size
0.4B params
Tensor type
F16
·
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Rybib/chatterbox-nano-fp16-mlx

Finetuned
(1)
this model