| --- |
| license: apache-2.0 |
| language: |
| - en |
| library_name: phoonnx |
| pipeline_tag: text-to-speech |
| tags: |
| - onnx |
| - tts |
| - orpheus |
| - snac |
| - codec-lm |
| base_model: |
| - canopylabs/orpheus-3b-0.1-ft |
| - hubertsiuzdak/snac_24khz |
| --- |
| |
| # phoonnx-orpheus β Orpheus 3B (English), ONNX |
|
|
| ONNX weights for [Orpheus TTS](https://github.com/canopyai/Orpheus-TTS) by |
| [Canopy Labs](https://canopylabs.ai), packaged for |
| [phoonnx](https://github.com/TigreGotico/phoonnx). This repository holds converted |
| weights only β no new training was done. Original model and code are Apache-2.0; |
| all credit for the model belongs to Canopy Labs. |
|
|
| Orpheus is a Llama-3.2-3B causal LM whose vocabulary carries 28 672 audio tokens. |
| It emits a flat token stream; every seven tokens form one |
| [SNAC](https://huggingface.co/hubertsiuzdak/snac_24khz) frame, which SNAC's decoder |
| turns into 2048 samples at 24 kHz. |
|
|
| ## Read this before you use it: Orpheus is not a CPU model |
|
|
| A 3B backbone needs about **0.37 s per decode step** on 12 CPU cores, and SNAC needs |
| ~82 tokens for every second of audio. That is roughly **37-41x slower than real time** |
| on CPU. Canopy Labs announced 1B / 400M / 150M tiers but **never released them** β |
| their own loader still raises `"not supported ... will be released very soon"` β so |
| there is no smaller Orpheus to fall back to. |
|
|
| Use this on a GPU. It is mirrored for completeness, not as an offline CPU default. |
|
|
| ## Files (`orpheus-3b-en-onnx/`) |
|
|
| | File | What it is | |
| |---|---| |
| | `model.onnx` + `model.onnx_data`..`model.onnx_data_6` | Llama backbone, fp32, KV-cached. **The graph needs all seven `.onnx_data` shards next to it.** | |
| | `snac_decoder.onnx` | SNAC 24 kHz decoder, fp32, copied from [onnx-community/snac_24khz-ONNX](https://huggingface.co/onnx-community/snac_24khz-ONNX) (Apache-2.0), unmodified. | |
| | `tokenizer.json` | The checkpoint's own BPE, copied from upstream. | |
| |
| `model.onnx` comes from |
| [onnx-community/orpheus-3b-0.1-ft-ONNX](https://huggingface.co/onnx-community/orpheus-3b-0.1-ft-ONNX), |
| re-verified against the torch reference before mirroring (see below). |
| |
| ## Voices |
| |
| `tara`, `leah`, `jess`, `leo`, `dan`, `mia`, `zac`, `zoe` β in Canopy Labs' own order of |
| conversational realism. The voice is a **name written into the prompt text**, not an |
| embedding or a speaker id. |
| |
| Emotive tags are ordinary text the same BPE encodes: `<laugh>`, `<chuckle>`, `<sigh>`, |
| `<cough>`, `<sniffle>`, `<groan>`, `<yawn>`, `<gasp>`. |
| |
| ## Verification |
| |
| Measured against the torch reference (`unsloth/orpheus-3b-0.1-ft`, an ungated copy of |
| the gated original) on a 19-token served prompt, fp32 both sides: |
| |
| | Variant | Size | Prefill max abs logit diff | Greedy agreement, 25 steps | ms / decode step | Verdict | |
| |---|---|---|---|---|---| |
| | `model` (fp32) | 12.7 GB | 0.166 | **25/25** | 369 | mirrored | |
| | `model_q4` | 2.4 GB | 8.65 | 23/25 | 371 | rejected β loses parity **and** is not faster | |
| | `model_q4f16` | 2.1 GB | 8.53 | 10/25 | 96 | rejected β fails parity | |
|
|
| Only fp32 is mirrored. The quantized variants either lose the greedy path or, in the |
| q4 case, buy nothing: its weights are int4 but its activations stay fp32, so |
| onnxruntime dequantizes on the fly and the step cost is unchanged. |
|
|
| SNAC's decoder is **stochastic** β it contains a noise block, so two decodes of the |
| same codes differ. Its ONNX export was therefore checked against the model's own |
| run-to-run spread (mean of 40 decodes, relative RMS) rather than by exact diff: |
|
|
| | SNAC decoder | Relative RMS vs torch mean | Ratio to noise floor | Verdict | |
| |---|---|---|---| |
| | `decoder_model` (fp32) | 0.0200 | 0.89x | pass | |
| | `decoder_model_fp16` | 0.0233 | 1.03x | pass | |
| | `decoder_model_int8` | 0.1356 | 6.02x | fail | |
| | `decoder_model_uint8` / `_quantized` | 0.1089 | 4.83x | fail | |
|
|
| `decoder_model_q4` and `decoder_model_bnb4` are byte-for-byte the same size as fp32 and |
| produce identical output β the quantizer skipped this convolutional graph entirely, so |
| those names are misleading rather than smaller. |
|
|
| ## Licensing |
|
|
| - Orpheus model and code: Apache-2.0, Β© Canopy Labs. |
| - SNAC: Apache-2.0, Β© Hubert Siuzdak. |
| - This repository: Apache-2.0. Converted weights only. |
|
|
| Canopy Labs ask that the model not be used for impersonation without consent, |
| misinformation, or deception. |
|
|