--- license: other license_name: mixed-see-below language: [en, ar, zh, nl, fr, de, it, ja, ko, lt, ru, es, pt, be, bn, ka, hu, lv, fa, pl, sw, ta, uk] pipeline_tag: text-to-speech library_name: phoonnx tags: [tts, onnx, outetts, dac, phoonnx] --- # phoonnx OuteTTS 1.0 mirror ONNX artifacts for the [OuteTTS 1.0](https://github.com/edwko/OuteTTS) family, laid out for the `outetts` engine in [phoonnx](https://github.com/OpenVoiceOS/phoonnx). All models are **redistributed unchanged** from the official OuteAI and IBM Research repositories. This mirror adds no weights of its own; it only guarantees a stable file layout and stable URLs for phoonnx voice entries. ## Contents | Path | Source | License | |---|---|---| | `0.6B/model.onnx` (+ `model.onnx_data`), `0.6B/tokenizer.json` | [OuteAI/OuteTTS-1.0-0.6B-ONNX](https://huggingface.co/OuteAI/OuteTTS-1.0-0.6B-ONNX) | Apache-2.0 | | `1B/model.onnx` (+ `model.onnx_data`) | **re-exported by phoonnx** from [OuteAI/Llama-OuteTTS-1.0-1B](https://huggingface.co/OuteAI/Llama-OuteTTS-1.0-1B) - see below | **CC-BY-NC-SA-4.0 (non-commercial)** | | `1B/tokenizer.json` | [OuteAI/Llama-OuteTTS-1.0-1B-ONNX](https://huggingface.co/OuteAI/Llama-OuteTTS-1.0-1B-ONNX) | **CC-BY-NC-SA-4.0 (non-commercial)** | | `dac/decoder_model.onnx`, `dac/encoder_model.onnx`, `dac/config.json` | [OuteAI/DAC-speech-v1.0-ONNX](https://huggingface.co/OuteAI/DAC-speech-v1.0-ONNX) (weights: [ibm-research/DAC.speech.v1.0](https://huggingface.co/ibm-research/DAC.speech.v1.0)) | CDLA-Permissive-2.0 | | `speakers/en-female-1-neutral.json` | the `outetts` Python package (`version/v3/default_speakers`) | Apache-2.0 | The two model sizes carry **different licenses**. The 0.6B is Apache-2.0 and covers 14 languages; phoonnx resolves those 14 to it. The 1B is CC-BY-NC-SA-4.0 - **not licensed for commercial use** - and is here because it is the only member of the family that covers Arabic, Belarusian, Bengali, Lithuanian, Persian, Portuguese, Swahili, Tamil and Ukrainian. Check the 1B license before you deploy it. Model authorship and all research credit belong to **OuteAI**. The audio codec is DAC.speech.v1.0 by **IBM Research**. ## Languages - **0.6B (Apache-2.0), 14 trained languages:** English, Chinese, Dutch, French, Georgian, German, Hungarian, Italian, Japanese, Korean, Latvian, Polish, Russian, Spanish. - **1B (CC-BY-NC-SA-4.0), 23 trained languages:** the 14 above plus Arabic, Belarusian, Bengali, Lithuanian, Persian, Portuguese, Swahili, Tamil, Ukrainian. ## Numeric parity Every export in this mirror was checked against the upstream torch weights on a fixed 1843-token prompt (greedy, 64 decode steps, CPU float32). | Export | max abs logit diff | greedy agreement | shipped | |---|---|---|---| | 0.6B `model.onnx` (fp32) | 3.8e-05 | 64/64 | **yes** | | 0.6B `model_fp16.onnx` | 0.036 | 63/64 | no | | 0.6B `model_quantized` / `uint8` | 5.7 | 62/64 | no | | 0.6B `model_bnb4` / `int8` / `q4` / `q4f16` | 3.9 - 5.9 | 32-34/64 | no | | 1B `model.onnx` - OuteAI's (fp32) | 12 | 61/64 | no, see below | | 1B `model.onnx` - this repo's re-export (fp32) | **1.78e-05** | **64/64** | **yes** | The 0.6B float32 export reproduces its torch weights exactly. Every quantized 0.6B export changes greedy decoding, so none is a safe default. ### The official 1B ONNX export is broken; this one is not `OuteAI/Llama-OuteTTS-1.0-1B-ONNX` **does not reproduce its own torch weights**, even in float32. On the prompt above its last-position logits are off by 12, the logit correlation is **0.48**, prefill argmax agrees on only 1580 of 1843 positions, and greedy decoding diverges from `OuteAI/Llama-OuteTTS-1.0-1B`. The error is already present on a 32-token prompt (0.08 max diff) and never closes (correlation 0.96-0.98 at 512-2048 tokens), so it is a property of the export, not of long-context accumulation. The 0.6B export on the same prompt differs by 3.8e-05. Every quantized 1B export inherits the defect and adds to it. The `1B/model.onnx` in **this** repository is therefore **not** OuteAI's file. It was re-exported from the torch checkpoint with [`scripts/conversion/outetts/export_outetts_onnx.py`](https://github.com/TigreGotico/phoonnx/blob/dev/scripts/conversion/outetts/export_outetts_onnx.py) and verified the same way: **max abs logit diff 1.78e-05, mean 3.74e-06, logit correlation 1.00000000, greedy 64/64**, and 64/64 again under the windowed repetition penalty OuteTTS actually samples with. The same script reproduces the 0.6B to 1.6e-05, which is how it is checked against a graph already known to be good. One difference to know about: the re-exported graph returns `logits` shaped `[1, 1, V]` - the final row only - where OuteAI's exports return every position. Any consumer that reads the last row, as transformers.js and phoonnx both do, works with either. ## Use it from phoonnx ```python from phoonnx.model_manager import TTSModelManager manager = TTSModelManager() manager.merge_default_voices() voice = manager.voices["outetts/0.6B/en"].load() for chunk in voice.synthesize("Hello from OuteTTS running in phoonnx."): ... # chunk.audio_float_array, 24 kHz mono ``` Voice ids run `outetts/0.6B/` for the 14 languages the 0.6B was trained on and `outetts/1B/` for the nine only the 1B covers (ar, be, bn, fa, lt, pt, sw, ta, uk). Every one of them resolves to this repository. The engine drives the LM with a KV cache, then decodes the two DAC codebooks to a waveform. ## How it works The LM emits interleaved `<|c1_N|><|c2_N|>` tokens. `dac/decoder_model.onnx` takes `audio_codes[1, 2, T]` and returns `audio_values[1, 1, T * 512]` at 24 kHz. A speaker profile (`speakers/*.json`) is an in-context audio prompt: a transcript whose words each carry a duration, three prosody buckets and their DAC codes. ## Ethical use OuteAIs guidance applies unchanged: do not clone a voice without the speakers explicit permission, and do not use this model to impersonate or to produce deceptive content.