phoonnx-spark-tts / README.md
Jarbas's picture
card: wav2vec2 parity on real audio
6f37844 verified
|
Raw
History Blame Contribute Delete
3.99 kB
---
license: cc-by-nc-sa-4.0
base_model: SparkAudio/Spark-TTS-0.5B
library_name: phoonnx
pipeline_tag: text-to-speech
tags:
- onnx
- text-to-speech
- phoonnx
- spark-tts
- bicodec
language:
- en
- zh
---
# Spark-TTS 0.5B — ONNX for phoonnx
ONNX weights for [Spark-TTS 0.5B](https://huggingface.co/SparkAudio/Spark-TTS-0.5B)
(SparkAudio), packaged for the [phoonnx](https://github.com/TigreGotico/phoonnx)
`sparktts` engine.
Spark-TTS is a decoder-only language model on a Qwen2.5-0.5B backbone. It does not
predict audio. It predicts **BiCodec** tokens, and BiCodec turns those tokens into a
waveform. BiCodec keeps two streams: 32 *global* tokens that carry the speaker, and one
*semantic* stream at 50 Hz that carries what is said. The language model reads the text
and the 32 global tokens, then writes the semantic stream.
## Files
| File | Stage | Input | Output |
| :--- | :--- | :--- | :--- |
| `model.onnx` (+ `model.onnx_data`) | Qwen2 language model, KV-cached | `input_ids`, `attention_mask`, `position_ids`, `past_key_values.*` | `logits`, `present.*` |
| `bicodec_vocoder.onnx` | BiCodec decoder | `semantic_tokens` `[1,T]`, `global_tokens` `[1,1,32]` | waveform `[1,1,N]` at 16 kHz |
| `wav2vec2_model.onnx` | cloning front end | `wav` `[1,N]` at 16 kHz, zero-mean unit-variance | `feat` `[1,T,1024]` |
| `bicodec_encoder_quantizer.onnx` | cloning: semantic stream | `feat` `[1,T,1024]` | `semantic_tokens` `[1,T]` |
| `speaker_encoder_tokenizer.onnx` | cloning: speaker stream | `spec` `[1,513,T]` magnitude STFT | `global_tokens` `[1,1,32]` |
| `tokenizer.json` | the model's own subword BPE | | |
| `voices/*.json` | preset speakers: 32 global tokens each | | |
The short-time Fourier transform in front of the speaker encoder is **not** in the
graphs. ONNX has no complex dtype, so neither torch exporter can lower `torch.stft`. The
mel filterbank projection *is* inside `speaker_encoder_tokenizer.onnx`; a caller only has
to supply the magnitude spectrogram of `torchaudio.transforms.Spectrogram(n_fft=1024,
win_length=640, hop_length=320, power=1)`.
## Provenance
* `model.onnx` is mirrored from
[Fhrozen/Spark-TTS-0.5B-ONNX](https://huggingface.co/Fhrozen/Spark-TTS-0.5B-ONNX),
verified against the SparkAudio torch weights (below).
* the four BiCodec / wav2vec2 graphs were exported from
[SparkAudio/Spark-TTS-0.5B](https://huggingface.co/SparkAudio/Spark-TTS-0.5B)
with opset 17.
* the preset voices were minted with Spark-TTS controllable-TTS mode
(gender, moderate pitch, moderate speed) and then frozen, so each voice is stable
across calls.
## Verified against the torch model
| Stage | Metric | Result |
| :--- | :--- | :--- |
| Language model, prefill | max absolute logit difference | 1.3e-3 |
| Language model, 8 decode steps | max absolute logit difference | 6.6e-4 |
| Language model, 8 decode steps | greedy token agreement | 8/8 |
| Speaker tokenizer | token mismatches | 0/32 |
| Encoder + quantizer | token mismatches | 0/149 |
| BiCodec vocoder | max absolute waveform difference | 1.2e-6 |
| Magnitude STFT (NumPy vs torchaudio) | max absolute difference | 4.8e-7 |
| wav2vec2 features (real clip) | max relative difference | 1.8e-5 |
| wav2vec2 → quantizer, end to end | semantic token mismatches | 0/199 |
The quantized language models published upstream (`model_q4`, `model_q4f16`,
`model_int8`) are **not** mirrored here. They diverge from the torch model by tens of
logits and disagree with greedy decoding, so they are not safe defaults.
## License and attribution
Spark-TTS 0.5B is released by **SparkAudio** under
[CC-BY-NC-SA-4.0](https://creativecommons.org/licenses/by-nc-sa/4.0/); the upstream code
is Apache-2.0. This repository redistributes converted weights under the same terms and
adds no license of its own. Cite the original work:
> Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled
> Speech Tokens — SparkAudio. https://github.com/SparkAudio/Spark-TTS