File size: 5,383 Bytes
05a6e70 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 | ---
license: cc-by-nc-4.0
language:
- en
- zh
library_name: phoonnx
pipeline_tag: text-to-speech
tags:
- onnx
- tts
- llasa
- xcodec2
- codec-lm
base_model:
- HKUSTAudio/Llasa-1B
- HKUSTAudio/xcodec2
---
# phoonnx-llasa — Llasa-1B + XCodec2, ONNX
ONNX conversion of [HKUSTAudio/Llasa-1B](https://huggingface.co/HKUSTAudio/Llasa-1B)
and its codec [HKUSTAudio/xcodec2](https://huggingface.co/HKUSTAudio/xcodec2), packaged
for [phoonnx](https://github.com/TigreGotico/phoonnx). This repository holds converted
weights only — no new training was done.
Llasa (arXiv [2502.04128](https://arxiv.org/abs/2502.04128)) is a LLaMA-3.2-1B backbone
whose vocabulary was extended with the 65,536 `<|s_N|>` speech tokens of XCodec2, a
single-codebook 16 kHz codec at 50 tokens per second. Prompt it with text and it emits a
run of speech tokens; the codec decoder turns those back into a waveform.
## Licence
Both upstream repositories are **CC BY-NC 4.0**, and so is this conversion:
non-commercial use only. The licence is upstream HKUST's, not phoonnx's. Model and
weights are by the HKUST Audio group; this repository redistributes them in ONNX form
and adds nothing but the graph rewrites described below.
## Files (`llasa-1b-onnx/`)
| File | What it is |
|---|---|
| `model.onnx` + `model.onnx_data` | LLaMA backbone, fp32, KV-cached. **The graph needs its `model.onnx_data` sidecar next to it.** |
| `xcodec2_decoder.onnx` | XCodec2 decoder, fp32. Codes in, waveform out. |
| `tokenizer.json` | The checkpoint's own BPE, copied from upstream. |
| `voices.json` | Four voice presets. Every one is machine-generated — see below. |
| `config.json` | Self-describing phoonnx voice config. |
| `samples/` | One rendering of each preset. |
There is **no quantized variant**. Dynamic int8 (per-tensor and per-channel) and 4-bit
`MatMulNBits` were all built and all failed the greedy-agreement gate against torch:
mean absolute logit error of 1.3 to 3.6 and 0 to 25 matching tokens out of 48, against
5e-6 and 48 out of 48 for fp32. Treat the quantized files in other Llasa ONNX
repositories with the same suspicion.
## Graph contract
`model.onnx` serves prefill and decode: the same graph with a different past length.
```
inputs input_ids int64 [1, S] prompt, or 1 token per step
attention_mask int64 [1, P + S] ones over past and current
position_ids int64 [1, S] absolute positions P .. P+S-1
past_key_values.<i>.key fp32 [1, 8, P, 64] i in 0..15
past_key_values.<i>.value fp32 [1, 8, P, 64]
outputs logits fp32 [1, 1, 193800] last position only
present.<i>.key / present.<i>.value fp32 [1, 8, P + S, 64]
```
`xcodec2_decoder.onnx` takes `codes` int64 `[1, 1, N]` and returns `audio` float32
`[1, 320 * N]` at 16 kHz.
Two rewrites were needed. Both preserve behaviour, and both were measured:
* **`lm_head` shares the embedding.** Llasa ties the two, but the exporter wrote the
193,800 x 2,048 matrix twice. The head is now `Reshape -> Gemm(transB=1) -> Unsqueeze`
over the embedding initialiser: 7.07 GB becomes 5.48 GB, with bit-identical logits.
* **The codec's ISTFT is real-valued.** ONNX cannot trace complex tensors, so the inverse
real FFT became two constant cosine/sine matmuls and the overlap-add became a
`conv_transpose1d` with an identity kernel. Against the complex path the largest sample
difference is 6e-7.
Only the last position's logits leave the graph. Over a 193,800-wide vocabulary,
returning a whole prefill would cost about 78 MB per 100 prompt tokens for a value the
sampler never reads.
## Parity against torch
`transformers` fp32 against this graph, English and Chinese prompts, 48 greedy steps each:
| Prompt | prefill max abs logit diff | decode max abs logit diff | greedy agreement |
|---|---|---|---|
| English | 4.1e-05 | 4.1e-05 | 48/48 |
| Chinese | 2.7e-05 | 3.7e-05 | 48/48 |
Codec decoder, 400 tokens (8 s), against upstream `decode_code`: largest sample
difference 1.2e-04 on a signal of RMS 0.258, correlation 0.9999999999.
## Voices
Llasa needs no reference audio: prompted with text alone it invents a speaker, and two
calls never sound like the same person. The presets in `voices.json` pin one down. Each
holds the transcript of an utterance the model generated **and the speech tokens it
emitted for it**; replaying those tokens as an in-context prefix continues that speaker.
Every preset is therefore machine-generated from text alone. **No preset is a recording
of any person**, and each is marked `"synthetic": true`.
| Preset | Language |
|---|---|
| `en_female_a` | English |
| `en_male_a` | English |
| `zh_female_a` | Chinese |
| `zh_male_a` | Chinese |
Cloning from a fresh clip is not supported by this bundle: tokenising one needs XCodec2's
encoder together with the w2v-BERT filterbank front end, which are not included.
## Usage
```python
from phoonnx.model_manager import TTSModelManager
from phoonnx.config import SynthesisConfig
voice = TTSModelManager().load_voice("llasa/HKUST/en/1b")
audio = b"".join(c.audio_int16_bytes for c in voice.synthesize(
"Dealing with family secrets is never easy.",
syn_config=SynthesisConfig(extra_params={"voice": "en_female_a"})))
```
Conversion scripts: `scripts/conversion/llasa/` in the phoonnx repository.
|