File size: 5,383 Bytes
05a6e70
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
---
license: cc-by-nc-4.0
language:
- en
- zh
library_name: phoonnx
pipeline_tag: text-to-speech
tags:
- onnx
- tts
- llasa
- xcodec2
- codec-lm
base_model:
- HKUSTAudio/Llasa-1B
- HKUSTAudio/xcodec2
---

# phoonnx-llasa — Llasa-1B + XCodec2, ONNX

ONNX conversion of [HKUSTAudio/Llasa-1B](https://huggingface.co/HKUSTAudio/Llasa-1B)
and its codec [HKUSTAudio/xcodec2](https://huggingface.co/HKUSTAudio/xcodec2), packaged
for [phoonnx](https://github.com/TigreGotico/phoonnx). This repository holds converted
weights only — no new training was done.

Llasa (arXiv [2502.04128](https://arxiv.org/abs/2502.04128)) is a LLaMA-3.2-1B backbone
whose vocabulary was extended with the 65,536 `<|s_N|>` speech tokens of XCodec2, a
single-codebook 16 kHz codec at 50 tokens per second. Prompt it with text and it emits a
run of speech tokens; the codec decoder turns those back into a waveform.

## Licence

Both upstream repositories are **CC BY-NC 4.0**, and so is this conversion:
non-commercial use only. The licence is upstream HKUST's, not phoonnx's. Model and
weights are by the HKUST Audio group; this repository redistributes them in ONNX form
and adds nothing but the graph rewrites described below.

## Files (`llasa-1b-onnx/`)

| File | What it is |
|---|---|
| `model.onnx` + `model.onnx_data` | LLaMA backbone, fp32, KV-cached. **The graph needs its `model.onnx_data` sidecar next to it.** |
| `xcodec2_decoder.onnx` | XCodec2 decoder, fp32. Codes in, waveform out. |
| `tokenizer.json` | The checkpoint's own BPE, copied from upstream. |
| `voices.json` | Four voice presets. Every one is machine-generated — see below. |
| `config.json` | Self-describing phoonnx voice config. |
| `samples/` | One rendering of each preset. |

There is **no quantized variant**. Dynamic int8 (per-tensor and per-channel) and 4-bit
`MatMulNBits` were all built and all failed the greedy-agreement gate against torch:
mean absolute logit error of 1.3 to 3.6 and 0 to 25 matching tokens out of 48, against
5e-6 and 48 out of 48 for fp32. Treat the quantized files in other Llasa ONNX
repositories with the same suspicion.

## Graph contract

`model.onnx` serves prefill and decode: the same graph with a different past length.

```
inputs   input_ids                  int64 [1, S]        prompt, or 1 token per step
         attention_mask             int64 [1, P + S]    ones over past and current
         position_ids               int64 [1, S]        absolute positions P .. P+S-1
         past_key_values.<i>.key    fp32  [1, 8, P, 64] i in 0..15
         past_key_values.<i>.value  fp32  [1, 8, P, 64]
outputs  logits                     fp32  [1, 1, 193800]  last position only
         present.<i>.key / present.<i>.value  fp32 [1, 8, P + S, 64]
```

`xcodec2_decoder.onnx` takes `codes` int64 `[1, 1, N]` and returns `audio` float32
`[1, 320 * N]` at 16 kHz.

Two rewrites were needed. Both preserve behaviour, and both were measured:

* **`lm_head` shares the embedding.** Llasa ties the two, but the exporter wrote the
  193,800 x 2,048 matrix twice. The head is now `Reshape -> Gemm(transB=1) -> Unsqueeze`
  over the embedding initialiser: 7.07 GB becomes 5.48 GB, with bit-identical logits.
* **The codec's ISTFT is real-valued.** ONNX cannot trace complex tensors, so the inverse
  real FFT became two constant cosine/sine matmuls and the overlap-add became a
  `conv_transpose1d` with an identity kernel. Against the complex path the largest sample
  difference is 6e-7.

Only the last position's logits leave the graph. Over a 193,800-wide vocabulary,
returning a whole prefill would cost about 78 MB per 100 prompt tokens for a value the
sampler never reads.

## Parity against torch

`transformers` fp32 against this graph, English and Chinese prompts, 48 greedy steps each:

| Prompt | prefill max abs logit diff | decode max abs logit diff | greedy agreement |
|---|---|---|---|
| English | 4.1e-05 | 4.1e-05 | 48/48 |
| Chinese | 2.7e-05 | 3.7e-05 | 48/48 |

Codec decoder, 400 tokens (8 s), against upstream `decode_code`: largest sample
difference 1.2e-04 on a signal of RMS 0.258, correlation 0.9999999999.

## Voices

Llasa needs no reference audio: prompted with text alone it invents a speaker, and two
calls never sound like the same person. The presets in `voices.json` pin one down. Each
holds the transcript of an utterance the model generated **and the speech tokens it
emitted for it**; replaying those tokens as an in-context prefix continues that speaker.

Every preset is therefore machine-generated from text alone. **No preset is a recording
of any person**, and each is marked `"synthetic": true`.

| Preset | Language |
|---|---|
| `en_female_a` | English |
| `en_male_a` | English |
| `zh_female_a` | Chinese |
| `zh_male_a` | Chinese |

Cloning from a fresh clip is not supported by this bundle: tokenising one needs XCodec2's
encoder together with the w2v-BERT filterbank front end, which are not included.

## Usage

```python
from phoonnx.model_manager import TTSModelManager
from phoonnx.config import SynthesisConfig

voice = TTSModelManager().load_voice("llasa/HKUST/en/1b")
audio = b"".join(c.audio_int16_bytes for c in voice.synthesize(
    "Dealing with family secrets is never easy.",
    syn_config=SynthesisConfig(extra_params={"voice": "en_female_a"})))
```

Conversion scripts: `scripts/conversion/llasa/` in the phoonnx repository.