license: cc-by-nc-4.0
language:
- en
- zh
library_name: phoonnx
pipeline_tag: text-to-speech
tags:
- onnx
- tts
- llasa
- xcodec2
- codec-lm
base_model:
- HKUSTAudio/Llasa-1B
- HKUSTAudio/xcodec2
phoonnx-llasa — Llasa-1B + XCodec2, ONNX
ONNX conversion of HKUSTAudio/Llasa-1B and its codec HKUSTAudio/xcodec2, packaged for phoonnx. This repository holds converted weights only — no new training was done.
Llasa (arXiv 2502.04128) is a LLaMA-3.2-1B backbone
whose vocabulary was extended with the 65,536 <|s_N|> speech tokens of XCodec2, a
single-codebook 16 kHz codec at 50 tokens per second. Prompt it with text and it emits a
run of speech tokens; the codec decoder turns those back into a waveform.
Licence
Both upstream repositories are CC BY-NC 4.0, and so is this conversion: non-commercial use only. The licence is upstream HKUST's, not phoonnx's. Model and weights are by the HKUST Audio group; this repository redistributes them in ONNX form and adds nothing but the graph rewrites described below.
Files (llasa-1b-onnx/)
| File | What it is |
|---|---|
model.onnx + model.onnx_data |
LLaMA backbone, fp32, KV-cached. The graph needs its model.onnx_data sidecar next to it. |
xcodec2_decoder.onnx |
XCodec2 decoder, fp32. Codes in, waveform out. |
tokenizer.json |
The checkpoint's own BPE, copied from upstream. |
voices.json |
Four voice presets. Every one is machine-generated — see below. |
config.json |
Self-describing phoonnx voice config. |
samples/ |
One rendering of each preset. |
There is no quantized variant. Dynamic int8 (per-tensor and per-channel) and 4-bit
MatMulNBits were all built and all failed the greedy-agreement gate against torch:
mean absolute logit error of 1.3 to 3.6 and 0 to 25 matching tokens out of 48, against
5e-6 and 48 out of 48 for fp32. Treat the quantized files in other Llasa ONNX
repositories with the same suspicion.
Graph contract
model.onnx serves prefill and decode: the same graph with a different past length.
inputs input_ids int64 [1, S] prompt, or 1 token per step
attention_mask int64 [1, P + S] ones over past and current
position_ids int64 [1, S] absolute positions P .. P+S-1
past_key_values.<i>.key fp32 [1, 8, P, 64] i in 0..15
past_key_values.<i>.value fp32 [1, 8, P, 64]
outputs logits fp32 [1, 1, 193800] last position only
present.<i>.key / present.<i>.value fp32 [1, 8, P + S, 64]
xcodec2_decoder.onnx takes codes int64 [1, 1, N] and returns audio float32
[1, 320 * N] at 16 kHz.
Two rewrites were needed. Both preserve behaviour, and both were measured:
lm_headshares the embedding. Llasa ties the two, but the exporter wrote the 193,800 x 2,048 matrix twice. The head is nowReshape -> Gemm(transB=1) -> Unsqueezeover the embedding initialiser: 7.07 GB becomes 5.48 GB, with bit-identical logits.- The codec's ISTFT is real-valued. ONNX cannot trace complex tensors, so the inverse
real FFT became two constant cosine/sine matmuls and the overlap-add became a
conv_transpose1dwith an identity kernel. Against the complex path the largest sample difference is 6e-7.
Only the last position's logits leave the graph. Over a 193,800-wide vocabulary, returning a whole prefill would cost about 78 MB per 100 prompt tokens for a value the sampler never reads.
Parity against torch
transformers fp32 against this graph, English and Chinese prompts, 48 greedy steps each:
| Prompt | prefill max abs logit diff | decode max abs logit diff | greedy agreement |
|---|---|---|---|
| English | 4.1e-05 | 4.1e-05 | 48/48 |
| Chinese | 2.7e-05 | 3.7e-05 | 48/48 |
Codec decoder, 400 tokens (8 s), against upstream decode_code: largest sample
difference 1.2e-04 on a signal of RMS 0.258, correlation 0.9999999999.
Voices
Llasa needs no reference audio: prompted with text alone it invents a speaker, and two
calls never sound like the same person. The presets in voices.json pin one down. Each
holds the transcript of an utterance the model generated and the speech tokens it
emitted for it; replaying those tokens as an in-context prefix continues that speaker.
Every preset is therefore machine-generated from text alone. No preset is a recording
of any person, and each is marked "synthetic": true.
| Preset | Language |
|---|---|
en_female_a |
English |
en_male_a |
English |
zh_female_a |
Chinese |
zh_male_a |
Chinese |
Cloning from a fresh clip is not supported by this bundle: tokenising one needs XCodec2's encoder together with the w2v-BERT filterbank front end, which are not included.
Usage
from phoonnx.model_manager import TTSModelManager
from phoonnx.config import SynthesisConfig
voice = TTSModelManager().load_voice("llasa/HKUST/en/1b")
audio = b"".join(c.audio_int16_bytes for c in voice.synthesize(
"Dealing with family secrets is never easy.",
syn_config=SynthesisConfig(extra_params={"voice": "en_female_a"})))
Conversion scripts: scripts/conversion/llasa/ in the phoonnx repository.