phoonnx-llasa β€” Llasa-1B + XCodec2, ONNX

ONNX conversion of HKUSTAudio/Llasa-1B and its codec HKUSTAudio/xcodec2, packaged for phoonnx. This repository holds converted weights only β€” no new training was done.

Llasa (arXiv 2502.04128) is a LLaMA-3.2-1B backbone whose vocabulary was extended with the 65,536 <|s_N|> speech tokens of XCodec2, a single-codebook 16 kHz codec at 50 tokens per second. Prompt it with text and it emits a run of speech tokens; the codec decoder turns those back into a waveform.

Licence

Both upstream repositories are CC BY-NC 4.0, and so is this conversion: non-commercial use only. The licence is upstream HKUST's, not phoonnx's. Model and weights are by the HKUST Audio group; this repository redistributes them in ONNX form and adds nothing but the graph rewrites described below.

Files (llasa-1b-onnx/)

File What it is
model.onnx + model.onnx_data LLaMA backbone, fp32, KV-cached. The graph needs its model.onnx_data sidecar next to it.
xcodec2_decoder.onnx XCodec2 decoder, fp32. Codes in, waveform out.
tokenizer.json The checkpoint's own BPE, copied from upstream.
voices.json Four voice presets. Every one is machine-generated β€” see below.
config.json Self-describing phoonnx voice config.
samples/ One rendering of each preset.

There is no quantized variant. Dynamic int8 (per-tensor and per-channel) and 4-bit MatMulNBits were all built and all failed the greedy-agreement gate against torch: mean absolute logit error of 1.3 to 3.6 and 0 to 25 matching tokens out of 48, against 5e-6 and 48 out of 48 for fp32. Treat the quantized files in other Llasa ONNX repositories with the same suspicion.

Graph contract

model.onnx serves prefill and decode: the same graph with a different past length.

inputs   input_ids                  int64 [1, S]        prompt, or 1 token per step
         attention_mask             int64 [1, P + S]    ones over past and current
         position_ids               int64 [1, S]        absolute positions P .. P+S-1
         past_key_values.<i>.key    fp32  [1, 8, P, 64] i in 0..15
         past_key_values.<i>.value  fp32  [1, 8, P, 64]
outputs  logits                     fp32  [1, 1, 193800]  last position only
         present.<i>.key / present.<i>.value  fp32 [1, 8, P + S, 64]

xcodec2_decoder.onnx takes codes int64 [1, 1, N] and returns audio float32 [1, 320 * N] at 16 kHz.

Two rewrites were needed. Both preserve behaviour, and both were measured:

  • lm_head shares the embedding. Llasa ties the two, but the exporter wrote the 193,800 x 2,048 matrix twice. The head is now Reshape -> Gemm(transB=1) -> Unsqueeze over the embedding initialiser: 7.07 GB becomes 5.48 GB, with bit-identical logits.
  • The codec's ISTFT is real-valued. ONNX cannot trace complex tensors, so the inverse real FFT became two constant cosine/sine matmuls and the overlap-add became a conv_transpose1d with an identity kernel. Against the complex path the largest sample difference is 6e-7.

Only the last position's logits leave the graph. Over a 193,800-wide vocabulary, returning a whole prefill would cost about 78 MB per 100 prompt tokens for a value the sampler never reads.

Parity against torch

transformers fp32 against this graph, English and Chinese prompts, 48 greedy steps each:

Prompt prefill max abs logit diff decode max abs logit diff greedy agreement
English 4.1e-05 4.1e-05 48/48
Chinese 2.7e-05 3.7e-05 48/48

Codec decoder, 400 tokens (8 s), against upstream decode_code: largest sample difference 1.2e-04 on a signal of RMS 0.258, correlation 0.9999999999.

Voices

Llasa needs no reference audio: prompted with text alone it invents a speaker, and two calls never sound like the same person. The presets in voices.json pin one down. Each holds the transcript of an utterance the model generated and the speech tokens it emitted for it; replaying those tokens as an in-context prefix continues that speaker.

Every preset is therefore machine-generated from text alone. No preset is a recording of any person, and each is marked "synthetic": true.

Preset Language
en_female_a English
en_male_a English
zh_female_a Chinese
zh_male_a Chinese

Cloning from a fresh clip is not supported by this bundle: tokenising one needs XCodec2's encoder together with the w2v-BERT filterbank front end, which are not included.

Usage

from phoonnx.model_manager import TTSModelManager
from phoonnx.config import SynthesisConfig

voice = TTSModelManager().load_voice("llasa/HKUST/en/1b")
audio = b"".join(c.audio_int16_bytes for c in voice.synthesize(
    "Dealing with family secrets is never easy.",
    syn_config=SynthesisConfig(extra_params={"voice": "en_female_a"})))

Conversion scripts: scripts/conversion/llasa/ in the phoonnx repository.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for OpenVoiceOS/phoonnx-llasa

Quantized
(5)
this model

Collection including OpenVoiceOS/phoonnx-llasa

Paper for OpenVoiceOS/phoonnx-llasa