phoonnx-llasa / README.md
Jarbas's picture
Upload README.md with huggingface_hub
05a6e70 verified
|
Raw
History Blame Contribute Delete
5.38 kB
---
license: cc-by-nc-4.0
language:
- en
- zh
library_name: phoonnx
pipeline_tag: text-to-speech
tags:
- onnx
- tts
- llasa
- xcodec2
- codec-lm
base_model:
- HKUSTAudio/Llasa-1B
- HKUSTAudio/xcodec2
---
# phoonnx-llasa — Llasa-1B + XCodec2, ONNX
ONNX conversion of [HKUSTAudio/Llasa-1B](https://huggingface.co/HKUSTAudio/Llasa-1B)
and its codec [HKUSTAudio/xcodec2](https://huggingface.co/HKUSTAudio/xcodec2), packaged
for [phoonnx](https://github.com/TigreGotico/phoonnx). This repository holds converted
weights only — no new training was done.
Llasa (arXiv [2502.04128](https://arxiv.org/abs/2502.04128)) is a LLaMA-3.2-1B backbone
whose vocabulary was extended with the 65,536 `<|s_N|>` speech tokens of XCodec2, a
single-codebook 16 kHz codec at 50 tokens per second. Prompt it with text and it emits a
run of speech tokens; the codec decoder turns those back into a waveform.
## Licence
Both upstream repositories are **CC BY-NC 4.0**, and so is this conversion:
non-commercial use only. The licence is upstream HKUST's, not phoonnx's. Model and
weights are by the HKUST Audio group; this repository redistributes them in ONNX form
and adds nothing but the graph rewrites described below.
## Files (`llasa-1b-onnx/`)
| File | What it is |
|---|---|
| `model.onnx` + `model.onnx_data` | LLaMA backbone, fp32, KV-cached. **The graph needs its `model.onnx_data` sidecar next to it.** |
| `xcodec2_decoder.onnx` | XCodec2 decoder, fp32. Codes in, waveform out. |
| `tokenizer.json` | The checkpoint's own BPE, copied from upstream. |
| `voices.json` | Four voice presets. Every one is machine-generated — see below. |
| `config.json` | Self-describing phoonnx voice config. |
| `samples/` | One rendering of each preset. |
There is **no quantized variant**. Dynamic int8 (per-tensor and per-channel) and 4-bit
`MatMulNBits` were all built and all failed the greedy-agreement gate against torch:
mean absolute logit error of 1.3 to 3.6 and 0 to 25 matching tokens out of 48, against
5e-6 and 48 out of 48 for fp32. Treat the quantized files in other Llasa ONNX
repositories with the same suspicion.
## Graph contract
`model.onnx` serves prefill and decode: the same graph with a different past length.
```
inputs input_ids int64 [1, S] prompt, or 1 token per step
attention_mask int64 [1, P + S] ones over past and current
position_ids int64 [1, S] absolute positions P .. P+S-1
past_key_values.<i>.key fp32 [1, 8, P, 64] i in 0..15
past_key_values.<i>.value fp32 [1, 8, P, 64]
outputs logits fp32 [1, 1, 193800] last position only
present.<i>.key / present.<i>.value fp32 [1, 8, P + S, 64]
```
`xcodec2_decoder.onnx` takes `codes` int64 `[1, 1, N]` and returns `audio` float32
`[1, 320 * N]` at 16 kHz.
Two rewrites were needed. Both preserve behaviour, and both were measured:
* **`lm_head` shares the embedding.** Llasa ties the two, but the exporter wrote the
193,800 x 2,048 matrix twice. The head is now `Reshape -> Gemm(transB=1) -> Unsqueeze`
over the embedding initialiser: 7.07 GB becomes 5.48 GB, with bit-identical logits.
* **The codec's ISTFT is real-valued.** ONNX cannot trace complex tensors, so the inverse
real FFT became two constant cosine/sine matmuls and the overlap-add became a
`conv_transpose1d` with an identity kernel. Against the complex path the largest sample
difference is 6e-7.
Only the last position's logits leave the graph. Over a 193,800-wide vocabulary,
returning a whole prefill would cost about 78 MB per 100 prompt tokens for a value the
sampler never reads.
## Parity against torch
`transformers` fp32 against this graph, English and Chinese prompts, 48 greedy steps each:
| Prompt | prefill max abs logit diff | decode max abs logit diff | greedy agreement |
|---|---|---|---|
| English | 4.1e-05 | 4.1e-05 | 48/48 |
| Chinese | 2.7e-05 | 3.7e-05 | 48/48 |
Codec decoder, 400 tokens (8 s), against upstream `decode_code`: largest sample
difference 1.2e-04 on a signal of RMS 0.258, correlation 0.9999999999.
## Voices
Llasa needs no reference audio: prompted with text alone it invents a speaker, and two
calls never sound like the same person. The presets in `voices.json` pin one down. Each
holds the transcript of an utterance the model generated **and the speech tokens it
emitted for it**; replaying those tokens as an in-context prefix continues that speaker.
Every preset is therefore machine-generated from text alone. **No preset is a recording
of any person**, and each is marked `"synthetic": true`.
| Preset | Language |
|---|---|
| `en_female_a` | English |
| `en_male_a` | English |
| `zh_female_a` | Chinese |
| `zh_male_a` | Chinese |
Cloning from a fresh clip is not supported by this bundle: tokenising one needs XCodec2's
encoder together with the w2v-BERT filterbank front end, which are not included.
## Usage
```python
from phoonnx.model_manager import TTSModelManager
from phoonnx.config import SynthesisConfig
voice = TTSModelManager().load_voice("llasa/HKUST/en/1b")
audio = b"".join(c.audio_int16_bytes for c in voice.synthesize(
"Dealing with family secrets is never easy.",
syn_config=SynthesisConfig(extra_params={"voice": "en_female_a"})))
```
Conversion scripts: `scripts/conversion/llasa/` in the phoonnx repository.