phoonnx-omnivoice / README.md
Jarbas's picture
ONNX build: bidirectional backbone re-export + verified Higgs codec graphs
4a90578 verified
|
Raw
History Blame Contribute Delete
5.67 kB
---
license: cc-by-nc-4.0
base_model:
- k2-fsa/OmniVoice
pipeline_tag: text-to-speech
library_name: phoonnx
tags:
- onnx
- phoonnx
- text-to-speech
- zero-shot
- voice-cloning
- multilingual
---
# phoonnx-omnivoice
ONNX build of [**OmniVoice**](https://github.com/k2-fsa/OmniVoice) by k2-fsa (Xiaomi
Corp., authors Han Zhu et al.), packaged for
[phoonnx](https://github.com/TigreGotico/phoonnx).
OmniVoice is a zero-shot text-to-speech model for 600+ languages. It is a **masked
diffusion language model**: a Qwen3-0.6B backbone writes eight streams of Higgs Audio V2
codec tokens, starting from an all-MASK grid and unmasking the most confident slots over
32 steps. Attention is **bidirectional**, and every step is a full-sequence forward.
## Why this build exists
An earlier community export, `onnx-community/OmniVoice-Onnx`, runs the backbone through
the ONNX Runtime contrib operator `com.microsoft::GroupQueryAttention`. That operator is
**unconditionally causal**. OmniVoice is not a causal model, so the export computes the
wrong hidden states. We measured this against upstream PyTorch on a fixed input:
| Graph | Reference | Result |
|---|---|---|
| `audio_embeddings_encoder` (community, fp32) | upstream torch | cos **1.0000000**, rel 6e-4 β€” correct |
| `audio_heads_decoder` (community, fp16) | upstream torch | cos **1.0000000**, rel 2e-4 β€” correct |
| `llm_decoder` (community, int4) | upstream torch, **bidirectional** | cos **0.954** β€” wrong |
| `llm_decoder` (community, int4) | upstream torch, **causal** | cos **0.99945** β€” matches a causal model |
| community chain, end to end | upstream torch | **18.11 %** greedy-token agreement |
The community `llm_decoder` reproduces a *causal* OmniVoice, which is a different model.
`omnivoice_backbone.onnx` here is a fresh export from the PyTorch checkpoint with the
bidirectional mask kept intact:
| Graph | Reference | Result |
|---|---|---|
| `omnivoice_backbone.onnx` (fp32) | upstream torch | cos **1.0000000**, rel 2.9e-7, **100.00 %** greedy-token agreement |
| full sampler, 32 steps, greedy | upstream `_generate_iterative` | **100.00 %** codec-token agreement |
The community **Higgs codec graphs are exact** and are mirrored here unchanged:
| Graph | Reference | Result |
|---|---|---|
| `acoustic_encoder` + `semantic_encoder` + `quantizer_encoder` | upstream torch encode | **100.00 %** codec-code agreement (all 8 codebooks) |
| `higgs_decoder` | upstream torch decode | max abs diff **0.0** (bit-identical) |
No quantized variant of the backbone is published: the int4 community build was rejected
on the agreement test above, and we have not yet produced a quantized export that passes.
## Files
| File | What it is |
|---|---|
| `omnivoice_backbone.onnx` (+ `.onnx_data`) | Qwen3 backbone + audio embeddings + audio heads, one graph. `(input_ids[B,8,S] int64, audio_mask[B,S] bool) -> logits[B,8,S,1025]`. Bidirectional; no KV cache. |
| `acoustic_encoder.onnx` | reference wav @24 kHz `(1,1,T)` -> acoustic features `(1,256,T')` |
| `semantic_encoder.onnx` | reference wav @16 kHz `(1,T)` -> semantic features `(1,768,T')` |
| `quantizer_encoder.onnx` | acoustic + semantic -> reference codes `(8,1,T')` |
| `higgs_decoder.onnx` | codes `(8,1,T')` -> waveform @24 kHz |
| `tokenizer.json`, `tokenizer_config.json` | the model's own Qwen3 subword BPE |
| `config.json` | names the phoonnx engine and the graph roles |
Classifier-free guidance runs the backbone twice per step β€” once over the full prompt and
once over the target span alone. Upstream batches both rows behind a `[2B,1,S,S]` block
mask; for a single item that is the same as two forwards of different lengths.
## Usage
```python
from phoonnx.model_manager import TTSModelManager
voice = TTSModelManager().load_voice("omnivoice/en")
audio = voice.synthesize(
"Machine learning models can now speak in hundreds of languages.",
speaker_reference="reference.wav",
speaker_reference_text="Transcription of the reference clip.",
)
```
`speaker_reference_text` is not optional in practice: OmniVoice joins the reference
transcription to the target text into a single prompt string, and cloning is noticeably
worse without it.
## Verified languages
Word/character error rate on FLEURS, scored with OpenVoiceOS ONNX ASR models and always
reported against a floor β€” the same ASR run over the **real human** FLEURS recording.
See the phoonnx pull request for the full table. Only a sample of the model's 600+
claimed languages has been measured; the rest are **untested**, not verified.
## Licensing and attribution
OmniVoice's **code** is Apache-2.0. Its **released weights** are **CC-BY-NC** because of
their training data (Emilia and others), and this repository inherits that: the ONNX
graphs are a format conversion of those weights. Non-commercial use only.
Model and method: Zhu Han, Ye Lingxuan, Kang Wei, Yao Zengwei, Guo Liyong, Kuang Fangjun,
Han Zhifeng, Zhuang Weiji, Lin Long and Povey Daniel β€” *OmniVoice: Towards Omnilingual
Zero-Shot Text-to-Speech with Diffusion Language Models*, arXiv 2604.00688, 2026.
The Higgs codec graphs are mirrored from `onnx-community/OmniVoice-Onnx`.
Do not use this model for voice cloning without the speaker's consent, or for
impersonation or fraud.
```bibtex
@article{zhu2026omnivoice,
title={OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models},
author={Zhu, Han and Ye, Lingxuan and Kang, Wei and Yao, Zengwei and Guo, Liyong and
Kuang, Fangjun and Han, Zhifeng and Zhuang, Weiji and Lin, Long and Povey, Daniel},
journal={arXiv preprint arXiv:2604.00688},
year={2026}
}
```