Maya1 β ONNX
These are ONNX exports of maya-research/maya1 by Maya Research. Maya1 is a 3B Llama-architecture speech model. It writes SNAC 24 kHz audio codes, and you steer it with a natural-language voice description plus inline emotion tags.
All credit for the model goes to Maya Research. This repo only converts it, under the same Apache-2.0 licence as the original. The exports were made for WinSTT's local TTS engine, but the graphs are plain ONNX Runtime and work from any language.
Files
| File | Rung | Size | Use |
|---|---|---|---|
onnx/model_q8.onnx + model_q8.onnx_data, _data_1, _data_2 |
q8 | 4.85 GB | int8 MatMulNBits (block 32, asymmetric), fp16 embedding table. Best quality (recommended). |
onnx/model_q4.onnx + model_q4.onnx_data, _data_1 |
q4 | 3.69 GB | int4 k-quant body with int8 lm_head and int8 "mixed" layers, plus an fp16 embedding table. About 20 % faster than q8 on CPU. |
tokenizer.json, tokenizer_config.json, special_tokens_map.json |
β | Llama-3 tokenizer plus Maya1's added tokens (unchanged) | |
config.json, generation_config.json, emotions.txt |
β | Original config, sampling defaults and tag list |
The SNAC decoder is not duplicated here. Use
onnx-community/snac_24khz-ONNX
onnx/decoder_model.onnx (52.6 MB, fp32).
Each graph keeps its weights in external-data shards next to it (model_X.onnx_data,
model_X.onnx_data_1, β¦), each under 2 GB. Keep the shards in the same folder as the graph.
Prompt format
Token ids (Llama-3 vocab + Maya1's 28 k SNAC code tokens):
[128000, 128259, 128000] # BOS, SOH, BOS
+ tokenize('<description="{DESCRIPTION}"> {TEXT}', add_special_tokens=False)
+ [128009, 128260, 128261, 128257] # EOT, EOH, SOA, SOS
This is exactly what the reference build_prompt string (soh + bos + '<description="β¦"> text' + eot + eoh + soa + sos) tokenizes to. The tokenizer's template adds the leading BOS.
Decoding
- Sample until token 128258 (end of speech).
- Audio tokens are
128266 + slot * 4096 + code. Every 7 tokens form one SNAC frame ([L1, L2a, L3a, L3b, L2b, L3c, L3d]), so about 12 frames make 1 s of audio. - Unpack each code with
(id - 128266) % 4096, decode the 3 code layers with SNAC, and drop the first 2048 output samples (warm-up). - Reference settings: temperature 0.4, top-p 0.9, repetition penalty 1.1, and keep 128258 masked for the first 28 new tokens.
- Restricting sampling to the SNAC band plus 128258 gives the same output as full-vocab sampling. We checked this token-for-token on q8.
Two sampled failure modes worth guarding against
- A run that never emits 128258.
- Reading the description aloud before the text.
We saw both only in 4-bit renders, about 1 in 10 per sentence before the fix. A per-sentence token budget catches both: WinSTT allows 210 tokens plus 7 per character of text, while clean renders use 3β5 tokens per character. When a render hits the budget, re-roll it with a new seed.
Voice descriptions
Describe the voice the way you would brief a voice actor: age, gender, accent, pitch, timbre, pacing and role. The model card's template works best:
Realistic female voice in the 30s age with an american accent. Normal pitch, warm timbre, conversational pacing, neutral tone at medium intensity.Realistic male voice in the 40s age with a british accent. Low pitch, deep timbre, slow pacing, storyteller role, neutral tone at medium intensity.Creative, ai_machine_voice character. Male voice in their 20s with a american accent. Normal pitch, robotic timbre, conversational pacing, neutral tone at med intensity.Dark villain character, Male voice in their 40s with a British accent. low pitch, gravelly timbre, slow pacing, angry tone at high intensity.
Keep the XML-attribute form <description="β¦">. Other wrappings make the model more likely to read
the description out loud.
Emotion tags
Put them inline, where the effect belongs: I cannot believe you actually did that <laugh> it is the funniest thingβ¦
<laugh> <laugh_harder> <sigh> <chuckle> <gasp> <angry> <excited> <whisper> <cry> <scream> <sing> <snort> <exhale> <gulp> <giggle> <sarcastic> <curious>. Each tag is a single token.
We measured each tag with q8 on the same sentence, untagged versus tagged (WinSTT sampler, seeded per prompt). The transcripts stay word-perfect in every case (whisper WER 0.000 over all 8 clips). The tag adds sound, not words:
| Tag | Duration | RMS |
|---|---|---|
<laugh> |
4.35 β 6.66 s | 0.111 β 0.093 |
<sigh> |
3.93 β 6.91 s | 0.107 β 0.076 |
<whisper> |
4.10 β 4.52 s | 0.103 β 0.071 (β31 %) |
<gasp> |
2.90 β 4.27 s | 0.110 β 0.096 |
Graph I/O (both rungs)
| Name | Type | Shape |
|---|---|---|
input_ids |
int64 | [batch, seq] |
attention_mask |
int64 | [batch, past + seq] |
past_key_values.{0..27}.key / .value |
float32 | [batch, 8, past, 128] |
logits |
float32 | [batch, seq, 156960] |
present.{0..27}.key / .value |
float32 | [batch, 8, past + seq, 128] |
There is no position_ids input: positions are derived from attention_mask, and rotary embeddings are fused into
GroupQueryAttention. The layout is the transformers.js / onnx-community decoder layout, with an ordinary growing KV
cache. Start with empty [1, 8, 0, 128] past tensors.
Parity vs PyTorch
The test is a forced KV loop: a 59-token prompt plus 32 decode steps, feeding the same token stream to both models.
The reference is transformers in bf16. We compare logits over the SNAC band:
| Graph | max-abs | cosine mean / min | argmax agreement |
|---|---|---|---|
| q8 | 0.71 | 0.99999 / 0.99989 | 28 / 33 |
| q4 | 6.83 | 0.99897 / 0.98731 | 16 / 33 |
Argmax agreement is low for q4 by design of the signal, not the graph: SNAC-band distributions are flat, and the model is sampled at temperature 0.4, not decoded greedily. Cosine is the meaningful figure.
An earlier plain round-to-nearest int4 export (4-bit embedding and lm_head) reached cosine 0.998. Even so, it
read the voice description aloud on the gate sentence (WER 2.1). The mixed-precision k-quant recipe above fixes
that.
Intelligibility (WER)
We ran 10 description Γ sentence cases (realistic and creative voices, US, UK and Indian accents) through WinSTT's
Rust engine, then transcribed each clip with faster-whisper base.en:
| Rung | WER | Word-perfect |
|---|---|---|
| q8 | 0.000 | 10 / 10 |
| q4 | 0.039 | 8 / 10 (single-word slips, e.g. "clock-breaker") |
Speed
These figures are from an i9-12900KF on the onnxruntime CPU EP, with other jobs running on the machine. Each sentence is about 3β6 s of audio. Real time needs about 86 tokens/s.
| Rung | RTF (compute s / audio s) |
|---|---|
| q4 | 6.9β9.1 |
| q8 | 10.0β11.2 |
On CPU this is a "render, then play" model, not a streaming one.
We also built a DirectML int4/fp16 graph (onnxruntime-genai -e dml) and did not publish it. Driven from plain
onnxruntime with IoBinding on an RTX 3080 Ti, it decoded at 0.3β2.4 tok/s, slower than CPU.
How it was exported
- The onnxruntime-genai 0.17.1 model builder generated both
graphs (opset 22, GroupQueryAttention with fused rotary, MatMulNBits block 32) from the original bf16
safetensors:
q8:-p int8 -e cpu --extra_options block_size=32 is_symmetric=false shared_embeddings=false. The embedding table was then cast to fp16.q4:-p int4 -e cpu --extra_options block_size=32 is_symmetric=false shared_embeddings=false algo_config=k_quant_mixed. This means a k-quant int4 body, withlm_headand llama.cpp's "mixed" sensitive layers (qkv/v/down in the first and last eighth of the layers, plus every third layer) at int8. The embedding table was then cast to fp16.
- Post-processing pinned the KV head dim (128) in the I/O shapes, renamed the graphs to
model_X.onnx, and streamed the initializers into external-data shards under 2 GB. - Validation covered:
- logits parity (above) against
transformers; - full synthesis, SNAC decode and faster-whisper WER through both a Python harness and WinSTT's Rust engine;
- the emotion-tag A/B.
- logits parity (above) against
Citation
@misc{maya1voice2025,
title={Maya1: Open Source Voice AI with Emotional Intelligence},
author={Maya Research},
year={2025},
publisher={Hugging Face},
howpublished={\url{https://huggingface.co/maya-research/maya1}},
}
- Downloads last month
- 52
Model tree for Masterx/maya1-ONNX
Base model
maya-research/maya1