Maya1 β€” ONNX

These are ONNX exports of maya-research/maya1 by Maya Research. Maya1 is a 3B Llama-architecture speech model. It writes SNAC 24 kHz audio codes, and you steer it with a natural-language voice description plus inline emotion tags.

All credit for the model goes to Maya Research. This repo only converts it, under the same Apache-2.0 licence as the original. The exports were made for WinSTT's local TTS engine, but the graphs are plain ONNX Runtime and work from any language.

Files

File Rung Size Use
onnx/model_q8.onnx + model_q8.onnx_data, _data_1, _data_2 q8 4.85 GB int8 MatMulNBits (block 32, asymmetric), fp16 embedding table. Best quality (recommended).
onnx/model_q4.onnx + model_q4.onnx_data, _data_1 q4 3.69 GB int4 k-quant body with int8 lm_head and int8 "mixed" layers, plus an fp16 embedding table. About 20 % faster than q8 on CPU.
tokenizer.json, tokenizer_config.json, special_tokens_map.json β€” Llama-3 tokenizer plus Maya1's added tokens (unchanged)
config.json, generation_config.json, emotions.txt β€” Original config, sampling defaults and tag list

The SNAC decoder is not duplicated here. Use onnx-community/snac_24khz-ONNX onnx/decoder_model.onnx (52.6 MB, fp32).

Each graph keeps its weights in external-data shards next to it (model_X.onnx_data, model_X.onnx_data_1, …), each under 2 GB. Keep the shards in the same folder as the graph.

Prompt format

Token ids (Llama-3 vocab + Maya1's 28 k SNAC code tokens):

[128000, 128259, 128000]                         # BOS, SOH, BOS
+ tokenize('<description="{DESCRIPTION}"> {TEXT}', add_special_tokens=False)
+ [128009, 128260, 128261, 128257]               # EOT, EOH, SOA, SOS

This is exactly what the reference build_prompt string (soh + bos + '<description="…"> text' + eot + eoh + soa + sos) tokenizes to. The tokenizer's template adds the leading BOS.

Decoding

  • Sample until token 128258 (end of speech).
  • Audio tokens are 128266 + slot * 4096 + code. Every 7 tokens form one SNAC frame ([L1, L2a, L3a, L3b, L2b, L3c, L3d]), so about 12 frames make 1 s of audio.
  • Unpack each code with (id - 128266) % 4096, decode the 3 code layers with SNAC, and drop the first 2048 output samples (warm-up).
  • Reference settings: temperature 0.4, top-p 0.9, repetition penalty 1.1, and keep 128258 masked for the first 28 new tokens.
  • Restricting sampling to the SNAC band plus 128258 gives the same output as full-vocab sampling. We checked this token-for-token on q8.

Two sampled failure modes worth guarding against

  • A run that never emits 128258.
  • Reading the description aloud before the text.

We saw both only in 4-bit renders, about 1 in 10 per sentence before the fix. A per-sentence token budget catches both: WinSTT allows 210 tokens plus 7 per character of text, while clean renders use 3–5 tokens per character. When a render hits the budget, re-roll it with a new seed.

Voice descriptions

Describe the voice the way you would brief a voice actor: age, gender, accent, pitch, timbre, pacing and role. The model card's template works best:

  • Realistic female voice in the 30s age with an american accent. Normal pitch, warm timbre, conversational pacing, neutral tone at medium intensity.
  • Realistic male voice in the 40s age with a british accent. Low pitch, deep timbre, slow pacing, storyteller role, neutral tone at medium intensity.
  • Creative, ai_machine_voice character. Male voice in their 20s with a american accent. Normal pitch, robotic timbre, conversational pacing, neutral tone at med intensity.
  • Dark villain character, Male voice in their 40s with a British accent. low pitch, gravelly timbre, slow pacing, angry tone at high intensity.

Keep the XML-attribute form <description="…">. Other wrappings make the model more likely to read the description out loud.

Emotion tags

Put them inline, where the effect belongs: I cannot believe you actually did that <laugh> it is the funniest thing…

<laugh> <laugh_harder> <sigh> <chuckle> <gasp> <angry> <excited> <whisper> <cry> <scream> <sing> <snort> <exhale> <gulp> <giggle> <sarcastic> <curious>. Each tag is a single token.

We measured each tag with q8 on the same sentence, untagged versus tagged (WinSTT sampler, seeded per prompt). The transcripts stay word-perfect in every case (whisper WER 0.000 over all 8 clips). The tag adds sound, not words:

Tag Duration RMS
<laugh> 4.35 β†’ 6.66 s 0.111 β†’ 0.093
<sigh> 3.93 β†’ 6.91 s 0.107 β†’ 0.076
<whisper> 4.10 β†’ 4.52 s 0.103 β†’ 0.071 (βˆ’31 %)
<gasp> 2.90 β†’ 4.27 s 0.110 β†’ 0.096

Graph I/O (both rungs)

Name Type Shape
input_ids int64 [batch, seq]
attention_mask int64 [batch, past + seq]
past_key_values.{0..27}.key / .value float32 [batch, 8, past, 128]
logits float32 [batch, seq, 156960]
present.{0..27}.key / .value float32 [batch, 8, past + seq, 128]

There is no position_ids input: positions are derived from attention_mask, and rotary embeddings are fused into GroupQueryAttention. The layout is the transformers.js / onnx-community decoder layout, with an ordinary growing KV cache. Start with empty [1, 8, 0, 128] past tensors.

Parity vs PyTorch

The test is a forced KV loop: a 59-token prompt plus 32 decode steps, feeding the same token stream to both models. The reference is transformers in bf16. We compare logits over the SNAC band:

Graph max-abs cosine mean / min argmax agreement
q8 0.71 0.99999 / 0.99989 28 / 33
q4 6.83 0.99897 / 0.98731 16 / 33

Argmax agreement is low for q4 by design of the signal, not the graph: SNAC-band distributions are flat, and the model is sampled at temperature 0.4, not decoded greedily. Cosine is the meaningful figure.

An earlier plain round-to-nearest int4 export (4-bit embedding and lm_head) reached cosine 0.998. Even so, it read the voice description aloud on the gate sentence (WER 2.1). The mixed-precision k-quant recipe above fixes that.

Intelligibility (WER)

We ran 10 description Γ— sentence cases (realistic and creative voices, US, UK and Indian accents) through WinSTT's Rust engine, then transcribed each clip with faster-whisper base.en:

Rung WER Word-perfect
q8 0.000 10 / 10
q4 0.039 8 / 10 (single-word slips, e.g. "clock-breaker")

Speed

These figures are from an i9-12900KF on the onnxruntime CPU EP, with other jobs running on the machine. Each sentence is about 3–6 s of audio. Real time needs about 86 tokens/s.

Rung RTF (compute s / audio s)
q4 6.9–9.1
q8 10.0–11.2

On CPU this is a "render, then play" model, not a streaming one.

We also built a DirectML int4/fp16 graph (onnxruntime-genai -e dml) and did not publish it. Driven from plain onnxruntime with IoBinding on an RTX 3080 Ti, it decoded at 0.3–2.4 tok/s, slower than CPU.

How it was exported

  1. The onnxruntime-genai 0.17.1 model builder generated both graphs (opset 22, GroupQueryAttention with fused rotary, MatMulNBits block 32) from the original bf16 safetensors:
    • q8: -p int8 -e cpu --extra_options block_size=32 is_symmetric=false shared_embeddings=false. The embedding table was then cast to fp16.
    • q4: -p int4 -e cpu --extra_options block_size=32 is_symmetric=false shared_embeddings=false algo_config=k_quant_mixed. This means a k-quant int4 body, with lm_head and llama.cpp's "mixed" sensitive layers (qkv/v/down in the first and last eighth of the layers, plus every third layer) at int8. The embedding table was then cast to fp16.
  2. Post-processing pinned the KV head dim (128) in the I/O shapes, renamed the graphs to model_X.onnx, and streamed the initializers into external-data shards under 2 GB.
  3. Validation covered:
    • logits parity (above) against transformers;
    • full synthesis, SNAC decode and faster-whisper WER through both a Python harness and WinSTT's Rust engine;
    • the emotion-tag A/B.

Citation

@misc{maya1voice2025,
  title={Maya1: Open Source Voice AI with Emotional Intelligence},
  author={Maya Research},
  year={2025},
  publisher={Hugging Face},
  howpublished={\url{https://huggingface.co/maya-research/maya1}},
}
Downloads last month
52
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for Masterx/maya1-ONNX

Quantized
(11)
this model