File size: 5,393 Bytes
a227e25 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 | ---
license: apache-2.0
language:
- zh
- en
- ja
- ko
- de
- fr
- ru
- pt
- es
- it
pipeline_tag: text-to-speech
library_name: llama.cpp
base_model: Qwen/Qwen3-TTS-12Hz-0.6B-Base
tags:
- qwen3-tts
- spacemit
- k1
- k3
- gguf
- onnx
- onnxruntime
- voice-cloning
---
# Qwen3-TTS-0.6B for SpaceMIT K1/K3
This repository contains a SpaceMIT-optimized split-runtime export of
`Qwen/Qwen3-TTS-12Hz-0.6B-Base`. The text embedding and codec stages run with
ONNX Runtime and the SpaceMIT Execution Provider, while the talker and code
predictor run through the SpaceMIT llama.cpp/ggml backend.
The bundle generates 24 kHz mono PCM16 WAV audio and includes one default
speaker embedding for evaluation.
## Validation Status
- K3: validated with Chinese, English, and mixed Chinese-English requests.
- K1: the current configuration is identical to K3, but has not yet completed
independent board acceptance.
- Functional runtime: SpaceMIT llama.cpp `v0.1.7` (`c9af964b5`).
- Recommended performance runtime: a release containing commit `787e5fcf9`
([spacemit-com/llama.cpp#29](https://github.com/spacemit-com/llama.cpp/pull/29)).
SpaceMIT llama.cpp `v0.1.7` is functionally compatible but does not contain the
persistent shared-threadpool optimization. In a controlled K3 A/B test, that
optimization reduced median warm RTF from approximately `1.1113` to `0.9006`
without changing the generated WAV bytes. Performance depends on text,
speaker, board configuration, system load, and runtime revision.
## Files
| File | Purpose | Precision/format |
| --- | --- | --- |
| `Qwen3-TTS-0.6B-talker-q8_0.gguf` | Autoregressive talker | GGUF Q8_0 |
| `Qwen3-TTS-0.6B-code-predictor-q4_0.gguf` | Multi-codebook predictor | GGUF Q4_0 |
| `Qwen3-TTS-0.6B-text-embed-proj.fp32.onnx` | Text embedding and projection | ONNX FP32 |
| `Qwen3-TTS-0.6B-codec-decoder-t50.dynq.onnx` | Codec decoder, bucket 50 | Dynamically quantized ONNX |
| `Qwen3-TTS-0.6B-tokenizer.gguf` | Tokenizer metadata | GGUF |
| `Qwen3-TTS-0.6B-aux.gguf` | Runtime auxiliary tensors | GGUF |
| `default.spk.bin` | Default speaker embedding | Raw float32 `[1024]` |
| `configs/K1/config.json` | K1 runtime configuration | JSON |
| `configs/K3/config.json` | K3 runtime configuration | JSON |
The FP32 text embedding model is intentionally named with its precision. It is
the largest file in the bundle and has a significant effect on resident memory
and cold-start cost.
## Runtime Requirements
- A SpaceMIT K1/K3 Linux image with ONNX Runtime and SpaceMIT EP installed.
- [SpaceMIT llama.cpp](https://github.com/spacemit-com/llama.cpp) `v0.1.7` or
newer.
- A future release containing commit `787e5fcf9` is recommended for the
validated shared-threadpool performance path.
The llama.cpp release archive contains llama/ggml/mtmd binaries and libraries.
It does not bundle ONNX Runtime or SpaceMIT EP.
## Run on K3
Download and extract the runtime:
```bash
wget https://github.com/spacemit-com/llama.cpp/releases/download/v0.1.7/spacemit-llama.cpp.riscv64.0.1.7.tar.gz
tar -xzf spacemit-llama.cpp.riscv64.0.1.7.tar.gz
```
Start the OpenAI-compatible speech service:
```bash
export LLAMA_ROOT=$PWD/spacemit-llama.cpp.riscv64.0.1.7
export MODEL_ROOT=$PWD/Qwen3-TTS-0.6B
export LD_LIBRARY_PATH=$LLAMA_ROOT/lib:${LD_LIBRARY_PATH:-}
$LLAMA_ROOT/bin/llama-server \
--media-backend smt \
--smt-config-dir $MODEL_ROOT/configs/K3 \
--host 127.0.0.1 \
--port 8090 \
--no-ui
```
Generate speech:
```bash
curl http://127.0.0.1:8090/v1/audio/speech \
-H 'Content-Type: application/json' \
-d '{
"model": "qwen3-tts",
"input": "你好,这是 Qwen3-TTS 在 K3 上的语音合成测试。",
"voice": "default",
"response_format": "wav"
}' \
-o qwen3_tts.wav
```
To use the K1 configuration, replace `configs/K3` with `configs/K1` after
validating the target board image and runtime packages.
## Speaker Selection
The default speaker is loaded from `default.spk.bin`. A compatible raw
float32 `[1024]` speaker embedding can be selected at server startup:
```bash
$LLAMA_ROOT/bin/llama-server \
--media-backend smt \
--smt-config-dir $MODEL_ROOT/configs/K3 \
--tts-speaker-file /path/to/custom.spk.bin \
--host 127.0.0.1 \
--port 8090 \
--no-ui
```
The bundled default embedding was derived from a generic female reference
approved by the project owner for redistribution. Obtain consent before
creating or distributing embeddings derived from another person's voice.
## Limitations
- The validated output format is 24 kHz mono PCM16 WAV.
- Runtime performance is not determined by model files alone; use the
recommended llama.cpp revision for the optimized path.
- The current text embedding model is FP32 rather than quantized.
- K1 uses the same initial configuration as K3 but still requires independent
acceptance testing.
## Revisions
- Export source revision: `58a4795010ba25b34966a56d07499d496ff9cd6f`
- Initial llama.cpp Qwen3-TTS runtime revision: `6ad6d85f1bc0967ece71506074f1f7e1c372e167`
- Shared-threadpool optimization revision: `787e5fcf956164d87e5bd0e2dc9d9f5d776a20ab`
## License and Attribution
This repository is distributed under the Apache License 2.0. See `LICENSE`.
The base model is provided by the Qwen team; review the upstream model card for
its complete usage guidance, limitations, and citation information.
|