File size: 2,980 Bytes
ff6d993 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 | ---
license: apache-2.0
base_model: OpenFormosa/BlueMagpie-TTS
tags:
- tts
- text-to-speech
- gguf
- llama.cpp
- llama.rn
- taiwanese-mandarin
language:
- zh
---
# BlueMagpie-TTS — GGUF
GGUF conversions of [OpenFormosa/BlueMagpie-TTS](https://huggingface.co/OpenFormosa/BlueMagpie-TTS),
a Taiwanese-Mandarin text-to-speech model, for use with
[llama.rn](https://github.com/mybigday/llama.rn) and
[codec.cpp](https://github.com/BricksDisplay/codec.cpp).
BlueMagpie is a **continuous-latent autoregressive-diffusion** TTS —
VoxCPM2 with its Text-Semantic LM swapped from MiniCPM-4 to
[Barbet](https://github.com/OpenFormosa/Barbet) (Mamba2 + attention hybrid,
1B params). The AudioVAE decodes the continuous latent sequence to a
48 kHz waveform.
## Files
### Text-Semantic LM (Barbet-1B backbone, runs in llama.cpp / llama.rn)
| File | Quant | Size |
|------|-------|------|
| `BlueMagpie-Barbet-1B-q4_k_m.gguf` | Q4_K_M | 661 MB |
| `BlueMagpie-Barbet-1B-q5_k_m.gguf` | Q5_K_M | 756 MB |
| `BlueMagpie-Barbet-1B-q6_k.gguf` | Q6_K | 857 MB |
| `BlueMagpie-Barbet-1B-q8_0.gguf` | Q8_0 | 1.08 GB |
| `BlueMagpie-Barbet-1B-f16.gguf` | F16 | 2.03 GB |
The **BPE (GPT2-family) tokenizer** is baked into every GGUF, so llama.cpp
can tokenize text natively — no external tokenizer runtime needed.
### Codec (AudioVAE + LM adaptor stack, runs in codec.cpp)
| File | Size |
|------|------|
| `BlueMagpie-AudioVAE.gguf` | 1.76 GB (F16) |
This bundle carries **all continuous-latent codec_lm components** the AR loop
needs: `tslm_adapter` + FSQ + RALM (MiniCPM4-8L) + LocEnc + LocDiT (12L CFM
diffusion) + `enc_to_lm_proj` + `enc_to_tslm_proj` + `lm_to_dit_proj` +
`res_to_dit_proj` + AudioVAE decoder + stop head. codec.cpp probes
`codec_common` `codec_lm_get_info().is_continuous == true` at load.
## Runtime
The llama.rn side loads Barbet as the backbone context and this codec.gguf
as the vocoder. `getFormattedAudioCompletion` returns
`flow = "continuous_embd"` + `embedding = true`; the standard `completion`
loop drives the codec_lm step machine per `llama_decode`, accumulating
latent patches into `result.embeddings`, which `decodeAudioEmbeddings`
turns into PCM at 48 kHz via the AudioVAE.
Requires **llama.rn ≥ codec branch** (adds `LLM_ARCH_BARBET`, the Mamba2/
attention hybrid graph builder, and the codec_common continuous-latent
completion-loop hook).
## License
Model weights follow the upstream Apache-2.0 license.
## Provenance
- Backbone converted via
[`scripts/vendor/convert_barbet_to_gguf.py`](https://github.com/mybigday/llama.rn/blob/codec/scripts/vendor/convert_barbet_to_gguf.py)
(llama.rn), which fuses the 5 Mamba2 in-projections + 3 conv1d into the
ssm_in / ssm_conv1d tensors llama.cpp expects and bakes the GPT2 BPE
tokenizer from the upstream `tokenizer.json`.
- Codec converted via
[`scripts/convert-to-gguf.py --model-type bluemagpie`](https://github.com/BricksDisplay/codec.cpp)
(codec.cpp).
|