File size: 7,766 Bytes
a8d652b 2f2f4ec 660f444 a8d652b 660f444 a8d652b 660f444 a8d652b 2f2f4ec a8d652b 660f444 35af8d7 2f2f4ec a8d652b 660f444 35af8d7 7d21d52 a8d652b d676522 a8d652b d676522 a8d652b d676522 a8d652b 660f444 a8d652b 660f444 a8d652b d676522 a8d652b 79892be a8d652b 660f444 a8d652b 660f444 a8d652b 660f444 28e721c 660f444 a8d652b 09a35a3 660f444 2f2f4ec a8d652b 660f444 a8d652b 660f444 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 | ---
language:
- nan
license: apache-2.0
library_name: mynahokkien
pipeline_tag: audio-to-audio
tags:
- hokkien
- singapore-hokkien
- speech-to-speech
- audio
---
# Myna-Hokkien
Myna-Hokkien is an open-source end-to-end conversational speech model for Singapore Hokkien.
Users speak to the model in Hokkien and the model replies in natural-sounding Hokkien speech -
no text round-trip required, though text input/output is also supported.
Hokkien is the native language of tens of millions of speakers across Asia,
yet it remains almost entirely absent from mainstream speech AI: no existing omni-style model (either open or closed) natively supports Hokkien.
Myna-Hokkien is an attempt to close that gap, and to leave behind a reusable recipe for other low-resource language communities to do the same.
Proudly built by [iNLP Lab](https://isakzhang.github.io/group.html) at SUTD.
## Model Details
* Languages: Hokkien / Minnan (闽南语), primarily Singapore-accent in this release
* Modalities: Audio in → Audio out: full spoken dialogue, text input/output also supported
* Architecture: Qwen3-Omni
* License: [Apache 2.0](https://www.apache.org/licenses/LICENSE-2.0)
## Demo Samples
Examples showcasing the model's Hokkien understanding & generation capabilities.
| Input | Myna-Hokkien | GPT Audio | Qwen3.5-Omni-Plus | Gemini Live | GLM-4-Voice |
|---|---|---|---|---|---|
| <audio controls src="https://huggingface.co/iNLP-Lab/Myna-Hokkien/resolve/main/demo_samples/release_csv/04/04_input.wav"></audio> | <audio controls src="https://huggingface.co/iNLP-Lab/Myna-Hokkien/resolve/main/demo_samples/release_csv/04/04_mynahokkien.wav"></audio> | <audio controls src="https://huggingface.co/iNLP-Lab/Myna-Hokkien/resolve/main/demo_samples/release_csv/04/04_gptaudio.wav"></audio> | <audio controls src="https://huggingface.co/iNLP-Lab/Myna-Hokkien/resolve/main/demo_samples/release_csv/04/04_qwen.wav"></audio> | <audio controls src="https://huggingface.co/iNLP-Lab/Myna-Hokkien/resolve/main/demo_samples/release_csv/04/04_gemini.wav"></audio> | <audio controls src="https://huggingface.co/iNLP-Lab/Myna-Hokkien/resolve/main/demo_samples/release_csv/04/04_glmvoice.wav"></audio> |
| <audio controls src="https://huggingface.co/iNLP-Lab/Myna-Hokkien/resolve/main/demo_samples/selected/05/05_input.wav"></audio> | <audio controls src="https://huggingface.co/iNLP-Lab/Myna-Hokkien/resolve/main/demo_samples/selected/05/05_mynahokkien.wav"></audio> | <audio controls src="https://huggingface.co/iNLP-Lab/Myna-Hokkien/resolve/main/demo_samples/selected/05/05_gptaudio.wav"></audio> | <audio controls src="https://huggingface.co/iNLP-Lab/Myna-Hokkien/resolve/main/demo_samples/selected/05/05_qwen.wav"></audio> | <audio controls src="https://huggingface.co/iNLP-Lab/Myna-Hokkien/resolve/main/demo_samples/selected/05/05_gemini.wav"></audio> | <audio controls src="https://huggingface.co/iNLP-Lab/Myna-Hokkien/resolve/main/demo_samples/selected/05/05_glmvoice.wav"></audio> |
| <audio controls src="https://huggingface.co/iNLP-Lab/Myna-Hokkien/resolve/main/demo_samples/selected/09/09_input.wav"></audio> | <audio controls src="https://huggingface.co/iNLP-Lab/Myna-Hokkien/resolve/main/demo_samples/selected/09/09_mynahokkien.wav"></audio> | <audio controls src="https://huggingface.co/iNLP-Lab/Myna-Hokkien/resolve/main/demo_samples/selected/09/09_gptaudio.wav"></audio> | <audio controls src="https://huggingface.co/iNLP-Lab/Myna-Hokkien/resolve/main/demo_samples/selected/09/09_qwen.wav"></audio> | <audio controls src="https://huggingface.co/iNLP-Lab/Myna-Hokkien/resolve/main/demo_samples/selected/09/09_gemini.wav"></audio> | <audio controls src="https://huggingface.co/iNLP-Lab/Myna-Hokkien/resolve/main/demo_samples/selected/09/09_glmvoice.wav"></audio> |
| How is the weather today? | <audio controls src="https://huggingface.co/iNLP-Lab/Myna-Hokkien/resolve/main/demo_samples/selected/12/12_mynahokkien.wav"></audio> | <audio controls src="https://huggingface.co/iNLP-Lab/Myna-Hokkien/resolve/main/demo_samples/selected/12/12_gptaudio.wav"></audio> | <audio controls src="https://huggingface.co/iNLP-Lab/Myna-Hokkien/resolve/main/demo_samples/selected/12/12_qwen.wav"></audio> | <audio controls src="https://huggingface.co/iNLP-Lab/Myna-Hokkien/resolve/main/demo_samples/selected/12/12_gemini.wav"></audio> | <audio controls src="https://huggingface.co/iNLP-Lab/Myna-Hokkien/resolve/main/demo_samples/selected/12/12_glmvoice.wav"></audio> |
| 我今天心情有点不好。 | <audio controls src="https://huggingface.co/iNLP-Lab/Myna-Hokkien/resolve/main/demo_samples/selected/13/13_mynahokkien.wav"></audio> | <audio controls src="https://huggingface.co/iNLP-Lab/Myna-Hokkien/resolve/main/demo_samples/selected/13/13_gptaudio.wav"></audio> | <audio controls src="https://huggingface.co/iNLP-Lab/Myna-Hokkien/resolve/main/demo_samples/selected/13/13_qwen.wav"></audio> | <audio controls src="https://huggingface.co/iNLP-Lab/Myna-Hokkien/resolve/main/demo_samples/selected/13/13_gemini.wav"></audio> | <audio controls src="https://huggingface.co/iNLP-Lab/Myna-Hokkien/resolve/main/demo_samples/selected/13/13_glmvoice.wav"></audio> |
## Installation
Download Myna-Hokkien into the standard Hugging Face cache, then install the included inference
runtime:
```bash
pip install --upgrade huggingface_hub
MODEL_DIR="$(hf download iNLP-Lab/Myna-Hokkien --quiet)"
pip install "$MODEL_DIR"
```
## Basic usage — spoken question
```python
import torch
import soundfile as sf
from mynahokkien import MynaHokkien
model = MynaHokkien.from_pretrained(
"iNLP-Lab/Myna-Hokkien",
device_map="cuda:0",
dtype=torch.float16,
)
output = model.generate(
audio="question.wav",
language="nan",
return_text=True,
return_audio=True,
)
print(output.text)
sf.write("output.wav", output.audio, output.sampling_rate)
```
## Text-query usage
Text is treated as a question or instruction to the Hokkien assistant; it is not treated as a TTS
transcript.
```python
output = model.generate(
text="講一個新加坡福建話的笑話。",
language="nan",
)
sf.write("output.wav", output.audio, output.sampling_rate)
print(output.text)
```
Exactly one of `audio=` and `text=` must be supplied. This release currently supports
`language="nan"` and the `Ethan` voice.
To return only one modality:
```python
text_only = model.generate(text="你會曉講福建話無?", return_text=True, return_audio=False)
audio_only = model.generate(text="講一句歡迎詞。", return_text=False, return_audio=True)
```
## Prompt behavior
If `prompt=` is omitted for audio input, Myna-Hokkien uses this built-in prompt:
```text
Listen to the spoken Hokkien and reply naturally in concise Singapore Hokkien. Always answer in colloquial Singapore Hokkien written in Hanji. Never answer in Mandarin or English. Do not repeat or transcribe the input; respond to it directly.
```
To override it, pass a different instruction through `prompt=`:
```python
output = model.generate(
audio="question.wav",
prompt="Listen to this audio and reply naturally in Singaporean Hokkien.",
language="nan",
)
```
For text input, put the instruction directly in `text=`.
## Citation
```
@misc{myna-hokkien-2026,
title = {Myna-Hokkien: An Open-Source End-to-End Hokkien Spoken Dialogue Model},
author = {Matthew Christopher Pohadi and Ryner Tan and Wenxuan Zhang},
year = {2026},
howpublished = {\url{https://huggingface.co/iNLP-Lab/Myna-Hokkien}}
```
## Contact & collaboration
This is an active, ongoing project — we're continuing to improve accent coverage, prosody, and expressiveness.
We'd love to hear from you if you want to collaborate, have feedback, or run into issues: [Wenxuan Zhang](https://isakzhang.github.io/).
|