File size: 7,766 Bytes
a8d652b
 
 
2f2f4ec
660f444
a8d652b
 
 
 
 
 
 
 
 
 
660f444
 
 
a8d652b
660f444
 
 
a8d652b
2f2f4ec
a8d652b
 
660f444
 
 
35af8d7
2f2f4ec
a8d652b
 
660f444
 
 
35af8d7
 
 
 
 
7d21d52
 
a8d652b
 
 
 
d676522
a8d652b
 
 
 
 
d676522
a8d652b
 
 
 
 
 
 
 
 
 
 
d676522
a8d652b
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
660f444
a8d652b
 
 
 
 
 
 
 
 
 
660f444
 
a8d652b
 
 
 
 
 
 
 
 
 
d676522
a8d652b
 
79892be
a8d652b
 
660f444
a8d652b
 
 
 
660f444
a8d652b
 
 
 
660f444
28e721c
 
 
660f444
a8d652b
 
09a35a3
660f444
2f2f4ec
a8d652b
 
660f444
a8d652b
660f444
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
---
language:
- nan
license: apache-2.0
library_name: mynahokkien
pipeline_tag: audio-to-audio
tags:
- hokkien
- singapore-hokkien
- speech-to-speech
- audio
---

# Myna-Hokkien

Myna-Hokkien is an open-source end-to-end conversational speech model for Singapore Hokkien. 
Users speak to the model in Hokkien and the model replies in natural-sounding Hokkien speech - 
no text round-trip required, though text input/output is also supported.

Hokkien is the native language of tens of millions of speakers across Asia, 
yet it remains almost entirely absent from mainstream speech AI: no existing omni-style model (either open or closed) natively supports Hokkien. 
Myna-Hokkien is an attempt to close that gap, and to leave behind a reusable recipe for other low-resource language communities to do the same.

Proudly built by [iNLP Lab](https://isakzhang.github.io/group.html) at SUTD.


## Model Details
* Languages: Hokkien / Minnan (闽南语), primarily Singapore-accent in this release
* Modalities: Audio in → Audio out: full spoken dialogue, text input/output also supported
* Architecture: Qwen3-Omni
* License: [Apache 2.0](https://www.apache.org/licenses/LICENSE-2.0)


## Demo Samples
Examples showcasing the model's Hokkien understanding & generation capabilities.

| Input | Myna-Hokkien | GPT Audio | Qwen3.5-Omni-Plus | Gemini Live | GLM-4-Voice |
|---|---|---|---|---|---|
| <audio controls src="https://huggingface.co/iNLP-Lab/Myna-Hokkien/resolve/main/demo_samples/release_csv/04/04_input.wav"></audio> | <audio controls src="https://huggingface.co/iNLP-Lab/Myna-Hokkien/resolve/main/demo_samples/release_csv/04/04_mynahokkien.wav"></audio> | <audio controls src="https://huggingface.co/iNLP-Lab/Myna-Hokkien/resolve/main/demo_samples/release_csv/04/04_gptaudio.wav"></audio> | <audio controls src="https://huggingface.co/iNLP-Lab/Myna-Hokkien/resolve/main/demo_samples/release_csv/04/04_qwen.wav"></audio> | <audio controls src="https://huggingface.co/iNLP-Lab/Myna-Hokkien/resolve/main/demo_samples/release_csv/04/04_gemini.wav"></audio> | <audio controls src="https://huggingface.co/iNLP-Lab/Myna-Hokkien/resolve/main/demo_samples/release_csv/04/04_glmvoice.wav"></audio> |
| <audio controls src="https://huggingface.co/iNLP-Lab/Myna-Hokkien/resolve/main/demo_samples/selected/05/05_input.wav"></audio> | <audio controls src="https://huggingface.co/iNLP-Lab/Myna-Hokkien/resolve/main/demo_samples/selected/05/05_mynahokkien.wav"></audio> | <audio controls src="https://huggingface.co/iNLP-Lab/Myna-Hokkien/resolve/main/demo_samples/selected/05/05_gptaudio.wav"></audio> | <audio controls src="https://huggingface.co/iNLP-Lab/Myna-Hokkien/resolve/main/demo_samples/selected/05/05_qwen.wav"></audio> | <audio controls src="https://huggingface.co/iNLP-Lab/Myna-Hokkien/resolve/main/demo_samples/selected/05/05_gemini.wav"></audio> | <audio controls src="https://huggingface.co/iNLP-Lab/Myna-Hokkien/resolve/main/demo_samples/selected/05/05_glmvoice.wav"></audio> |
| <audio controls src="https://huggingface.co/iNLP-Lab/Myna-Hokkien/resolve/main/demo_samples/selected/09/09_input.wav"></audio> | <audio controls src="https://huggingface.co/iNLP-Lab/Myna-Hokkien/resolve/main/demo_samples/selected/09/09_mynahokkien.wav"></audio> | <audio controls src="https://huggingface.co/iNLP-Lab/Myna-Hokkien/resolve/main/demo_samples/selected/09/09_gptaudio.wav"></audio> | <audio controls src="https://huggingface.co/iNLP-Lab/Myna-Hokkien/resolve/main/demo_samples/selected/09/09_qwen.wav"></audio> | <audio controls src="https://huggingface.co/iNLP-Lab/Myna-Hokkien/resolve/main/demo_samples/selected/09/09_gemini.wav"></audio> | <audio controls src="https://huggingface.co/iNLP-Lab/Myna-Hokkien/resolve/main/demo_samples/selected/09/09_glmvoice.wav"></audio> |
| How is the weather today? | <audio controls src="https://huggingface.co/iNLP-Lab/Myna-Hokkien/resolve/main/demo_samples/selected/12/12_mynahokkien.wav"></audio> | <audio controls src="https://huggingface.co/iNLP-Lab/Myna-Hokkien/resolve/main/demo_samples/selected/12/12_gptaudio.wav"></audio> | <audio controls src="https://huggingface.co/iNLP-Lab/Myna-Hokkien/resolve/main/demo_samples/selected/12/12_qwen.wav"></audio> | <audio controls src="https://huggingface.co/iNLP-Lab/Myna-Hokkien/resolve/main/demo_samples/selected/12/12_gemini.wav"></audio> | <audio controls src="https://huggingface.co/iNLP-Lab/Myna-Hokkien/resolve/main/demo_samples/selected/12/12_glmvoice.wav"></audio> |
| 我今天心情有点不好。 | <audio controls src="https://huggingface.co/iNLP-Lab/Myna-Hokkien/resolve/main/demo_samples/selected/13/13_mynahokkien.wav"></audio> | <audio controls src="https://huggingface.co/iNLP-Lab/Myna-Hokkien/resolve/main/demo_samples/selected/13/13_gptaudio.wav"></audio> | <audio controls src="https://huggingface.co/iNLP-Lab/Myna-Hokkien/resolve/main/demo_samples/selected/13/13_qwen.wav"></audio> | <audio controls src="https://huggingface.co/iNLP-Lab/Myna-Hokkien/resolve/main/demo_samples/selected/13/13_gemini.wav"></audio> | <audio controls src="https://huggingface.co/iNLP-Lab/Myna-Hokkien/resolve/main/demo_samples/selected/13/13_glmvoice.wav"></audio> |


## Installation

Download Myna-Hokkien into the standard Hugging Face cache, then install the included inference
runtime:

```bash
pip install --upgrade huggingface_hub

MODEL_DIR="$(hf download iNLP-Lab/Myna-Hokkien --quiet)"
pip install "$MODEL_DIR"
```

## Basic usage — spoken question

```python
import torch
import soundfile as sf
from mynahokkien import MynaHokkien

model = MynaHokkien.from_pretrained(
    "iNLP-Lab/Myna-Hokkien",
    device_map="cuda:0",
    dtype=torch.float16,
)

output = model.generate(
    audio="question.wav",
    language="nan",
    return_text=True,
    return_audio=True,
)

print(output.text)
sf.write("output.wav", output.audio, output.sampling_rate)
```

## Text-query usage

Text is treated as a question or instruction to the Hokkien assistant; it is not treated as a TTS
transcript.

```python
output = model.generate(
    text="講一個新加坡福建話的笑話。",
    language="nan",
)
sf.write("output.wav", output.audio, output.sampling_rate)
print(output.text)
```

Exactly one of `audio=` and `text=` must be supplied. This release currently supports
`language="nan"` and the `Ethan` voice.

To return only one modality:

```python
text_only = model.generate(text="你會曉講福建話無?", return_text=True, return_audio=False)
audio_only = model.generate(text="講一句歡迎詞。", return_text=False, return_audio=True)
```

## Prompt behavior

If `prompt=` is omitted for audio input, Myna-Hokkien uses this built-in prompt:

```text
Listen to the spoken Hokkien and reply naturally in concise Singapore Hokkien. Always answer in colloquial Singapore Hokkien written in Hanji. Never answer in Mandarin or English. Do not repeat or transcribe the input; respond to it directly.
```

To override it, pass a different instruction through `prompt=`:

```python
output = model.generate(
    audio="question.wav",
    prompt="Listen to this audio and reply naturally in Singaporean Hokkien.",
    language="nan",
)
```

For text input, put the instruction directly in `text=`.


## Citation
```
@misc{myna-hokkien-2026,
  title  = {Myna-Hokkien: An Open-Source End-to-End Hokkien Spoken Dialogue Model},
  author = {Matthew Christopher Pohadi and Ryner Tan and Wenxuan Zhang},
  year   = {2026},
  howpublished = {\url{https://huggingface.co/iNLP-Lab/Myna-Hokkien}}
```

## Contact & collaboration

This is an active, ongoing project — we're continuing to improve accent coverage, prosody, and expressiveness. 
We'd love to hear from you if you want to collaborate, have feedback, or run into issues: [Wenxuan Zhang](https://isakzhang.github.io/).