Instructions to use OpenVoiceOS/dolphin-small-onnx with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- ESPnet
How to use OpenVoiceOS/dolphin-small-onnx with ESPnet:
from espnet2.bin.asr_inference import Speech2Text model = Speech2Text.from_pretrained( "OpenVoiceOS/dolphin-small-onnx" ) speech, rate = soundfile.read("speech.wav") text, *_ = model(speech)[0] - Notebooks
- Google Colab
- Kaggle
license: apache-2.0
base_model: DataoceanAI/dolphin-small
library_name: onnx-asr
pipeline_tag: automatic-speech-recognition
language:
- ar
- az
- ba
- bn
- fa
- fil
- gu
- hi
- id
- ja
- jv
- kab
- kk
- km
- ko
- ks
- ky
- lo
- mn
- mr
- ms
- my
- ne
- or
- pa
- ps
- ru
- si
- su
- ta
- te
- tg
- th
- tl
- ug
- ur
- uz
- vi
- yue
- zh
tags:
- automatic-speech-recognition
- onnx
- espnet
- e-branchformer
- dolphin
dolphin-small-onnx
ONNX export of DataoceanAI/dolphin-small, a 372M parameter E-Branchformer encoder (12 blocks, width 768) with a 12 block transformer decoder, trained by DataoceanAI on 40 Eastern languages and 22 Chinese dialects. Apache-2.0, the same licence as the source model. All credit for the model goes to DataoceanAI; this repository only holds the converted graphs.
Files
| File | Size |
|---|---|
encoder.onnx + .data |
700 MB |
decoder.onnx + .data |
716 MB |
encoder.int8.onnx + .data |
219 MB |
decoder.int8.onnx + .data |
283 MB |
encoder.onnx takes the raw 16 kHz waveform: the ESPnet default frontend (STFT
512/400/160 plus 80 log-mel) and the global_mvn statistics are part of the graph, so
config.json declares "preprocessor": "identity".
Usage
Needs the feat/espnet-aed-prompts branch of the
TigreGotico/onnx-asr fork, which adds the
config-driven decode prompt that this model needs.
pip install "onnx-asr[cpu,hub] @ git+https://github.com/TigreGotico/onnx-asr@feat/espnet-aed-prompts"
import onnx_asr
model = onnx_asr.load_model("OpenVoiceOS/dolphin-small-onnx")
print(model.recognize("speech.wav", language="ja"))
# Chinese dialects use the full tag.
print(model.recognize("speech.wav", language="zh-SICHUAN"))
# int8
model = onnx_asr.load_model("OpenVoiceOS/dolphin-small-onnx", quantization="int8")
The language argument accepts a full tag (zh-CN), the same tag with an underscore
(zh_CN) or a bare language (zh), which maps to the first region of that language.
Leave language out and the model predicts the language and the region itself.
Languages
zh-CN, zh-TW, zh-WU, zh-SICHUAN, zh-SHANXI, zh-ANHUI, zh-TIANJIN, zh-NINGXIA, zh-SHAANXI, zh-HEBEI, zh-SHANDONG, zh-GUANGDONG, zh-SHANGHAI, zh-HUBEI, zh-LIAONING, zh-GANSU, zh-FUJIAN, zh-HUNAN, zh-HENAN, zh-YUNNAN, zh-MINNAN, zh-WENZHOU, ja-JP, th-TH, ru-RU, ko-KR, id-ID, vi-VN, ct-NULL, ct-HK, ct-GZ, hi-IN, ur-IN, ur-PK, ms-MY, uz-UZ, ar-MA, ar-GLA, ar-SA, ar-EG, ar-KW, ar-LY, ar-JO, ar-AE, ar-LVT, fa-IR, bn-BD, ta-SG, ta-LK, ta-IN, ta-MY, te-IN, ug-NULL, ug-CN, gu-IN, my-MM, tl-PH, kk-KZ, or-IN, ne-NP, mn-MN, km-KH, jv-ID, lo-LA, si-LK, fil-PH, ps-AF, pa-IN, kab-NULL, ba-NULL, ks-IN, tg-TJ, su-ID, mr-IN, ky-KG, az-AZ
Accuracy
Greedy decoding, four FLEURS test clips, against the native Dolphin implementation with the same greedy decoding and the same forced language and region.
| Clip | fp32 | int8 |
|---|---|---|
| zh (cmn_hans_cn) | character-exact | character-exact |
| ja (ja_jp) | character-exact | one word differs |
| ko (ko_kr) | character-exact | character-exact |
| th (th_th) | character-exact | one word differs, more word spacing |
Real-time factor on 12 idle cores of an AMD box, one clip at a time: 0.24 fp32, 0.13 int8.
Limitations
- Greedy decoding only. The native implementation defaults to beam search.
- No timestamps. The
<notimestamp>token is part of the baked prompt. - Hotword biasing of the source model is not exported.