How to use from the
Use from the
ESPnet library
from espnet2.bin.asr_inference import Speech2Text

model = Speech2Text.from_pretrained(
  "OpenVoiceOS/dolphin-small-onnx"
)

speech, rate = soundfile.read("speech.wav")
text, *_ = model(speech)[0]

dolphin-small-onnx

ONNX export of DataoceanAI/dolphin-small, a 372M parameter E-Branchformer encoder (12 blocks, width 768) with a 12 block transformer decoder, trained by DataoceanAI on 40 Eastern languages and 22 Chinese dialects. Apache-2.0, the same licence as the source model. All credit for the model goes to DataoceanAI; this repository only holds the converted graphs.

Files

File Size
encoder.onnx + .data 700 MB
decoder.onnx + .data 716 MB
encoder.int8.onnx + .data 219 MB
decoder.int8.onnx + .data 283 MB

encoder.onnx takes the raw 16 kHz waveform: the ESPnet default frontend (STFT 512/400/160 plus 80 log-mel) and the global_mvn statistics are part of the graph, so config.json declares "preprocessor": "identity".

Usage

Needs the feat/espnet-aed-prompts branch of the TigreGotico/onnx-asr fork, which adds the config-driven decode prompt that this model needs.

pip install "onnx-asr[cpu,hub] @ git+https://github.com/TigreGotico/onnx-asr@feat/espnet-aed-prompts"
import onnx_asr

model = onnx_asr.load_model("OpenVoiceOS/dolphin-small-onnx")
print(model.recognize("speech.wav", language="ja"))

# Chinese dialects use the full tag.
print(model.recognize("speech.wav", language="zh-SICHUAN"))

# int8
model = onnx_asr.load_model("OpenVoiceOS/dolphin-small-onnx", quantization="int8")

The language argument accepts a full tag (zh-CN), the same tag with an underscore (zh_CN) or a bare language (zh), which maps to the first region of that language. Leave language out and the model predicts the language and the region itself.

Languages

zh-CN, zh-TW, zh-WU, zh-SICHUAN, zh-SHANXI, zh-ANHUI, zh-TIANJIN, zh-NINGXIA, zh-SHAANXI, zh-HEBEI, zh-SHANDONG, zh-GUANGDONG, zh-SHANGHAI, zh-HUBEI, zh-LIAONING, zh-GANSU, zh-FUJIAN, zh-HUNAN, zh-HENAN, zh-YUNNAN, zh-MINNAN, zh-WENZHOU, ja-JP, th-TH, ru-RU, ko-KR, id-ID, vi-VN, ct-NULL, ct-HK, ct-GZ, hi-IN, ur-IN, ur-PK, ms-MY, uz-UZ, ar-MA, ar-GLA, ar-SA, ar-EG, ar-KW, ar-LY, ar-JO, ar-AE, ar-LVT, fa-IR, bn-BD, ta-SG, ta-LK, ta-IN, ta-MY, te-IN, ug-NULL, ug-CN, gu-IN, my-MM, tl-PH, kk-KZ, or-IN, ne-NP, mn-MN, km-KH, jv-ID, lo-LA, si-LK, fil-PH, ps-AF, pa-IN, kab-NULL, ba-NULL, ks-IN, tg-TJ, su-ID, mr-IN, ky-KG, az-AZ

Accuracy

Greedy decoding, four FLEURS test clips, against the native Dolphin implementation with the same greedy decoding and the same forced language and region.

Clip fp32 int8
zh (cmn_hans_cn) character-exact character-exact
ja (ja_jp) character-exact one word differs
ko (ko_kr) character-exact character-exact
th (th_th) character-exact one word differs, more word spacing

Real-time factor on 12 idle cores of an AMD box, one clip at a time: 0.24 fp32, 0.13 int8.

Limitations

  • Greedy decoding only. The native implementation defaults to beam search.
  • No timestamps. The <notimestamp> token is part of the baked prompt.
  • Hotword biasing of the source model is not exported.
Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for OpenVoiceOS/dolphin-small-onnx

Quantized
(1)
this model

Collections including OpenVoiceOS/dolphin-small-onnx