dolphin-small-onnx / README.md
Jarbas's picture
Language and region are predicted when no language is given
413fcf2 verified
|
Raw
History Blame Contribute Delete
3.74 kB
metadata
license: apache-2.0
base_model: DataoceanAI/dolphin-small
library_name: onnx-asr
pipeline_tag: automatic-speech-recognition
language:
  - ar
  - az
  - ba
  - bn
  - fa
  - fil
  - gu
  - hi
  - id
  - ja
  - jv
  - kab
  - kk
  - km
  - ko
  - ks
  - ky
  - lo
  - mn
  - mr
  - ms
  - my
  - ne
  - or
  - pa
  - ps
  - ru
  - si
  - su
  - ta
  - te
  - tg
  - th
  - tl
  - ug
  - ur
  - uz
  - vi
  - yue
  - zh
tags:
  - automatic-speech-recognition
  - onnx
  - espnet
  - e-branchformer
  - dolphin

dolphin-small-onnx

ONNX export of DataoceanAI/dolphin-small, a 372M parameter E-Branchformer encoder (12 blocks, width 768) with a 12 block transformer decoder, trained by DataoceanAI on 40 Eastern languages and 22 Chinese dialects. Apache-2.0, the same licence as the source model. All credit for the model goes to DataoceanAI; this repository only holds the converted graphs.

Files

File Size
encoder.onnx + .data 700 MB
decoder.onnx + .data 716 MB
encoder.int8.onnx + .data 219 MB
decoder.int8.onnx + .data 283 MB

encoder.onnx takes the raw 16 kHz waveform: the ESPnet default frontend (STFT 512/400/160 plus 80 log-mel) and the global_mvn statistics are part of the graph, so config.json declares "preprocessor": "identity".

Usage

Needs the feat/espnet-aed-prompts branch of the TigreGotico/onnx-asr fork, which adds the config-driven decode prompt that this model needs.

pip install "onnx-asr[cpu,hub] @ git+https://github.com/TigreGotico/onnx-asr@feat/espnet-aed-prompts"
import onnx_asr

model = onnx_asr.load_model("OpenVoiceOS/dolphin-small-onnx")
print(model.recognize("speech.wav", language="ja"))

# Chinese dialects use the full tag.
print(model.recognize("speech.wav", language="zh-SICHUAN"))

# int8
model = onnx_asr.load_model("OpenVoiceOS/dolphin-small-onnx", quantization="int8")

The language argument accepts a full tag (zh-CN), the same tag with an underscore (zh_CN) or a bare language (zh), which maps to the first region of that language. Leave language out and the model predicts the language and the region itself.

Languages

zh-CN, zh-TW, zh-WU, zh-SICHUAN, zh-SHANXI, zh-ANHUI, zh-TIANJIN, zh-NINGXIA, zh-SHAANXI, zh-HEBEI, zh-SHANDONG, zh-GUANGDONG, zh-SHANGHAI, zh-HUBEI, zh-LIAONING, zh-GANSU, zh-FUJIAN, zh-HUNAN, zh-HENAN, zh-YUNNAN, zh-MINNAN, zh-WENZHOU, ja-JP, th-TH, ru-RU, ko-KR, id-ID, vi-VN, ct-NULL, ct-HK, ct-GZ, hi-IN, ur-IN, ur-PK, ms-MY, uz-UZ, ar-MA, ar-GLA, ar-SA, ar-EG, ar-KW, ar-LY, ar-JO, ar-AE, ar-LVT, fa-IR, bn-BD, ta-SG, ta-LK, ta-IN, ta-MY, te-IN, ug-NULL, ug-CN, gu-IN, my-MM, tl-PH, kk-KZ, or-IN, ne-NP, mn-MN, km-KH, jv-ID, lo-LA, si-LK, fil-PH, ps-AF, pa-IN, kab-NULL, ba-NULL, ks-IN, tg-TJ, su-ID, mr-IN, ky-KG, az-AZ

Accuracy

Greedy decoding, four FLEURS test clips, against the native Dolphin implementation with the same greedy decoding and the same forced language and region.

Clip fp32 int8
zh (cmn_hans_cn) character-exact character-exact
ja (ja_jp) character-exact one word differs
ko (ko_kr) character-exact character-exact
th (th_th) character-exact one word differs, more word spacing

Real-time factor on 12 idle cores of an AMD box, one clip at a time: 0.24 fp32, 0.13 int8.

Limitations

  • Greedy decoding only. The native implementation defaults to beam search.
  • No timestamps. The <notimestamp> token is part of the baked prompt.
  • Hotword biasing of the source model is not exported.