--- license: apache-2.0 base_model: DataoceanAI/dolphin-small library_name: onnx-asr pipeline_tag: automatic-speech-recognition language: - ar - az - ba - bn - fa - fil - gu - hi - id - ja - jv - kab - kk - km - ko - ks - ky - lo - mn - mr - ms - my - ne - or - pa - ps - ru - si - su - ta - te - tg - th - tl - ug - ur - uz - vi - yue - zh tags: - automatic-speech-recognition - onnx - espnet - e-branchformer - dolphin --- # dolphin-small-onnx ONNX export of [DataoceanAI/dolphin-small](https://huggingface.co/DataoceanAI/dolphin-small), a 372M parameter E-Branchformer encoder (12 blocks, width 768) with a 12 block transformer decoder, trained by **DataoceanAI** on 40 Eastern languages and 22 Chinese dialects. Apache-2.0, the same licence as the source model. All credit for the model goes to DataoceanAI; this repository only holds the converted graphs. ## Files | File | Size | | --- | --- | | `encoder.onnx` + `.data` | 700 MB | | `decoder.onnx` + `.data` | 716 MB | | `encoder.int8.onnx` + `.data` | 219 MB | | `decoder.int8.onnx` + `.data` | 283 MB | `encoder.onnx` takes the raw 16 kHz waveform: the ESPnet `default` frontend (STFT 512/400/160 plus 80 log-mel) and the `global_mvn` statistics are part of the graph, so `config.json` declares `"preprocessor": "identity"`. ## Usage Needs the `feat/espnet-aed-prompts` branch of the [TigreGotico/onnx-asr](https://github.com/TigreGotico/onnx-asr) fork, which adds the config-driven decode prompt that this model needs. ```sh pip install "onnx-asr[cpu,hub] @ git+https://github.com/TigreGotico/onnx-asr@feat/espnet-aed-prompts" ``` ```py import onnx_asr model = onnx_asr.load_model("OpenVoiceOS/dolphin-small-onnx") print(model.recognize("speech.wav", language="ja")) # Chinese dialects use the full tag. print(model.recognize("speech.wav", language="zh-SICHUAN")) # int8 model = onnx_asr.load_model("OpenVoiceOS/dolphin-small-onnx", quantization="int8") ``` The `language` argument accepts a full tag (`zh-CN`), the same tag with an underscore (`zh_CN`) or a bare language (`zh`), which maps to the first region of that language. Leave `language` out and the model predicts the language and the region itself. ## Languages `zh-CN`, `zh-TW`, `zh-WU`, `zh-SICHUAN`, `zh-SHANXI`, `zh-ANHUI`, `zh-TIANJIN`, `zh-NINGXIA`, `zh-SHAANXI`, `zh-HEBEI`, `zh-SHANDONG`, `zh-GUANGDONG`, `zh-SHANGHAI`, `zh-HUBEI`, `zh-LIAONING`, `zh-GANSU`, `zh-FUJIAN`, `zh-HUNAN`, `zh-HENAN`, `zh-YUNNAN`, `zh-MINNAN`, `zh-WENZHOU`, `ja-JP`, `th-TH`, `ru-RU`, `ko-KR`, `id-ID`, `vi-VN`, `ct-NULL`, `ct-HK`, `ct-GZ`, `hi-IN`, `ur-IN`, `ur-PK`, `ms-MY`, `uz-UZ`, `ar-MA`, `ar-GLA`, `ar-SA`, `ar-EG`, `ar-KW`, `ar-LY`, `ar-JO`, `ar-AE`, `ar-LVT`, `fa-IR`, `bn-BD`, `ta-SG`, `ta-LK`, `ta-IN`, `ta-MY`, `te-IN`, `ug-NULL`, `ug-CN`, `gu-IN`, `my-MM`, `tl-PH`, `kk-KZ`, `or-IN`, `ne-NP`, `mn-MN`, `km-KH`, `jv-ID`, `lo-LA`, `si-LK`, `fil-PH`, `ps-AF`, `pa-IN`, `kab-NULL`, `ba-NULL`, `ks-IN`, `tg-TJ`, `su-ID`, `mr-IN`, `ky-KG`, `az-AZ` ## Accuracy Greedy decoding, four FLEURS test clips, against the native Dolphin implementation with the same greedy decoding and the same forced language and region. | Clip | fp32 | int8 | | --- | --- | --- | | zh (cmn_hans_cn) | character-exact | character-exact | | ja (ja_jp) | character-exact | one word differs | | ko (ko_kr) | character-exact | character-exact | | th (th_th) | character-exact | one word differs, more word spacing | Real-time factor on 12 idle cores of an AMD box, one clip at a time: **0.24 fp32**, **0.13 int8**. ## Limitations * Greedy decoding only. The native implementation defaults to beam search. * No timestamps. The `` token is part of the baked prompt. * Hotword biasing of the source model is not exported.