Discourse-STT / README.md
xfalcox's picture
Add Parakeet Ultra (4-bit) and Redux (2-bit) for in-browser captions
0862913 verified
|
Raw History Blame Contribute Delete
4.99 kB
---
license: cc-by-4.0
library_name: onnxruntime
pipeline_tag: automatic-speech-recognition
tags:
- parakeet
- tdt
- onnx
- webgpu
language:
- bg
- hr
- cs
- da
- nl
- en
- et
- fi
- fr
- de
- el
- hu
- it
- lv
- lt
- mt
- pl
- pt
- ro
- ru
- sk
- sl
- es
- sv
- uk
base_model:
- moondream/parakeet-ultra
- moondream/parakeet-redux
---
# Discourse STT models
Speech-to-text models for live captions and transcripts in [Discourse](https://www.discourse.org)'s voice plugin. They run entirely in the browser, with [onnxruntime-web](https://onnxruntime.ai) on WebGPU. The worker that loads them ships in the [`discourse_voice_assets`](https://github.com/discourse/discourse_voice_assets) gem.
Each folder holds one model in the same layout. The voice plugin points the worker at a folder's URL.
| Folder | Model | Encoder | Download |
|---|---|---|---|
| `ultra-q4/` | Parakeet Ultra | 4-bit MatMulNBits (block 32) | ~393 MB |
| `redux-w2a8/` | Parakeet Redux | 2-bit MatMulNBits (block 128), lossless for its ternary weights | ~199 MB |
Every folder contains:
- `encoder-model.onnx`: a FastConformer encoder in a single file with no external data. The input is 128-bin log-mel features at 16 kHz.
- `decoder_joint-model.int8.onnx`: the TDT prediction and joint network, dynamically quantized to int8.
- `vocab.txt`: an 8193-token SentencePiece vocabulary with blank id 8192. It is identical for both models.
Both models are post-trained versions of NVIDIA's [parakeet-tdt-0.6b-v3](https://huggingface.co/nvidia/parakeet-tdt-0.6b-v3). They keep its architecture, tokenizer and 25 European languages.
- **Ultra** is the default.
- **Redux** is a much smaller download, with worse accuracy in several languages (including French and German) and in noisy audio.
See Moondream's model cards for benchmarks.
## Modifications
The joint network's output bias for `<unk>` (token 0) is set to -1e4 in both decoders. Neither model can emit `<unk>`, which they otherwise produce for symbols missing from the vocabulary, such as `°` and `+`. Both decoders come from Olicorne's exports, which already carry this patch. The other files are byte-identical to their sources.
## Provenance
| File | Source repository | Revision | Original path |
|---|---|---|---|
| `ultra-q4/encoder-model.onnx` | [mrfakename/parakeet-ultra-ONNX](https://huggingface.co/mrfakename/parakeet-ultra-ONNX) | `590d2668ea7c80d7e487da67a2946a20cb50780f` | `encoder-model.q4.onnx` |
| `ultra-q4/decoder_joint-model.int8.onnx` | [Olicorne/parakeet-tdt-0.6b-v3-ultra-onnx](https://huggingface.co/Olicorne/parakeet-tdt-0.6b-v3-ultra-onnx) | `a1fe742758beb8ddc918bff00e1e319fe7e4d3e3` | `int8/decoder_joint-model.int8.onnx` |
| `ultra-q4/vocab.txt` | mrfakename/parakeet-ultra-ONNX | `590d2668ea7c80d7e487da67a2946a20cb50780f` | `vocab.txt` |
| `redux-w2a8/encoder-model.onnx` | [Olicorne/parakeet-tdt-0.6b-v3-redux-onnx](https://huggingface.co/Olicorne/parakeet-tdt-0.6b-v3-redux-onnx) | `7f2cf7ee13423501f912bf32e4624570d6cb6fcd` | `w2a8/encoder-model.w2a8.onnx` |
| `redux-w2a8/decoder_joint-model.int8.onnx` | Olicorne/parakeet-tdt-0.6b-v3-redux-onnx | `7f2cf7ee13423501f912bf32e4624570d6cb6fcd` | `int8/decoder_joint-model.int8.onnx` |
| `redux-w2a8/vocab.txt` | Olicorne/parakeet-tdt-0.6b-v3-redux-onnx | `7f2cf7ee13423501f912bf32e4624570d6cb6fcd` | `vocab.txt` |
SHA-256:
```
40265cd6754f44f3deed117fa551eed69b974233b1c28c9b5a46750d0e27f648 ultra-q4/encoder-model.onnx
667ffde2242a560923cba6bc7f69ef03d736f8a36ca9588f2e6f9fa818cf2fc1 ultra-q4/decoder_joint-model.int8.onnx
d58544679ea4bc6ac563d1f545eb7d474bd6cfa467f0a6e2c1dc1c7d37e3c35d ultra-q4/vocab.txt
24abd7ae5e82a3328de6c780ea21010460d14bc1c8a17e5c9c3bc031c381a712 redux-w2a8/encoder-model.onnx
c729ceaebd43ef6581890079505d04ab5b40c0b7a87ca1f78ad6054c1eae3144 redux-w2a8/decoder_joint-model.int8.onnx
d58544679ea4bc6ac563d1f545eb7d474bd6cfa467f0a6e2c1dc1c7d37e3c35d redux-w2a8/vocab.txt
```
## Self-hosting
Copy the folders to any HTTPS host that serves them with CORS enabled. Then set the voice plugin's `voice_stt_model_base_url` to the URL of the directory that contains them. Keep the layout and file names unchanged.
Browsers cache model files by URL. If you replace a file, publish it under a new URL instead of overwriting it in place.
## License and attribution
Released under [CC-BY-4.0](https://creativecommons.org/licenses/by/4.0/), like the models they derive from:
- [nvidia/parakeet-tdt-0.6b-v3](https://huggingface.co/nvidia/parakeet-tdt-0.6b-v3), © NVIDIA
- [moondream/parakeet-ultra](https://huggingface.co/moondream/parakeet-ultra) and [moondream/parakeet-redux](https://huggingface.co/moondream/parakeet-redux), © M87 Labs (Moondream)
- ONNX exports by [mrfakename](https://huggingface.co/mrfakename) and [Olicorne](https://huggingface.co/Olicorne), and through them [eschmidbauer/parakeet-redux-onnx](https://huggingface.co/eschmidbauer/parakeet-redux-onnx)