Discourse-STT / README.md
xfalcox's picture
Add Parakeet Ultra (4-bit) and Redux (2-bit) for in-browser captions
0862913 verified
|
Raw History Blame Contribute Delete
4.99 kB
metadata
license: cc-by-4.0
library_name: onnxruntime
pipeline_tag: automatic-speech-recognition
tags:
  - parakeet
  - tdt
  - onnx
  - webgpu
language:
  - bg
  - hr
  - cs
  - da
  - nl
  - en
  - et
  - fi
  - fr
  - de
  - el
  - hu
  - it
  - lv
  - lt
  - mt
  - pl
  - pt
  - ro
  - ru
  - sk
  - sl
  - es
  - sv
  - uk
base_model:
  - moondream/parakeet-ultra
  - moondream/parakeet-redux

Discourse STT models

Speech-to-text models for live captions and transcripts in Discourse's voice plugin. They run entirely in the browser, with onnxruntime-web on WebGPU. The worker that loads them ships in the discourse_voice_assets gem.

Each folder holds one model in the same layout. The voice plugin points the worker at a folder's URL.

Folder Model Encoder Download
ultra-q4/ Parakeet Ultra 4-bit MatMulNBits (block 32) ~393 MB
redux-w2a8/ Parakeet Redux 2-bit MatMulNBits (block 128), lossless for its ternary weights ~199 MB

Every folder contains:

  • encoder-model.onnx: a FastConformer encoder in a single file with no external data. The input is 128-bin log-mel features at 16 kHz.
  • decoder_joint-model.int8.onnx: the TDT prediction and joint network, dynamically quantized to int8.
  • vocab.txt: an 8193-token SentencePiece vocabulary with blank id 8192. It is identical for both models.

Both models are post-trained versions of NVIDIA's parakeet-tdt-0.6b-v3. They keep its architecture, tokenizer and 25 European languages.

  • Ultra is the default.
  • Redux is a much smaller download, with worse accuracy in several languages (including French and German) and in noisy audio.

See Moondream's model cards for benchmarks.

Modifications

The joint network's output bias for <unk> (token 0) is set to -1e4 in both decoders. Neither model can emit <unk>, which they otherwise produce for symbols missing from the vocabulary, such as ° and +. Both decoders come from Olicorne's exports, which already carry this patch. The other files are byte-identical to their sources.

Provenance

File Source repository Revision Original path
ultra-q4/encoder-model.onnx mrfakename/parakeet-ultra-ONNX 590d2668ea7c80d7e487da67a2946a20cb50780f encoder-model.q4.onnx
ultra-q4/decoder_joint-model.int8.onnx Olicorne/parakeet-tdt-0.6b-v3-ultra-onnx a1fe742758beb8ddc918bff00e1e319fe7e4d3e3 int8/decoder_joint-model.int8.onnx
ultra-q4/vocab.txt mrfakename/parakeet-ultra-ONNX 590d2668ea7c80d7e487da67a2946a20cb50780f vocab.txt
redux-w2a8/encoder-model.onnx Olicorne/parakeet-tdt-0.6b-v3-redux-onnx 7f2cf7ee13423501f912bf32e4624570d6cb6fcd w2a8/encoder-model.w2a8.onnx
redux-w2a8/decoder_joint-model.int8.onnx Olicorne/parakeet-tdt-0.6b-v3-redux-onnx 7f2cf7ee13423501f912bf32e4624570d6cb6fcd int8/decoder_joint-model.int8.onnx
redux-w2a8/vocab.txt Olicorne/parakeet-tdt-0.6b-v3-redux-onnx 7f2cf7ee13423501f912bf32e4624570d6cb6fcd vocab.txt

SHA-256:

40265cd6754f44f3deed117fa551eed69b974233b1c28c9b5a46750d0e27f648  ultra-q4/encoder-model.onnx
667ffde2242a560923cba6bc7f69ef03d736f8a36ca9588f2e6f9fa818cf2fc1  ultra-q4/decoder_joint-model.int8.onnx
d58544679ea4bc6ac563d1f545eb7d474bd6cfa467f0a6e2c1dc1c7d37e3c35d  ultra-q4/vocab.txt
24abd7ae5e82a3328de6c780ea21010460d14bc1c8a17e5c9c3bc031c381a712  redux-w2a8/encoder-model.onnx
c729ceaebd43ef6581890079505d04ab5b40c0b7a87ca1f78ad6054c1eae3144  redux-w2a8/decoder_joint-model.int8.onnx
d58544679ea4bc6ac563d1f545eb7d474bd6cfa467f0a6e2c1dc1c7d37e3c35d  redux-w2a8/vocab.txt

Self-hosting

Copy the folders to any HTTPS host that serves them with CORS enabled. Then set the voice plugin's voice_stt_model_base_url to the URL of the directory that contains them. Keep the layout and file names unchanged.

Browsers cache model files by URL. If you replace a file, publish it under a new URL instead of overwriting it in place.

License and attribution

Released under CC-BY-4.0, like the models they derive from: