Instructions to use mobilebytesensei/betterflow-en-streaming-fastconformer with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- NeMo
How to use mobilebytesensei/betterflow-en-streaming-fastconformer with NeMo:
import nemo.collections.asr as nemo_asr asr_model = nemo_asr.models.ASRModel.from_pretrained("mobilebytesensei/betterflow-en-streaming-fastconformer") transcriptions = asr_model.transcribe(["file.wav"]) - Notebooks
- Google Colab
- Kaggle
| license: cc-by-4.0 | |
| library_name: sherpa-onnx | |
| tags: | |
| - automatic-speech-recognition | |
| - streaming | |
| - cache-aware | |
| - onnx | |
| - nemo | |
| - transducer | |
| language: [en] | |
| # Betterflow — English streaming FastConformer, ONNX for sherpa-onnx | |
| An ONNX export of **`nvidia/stt_en_fastconformer_hybrid_large_streaming_multi`**, prepared so it | |
| loads in **`sherpa_onnx.OnlineRecognizer`** and produces **live partials** for English. | |
| **We are not the authors of the weights.** Upstream is NVIDIA; this repo is a format conversion | |
| plus quantization. | |
| ## Provenance and licence | |
| | | | | |
| |---|---| | |
| | Upstream | [`nvidia/stt_en_fastconformer_hybrid_large_streaming_multi`](https://huggingface.co/nvidia/stt_en_fastconformer_hybrid_large_streaming_multi) | | |
| | Upstream licence | **CC-BY-4.0** (read off the model card, not inferred) | | |
| | This repo | **CC-BY-4.0**, inherited — **attribution required** | | |
| | What changed | `.nemo` → ONNX via k2-fsa's own export path · `set_default_att_context_size([70,13])` · int8 quantization | | |
| | What did NOT change | the weights — no fine-tuning | | |
| Please cite NVIDIA for the underlying model. | |
| ## Contents — a transducer bundle (three graphs) | |
| ``` | |
| encoder.int8.onnx 131,507,640 B encoder.onnx 456,772,215 B | |
| decoder.int8.onnx 3,955,863 B decoder.onnx 15,753,087 B | |
| joiner.int8.onnx 1,408,183 B joiner.onnx 5,584,035 B | |
| tokens.txt 11,896 B | |
| ``` | |
| int8 total ≈ **137 MB**. | |
| ## Measured | |
| `librispeech-en`, n=50, through sherpa with a padded tail: | |
| | | | | |
| |---|---| | |
| | WER | **7.7% pooled · 5.3% median** | | |
| | RTF | **0.021** | | |
| | peak RSS | **662 MB** | | |
| | empty | **0/50** | | |
| | script | **100% Latin** | | |
| Sample decode (int8, 2 s tail pad): | |
| ``` | |
| 'concord returned to its place amidst the tents' | |
| 'congratulations were poured in upon the princess everywhere during her journey' | |
| ``` | |
| ## ⚠️ Three things worth knowing | |
| **1. Pad the tail — and this bundle tells you exactly how much.** sherpa's online recogniser only | |
| decodes when `num_frames_ready - num_processed >= window_size`, and `input_finished()` does **not** | |
| pad to a whole window, so up to `window_size - 1` frames of every utterance are never decoded. | |
| This encoder declares **`window_size = 121`, `chunk_shift = 112`, `subsampling_factor = 8`** in its | |
| ONNX metadata. At a 10 ms hop that is **1.21 s**, so **pad ≥ ~1.3 s**; we use 2,000 ms. Shorter pads | |
| lose words as *deletions*, which read as poor model quality rather than as a configuration error. | |
| > ‼️ **Read `window_size` off the graph rather than copying a number.** An earlier version of this | |
| > card quoted a `0 → 20.7% · 500 → 10.4% · 2,000 → 5.4%` sweep as if it were measured on this bundle. | |
| > **It was not** — those are third-party figures from a *streaming zipformer* on Android, a different | |
| > architecture whose chunk length we never read. The advice was right; the numbers were not ours. | |
| **2. `downloadMb` is not `peakRssMb`.** 137 MB on disk, **662 MB resident** — a 4.9× gap. Budget on | |
| the resident figure. | |
| **3. Peak RSS is FLAT in utterance length** — **671.2 MB at 5 s, 671.5 MB at 240 s**. **No | |
| utterance-length cap is needed** for this bundle. | |
| The reason is the **cache-aware architecture** — bounded left context plus a fixed cache — not the | |
| fact that it streams. ‼️ **"Streaming ⇒ bounded memory" is false as a general rule**: we measured a | |
| streaming decoder-only model whose peak RSS scales **T^1.49** and walls at ~10.6 s of audio. Flat | |
| memory is a property of *this family* (cache-aware Conformer), and an offline Conformer's attention | |
| is O(T²). **Check the scaling; do not infer it from the word "streaming."** | |
| ## Not evaluated | |
| Device/Android verification · lookaheads other than `[70,13]` (`0`/`80`/`480` ms are exportable via | |
| the same script) · languages other than English · dictation-register audio — the numbers above are | |
| read speech. | |