mobilebytesensei's picture
fix: the tail-pad sweep was a THIRD-PARTY zipformer measurement, not this bundle
d8ca24c verified
|
Raw
History Blame Contribute Delete
3.9 kB
---
license: cc-by-4.0
library_name: sherpa-onnx
tags:
- automatic-speech-recognition
- streaming
- cache-aware
- onnx
- nemo
- transducer
language: [en]
---
# Betterflow — English streaming FastConformer, ONNX for sherpa-onnx
An ONNX export of **`nvidia/stt_en_fastconformer_hybrid_large_streaming_multi`**, prepared so it
loads in **`sherpa_onnx.OnlineRecognizer`** and produces **live partials** for English.
**We are not the authors of the weights.** Upstream is NVIDIA; this repo is a format conversion
plus quantization.
## Provenance and licence
| | |
|---|---|
| Upstream | [`nvidia/stt_en_fastconformer_hybrid_large_streaming_multi`](https://huggingface.co/nvidia/stt_en_fastconformer_hybrid_large_streaming_multi) |
| Upstream licence | **CC-BY-4.0** (read off the model card, not inferred) |
| This repo | **CC-BY-4.0**, inherited — **attribution required** |
| What changed | `.nemo` → ONNX via k2-fsa's own export path · `set_default_att_context_size([70,13])` · int8 quantization |
| What did NOT change | the weights — no fine-tuning |
Please cite NVIDIA for the underlying model.
## Contents — a transducer bundle (three graphs)
```
encoder.int8.onnx 131,507,640 B encoder.onnx 456,772,215 B
decoder.int8.onnx 3,955,863 B decoder.onnx 15,753,087 B
joiner.int8.onnx 1,408,183 B joiner.onnx 5,584,035 B
tokens.txt 11,896 B
```
int8 total ≈ **137 MB**.
## Measured
`librispeech-en`, n=50, through sherpa with a padded tail:
| | |
|---|---|
| WER | **7.7% pooled · 5.3% median** |
| RTF | **0.021** |
| peak RSS | **662 MB** |
| empty | **0/50** |
| script | **100% Latin** |
Sample decode (int8, 2 s tail pad):
```
'concord returned to its place amidst the tents'
'congratulations were poured in upon the princess everywhere during her journey'
```
## ⚠️ Three things worth knowing
**1. Pad the tail — and this bundle tells you exactly how much.** sherpa's online recogniser only
decodes when `num_frames_ready - num_processed >= window_size`, and `input_finished()` does **not**
pad to a whole window, so up to `window_size - 1` frames of every utterance are never decoded.
This encoder declares **`window_size = 121`, `chunk_shift = 112`, `subsampling_factor = 8`** in its
ONNX metadata. At a 10 ms hop that is **1.21 s**, so **pad ≥ ~1.3 s**; we use 2,000 ms. Shorter pads
lose words as *deletions*, which read as poor model quality rather than as a configuration error.
> ‼️ **Read `window_size` off the graph rather than copying a number.** An earlier version of this
> card quoted a `0 → 20.7% · 500 → 10.4% · 2,000 → 5.4%` sweep as if it were measured on this bundle.
> **It was not** — those are third-party figures from a *streaming zipformer* on Android, a different
> architecture whose chunk length we never read. The advice was right; the numbers were not ours.
**2. `downloadMb` is not `peakRssMb`.** 137 MB on disk, **662 MB resident** — a 4.9× gap. Budget on
the resident figure.
**3. Peak RSS is FLAT in utterance length****671.2 MB at 5 s, 671.5 MB at 240 s**. **No
utterance-length cap is needed** for this bundle.
The reason is the **cache-aware architecture** — bounded left context plus a fixed cache — not the
fact that it streams. ‼️ **"Streaming ⇒ bounded memory" is false as a general rule**: we measured a
streaming decoder-only model whose peak RSS scales **T^1.49** and walls at ~10.6 s of audio. Flat
memory is a property of *this family* (cache-aware Conformer), and an offline Conformer's attention
is O(T²). **Check the scaling; do not infer it from the word "streaming."**
## Not evaluated
Device/Android verification · lookaheads other than `[70,13]` (`0`/`80`/`480` ms are exportable via
the same script) · languages other than English · dictation-register audio — the numbers above are
read speech.