Instructions to use mobilebytesensei/betterflow-indicconformer-ctc with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- NeMo
How to use mobilebytesensei/betterflow-indicconformer-ctc with NeMo:
import nemo.collections.asr as nemo_asr asr_model = nemo_asr.models.ASRModel.from_pretrained("mobilebytesensei/betterflow-indicconformer-ctc") transcriptions = asr_model.transcribe(["file.wav"]) - Notebooks
- Google Colab
- Kaggle
| license: mit | |
| library_name: sherpa-onnx | |
| tags: | |
| - automatic-speech-recognition | |
| - indic | |
| - onnx | |
| - sherpa-onnx | |
| - nemo | |
| language: | |
| - gu | |
| - pa | |
| - bn | |
| - mr | |
| - ml | |
| - te | |
| - ta | |
| - kn | |
| # Betterflow β IndicConformer CTC, ONNX int8 (8 Indic languages) | |
| ONNX int8 exports of **AI4Bharat's per-language IndicConformer** hybrid CTC/RNNT checkpoints, | |
| prepared for on-device offline ASR through `sherpa-onnx`'s `OfflineRecognizer.from_nemo_ctc`. | |
| **These are derived artefacts. We are not the authors of the weights.** | |
| ## Provenance and licence | |
| | | | | |
| |---|---| | |
| | Upstream | `ai4bharat/indicconformer_stt_<lang>_hybrid_ctc_rnnt_large` | | |
| | Upstream licence | **MIT** (verified on each source model card, not inferred) | | |
| | This repo's licence | **MIT**, inherited | | |
| | What changed | CTC head exported to ONNX; dynamic int8 quantization; sherpa metadata stamped | | |
| | What did NOT change | the weights themselves β no fine-tuning, no distillation | | |
| Each language directory carries `export-meta.json` with the **sha256 of the source `.nemo`**, so | |
| provenance is checkable rather than asserted. | |
| Please cite AI4Bharat for the underlying models. | |
| ## Contents | |
| Eight languages: `gu` `pa` `bn` `mr` `ml` `te` `ta` `kn`. | |
| ``` | |
| <lang>/model.int8.onnx 140,337,395 B (identical size across languages, DISTINCT weights) | |
| <lang>/tokens.txt 67,605 B 5,633 entries | |
| <lang>/export-meta.json source sha256, checksums, export settings | |
| ``` | |
| ### int8 sha256 β all eight distinct | |
| | lang | sha256 (first 16) | lang | sha256 (first 16) | | |
| |---|---|---|---| | |
| | `gu` | `9f8031996be1a4d7` | `ml` | `81c172f6ee3796f5` | | |
| | `pa` | `1ad447f477040bbb` | `te` | `b0b3d335f2505ecd` | | |
| | `bn` | `8b1bd509bb36ba86` | `ta` | `08b85596996429e1` | | |
| | `mr` | `af22de1897679e4a` | `kn` | `16d950f36cfdb01e` | | |
| The eight files are byte-identical **in size** because they share an architecture. They are | |
| **different models** β the checksums above are the check that matters. | |
| ## β οΈ Two things that will silently break this | |
| **1. `normalize_type` must be `per_feature`.** It is stamped on every graph. If it is stripped or | |
| set to `""`, the model **loads, runs, and returns EMPTY with no error**. Verify by decoding, never | |
| by a successful load. | |
| Required ONNX metadata: `vocab_size=5633`, `subsampling_factor=4`, | |
| `normalize_type=per_feature`, `feature_dim=80`. | |
| **2. `tokens.txt` opens with Bengali tokens in every language directory β this is correct.** The | |
| vocabulary is AI4Bharat's shared 22-language tokenizer (5,633 = 22 Γ 256 + 1) and the CTC head is | |
| multi-softmax. It looks like a packaging error and is not. The file is byte-identical across all | |
| eight (`ee60967630213f31β¦`); it is duplicated per directory so each bundle installs independently. | |
| Unlike the 600m multilingual checkpoint, these per-language models **did** learn to suppress the | |
| other 21 blocks β measured 100% own-script output on all eight, 0/400 wrong-alphabet characters β | |
| so **no `language_id` masking is required**. | |
| ## Measured quality | |
| Scored on held-out `ARTPARK-IISc/Vaani`, 50 clips Γ 50 speakers per language, against the | |
| general-purpose multilingual Whisper bundle these replace. | |
| | lang | incumbent pooled WER | **this export** | script purity | catastrophic | | |
| |---|---|---|---|---| | |
| | `pa` | 106.2% | **20.2%** | 0% β **100%** | 13 β **0** | | |
| | `gu` | 106.6% | **22.6%** | 0% β **100%** | 16 β **0** | | |
| | `bn` | 135.1% | **30.6%** | 0% β **100%** | 21 β **0** | | |
| | `mr` | 100.5% | **36.6%** | 83% β **100%** | 14 β **0** | | |
| | `ml` | 111.9% | **37.2%** | 0% β **100%** | 16 β **1** | | |
| | `te` | 115.5% | **45.0%** | 6% β **100%** | 8 β **1** | | |
| | `ta` | 87.0% | **52.1%** | 100% β **100%** | 3 β **1** | | |
| | `kn` | 100.3% | **59.5%** | 100% β **100%** | 2 β **1** | | |
| **Script purity is the headline, not WER.** A 0% score means the incumbent was not producing the | |
| language's alphabet at all. Catastrophic failures fell from **93/400 to 5/400**. | |
| `te` and `kn` remain insertion-bound and are **not** recommended without further evaluation. | |
| ## Runtime cost | |
| ### β οΈ Peak RSS scales with utterance length β it is not a constant | |
| Conformer self-attention is **O(TΒ²)** in sequence length, so memory is a function of how long the | |
| speaker talks. Measured in-process (Android arm64, sherpa-onnx 1.13.4, int8, 2 threads, one process | |
| per utterance, Β±1 MB reproducible): | |
| | utterance | peak RSS | | utterance | peak RSS | | |
| |---|---|---|---|---| | |
| | 5 s | **285 MB** | | 45 s | 868 MB | | |
| | 10 s | **350 MB** | | 60 s | 986 MB | | |
| | 20 s | **482 MB** | | 90 s | 1,628 MB | | |
| | 30 s | 606 MB | | 134 s | **2,758 MB** | | |
| **Budget accordingly**: a typical 3β10 s dictation utterance costs under 350 MB, but a two-minute | |
| monologue costs 2.7 GB. If you have a memory ceiling, cap utterance length β that is the cheapest | |
| control and it belongs in your app, not in the model. | |
| | | | | |
| |---|---| | |
| | RTF | 0.062β0.064 (~16Γ real time on an M4 laptop) | | |
| | RTF, 2Γ ARM Cortex-A55 @1.25 GHz | **1.33** (~13 s for a 10 s utterance) | | |
| | RTF, 4Γ Cortex-A78 | **0.255** | | |
| The int8 build quantizes **all 54 `Conv` nodes** to `ConvInteger`. Leaving them in fp32 β as some | |
| other re-exports do β makes ONNX Runtime allocate an im2col workspace that scales with clip length, | |
| pushing peak RSS to **2,099β2,635 MB** from the same nominal precision. If you are getting 4Γ the | |
| memory you expect from an int8 Conformer, check your `Conv` node dtypes. | |
| Cost of that choice: `ConvInteger` is ~3.6Γ slower than fp32 `Conv` on CPU. | |
| ## Not evaluated here | |
| Device-side accuracy, languages outside the eight, streaming/partial decoding (these are **offline** | |
| models), and the RNNT head (not exported β it needs numba/CUDA). | |