File size: 5,720 Bytes
8bdbc40
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
7f50eb3
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
8bdbc40
 
 
7f50eb3
 
8bdbc40
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
---
license: mit
library_name: sherpa-onnx
tags:
  - automatic-speech-recognition
  - indic
  - onnx
  - sherpa-onnx
  - nemo
language:
  - gu
  - pa
  - bn
  - mr
  - ml
  - te
  - ta
  - kn
---

# Betterflow β€” IndicConformer CTC, ONNX int8 (8 Indic languages)

ONNX int8 exports of **AI4Bharat's per-language IndicConformer** hybrid CTC/RNNT checkpoints,
prepared for on-device offline ASR through `sherpa-onnx`'s `OfflineRecognizer.from_nemo_ctc`.

**These are derived artefacts. We are not the authors of the weights.**

## Provenance and licence

| | |
|---|---|
| Upstream | `ai4bharat/indicconformer_stt_<lang>_hybrid_ctc_rnnt_large` |
| Upstream licence | **MIT** (verified on each source model card, not inferred) |
| This repo's licence | **MIT**, inherited |
| What changed | CTC head exported to ONNX; dynamic int8 quantization; sherpa metadata stamped |
| What did NOT change | the weights themselves β€” no fine-tuning, no distillation |

Each language directory carries `export-meta.json` with the **sha256 of the source `.nemo`**, so
provenance is checkable rather than asserted.

Please cite AI4Bharat for the underlying models.

## Contents

Eight languages: `gu` `pa` `bn` `mr` `ml` `te` `ta` `kn`.

```
<lang>/model.int8.onnx     140,337,395 B   (identical size across languages, DISTINCT weights)
<lang>/tokens.txt               67,605 B   5,633 entries
<lang>/export-meta.json                    source sha256, checksums, export settings
```

### int8 sha256 β€” all eight distinct

| lang | sha256 (first 16) | lang | sha256 (first 16) |
|---|---|---|---|
| `gu` | `9f8031996be1a4d7` | `ml` | `81c172f6ee3796f5` |
| `pa` | `1ad447f477040bbb` | `te` | `b0b3d335f2505ecd` |
| `bn` | `8b1bd509bb36ba86` | `ta` | `08b85596996429e1` |
| `mr` | `af22de1897679e4a` | `kn` | `16d950f36cfdb01e` |

The eight files are byte-identical **in size** because they share an architecture. They are
**different models** β€” the checksums above are the check that matters.

## ⚠️ Two things that will silently break this

**1. `normalize_type` must be `per_feature`.** It is stamped on every graph. If it is stripped or
set to `""`, the model **loads, runs, and returns EMPTY with no error**. Verify by decoding, never
by a successful load.

Required ONNX metadata: `vocab_size=5633`, `subsampling_factor=4`,
`normalize_type=per_feature`, `feature_dim=80`.

**2. `tokens.txt` opens with Bengali tokens in every language directory β€” this is correct.** The
vocabulary is AI4Bharat's shared 22-language tokenizer (5,633 = 22 Γ— 256 + 1) and the CTC head is
multi-softmax. It looks like a packaging error and is not. The file is byte-identical across all
eight (`ee60967630213f31…`); it is duplicated per directory so each bundle installs independently.

Unlike the 600m multilingual checkpoint, these per-language models **did** learn to suppress the
other 21 blocks β€” measured 100% own-script output on all eight, 0/400 wrong-alphabet characters β€”
so **no `language_id` masking is required**.

## Measured quality

Scored on held-out `ARTPARK-IISc/Vaani`, 50 clips Γ— 50 speakers per language, against the
general-purpose multilingual Whisper bundle these replace.

| lang | incumbent pooled WER | **this export** | script purity | catastrophic |
|---|---|---|---|---|
| `pa` | 106.2% | **20.2%** | 0% β†’ **100%** | 13 β†’ **0** |
| `gu` | 106.6% | **22.6%** | 0% β†’ **100%** | 16 β†’ **0** |
| `bn` | 135.1% | **30.6%** | 0% β†’ **100%** | 21 β†’ **0** |
| `mr` | 100.5% | **36.6%** | 83% β†’ **100%** | 14 β†’ **0** |
| `ml` | 111.9% | **37.2%** | 0% β†’ **100%** | 16 β†’ **1** |
| `te` | 115.5% | **45.0%** | 6% β†’ **100%** | 8 β†’ **1** |
| `ta` | 87.0% | **52.1%** | 100% β†’ **100%** | 3 β†’ **1** |
| `kn` | 100.3% | **59.5%** | 100% β†’ **100%** | 2 β†’ **1** |

**Script purity is the headline, not WER.** A 0% score means the incumbent was not producing the
language's alphabet at all. Catastrophic failures fell from **93/400 to 5/400**.

`te` and `kn` remain insertion-bound and are **not** recommended without further evaluation.

## Runtime cost

### ⚠️ Peak RSS scales with utterance length β€” it is not a constant

Conformer self-attention is **O(TΒ²)** in sequence length, so memory is a function of how long the
speaker talks. Measured in-process (Android arm64, sherpa-onnx 1.13.4, int8, 2 threads, one process
per utterance, Β±1 MB reproducible):

| utterance | peak RSS | | utterance | peak RSS |
|---|---|---|---|---|
| 5 s | **285 MB** | | 45 s | 868 MB |
| 10 s | **350 MB** | | 60 s | 986 MB |
| 20 s | **482 MB** | | 90 s | 1,628 MB |
| 30 s | 606 MB | | 134 s | **2,758 MB** |

**Budget accordingly**: a typical 3–10 s dictation utterance costs under 350 MB, but a two-minute
monologue costs 2.7 GB. If you have a memory ceiling, cap utterance length β€” that is the cheapest
control and it belongs in your app, not in the model.

| | |
|---|---|
| RTF | 0.062–0.064 (~16Γ— real time on an M4 laptop) |
| RTF, 2Γ— ARM Cortex-A55 @1.25 GHz | **1.33** (~13 s for a 10 s utterance) |
| RTF, 4Γ— Cortex-A78 | **0.255** |

The int8 build quantizes **all 54 `Conv` nodes** to `ConvInteger`. Leaving them in fp32 β€” as some
other re-exports do β€” makes ONNX Runtime allocate an im2col workspace that scales with clip length,
pushing peak RSS to **2,099–2,635 MB** from the same nominal precision. If you are getting 4Γ— the
memory you expect from an int8 Conformer, check your `Conv` node dtypes.

Cost of that choice: `ConvInteger` is ~3.6Γ— slower than fp32 `Conv` on CPU.

## Not evaluated here

Device-side accuracy, languages outside the eight, streaming/partial decoding (these are **offline**
models), and the RNNT head (not exported β€” it needs numba/CUDA).