Ear β 16MB on-device speech-to-text
Ear is a tiny English speech-to-text model designed for on-device inference. It prioritizes extreme smallness over raw accuracy: int8 weights, no tokenizer or vocab file, host-side mel frontend, greedy CTC decoder, published failure modes, and honest benchmarks against Whisper.
| Architecture | TinyConformer-CTC (9 blocks, d=256, 4 heads) |
| Params | 15.58M fp32 β 15.08 MB int8 |
| Alphabet | 29 symbols β blank, space, aβz, apostrophe. No BPE. No vocab file. |
| Training data | LibriSpeech train-clean-460 (100h + 360h, CC-BY-4.0) |
| Test WER (fp32) | 8.6% test-clean Β· 12.5% test-clean β€3s |
| Test WER (int8) | 8.6% β 0.0pp regression (lossless quantization) |
| Latency | 3.2ms/utt GPU Β· CPU RTF < 0.02 (target) |
| Runtime | onnxruntime + numpy β no PyTorch at inference |
Quick start
import numpy as np
import soundfile as sf
import onnxruntime as ort
import json, struct
# Load model artifacts
meta = json.load(open("artifacts/stt_meta.json"))
labels = json.load(open("artifacts/labels.json"))
ALPHABET = [e["symbol"] for e in labels]
sess = ort.InferenceSession("artifacts/stt.onnx",
providers=["CPUExecutionProvider"])
# Load torchaudio's exact filterbank (extract once and cache as .npy)
# See notebook Section 06 for the one-time extraction step.
MEL_FB = np.load("artifacts/mel_filterbank.npy") # (201, 80)
WIN = (0.5 * (1.0 - np.cos(2.0 * np.pi * np.arange(400) / 400))).astype(np.float32)
def log_mel(wav: np.ndarray) -> np.ndarray:
"""wav: float32 mono @16kHz β (80, T) log-mel."""
wav = np.pad(wav, 200, mode="reflect")
n_frames = 1 + max(0, (len(wav) - 400) // 160)
mel = np.empty((80, n_frames), dtype=np.float32)
for i in range(n_frames):
seg = wav[i*160:i*160+400] * WIN
spec = np.abs(np.fft.rfft(seg, n=400)) ** 2
mel[:, i] = MEL_FB.T @ spec
return 10.0 * np.log10(np.maximum(mel, 1e-10))
def greedy_ctc(log_probs, blank=0, thresh=0.35):
probs = np.exp(log_probs)
ids, conf = probs.argmax(-1), probs.max(-1)
out, prev, cs = [], -1, []
for t, i in enumerate(ids):
if i != blank:
cs.append(float(conf[t]))
if i != prev: out.append(ALPHABET[i])
prev = int(i)
text = "".join(out)
c = sum(cs)/len(cs) if cs else 0.0
return (text, c) if c >= thresh else ("[uncertain]", c)
def transcribe(path: str):
wav, sr = sf.read(path, dtype="float32")
assert sr == 16000, "resample to 16kHz first"
if wav.ndim > 1: wav = wav.mean(1)
mel = log_mel(wav).astype(np.float32)
m = mel.T[None, None] # (1,1,T,80)
logits, out_len = sess.run(None, {
"mel": m, "mel_len": np.array([mel.shape[1]])
})
T = int(out_len[0])
return greedy_ctc(logits[0, :T])
text, confidence = transcribe("audio.wav")
print(f"[{confidence:.2f}] {text}")
Artifacts
| File | Size | Description |
|---|---|---|
artifacts/stt_int8.bin |
15.08 MB | Shipped artifact. Raw int8 weights + fp32 scales. JSON header + blob. |
artifacts/stt_int4.bin |
9.66 MB | Experimental. Groupwise int4 (group=64). 1.1pp WER cost (9.7%). |
artifacts/stt.onnx |
β | ONNX graph (requires stt.onnx.data in same directory). |
artifacts/stt.onnx.data |
~60 MB | External weights for ONNX graph. |
artifacts/stt_meta.json |
0.6 KB | Mel constants, alphabet, blank id, decode params. |
artifacts/labels.json |
1.2 KB | 29 CTC symbols with ids. |
artifacts/mel_filterbank.npy |
~64 KB | Torchaudio's exact filterbank matrix β use in the numpy frontend. |
ckpts/tongue-stt-latest-checkpoint.pt |
~60 MB | Full fp32 PyTorch checkpoint for fine-tuning. |
int8 bin format
Any host can load stt_int8.bin in ~30 lines:
import struct, json, numpy as np
def load_int8_bin(path):
raw = open(path, "rb").read()
(hl,) = struct.unpack("<I", raw[:4])
header = json.loads(raw[4:4+hl])
blob = raw[4+hl:]
weights = {}
for name, t in header["tensors"].items():
b = blob[t["offset"]:]
if t["dtype"] == "int8":
n = int(np.prod(t["shape"]))
q = np.frombuffer(b[:n], dtype=np.int8).reshape(t["shape"])
weights[name] = q.astype(np.float32) * t["scale"]
else:
n = int(np.prod(t["shape"]))
weights[name] = np.frombuffer(b[:n*4], dtype=np.float32).reshape(t["shape"]).copy()
return weights
Architecture
16kHz waveform (host)
β 80-dim log-mel, 25ms window / 10ms hop (host, fixed constants)
β Conv2d 4Γ subsampling β 40ms frames
β 9 Γ Pre-LN ConformerBlock (d=256, 4 heads, conv kernel 15, FFN 4Γ)
β Linear β 29 CTC classes
β (host) greedy CTC decode β text + confidence
Key design choices:
- Manual QKV Linear layers (not
nn.MultiheadAttention) β every large matmul is int8-quantizable. - BatchNorm folded into depthwise convs before export β BN is never in the shipped artifact.
- Greedy CTC decoder fits in 20 lines of numpy β no inference runtime needed.
- Character-level alphabet, hardcoded β zero BPE files, zero vocab files.
Model definition: model_lib.py in this repo.
Training
| Data | LibriSpeech train-clean-460 (train-100 + train-360, ~460h, CC-BY-4.0) |
| Augmentation | SpecAugment (time + freq masks) |
| Optimizer | AdamW, lr=3e-4, weight_decay=1e-4 |
| Schedule | Linear warmup 4000 steps β cosine anneal |
| Epochs | 60 (step 426,174) |
| Hardware | Colab A100 (~46h wall clock across two sessions) |
| Precision | fp16 mixed (AMP) |
Training notebook: tongue_stt_v3-FINAL.ipynb. Checkpoints saved every 2000 steps to Google Drive; training survived a 24h session restart via resume.
Results
Gate scorecard
| Gate | Target | Result | Status |
|---|---|---|---|
| Artifact size (int8) | β€ 18 MB | 15.08 MB | β |
| test-clean WER (fp32 greedy) | β€ 15% | 8.6% | β |
| int8 WER regression | β€ 1pp | 0.0pp | β |
| int4 WER regression | β€ 1pp | 1.1pp | β 0.1pp over |
| Short-utterance β€3s WER gap | within 3pp | 3.9pp | β 0.9pp over |
| Golden-vector parity (20 utts) | 20/20 | 18/20 (cos=1.000 all 20) | β 2 int8 edge cases |
| Numpy mel parity | < 0.5 dB MAE | 0.0 dB | β |
| Zero vocab files shipped | β | β | β |
WER on LibriSpeech test-clean
| Model | Size | WER | CER | Latency |
|---|---|---|---|---|
| Ear fp32 | 59.4 MB | 8.6% | 2.6% | 3.2 ms/utt |
| Ear int8 | 15.08 MB | 8.6% | 2.7% | 2.9 ms/utt |
| Ear int4 | 9.66 MB | 9.7% | 3.0% | 2.9 ms/utt |
| Ear v1 (micro, 100h) | 2.11 MB | 28.7% | 9.2% | 1.7 ms/utt |
| Whisper tiny.en | 142 MB | 7.4%* | β | β |
| Whisper base.en | 244 MB | 6.8%* | β | β |
*Whisper measured on 200-utterance β€3s subset; Ear WER on same subset is 13.4%.
Ear is 9.4Γ smaller than Whisper tiny.en and 1.6Γ worse WER on the full test-clean. The win is size Γ latency, not raw accuracy.
Failure modes
These are published as part of the product β honesty about what the model can't do is as important as the benchmarks.
1. Proper nouns and names β structural No language model means no lexical anchoring. Irish, Greek, and archaic English names from LibriSpeech's literary corpus (Joyce, Plato, Shakespeare) fail badly.
- "stephanos dedalos" β "stuffronors daelos"
- "timaeus approves" β "to me as a proves"
2. Short utterances (β€3s) β harder by design 12.5% WER vs 8.6% full-set (3.9pp gap). Single phrases give less temporal context for error recovery.
3. Apostrophe handling Contractions occasionally produce spurious apostrophe placement.
- "he's swiftly punished" β "he 'is swiftly punished"
4. English clean speech only Trained on LibriSpeech clean read speech. Accents, noise, and far-field audio degrade sharply. Out of scope by design.
5. No language model in the shipped artifact Greedy CTC on characters β homophones and phonetically similar words are indistinguishable. Reported honestly, not hidden.
6. Confidence threshold is uncalibrated
uncertain=0.0% across all 2,620 test utterances. The confidence signal exists (mean non-blank CTC posterior) but has not been calibrated against accuracy. Do not rely on it for gating.
License
Code: Apache-2.0. Weights trained on LibriSpeech (CC-BY-4.0) β trained weights inherit CC-BY-4.0.
- Downloads last month
- -
Dataset used to train CDT5058/ear-stt
Space using CDT5058/ear-stt 1
Collection including CDT5058/ear-stt
Evaluation results
- Test WER (fp32 greedy) on LibriSpeech test-cleanself-reported8.600
- Test CER (fp32 greedy) on LibriSpeech test-cleanself-reported2.600
- Test WER (int8 greedy) on LibriSpeech test-cleanself-reported8.600
- Test CER (int8 greedy) on LibriSpeech test-cleanself-reported2.700
- Test WER (int4 groupwise greedy) on LibriSpeech test-cleanself-reported9.700
- Test CER (int4 groupwise greedy) on LibriSpeech test-cleanself-reported3.000