Ear β€” 16MB on-device speech-to-text

Ear is a tiny English speech-to-text model designed for on-device inference. It prioritizes extreme smallness over raw accuracy: int8 weights, no tokenizer or vocab file, host-side mel frontend, greedy CTC decoder, published failure modes, and honest benchmarks against Whisper.

Architecture TinyConformer-CTC (9 blocks, d=256, 4 heads)
Params 15.58M fp32 β†’ 15.08 MB int8
Alphabet 29 symbols β€” blank, space, a–z, apostrophe. No BPE. No vocab file.
Training data LibriSpeech train-clean-460 (100h + 360h, CC-BY-4.0)
Test WER (fp32) 8.6% test-clean Β· 12.5% test-clean ≀3s
Test WER (int8) 8.6% β€” 0.0pp regression (lossless quantization)
Latency 3.2ms/utt GPU Β· CPU RTF < 0.02 (target)
Runtime onnxruntime + numpy β€” no PyTorch at inference

Quick start

import numpy as np
import soundfile as sf
import onnxruntime as ort
import json, struct

# Load model artifacts
meta   = json.load(open("artifacts/stt_meta.json"))
labels = json.load(open("artifacts/labels.json"))
ALPHABET = [e["symbol"] for e in labels]

sess = ort.InferenceSession("artifacts/stt.onnx",
                             providers=["CPUExecutionProvider"])

# Load torchaudio's exact filterbank (extract once and cache as .npy)
# See notebook Section 06 for the one-time extraction step.
MEL_FB = np.load("artifacts/mel_filterbank.npy")          # (201, 80)
WIN    = (0.5 * (1.0 - np.cos(2.0 * np.pi * np.arange(400) / 400))).astype(np.float32)

def log_mel(wav: np.ndarray) -> np.ndarray:
    """wav: float32 mono @16kHz β†’ (80, T) log-mel."""
    wav = np.pad(wav, 200, mode="reflect")
    n_frames = 1 + max(0, (len(wav) - 400) // 160)
    mel = np.empty((80, n_frames), dtype=np.float32)
    for i in range(n_frames):
        seg = wav[i*160:i*160+400] * WIN
        spec = np.abs(np.fft.rfft(seg, n=400)) ** 2
        mel[:, i] = MEL_FB.T @ spec
    return 10.0 * np.log10(np.maximum(mel, 1e-10))

def greedy_ctc(log_probs, blank=0, thresh=0.35):
    probs = np.exp(log_probs)
    ids, conf = probs.argmax(-1), probs.max(-1)
    out, prev, cs = [], -1, []
    for t, i in enumerate(ids):
        if i != blank:
            cs.append(float(conf[t]))
            if i != prev: out.append(ALPHABET[i])
        prev = int(i)
    text = "".join(out)
    c = sum(cs)/len(cs) if cs else 0.0
    return (text, c) if c >= thresh else ("[uncertain]", c)

def transcribe(path: str):
    wav, sr = sf.read(path, dtype="float32")
    assert sr == 16000, "resample to 16kHz first"
    if wav.ndim > 1: wav = wav.mean(1)
    mel = log_mel(wav).astype(np.float32)
    m = mel.T[None, None]                          # (1,1,T,80)
    logits, out_len = sess.run(None, {
        "mel": m, "mel_len": np.array([mel.shape[1]])
    })
    T = int(out_len[0])
    return greedy_ctc(logits[0, :T])

text, confidence = transcribe("audio.wav")
print(f"[{confidence:.2f}] {text}")

Artifacts

File Size Description
artifacts/stt_int8.bin 15.08 MB Shipped artifact. Raw int8 weights + fp32 scales. JSON header + blob.
artifacts/stt_int4.bin 9.66 MB Experimental. Groupwise int4 (group=64). 1.1pp WER cost (9.7%).
artifacts/stt.onnx β€” ONNX graph (requires stt.onnx.data in same directory).
artifacts/stt.onnx.data ~60 MB External weights for ONNX graph.
artifacts/stt_meta.json 0.6 KB Mel constants, alphabet, blank id, decode params.
artifacts/labels.json 1.2 KB 29 CTC symbols with ids.
artifacts/mel_filterbank.npy ~64 KB Torchaudio's exact filterbank matrix β€” use in the numpy frontend.
ckpts/tongue-stt-latest-checkpoint.pt ~60 MB Full fp32 PyTorch checkpoint for fine-tuning.

int8 bin format

Any host can load stt_int8.bin in ~30 lines:

import struct, json, numpy as np

def load_int8_bin(path):
    raw = open(path, "rb").read()
    (hl,) = struct.unpack("<I", raw[:4])
    header = json.loads(raw[4:4+hl])
    blob   = raw[4+hl:]
    weights = {}
    for name, t in header["tensors"].items():
        b = blob[t["offset"]:]
        if t["dtype"] == "int8":
            n = int(np.prod(t["shape"]))
            q = np.frombuffer(b[:n], dtype=np.int8).reshape(t["shape"])
            weights[name] = q.astype(np.float32) * t["scale"]
        else:
            n = int(np.prod(t["shape"]))
            weights[name] = np.frombuffer(b[:n*4], dtype=np.float32).reshape(t["shape"]).copy()
    return weights

Architecture

16kHz waveform (host)
  β†’ 80-dim log-mel, 25ms window / 10ms hop (host, fixed constants)
  β†’ Conv2d 4Γ— subsampling β†’ 40ms frames
  β†’ 9 Γ— Pre-LN ConformerBlock (d=256, 4 heads, conv kernel 15, FFN 4Γ—)
  β†’ Linear β†’ 29 CTC classes
  β†’ (host) greedy CTC decode β†’ text + confidence

Key design choices:

  • Manual QKV Linear layers (not nn.MultiheadAttention) β€” every large matmul is int8-quantizable.
  • BatchNorm folded into depthwise convs before export β€” BN is never in the shipped artifact.
  • Greedy CTC decoder fits in 20 lines of numpy β€” no inference runtime needed.
  • Character-level alphabet, hardcoded β€” zero BPE files, zero vocab files.

Model definition: model_lib.py in this repo.


Training

Data LibriSpeech train-clean-460 (train-100 + train-360, ~460h, CC-BY-4.0)
Augmentation SpecAugment (time + freq masks)
Optimizer AdamW, lr=3e-4, weight_decay=1e-4
Schedule Linear warmup 4000 steps β†’ cosine anneal
Epochs 60 (step 426,174)
Hardware Colab A100 (~46h wall clock across two sessions)
Precision fp16 mixed (AMP)

Training notebook: tongue_stt_v3-FINAL.ipynb. Checkpoints saved every 2000 steps to Google Drive; training survived a 24h session restart via resume.


Results

Gate scorecard

Gate Target Result Status
Artifact size (int8) ≀ 18 MB 15.08 MB βœ…
test-clean WER (fp32 greedy) ≀ 15% 8.6% βœ…
int8 WER regression ≀ 1pp 0.0pp βœ…
int4 WER regression ≀ 1pp 1.1pp ❌ 0.1pp over
Short-utterance ≀3s WER gap within 3pp 3.9pp ❌ 0.9pp over
Golden-vector parity (20 utts) 20/20 18/20 (cos=1.000 all 20) ❌ 2 int8 edge cases
Numpy mel parity < 0.5 dB MAE 0.0 dB βœ…
Zero vocab files shipped β€” βœ“ βœ…

WER on LibriSpeech test-clean

Model Size WER CER Latency
Ear fp32 59.4 MB 8.6% 2.6% 3.2 ms/utt
Ear int8 15.08 MB 8.6% 2.7% 2.9 ms/utt
Ear int4 9.66 MB 9.7% 3.0% 2.9 ms/utt
Ear v1 (micro, 100h) 2.11 MB 28.7% 9.2% 1.7 ms/utt
Whisper tiny.en 142 MB 7.4%* β€” β€”
Whisper base.en 244 MB 6.8%* β€” β€”

*Whisper measured on 200-utterance ≀3s subset; Ear WER on same subset is 13.4%.

Ear is 9.4Γ— smaller than Whisper tiny.en and 1.6Γ— worse WER on the full test-clean. The win is size Γ— latency, not raw accuracy.


Failure modes

These are published as part of the product β€” honesty about what the model can't do is as important as the benchmarks.

1. Proper nouns and names β€” structural No language model means no lexical anchoring. Irish, Greek, and archaic English names from LibriSpeech's literary corpus (Joyce, Plato, Shakespeare) fail badly.

  • "stephanos dedalos" β†’ "stuffronors daelos"
  • "timaeus approves" β†’ "to me as a proves"

2. Short utterances (≀3s) β€” harder by design 12.5% WER vs 8.6% full-set (3.9pp gap). Single phrases give less temporal context for error recovery.

3. Apostrophe handling Contractions occasionally produce spurious apostrophe placement.

  • "he's swiftly punished" β†’ "he 'is swiftly punished"

4. English clean speech only Trained on LibriSpeech clean read speech. Accents, noise, and far-field audio degrade sharply. Out of scope by design.

5. No language model in the shipped artifact Greedy CTC on characters β€” homophones and phonetically similar words are indistinguishable. Reported honestly, not hidden.

6. Confidence threshold is uncalibrated uncertain=0.0% across all 2,620 test utterances. The confidence signal exists (mean non-blank CTC posterior) but has not been calibrated against accuracy. Do not rely on it for gating.


License

Code: Apache-2.0. Weights trained on LibriSpeech (CC-BY-4.0) β€” trained weights inherit CC-BY-4.0.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Dataset used to train CDT5058/ear-stt

Space using CDT5058/ear-stt 1

Collection including CDT5058/ear-stt

Evaluation results

  • Test WER (fp32 greedy) on LibriSpeech test-clean
    self-reported
    8.600
  • Test CER (fp32 greedy) on LibriSpeech test-clean
    self-reported
    2.600
  • Test WER (int8 greedy) on LibriSpeech test-clean
    self-reported
    8.600
  • Test CER (int8 greedy) on LibriSpeech test-clean
    self-reported
    2.700
  • Test WER (int4 groupwise greedy) on LibriSpeech test-clean
    self-reported
    9.700
  • Test CER (int4 groupwise greedy) on LibriSpeech test-clean
    self-reported
    3.000