File size: 4,728 Bytes
39abe52
545531b
 
 
 
 
 
 
 
 
 
 
 
 
39abe52
7a9ba32
545531b
7a9ba32
545531b
7a9ba32
545531b
7a9ba32
545531b
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
7a9ba32
545531b
7a9ba32
545531b
7a9ba32
545531b
 
 
 
7a9ba32
545531b
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
---
language: en
tags:
  - speech-recognition
  - asr
  - speaker-diarization
  - audio-tagging
  - speech-enhancement
  - onnx
  - sherpa-onnx
  - flutter
  - android
library_name: sherpa-onnx
pipeline_tag: automatic-speech-recognition
---

# Mum's Ears β€” ONNX Model Repository

Central model store for Mum's Ears, an Android speech-to-text app built with Flutter and [sherpa-onnx](https://github.com/k2-fsa/sherpa-onnx).

All models are downloaded at runtime with SHA-256 checksum verification.

---

## Folder Structure
mums-ears-models/ β”œβ”€β”€ en/ # English streaming ASR + punctuation β”œβ”€β”€ de/ # German streaming ASR β”œβ”€β”€ es/ # Spanish streaming ASR β”œβ”€β”€ fr/ # French streaming ASR β”œβ”€β”€ zh/ # Chinese streaming ASR β”œβ”€β”€ qwen3-asr/ # Qwen3-ASR (offline, 938 MB) β”œβ”€β”€ audio_tagging/ # CED Tiny audio event classification β”œβ”€β”€ denoiser/ # DPDFNet8 speech enhancement β”œβ”€β”€ speaker_embedder/ # CAM++ speaker embedding └── pyannote_model.onnx # Speaker diarization segmentation


---

## Model Inventory

### Streaming ASR (per-language)

Each language folder contains a streaming Zipformer transducer model fine-tuned for real-time captioning.

| File | Size | Purpose |
|------|------|---------|
| `encoder.onnx` | 70 MB (154 MB zh) | Acoustic encoder |
| `decoder.onnx` | 1 MB (5 MB zh) | Prediction decoder |
| `joiner.onnx` | 1 MB | Joint network |
| `tokens.txt` | < 1 MB | Token vocabulary |

**Languages:** `en`, `de`, `es`, `fr`, `zh`

### Punctuation (English)

| File | Size | Purpose |
|------|------|---------|
| `en/model_punct.int8.onnx` | 7 MB | Online punctuation model (int8 quantized) |
| `en/bpe_punct.vocab` | < 1 MB | BPE vocabulary |

### Speaker Diarization

| File | Size | Purpose | Original Source |
|------|------|---------|-----------------|
| `pyannote_model.onnx` | 14 MB | Speaker segmentation (pyannote 3.0) | [csukuangfj/sherpa-onnx-pyannote-segmentation-3-0](https://huggingface.co/csukuangfj/sherpa-onnx-pyannote-segmentation-3-0) |
| `speaker_embedder/<filename>` | 27 MB | CAM++ speaker embeddings (zh+en, 192-dim) | [k2-fsa/sherpa-onnx](https://github.com/k2-fsa/sherpa-onnx/releases/tag/speaker-recongition-models) |

### Audio Tagging

| File | Size | Purpose | Original Source |
|------|------|---------|-----------------|
| `audio_tagging/model.int8.onnx` | 27 MB | CED Tiny audio event classifier (int8) | [k2-fsa/sherpa-onnx-ced-tiny-audio-tagging-2024-04-19](https://huggingface.co/k2-fsa/sherpa-onnx-ced-tiny-audio-tagging-2024-04-19) |
| `audio_tagging/class_labels_indices.csv` | < 1 MB | Audio event class labels | Same |

### Speech Enhancement (Denoiser)

| File | Size | Purpose | Original Source |
|------|------|---------|-----------------|
| `denoiser/dpdfnet8.onnx` | 14 MB | DPDFNet8 speech denoiser | [k2-fsa/sherpa-onnx](https://github.com/k2-fsa/sherpa-onnx/releases/download/speech-enhancement-models/dpdfnet8.onnx) |

### Qwen3-ASR (Offline)

Large offline ASR model (int4 quantized, ~938 MB total). Optional β€” requires β‰₯ 4 GB free RAM.

| File | Size | Purpose |
|------|------|---------|
| `qwen3-asr/encoder.int8.onnx` | ~400 MB | Audio encoder |
| `qwen3-asr/decoder.int8.onnx` | ~400 MB | Language model decoder |
| `qwen3-asr/conv_frontend.onnx` | ~100 MB | Convolutional frontend |
| `qwen3-asr/tokenizer/vocab.json` | < 1 MB | Tokenizer vocabulary |
| `qwen3-asr/tokenizer/merges.txt` | < 1 MB | BPE merge rules |
| `qwen3-asr/tokenizer/tokenizer_config.json` | < 1 MB | Tokenizer config |

---

## Attribution

Models in this repository are sourced from the following open-source projects:

- **[k2-fsa/sherpa-onnx](https://github.com/k2-fsa/sherpa-onnx)** β€” Streaming ASR, speaker embedding, diarization, punctuation, denoiser, audio tagging (Apache-2.0)
- **[pyannote/segmentation-3.0](https://huggingface.co/pyannote/segmentation-3.0)** β€” Speaker diarization segmentation (MIT)
- **[Qwen/Qwen3-ASR](https://huggingface.co/Qwen/Qwen3-ASR)** β€” Offline ASR model (Apache-2.0)
- **[frankyoujian/Edge-Punct-Casing](https://github.com/frankyoujian/Edge-Punct-Casing)** β€” Punctuation model ([paper](https://arxiv.org/abs/2407.13142))

The "kroko" streaming ASR models (en/de/es/fr) are custom fine-tuned variants. The Chinese streaming model is int8-quantized.

---

## Download & Verification

Models are downloaded by the app's built-in download manager with:

- **Resumable downloads** via HTTP Range headers (`.part` files)
- **SHA-256 checksum verification** before accepting any file
- **Mirror fallback** β€” if the primary URL fails, the original source URL is tried automatically
- **WiFi-only mode** (default, toggleable)

All checksums are defined in `lib/services/model_checksums.dart`.