VoiceDictation model repository
Model artifacts for a privacy-focused, on-device dictation Android app. This repo hosts three model families:
- Roman-Urdu whisper.cpp GGML conversions (
.bin), - Roman-Urdu sherpa-onnx ONNX pairs (
.onnx), and - Dolphin attention ASR ONNX pairs (
.onnx), built for pinned-language (ur/PK) decoding on device.
1. Roman-Urdu whisper.cpp (GGML)
whisper.cpp (GGML) conversions of the
cheetos18/whisper-small-roman-urdu
fine-tune, consumed by the app's whisper.cpp backend.
The upstream repo hosts the Transformers/safetensors checkpoint; this repo hosts
the same model converted to the GGML .bin format that whisper.cpp consumes, so
the app can download it directly at runtime.
Files
| File | Format | Size | Notes |
|---|---|---|---|
ggml-model-q4_0.bin |
q4_0 quantized | ~139 MB | Default: fastest on device |
ggml-model-f16.bin |
f16 | ~466 MB | Unquantized: higher quality |
Model is small (n_vocab 51865), multilingual.
Usage
whisper-cli -m ggml-model-q4_0.bin -f audio.wav -l auto
Important: always run with language auto-detect (-l auto), NEVER force
-l ur. Forcing the native language token leaks native-script (Urdu-script)
output; -l auto keeps the transcript in clean Roman/Latin script.
Derivation
- Source:
cheetos18/whisper-small-roman-urdu(model.safetensors, verified) - Conversion: safetensors โ GGML via whisper.cpp's legacy converter
- Quantization:
whisper-quantize ggml-model.bin ggml-model-q4_0.bin q4_0 - Verified: both files transcribe a Roman-Urdu clip as
aap ka kya haal hai?under-l auto.
2. Roman-Urdu sherpa-onnx (ONNX)
Sherpa-onnxโformat ONNX exports of the same
cheetos18/whisper-small-roman-urdu
fine-tune, consumed by the app's sherpa-onnx backend
(OfflineRecognizer.from_whisper). Two precision tiers share one token file.
Files
| File | Tier | Size | Notes |
|---|---|---|---|
roman-urdu-encoder.onnx |
fp32 | ~409 MB | whisper audio encoder |
roman-urdu-decoder.onnx |
fp32 | ~559 MB | whisper text decoder (cross-attention KV caches) |
roman-urdu-encoder.int8.onnx |
int8 | ~112 MB | MatMul weights int8, activations fp32 |
roman-urdu-decoder.int8.onnx |
int8 | ~262 MB | MatMul weights int8, activations fp32 |
roman-urdu-tokens.txt |
โ | ~0.8 MB | token ID โ token list (shared) |
Tier totals: fp32 โ 969 MB, int8 โ 375 MB.
Usage
Load with sherpa-onnx using the ONNX (whisper) recognizer and the shared token file, e.g.:
sherpa-onnx-offline \
--whisper-encoder=roman-urdu-encoder.onnx \
--whisper-decoder=roman-urdu-decoder.onnx \
--tokens=roman-urdu-tokens.txt \
audio.wav
Important: always decode with fixed language en (never ur). Forcing the
native-language token leaks Urdu-script output; English keeps the transcript in clean
Roman/Latin script.
Derivation
- Source:
cheetos18/whisper-small-roman-urdu(model.safetensors) - Export: openai-whisper audio-encoder / text-decoder wrappers exported to sherpa-onnx's custom ONNX graph format (torchscript), which the stock sherpa-onnx converter produces.
- int8: onnxruntime dynamic quantization restricted to MatMul weights
(
quantize_dynamic(..., op_types_to_quantize=["MatMul"], weight_type=QInt8)). - Verified: fp32 transcribes a Roman-Urdu clip as
aap ka kya haal hai?; int8 asaap ka kya haal hai.
3. Dolphin attention ASR (ONNX)
On-device Urdu recognition that honors an explicit ur/PK language pin โ
unlike Dolphin's CTC model, which auto-detects and can leak into the wrong
script (e.g. Devanagari for Urdu speech). These are the ONNX encoder+decoder
pairs used by the app's onnxruntime-based Dolphin attention engine.
Files
dolphin-attn/ (fp16/arm tier)
units.txt token ID โ subword list (shared)
bpe.model sentencepiece BPE model (shared; DecodePieces)
base/
encoder.onnx fp16 arm, fused STFT+mel+CMVN in-graph, raw int16 in
decoder.onnx fp16 arm, attention decoder with FULL-logits output
small/
encoder.onnx
decoder.onnx
dolphin-attn-int8/ (int8 tier โ see "int8 tier" below)
base/
encoder.onnx
decoder.onnx
units.txt
small/
encoder.onnx
decoder.onnx
units.txt
(The standalone units.txt at repo root is the same Dolphin subword vocab.)
| Tier | Model | encoder | decoder | total | notes |
|---|---|---|---|---|---|
| fp16 | base |
~124 MB | ~126 MB | ~250 MB | 6 layers, 8 heads; lower latency |
| fp16 | small |
~378 MB | ~322 MB | ~699 MB | 12 layers, 12 heads; higher quality |
| int8 | base |
~110 MB | ~153 MB | ~263 MB | int8 MatMul + output layer; see below |
| int8 | small |
~367 MB | ~381 MB | ~748 MB | int8 MatMul + output layer; see below |
Notes
- Decoder
decoder.onnxis graph-surgeried: it exposes/output_layer/Gemm_output_0(the full logits over the whole 40,002-token vocabulary) in addition tomax_logit_id, so attention beam search can read the real distribution. Stock DakeQQ decoders only outputmax_logit_id(argmax), which is insufficient for beam search. - Use the FULL-logits output for content generation. The decoder also has a
language_start/language_end-sliced logits output, but that is only for LID (auto-detection); beam-searching it yields language-token garbage. - Decoding method must be
attention(beam search), length-normalized, neverattention_rescoringor greedy โ the rescoring/greedy paths override theur/PKpin and flip pinned Urdu to Devanagari. - Language pin: prepend
<sos> <ur> <PK> <asr> <notimestamp>and passlanguage_start=136/language_end=269. - fp16 tiers keep the KV cache in float16; the encoder output is already fp16.
- int8 tiers keep the KV cache in float32 (the consuming engine derives the KV cache dtype from the encoder output at load time, so it serves both tiers correctly).
- Base decoder round-trip download verified byte-identical.
int8 tier
Where dolphin-attn/ is the upstream fp16/arm conversion (+ graph surgery),
dolphin-attn-int8/ was produced locally in this project because the upstream
stock int8 tier cannot satisfy the beam-search requirement:
- The upstream repo ships three tiers: fp32 (
.onnx), fp16/arm (.onnx), and int8 (.ort). The.ortint8 files are onnxruntime flatbuffers; the full-logits graph surgery operates on the ONNX proto and cannot modify them โ so the stock int8 tier can never expose/output_layer/Gemm_output_0for beam search. - These int8
.onnxfiles are therefore the fp32 tier run through dynamic int8 quantization locally, followed by the same full-logits surgery as the fp16 tier:quantize_dynamic(..., op_types_to_quantize=["MatMul"], weight_type=QInt8). - Resulting dtype layout: attention/FFN MatMul weights and the output-layer Gemm
projection are int8; activations stay fp32. The input token embedding (a
Gather, which dynamic quantization does not touch) stays fp32 โ that one123 MB fp32 table is why int8-748 MB) is larger than fp16-small(small(~699 MB). - Verified like the fp16 tier: both int8 variants beam-decode the Urdu check clip to the same clean transcript as the fp32/fp16 tiers.
Source / provenance
- fp16 tier upstream:
onnx-community/dataocean-dolphin-asr(fp16/arm tier). - int8 tier: produced in this project from the fp32 tier of the same upstream repo
via onnxruntime dynamic int8 quantization (see "int8 tier" above); the upstream's
own int8 is
.ortand not graph-surgeriable. - These files add only the full-logits graph surgery on the stock decoder; accuracy is unchanged from the upstream fp16/arm model.