VoiceDictation model repository

Model artifacts for a privacy-focused, on-device dictation Android app. This repo hosts three model families:

  1. Roman-Urdu whisper.cpp GGML conversions (.bin),
  2. Roman-Urdu sherpa-onnx ONNX pairs (.onnx), and
  3. Dolphin attention ASR ONNX pairs (.onnx), built for pinned-language (ur/PK) decoding on device.

1. Roman-Urdu whisper.cpp (GGML)

whisper.cpp (GGML) conversions of the cheetos18/whisper-small-roman-urdu fine-tune, consumed by the app's whisper.cpp backend.

The upstream repo hosts the Transformers/safetensors checkpoint; this repo hosts the same model converted to the GGML .bin format that whisper.cpp consumes, so the app can download it directly at runtime.

Files

File Format Size Notes
ggml-model-q4_0.bin q4_0 quantized ~139 MB Default: fastest on device
ggml-model-f16.bin f16 ~466 MB Unquantized: higher quality

Model is small (n_vocab 51865), multilingual.

Usage

whisper-cli -m ggml-model-q4_0.bin -f audio.wav -l auto

Important: always run with language auto-detect (-l auto), NEVER force -l ur. Forcing the native language token leaks native-script (Urdu-script) output; -l auto keeps the transcript in clean Roman/Latin script.

Derivation

  • Source: cheetos18/whisper-small-roman-urdu (model.safetensors, verified)
  • Conversion: safetensors โ†’ GGML via whisper.cpp's legacy converter
  • Quantization: whisper-quantize ggml-model.bin ggml-model-q4_0.bin q4_0
  • Verified: both files transcribe a Roman-Urdu clip as aap ka kya haal hai? under -l auto.

2. Roman-Urdu sherpa-onnx (ONNX)

Sherpa-onnxโ€“format ONNX exports of the same cheetos18/whisper-small-roman-urdu fine-tune, consumed by the app's sherpa-onnx backend (OfflineRecognizer.from_whisper). Two precision tiers share one token file.

Files

File Tier Size Notes
roman-urdu-encoder.onnx fp32 ~409 MB whisper audio encoder
roman-urdu-decoder.onnx fp32 ~559 MB whisper text decoder (cross-attention KV caches)
roman-urdu-encoder.int8.onnx int8 ~112 MB MatMul weights int8, activations fp32
roman-urdu-decoder.int8.onnx int8 ~262 MB MatMul weights int8, activations fp32
roman-urdu-tokens.txt โ€” ~0.8 MB token ID โ†’ token list (shared)

Tier totals: fp32 โ‰ˆ 969 MB, int8 โ‰ˆ 375 MB.

Usage

Load with sherpa-onnx using the ONNX (whisper) recognizer and the shared token file, e.g.:

sherpa-onnx-offline \
  --whisper-encoder=roman-urdu-encoder.onnx \
  --whisper-decoder=roman-urdu-decoder.onnx \
  --tokens=roman-urdu-tokens.txt \
  audio.wav

Important: always decode with fixed language en (never ur). Forcing the native-language token leaks Urdu-script output; English keeps the transcript in clean Roman/Latin script.

Derivation

  • Source: cheetos18/whisper-small-roman-urdu (model.safetensors)
  • Export: openai-whisper audio-encoder / text-decoder wrappers exported to sherpa-onnx's custom ONNX graph format (torchscript), which the stock sherpa-onnx converter produces.
  • int8: onnxruntime dynamic quantization restricted to MatMul weights (quantize_dynamic(..., op_types_to_quantize=["MatMul"], weight_type=QInt8)).
  • Verified: fp32 transcribes a Roman-Urdu clip as aap ka kya haal hai?; int8 as aap ka kya haal hai.

3. Dolphin attention ASR (ONNX)

On-device Urdu recognition that honors an explicit ur/PK language pin โ€” unlike Dolphin's CTC model, which auto-detects and can leak into the wrong script (e.g. Devanagari for Urdu speech). These are the ONNX encoder+decoder pairs used by the app's onnxruntime-based Dolphin attention engine.

Files

dolphin-attn/                       (fp16/arm tier)
  units.txt          token ID โ†’ subword list (shared)
  bpe.model          sentencepiece BPE model (shared; DecodePieces)
  base/
    encoder.onnx     fp16 arm, fused STFT+mel+CMVN in-graph, raw int16 in
    decoder.onnx     fp16 arm, attention decoder with FULL-logits output
  small/
    encoder.onnx
    decoder.onnx
dolphin-attn-int8/                  (int8 tier โ€” see "int8 tier" below)
  base/
    encoder.onnx
    decoder.onnx
    units.txt
  small/
    encoder.onnx
    decoder.onnx
    units.txt

(The standalone units.txt at repo root is the same Dolphin subword vocab.)

Tier Model encoder decoder total notes
fp16 base ~124 MB ~126 MB ~250 MB 6 layers, 8 heads; lower latency
fp16 small ~378 MB ~322 MB ~699 MB 12 layers, 12 heads; higher quality
int8 base ~110 MB ~153 MB ~263 MB int8 MatMul + output layer; see below
int8 small ~367 MB ~381 MB ~748 MB int8 MatMul + output layer; see below

Notes

  • Decoder decoder.onnx is graph-surgeried: it exposes /output_layer/Gemm_output_0 (the full logits over the whole 40,002-token vocabulary) in addition to max_logit_id, so attention beam search can read the real distribution. Stock DakeQQ decoders only output max_logit_id (argmax), which is insufficient for beam search.
  • Use the FULL-logits output for content generation. The decoder also has a language_start/language_end-sliced logits output, but that is only for LID (auto-detection); beam-searching it yields language-token garbage.
  • Decoding method must be attention (beam search), length-normalized, never attention_rescoring or greedy โ€” the rescoring/greedy paths override the ur/PK pin and flip pinned Urdu to Devanagari.
  • Language pin: prepend <sos> <ur> <PK> <asr> <notimestamp> and pass language_start=136 / language_end=269.
  • fp16 tiers keep the KV cache in float16; the encoder output is already fp16.
  • int8 tiers keep the KV cache in float32 (the consuming engine derives the KV cache dtype from the encoder output at load time, so it serves both tiers correctly).
  • Base decoder round-trip download verified byte-identical.

int8 tier

Where dolphin-attn/ is the upstream fp16/arm conversion (+ graph surgery), dolphin-attn-int8/ was produced locally in this project because the upstream stock int8 tier cannot satisfy the beam-search requirement:

  • The upstream repo ships three tiers: fp32 (.onnx), fp16/arm (.onnx), and int8 (.ort). The .ort int8 files are onnxruntime flatbuffers; the full-logits graph surgery operates on the ONNX proto and cannot modify them โ€” so the stock int8 tier can never expose /output_layer/Gemm_output_0 for beam search.
  • These int8 .onnx files are therefore the fp32 tier run through dynamic int8 quantization locally, followed by the same full-logits surgery as the fp16 tier: quantize_dynamic(..., op_types_to_quantize=["MatMul"], weight_type=QInt8).
  • Resulting dtype layout: attention/FFN MatMul weights and the output-layer Gemm projection are int8; activations stay fp32. The input token embedding (a Gather, which dynamic quantization does not touch) stays fp32 โ€” that one 123 MB fp32 table is why int8-small (748 MB) is larger than fp16-small (~699 MB).
  • Verified like the fp16 tier: both int8 variants beam-decode the Urdu check clip to the same clean transcript as the fp32/fp16 tiers.

Source / provenance

  • fp16 tier upstream: onnx-community/dataocean-dolphin-asr (fp16/arm tier).
  • int8 tier: produced in this project from the fp32 tier of the same upstream repo via onnxruntime dynamic int8 quantization (see "int8 tier" above); the upstream's own int8 is .ort and not graph-surgeriable.
  • These files add only the full-logits graph surgery on the stock decoder; accuracy is unchanged from the upstream fp16/arm model.
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support