whistle-he

A small Hebrew speech-to-text model: Cactus Whistle fine-tuned on Hebrew. 55M parameters, quantization-aware trained, 24.7 MB as a single on-device file. It takes 16 kHz mono audio and writes unvocalised Hebrew with punctuation.

This is an independent fine-tune. It is not made or endorsed by Cactus Compute or by ivrit.ai.

What is in this repository

Path Size Use
whistle-he.cact 24.7 MB the on-device file for the Cactus engine (pip install cactus-needle)
onnx/ 396 MB ONNX Runtime export: encoder, a KV-cached decoder (he_decoder_init + he_decoder_step) and a plain decoder
pytorch/whistle-he.pt 221 MB the float (latent) weights, for fine-tuning or re-export
tokenizer/he4096.model 73 KB SentencePiece BPE, 4,096 pieces
hewhistle/, transcribe.py, serve.py, web/ inference code for the ONNX and PyTorch files, and a local test page

Results

Word error rate measured with ivrit.ai's standard evaluation script (evaluate_model.py, unmodified: its Hebrew text normaliser and jiwer), on four Hebrew test sets.

Test set Cactus engine, whistle-he.cact ONNX Runtime, transcribe.py
ivrit-ai/eval-d1 (informal podcast, 5-minute excerpts) 0.102 0.088
ivrit-ai/eval-whatsapp (voice notes) 0.166 0.157
upai-inc/saspeech (single professional speaker) 0.103 0.097
imvladikon/hebrew_speech_kan (broadcast, short clips) 0.203 0.188

How to read these numbers:

  • The ONNX column is the full pipeline of this repository: beam 5 with CTC scoring in the search, guards, pauses shortened, long audio cut at quiet moments. The engine column is the bare model with the engine's own beam search, fed the same pieces (the engine takes at most 30 s per call and does not split audio itself, see "Limitations").
  • No language model and no keyword list were used. The eval-d1 episode was checked not to be among the training sources.
  • On informal speech about half of the word substitutions differ from the reference by a single letter (a one-letter prefix, or yod / vav spelling).

Use

On-device, with the Cactus engine

# pip install cactus-needle
import needle

model = needle.Whistle(weights="whistle-he.cact")
print(model.transcribe("clip.wav")["text"])        # 16 kHz mono WAV, at most 30 s

No language argument is needed (the file reports he; language="he" also works). The engine is Cactus Compute's runtime; this repository only provides a model file in its format.

With ONNX Runtime (best accuracy, any length)

pip install -r requirements.txt
python transcribe.py recording.mp3                      # any length, any sample rate
python transcribe.py clip.wav --keywords "שם פרטי, מונח"   # optional: names and terms to favour
python serve.py                                         # local test page with record / upload / live modes: http://localhost:8765

transcribe.py and serve.py add what the bare model lacks: CTC scoring inside the beam search, a guard that returns nothing on silence, noise or music instead of an invented sentence, a guard against repetition loops, pause shortening, cutting long audio at quiet moments, and parallel decoding of the pieces. Audio stays on your machine.

With PyTorch

import torch, sentencepiece as spm
from hewhistle.qtrain import HebrewWhistle

sp = spm.SentencePieceProcessor(model_file="tokenizer/he4096.model")
model = HebrewWhistle(sp.GetPieceSize(), 2)
model.enable_qat()
model.load_state_dict(torch.load("pytorch/whistle-he.pt", map_location="cpu")["model"], strict=False)
model.eval().freeze_quant()        # bake the quantised weights in; the float weights alone do not give this model's output

Model details

  • Architecture (Whistle): log-mel front end, convolutional stem (one frame per 80 ms), 8 conformer-style encoder blocks, 8 decoder blocks with gated cross-attention, tied embeddings, 4-lane hyper-connections. The encoder and stem layout is not published by Cactus; it was reconstructed and checked against the engine (encoder output cosine 0.999).
  • Quantization: trained quantization-aware. Encoder weights 2 bits, decoder weights, token embedding and n-gram tables 4 bits, activations 8 bits. The original English Whistle stores its decoder at 2 bits, which is why it is 16.9 MB and this file is 24.7 MB.
  • Vocabulary: a new Hebrew SentencePiece BPE (4,096 pieces, about 1.7 tokens per word) replaces Whistle's 8,192 pieces.
  • Training: 134,000 steps on one RTX 3090, with an auxiliary CTC head on the encoder and SpecAugment. The released weights are the average of the last ten checkpoints (steps 125,000 to 134,000).
  • Data: ivrit.ai datasets. audio-v2 (about 19,500 hours with machine transcripts, filtered by the transcripts' own confidence fields), a slice of the Knesset plenums (638 hours), and the human-transcribed crowd-transcribe-v5 (295 hours) and crowd-recital (45 hours), the last two repeated four times per pass.

Limitations

  • Accuracy is modest. Between one word in eleven and one in five is wrong on the test sets above, and it is clearly worse on formal read text dense with names, numbers and foreign words. Larger Hebrew Whisper models are more accurate.
  • Background noise hurts quickly. An earlier checkpoint of the same run went from 0.13 to 0.32 WER with white noise at 10 dB SNR. The training used no noise augmentation.
  • Small numerical differences change outputs. The 8-bit activation rounding amplifies float differences, so the engine, ONNX Runtime and PyTorch can produce slightly different transcripts for the same audio. Their error rates agree.
  • The engine file is the bare model. At most 30 s per call, the engine's own beam search, and none of the guards or long-audio handling listed above; on silence, noise or music it may invent a short sentence. Feed it pieces prepared the way hewhistle/longform.py does.
  • The compact file is larger than the original Whistle (24.7 MB against 16.9 MB) for the reason given under "Quantization". Getting to 16 MB needs retraining with a 2-bit decoder; re-quantising these weights visibly degrades output.
  • Hebrew only. The second language token is an untrained copy of the Hebrew one.
  • Machine transcripts in the training data mean the model can inherit the habits and mistakes of the system that made them.

Licence and attribution

  • Released under Apache 2.0, the licence of the base model. The weights are a modified version of Cactus Whistle (fine-tuned on Hebrew, new vocabulary).
  • Trained on data from ivrit.ai, used under the ivrit.ai License, which allows training AI models, including for commercial use, and forbids using the data to imitate a person's voice or likeness.
  • "Whistle" and "Cactus" are names of Cactus Compute, used here only to say what this model is based on.

If you use Whistle, cite it:

@misc{whistle_2026,
  title        = {Whistle: Speech Recognition for Tiny Devices},
  author       = {Mroz, Jakub and Ndubuaku, Henry and Mosoyan, Karen and Cylich, Noah and
                  Kumar, Satyajit and Sandhu, Parkirat and Shemet, Roman and Lee, Justin H.},
  year         = {2026},
  organization = {Cactus Compute, Inc.},
  howpublished = {\url{https://github.com/cactus-compute/needle}}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for MaorB/whistle-he

Quantized
(3)
this model

Datasets used to train MaorB/whistle-he