whistle-he
A small Hebrew speech-to-text model: Cactus Whistle fine-tuned on Hebrew. 55M parameters, quantization-aware trained, 24.7 MB as a single on-device file. It takes 16 kHz mono audio and writes unvocalised Hebrew with punctuation.
This is an independent fine-tune. It is not made or endorsed by Cactus Compute or by ivrit.ai.
What is in this repository
| Path | Size | Use |
|---|---|---|
whistle-he.cact |
24.7 MB | the on-device file for the Cactus engine (pip install cactus-needle) |
onnx/ |
396 MB | ONNX Runtime export: encoder, a KV-cached decoder (he_decoder_init + he_decoder_step) and a plain decoder |
pytorch/whistle-he.pt |
221 MB | the float (latent) weights, for fine-tuning or re-export |
tokenizer/he4096.model |
73 KB | SentencePiece BPE, 4,096 pieces |
hewhistle/, transcribe.py, serve.py, web/ |
inference code for the ONNX and PyTorch files, and a local test page |
Results
Word error rate measured with ivrit.ai's standard evaluation script
(evaluate_model.py, unmodified: its Hebrew text
normaliser and jiwer), on four Hebrew test sets.
| Test set | Cactus engine, whistle-he.cact |
ONNX Runtime, transcribe.py |
|---|---|---|
ivrit-ai/eval-d1 (informal podcast, 5-minute excerpts) |
0.102 | 0.088 |
ivrit-ai/eval-whatsapp (voice notes) |
0.166 | 0.157 |
upai-inc/saspeech (single professional speaker) |
0.103 | 0.097 |
imvladikon/hebrew_speech_kan (broadcast, short clips) |
0.203 | 0.188 |
How to read these numbers:
- The ONNX column is the full pipeline of this repository: beam 5 with CTC scoring in the search, guards, pauses shortened, long audio cut at quiet moments. The engine column is the bare model with the engine's own beam search, fed the same pieces (the engine takes at most 30 s per call and does not split audio itself, see "Limitations").
- No language model and no keyword list were used. The
eval-d1episode was checked not to be among the training sources. - On informal speech about half of the word substitutions differ from the reference by a single letter (a one-letter prefix, or yod / vav spelling).
Use
On-device, with the Cactus engine
# pip install cactus-needle
import needle
model = needle.Whistle(weights="whistle-he.cact")
print(model.transcribe("clip.wav")["text"]) # 16 kHz mono WAV, at most 30 s
No language argument is needed (the file reports he; language="he" also works). The engine is Cactus Compute's runtime;
this repository only provides a model file in its format.
With ONNX Runtime (best accuracy, any length)
pip install -r requirements.txt
python transcribe.py recording.mp3 # any length, any sample rate
python transcribe.py clip.wav --keywords "שם פרטי, מונח" # optional: names and terms to favour
python serve.py # local test page with record / upload / live modes: http://localhost:8765
transcribe.py and serve.py add what the bare model lacks: CTC scoring inside the beam search, a guard that returns
nothing on silence, noise or music instead of an invented sentence, a guard against repetition loops, pause shortening,
cutting long audio at quiet moments, and parallel decoding of the pieces. Audio stays on your machine.
With PyTorch
import torch, sentencepiece as spm
from hewhistle.qtrain import HebrewWhistle
sp = spm.SentencePieceProcessor(model_file="tokenizer/he4096.model")
model = HebrewWhistle(sp.GetPieceSize(), 2)
model.enable_qat()
model.load_state_dict(torch.load("pytorch/whistle-he.pt", map_location="cpu")["model"], strict=False)
model.eval().freeze_quant() # bake the quantised weights in; the float weights alone do not give this model's output
Model details
- Architecture (Whistle): log-mel front end, convolutional stem (one frame per 80 ms), 8 conformer-style encoder blocks, 8 decoder blocks with gated cross-attention, tied embeddings, 4-lane hyper-connections. The encoder and stem layout is not published by Cactus; it was reconstructed and checked against the engine (encoder output cosine 0.999).
- Quantization: trained quantization-aware. Encoder weights 2 bits, decoder weights, token embedding and n-gram tables 4 bits, activations 8 bits. The original English Whistle stores its decoder at 2 bits, which is why it is 16.9 MB and this file is 24.7 MB.
- Vocabulary: a new Hebrew SentencePiece BPE (4,096 pieces, about 1.7 tokens per word) replaces Whistle's 8,192 pieces.
- Training: 134,000 steps on one RTX 3090, with an auxiliary CTC head on the encoder and SpecAugment. The released weights are the average of the last ten checkpoints (steps 125,000 to 134,000).
- Data: ivrit.ai datasets.
audio-v2(about 19,500 hours with machine transcripts, filtered by the transcripts' own confidence fields), a slice of the Knesset plenums (638 hours), and the human-transcribedcrowd-transcribe-v5(295 hours) andcrowd-recital(45 hours), the last two repeated four times per pass.
Limitations
- Accuracy is modest. Between one word in eleven and one in five is wrong on the test sets above, and it is clearly worse on formal read text dense with names, numbers and foreign words. Larger Hebrew Whisper models are more accurate.
- Background noise hurts quickly. An earlier checkpoint of the same run went from 0.13 to 0.32 WER with white noise at 10 dB SNR. The training used no noise augmentation.
- Small numerical differences change outputs. The 8-bit activation rounding amplifies float differences, so the engine, ONNX Runtime and PyTorch can produce slightly different transcripts for the same audio. Their error rates agree.
- The engine file is the bare model. At most 30 s per call, the engine's own beam search, and none of the guards or
long-audio handling listed above; on silence, noise or music it may invent a short sentence. Feed it pieces prepared the
way
hewhistle/longform.pydoes. - The compact file is larger than the original Whistle (24.7 MB against 16.9 MB) for the reason given under "Quantization". Getting to 16 MB needs retraining with a 2-bit decoder; re-quantising these weights visibly degrades output.
- Hebrew only. The second language token is an untrained copy of the Hebrew one.
- Machine transcripts in the training data mean the model can inherit the habits and mistakes of the system that made them.
Licence and attribution
- Released under Apache 2.0, the licence of the base model. The weights are a modified version of Cactus Whistle (fine-tuned on Hebrew, new vocabulary).
- Trained on data from ivrit.ai, used under the ivrit.ai License, which allows training AI models, including for commercial use, and forbids using the data to imitate a person's voice or likeness.
- "Whistle" and "Cactus" are names of Cactus Compute, used here only to say what this model is based on.
If you use Whistle, cite it:
@misc{whistle_2026,
title = {Whistle: Speech Recognition for Tiny Devices},
author = {Mroz, Jakub and Ndubuaku, Henry and Mosoyan, Karen and Cylich, Noah and
Kumar, Satyajit and Sandhu, Parkirat and Shemet, Roman and Lee, Justin H.},
year = {2026},
organization = {Cactus Compute, Inc.},
howpublished = {\url{https://github.com/cactus-compute/needle}}
}
Model tree for MaorB/whistle-he
Base model
Cactus-Compute/whistle