- WakePhoneHuBERT
- What it is good for
- What it is not
- Quick start
- Outputs
- Zero-shot keyword spotting
- Voice activity
- Phones
- Aligning audio to phones
- Language identification
- Trained wake-word heads
- Fine-tuning in PyTorch
- Benchmarks: probes on frozen featurizers
- Training
- Validation
- Size and speed
- Limitations
- Files
- Licences
- What it is good for
🤖 Model card written by Claude Opus 5.5 (claude-opus-5-5) via Claude Code. Auto-generated and NOT human-reviewed.
WakePhoneHuBERT
WakePhoneHuBERT is a 3.5 MB streaming, causal speech featurizer. It turns 16 kHz mono audio into 50 feature frames per second. It is a pretrained base and feature extractor for small speech models, and two of its outputs are usable as they are:
- active-speaker voice activity, one probability per 20 ms frame;
- IPA phone posteriors over 392 eSpeak symbols, from a head trained on speech in 24 languages.
Try it in your browser: WakePhoneHuBERT Space shows voice activity, live phones and keyword spotting from typed text, from your microphone or an uploaded file. Everything runs locally in the browser.
It is the frozen int8 trunk of WakeHuBERT-tiny with these two heads attached. This is version 1.1.
What it is good for
| Use | How well, measured |
|---|---|
| Feature extraction for small heads | 521 values per frame, plus nine 256-wide trunk layers |
| Phone posteriors | 37.46% phone error rate on held-out LibriSpeech test-clean speakers; 58.79% on Multilingual LibriSpeech in seven languages |
| Alignment of a known transcript | 82.0% of word boundaries within 20 ms, with the aligner head |
| Language identification with a small probe | 72.33% on 10 VoxLingua107 languages, linear layer (chance 10%) |
| Voice activity | F1 85.92% zero-shot on LibriParty; Silero VAD v5 gets 90.51% |
| Zero-shot keyword spotting from typed phones | 87.0% recall at 1 false activation per hour for "jarvis"; fails for "alexa" |
What it is not
- Not a transcriber. It emits phones, not words, and about one phone in three is wrong (37.46% phone error rate, against 12.94% for its teacher, wav2vec2-xlsr-53-espeak-cv-ft).
- Not a better wake-word featurizer than WakeHuBERT-tiny. On three words with real-speaker tests, heads trained on its features were not better on average over two seeds than heads trained on WakeHuBERT-tiny's features (table under Trained wake-word heads). For trained wake-word heads, use WakeHuBERT-tiny, which is 1.46 MB and a third of the compute. Ready-made heads are in OpenVoiceOS/wakehubert-wakewords.
- Not a sound-event model. A probe on its features scores 52.65% on ESC-50, below WakeHuBERT-tiny (54.83%) and HuBERT-base (68.40%).
- Weak outside its training corpus. The phone error rate is 37.46% on LibriSpeech test-clean, whose training split the IPA head was trained on, but 78.74% on English FLEURS and 47.96% to 68.59% on Multilingual LibriSpeech (see Phones). Alignment and voice-activity accuracy in other languages, in conversation and in far-field audio are not measured.
OVOS plugin support is coming.
Quick start
Install the packages and download the repository:
pip install numpy onnxruntime soundfile huggingface_hub pyarrow
hf download TigreGotico/wakephonehubert --local-dir wakephonehubert
cd wakephonehubert
Every example below runs from that directory. The first one fetches a LibriSpeech dev-clean utterance (CC BY 4.0) from a small public file and saves it as example.wav and example.txt; the later examples read those two files. Any 16 kHz mono WAV works in place of example.wav.
import io
import onnxruntime as ort
import pyarrow.parquet as pq
import soundfile as sf
from huggingface_hub import hf_hub_download
# One LibriSpeech dev-clean utterance (CC BY 4.0), from a small public test file; saved for the later examples.
table = pq.read_table(hf_hub_download("hf-internal-testing/librispeech_asr_dummy",
"clean/validation-00000-of-00001.parquet", repo_type="dataset"))
row = table.slice(0, 1).to_pylist()[0]
wav, sr = sf.read(io.BytesIO(row["audio"]["bytes"]), dtype="float32")
sf.write("example.wav", wav, sr)
open("example.txt", "w").write(row["text"].lower())
print(row["id"], sr, "Hz,", len(wav) / sr, "s:", row["text"])
sess = ort.InferenceSession("wakephonehubert_int8.onnx", providers=["CPUExecutionProvider"])
names = [o.name for o in sess.get_outputs()]
out = dict(zip(names, sess.run(None, {"waveform": wav[None]})))
for name in names:
print(f"{name:9s} {out[name].shape}")
f = out["features"][0] # [frames, 521], 50 frames per second
hubert, vad, ipa = f[:, :128], f[:, 128], f[:, 129:]
print("speech frames:", int((vad >= 0.5).sum()), "of", len(vad))
Output:
1272-128104-0000 16000 Hz, 5.855 s: MISTER QUILTER IS THE APOSTLE OF THE MIDDLE CLASSES AND WE ARE GLAD TO WELCOME HIS GOSPEL
hubert (1, 292, 128)
layers (1, 292, 9, 256)
logmel (1, 584, 64)
mfcc (1, 584, 20)
vad (1, 292, 1)
ipa (1, 292, 392)
features (1, 292, 521)
speech frames: 245 of 292
The examples were run with Python 3.12, numpy 2.5.3, onnxruntime 1.30.0, soundfile 0.14.0, huggingface_hub 2.1.1 and pyarrow 25.0.1.
Outputs
The input is waveform: float32, shape [batch, samples], 16 kHz mono, scaled to -1..1. frames is samples // 320.
| Output | Shape | Rate | Meaning |
|---|---|---|---|
features |
frames × 521 | 50 Hz | hubert, vad and ipa joined: slices 0:128, 128:129, 129:521 |
hubert |
frames × 128 | 50 Hz | WakeHuBERT-tiny int8 features, bit-identical to that model |
layers |
frames × 9 × 256 | 50 Hz | trunk taps: the stem, then the eight blocks, after ReLU |
vad |
frames × 1 | 50 Hz | active-speaker probability |
ipa |
frames × 392 | 50 Hz | CTC posteriors over vocab.json; id 0 is the blank, ids 1 to 3 are never emitted |
logmel |
mel_frames × 64 | 100 Hz | normalised log-mel |
mfcc |
mel_frames × 20 | 100 Hz | DCT-II of the log-mel |
Every output has a leading batch axis. wakephonehubert_int8.onnx returns all of them; wakephonehubert_features_int8.onnx returns only features and is slightly smaller and faster.
The model is causal: frame t depends only on audio before sample 320·(t+1), over a receptive field of 41,920 samples (2.6 s). A streaming consumer that keeps 41,920 samples of context gets the same frames as an offline run. The ipa output describes the audio about 100 ms (5 frames) late.
Zero-shot keyword spotting
The
ipaoutput can spot a word from its phones alone, with no recordings of the word and no training. It works well for some words and badly for others: measure it on your word before you rely on it.
Recall on real speakers from the Picovoice wake-word benchmark, at the threshold that allows 1 false activation per hour on 31.1 h of other audio:
| Word | Phones | Best path | Sum over paths |
|---|---|---|---|
| jarvis | dʒ ɑːɹ v ɪ s | 79.4% | 87.0% |
| computer | k ə m p j uː ɾ ɚ | 74.0% | 76.4% |
| alexa | ɐ l ɛ k s ə | 33.0% | 41.6% |
On "computer", zero-shot spotting beat every trained head we measured for the word (64.5% to 73.5%). On "jarvis", trained heads lead it by 8 to 12 points. On "alexa" it is not usable, and a trained head (79.4% to 91.1%) is needed.
The test: 384 "jarvis", 411 "computer" and 315 "alexa" clips, each with 1 s of silence on both sides. The 31.1 h of negatives are LibriSpeech test-clean, test-other and a 19.9 h sample of train-other-500, plus 0.5 h of public non-speech clips (environmental sounds including ESC-50, 3 s music excerpts from FMA, kitchen recordings and ambient noise), used for evaluation only; the licences of the music, kitchen and ambient-noise sets are not stated here. Each 1.5 s window is scored after every 80 ms block, and activations closer than 2 s count once.
The score is the best CTC path of the word's phones anywhere in the window, measured against each frame's most likely symbol and divided by the number of phones; 0 is a perfect match. The example runs it on one "jarvis" clip from the benchmark (Apache-2.0):
import io
import json
import urllib.request
import numpy as np
import onnxruntime as ort
import soundfile as sf
def keyword_score(ipa, kw, combine=np.maximum):
"""Score of the keyword's phones (vocab ids) in one window of ipa posteriors [frames, 392]: the best CTC path
(combine=np.maximum) or the sum over paths (combine=np.logaddexp), starting and ending on any frame, as a log
likelihood ratio against each frame's best symbol, divided by the number of phones. A best path scores
at most 0, a perfect match."""
r = np.log(ipa + 1e-10)
r -= r.max(1, keepdims=True)
lab = np.array([0] + [x for k in kw for x in (k, 0)])
skip = np.r_[False, False, [lab[s] != 0 and lab[s] != lab[s - 2] for s in range(2, len(lab))]]
a = np.full(len(lab), -np.inf)
best = -np.inf
for t in range(len(r)):
prev = combine(a, np.r_[-np.inf, a[:-1]])
prev = combine(prev, np.where(skip, np.r_[-np.inf, -np.inf, a[:-2]], -np.inf))
prev[:2] = combine(prev[:2], 0.0) # the keyword may start on any frame
a = prev + r[t, lab]
best = max(best, a[-1], a[-2])
return best / len(kw)
symbols = json.load(open("vocab.json"))["symbols"]
# eSpeak NG en-us phones of each word, in the units of vocab.json, and the best-path threshold that gave
# 1 false activation per hour on 31.1 h of speech and non-speech audio
KEYWORDS = {
"jarvis": (["dʒ", "ɑːɹ", "v", "ɪ", "s"], -1.219),
"alexa": (["ɐ", "l", "ɛ", "k", "s", "ə"], -0.804),
"computer": (["k", "ə", "m", "p", "j", "uː", "ɾ", "ɚ"], -1.992),
}
# A real "jarvis" from the Picovoice wake-word benchmark (Apache-2.0), with 1 s of silence on each side
url = ("https://raw.githubusercontent.com/Picovoice/wake-word-benchmark/master/audio/jarvis/"
"008a6329-b20c-4cfc-9ad4-9e7034bc5148.wav")
clip, sr = sf.read(io.BytesIO(urllib.request.urlopen(url).read()), dtype="float32")
assert sr == 16000
audio = np.concatenate([np.zeros(16000, np.float32), clip, np.zeros(16000, np.float32)])
sess = ort.InferenceSession("wakephonehubert_features_int8.onnx", providers=["CPUExecutionProvider"])
buf = np.zeros(24000, np.float32) # 1.5 s window, scored after every 80 ms block
peak = {w: -np.inf for w in KEYWORDS}
for i in range(len(audio) // 1280):
buf = np.concatenate([buf[1280:], audio[i * 1280:(i + 1) * 1280]])
ipa = sess.run(["features"], {"waveform": buf[None]})[0][0, :, 129:]
for word, (phones, threshold) in KEYWORDS.items():
s = keyword_score(ipa, [symbols.index(p) for p in phones])
if s >= threshold and peak[word] < threshold:
print(f"{word} detected at {(i + 1) * 0.08:.2f} s, score {s:.3f}")
peak[word] = max(peak[word], s)
for word, (_, threshold) in KEYWORDS.items():
print(f"{word:9s} peak score {peak[word]:.3f}, threshold {threshold}")
Output:
jarvis detected at 2.24 s, score -0.799
jarvis peak score -0.007, threshold -1.219
alexa peak score -4.605, threshold -0.804
computer peak score -5.421, threshold -1.992
For the "sum over paths" column, call keyword_score(ipa, kw, combine=np.logaddexp); its thresholds at 1 false activation per hour are -0.536 ("jarvis"), -1.423 ("computer") and -0.220 ("alexa").
To spot another word, write it in the phone units of vocab.json. The phones above come from eSpeak NG 1.52.0 (en-us), as bundled by espeakng-loader 0.2.4. Other eSpeak builds can give other symbols for the same word (for example oːɹ or ɔːɹ in "more"), so pin the build you calibrate with.
pip install phonemizer==3.4.0 espeakng-loader==0.2.4
import json
import espeakng_loader
from phonemizer.backend import EspeakBackend
from phonemizer.backend.espeak.wrapper import EspeakWrapper
from phonemizer.separator import Separator
EspeakWrapper.set_library(espeakng_loader.get_library_path())
EspeakWrapper.set_data_path(espeakng_loader.get_data_path())
espeak = EspeakBackend("en-us")
print("eSpeak NG", ".".join(map(str, espeak.version())))
symbols = json.load(open("vocab.json"))["symbols"]
for word in ["jarvis", "alexa", "computer", "robot"]:
phones = espeak.phonemize([word], separator=Separator(phone=" ", word="", syllable=""), strip=True)[0].split()
missing = [p for p in phones if p not in symbols[4:]]
print(f"{word:10s} {phones}" + (f" not in vocab.json: {missing}" if missing else ""))
Output:
eSpeak NG 1.52.0
jarvis ['dʒ', 'ɑːɹ', 'v', 'ɪ', 's']
alexa ['ɐ', 'l', 'ɛ', 'k', 's', 'ə']
computer ['k', 'ə', 'm', 'p', 'j', 'uː', 'ɾ', 'ɚ']
robot ['ɹ', 'oʊ', 'b', 'ɑː', 't']
A threshold for a new word has to be set on audio without the word; the three above do not transfer.
Voice activity
The vad value (slice 128 of features) is the probability that someone in the foreground is speaking in that 20 ms frame. Background talkers, music and singing were trained towards 0.
import numpy as np
import onnxruntime as ort
import soundfile as sf
wav, sr = sf.read("example.wav", dtype="float32") # any 16 kHz mono clip
assert sr == 16000
sess = ort.InferenceSession("wakephonehubert_features_int8.onnx", providers=["CPUExecutionProvider"])
p = sess.run(["features"], {"waveform": wav[None]})[0][0, :, 128] # speech probability per 20 ms frame
speech = p >= 0.5
edges = np.flatnonzero(np.diff(np.r_[0, speech.astype(int), 0]))
for start, stop in zip(edges[::2], edges[1::2]):
print(f"speech {start * 0.02:5.2f} s to {stop * 0.02:5.2f} s")
Output:
speech 0.58 s to 5.48 s
LibriParty eval, 50 sessions (4.0 h), speech-frame F1 on a 10 ms grid. The reference marks each utterance's whole span as speech. The tuned threshold maximises F1 on the 50 dev sessions over a grid from 0.05 to 0.95 in steps of 0.05. For the vad output and for Silero the best value is 0.05, the lowest threshold searched, so their tuned figures sit at the edge of the grid and a lower threshold may score higher.
| System | F1 at 0.5 | F1 at tuned threshold |
|---|---|---|
vad, zero-shot |
67.03 | 85.92 (0.05) |
| Silero VAD v5 | 81.98 | 90.51 (0.05) |
| WebRTC VAD, mode 0 | 81.01 | binary output |
head trained on features, seeds 0 / 1 |
93.83 / 94.50 | 93.76 / 94.46 (0.65 / 0.6) |
The zero-shot output is precise (99.47% at threshold 0.05) but marks pauses inside an utterance as non-speech, which this reference counts as speech; Silero behaves the same way. The trained head is two causal convolutions (166,977 parameters) trained on the 250 LibriParty train sessions, so part of its lead is learning the reference's convention. It is not included in this repository.
Phones
Greedy decoding takes each frame's most likely symbol, merges repeats and drops ids 0 to 3. The output favours the blank; lowering the blank's log probability by 1 gives more phones and fewer errors.
import json
import numpy as np
import onnxruntime as ort
import soundfile as sf
symbols = json.load(open("vocab.json"))["symbols"] # id 0 is the CTC blank; ids 1 to 3 are never emitted
wav, sr = sf.read("example.wav", dtype="float32")
wav = np.concatenate([wav, np.zeros(3200, np.float32)]) # 200 ms of silence: the ipa output runs 100 ms late
sess = ort.InferenceSession("wakephonehubert_int8.onnx", providers=["CPUExecutionProvider"])
ipa = sess.run(["ipa"], {"waveform": wav[None]})[0][0] # [frames, 392] posteriors
logp = np.log(ipa + 1e-10)
logp[:, 0] -= 1.0 # lower the blank by 1 (tuned on LibriSpeech dev-clean); 0 gives plain greedy decoding
best = logp.argmax(1)
print(" ".join(symbols[t] for i, t in enumerate(best) if t > 3 and (i == 0 or t != best[i - 1])))
Output for "mister quilter is the apostle of the middle classes and we are glad to welcome his gospel":
v ɪ s t ə k oʊ l d ɚ ɹ ɪ z i ɐ p ɔː s əl v f m ɪ ɾ o ɔ k l a s ə z p n ɡ d l æ d t l ɛ l k ə m h ɪ z ɡ ɑː s p əl oʊ
heads/phones_probe.npz is a trained probe: a learned mix of layers, hubert, vad and ipa, then one linear CTC layer, trained on LibriSpeech train-clean-100.
import json
import numpy as np
import onnxruntime as ort
import soundfile as sf
symbols = json.load(open("vocab.json"))["symbols"]
wav, sr = sf.read("example.wav", dtype="float32")
sess = ort.InferenceSession("wakephonehubert_int8.onnx", providers=["CPUExecutionProvider"])
head = np.load("heads/phones_probe.npz") # learned mix of the outputs, then one linear CTC layer
o = dict(zip(["layers", "hubert", "vad", "ipa"], sess.run(["layers", "hubert", "vad", "ipa"], {"waveform": wav[None]})))
entries = {f"tap{i}": o["layers"][0, :, i] for i in range(9)} | {"out": o["hubert"][0], "vad": o["vad"][0], "ipa": o["ipa"][0]}
w = np.exp(head["mix"] - head["mix"].max())
w /= w.sum()
mixed = 0
for i, name in enumerate(head["names"]):
x = (entries[name] - head[f"mu_{name}"]) / head[f"sd_{name}"]
if f"proj_{name}_weight" in head:
x = x @ head[f"proj_{name}_weight"].T + head[f"proj_{name}_bias"]
mixed = mixed + w[i] * x
best = (mixed @ head["lin_weight"].T + head["lin_bias"]).argmax(1)
print(" ".join(symbols[t] for i, t in enumerate(best) if t != 0 and (i == 0 or t != best[i - 1])))
Output:
w ɪ s t ə k oʊ l d ɚ ɹ z iː ɐ p ɔː s əl v f m ɪ t əl ʌ k l s ᵻ z p d l æ d t l l k ə m h ɪ z ɡ ɔː ɑː s p oʊ
LibriSpeech test-clean, 2,620 utterances, 187,775 reference phones (eSpeak en-us through the wav2vec2-xlsr-53-espeak-cv-ft tokenizer). Phone error rate, lower is better. This is not a zero-shot result: the IPA head was trained with CTC on 40 h of LibriSpeech train-clean-100 with eSpeak en-us targets, so these are held-out speakers from the training corpus.
| System | PER % |
|---|---|
ipa, greedy |
38.76 |
ipa, greedy, blank lowered by 1 (tuned on dev-clean) |
37.46 |
phones_probe.npz recipe, seeds 0 / 1 |
35.93 / 35.95 |
| same probe on WakeHuBERT-tiny, seeds 0 / 1 | 53.74 / 54.12 |
| teacher, wav2vec2-xlsr-53-espeak-cv-ft (316 M parameters) | 12.94 |
The probe rows were measured with the probe on the PyTorch model; the example runs the seed-0 probe on the int8 ONNX outputs.
Outside that corpus the error rate is much higher. Greedy decoding of the shipped wakephonehubert_int8.onnx, blank not lowered:
| Test set | Utterances | PER % |
|---|---|---|
| FLEURS English (en_us) test | 300 | 78.74 |
| Multilingual LibriSpeech test, Dutch | 294 | 68.59 |
| Multilingual LibriSpeech test, French | 300 | 62.27 |
| Multilingual LibriSpeech test, German | 300 | 63.50 |
| Multilingual LibriSpeech test, Italian | 292 | 47.96 |
| Multilingual LibriSpeech test, Polish | 297 | 57.70 |
| Multilingual LibriSpeech test, Portuguese | 300 | 60.52 |
| Multilingual LibriSpeech test, Spanish | 300 | 50.26 |
| Multilingual LibriSpeech, seven languages pooled | 2,083 | 58.79 |
On held-out isolated words from MSWC, at the training epoch that was kept, the phone error rate against eSpeak targets is 62.37% averaged over languages and 82.0% for English (heads/report.json).
Aligning audio to phones
Given a clip and its transcript, two methods return when each word and phone starts and ends. Both work offline on a whole clip.
- CTC forced alignment over
ipaneeds nothing beyond this model and gives IPA phones. It places phone spikes rather than spans: words come out about 40 ms late at the start and 30 ms early at the end (medians). - The aligner head (
aligner/aligner_head.onnx, 321,284 parameters) readslayersandfeaturesand scores 41 ARPAbet classes per frame; a Viterbi search places the transcript's phones. It is the method to use when timing matters. It knows English phones only.
CTC forced alignment
This needs PyTorch and torchaudio 2.8 (forced_align is removed in torchaudio 2.9), and the eSpeak packages above:
pip install torch==2.8.0 torchaudio==2.8.0 --index-url https://download.pytorch.org/whl/cpu
pip install phonemizer==3.4.0 espeakng-loader==0.2.4
import json
import espeakng_loader
import numpy as np
import onnxruntime as ort
import soundfile as sf
import torch
import torchaudio.functional as F
from phonemizer.backend import EspeakBackend
from phonemizer.backend.espeak.wrapper import EspeakWrapper
from phonemizer.separator import Separator
HOP, DELAY = 0.02, 5 # 50 Hz frames; each ipa frame describes the audio 5 frames (100 ms) earlier
audio, sr = sf.read("example.wav", dtype="float32")
words = open("example.txt").read().split()
EspeakWrapper.set_library(espeakng_loader.get_library_path())
EspeakWrapper.set_data_path(espeakng_loader.get_data_path())
espeak = EspeakBackend("en-us", language_switch="remove-flags")
symbols = json.load(open("vocab.json"))["symbols"]
ids = {p: i for i, p in enumerate(symbols) if i > 3}
phonemes = espeak.phonemize(words, separator=Separator(phone=" ", word="", syllable=""), strip=True)
targets = [[ids[p] for p in w.split() if p in ids] for w in phonemes]
sess = ort.InferenceSession("wakephonehubert_int8.onnx", providers=["CPUExecutionProvider"])
ipa = sess.run(["ipa"], {"waveform": audio[None]})[0][0]
log_probs = torch.from_numpy(np.log(np.maximum(ipa, 1e-10)))
flat = torch.tensor([[i for w in targets for i in w]], dtype=torch.int32)
labels, scores = F.forced_align(log_probs[None], flat, blank=0)
spans = F.merge_tokens(labels[0], scores[0].exp())
def seconds(frame):
return max(frame - DELAY, 0) * HOP
k = 0
for word, t in zip(words, targets):
s = spans[k:k + len(t)]
k += len(t)
phones = " ".join(f"{symbols[x.token]}[{seconds(x.start):.2f}]" for x in s)
print(f"{word:9s} {seconds(s[0].start):5.2f} {seconds(s[-1].end):5.2f} {phones}")
Output (word, start, end, and each phone's spike):
mister 0.56 0.76 m[0.56] ɪ[0.60] s[0.64] t[0.70] ɚ[0.74]
quilter 0.84 1.14 k[0.84] w[0.92] ɪ[0.96] l[1.00] t[1.08] ɚ[1.12]
is 1.28 1.42 ɪ[1.28] z[1.40]
the 1.44 1.50 ð[1.44] ə[1.48]
apostle 1.54 2.02 ɐ[1.54] p[1.66] ɑː[1.76] s[1.90] əl[2.00]
of 2.16 2.26 ʌ[2.16] v[2.24]
the 2.28 2.34 ð[2.28] ə[2.32]
middle 2.40 2.56 m[2.40] ɪ[2.44] d[2.48] əl[2.54]
classes 2.66 3.22 k[2.66] l[2.70] æ[2.78] s[2.96] ᵻ[3.02] z[3.20]
and 3.36 3.44 æ[3.36] n[3.40] d[3.42]
we 3.48 3.56 w[3.48] iː[3.54]
are 3.62 3.64 ɑːɹ[3.62]
glad 3.70 4.04 ɡ[3.70] l[3.78] æ[3.84] d[4.02]
to 4.10 4.18 t[4.10] uː[4.16]
welcome 4.24 4.60 w[4.24] ɛ[4.28] l[4.36] k[4.44] ʌ[4.48] m[4.58]
his 4.66 4.82 h[4.66] ɪ[4.70] z[4.80]
gospel 4.90 5.28 ɡ[4.90] ɑː[4.98] s[5.10] p[5.22] əl[5.26]
The aligner head
Pronunciations come from CMUdict, downloaded by the example; a word missing from it becomes one spn frame and the next word absorbs its time.
import urllib.request
import numpy as np
import onnxruntime as ort
import soundfile as sf
HOP = 0.02
PHONES = ["sil", "AA", "AE", "AH", "AO", "AW", "AY", "B", "CH", "D", "DH", "EH", "ER", "EY", "F", "G",
"HH", "IH", "IY", "JH", "K", "L", "M", "N", "NG", "OW", "OY", "P", "R", "S", "SH", "T", "TH",
"UH", "UW", "V", "W", "Y", "Z", "ZH", "spn"] # the head's 41 classes, ARPAbet without stress
INDEX = {p: i for i, p in enumerate(PHONES)}
def viterbi(log_probs, words):
"""Align one list of class ids per word to [frames, 41] log posteriors. Silence is optional before, between
and after words; each phone takes at least one frame. Returns, per word, a list of (phone, start_s, end_s)."""
cls, word_of, first = [0], [-1], [False]
for k, w in enumerate(words):
for j, c in enumerate(w):
cls.append(c); word_of.append(k); first.append(j == 0)
cls.append(0); word_of.append(-1); first.append(False)
cls, first = np.array(cls), np.array(first)
T, S = len(log_probs), len(cls)
emit = log_probs[:, cls]
score = np.full(S, -1e30)
score[:2] = emit[0, :2]
back = np.zeros((T, S), np.int8) # 0 stay, 1 from the previous state, 2 skipping a silence
for t in range(1, T):
cand = np.full((3, S), -1e30)
cand[0], cand[1, 1:] = score, score[:-1]
cand[2, 2:] = np.where(first[2:], score[:-2], -1e30)
back[t] = cand.argmax(0)
score = cand[back[t], np.arange(S)] + emit[t]
s, path = (S - 1 if score[-1] >= score[-2] else S - 2), np.empty(T, int)
for t in range(T - 1, -1, -1):
path[t] = s
s -= int(back[t, s])
out = [[] for _ in words]
for s in range(S):
if word_of[s] >= 0:
frames = np.nonzero(path == s)[0]
out[word_of[s]].append((PHONES[cls[s]], frames[0] * HOP, (frames[-1] + 1) * HOP))
return out
# Pronunciations: CMUdict (BSD-style licence), first variant of each word, stress removed
url = "https://raw.githubusercontent.com/cmusphinx/cmudict/master/cmudict.dict"
lexicon = {}
for line in urllib.request.urlopen(url).read().decode().splitlines():
word, *phones = line.split("#")[0].split()
if "(" not in word:
lexicon.setdefault(word, [INDEX[p.rstrip("012")] for p in phones])
audio, sr = sf.read("example.wav", dtype="float32")
words = open("example.txt").read().split()
featurizer = ort.InferenceSession("wakephonehubert_int8.onnx", providers=["CPUExecutionProvider"])
head = ort.InferenceSession("aligner/aligner_head.onnx", providers=["CPUExecutionProvider"])
layers, features = featurizer.run(["layers", "features"], {"waveform": audio[None]})
log_probs = head.run(["logp"], {"layers": layers, "features": features})[0][0] # [frames, 41]
sequence = [lexicon.get(w, [INDEX["spn"]]) for w in words] # spn: a word the lexicon lacks
for word, phones in zip(words, viterbi(log_probs, sequence)):
detail = " ".join(f"{p}[{a:.2f}-{b:.2f}]" for p, a, b in phones)
print(f"{word:9s} {phones[0][1]:5.2f} {phones[-1][2]:5.2f} {detail}")
Output:
mister 0.50 0.78 M[0.50-0.58] IH[0.58-0.64] S[0.64-0.66] T[0.66-0.74] ER[0.74-0.78]
quilter 0.78 1.28 K[0.78-0.92] W[0.92-0.94] IH[0.94-0.96] L[0.96-1.04] T[1.04-1.12] ER[1.12-1.28]
is 1.28 1.40 IH[1.28-1.32] Z[1.32-1.40]
the 1.40 1.48 DH[1.40-1.46] AH[1.46-1.48]
apostle 1.48 2.14 AH[1.48-1.58] P[1.58-1.76] AA[1.76-1.84] S[1.84-1.96] AH[1.96-2.00] L[2.00-2.14]
of 2.14 2.28 AH[2.14-2.18] V[2.18-2.28]
the 2.28 2.34 DH[2.28-2.30] AH[2.30-2.34]
middle 2.34 2.60 M[2.34-2.42] IH[2.42-2.46] D[2.46-2.48] AH[2.48-2.54] L[2.54-2.60]
classes 2.60 3.26 K[2.60-2.70] L[2.70-2.76] AE[2.76-2.90] S[2.90-3.00] AH[3.00-3.12] Z[3.12-3.26]
and 3.30 3.44 AH[3.30-3.40] N[3.40-3.42] D[3.42-3.44]
we 3.44 3.58 W[3.44-3.54] IY[3.54-3.58]
are 3.58 3.68 AA[3.58-3.60] R[3.60-3.68]
glad 3.68 4.08 G[3.68-3.72] L[3.72-3.84] AE[3.84-4.00] D[4.00-4.08]
to 4.08 4.22 T[4.08-4.16] UW[4.16-4.22]
welcome 4.22 4.60 W[4.22-4.26] EH[4.26-4.32] L[4.32-4.38] K[4.38-4.46] AH[4.46-4.50] M[4.50-4.60]
his 4.60 4.84 HH[4.60-4.70] IH[4.70-4.74] Z[4.74-4.84]
gospel 4.84 5.52 G[4.84-4.94] AA[4.94-5.06] S[5.06-5.16] P[5.16-5.22] AH[5.22-5.26] L[5.26-5.52]
Accuracy
LibriSpeech test-clean: 2,620 utterances, 40 speakers, 5.4 h, 105,196 word boundaries. The reference is the Montreal Forced Aligner alignment of the corpus (gilkeyio/librispeech-alignments, CC BY 4.0). The aligner head was trained on 30 h of train-clean-100, from speakers not in test-clean.
| Method | Within 20 ms | Within 50 ms | Median error |
|---|---|---|---|
| words spread evenly over the speech span | 9.1% | 15.3% | 237 ms |
| teacher, CTC forced alignment | 34.8% | 69.9% | 40 ms |
CTC forced alignment over ipa |
37.0% | 70.5% | 30 ms |
| aligner head, CMUdict | 82.0% | 94.7% | 10 ms |
| aligner head, reference pronunciations | 85.0% | 96.6% | 10 ms |
Phone boundaries, aligner head with reference pronunciations: 90.8% within 20 ms and 98.2% within 50 ms, over 379,080 boundaries. The aligner head learned from Montreal Forced Aligner output and is scored against it, so these tables measure agreement with that aligner, not with hand-placed boundaries. The even-spread row knows where speech starts and ends; it is a floor, not a method.
Language identification
The mean of features over a clip separates languages. heads/lid10_linear.npz is a linear layer on that vector for ten languages (sv, nl, hy, fr, fa, ar, da, lv, fi, et), trained on VoxLingua107 (CC BY 4.0). The example uses a French utterance from the Multilingual LibriSpeech dev split (CC BY 4.0), which is not in the IPA head's training data:
import io
import tarfile
import numpy as np
import onnxruntime as ort
import soundfile as sf
from huggingface_hub import hf_hub_download
# One French utterance from the Multilingual LibriSpeech dev split (CC BY 4.0); any 16 kHz mono clip works
tar = tarfile.open(hf_hub_download("facebook/multilingual_librispeech", "data/mls_french/dev/audio/6318_5203_000.tar.gz",
repo_type="dataset"))
member = tar.getmembers()[0]
wav, sr = sf.read(io.BytesIO(tar.extractfile(member).read()), dtype="float32")
assert sr == 16000
sess = ort.InferenceSession("wakephonehubert_features_int8.onnx", providers=["CPUExecutionProvider"])
x = sess.run(["features"], {"waveform": wav[None]})[0][0].mean(0) # mean of the 521 values over the clip
head = np.load("heads/lid10_linear.npz") # linear layer on VoxLingua107 training clips, ten languages
logits = ((x - head["mu"]) / head["sd"]) @ head["W"].T + head["b"]
p = np.exp(logits - logits.max())
p /= p.sum()
print(member.name, {str(head["classes"][k]): round(float(p[k]), 3) for k in np.argsort(-p)[:3]})
Output:
6318_5203_000000.flac {'fr': 0.66, 'da': 0.147, 'nl': 0.069}
VoxLingua107, the ten languages above: 8,975 training clips, 1,025 for early stopping, 972 test clips from the official dev set. Accuracy; chance is 10%.
| System | Accuracy % |
|---|---|
nearest class centroid (heads/lid10_centroids.npz), no training |
45.47 |
heads/lid10_linear.npz |
72.33 |
| learned mix of all outputs, then linear, seeds 0 / 1 | 71.81 / 73.05 |
| same probe on WakeHuBERT-tiny, seeds 0 / 1 | 58.23 / 58.85 |
| same probe on HuBERT-base, seeds 0 / 1 | 87.55 / 87.76 |
Trained wake-word heads
heads/wakephonehubert_jarvis.onnx is an example GRU head: it reads the 75 frames of features in a 1.5 s window and returns one logit. Its calibrated threshold is in the ONNX metadata. Its positives are synthetic only: "jarvis" spoken by Microsoft Edge neural TTS voices and Google Translate TTS voices, voice-converted copies of those clips, and OmniVoice clips. Its negatives are AudioSet noise and synthetic sound-alike phrases.
import io
import urllib.request
import numpy as np
import onnxruntime as ort
import soundfile as sf
url = ("https://raw.githubusercontent.com/Picovoice/wake-word-benchmark/master/audio/jarvis/"
"008a6329-b20c-4cfc-9ad4-9e7034bc5148.wav")
clip, sr = sf.read(io.BytesIO(urllib.request.urlopen(url).read()), dtype="float32")
audio = np.concatenate([np.zeros(16000, np.float32), clip, np.zeros(16000, np.float32)])
feat = ort.InferenceSession("wakephonehubert_features_int8.onnx", providers=["CPUExecutionProvider"])
head = ort.InferenceSession("heads/wakephonehubert_jarvis.onnx", providers=["CPUExecutionProvider"])
threshold = float(head.get_modelmeta().custom_metadata_map["default_threshold"]) # calibrated per head
buf = np.zeros(24000, np.float32) # 1.5 s window, scored after every 80 ms block
for i in range(len(audio) // 1280):
buf = np.concatenate([buf[1280:], audio[i * 1280:(i + 1) * 1280]])
f = feat.run(["features"], {"waveform": buf[None]})[0] # [1, 75, 521]
p = 1 / (1 + np.exp(-head.run(["logit"], {"features": f})[0][0]))
if p >= threshold:
print(f"jarvis at {(i + 1) * 0.08:.2f} s, probability {p:.3f} (threshold {threshold})")
break
else:
print("no detection")
Output:
jarvis at 2.16 s, probability 0.398 (threshold 0.1)
Heads trained on this model's features against heads trained on WakeHuBERT-tiny's features, same recipe, two seeds each (42 and 43). Recall on the Picovoice real speakers at 1 false activation per hour, on the 31.1 h of negatives described above:
| Word | WakePhoneHuBERT heads | WakeHuBERT-tiny heads | Zero-shot, sum over paths |
|---|---|---|---|
| jarvis | 95.6%, 95.3% | 95.6%, 98.7% | 87.0% |
| alexa | 91.1%, 79.4% | 89.5%, 89.2% | 41.6% |
| computer | 64.5%, 67.6% | 72.5%, 73.5% | 76.4% |
The shipped head is the seed-42 "jarvis" head; at its calibrated threshold of 0.10 it gives 96.1% recall at 1.77 false activations per hour on this test.
Fine-tuning in PyTorch
wakephonehubert.safetensors holds float32 weights: the WakeHuBERT-tiny trunk under trunk.* (unchanged from its release at revision 3caa605725a932440442692f644b7217f54a29c3) and the heads under vad.* and ipa.*. wph_model.py defines the PyTorch classes and fetches the trunk's student.py from the WakeHuBERT-tiny repository on first use.
pip install torch==2.8.0 --index-url https://download.pytorch.org/whl/cpu
pip install safetensors huggingface_hub
The example freezes the model and trains a new clip-level head for three steps on a dummy batch:
import torch
import torch.nn as nn
from wph_model import load_wakephonehubert
model = load_wakephonehubert("wakephonehubert.safetensors") # float32 trunk and heads, in eval mode
for p in model.parameters():
p.requires_grad_(False)
class Probe(nn.Module):
"""A new clip-level head on the frozen model: a learned mix of the nine trunk taps, and the 521 features."""
def __init__(self, n_classes):
super().__init__()
self.mix = nn.Parameter(torch.zeros(9))
self.on_layers = nn.Linear(256, n_classes)
self.on_features = nn.Linear(521, n_classes)
def forward(self, wav):
with torch.no_grad():
taps, _, _ = model.taps(wav) # [batch, 9, 256, frames]
feats = model.features(wav) # [batch, frames, 521]
w = torch.softmax(self.mix, 0)
mixed = (taps * w[None, :, None, None]).sum(1).mean(-1)
return self.on_layers(mixed) + self.on_features(feats.mean(1))
torch.manual_seed(0)
probe = Probe(n_classes=4)
opt = torch.optim.Adam(probe.parameters(), lr=1e-3)
wav, labels = torch.randn(4, 32000) * 0.1, torch.tensor([0, 1, 2, 3]) # replace with your clips and labels
for step in range(3):
loss = nn.functional.cross_entropy(probe(wav), labels)
opt.zero_grad()
loss.backward()
opt.step()
print(f"step {step} loss {loss.item():.4f}")
Output:
step 0 loss 1.4156
step 1 loss 1.3933
step 2 loss 1.3844
The float32 model's outputs differ slightly from the int8 ONNX file, whose trunk is quantised. The model loads in evaluation mode, which keeps batch norm fixed; to train a head, set requires_grad_(True) on it and call its .train(). Training the trunk changes hubert, which then no longer matches WakeHuBERT-tiny.
Each head is also stored alone: heads/vad_head.safetensors and heads/ipa_head.safetensors (one state dict each; the IPA file's metadata records the blank id, the dropped ids and the 5-frame delay), with the original PyTorch files beside them.
checkpoints/ holds the full training state of the heads, with optimiser state, for continuing training. They are PyTorch pickles: load them with torch.load(path, map_location="cpu", weights_only=False) only if you trust this repository. The trunk is not in them, because it is never trained.
| File | State |
|---|---|
training_checkpoint_round3.pt |
16 epochs in three stages; the released voice-activity head |
training_checkpoint_round4_final_ipa.pt |
2 more epochs of the IPA head on isolated words only |
training_checkpoint_v1.1_ipa.pt |
7 more epochs of the IPA head on sentences, words and non-speech; the released IPA head |
The voice-activity tensors are identical in all three files.
Benchmarks: probes on frozen featurizers
Each featurizer gets the same probe: a learned softmax-weighted sum of its layers, each standardised, then one linear layer; two seeds (0 and 1) set the probe's initialisation. WakePhoneHuBERT's entries are the nine trunk taps, hubert, vad and ipa, read from the float32 PyTorch model.
Phone recognition, LibriSpeech test-clean, phone error rate (lower is better); up to 100 epochs on train-clean-100, early stopping on dev-clean:
| Featurizer | Seed 0 | Seed 1 | Mean |
|---|---|---|---|
| log-mel | 82.96 | 82.94 | 82.95 |
| WakeHuBERT-tiny | 53.81 | 54.00 | 53.90 |
| WakePhoneHuBERT | 35.76 | 35.59 | 35.68 |
With 39 ARPAbet targets (LibriSpeech lexicon, stress removed, utterances with an unknown word dropped: 7.4% of test-clean) the means are 82.03, 51.59 and 33.75. The recipe is not SUPERB's (different training length, no projection layer, no gradient clipping, a smaller lexicon), so do not set these beside published SUPERB figures.
Sound events, ESC-50, accuracy over five folds (used for evaluation only):
| Featurizer | Seed 0 | Seed 1 | Mean |
|---|---|---|---|
| log-mel | 20.00 | 20.30 | 20.15 |
| WakeHuBERT-tiny | 54.80 | 54.85 | 54.83 |
| WakePhoneHuBERT | 52.55 | 52.75 | 52.65 |
| HuBERT-base | 68.05 | 68.75 | 68.40 |
| wav2vec2-xlsr-53-espeak | 74.25 | 73.35 | 73.80 |
Language identification, VoxLingua107, ten languages, accuracy:
| Featurizer | Seed 0 | Seed 1 | Mean |
|---|---|---|---|
| log-mel | 22.43 | 21.81 | 22.12 |
| WakeHuBERT-tiny | 58.23 | 58.85 | 58.54 |
| WakePhoneHuBERT | 71.81 | 73.05 | 72.43 |
| HuBERT-base | 87.55 | 87.76 | 87.65 |
| wav2vec2-xlsr-53-espeak | 92.70 | 92.80 | 92.75 |
Differences of a point or two between rows are within what two seeds can show.
Training
The trunk is never updated. Each head reads a learned mix of the nine trunk taps, the 128-dimensional output and the log-mel, through a small causal convolutional adapter. The heads were trained on the float32 trunk and run here on the int8 trunk.
| Head | Teacher | Training data |
|---|---|---|
| Voice activity | Silero VAD v5 | MSWC words, LibriSpeech, AudioSet, FMA, mixed with other talkers, music, noise and room impulse responses |
| IPA | wav2vec2-xlsr-53-espeak-cv-ft | sentences in 24 languages, MSWC single words in 50 languages, AudioSet non-speech |
The voice-activity target is Silero's speech probability on the clean foreground speech, computed before any background is mixed in; background talkers, music and noise are targets of 0. The IPA target is the teacher's top-8 posteriors (each student frame against the teacher frame 100 ms earlier) plus CTC on phone transcripts: eSpeak en-us for English and orthography2ipa for the other languages. Arabic, Persian, Guarani, Vietnamese and Mandarin train with the teacher term only. Audio known to be non-speech is trained towards the blank.
The 24 sentence languages: English (40 h of LibriSpeech train-clean-100); Dutch, French, German, Italian, Polish, Portuguese and Spanish (Multilingual LibriSpeech); Catalan, Welsh, Russian, Persian, Turkish, Ukrainian, Arabic, Czech, Greek, Estonian, Indonesian, Kyrgyz, Mongolian, Romanian, Swedish and Latvian (FLEURS); about 1.7 h each outside English.
| Head | Parameters | Stored as |
|---|---|---|
| Voice activity | 85,098 | float16 |
| IPA | 903,953 | float16 |
Validation
VALIDATION.md and validation_report.json hold the full tables; validation_clips.json lists the clips.
- Trunk identity. On 24 held-out inputs,
hubertequals the output of WakeHuBERT-tiny'swakehubert_int8.onnx(sha2560b6a7047…e064) exactly. - Graph against PyTorch. The shipped graph matches the PyTorch heads to 4.5e-4 for voice activity and 2.6e-3 for IPA probabilities, with every decision the same.
- Effect of the int8 trunk. Against the float32 trunk the heads were trained on, voice-activity decisions agree on 99.90% of frames and the IPA argmax on 99.75%.
- Non-speech. The voice-activity output is below 0.5 on every frame of the 3 AudioSet and 3 FMA held-out segments.
The real clips in this check are 14 LibriSpeech utterances from one speaker and 6 non-speech segments, so it tests the export, not accuracy.
Size and speed
| Model | File size | Median time per call |
|---|---|---|
wakephonehubert_int8.onnx |
3,482,985 bytes | 3.16 ms |
wakephonehubert_features_int8.onnx |
3,467,291 bytes | 3.06 to 3.07 ms |
WakeHuBERT-tiny wakehubert_int8.onnx |
1,462,845 bytes | 0.97 ms |
One call is a 1.5 s window, as a streaming consumer runs it every 80 ms: one CPU thread, onnxruntime's CPU provider, on a shared AMD EPYC 9554. That is about 4% of real time. Embedded CPUs were not measured.
Limitations
- Voice activity was trained with English-centric foreground speech, and the labels on background speech come from how the mixes were built, not from annotation. Other languages, a television at close range and a quiet far talker are not measured.
- The IPA head was trained on isolated words and read sentences. Its phone error rate is 37.46% on held-out LibriSpeech speakers but 47.96% to 78.74% on other read-speech corpora and 62.37% on isolated words; use it as a phonetic feature or a phone spotter, not as a recogniser.
- Zero-shot keyword spotting depends on the word: measured recall at 1 false activation per hour ranges from 41.6% to 87.0% over three words.
- Alignment is offline, English-only for the aligner head, at 20 ms resolution, and needs the transcript.
Files
| File | Contents |
|---|---|
wakephonehubert_int8.onnx |
the full graph with every output |
wakephonehubert_features_int8.onnx |
the same graph with only features |
wakephonehubert.safetensors, wph_model.py |
float32 weights and PyTorch classes |
vocab.json, config.json |
the 392 IPA symbols in output order; output layout and head details |
heads/vad_head.*, heads/ipa_head.* |
each head alone, safetensors and PyTorch |
heads/report.json |
training data counts and history of the IPA head |
heads/phones_probe.npz |
trained phone probe |
heads/lid10_linear.npz, heads/lid10_centroids.npz |
ten-language identification heads |
heads/wakephonehubert_jarvis.onnx |
example "jarvis" wake-word head |
aligner/ |
the aligner head, ONNX and safetensors |
checkpoints/ |
training checkpoints of the heads |
VALIDATION.md, validation_*.json |
export validation |
Licences
The model is released under the Apache License 2.0.
| Item | Licence | Role |
|---|---|---|
| WakeHuBERT-tiny | Apache-2.0 | frozen trunk |
| wav2vec2-xlsr-53-espeak-cv-ft | Apache-2.0 | IPA teacher |
| Silero VAD v5 | MIT | voice-activity teacher |
| EfficientAT mn10_as | MIT | used only to flag non-speech windows |
| LibriSpeech | CC BY 4.0 | training, evaluation |
| MSWC, from Common Voice | CC BY 4.0 | training |
| Multilingual LibriSpeech | CC BY 4.0 | training |
| FLEURS | CC BY 4.0 | training |
| AudioSet | labels CC BY 4.0; audio under YouTube uploaders' terms | training |
| FMA | only tracks under CC BY, CC BY-SA, CC0, Free Art or public domain (commercial use allowed) | training |
| MIT reverberation survey impulse responses | no licence stated by the source | training augmentation |
| VoxLingua107 | CC BY 4.0 | language heads, evaluation |
| Picovoice wake-word benchmark | Apache-2.0 | evaluation |
| LibriParty | built from LibriSpeech and noise corpora; no licence of its own stated | evaluation |
| ESC-50 | CC BY-NC 3.0 | evaluation only |
| Montreal Forced Aligner alignments | CC BY 4.0 | aligner training and reference |
| CMUdict | BSD-style | aligner example |
| Synthetic speech: Microsoft Edge neural TTS, Google Translate TTS, voice-converted copies, OmniVoice | generated audio; each service's terms apply | example "jarvis" head |
- Downloads last month
- 206
Model tree for TigreGotico/wakephonehubert
Base model
facebook/hubert-base-ls960