wav2vec2 XLSR-53 Russian — ExecuTorch (CTC speech recognition)

Ten seconds of Russian audio in, a transcript out, in one forward pass. No decoder loop, no beam, no KV cache, nothing to carry between calls — the encoder is the model and the transcript falls out of an argmax per 20 ms frame.

That is the difference from the Whisper conversions on this shelf, which are encoder-decoder: one encoder pass per window plus one decoder pass per generated token. A CTC model trades that machinery — and the ability to punctuate — for a much simpler thing to run on a device.

(waveform (1,160000) fp32 16 kHz mono, attention_mask (1,160000) int64)
    -> logits (1, 499, 39)
build file MB corr vs fp32 eager transcripts identical Mac ms*
fp32 wav2vec2_xlsr53_ru_xnnpack_fp32.pte 1262.3 1.000000 5/5 276.0
fp16 wav2vec2_xlsr53_ru_xnnpack_fp16.pte 656.9 0.999914 5/5 682.4
int8 (dynamic) wav2vec2_xlsr53_ru_xnnpack_int8.pte 355.6 0.999034 5/5 273.0

*Mac arm64, single process, median of 5, for a 10-second window — a reference point for relative cost, not a device number. Torch eager fp32 on the same machine: 227.9 ms. XNNPACK delegate coverage 64.5%.

There is no Core ML build. coremltools rejects the graph on a stack whose operands are a mix of int and float. This one is XNNPACK only, and the same file runs on Android and on the Mac.

Running it

1. The audio. 16 kHz mono in [-1, 1], exactly 160 000 samples. Zero-pad a shorter clip; cut a longer one. Do not normalise it — the per-utterance normalisation wav2vec2 expects is inside the graph, and it needs the mask to be correct (below).

2. The mask. attention_mask is int64 (1, 160000), 1 on the real samples and 0 on the padding. It is not optional and it is not cosmetic: padding a 3.7-second clip to 10 seconds without it changed the transcript on 4 of 5 test sentences in the Japanese build, where the zeros drag the mean and shrink the variance the normalisation divides by, and the encoder attends to the silence.

3. The decode. CTC greedy, over logits[0]:

ids = argmax per frame
collapse runs of the same id
drop the blank (id 0)
map through the repo's vocab.json, turn the word delimiter | into a space, strip

That last clause matters: without the strip this recipe returns a trailing space where the model's own tokenizer does not. The checker in the conversion repo runs the hand-written recipe and asserts it against processor.batch_decode on the same ids, which is how that was caught.

There is no punctuation in the output — the vocabulary has 39 entries and none of them are sentence marks.

Verification

Five sentences spoken by macOS say -v Milena, run through the fp32 .pte and through the eager model on the identical padded window and mask. This is what both return — the model's own reading of synthetic speech, warts and all:

сегодня хорошая погода
где находится вокзал
один кофе пожалуйста
встреча начинается в десять часов
этот поезд идет до москвы

5/5 identical to eager, worst character error rate 0.000. Synthetic speech is a control, not a benchmark — it says the conversion transcribes what the fp32 model transcribes, and nothing about word error rate on real speakers.

python convert/check_wav2vec2.py ru fp32     # or int8
python convert/audit_int8.py wav2vec2_xlsr53_ru

The conversion note worth reading

The first build of this architecture was correct and 13x slower than it should have been: 3485 ms against eager's 215 ms, at correlation 1.000000. The cost was one module.

wav2vec2 puts a positional convolution in front of its encoder, Conv1d(1024, 1024, kernel_size=128, groups=16), and it carries weight_norm — a live torch.nn.utils.parametrize parametrization. Under torch.export that makes the weight a computed tensor: the graph gets a pow and a sum where a constant should be. XNNPACK's partitioner requires a weight that is a static parameter, so it declines the convolution, and the convolution runs on the portable kernels.

Measured on that module alone at sequence 499:

as loaded                          delegated  70.0%   3206.7 ms
after remove_parametrizations      delegated 100.0%      3.6 ms

with the output identical to the last bit. Materialising the weight before export is what brought the whole model back to its proper speed. It is reported upstream as pytorch/executorch#22078, because ExecuTorch already ships the pass that would fix it without touching the model — it just does not run it during lowering.

torch.export -> to_edge_transform_and_lower(partitioner) -> .pte (conversion scripts: executorch-models)

Downloads last month
8
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for mlboydaisuke/wav2vec2-XLSR53-Russian-ExecuTorch

Quantized
(3)
this model