--- license: apache-2.0 tags: - executorch - xnnpack - pte - on-device - automatic-speech-recognition base_model: - Qwen/Qwen3-ASR-0.6B --- # Qwen3-ASR-0.6B — ExecuTorch (30-language speech recognition) Speech in, a transcript out, in one `.pte` with several methods. An 18-layer audio tower turns mel frames into one embedding per 77 ms, those rows are spliced into the prompt where `<|audio_pad|>` sits, and a 28-layer Qwen3 decoder generates the text — the same decoder architecture as [Qwen3-Embedding-0.6B](https://huggingface.co/mlboydaisuke/Qwen3-Embedding-0.6B-ExecuTorch) elsewhere on this shelf. ``` audio_encoder_s input_features (1, 128, 100N) fp32 -> audio rows (13N, 1024) token_embeddings input_ids (1, L) int64 -> rows (1, L, 1024) text_model rows, positions, slots -> logits (1, L, 151936) ``` - **File**: `qwen3_asr_0_6b_xnnpack_8da4w.pte` — **697.0 MB**, fifteen methods - **Source**: [Qwen/Qwen3-ASR-0.6B](https://huggingface.co/Qwen/Qwen3-ASR-0.6B) — 782.4M parameters - **License**: apache-2.0 - **Languages**: 30, auto-detected — the model emits `language ` before the text ## The recipe int8 on the audio tower, int4 on the decoder, and the token table quantised to int8 by hand. That split is measured rather than assumed: at the 30-second window, int8 on the tower alone changes **no** transcript (table below), while the decoder is most of the weight and is where four bits buys the most. The table needs doing by hand because `quantize_`'s Linear-only filters skip an `nn.Embedding`, and `tie_word_embeddings` does not help — `quantize_` replaces `lm_head.weight` and the embedding keeps pointing at the old fp32 tensor. 151936 x 1024 is 622 MB, and in this bundle it appears **twice**: once as the table, once as the decoder's output projection. Tying weights in PyTorch does not tie them here. Unquantised, the same bundle is **3757.2 MB**. ## Running it **1. The mel.** `WhisperFeatureExtractor` settings: 16 kHz mono, 128 bins, `n_fft` 400, hop 160. One second of audio is 100 frames. **2. Pick the window, and pick it tight.** The file carries an audio encoder for each of 2, 3, 4, 5, 6, 7, 8, 10, 12, 15, 20, 25 and 30 seconds. Use the smallest one your clip fits and zero-pad the mel to exactly `100 x N` frames. **Do not reach for the 30-second window because it is there** — see below. **3. Size the prompt to the window.** The chat template puts a run of `<|audio_pad|>` where the audio goes, and the processor sizes that run from the clip it was given. The graph always emits `13 x N` rows, and the two have to agree. Rewrite the run to `13 x N` placeholders, then overwrite those rows with the encoder's output. **4. Decode.** `text_model` takes a dynamic sequence length, so the whole prompt goes through in one call and each generated token in a call of its own. `positions` and `slots` are both the absolute row index; the cache is causal against the full buffer. Stop at 151643 or 151645. ## The window is the thing to get right The upstream encoder takes a mask so it can ignore padding, and it uses that mask to pack only the valid positions into its output — through a `nonzero()`. That is what makes the output length depend on the data, and a fixed-shape graph cannot have it. So this conversion pins the window and treats it as fully valid. **That is not free, and here is the price.** Fifteen clips in three languages, each transcribed through the model's own masked pipeline and then again with the window padded out and the mask dropped: | padding | transcripts identical to the masked pipeline | |---|---| | tight — the next whole second | **14 of 15** | | every clip padded to 5 s | 8 of 15 | | every clip padded to 10 s | 11 of 15 | | every clip padded to 30 s | 8 of 15 | It is not a dose-response — 10 s beats 5 s — which says the failure is a near-tie tipping rather than silence gradually swamping the signal. The errors look like 停まりますか becoming 泊まりますか: same reading, wrong character. **Hence the ladder.** Thirteen encoder methods cost **3.8 MB** in total, because methods in one `.pte` share their constants: one window is 3753.4 MB in fp32 and thirteen are 3757.2 MB. There is no reason to make a caller over-pad. (Those two fp32 figures are from the pre-rewrite graph; the rewrite moved the published 8da4w file by 0.5 MB, and the +3.8 MB it is making a point about is unaffected.) ## Verification Greedy decode driven entirely by the `.pte` — its own encoder, its own token table, its own decoder and KV cache — against the whole model in eager on the identical clip and window. Nine clips, Japanese / Portuguese / Russian, windows 2–4 s: | bundle | worst CER | ignoring case and terminal punctuation | |---|---|---| | fp32 (3757.2 MB, not published) | **0.0000** | 0.0000 | | 8da4w (697.0 MB, this file) | 0.1333 | **0.0667** | Both arms run at the same window on purpose, so this measures the conversion and not the window choice, which is measured separately above. **The unquantised bundle is exact** — nine of nine, character for character. That is the number that says the graph surgery is right. **Quantised, eight of nine survive**, and the two kinds of difference are worth separating because they are not the same thing: ``` eager hoje está um dia bonito один кофе, пожалуйста pte Hoje está um dia bonito. Один кофе, пожалуйста. ``` Four of the five differing clips are **only** a capital first letter and a full stop. That is systematic — the quantised bundle punctuates where fp32 does not — and it is a real difference in what the file emits, so it stays in the raw number. It is not a misrecognition. The ninth clip is: ``` eager この電車は東京駅に停まりますか (does this train stop at Tokyo Station) pte この電車は東京駅に泊まりますか。 (does this train stay overnight at ...) ``` One character, same reading, wrong kanji — a genuine word error. It is also the same clip and the same substitution that the window table above turns up, and **the fp32 model flips on it too**: tight padding gives 停, padding to 5 s or 10 s or 30 s gives 泊, with no quantisation anywhere. This clip is a near-tie in the model itself, and anything that perturbs it — four-bit weights, a second of silence — tips it. Nine clips of synthetic speech are a control, not a benchmark: they say the conversion reproduces the fp32 model, and nothing about word error rate on real speakers. ### What the audio encoder alone costs Measured standalone at the 30-second window, swapping only the audio half and reading the transcript: | build | size (MB) | Mac ms* | XNNPACK takes | corr vs eager | worst CER† | |---|---|---|---|---|---| | XNNPACK fp32 | 746.0 | 143.6 | 71.9% | 0.999999 | 0.0000 | | XNNPACK fp16 | 381.7 | 237.6 | 64.5% | 0.999995 | 0.0000 | | XNNPACK int8 | 200.6 | 123.5 | 78.0% | 0.998266 | **0.0000** | | Core ML | 374.1 | **40.6** | 100% | 0.972872 | **0.0690** | \*Mac arm64, median of 10 — a reference point for relative cost, not a device number. Torch eager fp32 on the same machine is 600.6 ms, so the XNNPACK build is 4.2x faster than eager. †The CER column was measured on the pre-rewrite graph (below). The rewrite agrees with it to 1.3e-06 at the tensor level and leaves every transcript in this card unchanged, but the transcripts behind that column were not re-read. **Those speeds moved because the windowed attention stopped calling `F.scaled_dot_product_attention`.** It decomposes to `_safe_softmax`, whose guard for fully-masked rows is six operations XNNPACK cannot take — and the call has no mask and `is_causal=False`, so the guard can never fire. There are 18 layers x 4 windows of it at 30 seconds. Writing the softmax out took the fp32 tower from **55.8% delegated at 170.0 ms to 71.9% at 143.6 ms**, and int8 from 149.4 ms to 123.5 ms, agreeing with the old graph to 7.1e-07. Two things worth taking from that table. **int8 on the audio tower is free** in the only unit that matters, which is why the shipped bundle uses it there. And **the Core ML build is four times faster and does change transcripts** — correlation flagged it at 0.973 and the transcripts agree with the flag, so it is not published. ## Converting it yourself: the checkpoint does not load `Qwen3ASRForConditionalGeneration.from_pretrained("Qwen/Qwen3-ASR-0.6B")` on transformers 5.15.0 returns a **fully randomly initialised model** and only warns. `generate` still runs; only the text is meaningless. Four things drifted: | | checkpoint | transformers 5.15.0 | |---|---|---| | config | `thinker_config.audio_config` / `.text_config` | both at the top level | | encoder type | `qwen3_asr_audio_encoder` | `qwen3_asr_encoder` | | tensors | `thinker.model.*`, `thinker.audio_tower.*` | `model.language_model.*`, `model.audio_tower.*` | | projector | `audio_tower.proj1` / `proj2` | `multi_modal_projector.linear_1` / `linear_2` | Because the config is not read, the audio tower is built at its default 24 layers against an 18-layer checkpoint and nothing matches at all. Qwen's model card points at their own `qwen-asr` package; transformers' built-in implementation targets a different layout. `convert/export_qwen3_asr.py` rebuilds the config from `thinker_config`, renames the tensors, and asserts that nothing was left behind — and then checks the result by transcribing known speech, because a remap that loads cleanly can still be wrong. ```bash python convert/fetch_hf.py Qwen/Qwen3-ASR-0.6B python convert/export_qwen3_asr.py bundle 8da4w 2 3 4 5 6 7 8 10 12 15 20 25 30 python convert/check_qwen3_asr.py ``` torch.export -> to_edge_transform_and_lower(partitioner) -> .pte (conversion scripts: [executorch-models](https://github.com/john-rocky/executorch-models))