mlboydaisuke's picture
Upload README.md with huggingface_hub
743aa9e verified
|
Raw
History Blame Contribute Delete
9.93 kB
---
license: apache-2.0
tags:
- executorch
- xnnpack
- pte
- on-device
- automatic-speech-recognition
base_model:
- Qwen/Qwen3-ASR-0.6B
---
# Qwen3-ASR-0.6B — ExecuTorch (30-language speech recognition)
Speech in, a transcript out, in one `.pte` with several methods. An 18-layer audio
tower turns mel frames into one embedding per 77 ms, those rows are spliced into the
prompt where `<|audio_pad|>` sits, and a 28-layer Qwen3 decoder generates the text —
the same decoder architecture as
[Qwen3-Embedding-0.6B](https://huggingface.co/mlboydaisuke/Qwen3-Embedding-0.6B-ExecuTorch)
elsewhere on this shelf.
```
audio_encoder_<N>s input_features (1, 128, 100N) fp32 -> audio rows (13N, 1024)
token_embeddings input_ids (1, L) int64 -> rows (1, L, 1024)
text_model rows, positions, slots -> logits (1, L, 151936)
```
- **File**: `qwen3_asr_0_6b_xnnpack_8da4w.pte`**697.0 MB**, fifteen methods
- **Source**: [Qwen/Qwen3-ASR-0.6B](https://huggingface.co/Qwen/Qwen3-ASR-0.6B) — 782.4M parameters
- **License**: apache-2.0
- **Languages**: 30, auto-detected — the model emits `language <Name>` before the text
## The recipe
int8 on the audio tower, int4 on the decoder, and the token table quantised to int8
by hand. That split is measured rather than assumed: at the 30-second window, int8 on
the tower alone changes **no** transcript (table below), while the decoder is most of
the weight and is where four bits buys the most.
The table needs doing by hand because `quantize_`'s Linear-only filters skip an
`nn.Embedding`, and `tie_word_embeddings` does not help — `quantize_` replaces
`lm_head.weight` and the embedding keeps pointing at the old fp32 tensor. 151936 x 1024
is 622 MB, and in this bundle it appears **twice**: once as the table, once as the
decoder's output projection. Tying weights in PyTorch does not tie them here.
Unquantised, the same bundle is **3757.2 MB**.
## Running it
**1. The mel.** `WhisperFeatureExtractor` settings: 16 kHz mono, 128 bins, `n_fft` 400,
hop 160. One second of audio is 100 frames.
**2. Pick the window, and pick it tight.** The file carries an audio encoder for each
of 2, 3, 4, 5, 6, 7, 8, 10, 12, 15, 20, 25 and 30 seconds. Use the smallest one your
clip fits and zero-pad the mel to exactly `100 x N` frames. **Do not reach for the
30-second window because it is there** — see below.
**3. Size the prompt to the window.** The chat template puts a run of `<|audio_pad|>`
where the audio goes, and the processor sizes that run from the clip it was given. The
graph always emits `13 x N` rows, and the two have to agree. Rewrite the run to
`13 x N` placeholders, then overwrite those rows with the encoder's output.
**4. Decode.** `text_model` takes a dynamic sequence length, so the whole prompt goes
through in one call and each generated token in a call of its own. `positions` and
`slots` are both the absolute row index; the cache is causal against the full buffer.
Stop at 151643 or 151645.
## The window is the thing to get right
The upstream encoder takes a mask so it can ignore padding, and it uses that mask to
pack only the valid positions into its output — through a `nonzero()`. That is what
makes the output length depend on the data, and a fixed-shape graph cannot have it. So
this conversion pins the window and treats it as fully valid.
**That is not free, and here is the price.** Fifteen clips in three languages, each
transcribed through the model's own masked pipeline and then again with the window
padded out and the mask dropped:
| padding | transcripts identical to the masked pipeline |
|---|---|
| tight — the next whole second | **14 of 15** |
| every clip padded to 5 s | 8 of 15 |
| every clip padded to 10 s | 11 of 15 |
| every clip padded to 30 s | 8 of 15 |
It is not a dose-response — 10 s beats 5 s — which says the failure is a near-tie
tipping rather than silence gradually swamping the signal. The errors look like
停まりますか becoming 泊まりますか: same reading, wrong character.
**Hence the ladder.** Thirteen encoder methods cost **3.8 MB** in total, because
methods in one `.pte` share their constants: one window is 3753.4 MB in fp32 and
thirteen are 3757.2 MB. There is no reason to make a caller over-pad. (Those two fp32
figures are from the pre-rewrite graph; the rewrite moved the published 8da4w file by
0.5 MB, and the +3.8 MB it is making a point about is unaffected.)
## Verification
Greedy decode driven entirely by the `.pte` — its own encoder, its own token table,
its own decoder and KV cache — against the whole model in eager on the identical clip
and window. Nine clips, Japanese / Portuguese / Russian, windows 2–4 s:
| bundle | worst CER | ignoring case and terminal punctuation |
|---|---|---|
| fp32 (3757.2 MB, not published) | **0.0000** | 0.0000 |
| 8da4w (697.0 MB, this file) | 0.1333 | **0.0667** |
Both arms run at the same window on purpose, so this measures the conversion and not
the window choice, which is measured separately above.
**The unquantised bundle is exact** — nine of nine, character for character. That is
the number that says the graph surgery is right.
**Quantised, eight of nine survive**, and the two kinds of difference are worth
separating because they are not the same thing:
```
eager hoje está um dia bonito один кофе, пожалуйста
pte Hoje está um dia bonito. Один кофе, пожалуйста.
```
Four of the five differing clips are **only** a capital first letter and a full stop.
That is systematic — the quantised bundle punctuates where fp32 does not — and it is a
real difference in what the file emits, so it stays in the raw number. It is not a
misrecognition.
The ninth clip is:
```
eager この電車は東京駅に停まりますか (does this train stop at Tokyo Station)
pte この電車は東京駅に泊まりますか。 (does this train stay overnight at ...)
```
One character, same reading, wrong kanji — a genuine word error. It is also the same
clip and the same substitution that the window table above turns up, and **the fp32
model flips on it too**: tight padding gives 停, padding to 5 s or 10 s or 30 s gives
泊, with no quantisation anywhere. This clip is a near-tie in the model itself, and
anything that perturbs it — four-bit weights, a second of silence — tips it. Nine
clips of synthetic speech are a control, not a benchmark: they say the conversion
reproduces the fp32 model, and nothing about word error rate on real speakers.
### What the audio encoder alone costs
Measured standalone at the 30-second window, swapping only the audio half and reading
the transcript:
| build | size (MB) | Mac ms* | XNNPACK takes | corr vs eager | worst CER† |
|---|---|---|---|---|---|
| XNNPACK fp32 | 746.0 | 143.6 | 71.9% | 0.999999 | 0.0000 |
| XNNPACK fp16 | 381.7 | 237.6 | 64.5% | 0.999995 | 0.0000 |
| XNNPACK int8 | 200.6 | 123.5 | 78.0% | 0.998266 | **0.0000** |
| Core ML | 374.1 | **40.6** | 100% | 0.972872 | **0.0690** |
\*Mac arm64, median of 10 — a reference point for relative cost, not a device number.
Torch eager fp32 on the same machine is 600.6 ms, so the XNNPACK build is 4.2x faster
than eager.
†The CER column was measured on the pre-rewrite graph (below). The rewrite agrees with
it to 1.3e-06 at the tensor level and leaves every transcript in this card unchanged,
but the transcripts behind that column were not re-read.
**Those speeds moved because the windowed attention stopped calling
`F.scaled_dot_product_attention`.** It decomposes to `_safe_softmax`, whose guard for
fully-masked rows is six operations XNNPACK cannot take — and the call has no mask and
`is_causal=False`, so the guard can never fire. There are 18 layers x 4 windows of it at
30 seconds. Writing the softmax out took the fp32 tower from **55.8% delegated at
170.0 ms to 71.9% at 143.6 ms**, and int8 from 149.4 ms to 123.5 ms, agreeing with the
old graph to 7.1e-07.
Two things worth taking from that table. **int8 on the audio tower is free** in the
only unit that matters, which is why the shipped bundle uses it there. And **the Core
ML build is four times faster and does change transcripts** — correlation flagged it
at 0.973 and the transcripts agree with the flag, so it is not published.
## Converting it yourself: the checkpoint does not load
`Qwen3ASRForConditionalGeneration.from_pretrained("Qwen/Qwen3-ASR-0.6B")` on
transformers 5.15.0 returns a **fully randomly initialised model** and only warns.
`generate` still runs; only the text is meaningless. Four things drifted:
| | checkpoint | transformers 5.15.0 |
|---|---|---|
| config | `thinker_config.audio_config` / `.text_config` | both at the top level |
| encoder type | `qwen3_asr_audio_encoder` | `qwen3_asr_encoder` |
| tensors | `thinker.model.*`, `thinker.audio_tower.*` | `model.language_model.*`, `model.audio_tower.*` |
| projector | `audio_tower.proj1` / `proj2` | `multi_modal_projector.linear_1` / `linear_2` |
Because the config is not read, the audio tower is built at its default 24 layers
against an 18-layer checkpoint and nothing matches at all. Qwen's model card points at
their own `qwen-asr` package; transformers' built-in implementation targets a
different layout. `convert/export_qwen3_asr.py` rebuilds the config from
`thinker_config`, renames the tensors, and asserts that nothing was left behind — and
then checks the result by transcribing known speech, because a remap that loads
cleanly can still be wrong.
```bash
python convert/fetch_hf.py Qwen/Qwen3-ASR-0.6B
python convert/export_qwen3_asr.py bundle 8da4w 2 3 4 5 6 7 8 10 12 15 20 25 30
python convert/check_qwen3_asr.py
```
torch.export -> to_edge_transform_and_lower(partitioner) -> .pte
(conversion scripts: [executorch-models](https://github.com/john-rocky/executorch-models))