| --- |
| license: apache-2.0 |
| tags: |
| - executorch |
| - xnnpack |
| - pte |
| - on-device |
| - image-text-to-text |
| base_model: |
| - Qwen/Qwen3-VL-2B-Instruct |
| --- |
| # Qwen3-VL-2B-Instruct β ExecuTorch |
|
|
| A vision-language model as one `.pte` with three entry points. Ask it about a photograph on |
| the device; nothing leaves it. |
|
|
| - **Source**: Qwen/Qwen3-VL-2B-Instruct β 2.13B parameters, 28-layer decoder, 24-layer vision tower |
| - **License**: apache-2.0 |
| - **Input**: a 512x512 RGB photograph and a question |
| - **Output**: text |
|
|
| ## Methods |
|
|
| | method | takes | returns | |
| |---|---|---| |
| | `vision_encoder` | `pixel_values [1024, 1536]` | `merged [256, 2048]` **and three deepstack maps `[256, 2048]`** | |
| | `token_embeddings` | `ids [1, seq]` | `rows [1, seq, 2048]` | |
| | `text_model` | `rows`, `positions [3, 1, seq]`, `slots [1, seq]`, three deepstack maps `[1, seq, 2048]` | `logits [1, seq, 151936]` | |
|
|
| Three methods rather than three files because ExecuTorch copies a mutable buffer into every |
| method that names it, so a prefill method and a decode method would each get their own KV |
| cache and neither would see the other's writes. One graph serves a 277-token prompt and a |
| single token. |
|
|
| ## What the caller owns |
|
|
| Three things, and getting any of them wrong does not throw β it answers about the wrong |
| thing: |
|
|
| 1. **Splice** the 256 vision rows into the prompt embedding at the image token positions. |
| 2. **Scatter** the three deepstack maps to those same positions, zero everywhere else. |
| Qwen3-VL adds features from vision blocks 5, 11 and 17 into decoder layers 0, 1 and 2, |
| and only at image positions. Upstream writes that as |
| `hidden_states[visual_pos_masks, :] += embeds`, which does not export; here the maps |
| arrive already placed, so the graph just adds them. |
| 3. **Pass Qwen's own 3-D positions** from `get_rope_index`, not `arange`. The rope is |
| interleaved M-RoPE with `mrope_section [24, 20, 20]` β time, height, width. A 1-D |
| position array gives every image patch the same angle and reads as a mild degradation |
| rather than an error. |
|
|
| `convert/run_qwen3_vl.py` does all three in 40 lines. |
|
|
| ## Verification (Mac arm64, executorch 1.4.0, torch 2.13.0) |
|
|
| Parity is the last-row logits of a real prompt against `Qwen3VLForConditionalGeneration` in |
| eager fp32, per question, because a model of this size can hold a scene and lose the small |
| print. Both questions are about the same photograph. |
|
|
| | build | file | size (MB) | corr, "name on the sign" | corr, "describe it" | vision (ms) | prefill (ms) | decode | |
| |---|---|---|---|---|---|---|---| |
| | 8da8w | `qwen3_vl_2b_instruct_xnnpack_8da8w.pte` | 2453.1 | 0.993435 | 0.992611 | 568.7 | 590.6 | 17.1 tok/s | |
| | 8da4w *(not shipped)* | `qwen3_vl_2b_instruct_xnnpack_8da4w.pte` | 1697.5 | 0.739964 | 0.948887 | 566.7 | 739.4 | 19.8 tok/s | |
|
|
| The prompt is 277 tokens, of which 256 are the picture. `vision_encoder` runs |
| once per image, `text_model` once per prompt and once per generated token. |
|
|
| ### What it says |
|
|
| Asked for the name on a pub sign, against the eager model's *"Lamb & Flag"*: |
|
|
| - **8da8w** β Based on the image provided, the name of the pub is **The Lamb & Flag**. |
|
|
| This name is visible on the sign above the entrance to the establishment. The sign also includes the address " |
| - **8da4w** β Based on the text visible on the sign above the entrance of the establishment, the name of the pub is **Pine Purpose**. |
|
|
| The sign is located on the front of the building, just above |
|
|
| ## The rewrites, and what each was checked against |
|
|
| Neither this repo's generic VL bundle nor `export_llm` reaches Qwen3-VL. Five things had to |
| change, each measured against the module it replaced: |
|
|
| | what | why | checked | |
| |---|---|---| |
| | vision constants precomputed | position embeddings are bilinearly interpolated from `grid_thw` on every call | part of the 0.000e+00 below | |
| | variable-length attention β plain SDPA | one image is one sequence, so `cu_seqlens` is `[0, 1024]` | exact | |
| | patch embedding `Conv3d` β `F.linear` | portable convolution takes 3-D or 4-D only: `Expect input tensor to be 3-D or 4-D, but got, 5` | max_abs_diff 4.9e-04 | |
| | deepstack mask-assignment β add | boolean-mask assignment does not export | first-step logits corr 1.000000 | |
| | embedding table quantised by hand | see below | corr 0.999965 | |
|
|
| With the first two in place and no quantisation, the vision wrapper reproduces |
| `Qwen3VLVisionModel` at **max_abs_diff 0.000e+00** across all four outputs, and the decoder |
| wrapper reproduces the full model's first-step logits at **corr 1.000000**. |
|
|
| ## Two traps worth knowing |
|
|
| **The token embedding table escapes quantisation.** It is an `nn.Embedding`, so a filter |
| written as `isinstance(module, nn.Linear)` skips it β and `tie_word_embeddings` does not |
| save it, because `quantize_` swaps `lm_head.weight` for a new tensor while the embedding |
| keeps pointing at the old fp32 one. At 151,936 x 2048 that is **1.24 GB left in fp32**: the |
| first export of this model came out at 2630 MB. torchao's own `IntxWeightOnlyConfig` does |
| reach an `nn.Embedding` but will not lower β `Missing out variants: |
| {'torchao::dequantize_affine'}` β so the quantisation is written out by hand, one scale |
| per row. |
| |
| **512x512 is where small text lives or dies, and the resampling filter decides.** Asked |
| for the name on the sign, **the eager fp32 model** answers: |
| |
| | resize filter | eager fp32 | |
| |---|---| |
| | BICUBIC | "Lamb & Flag" β right | |
| | BILINEAR | "The Lamb & Flag" β right | |
| | LANCZOS | "The Pigeon House" β wrong | |
| | NEAREST | "Pompadour" β wrong | |
| |
| That is before any quantisation. If fine print matters for your use, the preprocessing is |
| not a detail; `convert/export_qwen3_vl.py` puts the resize behind one helper for that reason. |
| |
| ## Not shipped |
| |
| **8da4w** β the same bundle with a 4-bit decoder β is 1697.5 MB against |
| 2453.1 MB, and it does not hold: logits correlation 0.740 on the sign question |
| against 0.993, and it reads the sign as *"Pine Purpose"*. Measured where the loss is, |
| one arm at a time: the vision tower at int8 costs corr 0.990 against fp32's 0.99982, the |
| hand-quantised table costs 0.99997, and the 4-bit decoder is the rest. Reported here rather |
| than dropped, because a 1.7 GB 2B vision-language model is worth wanting and this is what it |
| costs. |
| |
| ## Conversion |
| |
| ```bash |
| python convert/export_qwen3_vl.py # check both halves against eager |
| python convert/export_qwen3_vl.py --export # write the bundle |
| python convert/check_qwen3_vl.py 8da8w # parity and speed |
| python convert/run_qwen3_vl.py photo.jpg "what is on the sign?" |
| ``` |
| |
| (conversion scripts: [executorch-models](https://github.com/john-rocky/executorch-models)) |
| |