File size: 7,073 Bytes
c488bc7 acfbc1e c488bc7 acfbc1e c488bc7 acfbc1e c488bc7 acfbc1e c488bc7 acfbc1e c488bc7 acfbc1e c488bc7 acfbc1e c488bc7 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 | ---
license: apache-2.0
tags:
- executorch
- xnnpack
- pte
- on-device
- feature-extraction
- sentence-similarity
base_model:
- Qwen/Qwen3-Embedding-0.6B
---
# Qwen3-Embedding-0.6B β ExecuTorch
Multilingual text embeddings from a causal LM: text in, one 1024-dimensional vector out.
The largest and most capable embedding model on this shelf, and the only one that is a
decoder rather than an encoder.
- **Source**: Qwen/Qwen3-Embedding-0.6B β 595.8M parameters, 28 Qwen3 layers, hidden 1024,
151669 vocabulary
- **License**: apache-2.0
- **Input**: `input_ids` and `attention_mask`, both `[1, 256]` int64
- **Output**: `[1, 1024]`, last-token pooled over the mask and L2-normalised inside the graph
## The pooling is in the graph, and it is a different pooling
`1_Pooling/config.json` says `pooling_mode_lasttoken`. Every other embedding model on this
shelf averages over the sequence or takes CLS; this one takes the hidden state of the
**last real token**, because a causal model's last position is the only one that has
attended to the whole input.
| | pooling | normalised |
|---|---|---|
| all-MiniLM-L6-v2 | mean | yes |
| all-MiniLM-L12-v2 | mean | yes |
| bge-small-en-v1.5 | **CLS** | yes |
| multilingual-e5-base | mean | yes |
| multilingual-e5-large | mean | yes |
| **Qwen3-Embedding-0.6B** | **last token** | **yes** |
| paraphrase-multilingual-L12 | mean | **no** |
**Pad on either side.** Qwen's reference `last_token_pool` branches on
`attention_mask[:, -1]` to tell left padding from right. That branch is data-dependent, so
`torch.export` would freeze whichever way the example input happened to be padded and then
silently return a pad token's hidden state for the other. This graph instead locates the
last set position from the mask itself:
```python
idx = (attention_mask * torch.arange(L)).argmax(dim=1)
```
Measured against the reference implementation on the same input, both padding sides,
max_abs_diff **0.000e+00** β it selects the same token either way.
**And the padding is free on both sides.** Right padding cannot matter to a causal model:
a pad at the end attends to nothing earlier. Left padding shifts every real token's RoPE
position, which sounds like it should matter and does not β RoPE encodes *differences*, so
a constant offset cancels. Measured: `position_ids = arange + 229` and `+ 1000` both leave
the vector unchanged to 5e-7, while a **non-uniform** shift (a gap inserted at position 10)
moves it to cosine 0.9957. The null result is real, not a disconnected knob.
## The prefix is not in the graph, and it is asymmetric
`config_sentence_transformers.json` gives the query a whole task instruction and the
document the empty string:
```
query: Instruct: Given a web search query, retrieve relevant passages that answer the query
Query:{your query}
document: {your document, with nothing in front of it}
```
Unlike E5's `query: `/`passage: `, there is no symmetric mode documented. Putting the
instruction on both sides, or on neither, is not what the model was trained for. It does
not throw and does not look wrong β it retrieves worse.
## Verification
| build | file | size (MB) | Mac median (ms)* | XNNPACK takes | worst cosine vs eager | retrieval budget spent |
|---|---|---|---|---|---|---|
| XNNPACK fp32 | `embed_qwen3_embedding_0_6b_xnnpack_fp32.pte` | 2383.7 | 184.0 | 80.5% | 1.000000 | 0% |
| XNNPACK fp16 | `embed_qwen3_embedding_0_6b_xnnpack_fp16.pte` | 1192.4 | 521.5 | 70.2% | 0.999992 | 2% |
| Core ML fp32 | `embed_qwen3_embedding_0_6b_coreml_all.pte` | 1195.1 | **30.0** | 100% | 0.999910 | 17% |
\*Mac arm64, median of 10, one 256-token sequence β a reference point for relative cost,
not a device number. Eager fp32 on the same input is 179.7 ms, so the Core ML build is
**6.0x eager**, 100% delegated in one subgraph.
**The XNNPACK column moved because the export stopped going through
`F.scaled_dot_product_attention`.** It decomposes to `_safe_softmax`, whose guard for
fully-masked rows is six operations XNNPACK cannot take, once per attention block. This
model is right-padded, so no query row can lose every key and the guard can never fire:
with 249 of 256 positions masked, `attn_implementation="eager"` agrees with the sdpa graph
at cosine 1.00000000 while taking delegation from **71.8% to 80.5%** and fp32 from 203.9 ms
to 184.0 ms.
Agreement with eager is the first check and it separates nothing here. The one that decides
is the last column: rank all eight test sentences against each other, then ask whether this
build's score error is smaller than the gap between the document a query retrieves and the
runner-up. Every shipped build keeps all 8 top-1 results and spends well under half the
available room.
The vectors also have to be useful, which agreement cannot show:
```
same language: 0.596 same meaning vs 0.123 unrelated
across languages: 0.521 same meaning vs 0.157 unrelated
```
Note how much wider that is than the E5 models on this shelf, which compress everything
into 0.66β0.85. Both work; this one is easier to threshold against.
```bash
python convert/check_embed.py qwen3_embedding_0_6b fp32 # or fp16, int8, coreml
```
## Not shipped: int8, and this is the first time it was close
**int8 is 1063.9 MB β smaller than fp16's 1192.4 MB.** No other embedding model on this
shelf manages that. Dynamic int8 quantises the linear weights and leaves the token
embedding table in fp32, so it only wins when that table is a small share of the file, and
here it is: 151669 x 1024 = **621 MB of the 2384 MB model, 26%**. Every other embedding
model on the shelf is 46β69% table and loses to fp16 on size alone.
It is still not published, because the size question and the quality question are
different questions:
```
worst retrieval score shift 0.0258
closest fp32 top-1 decision 0.0234
```
The error is **110% of the room available**. On these eight sentences it happens not to
change any result β but eight deliberately dissimilar sentences are the friendliest corpus
a retrieval model will ever see, and that 0.0234 is an upper bound on what a real corpus
leaves. For scale, the same eight sentences give multilingual-e5-large a top-1 decision of
**0.0014**, and *its* int8 build β whose score error is smaller than this one's β does
change what one of the eight queries retrieves.
So the bar here is half the decision distance rather than all of it, and int8 misses it.
Correlation reads 0.990459, which would also have flagged it; the retrieval measurement is
what says by how much and why.
The attention rewrite above made int8 **13% faster (174.8 ms to 151.7 ms) and 7 points
more delegated (77.9% to 85.1%)** and did not change this verdict β it moved the score
error from 100% of the room to 110%. Speed and quality are separate questions here, and
only one of them was ever the reason this build is unpublished.
torch.export -> to_edge_transform_and_lower(partitioner) -> .pte
(conversion scripts: [executorch-models](https://github.com/john-rocky/executorch-models))
|