| --- |
| license: mit |
| tags: |
| - executorch |
| - xnnpack |
| - pte |
| - on-device |
| - sentence-similarity |
| - feature-extraction |
| base_model: |
| - intfloat/multilingual-e5-large |
| --- |
| # multilingual-e5-large β ExecuTorch |
|
|
| Multilingual E5 at 24 layers. Text in, one |
| 1024-dimensional vector out, for search and retrieval that never leaves the |
| device. |
|
|
| - **Source**: intfloat/multilingual-e5-large β 24 layers, 1024 dimensions, 250,002 vocabulary |
| - **License**: mit |
| - **Input**: `input_ids` and `attention_mask`, both `[1, 256]` int64 |
| - **Output**: `[1, 1024]`, mean-pooled and L2-normalised inside the graph |
|
|
| ## The recipe is in the graph, and it was read off this repo |
|
|
| sentence-transformers stores it per model, and the shelf's seven embedding models do |
| not agree. This one pools **mean** and |
| **normalises**, read from |
| `1_Pooling/config.json` and `modules.json` rather than inferred from the family name. |
| Getting it wrong does not throw; it returns vectors that look fine and rank wrong. |
|
|
| ## The prefix is not in the graph |
|
|
| This model is trained with `query: ` in front of the text and expects it at |
| inference. That happens before tokenisation, so the `.pte` never sees it as anything |
| but tokens β and leaving it out does not throw. It returns a plausible vector that |
| retrieves worse. |
|
|
| ## Verification |
|
|
| | build | file | size (MB) | Mac ms* | backend takes | worst cosine vs eager | retrieval budget | |
| |---|---|---|---|---|---|---| |
| | fp32 | `embed_multilingual_e5_large_xnnpack_fp32.pte` | 2235.7 | 85.9 | 80.7% | 1.000000 | 0% | |
| | fp16 | `embed_multilingual_e5_large_xnnpack_fp16.pte` | 1118.2 | 142.2 | 69.6% | 0.999999 | 6% | |
| | Core ML (fp16, iOS) | `embed_multilingual_e5_large_coreml_all.pte` | 1119.5 | 27.2 | 100.0% | 0.999992 | 40% | |
|
|
| \*Mac arm64, one 256-token sequence, **fastest of five medians of ten** β a reference |
| point for relative cost, not a device number. The host shares its cores with other work, |
| and a single median does not survive that: the same eager model here measured 19.6 ms and |
| 182.8 ms twenty minutes apart. Contention only ever adds time, so the fastest repetition is |
| the one that means something. Torch eager fp32, measured the same way, is |
| 108.5 ms. |
| |
| Cosine is measured against the model run in eager through its own pooling, over eight |
| sentences. The last column is the one that decides: rank those eight against each |
| other, and ask whether this build's score error is smaller than the gap between the |
| document a query retrieves and the runner-up. Every shipped build keeps all eight |
| top-1 results. |
| |
| ## The attention is eager, and that is the faster export |
| |
| `F.scaled_dot_product_attention` does not survive export as one operation. The edge |
| dialect lowers it through `_safe_softmax`, whose guard against a row with no unmasked key |
| at all leaves **11 operations XNNPACK cannot take, in every attention |
| block** β `scalar_tensor`, `where`, `mul.Scalar`, `logical_not`, `eq`, `full_like`, `any.dim`. Each one cuts the subgraph in two. |
| |
| The switch is `attn_implementation="eager"`: transformers then builds the mask |
| itself, as `torch.finfo(dtype).min`, instead of handing `F.sdpa` a **boolean** mask |
| for PyTorch to fill with `-inf`. |
| |
| The guard is emitted whether or not it can ever fire, and here it cannot: it triggers only |
| on `-inf`, and this arm never produces one. So the two differ only about rows that have no |
| unmasked key at all β sdpa zeroes them, this one gives them a uniform row β and those are |
| padding rows, which the pooling discards and which every real query row masks out anyway. |
| Measured with all but eight positions masked, as adversarial as this shape gets, the two |
| graphs agree to 1.4e-07. |
| |
| XNNPACK fp32 goes from **63.8% to 80.7%** delegated. |
| |
| ## Not shipped: int8 |
| |
| `embed_multilingual_e5_large_xnnpack_int8.pte` is **1330.6 MB** against fp16's 1118.2 MB. Dynamic int8 quantises the |
| linear weights and leaves the token embedding table in fp32, and here that table is |
| 1024 MB of the 2235.7 MB model β **46%**. The size a build comes out at is |
| `0.5 + 1.5 x (table share)` times the fp16 build; at 46% that is |
| 1.19, so there was never a smaller file to be had. |
| |
| It is withheld on the number that decides. Ranking the eight test sentences against |
| each other, this build moves a pair score by at most **0.0083** while the |
| closest fp32 decision β the gap between the document a query retrieves and the |
| runner-up β is **0.0014**. That is **590%** of |
| the room available, against a bar of 50%. |
| |
| Correlation reads 0.997579 for this build, which no correlation gate |
| would stop. |
| |
| torch.export -> to_edge_transform_and_lower(partitioner) -> .pte |
| (conversion scripts: [executorch-models](https://github.com/john-rocky/executorch-models)) |
| |