all-MiniLM-L12-v2 β ExecuTorch
all-MiniLM at twice the depth, on the same recipe. Text in, one 384-dimensional vector out, for search and retrieval that never leaves the device.
- Source: sentence-transformers/all-MiniLM-L12-v2 β 12 layers, 384 dimensions, 30,522 vocabulary
- License: apache-2.0
- Input:
input_idsandattention_mask, both[1, 256]int64 - Output:
[1, 384], mean-pooled and L2-normalised inside the graph
The recipe is in the graph, and it was read off this repo
sentence-transformers stores it per model, and the shelf's seven embedding models do
not agree. This one pools mean and
normalises, read from
1_Pooling/config.json and modules.json rather than inferred from the family name.
Getting it wrong does not throw; it returns vectors that look fine and rank wrong.
Verification
| build | file | size (MB) | Mac ms* | backend takes | worst cosine vs eager | retrieval budget |
|---|---|---|---|---|---|---|
| fp32 | embed_all_minilm_l12_xnnpack_fp32.pte |
133.0 | 22.2 | 78.5% | 1.000000 | 0% |
| fp16 | embed_all_minilm_l12_xnnpack_fp16.pte |
66.7 | 33.5 | 67.6% | 0.999999 | 4% |
| Core ML (fp16, iOS) | embed_all_minilm_l12_coreml_all.pte |
67.2 | 3.6 | 100.0% | 0.999969 | 23% |
*Mac arm64, median of 10, one 256-token sequence β a reference point for relative cost, not a device number. Torch eager fp32 on the same machine is 19.6 ms.
Cosine is measured against the model run in eager through its own pooling, over eight sentences. The last column is the one that decides: rank those eight against each other, and ask whether this build's score error is smaller than the gap between the document a query retrieves and the runner-up. Every shipped build keeps all eight top-1 results.
The attention is eager, and that is the faster export
F.scaled_dot_product_attention decomposes in the edge dialect to _safe_softmax,
whose guard against a fully-masked row costs six operations XNNPACK cannot take β
scalar_tensor, where, logical_not, eq, full_like, any.dim β once per
attention block, and each one cuts the subgraph in two. The guard can only ever fire
when some query row loses every key, which needs left padding or an empty
sequence. This model is right-padded, so even a row that is all padding still sees the
real tokens and the guard protects nothing.
Exporting with attn_implementation="eager" removes it. Measured with 251 of 256
positions masked β as adversarial as this shape gets β the two graphs agree to 1.4e-07,
and XNNPACK fp32 goes from 62.6% to 78.5% delegated.
Not shipped: int8
embed_all_minilm_l12_xnnpack_int8.pte is 69.5 MB against fp16's 66.7 MB. Dynamic int8 quantises the
linear weights and leaves the token embedding table in fp32, and here that table is
47 MB of the 133.0 MB model β 35%. The size a build comes out at is
0.5 + 1.5 x (table share) times the fp16 build; at 35% that is
1.03, so there was never a smaller file to be had.
It is withheld on the number that decides. Ranking the eight test sentences against each other, this build moves a pair score by at most 0.0107 while the closest fp32 decision β the gap between the document a query retrieves and the runner-up β is 0.0078. That is 138% of the room available, against a bar of 50%.
Correlation reads 0.998465 for this build, which no correlation gate would stop.
torch.export -> to_edge_transform_and_lower(partitioner) -> .pte (conversion scripts: executorch-models)
- Downloads last month
- 9
Model tree for mlboydaisuke/all-MiniLM-L12-v2-ExecuTorch
Base model
microsoft/MiniLM-L12-H384-uncased