ms-marco-MiniLM-L4-v2 โ€” ExecuTorch

A cross-encoder reranker: a query and one document in, one relevance score out. The second stage of on-device retrieval โ€” an embedding model fetches candidates cheaply, this reads each candidate together with the query and scores it properly.

  • Source: cross-encoder/ms-marco-MiniLM-L4-v2 โ€” 19.2M parameters, 4 BERT layers, hidden 384
  • License: Apache-2.0
  • Input: input_ids, attention_mask and token_type_ids, each [1, 512] int64
  • Output: [1, 1] fp32 โ€” the raw logit. sigmoid(x) maps it to 0..1 and does not change the ordering.

Variants

build file size (MB) worst score error vs eager Mac median (ms)* backend takes
fp32 rerank_ms_marco_minilm_l4_xnnpack_fp32.pte 76.7 0.0000 logits 11.4 72.5%
fp16 rerank_ms_marco_minilm_l4_xnnpack_fp16.pte 38.4 0.0049 logits 19.2 63.0%
Core ML (fp16, iOS) rerank_ms_marco_minilm_l4_coreml_all.pte 39.6 0.0501 logits 4.6 100.0%

*Mac arm64, single process, median of 10, one query-document pair at 512 tokens โ€” a reference point for relative cost, not a device number. PyTorch eager fp32 on the same machine: 11.9 ms.

Correlation is not reported because it cannot be: the output is a single number, and the correlation of a one-element vector is undefined. The column above is the error in the units the model is used in โ€” logits โ€” over 6 real query-document pairs, and every build listed reproduces eager's ranking order exactly.

What it does, on the shipped fp32 build

Query: "How many people live in Berlin?"

rank score document
1 +9.147 Berlin has a population of 3,520,031 registered inhabitants in an area of 891.82 kmยฒ.
2 +2.275 In 2019 the city recorded 3.7 million residents within its metropolitan area.
3 -3.964 Berlin is well known for its museums, its nightlife and its history.
4 -5.041 The capital of France is Paris, a city of about 2.1 million people.
5 -11.677 ใƒ™ใƒซใƒชใƒณใฎไบบๅฃใฏใŠใ‚ˆใ350ไธ‡ไบบใงใ™ใ€‚
6 -11.712 Water boils at 100 degrees Celsius at sea level.

The narrowest gap between adjacent ranks here is 0.0356 logits โ€” the top two both answer the question, so their order is a coin toss and a build that swapped them would not be wrong.

token_type_ids is not optional

A BERT cross-encoder marks the document half of the pair with segment id 1, and the segment embedding is doing real work. Measured on this model with the query above and the passage that answers it, feeding zeros instead of the real segment ids moves the score from +9.15 to -4.04, a drop of 13.19 logits โ€” the best document in the list crosses into negative territory, where any threshold rejects it. The graph therefore takes three inputs. XLM-R rerankers (type_vocab_size: 1) have no second segment and take two; the signature follows the model rather than being made uniform.

The attention is eager, and that is the faster export

F.scaled_dot_product_attention does not survive export as one operation. The edge dialect lowers it through _safe_softmax, whose guard against a row with no unmasked key at all leaves eleven operations XNNPACK cannot take, in every attention block โ€” scalar_tensor, where, mul.Scalar, logical_not, eq, full_like, any.dim. Each one cuts the subgraph in two. The count is exact and does not vary by family: measured across this shelf, from a 4-layer cross-encoder to a 28-layer causal reranker, it is 11 per block every time.

The guard is emitted whether or not it can ever fire, and here it cannot. It triggers only on -inf, which reaches the graph only because the sdpa path hands F.sdpa a boolean mask for PyTorch to fill; attn_implementation="eager" masks with torch.finfo(dtype).min, a large finite number, and never produces one. So the two arms differ only on rows that have no unmasked key โ€” sdpa zeroes them, eager gives them a uniform row โ€” and those are padding rows, which the pooling discards and every real query row masks out. Measured with all but eight positions masked, as adversarial as this shape gets, the two graphs agree to 1.1e-05.

XNNPACK fp32 goes from 59.4% to 72.5% delegated.

Not shipped

  • int8 (dynamic) is not shipped: at 55.1 MB it is larger than the fp16 build's 38.4 MB, and its score error is 0.1019 logits. Dynamic int8 quantizes the linear weights and leaves the token embedding table in fp32, while fp16 halves that table too. The table here is 47 MB of a 77 MB model, and the arithmetic says int8 only comes out smaller when the table is under a third of the weights (26 MB) โ€” measured on ten models on this shelf, the rule called all ten correctly.

Verification

python convert/export_rerank.py ms_marco_minilm_l4
python convert/check_rerank.py ms_marco_minilm_l4 fp32

The check has two halves. One is agreement with the model run in eager, in logits and in ranking order. The other is that the ranking is useful at all: the passage that answers the question has to outscore a passage about the same subject that does not โ€” agreement alone would pass a build that ranked by document length in both arms.

(conversion scripts: executorch-models)

Downloads last month
9
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for mlboydaisuke/ms-marco-MiniLM-L4-v2-ExecuTorch