bge-reranker-v2-m3 โ€” ExecuTorch

A cross-encoder reranker: a query and one document in, one relevance score out. The second stage of on-device retrieval โ€” an embedding model fetches candidates cheaply, this reads each candidate together with the query and scores it properly.

  • Source: BAAI/bge-reranker-v2-m3 โ€” 568M parameters, 24 XLM-RoBERTa layers, hidden 1024, 250k vocabulary
  • License: Apache-2.0
  • Input: input_ids and attention_mask, both [1, 512] int64
  • Output: [1, 1] fp32 โ€” the raw logit. sigmoid(x) maps it to 0..1 and does not change the ordering.

Variants

build file size (MB) worst score error vs eager Mac median (ms)* backend takes
fp32 rerank_bge_reranker_v2_m3_xnnpack_fp32.pte 2271.5 0.0000 logits 162.6 81.3%
fp16 rerank_bge_reranker_v2_m3_xnnpack_fp16.pte 1136.2 0.0061 logits 267.4 70.1%
Core ML (fp16, iOS) rerank_bge_reranker_v2_m3_coreml_all.pte 1138.2 0.0143 logits 63.6 100.0%

*Mac arm64, one query-document pair at 512 tokens, fastest of five medians of ten โ€” a reference point for relative cost, not a device number. The host shares its cores with other work, and a single median does not survive that; contention only ever adds time, so the fastest repetition is the one that means something. PyTorch eager fp32, measured the same way: 205.6 ms.

Correlation is not reported because it cannot be: the output is a single number, and the correlation of a one-element vector is undefined. The column above is the error in the units the model is used in โ€” logits โ€” over 6 real query-document pairs, and every build listed reproduces eager's ranking order exactly.

What it does, on the shipped fp32 build

Query: "How many people live in Berlin?"

rank score document
1 +7.049 Berlin has a population of 3,520,031 registered inhabitants in an area of 891.82 kmยฒ.
2 +6.728 ใƒ™ใƒซใƒชใƒณใฎไบบๅฃใฏใŠใ‚ˆใ350ไธ‡ไบบใงใ™ใ€‚
3 -0.791 In 2019 the city recorded 3.7 million residents within its metropolitan area.
4 -6.602 The capital of France is Paris, a city of about 2.1 million people.
5 -7.570 Berlin is well known for its museums, its nightlife and its history.
6 -11.036 Water boils at 100 degrees Celsius at sea level.

The narrowest gap between adjacent ranks here is 0.3211 logits.

It ranks across languages

The candidate list above includes a Japanese passage that answers the English query. This model puts it second at +6.73; ms-marco-MiniLM-L6, the English-only reranker on this shelf, scores the same passage -10.96 and puts it fifth of 6. That is what the 250k-token vocabulary buys, and it is also why this file is 25 times larger.

The attention is eager, and that is the faster export

F.scaled_dot_product_attention does not survive export as one operation. The edge dialect lowers it through _safe_softmax, whose guard against a row with no unmasked key at all leaves 11 operations XNNPACK cannot take, in every attention block โ€” scalar_tensor, where, mul.Scalar, logical_not, eq, full_like, any.dim. Each one cuts the subgraph in two.

The switch is attn_implementation="eager": transformers then builds the mask itself, as torch.finfo(dtype).min, instead of handing F.sdpa a boolean mask for PyTorch to fill with -inf.

The guard is emitted whether or not it can ever fire, and here it cannot: it triggers only on -inf, and this arm never produces one. So the two differ only about rows that have no unmasked key at all โ€” sdpa zeroes them, this one gives them a uniform row โ€” and those are padding rows, which the pooling discards and which every real query row masks out anyway. Measured with all but eight positions masked, as adversarial as this shape gets, the two graphs agree to 1.1e-05.

XNNPACK fp32 goes from 64.2% to 81.3% delegated.

Not shipped

  • int8 (dynamic) is not shipped: at 1363.2 MB it is larger than the fp16 build's 1136.2 MB, and its score error is 0.3721 logits. Dynamic int8 quantizes the linear weights and leaves the token embedding table in fp32, while fp16 halves that table too. The table here is 1024 MB of a 2272 MB model, and the arithmetic says int8 only comes out smaller when the table is under a third of the weights (757 MB) โ€” measured on ten models on this shelf, the rule called all ten correctly.

Verification

python convert/export_rerank.py bge_reranker_v2_m3
python convert/check_rerank.py bge_reranker_v2_m3 fp32

The check has two halves. One is agreement with the model run in eager, in logits and in ranking order. The other is that the ranking is useful at all: the passage that answers the question has to outscore a passage about the same subject that does not โ€” agreement alone would pass a build that ranked by document length in both arms.

(conversion scripts: executorch-models)

Downloads last month
12
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for mlboydaisuke/bge-reranker-v2-m3-ExecuTorch

Quantized
(60)
this model