Mizan-Rerank-v3-Turbo

The accuracy of Mizan-Rerank-v3 at half the size and twice the speed: a compact Arabic reranker that tells the passage that answers the question apart from passages that only look like they do.

Parameters Speed Language Context Demo License

Try it: live demo — rerank your own passages, run the eleven adversarial trap cases, and switch between this model and Mizan-Rerank-v3 to compare.

The Mizan reranker family

Model Parameters Throughput Held-out mean nDCG@10 Fiqh nDCG@10
Mizan-Rerank-v3-Turbo (this model) 150M 2,487 pairs/s 0.875 0.761
Mizan-Rerank-v3 306M 1,238 pairs/s 0.826 0.669

Pick Turbo for the best accuracy per millisecond, v3 if you prefer the larger model with the same behaviour as the previous release.

Held-out mean nDCG@10 is the mean over the four held-out sets both cards share (NamaaMrTydi unseen, short adversarial, long-context adversarial, Multi-LLM). Mizan-Rerank-v3's own card quotes 0.845, which averages a fifth set (Arabic Hard Negatives unseen) that cannot be used here: only 10 of its 12,373 cases are unseen for Turbo.

Mizan-Rerank-v3-Turbo is a 150M-parameter Arabic cross-encoder. It starts from the Mizan-Embeddings-0.1B encoder and is trained in three stages: broad Arabic relevance with hard negatives and teacher scores, in-domain distillation on production retrieval candidates, then v3's adversarial and long-context recipe. It learns to separate the correct passage from adversarial "trap" passages that share most of its words but change its meaning: a negated ruling, a swapped entity, a shifted number or date, an exception applied to the wrong case.

The result is the most accurate model in the family while being the smallest: half the parameters of Mizan-Rerank-v3, twice its throughput, 3.2× less peak GPU memory, and higher scores on every held-out benchmark we measure, including the real production fiqh queries.

Summary

Highlights

  • Best accuracy in the family. Held-out mean nDCG@10 is 0.875 at 150M parameters: Mizan-Rerank-v3 scores 0.826, bge-reranker-v2-m3 (568M) 0.777, gte-multilingual-reranker-base 0.752 and Mizan-Rerank-V2 0.685.
  • Fast and small. On one RTX 4090 in fp16 it scores 2,487 pairs/s (2.0× v3, 4.1× bge-reranker-v2-m3) with 0.34 GiB peak GPU memory (see Efficiency).
  • Best on real production queries. On 224 real fiqh search queries with the production retriever's candidate lists, nDCG@10 is 0.761 against 0.669 for v3 and 0.737 for bge-reranker-v2-m3 (+0.092 over v3, 95% CI [+0.069, +0.115]). It also keeps more of the relevant rulings in production's top-8 filter (0.749).
  • Rarely puts a contradiction on top. Across 2,666 held-out adversarial queries, the top-ranked passage is a contradicting trap 16.9% of the time. For Mizan-Rerank-v3 the figure is 32.5%, for bge-reranker-v2-m3 56.8% and for Mizan-Rerank-V2 73.0%.
  • Calibrated scores. The sigmoid score works directly with a fixed cut-off: the scores were affine-calibrated after distillation (see Practical notes).

Usage

Sentence Transformers

pip install -U sentence-transformers
from sentence_transformers import CrossEncoder

model = CrossEncoder("ALJIACHI/Mizan-Rerank-v3-Turbo", max_length=3072)

query = "ما هي فوائد فيتامين د؟"
passages = [
    "يساعد فيتامين د على امتصاص الكالسيوم وتقوية العظام، كما يدعم عمل الجهاز المناعي.",
    "يستخدم فيتامين د في بعض الصناعات الغذائية كمادة مضافة لتدعيم الحليب.",
    "أطلقت وزارة الزراعة حملة وطنية لزيادة الوعي بأهمية الزراعة العضوية.",
]

scores = model.predict([(query, passage) for passage in passages])
print(scores)  # ≈ [0.98, 0.60, 0.01]  (sigmoid relevance scores; the second passage is on-topic but does not answer
#               the question, so it keeps a middling score while the answer is ranked first)

for hit in model.rank(query, passages, return_documents=True):
    print(f"{hit['score']:.3f}  {hit['text']}")

Transformers

import torch
from transformers import AutoModelForSequenceClassification, AutoTokenizer

tokenizer = AutoTokenizer.from_pretrained("ALJIACHI/Mizan-Rerank-v3-Turbo")
model = AutoModelForSequenceClassification.from_pretrained(
    "ALJIACHI/Mizan-Rerank-v3-Turbo", torch_dtype=torch.float16
).to("cuda").eval()

query = "ما هي فوائد فيتامين د؟"
passages = [
    "يساعد فيتامين د على امتصاص الكالسيوم وتقوية العظام، كما يدعم عمل الجهاز المناعي.",
    "أطلقت وزارة الزراعة حملة وطنية لزيادة الوعي بأهمية الزراعة العضوية.",
]

with torch.inference_mode():
    features = tokenizer([query] * len(passages), passages, padding=True, truncation=True,
                         max_length=3072, return_tensors="pt").to("cuda")
    scores = torch.sigmoid(model(**features).logits.squeeze(-1).float())

for passage, score in sorted(zip(passages, scores.tolist()), key=lambda item: -item[1]):
    print(f"{score:.3f}  {passage}")

Practical notes

  • No trust_remote_code needed. The architecture is plain ModernBERT, so the model loads with any Transformers version that supports ModernBertForSequenceClassification (4.48+). Mizan-Rerank-v3 ships custom code and needs trust_remote_code=True.
  • Sequence length. The model was trained on query + passage inputs of up to 3,072 tokens and is benchmarked here at 2,048. The architecture accepts 8,192 positions, but quality beyond 3,072 tokens has not been measured. If your query + passage inputs fit in 512 tokens, max_length=512 is much faster (it is what the fiqh evaluation and the speed benchmark use).
  • Scores. The sigmoid of the single output logit is a relevance score. It is calibrated to work with coarse cut-offs (production uses sigmoid > 0.1 to keep the top 8), but it is not a probability of correctness: tune any threshold on your own data.
  • Small and fast. fp16 throughput is 2,487 pairs/s at 0.34 GiB, so it can rerank about twice as many candidates as Mizan-Rerank-v3 in the same latency budget.

Evaluation

All five models were scored by the same script with the same inputs: float32, max_length=2048, and the full candidate list of every query. Metrics use graded gains (correct passage 3, partially correct passage 1, trap or irrelevant passage 0).

Metric Meaning
nDCG@10 quality of the whole ranking
MRR@10 1 / rank of the correct passage
Hit@1 share of queries where the correct passage is ranked first

Held-out benchmarks

Benchmark Queries Mizan-Rerank-v3-Turbo Mizan-Rerank-v3 Mizan-Rerank-V2 gte-multilingual-reranker-base bge-reranker-v2-m3
MTEB NamaaMrTydi, unseen subset¹ 572 0.9023 0.8788 0.8224 0.8775 0.8902
Adversarial short (test)² 602 0.9188 0.8432 0.6965 0.7565 0.7938
Adversarial long-context (test)² 1,703 0.8973 0.8280 0.6203 0.6747 0.6873
Multi-LLM benchmark² 361 0.7811 0.7536 0.5998 0.6992 0.7365
Average nDCG@10 0.8749 0.8259 0.6847 0.7520 0.7770
Average MRR@10 0.8308 0.7655 0.5801 0.6766 0.7108
Average Hit@1 0.7132 0.6049 0.3496 0.4679 0.5242

¹ Queries whose question and correct passage never occur in the Turbo training pool (checked by exact query/passage hash match against the full pool). ² Internal test sets, not released. They share no queries or passages with the training data (verified by exact match), and the long-context split is by source article.

The averages cover the four held-out sets above. Arabic Hard Negatives is reported separately and never averaged: 12,363 of its 12,373 queries occur in the Turbo training pool (Mr. TyDi replay and Arabic MS MARCO-style triplets), so only 10 cases are unseen and a score on it would not measure generalisation. Its full-set nDCG@10 is 0.9265 for Turbo, 0.9103 for v3 and 0.9082 for bge-reranker-v2-m3, which is why it is excluded here.

Public benchmarks

Adversarial benchmarks

Where the gain comes from: ranking the correct passage first. Hit@1 is where Turbo differs most from the other models, especially when traps share most of their text with the answer:

Hit@1

Real production queries (fiqh)

224 real fatwa-search queries from the production trace database, each with the production retriever's top-25 candidates and graded judgments (0-3) from two LLM judges (weighted kappa 0.81). Scoring matches production: fp16, max_length=512, sigmoid, keep the top 8 above 0.1.

Model nDCG@5 nDCG@10 MRR@10 p@1 relevant_kept@8 misleading@3
Mizan-Rerank-v3-Turbo 0.7174 0.7612 0.6751 0.7366 0.7494 0.2723
bge-reranker-v2-m3 0.6831 0.7372 0.6601 0.7098 0.7401 0.2946
Mizan-Rerank-v3 0.6151 0.6693 0.5807 0.6429 0.6764 0.3080
gte-multilingual-reranker-base 0.5932 0.6567 0.5346 0.5893 0.6661 0.3259
Mizan-Rerank-V2 0.5471 0.6176 — — — —
Retriever only (no reranker) 0.5067 0.5732 0.4691 0.5045 0.5740 0.3482

misleading@3 counts queries with a ruling that answers a different case in the top 3 (lower is better). relevant_kept@8 is the share of relevant rulings that survive production's top-8 + sigmoid > 0.1 filter.

Contradiction ranked first

For RAG the worst failure is not a missing answer but a passage that says the opposite ranked at the top: a negated ruling, a reversed condition, a wrong number. Every adversarial test query has one correct passage and several contradicting traps. The table shows how often a trap was ranked #1 (lower is better):

Test set Queries Mizan-Rerank-v3-Turbo Mizan-Rerank-v3 Mizan-Rerank-V2 gte-multilingual-reranker-base bge-reranker-v2-m3
Adversarial short 602 11.0% (66) 20.1% (121) 62.1% (374) 51.3% (309) 43.5% (262)
Adversarial long-context 1,703 14.1% (240) 33.4% (568) 76.4% (1,301) 68.2% (1,161) 62.0% (1,056)
Multi-LLM 361 39.9% (144) 49.3% (178) 75.3% (272) 64.0% (231) 54.0% (195)

Across all 2,666 queries, Turbo ranks a trap first 16.9% of the time. For Mizan-Rerank-v3 the figure is 32.5%; for bge-reranker-v2-m3, 56.8%; for gte base, 63.8%; for Mizan-Rerank-V2, 73.0%:

Contradiction ranked first

Reproduce with python compare_generated_test_models.py --dataset-file <test.jsonl> ... --chart contradiction_first.png (in the training repository; the test sets are internal).

Full per-benchmark tables (nDCG@10 / MRR@10 / Hit@1), including full-set public scores

nDCG@10

Model NamaaMrTydi (full, 918) NamaaMrTydi (unseen, 572) Arabic Hard Neg. (full, 12,373)³ Arabic Hard Neg. (unseen, 10) Adversarial short Adversarial long-ctx Multi-LLM
Mizan-Rerank-v3-Turbo 0.9116 0.9023 0.9265 0.7210 0.9188 0.8973 0.7811
Mizan-Rerank-v3 0.8885 0.8788 0.9103 0.7454 0.8432 0.8280 0.7536
Mizan-Rerank-V2 0.8312 0.8224 0.8447 0.7085 0.6965 0.6203 0.5998
gte-multilingual-reranker-base 0.8874 0.8775 0.8960 0.7385 0.7565 0.6747 0.6992
bge-reranker-v2-m3 0.8995 0.8902 0.9082 0.7948 0.7938 0.6873 0.7365

MRR@10

Model NamaaMrTydi (full) NamaaMrTydi (unseen) Arabic Hard Neg. (full)³ Arabic Hard Neg. (unseen) Adversarial short Adversarial long-ctx Multi-LLM
Mizan-Rerank-v3-Turbo 0.8824 0.8701 0.9012 0.6283 0.8843 0.8552 0.7137
Mizan-Rerank-v3 0.8518 0.8391 0.8796 0.6583 0.7731 0.7658 0.6841
Mizan-Rerank-V2 0.7755 0.7640 0.7919 0.6083 0.5952 0.4983 0.4629
gte-multilingual-reranker-base 0.8498 0.8367 0.8607 0.6500 0.6826 0.5765 0.6105
bge-reranker-v2-m3 0.8662 0.8537 0.8771 0.7283 0.7374 0.5879 0.6643

Hit@1

Model NamaaMrTydi (full) NamaaMrTydi (unseen) Arabic Hard Neg. (full)³ Arabic Hard Neg. (unseen) Adversarial short Adversarial long-ctx Multi-LLM
Mizan-Rerank-v3-Turbo 0.8159 0.7972 0.8252 0.4000 0.7940 0.7352 0.5263
Mizan-Rerank-v3 0.7691 0.7535 0.7882 0.4000 0.6146 0.5696 0.4820
Mizan-Rerank-V2 0.6492 0.6346 0.6433 0.3000 0.3405 0.2267 0.1967
gte-multilingual-reranker-base 0.7549 0.7360 0.7600 0.4000 0.4618 0.3136 0.3601
bge-reranker-v2-m3 0.7865 0.7640 0.7916 0.6000 0.5449 0.3365 0.4515

³ 12,363 of these queries occur in the Turbo training pool (and 97% in v3's), so the full-set score is not a fair measure for either Mizan model and is not bolded or used in any average. The unseen subset is only 10 cases.

Efficiency

Every model scored the same 2,048 Arabic query–passage pairs from MTEB NamaaMrTydi. The run used one NVIDIA RTX 4090, float16, batch size 32 and max_length=512, after a warm-up:

Model Parameters Throughput (pairs/s) Peak GPU memory Held-out mean nDCG@10
Mizan-Rerank-v3-Turbo 150M 2,487 0.34 GiB 0.875
Mizan-Rerank-v3 306M 1,238 1.10 GiB 0.826
Mizan-Rerank-V2 306M 1,191 1.10 GiB 0.685
gte-multilingual-reranker-base 306M 1,192 1.10 GiB 0.752
bge-reranker-v2-m3 568M 600 1.43 GiB 0.777

Accuracy vs. speed

At the same latency budget, Turbo can rerank twice as many candidates as Mizan-Rerank-v3 and four times as many as bge-reranker-v2-m3, at the highest held-out accuracy of the five.

Training curve

Framework versions

Python 3.10.14 · PyTorch 2.8.0+cu126 · Transformers 4.55.4 · Sentence Transformers 5.4.1 · Accelerate 1.10.0 · Tokenizers 0.21.0

Citation

@software{Mizan_Rerank_v3_Turbo_2026,
  author    = {Ali Aljiachi},
  title     = {Mizan-Rerank-v3-Turbo: Compact Arabic Long-Context Reranker},
  year      = {2026},
  publisher = {Hugging Face},
  url       = {https://huggingface.co/ALJIACHI/Mizan-Rerank-v3-Turbo}
}

License

Apache 2.0.

Downloads last month
57
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ALJIACHI/Mizan-Rerank-v3-Turbo

Finetuned
(1)
this model

Space using ALJIACHI/Mizan-Rerank-v3-Turbo 1

Collections including ALJIACHI/Mizan-Rerank-v3-Turbo