Instructions to use ALJIACHI/Mizan-Rerank-v3-Turbo with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- sentence-transformers
How to use ALJIACHI/Mizan-Rerank-v3-Turbo with sentence-transformers:
from sentence_transformers import CrossEncoder model = CrossEncoder("ALJIACHI/Mizan-Rerank-v3-Turbo") query = "Which planet is known as the Red Planet?" passages = [ "Venus is often called Earth's twin because of its similar size and proximity.", "Mars, known for its reddish appearance, is often referred to as the Red Planet.", "Jupiter, the largest planet in our solar system, has a prominent red spot.", "Saturn, famous for its rings, is sometimes mistaken for the Red Planet." ] scores = model.predict([(query, passage) for passage in passages]) print(scores) - Notebooks
- Google Colab
- Kaggle
Mizan-Rerank-v3-Turbo
The accuracy of Mizan-Rerank-v3 at half the size and twice the speed: a compact Arabic reranker that tells the passage that answers the question apart from passages that only look like they do.
Try it: live demo — rerank your own passages, run the eleven adversarial trap cases, and switch between this model and Mizan-Rerank-v3 to compare.
The Mizan reranker family
Model Parameters Throughput Held-out mean nDCG@10 Fiqh nDCG@10 Mizan-Rerank-v3-Turbo (this model) 150M 2,487 pairs/s 0.875 0.761 Mizan-Rerank-v3 306M 1,238 pairs/s 0.826 0.669 Pick Turbo for the best accuracy per millisecond, v3 if you prefer the larger model with the same behaviour as the previous release.
Held-out mean nDCG@10 is the mean over the four held-out sets both cards share (NamaaMrTydi unseen, short adversarial, long-context adversarial, Multi-LLM). Mizan-Rerank-v3's own card quotes 0.845, which averages a fifth set (Arabic Hard Negatives unseen) that cannot be used here: only 10 of its 12,373 cases are unseen for Turbo.
Mizan-Rerank-v3-Turbo is a 150M-parameter Arabic cross-encoder. It starts from the Mizan-Embeddings-0.1B encoder and is trained in three stages: broad Arabic relevance with hard negatives and teacher scores, in-domain distillation on production retrieval candidates, then v3's adversarial and long-context recipe. It learns to separate the correct passage from adversarial "trap" passages that share most of its words but change its meaning: a negated ruling, a swapped entity, a shifted number or date, an exception applied to the wrong case.
The result is the most accurate model in the family while being the smallest: half the parameters of Mizan-Rerank-v3, twice its throughput, 3.2× less peak GPU memory, and higher scores on every held-out benchmark we measure, including the real production fiqh queries.
Highlights
- Best accuracy in the family. Held-out mean nDCG@10 is 0.875 at 150M parameters: Mizan-Rerank-v3 scores 0.826, bge-reranker-v2-m3 (568M) 0.777, gte-multilingual-reranker-base 0.752 and Mizan-Rerank-V2 0.685.
- Fast and small. On one RTX 4090 in fp16 it scores 2,487 pairs/s (2.0× v3, 4.1× bge-reranker-v2-m3) with 0.34 GiB peak GPU memory (see Efficiency).
- Best on real production queries. On 224 real fiqh search queries with the production retriever's candidate lists, nDCG@10 is 0.761 against 0.669 for v3 and 0.737 for bge-reranker-v2-m3 (+0.092 over v3, 95% CI [+0.069, +0.115]). It also keeps more of the relevant rulings in production's top-8 filter (0.749).
- Rarely puts a contradiction on top. Across 2,666 held-out adversarial queries, the top-ranked passage is a contradicting trap 16.9% of the time. For Mizan-Rerank-v3 the figure is 32.5%, for bge-reranker-v2-m3 56.8% and for Mizan-Rerank-V2 73.0%.
- Calibrated scores. The sigmoid score works directly with a fixed cut-off: the scores were affine-calibrated after distillation (see Practical notes).
Usage
Sentence Transformers
pip install -U sentence-transformers
from sentence_transformers import CrossEncoder
model = CrossEncoder("ALJIACHI/Mizan-Rerank-v3-Turbo", max_length=3072)
query = "ما هي فوائد فيتامين د؟"
passages = [
"يساعد فيتامين د على امتصاص الكالسيوم وتقوية العظام، كما يدعم عمل الجهاز المناعي.",
"يستخدم فيتامين د في بعض الصناعات الغذائية كمادة مضافة لتدعيم الحليب.",
"أطلقت وزارة الزراعة حملة وطنية لزيادة الوعي بأهمية الزراعة العضوية.",
]
scores = model.predict([(query, passage) for passage in passages])
print(scores) # ≈ [0.98, 0.60, 0.01] (sigmoid relevance scores; the second passage is on-topic but does not answer
# the question, so it keeps a middling score while the answer is ranked first)
for hit in model.rank(query, passages, return_documents=True):
print(f"{hit['score']:.3f} {hit['text']}")
Transformers
import torch
from transformers import AutoModelForSequenceClassification, AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("ALJIACHI/Mizan-Rerank-v3-Turbo")
model = AutoModelForSequenceClassification.from_pretrained(
"ALJIACHI/Mizan-Rerank-v3-Turbo", torch_dtype=torch.float16
).to("cuda").eval()
query = "ما هي فوائد فيتامين د؟"
passages = [
"يساعد فيتامين د على امتصاص الكالسيوم وتقوية العظام، كما يدعم عمل الجهاز المناعي.",
"أطلقت وزارة الزراعة حملة وطنية لزيادة الوعي بأهمية الزراعة العضوية.",
]
with torch.inference_mode():
features = tokenizer([query] * len(passages), passages, padding=True, truncation=True,
max_length=3072, return_tensors="pt").to("cuda")
scores = torch.sigmoid(model(**features).logits.squeeze(-1).float())
for passage, score in sorted(zip(passages, scores.tolist()), key=lambda item: -item[1]):
print(f"{score:.3f} {passage}")
Practical notes
- No
trust_remote_codeneeded. The architecture is plain ModernBERT, so the model loads with any Transformers version that supportsModernBertForSequenceClassification(4.48+). Mizan-Rerank-v3 ships custom code and needstrust_remote_code=True. - Sequence length. The model was trained on query + passage inputs of up to 3,072 tokens and is benchmarked
here at 2,048. The architecture accepts 8,192 positions, but quality beyond 3,072 tokens has not been measured.
If your query + passage inputs fit in 512 tokens,
max_length=512is much faster (it is what the fiqh evaluation and the speed benchmark use). - Scores. The sigmoid of the single output logit is a relevance score. It is calibrated to work with coarse
cut-offs (production uses
sigmoid > 0.1to keep the top 8), but it is not a probability of correctness: tune any threshold on your own data. - Small and fast. fp16 throughput is 2,487 pairs/s at 0.34 GiB, so it can rerank about twice as many candidates as Mizan-Rerank-v3 in the same latency budget.
Evaluation
All five models were scored by the same script with the same inputs: float32, max_length=2048, and the full
candidate list of every query. Metrics use graded gains (correct passage 3, partially correct passage 1, trap or
irrelevant passage 0).
| Metric | Meaning |
|---|---|
| nDCG@10 | quality of the whole ranking |
| MRR@10 | 1 / rank of the correct passage |
| Hit@1 | share of queries where the correct passage is ranked first |
Held-out benchmarks
| Benchmark | Queries | Mizan-Rerank-v3-Turbo | Mizan-Rerank-v3 | Mizan-Rerank-V2 | gte-multilingual-reranker-base | bge-reranker-v2-m3 |
|---|---|---|---|---|---|---|
| MTEB NamaaMrTydi, unseen subset¹ | 572 | 0.9023 | 0.8788 | 0.8224 | 0.8775 | 0.8902 |
| Adversarial short (test)² | 602 | 0.9188 | 0.8432 | 0.6965 | 0.7565 | 0.7938 |
| Adversarial long-context (test)² | 1,703 | 0.8973 | 0.8280 | 0.6203 | 0.6747 | 0.6873 |
| Multi-LLM benchmark² | 361 | 0.7811 | 0.7536 | 0.5998 | 0.6992 | 0.7365 |
| Average nDCG@10 | 0.8749 | 0.8259 | 0.6847 | 0.7520 | 0.7770 | |
| Average MRR@10 | 0.8308 | 0.7655 | 0.5801 | 0.6766 | 0.7108 | |
| Average Hit@1 | 0.7132 | 0.6049 | 0.3496 | 0.4679 | 0.5242 |
¹ Queries whose question and correct passage never occur in the Turbo training pool (checked by exact query/passage hash match against the full pool). ² Internal test sets, not released. They share no queries or passages with the training data (verified by exact match), and the long-context split is by source article.
The averages cover the four held-out sets above. Arabic Hard Negatives is reported separately and never averaged: 12,363 of its 12,373 queries occur in the Turbo training pool (Mr. TyDi replay and Arabic MS MARCO-style triplets), so only 10 cases are unseen and a score on it would not measure generalisation. Its full-set nDCG@10 is 0.9265 for Turbo, 0.9103 for v3 and 0.9082 for bge-reranker-v2-m3, which is why it is excluded here.
Where the gain comes from: ranking the correct passage first. Hit@1 is where Turbo differs most from the other models, especially when traps share most of their text with the answer:
Real production queries (fiqh)
224 real fatwa-search queries from the production trace database, each with the production retriever's top-25
candidates and graded judgments (0-3) from two LLM judges (weighted kappa 0.81). Scoring matches production:
fp16, max_length=512, sigmoid, keep the top 8 above 0.1.
| Model | nDCG@5 | nDCG@10 | MRR@10 | p@1 | relevant_kept@8 | misleading@3 |
|---|---|---|---|---|---|---|
| Mizan-Rerank-v3-Turbo | 0.7174 | 0.7612 | 0.6751 | 0.7366 | 0.7494 | 0.2723 |
| bge-reranker-v2-m3 | 0.6831 | 0.7372 | 0.6601 | 0.7098 | 0.7401 | 0.2946 |
| Mizan-Rerank-v3 | 0.6151 | 0.6693 | 0.5807 | 0.6429 | 0.6764 | 0.3080 |
| gte-multilingual-reranker-base | 0.5932 | 0.6567 | 0.5346 | 0.5893 | 0.6661 | 0.3259 |
| Mizan-Rerank-V2 | 0.5471 | 0.6176 | — | — | — | — |
| Retriever only (no reranker) | 0.5067 | 0.5732 | 0.4691 | 0.5045 | 0.5740 | 0.3482 |
misleading@3 counts queries with a ruling that answers a different case in the top 3 (lower is better).
relevant_kept@8 is the share of relevant rulings that survive production's top-8 + sigmoid > 0.1 filter.
Contradiction ranked first
For RAG the worst failure is not a missing answer but a passage that says the opposite ranked at the top: a negated ruling, a reversed condition, a wrong number. Every adversarial test query has one correct passage and several contradicting traps. The table shows how often a trap was ranked #1 (lower is better):
| Test set | Queries | Mizan-Rerank-v3-Turbo | Mizan-Rerank-v3 | Mizan-Rerank-V2 | gte-multilingual-reranker-base | bge-reranker-v2-m3 |
|---|---|---|---|---|---|---|
| Adversarial short | 602 | 11.0% (66) | 20.1% (121) | 62.1% (374) | 51.3% (309) | 43.5% (262) |
| Adversarial long-context | 1,703 | 14.1% (240) | 33.4% (568) | 76.4% (1,301) | 68.2% (1,161) | 62.0% (1,056) |
| Multi-LLM | 361 | 39.9% (144) | 49.3% (178) | 75.3% (272) | 64.0% (231) | 54.0% (195) |
Across all 2,666 queries, Turbo ranks a trap first 16.9% of the time. For Mizan-Rerank-v3 the figure is 32.5%; for bge-reranker-v2-m3, 56.8%; for gte base, 63.8%; for Mizan-Rerank-V2, 73.0%:
Reproduce with python compare_generated_test_models.py --dataset-file <test.jsonl> ... --chart contradiction_first.png
(in the training repository; the test sets are internal).
Full per-benchmark tables (nDCG@10 / MRR@10 / Hit@1), including full-set public scores
nDCG@10
| Model | NamaaMrTydi (full, 918) | NamaaMrTydi (unseen, 572) | Arabic Hard Neg. (full, 12,373)³ | Arabic Hard Neg. (unseen, 10) | Adversarial short | Adversarial long-ctx | Multi-LLM |
|---|---|---|---|---|---|---|---|
| Mizan-Rerank-v3-Turbo | 0.9116 | 0.9023 | 0.9265 | 0.7210 | 0.9188 | 0.8973 | 0.7811 |
| Mizan-Rerank-v3 | 0.8885 | 0.8788 | 0.9103 | 0.7454 | 0.8432 | 0.8280 | 0.7536 |
| Mizan-Rerank-V2 | 0.8312 | 0.8224 | 0.8447 | 0.7085 | 0.6965 | 0.6203 | 0.5998 |
| gte-multilingual-reranker-base | 0.8874 | 0.8775 | 0.8960 | 0.7385 | 0.7565 | 0.6747 | 0.6992 |
| bge-reranker-v2-m3 | 0.8995 | 0.8902 | 0.9082 | 0.7948 | 0.7938 | 0.6873 | 0.7365 |
MRR@10
| Model | NamaaMrTydi (full) | NamaaMrTydi (unseen) | Arabic Hard Neg. (full)³ | Arabic Hard Neg. (unseen) | Adversarial short | Adversarial long-ctx | Multi-LLM |
|---|---|---|---|---|---|---|---|
| Mizan-Rerank-v3-Turbo | 0.8824 | 0.8701 | 0.9012 | 0.6283 | 0.8843 | 0.8552 | 0.7137 |
| Mizan-Rerank-v3 | 0.8518 | 0.8391 | 0.8796 | 0.6583 | 0.7731 | 0.7658 | 0.6841 |
| Mizan-Rerank-V2 | 0.7755 | 0.7640 | 0.7919 | 0.6083 | 0.5952 | 0.4983 | 0.4629 |
| gte-multilingual-reranker-base | 0.8498 | 0.8367 | 0.8607 | 0.6500 | 0.6826 | 0.5765 | 0.6105 |
| bge-reranker-v2-m3 | 0.8662 | 0.8537 | 0.8771 | 0.7283 | 0.7374 | 0.5879 | 0.6643 |
Hit@1
| Model | NamaaMrTydi (full) | NamaaMrTydi (unseen) | Arabic Hard Neg. (full)³ | Arabic Hard Neg. (unseen) | Adversarial short | Adversarial long-ctx | Multi-LLM |
|---|---|---|---|---|---|---|---|
| Mizan-Rerank-v3-Turbo | 0.8159 | 0.7972 | 0.8252 | 0.4000 | 0.7940 | 0.7352 | 0.5263 |
| Mizan-Rerank-v3 | 0.7691 | 0.7535 | 0.7882 | 0.4000 | 0.6146 | 0.5696 | 0.4820 |
| Mizan-Rerank-V2 | 0.6492 | 0.6346 | 0.6433 | 0.3000 | 0.3405 | 0.2267 | 0.1967 |
| gte-multilingual-reranker-base | 0.7549 | 0.7360 | 0.7600 | 0.4000 | 0.4618 | 0.3136 | 0.3601 |
| bge-reranker-v2-m3 | 0.7865 | 0.7640 | 0.7916 | 0.6000 | 0.5449 | 0.3365 | 0.4515 |
³ 12,363 of these queries occur in the Turbo training pool (and 97% in v3's), so the full-set score is not a fair measure for either Mizan model and is not bolded or used in any average. The unseen subset is only 10 cases.
Efficiency
Every model scored the same 2,048 Arabic query–passage pairs from MTEB NamaaMrTydi. The run used one NVIDIA RTX
4090, float16, batch size 32 and max_length=512, after a warm-up:
| Model | Parameters | Throughput (pairs/s) | Peak GPU memory | Held-out mean nDCG@10 |
|---|---|---|---|---|
| Mizan-Rerank-v3-Turbo | 150M | 2,487 | 0.34 GiB | 0.875 |
| Mizan-Rerank-v3 | 306M | 1,238 | 1.10 GiB | 0.826 |
| Mizan-Rerank-V2 | 306M | 1,191 | 1.10 GiB | 0.685 |
| gte-multilingual-reranker-base | 306M | 1,192 | 1.10 GiB | 0.752 |
| bge-reranker-v2-m3 | 568M | 600 | 1.43 GiB | 0.777 |
At the same latency budget, Turbo can rerank twice as many candidates as Mizan-Rerank-v3 and four times as many as bge-reranker-v2-m3, at the highest held-out accuracy of the five.
Framework versions
Python 3.10.14 · PyTorch 2.8.0+cu126 · Transformers 4.55.4 · Sentence Transformers 5.4.1 · Accelerate 1.10.0 · Tokenizers 0.21.0
Citation
@software{Mizan_Rerank_v3_Turbo_2026,
author = {Ali Aljiachi},
title = {Mizan-Rerank-v3-Turbo: Compact Arabic Long-Context Reranker},
year = {2026},
publisher = {Hugging Face},
url = {https://huggingface.co/ALJIACHI/Mizan-Rerank-v3-Turbo}
}
License
Apache 2.0.
- Downloads last month
- 57
Model tree for ALJIACHI/Mizan-Rerank-v3-Turbo
Base model
ALJIACHI/Mizan-Embeddings-0.1B





