mentee-embed-v3 / README.md
SyedSyab's picture
Fix: correct Protocol A avgs, usage code, DOI, remove val column
60da417 verified
|
Raw
History Blame Contribute Delete
8.07 kB
---
license: apache-2.0
pipeline_tag: sentence-similarity
library_name: transformers
language:
- en
- ar
- ur
tags:
- embeddings
- sentence-similarity
- retrieval
- contrastive-learning
- multilingual
- from-scratch
datasets:
- sentence-transformers/all-NLI
- facebook/xnli
- Helsinki-NLP/opus-100
- sentence-transformers/msmarco-bm25
- nlpai-lab/miracl-multilingual-triplets
metrics:
- mrr@10
- spearman_correlation
---
# mentee-embed-v3
**A 41M-parameter trilingual text embedding model trained entirely from scratch — no pretrained backbone, no BERT, no RoBERTa. Random initialization only.**
Developed by **Team MenteE AI** (menteeai.org) as part of an ongoing research effort to build competitive multilingual embeddings from the ground up for Arabic, English, and Urdu.
---
## Key Facts
| Property | Value |
|----------|-------|
| Parameters | **41M** |
| Embedding dimension | 384 |
| Max sequence length | 128 tokens |
| Languages | Arabic 🇸🇦 · English 🇬🇧 · Urdu 🇵🇰 |
| Architecture | 12-layer Transformer, custom BPE tokenizer (50K vocab) |
| Initialization | **Random** (trained from scratch) |
| Training data | ~2.1M triplets (NLI + MS-MARCO + OPUS parallel + MIRACL) |
| Pooling | Mean pooling |
| Training objective | Relational distillation + InfoNCE contrastive (teacher: `multilingual-e5-base`) |
---
## What "From Scratch" Means
Most embedding models fine-tune an existing pretrained encoder (BERT, RoBERTa, MPNet). **mentee-embed-v3 does not.** We:
1. Trained a **custom BPE tokenizer** on Arabic, English, and Urdu text
2. Initialized a **12-layer Transformer with random weights**
3. Ran **masked language modeling** pretraining on 2.1M+ sentences
4. Applied **two-round contrastive distillation** with hard negative mining
No pretrained checkpoint was used at any stage.
---
## Benchmark Results
All baselines evaluated under identical conditions on the same hardware.
### Protocol A — In-batch Retrieval (pool ≈ 97 candidates)
*Format: acc@1 / R@5 / MRR@10*
*Avg MRR@10 computed over 5 datasets: MIRACL-EN, MIRACL-AR, MIRACL-UR, xling EN-UR, MS-MARCO (val excluded from avg).*
| Model | MIRACL-EN | MIRACL-AR | MIRACL-UR | xling EN↔UR | MS-MARCO | **Avg MRR@10** |
|-------|-----------|-----------|-----------|-------------|----------|----------------|
| **mentee-embed-v3 (ours)** | 0.636/0.920/0.766 | 0.326/0.604/0.475 | 0.290/0.568/0.443 | 0.781/0.925/0.848 | 0.517/0.981/0.742 | **0.655** |
| paraphrase-multilingual-mpnet-base-v2 | 0.864/1.000/0.931 | 0.722/0.975/0.839 | 0.686/0.950/0.806 | 0.831/0.937/0.880 | 0.665/0.998/0.830 | **0.857** |
| paraphrase-multilingual-MiniLM-L12-v2 | 0.854/0.997/0.924 | 0.696/0.964/0.819 | 0.621/0.908/0.753 | 0.782/0.907/0.841 | 0.600/0.998/0.796 | **0.827** |
| all-MiniLM-L6-v2 | 0.856/0.999/0.927 | 0.025/0.109/0.144 | 0.028/0.088/0.140 | 0.065/0.172/0.186 | 0.696/1.000/0.848 | **0.449** |
> ✅ **mentee-embed-v3 beats all-MiniLM-L6-v2** (0.655 vs 0.449) on Protocol A avg MRR@10 — despite all-MiniLM being a pretrained model.
---
### Protocol B — MIRACL Wikipedia Corpus Retrieval (full ranking, ~5K–15K passages)
*Format: MRR@10 · R@5 · R@100*
| Model | EN | AR | UR | **Avg MRR@10** |
|-------|----|----|----|----|
| paraphrase-multilingual-mpnet-base-v2 | 0.853 · 0.923 · 0.997 | 0.622 · 0.757 · 0.947 | 0.534 · 0.680 · 0.903 | **0.670** |
| paraphrase-multilingual-MiniLM-L12-v2 | 0.840 · 0.933 · 0.990 | 0.591 · 0.710 · 0.943 | 0.469 · 0.557 · 0.847 | **0.633** |
| all-MiniLM-L6-v2 | 0.867 · 0.967 · 1.000 | 0.100 · 0.000 · 0.010 | 0.106 · 0.007 · 0.013 | **0.358** |
| **mentee-embed-v3 (ours)** | 0.418 · 0.473 · 0.850 | 0.182 · 0.167 · 0.500 | 0.180 · 0.143 · 0.440 | **0.260** |
> 📌 Protocol B uses Wikipedia passages (MIRACL) — a challenging out-of-domain test for a model trained primarily on NLI and MS-MARCO data.
---
### Protocol C — MS-MARCO Corpus Retrieval (10K passages, in-domain)
| Model | MRR@10 | R@5 | R@100 |
|-------|--------|-----|-------|
| all-MiniLM-L6-v2 | 0.951 | 0.993 | 1.000 |
| paraphrase-multilingual-mpnet-base-v2 | 0.882 | 0.970 | 0.993 |
| paraphrase-multilingual-MiniLM-L12-v2 | 0.839 | 0.910 | 0.997 |
| **mentee-embed-v3 (ours)** | **0.645** | **0.760** | **0.957** |
> 🔥 **0.645 MRR@10 from a randomly initialized 41M model** — trained on 2.1M triplets vs billions for the baselines.
> Gap to MiniLM-multilingual: only **0.194**.
---
## Training Data
| Source | Language | Triplets | Type |
|--------|----------|----------|------|
| all-NLI (sentence-transformers) | EN | 558K | NLI triplets |
| XNLI | AR | 128K | NLI triplets |
| XNLI | UR | 125K | NLI triplets |
| OPUS-100 EN↔UR | EN/UR | 300K | Parallel translation |
| OPUS-100 AR↔EN | AR/EN | 300K | Parallel translation |
| MS-MARCO BM25 triplets | EN | 500K | Passage retrieval |
| MS-MARCO hard negatives | EN | 200K | Hard retrieval |
| MIRACL | EN/AR/UR | ~9K | Wikipedia retrieval |
| **Total** | | **~2.1M** | |
---
## Training Pipeline
```
Stage 1 — MLM Pretraining
Random init → masked language modeling on 2.1M sentences
8,000 steps · batch=32 · vocab=50K BPE
Stage 2 — Distillation Round 1 (no hard negatives)
Teacher: intfloat/multilingual-e5-base (768-dim)
InfoNCE contrastive + relational distillation
4,000 steps · batch=512 · temp=0.05
Stage 3 — Hard Negative Mining
GPU-accelerated top-5 mining across full 2.1M corpus
Stage 4 — Distillation Round 2 (with hard negatives)
Same objective + mined hard negatives per anchor
10,000 steps · batch=512 · temp=0.05
```
---
## Usage
```python
# pip install torch transformers tokenizers huggingface_hub
from transformers import AutoModel, AutoTokenizer
tok = AutoTokenizer.from_pretrained("MenteEAI/mentee-embed-v3", trust_remote_code=True)
model = AutoModel.from_pretrained("MenteEAI/mentee-embed-v3", trust_remote_code=True)
sentences = [
"Hello, how are you?",
"مرحبا، كيف حالك؟",
"ہیلو، آپ کیسے ہیں؟"
]
embeddings = model.encode(sentences, tokenizer=tok)
print(embeddings.shape) # torch.Size([3, 384])
```
> `trust_remote_code=True` is required — standard for custom-architecture models on HuggingFace. The code runs entirely on your machine.
### Similarity search
```python
from transformers import AutoModel, AutoTokenizer
tok = AutoTokenizer.from_pretrained("MenteEAI/mentee-embed-v3", trust_remote_code=True)
model = AutoModel.from_pretrained("MenteEAI/mentee-embed-v3", trust_remote_code=True)
query = model.encode(["What is machine learning?"], tokenizer=tok)
passages = model.encode([
"Machine learning is a subset of artificial intelligence.",
"The weather today is sunny.",
"تعلم الآلة هو فرع من فروع الذكاء الاصطناعي.",
], tokenizer=tok)
scores = query @ passages.T
print(scores) # tensor([[0.81, 0.60, 0.79]])
```
---
## Limitations
- **Protocol B (Wikipedia retrieval)** scores are lower than pretrained baselines — the model was not trained on Wikipedia-style passages
- **Arabic and Urdu** lag behind English due to less retrieval-specific training data (mMARCO Arabic/Urdu was unavailable in a compatible format)
- Vocabulary limited to 50K tokens trained on ~2.1M sentences — rare scripts and dialects may tokenize poorly
- Not evaluated on MTEB full suite yet
---
## Citation
```bibtex
@misc{mentee-embed-v3-2026,
title = {How Far Can Multilingual Text Embeddings Be Trained From Scratch?
A Compute-Efficient Study of Arabic, English, and Urdu},
author = {Shah, Syed Syab Ahmad and Sania, Shakeel and Hamza, Rustam and Mahboob, Iqbal},
year = {2026},
doi = {10.5281/zenodo.22117673},
url = {https://doi.org/10.5281/zenodo.22117673},
note = {MenteE AI. Apache-2.0 License}
}
```
---
## About MenteE AI
Built by **Syed Syab Ahmad Shah** and Team MenteE AI.
🌐 [menteeai.org](https://menteeai.org) · 📧 syab@menteeai.org
Research paper: [10.5281/zenodo.22117673](https://doi.org/10.5281/zenodo.22117673)