Sentence Similarity
sentence-transformers
Safetensors
Turkish
xlm-roberta
feature-extraction
embeddings
turkish
türkçe
e5
retrieval
semantic-search
mteb
tr-mteb
Eval Results (legacy)
text-embeddings-inference
Instructions to use thealper2/intfloat-multilingual-e5-base-tr-nli with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- sentence-transformers
How to use thealper2/intfloat-multilingual-e5-base-tr-nli with sentence-transformers:
from sentence_transformers import SentenceTransformer model = SentenceTransformer("thealper2/intfloat-multilingual-e5-base-tr-nli") sentences = [ "The weather is lovely today.", "It's so sunny outside!", "He drove to the stadium." ] embeddings = model.encode(sentences) similarities = model.similarity(embeddings, embeddings) print(similarities.shape) # [3, 3] - Notebooks
- Google Colab
- Kaggle
File size: 14,129 Bytes
cf90e06 b525f45 cf90e06 b525f45 cf90e06 b525f45 cf90e06 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 390 | ---
language:
- tr
license: cc-by-sa-4.0
library_name: sentence-transformers
pipeline_tag: sentence-similarity
tags:
- sentence-transformers
- sentence-similarity
- feature-extraction
- embeddings
- turkish
- türkçe
- e5
- retrieval
- semantic-search
- mteb
- tr-mteb
base_model: intfloat/multilingual-e5-base
datasets:
- mertcobanov/all-nli-triplets-turkish
- emrecan/stsb-mt-turkish
metrics:
- spearmanr
- ndcg
- accuracy
- v_measure
model-index:
- name: e5-tr-nli
results:
- task:
type: semantic-textual-similarity
name: Semantic Textual Similarity
dataset:
name: STSb-TR (test)
type: trmteb/stsb-tr
metrics:
- type: spearman_cosine
value: 0.7907
- task:
type: retrieval
name: Retrieval
dataset:
name: TQuad
type: trmteb/tquad
metrics:
- type: ndcg_at_10
value: 0.8222
- task:
type: retrieval
name: Retrieval
dataset:
name: Quora-TR
type: trmteb/quora-tr
metrics:
- type: ndcg_at_10
value: 0.7323
- task:
type: retrieval
name: Retrieval
dataset:
name: MS MARCO-TR
type: trmteb/msmarco-tr
metrics:
- type: ndcg_at_10
value: 0.2734
- task:
type: classification
name: Classification
dataset:
name: News-Cat
type: trmteb/news-cat
metrics:
- type: accuracy
value: 0.9600
---
# e5-tr-nli — Turkish Sentence Embedding Model (Bi-Encoder)
A **Turkish sentence-embedding (bi-encoder) model** for **retrieval and semantic
search**, fine-tuned from [`intfloat/multilingual-e5-base`](https://huggingface.co/intfloat/multilingual-e5-base)
with contrastive learning (`MultipleNegativesRankingLoss`) on Turkish NLI triplets.
- **Embedding dimension:** 768
- **Max sequence length:** 256 tokens (trained); backbone supports 512
- **Pooling:** mean pooling
- **Similarity function:** cosine
- **Parameters:** ~278M (XLM-RoBERTa-base backbone)
- **Prefix convention:** e5-style `query:` / `passage:` (see [Usage](#usage) — **required**)
---
## Table of Contents
1. [Intended Use](#intended-use)
2. [Usage](#usage)
3. [Training Data](#training-data)
4. [Training Procedure](#training-procedure)
5. [Evaluation](#evaluation)
6. [TR-MTEB Results](#tr-mteb-results-full)
7. [Limitations & Biases](#limitations--biases)
8. [Compute & Environmental Impact](#compute--environmental-impact)
9. [Licensing](#licensing)
10. [Citation](#citation)
11. [Reproduction](#reproduction)
---
## Intended Use
**In scope**
- Turkish semantic search / passage retrieval (dense retrieval, RAG)
- Semantic textual similarity, paraphrase & duplicate detection
- Clustering and topic grouping of Turkish text
- Feature extraction for downstream Turkish NLP classifiers
**Out of scope**
- Re-ranking with query–document cross-attention (use a cross-encoder instead)
- Long-document embedding beyond 256 tokens without chunking
- High-stakes decisions (legal, medical, hiring) without human oversight
- Non-Turkish text (backbone is multilingual, but this model is tuned for Turkish)
---
## Usage
### ⚠️ e5 prefixes are mandatory
The model was trained with the e5 asymmetric prefixes and **must** be used with them,
consistently at training / evaluation / inference:
| Role | Prefix |
|------|--------|
| Query / anchor / any symmetric-task sentence (STS, classification, clustering) | `query: ` |
| Passage / document / positive & negative candidates | `passage: ` |
Omitting the prefixes degrades quality noticeably.
### With `sentence-transformers`
```python
from sentence_transformers import SentenceTransformer
from sentence_transformers.util import cos_sim
model = SentenceTransformer("thealper2/intfloat-multilingual-e5-base-tr-nli")
# Asymmetric retrieval: query vs. candidate passages
query = "query: Türkiye'nin başkenti neresidir?"
passages = [
"passage: Ankara, Türkiye'nin başkentidir.",
"passage: İstanbul Türkiye'nin en kalabalık şehridir.",
"passage: Muz tropikal bir meyvedir.",
]
q = model.encode(query, normalize_embeddings=True)
p = model.encode(passages, normalize_embeddings=True)
print(cos_sim(q, p)) # highest score -> the Ankara passage
```
### Semantic similarity (symmetric — use `query:` on both sides)
```python
a = model.encode("query: Bugün hava çok güzel.", normalize_embeddings=True)
b = model.encode("query: Hava bugün oldukça güzel.", normalize_embeddings=True)
print(float(cos_sim(a, b))) # ~0.9
```
### With 🤗 Transformers (mean pooling)
```python
import torch, torch.nn.functional as F
from transformers import AutoTokenizer, AutoModel
tok = AutoTokenizer.from_pretrained("thealper2/intfloat-multilingual-e5-base-tr-nli")
mdl = AutoModel.from_pretrained("thealper2/intfloat-multilingual-e5-base-tr-nli")
def encode(texts):
batch = tok(texts, padding=True, truncation=True, max_length=256, return_tensors="pt")
with torch.no_grad():
out = mdl(**batch)
mask = batch["attention_mask"].unsqueeze(-1).float()
emb = (out.last_hidden_state * mask).sum(1) / mask.sum(1) # mean pooling
return F.normalize(emb, p=2, dim=1)
emb = encode(["query: örnek cümle", "passage: örnek pasaj"])
```
---
## Training Data
- **Source:** [`mertcobanov/all-nli-triplets-turkish`](https://huggingface.co/datasets/mertcobanov/all-nli-triplets-turkish)
— a machine-translated Turkish version of the AllNLI (SNLI + MultiNLI) triplet set.
- **Format:** `(anchor, positive, negative)` triplets, where the `negative` acts as a
**hard negative** for the contrastive objective.
- **Column handling:** **Only the Turkish columns** (`anchor_translated`,
`positive_translated`, `negative_translated`) were used and renamed to
`anchor` / `positive` / `negative`. **All English columns were discarded.**
- **Cleaning:** rows with `None` / empty / whitespace-only fields were filtered out.
- **Resulting sizes (after filtering):**
| Split | Triplets |
|-------|---------:|
| train | **277,167** |
| dev | 6,584 |
| test | 6,609 |
---
## Training Procedure
### Objective
- **Loss:** `CachedMultipleNegativesRankingLoss` (MNRL with in-batch + hard negatives;
the cached variant is used to simulate a large effective batch on limited VRAM).
- **Batch sampler:** `NO_DUPLICATES` (required for MNRL to avoid trivial in-batch collisions).
- **In-batch negatives:** every other positive/negative in the batch serves as a negative
for a given anchor, so larger batches yield a stronger contrastive signal.
### Hyperparameter search
An **Optuna** search (TPE sampler) was run **before** the final training:
| Setting | Value |
|--------|-------|
| Trials | 10 |
| Subset | 40,000 random training triplets |
| Steps / trial | 600 |
| Search space | `lr ∈ {1e-5, 2e-5, 3e-5}`, `batch ∈ {32, 64, 128}`, `warmup_ratio ∈ {0.0, 0.1}` |
| Selection metric | `stsb-tr` **dev** `spearman_cosine` |
| **Best config** | **lr = 2e-5, batch = 32, warmup_ratio = 0.1** (dev spearman = **0.8285**) |
### Final training
| Setting | Value |
|--------|-------|
| Base model | `intfloat/multilingual-e5-base` |
| Epochs | 1 |
| Learning rate | 2e-5 |
| Batch size | 32 (cached MNRL) |
| Warmup ratio | 0.1 |
| Max sequence length | 256 |
| Precision | bf16 |
| Seed | 42 |
| Eval strategy | steps, monitoring `eval_stsb-tr-dev_spearman_cosine` |
| Checkpoint selection | `load_best_model_at_end=True` (best dev spearman) |
### Frameworks & hardware
- `sentence-transformers` 5.2.2, `transformers` 5.0.0, `torch` 2.11.0+cu128,
`datasets` 4.4.1, Python 3.12
- 1× NVIDIA GeForce RTX 5050 Laptop GPU (8 GB), CUDA 12.8
---
## Evaluation
### Baseline (base model) vs. Final (this model)
Measured on the dataset's own test triplets (`TripletEvaluator`, cosine accuracy) and
on the external Turkish STS set [`emrecan/stsb-mt-turkish`](https://huggingface.co/datasets/emrecan/stsb-mt-turkish)
(`EmbeddingSimilarityEvaluator`, Spearman; scores normalised 0–5 → 0–1):
| Metric | Baseline (e5-base) | Final | Δ |
|--------|:---:|:---:|:---:|
| stsb-tr **dev** spearman_cosine | 0.8092 | **0.8312** | +0.0220 |
| stsb-tr **test** spearman_cosine | 0.7761 | **0.7814** | +0.0053 |
| nli-tr **test** cosine_accuracy | 0.8968 | **0.9260** | +0.0292 |
### Methodology notes for TR-MTEB
- Datasets pulled **directly from HuggingFace** (`trmteb/*`); metrics computed with a
standalone script (no `mteb` package dependency), following standard MTEB protocols:
- **STS** → Spearman of cosine vs. gold; **Retrieval** → nDCG@10 / Recall@10 / MAP@10
via exact cosine search; **Classification** → logistic regression on frozen
embeddings (accuracy + macro-F1); **PairClassification** → average precision of cosine;
**Clustering** → V-measure (MiniBatchKMeans, k = #labels); **BitextMining** → top-1
nearest-neighbour accuracy.
- e5 prefixes applied throughout (`query:` for queries/symmetric sides, `passage:` for corpus).
- Retrieval search runs on-GPU in fp16.
- `75haber`, `thy_sa`, `irony-tr` have **no train split** in the `trmteb` org, so a seeded
**70/30 stratified self-split** was used for those classification tasks (marked `note` in
`trmteb_results.json`). All other classification sets use their official train/test.
### Category summary (macro averages)
| Task type | # datasets | Avg. main metric |
|-----------|:---:|:---:|
| STS (spearman_cosine) | 1 | **0.791** |
| BitextMining (accuracy) | 1 | **0.975** |
| Classification (accuracy) | 6 | **0.817** |
| PairClassification (AP) | 3 | **0.573** |
| Clustering (v-measure) | 2 | **0.487** |
| Retrieval (nDCG@10) | 10 | **0.407** |
| **Overall (mean of category means)** | 23 | **0.675** |
---
## TR-MTEB Results (full)
Evaluated on the [TR-MTEB](https://huggingface.co/trmteb) datasets
(Baysan & Güngör, *TR-MTEB*, Findings of EMNLP 2025).
| Task type | Dataset | Main metric | Value |
|---|---|---|---:|
| STS | stsb-tr | spearman_cosine | 0.7907 |
| BitextMining | wmt16_en_tr | accuracy | 0.9753 |
| Classification | news-cat | accuracy | 0.9600 |
| Classification | 75haber ᵃ | accuracy | 0.8696 |
| Classification | ts_timeline_news_category | accuracy | 0.7250 |
| Classification | thy_sa ᵃ | accuracy | 0.8386 |
| Classification | offenseval | accuracy | 0.8427 |
| Classification | irony-tr ᵃ | accuracy | 0.6667 |
| PairClassification | snli_tr | ap_cosine | 0.5280 |
| PairClassification | multinli_tr | ap_cosine | 0.6324 |
| PairClassification | xnli_tr | ap_cosine | 0.5591 |
| Clustering | ts_abstract_corpus | v_measure | 0.5762 |
| Clustering | 630koseyazisi | v_measure | 0.3978 |
| Retrieval | tquad | ndcg@10 | 0.8222 |
| Retrieval | quora-tr | ndcg@10 | 0.7323 |
| Retrieval | squad-tr | ndcg@10 | 0.6584 |
| Retrieval | scifact-tr | ndcg@10 | 0.5309 |
| Retrieval | cqadupstack-gaming-tr | ndcg@10 | 0.3584 |
| Retrieval | msmarco-tr | ndcg@10 | 0.2734 |
| Retrieval | arguana-tr | ndcg@10 | 0.2583 |
| Retrieval | nfcorpus-tr | ndcg@10 | 0.2272 |
| Retrieval | fiqa-tr | ndcg@10 | 0.1237 |
| Retrieval | scidocs-tr | ndcg@10 | 0.0887 |
---
## Limitations & Biases
- **Translated training data.** The NLI triplets are machine-translated; translation
artefacts and noise are inherited, which can cap fine-grained semantic precision.
- **Short-text bias.** Trained on NLI-style short sentences at 256 tokens; long-document
retrieval requires chunking and may underperform (see the lower `fiqa`/`scidocs` scores).
- **Prefix sensitivity.** Using the model without `query:`/`passage:` prefixes degrades results.
- **Domain gaps.** Scientific/financial retrieval (scidocs, fiqa) is weak; the model is
strongest on general-domain QA-style retrieval (tquad, squad-tr, quora-tr).
- **Societal bias.** The backbone and NLI data may encode social biases; validate before
deployment in sensitive settings.
- **Single epoch.** Trained for 1 epoch; longer schedules or larger effective batches may
improve retrieval further.
## Compute & Environmental Impact
- Single consumer laptop GPU (RTX 5050, 8 GB). Full pipeline (Optuna sweep + 1-epoch
training on 277k triplets + evaluation) ran in a few GPU-hours. No multi-GPU/cluster
training was used.
## Licensing
- **Backbone** `intfloat/multilingual-e5-base`: MIT.
- **Training data** derives from **AllNLI** (SNLI is **CC BY-SA 4.0**; MultiNLI has mixed
source licenses), translated to Turkish. Because the most restrictive component is
CC BY-SA 4.0, this model card is released under **CC BY-SA 4.0**; verify that this suits
your use case and comply with the ShareAlike terms.
- **Evaluation datasets** belong to their respective authors (see the `trmteb` org and
`emrecan/stsb-mt-turkish`).
## Citation
This model:
```bibtex
@misc{e5-tr-nli,
title = {e5-tr-nli: A Turkish Sentence Embedding Model},
note = {Fine-tuned from intfloat/multilingual-e5-base on Turkish NLI triplets},
year = {2026}
}
```
TR-MTEB benchmark:
```bibtex
@inproceedings{baysan-gungor-2025-trmteb,
title = {{TR-MTEB}: A Comprehensive Benchmark and Embedding Model Suite for {T}urkish Sentence Representations},
author = {Baysan, Mehmet Selman and G{\"u}ng{\"o}r, Tunga},
booktitle = {Findings of the Association for Computational Linguistics: EMNLP 2025},
year = {2025}
}
```
Base model (E5):
```bibtex
@article{wang2024multilingual,
title = {Multilingual E5 Text Embeddings: A Technical Report},
author = {Wang, Liang and Yang, Nan and Huang, Xiaolong and Yang, Linjun and Majumder, Rangan and Wei, Furu},
journal = {arXiv preprint arXiv:2402.05672},
year = {2024}
}
```
## Reproduction
The full pipeline (data prep, Optuna sweep, training, and TR-MTEB evaluation) is scripted:
- `train_embedding_tr.py` — `--mode sweep | train | eval | mteb`
- `eval_trmteb_hf.py` — TR-MTEB evaluation straight from HuggingFace `trmteb/*` datasets
- `config.py`, `Makefile`, `requirements.txt`
To publish (disabled by default — no secrets are used):
```python
from sentence_transformers import SentenceTransformer
SentenceTransformer("models/e5-tr-nli-final").push_to_hub("thealper2/intfloat-multilingual-e5-base-tr-nli")
```
|