File size: 14,129 Bytes
cf90e06
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
b525f45
cf90e06
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
b525f45
 
cf90e06
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
b525f45
cf90e06
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
---
language:
- tr
license: cc-by-sa-4.0
library_name: sentence-transformers
pipeline_tag: sentence-similarity
tags:
- sentence-transformers
- sentence-similarity
- feature-extraction
- embeddings
- turkish
- türkçe
- e5
- retrieval
- semantic-search
- mteb
- tr-mteb
base_model: intfloat/multilingual-e5-base
datasets:
- mertcobanov/all-nli-triplets-turkish
- emrecan/stsb-mt-turkish
metrics:
- spearmanr
- ndcg
- accuracy
- v_measure
model-index:
- name: e5-tr-nli
  results:
  - task:
      type: semantic-textual-similarity
      name: Semantic Textual Similarity
    dataset:
      name: STSb-TR (test)
      type: trmteb/stsb-tr
    metrics:
    - type: spearman_cosine
      value: 0.7907
  - task:
      type: retrieval
      name: Retrieval
    dataset:
      name: TQuad
      type: trmteb/tquad
    metrics:
    - type: ndcg_at_10
      value: 0.8222
  - task:
      type: retrieval
      name: Retrieval
    dataset:
      name: Quora-TR
      type: trmteb/quora-tr
    metrics:
    - type: ndcg_at_10
      value: 0.7323
  - task:
      type: retrieval
      name: Retrieval
    dataset:
      name: MS MARCO-TR
      type: trmteb/msmarco-tr
    metrics:
    - type: ndcg_at_10
      value: 0.2734
  - task:
      type: classification
      name: Classification
    dataset:
      name: News-Cat
      type: trmteb/news-cat
    metrics:
    - type: accuracy
      value: 0.9600
---

# e5-tr-nli — Turkish Sentence Embedding Model (Bi-Encoder)

A **Turkish sentence-embedding (bi-encoder) model** for **retrieval and semantic
search**, fine-tuned from [`intfloat/multilingual-e5-base`](https://huggingface.co/intfloat/multilingual-e5-base)
with contrastive learning (`MultipleNegativesRankingLoss`) on Turkish NLI triplets.

- **Embedding dimension:** 768
- **Max sequence length:** 256 tokens (trained); backbone supports 512
- **Pooling:** mean pooling
- **Similarity function:** cosine
- **Parameters:** ~278M (XLM-RoBERTa-base backbone)
- **Prefix convention:** e5-style `query:` / `passage:` (see [Usage](#usage) — **required**)

---

## Table of Contents
1. [Intended Use](#intended-use)
2. [Usage](#usage)
3. [Training Data](#training-data)
4. [Training Procedure](#training-procedure)
5. [Evaluation](#evaluation)
6. [TR-MTEB Results](#tr-mteb-results-full)
7. [Limitations & Biases](#limitations--biases)
8. [Compute & Environmental Impact](#compute--environmental-impact)
9. [Licensing](#licensing)
10. [Citation](#citation)
11. [Reproduction](#reproduction)

---

## Intended Use

**In scope**
- Turkish semantic search / passage retrieval (dense retrieval, RAG)
- Semantic textual similarity, paraphrase & duplicate detection
- Clustering and topic grouping of Turkish text
- Feature extraction for downstream Turkish NLP classifiers

**Out of scope**
- Re-ranking with query–document cross-attention (use a cross-encoder instead)
- Long-document embedding beyond 256 tokens without chunking
- High-stakes decisions (legal, medical, hiring) without human oversight
- Non-Turkish text (backbone is multilingual, but this model is tuned for Turkish)

---

## Usage

### ⚠️ e5 prefixes are mandatory
The model was trained with the e5 asymmetric prefixes and **must** be used with them,
consistently at training / evaluation / inference:

| Role | Prefix |
|------|--------|
| Query / anchor / any symmetric-task sentence (STS, classification, clustering) | `query: ` |
| Passage / document / positive & negative candidates | `passage: ` |

Omitting the prefixes degrades quality noticeably.

### With `sentence-transformers`
```python
from sentence_transformers import SentenceTransformer
from sentence_transformers.util import cos_sim

model = SentenceTransformer("thealper2/intfloat-multilingual-e5-base-tr-nli")

# Asymmetric retrieval: query vs. candidate passages
query = "query: Türkiye'nin başkenti neresidir?"
passages = [
    "passage: Ankara, Türkiye'nin başkentidir.",
    "passage: İstanbul Türkiye'nin en kalabalık şehridir.",
    "passage: Muz tropikal bir meyvedir.",
]
q = model.encode(query, normalize_embeddings=True)
p = model.encode(passages, normalize_embeddings=True)
print(cos_sim(q, p))   # highest score -> the Ankara passage
```

### Semantic similarity (symmetric — use `query:` on both sides)
```python
a = model.encode("query: Bugün hava çok güzel.", normalize_embeddings=True)
b = model.encode("query: Hava bugün oldukça güzel.", normalize_embeddings=True)
print(float(cos_sim(a, b)))   # ~0.9
```

### With 🤗 Transformers (mean pooling)
```python
import torch, torch.nn.functional as F
from transformers import AutoTokenizer, AutoModel

tok = AutoTokenizer.from_pretrained("thealper2/intfloat-multilingual-e5-base-tr-nli")
mdl = AutoModel.from_pretrained("thealper2/intfloat-multilingual-e5-base-tr-nli")

def encode(texts):
    batch = tok(texts, padding=True, truncation=True, max_length=256, return_tensors="pt")
    with torch.no_grad():
        out = mdl(**batch)
    mask = batch["attention_mask"].unsqueeze(-1).float()
    emb = (out.last_hidden_state * mask).sum(1) / mask.sum(1)   # mean pooling
    return F.normalize(emb, p=2, dim=1)

emb = encode(["query: örnek cümle", "passage: örnek pasaj"])
```

---

## Training Data

- **Source:** [`mertcobanov/all-nli-triplets-turkish`](https://huggingface.co/datasets/mertcobanov/all-nli-triplets-turkish)
  — a machine-translated Turkish version of the AllNLI (SNLI + MultiNLI) triplet set.
- **Format:** `(anchor, positive, negative)` triplets, where the `negative` acts as a
  **hard negative** for the contrastive objective.
- **Column handling:** **Only the Turkish columns** (`anchor_translated`,
  `positive_translated`, `negative_translated`) were used and renamed to
  `anchor` / `positive` / `negative`. **All English columns were discarded.**
- **Cleaning:** rows with `None` / empty / whitespace-only fields were filtered out.
- **Resulting sizes (after filtering):**

  | Split | Triplets |
  |-------|---------:|
  | train | **277,167** |
  | dev   | 6,584 |
  | test  | 6,609 |

---

## Training Procedure

### Objective
- **Loss:** `CachedMultipleNegativesRankingLoss` (MNRL with in-batch + hard negatives;
  the cached variant is used to simulate a large effective batch on limited VRAM).
- **Batch sampler:** `NO_DUPLICATES` (required for MNRL to avoid trivial in-batch collisions).
- **In-batch negatives:** every other positive/negative in the batch serves as a negative
  for a given anchor, so larger batches yield a stronger contrastive signal.

### Hyperparameter search
An **Optuna** search (TPE sampler) was run **before** the final training:

| Setting | Value |
|--------|-------|
| Trials | 10 |
| Subset | 40,000 random training triplets |
| Steps / trial | 600 |
| Search space | `lr ∈ {1e-5, 2e-5, 3e-5}`, `batch ∈ {32, 64, 128}`, `warmup_ratio ∈ {0.0, 0.1}` |
| Selection metric | `stsb-tr` **dev** `spearman_cosine` |
| **Best config** | **lr = 2e-5, batch = 32, warmup_ratio = 0.1** (dev spearman = **0.8285**) |

### Final training

| Setting | Value |
|--------|-------|
| Base model | `intfloat/multilingual-e5-base` |
| Epochs | 1 |
| Learning rate | 2e-5 |
| Batch size | 32 (cached MNRL) |
| Warmup ratio | 0.1 |
| Max sequence length | 256 |
| Precision | bf16 |
| Seed | 42 |
| Eval strategy | steps, monitoring `eval_stsb-tr-dev_spearman_cosine` |
| Checkpoint selection | `load_best_model_at_end=True` (best dev spearman) |

### Frameworks & hardware
- `sentence-transformers` 5.2.2, `transformers` 5.0.0, `torch` 2.11.0+cu128,
  `datasets` 4.4.1, Python 3.12
- 1× NVIDIA GeForce RTX 5050 Laptop GPU (8 GB), CUDA 12.8

---

## Evaluation

### Baseline (base model) vs. Final (this model)
Measured on the dataset's own test triplets (`TripletEvaluator`, cosine accuracy) and
on the external Turkish STS set [`emrecan/stsb-mt-turkish`](https://huggingface.co/datasets/emrecan/stsb-mt-turkish)
(`EmbeddingSimilarityEvaluator`, Spearman; scores normalised 0–5 → 0–1):

| Metric | Baseline (e5-base) | Final | Δ |
|--------|:---:|:---:|:---:|
| stsb-tr **dev** spearman_cosine | 0.8092 | **0.8312** | +0.0220 |
| stsb-tr **test** spearman_cosine | 0.7761 | **0.7814** | +0.0053 |
| nli-tr **test** cosine_accuracy | 0.8968 | **0.9260** | +0.0292 |

### Methodology notes for TR-MTEB
- Datasets pulled **directly from HuggingFace** (`trmteb/*`); metrics computed with a
  standalone script (no `mteb` package dependency), following standard MTEB protocols:
  - **STS** → Spearman of cosine vs. gold; **Retrieval** → nDCG@10 / Recall@10 / MAP@10
    via exact cosine search; **Classification** → logistic regression on frozen
    embeddings (accuracy + macro-F1); **PairClassification** → average precision of cosine;
    **Clustering** → V-measure (MiniBatchKMeans, k = #labels); **BitextMining** → top-1
    nearest-neighbour accuracy.
- e5 prefixes applied throughout (`query:` for queries/symmetric sides, `passage:` for corpus).
- Retrieval search runs on-GPU in fp16.
- `75haber`, `thy_sa`, `irony-tr` have **no train split** in the `trmteb` org, so a seeded
  **70/30 stratified self-split** was used for those classification tasks (marked `note` in
  `trmteb_results.json`). All other classification sets use their official train/test.

### Category summary (macro averages)

| Task type | # datasets | Avg. main metric |
|-----------|:---:|:---:|
| STS (spearman_cosine) | 1 | **0.791** |
| BitextMining (accuracy) | 1 | **0.975** |
| Classification (accuracy) | 6 | **0.817** |
| PairClassification (AP) | 3 | **0.573** |
| Clustering (v-measure) | 2 | **0.487** |
| Retrieval (nDCG@10) | 10 | **0.407** |
| **Overall (mean of category means)** | 23 | **0.675** |

---

## TR-MTEB Results (full)

Evaluated on the [TR-MTEB](https://huggingface.co/trmteb) datasets
(Baysan & Güngör, *TR-MTEB*, Findings of EMNLP 2025).

| Task type | Dataset | Main metric | Value |
|---|---|---|---:|
| STS | stsb-tr | spearman_cosine | 0.7907 |
| BitextMining | wmt16_en_tr | accuracy | 0.9753 |
| Classification | news-cat | accuracy | 0.9600 |
| Classification | 75haber ᵃ | accuracy | 0.8696 |
| Classification | ts_timeline_news_category | accuracy | 0.7250 |
| Classification | thy_sa ᵃ | accuracy | 0.8386 |
| Classification | offenseval | accuracy | 0.8427 |
| Classification | irony-tr ᵃ | accuracy | 0.6667 |
| PairClassification | snli_tr | ap_cosine | 0.5280 |
| PairClassification | multinli_tr | ap_cosine | 0.6324 |
| PairClassification | xnli_tr | ap_cosine | 0.5591 |
| Clustering | ts_abstract_corpus | v_measure | 0.5762 |
| Clustering | 630koseyazisi | v_measure | 0.3978 |
| Retrieval | tquad | ndcg@10 | 0.8222 |
| Retrieval | quora-tr | ndcg@10 | 0.7323 |
| Retrieval | squad-tr | ndcg@10 | 0.6584 |
| Retrieval | scifact-tr | ndcg@10 | 0.5309 |
| Retrieval | cqadupstack-gaming-tr | ndcg@10 | 0.3584 |
| Retrieval | msmarco-tr | ndcg@10 | 0.2734 |
| Retrieval | arguana-tr | ndcg@10 | 0.2583 |
| Retrieval | nfcorpus-tr | ndcg@10 | 0.2272 |
| Retrieval | fiqa-tr | ndcg@10 | 0.1237 |
| Retrieval | scidocs-tr | ndcg@10 | 0.0887 |

---

## Limitations & Biases
- **Translated training data.** The NLI triplets are machine-translated; translation
  artefacts and noise are inherited, which can cap fine-grained semantic precision.
- **Short-text bias.** Trained on NLI-style short sentences at 256 tokens; long-document
  retrieval requires chunking and may underperform (see the lower `fiqa`/`scidocs` scores).
- **Prefix sensitivity.** Using the model without `query:`/`passage:` prefixes degrades results.
- **Domain gaps.** Scientific/financial retrieval (scidocs, fiqa) is weak; the model is
  strongest on general-domain QA-style retrieval (tquad, squad-tr, quora-tr).
- **Societal bias.** The backbone and NLI data may encode social biases; validate before
  deployment in sensitive settings.
- **Single epoch.** Trained for 1 epoch; longer schedules or larger effective batches may
  improve retrieval further.

## Compute & Environmental Impact
- Single consumer laptop GPU (RTX 5050, 8 GB). Full pipeline (Optuna sweep + 1-epoch
  training on 277k triplets + evaluation) ran in a few GPU-hours. No multi-GPU/cluster
  training was used.

## Licensing
- **Backbone** `intfloat/multilingual-e5-base`: MIT.
- **Training data** derives from **AllNLI** (SNLI is **CC BY-SA 4.0**; MultiNLI has mixed
  source licenses), translated to Turkish. Because the most restrictive component is
  CC BY-SA 4.0, this model card is released under **CC BY-SA 4.0**; verify that this suits
  your use case and comply with the ShareAlike terms.
- **Evaluation datasets** belong to their respective authors (see the `trmteb` org and
  `emrecan/stsb-mt-turkish`).

## Citation

This model:
```bibtex
@misc{e5-tr-nli,
  title  = {e5-tr-nli: A Turkish Sentence Embedding Model},
  note   = {Fine-tuned from intfloat/multilingual-e5-base on Turkish NLI triplets},
  year   = {2026}
}
```

TR-MTEB benchmark:
```bibtex
@inproceedings{baysan-gungor-2025-trmteb,
  title     = {{TR-MTEB}: A Comprehensive Benchmark and Embedding Model Suite for {T}urkish Sentence Representations},
  author    = {Baysan, Mehmet Selman and G{\"u}ng{\"o}r, Tunga},
  booktitle = {Findings of the Association for Computational Linguistics: EMNLP 2025},
  year      = {2025}
}
```

Base model (E5):
```bibtex
@article{wang2024multilingual,
  title   = {Multilingual E5 Text Embeddings: A Technical Report},
  author  = {Wang, Liang and Yang, Nan and Huang, Xiaolong and Yang, Linjun and Majumder, Rangan and Wei, Furu},
  journal = {arXiv preprint arXiv:2402.05672},
  year    = {2024}
}
```

## Reproduction
The full pipeline (data prep, Optuna sweep, training, and TR-MTEB evaluation) is scripted:
- `train_embedding_tr.py``--mode sweep | train | eval | mteb`
- `eval_trmteb_hf.py` — TR-MTEB evaluation straight from HuggingFace `trmteb/*` datasets
- `config.py`, `Makefile`, `requirements.txt`

To publish (disabled by default — no secrets are used):
```python
from sentence_transformers import SentenceTransformer
SentenceTransformer("models/e5-tr-nli-final").push_to_hub("thealper2/intfloat-multilingual-e5-base-tr-nli")
```