--- tags: - sentence-transformers - sentence-similarity - feature-extraction - dense - generated_from_trainer - dataset_size:137745 - loss:CosineSimilarityLoss - loss:MultipleNegativesRankingLoss base_model: - intfloat/multilingual-e5-large widget: - source_sentence: Mor a epidémia sa očividne vymkli spod kontroly . sentences: - Choroba bola nekontrolovateľná a ohrozovala všetok život . - Tieto vylúčenia sú určené na iné cieľové skupiny ako obchodné štvrte . - >- Autobus National Trust Tour odchádza každý deň o 9:00 z National Trust Information Centre . - source_sentence: Bombové útoky v Iraku zabili najmenej 26 ľudí sentences: - Nová sláva pre Čínu v snahe o obnovu - Samovražedný útočník v Iraku zabil najmenej 12 ľudí - Muž lúpe banán. - source_sentence: Chlapec a dievča tancujú v daždi. sentences: - Turecko sa snaží premiestniť časť sýrskych utečencov - Záber zblízka na bielu ovcu s čiernou hlavou. - Muž a žena tancujú v daždi. - source_sentence: Niekedy fantazírujem o tom, že mám moc. sentences: - Egyptského Mursího čaká súd za útek z väzenia - >- Aj ja niekedy fantazírujem o tom, ako ohýbam ľudí pred svojou dementnou vôľou. - Niekto lúpe banán. - source_sentence: >- Vidiecke mydlo „Emmerdale“ vo štvrtok sledovalo 8,3 milióna ľudí, zatiaľ čo živá epizóda „The Bill“ mala 7,9 milióna divákov. sentences: - >- Hodinové špeciálne epizódy seriálov ITV „Emmerdale“ a „The Bill“ boli najlepšie hodnotené britské televízne programy pri príležitosti 50. výročia komerčného kanála. - 4. mája 88 bolo unesených niekoľko honorárnych konzulov. - >- Existujú spoločnosti, ktoré majú rozdelené riadenie procesov výskumu a vývoja. pipeline_tag: sentence-similarity library_name: sentence-transformers metrics: - pearson_cosine - spearman_cosine - cosine_accuracy - cosine_accuracy_threshold - cosine_f1 - cosine_f1_threshold - cosine_precision - cosine_recall - cosine_ap - cosine_mcc - validation_sts_pearson_cosine - validation_sts_spearman_cosine - validation_nli_cosine_accuracy - validation_nli_cosine_accuracy_threshold - validation_nli_cosine_f1 - validation_nli_cosine_f1_threshold - validation_nli_cosine_precision - validation_nli_cosine_recall - validation_nli_cosine_ap - validation_nli_cosine_mcc - validation_rte_cosine_accuracy - validation_rte_cosine_accuracy_threshold - validation_rte_cosine_f1 - validation_rte_cosine_f1_threshold - validation_rte_cosine_precision - validation_rte_cosine_recall - validation_rte_cosine_ap - validation_rte_cosine_mcc - validation_dev_overall model-index: - name: E5-sk-large results: - task: type: semantic-similarity name: Semantic Similarity dataset: name: validation sts type: validation_sts metrics: - type: pearson_cosine value: 0.8404018187290129 name: Pearson Cosine - type: spearman_cosine value: 0.8424145819561701 name: Spearman Cosine - task: type: binary-classification name: Binary Classification dataset: name: validation nli type: validation_nli metrics: - type: cosine_accuracy value: 0.6662650602409639 name: Cosine Accuracy - type: cosine_accuracy_threshold value: 0.9897347688674927 name: Cosine Accuracy Threshold - type: cosine_f1 value: 0.49969824984912486 name: Cosine F1 - type: cosine_f1_threshold value: 0.7573917508125305 name: Cosine F1 Threshold - type: cosine_precision value: 0.3333333333333333 name: Cosine Precision - type: cosine_recall value: 0.9975903614457832 name: Cosine Recall - type: cosine_ap value: 0.2791130922755554 name: Cosine Ap - type: cosine_mcc value: 0 name: Cosine Mcc - task: type: binary-classification name: Binary Classification dataset: name: validation rte type: validation_rte metrics: - type: cosine_accuracy value: 0.5234657039711191 name: Cosine Accuracy - type: cosine_accuracy_threshold value: 0.9790897369384766 name: Cosine Accuracy Threshold - type: cosine_f1 value: 0.645320197044335 name: Cosine F1 - type: cosine_f1_threshold value: 0.7796555757522583 name: Cosine F1 Threshold - type: cosine_precision value: 0.4763636363636364 name: Cosine Precision - type: cosine_recall value: 1 name: Cosine Recall - type: cosine_ap value: 0.37950776315755064 name: Cosine Ap - type: cosine_mcc value: 0.08078074811027249 name: Cosine Mcc - task: type: multi-task-dev name: Multi Task Dev dataset: name: Unknown type: unknown metrics: - type: validation_sts_pearson_cosine value: 0.8404018187290129 name: Validation Sts Pearson Cosine - type: validation_sts_spearman_cosine value: 0.8424145819561701 name: Validation Sts Spearman Cosine - type: validation_nli_cosine_accuracy value: 0.6662650602409639 name: Validation Nli Cosine Accuracy - type: validation_nli_cosine_accuracy_threshold value: 0.9897347688674927 name: Validation Nli Cosine Accuracy Threshold - type: validation_nli_cosine_f1 value: 0.49969824984912486 name: Validation Nli Cosine F1 - type: validation_nli_cosine_f1_threshold value: 0.7573917508125305 name: Validation Nli Cosine F1 Threshold - type: validation_nli_cosine_precision value: 0.3333333333333333 name: Validation Nli Cosine Precision - type: validation_nli_cosine_recall value: 0.9975903614457832 name: Validation Nli Cosine Recall - type: validation_nli_cosine_ap value: 0.2791130922755554 name: Validation Nli Cosine Ap - type: validation_nli_cosine_mcc value: 0 name: Validation Nli Cosine Mcc - type: validation_rte_cosine_accuracy value: 0.5234657039711191 name: Validation Rte Cosine Accuracy - type: validation_rte_cosine_accuracy_threshold value: 0.9790897369384766 name: Validation Rte Cosine Accuracy Threshold - type: validation_rte_cosine_f1 value: 0.645320197044335 name: Validation Rte Cosine F1 - type: validation_rte_cosine_f1_threshold value: 0.7796555757522583 name: Validation Rte Cosine F1 Threshold - type: validation_rte_cosine_precision value: 0.4763636363636364 name: Validation Rte Cosine Precision - type: validation_rte_cosine_recall value: 1 name: Validation Rte Cosine Recall - type: validation_rte_cosine_ap value: 0.37950776315755064 name: Validation Rte Cosine Ap - type: validation_rte_cosine_mcc value: 0.08078074811027249 name: Validation Rte Cosine Mcc - type: validation_dev_overall value: 0.5003451457964254 name: Validation Dev Overall license: mit datasets: - slovak-nlp/sklep language: - sk --- # e5-sk-large **e5-sk-large** is a Slovak text embedding model (365M parameters, 1024-dimensional embeddings) built by applying vocabulary trimming and fine-tuning to [multilingual-e5-large](https://huggingface.co/intfloat/multilingual-e5-large). It achieves competitive performance with proprietary embedding APIs on [SkMTEB](https://huggingface.co/collections/slovak-nlp/skmteb) — the first comprehensive Slovak text embedding benchmark — while being 35% smaller than the original model and fully locally deployable. Released as part of the **SkMTEB** project ([paper](https://arxiv.org/abs/2606.13647) · [GitHub](https://github.com/slovak-nlp/skmteb) · [collection](https://huggingface.co/collections/slovak-nlp/skmteb)). For a smaller, faster variant, see [e5-sk-small](https://huggingface.co/slovak-nlp/e5-sk-small) (45M parameters). --- ## Model Details | | | |---|---| | **Base model** | [intfloat/multilingual-e5-large](https://huggingface.co/intfloat/multilingual-e5-large) | | **Parameters** | 365M (vs. 560M original — 35% reduction) | | **Embedding dimension** | 1024 | | **Max sequence length** | 256 tokens | | **Pooling** | Mean pooling | | **Languages** | Slovak (primary); Slovak–English and Slovak–Czech cross-lingual tasks preserved | ## Usage This model follows the standard E5 prefix convention: prepend `query:` to queries and `passage:` to documents during retrieval. For symmetric tasks (STS, clustering, classification), no prefix is needed. ### With `sentence-transformers` First install the Sentence Transformers library: ```bash pip install -U sentence-transformers ``` Then you can load this model and run inference. ```python from sentence_transformers import SentenceTransformer model = SentenceTransformer("slovak-nlp/e5-sk-large") # Retrieval query_embedding = model.encode("query: Čo je hlavné mesto Slovenska?") passage_embedding = model.encode("passage: Bratislava je hlavné a najväčšie mesto Slovenska.") similarity = model.similarity(query_embedding, passage_embedding) print(similarity) # tensor([[0.9269]]) # Batch encoding sentences = [ "query: Aké je počasie v Bratislave?", "passage: V Bratislave je dnes slnečno a teplo.", "passage: Bratislava leží na brehu Dunaja.", ] embeddings = model.encode(sentences) print(embeddings.shape) # (3, 1024) ``` ### With `transformers` directly ```python import torch import torch.nn.functional as F from transformers import AutoTokenizer, AutoModel def average_pool(last_hidden_states, attention_mask): last_hidden = last_hidden_states.masked_fill(~attention_mask[..., None].bool(), 0.0) return last_hidden.sum(dim=1) / attention_mask.sum(dim=1)[..., None] tokenizer = AutoTokenizer.from_pretrained("slovak-nlp/e5-sk-large") model = AutoModel.from_pretrained("slovak-nlp/e5-sk-large") texts = [ "query: Čo je hlavné mesto Slovenska?", "passage: Bratislava je hlavné a najväčšie mesto Slovenska.", ] batch_dict = tokenizer(texts, max_length=512, padding=True, truncation=True, return_tensors="pt") with torch.no_grad(): outputs = model(**batch_dict) embeddings = average_pool(outputs.last_hidden_state, batch_dict["attention_mask"]) embeddings = F.normalize(embeddings, p=2, dim=1) print((embeddings[0] @ embeddings[1]).item()) ``` ### Prefix guide | Task | Input prefix | |---|---| | Search / retrieval query | `query: ` | | Document / passage to index | `passage: ` | | STS, clustering, classification | *(no prefix)* | --- ## Training ### How it was built **Step 1 — Vocabulary Trimming.** Before fine-tuning, [Vocabulary Trimming](https://aclanthology.org/2023.findings-emnlp.1029/) (Ushio et al., 2023) was applied to `multilingual-e5-large` to remove tokens irrelevant to Slovak. Token frequencies were computed on [FineWeb2-Slovak](https://huggingface.co/datasets/ivykopal/fineweb2-slovak), a quality-filtered Slovak web corpus, and the top 60K tokens (out of 250K) were retained. This reduced the model from 560M to **365M parameters** (35% reduction) without meaningful performance loss. **Step 2 — Fine-tuning.** The trimmed model was fine-tuned on curated Slovak datasets from the [skLEP benchmark](https://huggingface.co/slovak-nlp/sklep): | Dataset | Task | Pairs | |---|---|---| | SK-SQuAD | Question–context retrieval | ~72K | | Slovak NLI (from XNLI) | Entailment | ~393K | | Slovak STS (from GLUE STSb) | Similarity scoring | ~6K | | Slovak RTE (from GLUE) | Textual entailment | ~2.5K | ### Training configuration | Hyperparameter | Value | |---|---| | Pooling | Mean pooling | | Max sequence length | 256 tokens | | Batch size | 32 | | Learning rate | 2 × 10⁻⁵ | | LR scheduler | Linear with 10% warmup | | Epochs | 3 | | Loss (STS) | Cosine Similarity Loss | | Loss (other) | Multiple Negatives Ranking Loss | | Hardware | 1× NVIDIA H100 (~50 min) | | Seed | 42 | --- ## Evaluation: SkMTEB Results Evaluated on [SkMTEB](https://huggingface.co/collections/slovak-nlp/skmteb) — 31 datasets across 7 task types. Scores are percentages (higher is better). | Model | Params | **All** | Bitext | Classif. | Clustering | Pair Clf. | Reranking | Retrieval | STS | |---|---|---|---|---|---|---|---|---|---| | **e5-sk-large** | **365M** | **74.70** | 96.39 | 66.34 | 41.43 | 67.32 | 87.81 | 85.60 | 86.25 | | multilingual-e5-large | 560M | 74.25 | 96.29 | 65.34 | 40.35 | 66.78 | 87.96 | 85.80 | 85.90 | | text-embedding-3-large (API) | — | 75.07 | 96.79 | 66.91 | 44.22 | 66.58 | 86.96 | 85.55 | 84.21 | | multilingual-e5-large-instruct | 560M | 77.49 | 97.09 | 70.28 | 49.69 | 70.55 | 86.49 | 86.08 | 88.86 | | e5-sk-small (ours) | 45M | 70.56 | 91.34 | 60.84 | 40.95 | 66.05 | 84.94 | 78.64 | 81.32 | **e5-sk-large is practically equivalent to `text-embedding-3-large`** (TOST equivalence test: 90% CI within ±2 points), while being open-weight, locally deployable, and free to run. Cross-lingual Slovak–English and Slovak–Czech bitext mining performance is preserved within 1 F1 point compared to the original `multilingual-e5-large`. --- ## Intended Uses - Semantic search and retrieval-augmented generation (RAG) over Slovak text - Semantic textual similarity (STS) - Text clustering and classification via embedding features - Cross-lingual retrieval (Slovak–English, Slovak–Czech) - Local deployment where API latency or cost is a concern ## Limitations - Optimised for Slovak; cross-lingual transfer to non-Slavic languages is not evaluated. - Vocabulary trimming removes non-Slovak tokens; performance on heavily code-mixed text may be reduced. - Training data skews toward news, parliamentary, and encyclopedic domains. - Max sequence length during fine-tuning is 256 tokens (underlying architecture supports up to 512). --- ## Citation ```bibtex @inproceedings{suppa2025skmteb, title = {{SkMTEB}: {Slovak} Massive Text Embedding Benchmark and Model Adaptation}, author = {{\v{S}}uppa, Marek and Ridzik, Andrej and Hl{\'a}dek, Daniel and Kna{\v{z}}ekov{\'a}, Nat{\'a}lia and Ondrejov{\'a}, Vikt{\'o}ria}, year = {2025}, eprint = {2606.13647}, archivePrefix = {arXiv}, url = {https://arxiv.org/abs/2606.13647} } @inproceedings{reimers-2019-sentence-bert, title = {Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks}, author = {Reimers, Nils and Gurevych, Iryna}, booktitle = {Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing}, year = {2019}, publisher = {Association for Computational Linguistics}, url = {https://arxiv.org/abs/1908.10084} } ``` --- ## License MIT