Sentence Similarity
sentence-transformers
Safetensors
English
bert
feature-extraction
Generated from Trainer
dataset_size:13284
loss:MultipleNegativesRankingLoss
domain:agri-food
task:ontology-alignment
Eval Results (legacy)
text-embeddings-inference
Instructions to use P2M2/Ows2-MiniLM_1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- sentence-transformers
How to use P2M2/Ows2-MiniLM_1 with sentence-transformers:
from sentence_transformers import SentenceTransformer model = SentenceTransformer("P2M2/Ows2-MiniLM_1") sentences = [ "That is a happy person", "That is a happy dog", "That is a very happy person", "Today is a sunny day" ] embeddings = model.encode(sentences) similarities = model.similarity(embeddings, embeddings) print(similarities.shape) # [4, 4] - Notebooks
- Google Colab
- Kaggle
| tags: | |
| - sentence-transformers | |
| - sentence-similarity | |
| - feature-extraction | |
| - generated_from_trainer | |
| - dataset_size:13284 | |
| - loss:MultipleNegativesRankingLoss | |
| - domain:agri-food | |
| - task:ontology-alignment | |
| base_model: sentence-transformers/all-MiniLM-L6-v2 | |
| language: | |
| - en | |
| license: apache-2.0 | |
| pipeline_tag: sentence-similarity | |
| library_name: sentence-transformers | |
| metrics: | |
| - cosine_accuracy@1 | |
| - cosine_accuracy@3 | |
| - cosine_mrr@10 | |
| - cosine_ndcg@10 | |
| model-index: | |
| - name: Ows2-MiniLM | |
| results: | |
| - task: | |
| type: information-retrieval | |
| name: Information Retrieval | |
| dataset: | |
| name: Validation set (InformationRetrievalEvaluator — 2238 queries, ~14 400 corpus terms) | |
| type: val_evaluator | |
| metrics: | |
| - type: cosine_accuracy@1 | |
| value: 0.3968 | |
| name: Cosine Accuracy@1 | |
| - type: cosine_accuracy@3 | |
| value: 0.5366 | |
| name: Cosine Accuracy@3 | |
| - type: cosine_accuracy@5 | |
| value: 0.5920 | |
| name: Cosine Accuracy@5 | |
| - type: cosine_ndcg@10 | |
| value: 0.5185 | |
| name: Cosine NDCG@10 | |
| - type: cosine_mrr@10 | |
| value: 0.4808 | |
| name: Cosine MRR@10 | |
| - task: | |
| type: information-retrieval | |
| name: Goldtest — Standard IR (goldtest_evaluator — 100 expert queries, ~14 600 corpus terms) | |
| dataset: | |
| name: OWS2 Goldtest | |
| type: goldtest | |
| metrics: | |
| - type: cosine_accuracy@1 | |
| value: 0.55 | |
| name: Cosine Accuracy@1 | |
| - type: cosine_accuracy@3 | |
| value: 0.75 | |
| name: Cosine Accuracy@3 | |
| - type: cosine_mrr@10 | |
| value: 0.6663 | |
| name: Cosine MRR@10 | |
| - task: | |
| type: information-retrieval | |
| name: Goldtest — Exact/Taboo Benchmark (CombinedGoldtestEvaluator — 100 expert queries, 14 544 unique terms) | |
| dataset: | |
| name: OWS2 Goldtest (Deduplicated) | |
| type: goldtest_combined | |
| metrics: | |
| - type: exact_match_accuracy@1 | |
| value: 0.6618 | |
| name: Exact Match Accuracy@1 | |
| - type: exact_match_mrr@10 | |
| value: 0.7394 | |
| name: Exact Match MRR@10 | |
| - type: taboo_match_mrr@10 | |
| value: 0.5448 | |
| name: Taboo Match MRR@10 | |
| - type: semantic_margin | |
| value: 0.1780 | |
| name: Semantic Discrimination Margin | |
| # Ows2-MiniLM — Ontology Alignment Model for Agri-Food Domains | |
| **Ows2-MiniLM** is a sentence embedding model fine-tuned for **semantic alignment between scientific text and agri-food ontology terms**. | |
| It is part of the **OWS² (Ontology-based Workflow for Semantic Specialisation)** project developed at [INRAE](https://www.inrae.fr/) (French National Research Institute for Agriculture, Food and Environment). | |
| Given a sentence extracted from a scientific abstract or publication, the model retrieves the most semantically relevant ontology terms from a reference vocabulary of ~14 600 concepts covering plant science, food science, and agronomy domains. | |
| The reference vocabulary incorporates concepts from: | |
| - **TransformON**: [Nature Publication](https://www.nature.com/articles/s41538-023-00221-2) | |
| - **Planteome**: [PubMed Publication](https://pubmed.ncbi.nlm.nih.gov/38055832/) | |
| > **Acknowledgement:** This work was supported by **INRAE Metaprogramme DIGIT-BIO** within the framework of the **SEED** project, and was carried out within the **Plateforme de Profilage Métabolique et de Métabolomique (P2M2)** — UMR 1349 IGEPP, INRAE. | |
| --- | |
| ## Model Details | |
| | Property | Value | | |
| |:---|:---| | |
| | **Model name** | Ows2-MiniLM | | |
| | **Base model** | [sentence-transformers/all-MiniLM-L6-v2](https://huggingface.co/sentence-transformers/all-MiniLM-L6-v2) | | |
| | **Architecture** | Bi-encoder (Transformer + Mean Pooling + L2 Normalize) | | |
| | **Output dimensionality** | 384 dimensions | | |
| | **Maximum sequence length** | 512 tokens | | |
| | **Similarity function** | Cosine similarity | | |
| | **License** | [Apache 2.0](https://www.apache.org/licenses/LICENSE-2.0) (inherits from base model) | | |
| | **Training loss** | MultipleNegativesRankingLoss (asymmetric, in-batch negatives) | | |
| | **Language** | English | | |
| | **Domain** | Agri-food science (plant biology, food science, agronomy) | | |
| ### Full Model Architecture | |
| ``` | |
| SentenceTransformer( | |
| (0): Transformer({'architecture': 'BertModel'}) | |
| (1): Pooling({'pooling_mode': 'mean'}) | |
| (2): Normalize({}) | |
| ) | |
| ``` | |
| --- | |
| ## Intended Use | |
| This model is designed for **ontology-based semantic search** in scientific literature, specifically: | |
| - **Automatic annotation** of scientific abstracts with ontology concepts | |
| - **Semantic retrieval** of the most relevant agri-food ontology terms given a sentence | |
| - **Alignment** between free-text descriptions and controlled vocabularies (e.g., PPDO, AgroVoc-related thesauri) | |
| > **Out-of-scope use:** This model is specialized for agri-food science. Performance on general-purpose or non-scientific text may be significantly lower. | |
| --- | |
| ## Usage | |
| ```python | |
| from sentence_transformers import SentenceTransformer, util | |
| model = SentenceTransformer("tcolombu/Ows2-MiniLM_1") | |
| # A sentence from a scientific abstract | |
| query = "Fruit-based sauces exhibited significant variations in antioxidant activity among different cultivars." | |
| # Ontology terms to rank | |
| ontology_terms = [ | |
| "fruit sauce", | |
| "antioxidant activity", | |
| "barley powdery mildew disease response", | |
| "rootless", | |
| ] | |
| # Encode and compute cosine similarities | |
| query_embedding = model.encode(query, convert_to_tensor=True) | |
| term_embeddings = model.encode(ontology_terms, convert_to_tensor=True) | |
| scores = util.cos_sim(query_embedding, term_embeddings)[0] | |
| ranked = sorted(zip(ontology_terms, scores.tolist()), key=lambda x: x[1], reverse=True) | |
| for term, score in ranked: | |
| print(f"{score:.4f} {term}") | |
| # Expected top result: fruit sauce | |
| ``` | |
| --- | |
| ## Evaluation | |
| Two distinct evaluation protocols were used, each measuring different aspects of model performance. **Both are computed on the same 100-query expert goldtest**, but differ in corpus construction and scoring logic. | |
| ### Evaluator 1 — Standard IR (`InformationRetrievalEvaluator`) | |
| > Called during training via `python main.py finetune --config <yaml>` | |
| This is the standard `InformationRetrievalEvaluator` from the `sentence-transformers` library. Each ontology term is registered with a composite key `"{ontology_id}_{term}"`, meaning the **same surface form can appear multiple times** in the corpus if it has multiple ontology IDs. Relevance is checked by exact `doc_id` match only. | |
| - **Corpus size:** ~14 600 entries (with duplicates) | |
| - **Scoring:** Single exact `doc_id` match required | |
| - **Stored in:** `metrics_comparison.json` | |
| | Dataset | Acc@1 | Acc@3 | MRR@10 | NDCG@10 | | |
| |:---|:---:|:---:|:---:|:---:| | |
| | Validation set (15% of training data) | 39.68% | 53.66% | 48.08% | 51.85% | | |
| | **Goldtest** (100 expert queries) | **55.00%** | **75.00%** | **66.63%** | 44.35% | | |
| --- | |
| ### Evaluator 2 — Combined Goldtest Benchmark (`CombinedGoldtestEvaluator`) | |
| > Called post-training via `python main.py eval_model <model_path>` | |
| A custom three-component evaluator. Terms are **deduplicated** to 14 544 unique surface forms. Synonymous ontology IDs sharing the same term are **merged**: a prediction is correct if any valid ID for the target term is retrieved. Queries are split into two subsets: | |
| - **Exact Match**: the target term appears verbatim in the query sentence | |
| - **Taboo Match**: the target term is *absent* from the query — pure semantic understanding required | |
| | Subset | Acc@1 | Acc@3 | MRR@10 | | |
| |:---|:---:|:---:|:---:| | |
| | **Exact Match** | **66.18%** | **79.41%** | **73.94%** | | |
| | **Taboo Match** | 37.50% | 68.75% | 54.48% | | |
| | Semantic metric | Value | | |
| |:---|:---:| | |
| | Mean similarity (positive targets) | 0.5580 | | |
| | Mean similarity (distractors) | 0.3801 | | |
| | **Discrimination margin** | **0.1780** | | |
| | Discrimination accuracy | 86.36% | | |
| *Stored in: `goldtest_eval/combined_goldtest_metrics.json`* | |
| > **Reading the two scores:** Evaluator 1 (66.63% MRR) is stricter because duplicated corpus entries act as additional distractors. Evaluator 2 (73.94% MRR) is more semantically faithful because synonyms of the correct answer are also accepted as valid. | |
| --- | |
| ### Baseline Comparison — `all-MiniLM-L6-v2` (no fine-tuning) vs `Ows2-MiniLM` | |
| All metrics below are computed on **our agri-food ontology corpus** (not the standard sentence-transformers benchmarks). The base model has never seen this vocabulary. | |
| #### Validation set — InformationRetrievalEvaluator (2 238 queries, ~14 400 corpus terms) | |
| | Metric | Base `all-MiniLM-L6-v2` | **Ows2-MiniLM** | Δ Gain | | |
| |:---|:---:|:---:|:---:| | |
| | Accuracy@1 | 23.15% | **39.68%** | +16.5 pts | | |
| | Accuracy@3 | 35.17% | **53.66%** | +18.5 pts | | |
| | Accuracy@5 | 41.51% | **59.20%** | +17.7 pts | | |
| | NDCG@10 | 34.46% | **51.85%** | +17.4 pts | | |
| | MRR@10 | 30.76% | **48.08%** | +17.3 pts | | |
| #### Goldtest — Expert benchmark (100 queries, ~14 600 corpus terms) | |
| | Metric | Base `all-MiniLM-L6-v2` | **Ows2-MiniLM** | Δ Gain | | |
| |:---|:---:|:---:|:---:| | |
| | Accuracy@1 | 22.00% | **55.00%** | **+33.0 pts** | | |
| | Accuracy@3 | 41.00% | **75.00%** | **+34.0 pts** | | |
| | Accuracy@5 | 52.00% | **84.00%** | +32.0 pts | | |
| | NDCG@10 | 24.16% | **44.35%** | +20.2 pts | | |
| | MRR@10 | 34.81% | **66.63%** | **+31.8 pts** | | |
| Fine-tuning on the OWS² mixed dataset (real + synthetic pairs) provides a **~+32 points absolute gain on Acc@1** and **~+32 points on MRR@10** on the expert goldtest, demonstrating the strong impact of domain specialisation even with limited real annotated data (~491 expert examples). | |
| --- | |
| ## Training Details | |
| ### Training Data | |
| The model was trained on a **mixed dataset** combining real annotated data and LLM-generated synthetic pairs: | |
| | Source | Examples | Description | | |
| |:---|:---:|:---| | |
| | Real annotations | ~491 | Expert-annotated (sentence, ontology term) pairs from scientific abstracts | | |
| | Synthetic data | ~14 428 | Sentence/term pairs generated with **[Qwen](https://huggingface.co/Qwen)** (local inference), covering the full ontology vocabulary | | |
| | **Total (85% train split)** | **~13 284** | Used for training | | |
| | Validation (15%) | ~2 341 | Used for in-training evaluation | | |
| - **Format:** `(anchor_sentence, positive_term, hard_negative_term)` triplets | |
| - **Loss:** `MultipleNegativesRankingLoss` — uses all other items in the batch as implicit negatives | |
| - **Synthetic generation LLM:** [Qwen](https://huggingface.co/Qwen) — run locally (Qwen2+ is Apache 2.0; Qwen1.x Tongyi license is compatible with domain-specific fine-tuned model publication) | |
| - **Dataset file:** `mixed_dataset_v0_full.jsonl` | |
| ### Training Hyperparameters | |
| | Parameter | Value | | |
| |:---|:---:| | |
| | Batch size | **16** | | |
| | Learning rate | **2e-5** | | |
| | Epochs | **15** | | |
| | Warmup ratio | 0.1 | | |
| | LR scheduler | Linear | | |
| | Optimizer | AdamW | | |
| | Freeze layers | 0 (full fine-tuning) | | |
| | Seed | 42 | | |
| > **Note on hyperparameter selection:** Batch size 16 with LR 2e-5 over 15 epochs was identified as the optimal configuration after a systematic grid search across 44 experiments. Increasing epochs beyond 15 showed no further gain (convergence plateau confirmed). | |
| ### Training Logs (validation NDCG@10 over 15 epochs) | |
| <details><summary>Click to expand</summary> | |
| | Epoch | Step | Training Loss | val_evaluator_cosine_ndcg@10 | | |
| |:---:|:---:|:---:|:---:| | |
| | 0.12 | 100 | 1.0499 | 0.3849 | | |
| | 0.96 | 800 | 0.6236 | 0.4861 | | |
| | 1.93 | 1600 | 0.4212 | 0.5072 | | |
| | 3.01 | 2500 | 0.3228 | 0.5107 | | |
| | 5.05 | 4200 | 0.1608 | 0.5094 | | |
| | 7.10 | 5900 | 0.1085 | 0.5137 | | |
| | 9.03 | 7500 | 0.0890 | 0.5177 | | |
| | 11.07 | 9200 | 0.0667 | 0.5170 | | |
| | 13.00 | 10800 | 0.0537 | 0.5173 | | |
| | **15.00** | **12465** | — | **0.5185** | | |
| </details> | |
| ### Training Time | |
| - **Training:** ~10.7 minutes (NVIDIA GPU, SLURM cluster) | |
| - **Evaluation:** ~7.8 minutes | |
| - **Total:** ~18.5 minutes | |
| --- | |
| ## Limitations | |
| - **Domain specificity:** The model is fine-tuned on agri-food science vocabulary. Performance on out-of-domain text (e.g., clinical, legal) is expected to be significantly lower. | |
| - **Ontology coverage:** The reference ontology contains ~14 600 terms. Terms absent from this vocabulary cannot be retrieved. | |
| - **Real data scarcity:** Only ~491 real annotated examples were available. Performance is expected to improve as the real annotation corpus grows. | |
| - **Taboo match gap:** The 19-point gap between Exact Match MRR (73.94%) and Taboo Match MRR (54.48%) indicates the model still relies partially on lexical overlap. Improving semantic generalization is an identified next step. | |
| --- | |
| ## Framework Versions | |
| - Python: 3.10.20 | |
| - Sentence Transformers: 5.6.0 | |
| - Transformers: 5.14.1 | |
| - PyTorch: 2.10.0+cu128 | |
| - Accelerate: 1.14.0 | |
| - Datasets: 5.0.0 | |
| --- | |
| ## Citation | |
| If you use this model, please cite: | |
| ```bibtex | |
| @misc{ows2-minilm-2025, | |
| title = {Ows2-MiniLM: An Ontology Alignment Model for Agri-Food Science}, | |
| institution = {INRAE, P2M2}, | |
| note = {OWS\textsuperscript{2} -- Ontology-based Workflow for Semantic Specialisation}, | |
| howpublished = {\url{https://huggingface.co/tcolombu/Ows2-MiniLM_1}} | |
| } | |
| ``` | |
| ```bibtex | |
| @inproceedings{reimers-2019-sentence-bert, | |
| title = {Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks}, | |
| author = {Reimers, Nils and Gurevych, Iryna}, | |
| booktitle = {Proceedings of EMNLP 2019}, | |
| year = {2019}, | |
| url = {https://arxiv.org/abs/1908.10084} | |
| } | |
| ``` | |
| ### Related Poster | |
| If you refer to this work, please also cite the related poster: | |
| * **HAL Id:** `hal-05681918` | |
| * **URL:** [https://hal.science/hal-05681918v1](https://hal.science/hal-05681918v1) |