Sentence Similarity
sentence-transformers
Safetensors
English
esm
protein
esm2
contrastive-learning
protein-embeddings
biology
Instructions to use GrimSqueaker/ProtSent-V2-150M with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- sentence-transformers
How to use GrimSqueaker/ProtSent-V2-150M with sentence-transformers:
from sentence_transformers import SentenceTransformer model = SentenceTransformer("GrimSqueaker/ProtSent-V2-150M") sentences = [ "That is a happy person", "That is a happy dog", "That is a very happy person", "Today is a sunny day" ] embeddings = model.encode(sentences) similarities = model.similarity(embeddings, embeddings) print(similarities.shape) # [4, 4] - Notebooks
- Google Colab
- Kaggle
File size: 5,756 Bytes
38ce085 d8710fa 38ce085 d8710fa 38ce085 d8710fa 38ce085 d8710fa 38ce085 d8710fa 38ce085 d8710fa 38ce085 d8710fa 38ce085 efb72b0 38ce085 efb72b0 38ce085 efb72b0 38ce085 d8710fa 38ce085 d8710fa 38ce085 d8710fa 38ce085 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 | ---
library_name: sentence-transformers
tags:
- sentence-transformers
- protein
- esm2
- contrastive-learning
- protein-embeddings
- biology
license: mit
language:
- en
pipeline_tag: sentence-similarity
---
# ProtSent-V2 ESM-2 150M
Contrastively fine-tuned [ESM-2 150M](https://huggingface.co/facebook/esm2_t30_150M_UR50D) producing fixed-length protein
embeddings where biological similarity maps to embedding proximity. Intended for
retrieval, clustering, and nearest-neighbour transfer.
Retrained on a corpus decontaminated against the benchmark test sets. Predecessor:
[oriel9p/protsent-esm2-150M](https://huggingface.co/oriel9p/protsent-esm2-150M).
Other scale: [GrimSqueaker/ProtSent-V2-35M](https://huggingface.co/GrimSqueaker/ProtSent-V2-35M).
## Usage
```python
from sentence_transformers import SentenceTransformer
from sentence_transformers.util import cos_sim
model = SentenceTransformer("GrimSqueaker/ProtSent-V2-150M")
emb = model.encode([
"MKTLLLTLVVVTIVCLDLGYT",
"MKTLLLTLVVVTIVCLDLGYN",
"AGWYRSPQEGLKPVDTFKDIV",
])
print(cos_sim(emb[0], emb[1:]))
```
Embeddings are mean-pooled over the final layer, dimension 640. Matryoshka heads at
64/128/256 are available by truncating the embedding.
## Training data
Three sources, all decontaminated (see below). **No DMS/ProteinGym component** — unlike
the V1 release, which included ProteinGym DMS pairs under a CoSENT loss.
| source | pairs used |
|---|---:|
| Pfam families | 284,683 |
| AlphaFold DB (Foldseek clusters) | 8,612,331 |
| STRING-DB v12 PPI | 15,000,000 |
| **total** | **23,897,014** |
Pfam and AFDB pairs are sampled within clusters; STRING is a fixed 15M-pair subsample
(seed 42) of the filtered pair table.
## Decontamination
Every source was searched against the benchmark test sequences with MMseqs2 `easy-search`
(corpus as query, 40% identity, 80% coverage, `--cov-mode 1`) and matching sequences
removed before training.
| corpus | rows before | rows after | removed |
|---|---:|---:|---:|
| Pfam | 28,530,684 | 27,929,772 | 2.11% |
| AlphaFold DB | 135,404,259 | 126,301,607 | 6.72% |
| STRING | 76,070,154 | 71,891,417 | 5.49% |
Filter targets: `biomap-research/fold_prediction` (remote homology) and
`Synthyra/bernett_gold_ppi` (PPI) test splits. The result was verified by semi-joining each
training file against the removal lists: zero flagged sequences remained.
SCOPe-40 was not a filter target — it has no train/test split, so filtering against it
would remove nearly all domain sequences from the corpus.
## Training configuration
| setting | value |
|---|---|
| backbone | ESM-2 150M (640 hidden, 30 layers) |
| loss | CachedMultipleNegativesRankingLoss |
| contrastive batch | 1024 per device |
| gather across devices | off |
| synthetic hard negatives | off |
| multi-dataset sampler | proportional |
| Matryoshka dims | 64 / 128 / 256 |
| max sequence length | 512 |
| optimiser | AdamW, LR 2e-4, `cosine_with_min_lr` |
| precision / attention | bf16, flash-attention-2 |
| hardware | 6x NVIDIA B300 |
| steps | 3,890 (one epoch) |
| gradient-cache mini-batch | 64 |
| warmup | 300 steps |
Training code: [github.com/oriel9p/ProtSent](https://github.com/oriel9p/ProtSent),
`train_esm2_150m.sh`.
## Results
SCOPe-40 structural retrieval, test split, self excluded, no-hit queries counted as
failures. Restricted to the 1,693 of 2,207 queries that have a non-self same-family protein
in the gallery.
| method | R@1 | R@10 | MAP |
|---|---:|---:|---:|
| ESM-2 150M | 0.5535 | 0.7702 | 0.4236 |
| MMseqs2 (`-s 7.5`) | 0.6556 | 0.7401 | 0.4098 |
| HMMER (phmmer, `-E 10`, default filters) | 0.6970 | 0.7809 | 0.4747 |
| HMMER (phmmer, filters off — strongest) | **0.7525** | 0.8978 | 0.6067 |
| ProtSent-V1 150M | 0.6615 | 0.8943 | 0.6431 |
| **ProtSent-V2 150M** | 0.7431 | **0.9368** | **0.7042** |
Paired bootstrap over queries (10,000 resamples): V2 − V1 R@1 +0.0809 [+0.0602, +0.1022];
V2 − MMseqs2 R@1 +0.0868 [+0.0620, +0.1116], both excluding zero.
**Against a maximally sensitive profile search, top-1 is not ours.** With filters off,
phmmer reaches R@1 0.7525 against this model's 0.7431. The embedding advantage is in
ranking depth and MAP (R@10 0.9368 vs 0.8978, MAP 0.7042 vs 0.6067), at one forward pass
per sequence and with indexable sub-linear search, rather than an all-vs-all profile
comparison.
Remote homology (the task the corpus was filtered against), test split:
| model | 3-NN accuracy | linear-probe accuracy |
|---|---:|---:|
| ESM-2 150M | 0.5194 | 0.7500 |
| ProtSent-V1 150M | 0.7047 | 0.7401 |
| **ProtSent-V2 150M** | 0.6612 | **0.7503** |
Under 3-NN this model scores below V1 here. V1 was trained on a corpus containing sequences
at ≥40% identity to this test set; removing them removed that advantage. Under a linear
probe the ordering reverses.
## Limitations
- Under a **trained linear probe** on the final layer, this model is roughly neutral to
slightly worse than the stock ESM-2 backbone across a 23-task suite. The advantage is in
nearest-neighbour geometry, not in information a trained readout can extract.
- The final layer is not the best pooling layer for property prediction. In a layer sweep
on remote homology, an intermediate layer (~2/3 depth) scored higher for every model
tested, including the stock backbone.
- V2 differs from V1 in more than decontamination: no hard negatives, proportional
sampling, no DMS source, larger effective batch. It is not a controlled ablation of
filtering alone.
- Only the remote-homology and PPI test sets were decontamination targets; other benchmark
test sets were not filtered against.
## Citation
Paper: [ProtSent: Protein Sentence Transformers](https://doi.org/10.48550/arXiv.2605.06830)
|