Sentence Similarity
sentence-transformers
Safetensors
English
esm
protein
esm2
contrastive-learning
protein-embeddings
biology
Instructions to use GrimSqueaker/ProtSent-V2-150M with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- sentence-transformers
How to use GrimSqueaker/ProtSent-V2-150M with sentence-transformers:
from sentence_transformers import SentenceTransformer model = SentenceTransformer("GrimSqueaker/ProtSent-V2-150M") sentences = [ "That is a happy person", "That is a happy dog", "That is a very happy person", "Today is a sunny day" ] embeddings = model.encode(sentences) similarities = model.similarity(embeddings, embeddings) print(similarities.shape) # [4, 4] - Notebooks
- Google Colab
- Kaggle
| library_name: sentence-transformers | |
| tags: | |
| - sentence-transformers | |
| - protein | |
| - esm2 | |
| - contrastive-learning | |
| - protein-embeddings | |
| - biology | |
| license: mit | |
| language: | |
| - en | |
| pipeline_tag: sentence-similarity | |
| # ProtSent-V2 ESM-2 150M | |
| Contrastively fine-tuned [ESM-2 150M](https://huggingface.co/facebook/esm2_t30_150M_UR50D) producing fixed-length protein | |
| embeddings where biological similarity maps to embedding proximity. Intended for | |
| retrieval, clustering, and nearest-neighbour transfer. | |
| Retrained on a corpus decontaminated against the benchmark test sets. Predecessor: | |
| [oriel9p/protsent-esm2-150M](https://huggingface.co/oriel9p/protsent-esm2-150M). | |
| Other scale: [GrimSqueaker/ProtSent-V2-35M](https://huggingface.co/GrimSqueaker/ProtSent-V2-35M). | |
| ## Usage | |
| ```python | |
| from sentence_transformers import SentenceTransformer | |
| from sentence_transformers.util import cos_sim | |
| model = SentenceTransformer("GrimSqueaker/ProtSent-V2-150M") | |
| emb = model.encode([ | |
| "MKTLLLTLVVVTIVCLDLGYT", | |
| "MKTLLLTLVVVTIVCLDLGYN", | |
| "AGWYRSPQEGLKPVDTFKDIV", | |
| ]) | |
| print(cos_sim(emb[0], emb[1:])) | |
| ``` | |
| Embeddings are mean-pooled over the final layer, dimension 640. Matryoshka heads at | |
| 64/128/256 are available by truncating the embedding. | |
| ## Training data | |
| Three sources, all decontaminated (see below). **No DMS/ProteinGym component** — unlike | |
| the V1 release, which included ProteinGym DMS pairs under a CoSENT loss. | |
| | source | pairs used | | |
| |---|---:| | |
| | Pfam families | 284,683 | | |
| | AlphaFold DB (Foldseek clusters) | 8,612,331 | | |
| | STRING-DB v12 PPI | 15,000,000 | | |
| | **total** | **23,897,014** | | |
| Pfam and AFDB pairs are sampled within clusters; STRING is a fixed 15M-pair subsample | |
| (seed 42) of the filtered pair table. | |
| ## Decontamination | |
| Every source was searched against the benchmark test sequences with MMseqs2 `easy-search` | |
| (corpus as query, 40% identity, 80% coverage, `--cov-mode 1`) and matching sequences | |
| removed before training. | |
| | corpus | rows before | rows after | removed | | |
| |---|---:|---:|---:| | |
| | Pfam | 28,530,684 | 27,929,772 | 2.11% | | |
| | AlphaFold DB | 135,404,259 | 126,301,607 | 6.72% | | |
| | STRING | 76,070,154 | 71,891,417 | 5.49% | | |
| Filter targets: `biomap-research/fold_prediction` (remote homology) and | |
| `Synthyra/bernett_gold_ppi` (PPI) test splits. The result was verified by semi-joining each | |
| training file against the removal lists: zero flagged sequences remained. | |
| SCOPe-40 was not a filter target — it has no train/test split, so filtering against it | |
| would remove nearly all domain sequences from the corpus. | |
| ## Training configuration | |
| | setting | value | | |
| |---|---| | |
| | backbone | ESM-2 150M (640 hidden, 30 layers) | | |
| | loss | CachedMultipleNegativesRankingLoss | | |
| | contrastive batch | 1024 per device | | |
| | gather across devices | off | | |
| | synthetic hard negatives | off | | |
| | multi-dataset sampler | proportional | | |
| | Matryoshka dims | 64 / 128 / 256 | | |
| | max sequence length | 512 | | |
| | optimiser | AdamW, LR 2e-4, `cosine_with_min_lr` | | |
| | precision / attention | bf16, flash-attention-2 | | |
| | hardware | 6x NVIDIA B300 | | |
| | steps | 3,890 (one epoch) | | |
| | gradient-cache mini-batch | 64 | | |
| | warmup | 300 steps | | |
| Training code: [github.com/oriel9p/ProtSent](https://github.com/oriel9p/ProtSent), | |
| `train_esm2_150m.sh`. | |
| ## Results | |
| SCOPe-40 structural retrieval, test split, self excluded, no-hit queries counted as | |
| failures. Restricted to the 1,693 of 2,207 queries that have a non-self same-family protein | |
| in the gallery. | |
| | method | R@1 | R@10 | MAP | | |
| |---|---:|---:|---:| | |
| | ESM-2 150M | 0.5535 | 0.7702 | 0.4236 | | |
| | MMseqs2 (`-s 7.5`) | 0.6556 | 0.7401 | 0.4098 | | |
| | HMMER (phmmer, `-E 10`, default filters) | 0.6970 | 0.7809 | 0.4747 | | |
| | HMMER (phmmer, filters off — strongest) | **0.7525** | 0.8978 | 0.6067 | | |
| | ProtSent-V1 150M | 0.6615 | 0.8943 | 0.6431 | | |
| | **ProtSent-V2 150M** | 0.7431 | **0.9368** | **0.7042** | | |
| Paired bootstrap over queries (10,000 resamples): V2 − V1 R@1 +0.0809 [+0.0602, +0.1022]; | |
| V2 − MMseqs2 R@1 +0.0868 [+0.0620, +0.1116], both excluding zero. | |
| **Against a maximally sensitive profile search, top-1 is not ours.** With filters off, | |
| phmmer reaches R@1 0.7525 against this model's 0.7431. The embedding advantage is in | |
| ranking depth and MAP (R@10 0.9368 vs 0.8978, MAP 0.7042 vs 0.6067), at one forward pass | |
| per sequence and with indexable sub-linear search, rather than an all-vs-all profile | |
| comparison. | |
| Remote homology (the task the corpus was filtered against), test split: | |
| | model | 3-NN accuracy | linear-probe accuracy | | |
| |---|---:|---:| | |
| | ESM-2 150M | 0.5194 | 0.7500 | | |
| | ProtSent-V1 150M | 0.7047 | 0.7401 | | |
| | **ProtSent-V2 150M** | 0.6612 | **0.7503** | | |
| Under 3-NN this model scores below V1 here. V1 was trained on a corpus containing sequences | |
| at ≥40% identity to this test set; removing them removed that advantage. Under a linear | |
| probe the ordering reverses. | |
| ## Limitations | |
| - Under a **trained linear probe** on the final layer, this model is roughly neutral to | |
| slightly worse than the stock ESM-2 backbone across a 23-task suite. The advantage is in | |
| nearest-neighbour geometry, not in information a trained readout can extract. | |
| - The final layer is not the best pooling layer for property prediction. In a layer sweep | |
| on remote homology, an intermediate layer (~2/3 depth) scored higher for every model | |
| tested, including the stock backbone. | |
| - V2 differs from V1 in more than decontamination: no hard negatives, proportional | |
| sampling, no DMS source, larger effective batch. It is not a controlled ablation of | |
| filtering alone. | |
| - Only the remote-homology and PPI test sets were decontamination targets; other benchmark | |
| test sets were not filtered against. | |
| ## Citation | |
| Paper: [ProtSent: Protein Sentence Transformers](https://doi.org/10.48550/arXiv.2605.06830) | |