Sentence Similarity
sentence-transformers
Safetensors
English
ESMplusplus
protein
esm-c
contrastive-learning
protein-embeddings
biology
custom_code
Instructions to use GrimSqueaker/ProtSent-V2-ESMC-300M with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- sentence-transformers
How to use GrimSqueaker/ProtSent-V2-ESMC-300M with sentence-transformers:
from sentence_transformers import SentenceTransformer model = SentenceTransformer("GrimSqueaker/ProtSent-V2-ESMC-300M", trust_remote_code=True) sentences = [ "That is a happy person", "That is a happy dog", "That is a very happy person", "Today is a sunny day" ] embeddings = model.encode(sentences) similarities = model.similarity(embeddings, embeddings) print(similarities.shape) # [4, 4] - Notebooks
- Google Colab
- Kaggle
| library_name: sentence-transformers | |
| tags: | |
| - sentence-transformers | |
| - protein | |
| - esm-c | |
| - contrastive-learning | |
| - protein-embeddings | |
| - biology | |
| license: mit | |
| language: | |
| - en | |
| pipeline_tag: sentence-similarity | |
| datasets: | |
| - fredzzp/afdbseq | |
| - OATML-Markslab/ProteinGym_v1 | |
| - Synthyra/StringDBSeqsv12 | |
| base_model: | |
| - biohub/ESMC-300M | |
| # ProtSent-V2 ESM-C 300M | |
| Contrastively fine-tuned [ESM-C 300M](https://huggingface.co/Synthyra/ESMplusplus_small) | |
| producing fixed-length protein embeddings where biological similarity maps to embedding | |
| proximity. Intended for retrieval, clustering, and nearest-neighbour transfer. | |
| This is the ProtSent V2 recipe applied to an ESM-C backbone. | |
| Continued training from the `Synthyra/ESMplusplus_small` (ESMC-300M) weights. | |
| ## Usage | |
| ```python | |
| from sentence_transformers import SentenceTransformer | |
| from sentence_transformers.util import cos_sim | |
| model = SentenceTransformer("GrimSqueaker/ProtSent-V2-ESMC-300M", trust_remote_code=True) | |
| emb = model.encode([ | |
| "MKTLLLTLVVVTIVCLDLGYT", | |
| "MKTLLLTLVVVTIVCLDLGYN", | |
| "AGWYRSPQEGLKPVDTFKDIV", | |
| ]) | |
| print(cos_sim(emb[0], emb[1:])) | |
| ``` | |
| Embeddings are mean-pooled over the final layer, dimension 960. No Matryoshka heads on | |
| this release (see Training configuration). | |
| ## Training data | |
| Four sources. Pfam/AFDB/STRING are decontaminated against benchmark test sets (see | |
| below); the DMS source is not — see Limitations. | |
| | source | pairs used | | |
| |---|---:| | |
| | Pfam families | 1,230,551 | | |
| | AlphaFold DB (Foldseek clusters) | 27,105,179 | | |
| | STRING-DB v12 PPI | 15,000,000 | | |
| | ProteinGym DMS (CoSENT target) | 1,000,000 | | |
| | **total** | **44,335,730** | | |
| Pfam/AFDB pairs are drawn with the disjoint-pair sampler at up to 10 pairs per cluster | |
| (`k=10`), which spends the pair budget on distinct sequence pairs rather than repeating | |
| members — covers up to 90 distinct sequences per cluster instead of 10. STRING is a fixed | |
| 15M-pair subsample (seed 42). | |
| ## Decontamination | |
| Pfam, AlphaFold DB and STRING were searched against the benchmark test sequences with | |
| MMseqs2 `easy-search` (40% identity, 80% coverage, `--cov-mode 1`) and matching sequences | |
| removed before training — the same `protsent-data-dc40` corpus used by the ProtSent-V2.5 | |
| ESM-2 releases. | |
| ## Training configuration | |
| | setting | value | | |
| |---|---| | |
| | backbone | ESM-C 300M (`Synthyra/ESMplusplus_small`, 960 hidden) | | |
| | loss | CachedMultipleNegativesRankingLoss (Pfam/AFDB/STRING) + CoSENT (DMS) | | |
| | contrastive batch | 2048 per device | | |
| | gather across devices | off | | |
| | multi-dataset sampler | proportional, world-size-aligned (see below) | | |
| | GOR weight | 0 (off) | | |
| | Matryoshka | off | | |
| | max sequence length | 512 | | |
| | optimiser | AdamW, LR 1e-4, warmup 300 steps | | |
| | precision / attention | bf16, flash-attention-2 (FA3 unavailable on this hardware) | | |
| | hardware | 3x NVIDIA B300 | | |
| | steps | 7,216 (one epoch) | | |
| | gradient-cache mini-batch | 128 | | |
| Training code: [github.com/oriel9p/ProtSent](https://github.com/oriel9p/ProtSent), | |
| `train_esmc_300m_v2.sh`. | |
| ## Results | |
| 23-task suite, `-e test`, seed 42, both kNN and linear-probe transfer. Full per-task CSVs: | |
| `results/benchmarks/ism/protsent_esmc_300m_v2_{knn,linear}/` in the repo above. | |
| ### SCOPe-40 structural retrieval, eligible queries only | |
| | method | R@1 | R@10 | MAP | | |
| |---|---:|---:|---:| | |
| | ESM-C 300M (this model's backbone, untrained) | 0.3709 | 0.5794 | 0.2212 | | |
| | ISM-C 300M (structure-distilled ESM-C 300M) | 0.4300 | 0.6592 | 0.2733 | | |
| | HMMER (phmmer, filters off) | 0.7525 | 0.8978 | 0.6067 | | |
| | **ProtSent-V2 ESM-C 300M** | **0.7974** | **0.9539** | **0.7692** | | |
| ### 23-task suite vs the model's own untrained backbone and vs ISM-C | |
| Win/tie/loss, tie tolerance 0.005: | |
| | probe | vs ESM-C 300M (this backbone, untrained) | vs ISM-C 300M | | |
| |---|---|---| | |
| | kNN | 16W / 0T / 7L | 14W / 1T / 8L | | |
| | linear | 6W / 3T / 14L | 6W / 4T / 13L | | |
| ## Citation | |
| Paper: [ProtSent: Protein Sentence Transformers](https://doi.org/10.48550/arXiv.2605.06830) |