ProtSent-V2 ESM-C 300M

Contrastively fine-tuned ESM-C 300M producing fixed-length protein embeddings where biological similarity maps to embedding proximity. Intended for retrieval, clustering, and nearest-neighbour transfer.

This is the ProtSent V2 recipe applied to an ESM-C backbone. Continued training from the Synthyra/ESMplusplus_small (ESMC-300M) weights.

Usage

from sentence_transformers import SentenceTransformer
from sentence_transformers.util import cos_sim

model = SentenceTransformer("GrimSqueaker/ProtSent-V2-ESMC-300M", trust_remote_code=True)
emb = model.encode([
    "MKTLLLTLVVVTIVCLDLGYT",
    "MKTLLLTLVVVTIVCLDLGYN",
    "AGWYRSPQEGLKPVDTFKDIV",
])
print(cos_sim(emb[0], emb[1:]))

Embeddings are mean-pooled over the final layer, dimension 960. No Matryoshka heads on this release (see Training configuration).

Training data

Four sources. Pfam/AFDB/STRING are decontaminated against benchmark test sets (see below); the DMS source is not — see Limitations.

source pairs used
Pfam families 1,230,551
AlphaFold DB (Foldseek clusters) 27,105,179
STRING-DB v12 PPI 15,000,000
ProteinGym DMS (CoSENT target) 1,000,000
total 44,335,730

Pfam/AFDB pairs are drawn with the disjoint-pair sampler at up to 10 pairs per cluster (k=10), which spends the pair budget on distinct sequence pairs rather than repeating members — covers up to 90 distinct sequences per cluster instead of 10. STRING is a fixed 15M-pair subsample (seed 42).

Decontamination

Pfam, AlphaFold DB and STRING were searched against the benchmark test sequences with MMseqs2 easy-search (40% identity, 80% coverage, --cov-mode 1) and matching sequences removed before training — the same protsent-data-dc40 corpus used by the ProtSent-V2.5 ESM-2 releases.

Training configuration

setting value
backbone ESM-C 300M (Synthyra/ESMplusplus_small, 960 hidden)
loss CachedMultipleNegativesRankingLoss (Pfam/AFDB/STRING) + CoSENT (DMS)
contrastive batch 2048 per device
gather across devices off
multi-dataset sampler proportional, world-size-aligned (see below)
GOR weight 0 (off)
Matryoshka off
max sequence length 512
optimiser AdamW, LR 1e-4, warmup 300 steps
precision / attention bf16, flash-attention-2 (FA3 unavailable on this hardware)
hardware 3x NVIDIA B300
steps 7,216 (one epoch)
gradient-cache mini-batch 128

Training code: github.com/oriel9p/ProtSent, train_esmc_300m_v2.sh.

Results

23-task suite, -e test, seed 42, both kNN and linear-probe transfer. Full per-task CSVs: results/benchmarks/ism/protsent_esmc_300m_v2_{knn,linear}/ in the repo above.

SCOPe-40 structural retrieval, eligible queries only

method R@1 R@10 MAP
ESM-C 300M (this model's backbone, untrained) 0.3709 0.5794 0.2212
ISM-C 300M (structure-distilled ESM-C 300M) 0.4300 0.6592 0.2733
HMMER (phmmer, filters off) 0.7525 0.8978 0.6067
ProtSent-V2 ESM-C 300M 0.7974 0.9539 0.7692

23-task suite vs the model's own untrained backbone and vs ISM-C

Win/tie/loss, tie tolerance 0.005:

probe vs ESM-C 300M (this backbone, untrained) vs ISM-C 300M
kNN 16W / 0T / 7L 14W / 1T / 8L
linear 6W / 3T / 14L 6W / 4T / 13L

Citation

Paper: ProtSent: Protein Sentence Transformers

Downloads last month
11
Safetensors
Model size
0.3B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for GrimSqueaker/ProtSent-V2-ESMC-300M

Base model

biohub/ESMC-300M
Finetuned
(1)
this model

Datasets used to train GrimSqueaker/ProtSent-V2-ESMC-300M