Instructions to use GrimSqueaker/ProtSent-V2-ESMC-300M with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- sentence-transformers
How to use GrimSqueaker/ProtSent-V2-ESMC-300M with sentence-transformers:
from sentence_transformers import SentenceTransformer model = SentenceTransformer("GrimSqueaker/ProtSent-V2-ESMC-300M", trust_remote_code=True) sentences = [ "That is a happy person", "That is a happy dog", "That is a very happy person", "Today is a sunny day" ] embeddings = model.encode(sentences) similarities = model.similarity(embeddings, embeddings) print(similarities.shape) # [4, 4] - Notebooks
- Google Colab
- Kaggle
ProtSent-V2 ESM-C 300M
Contrastively fine-tuned ESM-C 300M producing fixed-length protein embeddings where biological similarity maps to embedding proximity. Intended for retrieval, clustering, and nearest-neighbour transfer.
This is the ProtSent V2 recipe applied to an ESM-C backbone.
Continued training from the Synthyra/ESMplusplus_small (ESMC-300M) weights.
Usage
from sentence_transformers import SentenceTransformer
from sentence_transformers.util import cos_sim
model = SentenceTransformer("GrimSqueaker/ProtSent-V2-ESMC-300M", trust_remote_code=True)
emb = model.encode([
"MKTLLLTLVVVTIVCLDLGYT",
"MKTLLLTLVVVTIVCLDLGYN",
"AGWYRSPQEGLKPVDTFKDIV",
])
print(cos_sim(emb[0], emb[1:]))
Embeddings are mean-pooled over the final layer, dimension 960. No Matryoshka heads on this release (see Training configuration).
Training data
Four sources. Pfam/AFDB/STRING are decontaminated against benchmark test sets (see below); the DMS source is not — see Limitations.
| source | pairs used |
|---|---|
| Pfam families | 1,230,551 |
| AlphaFold DB (Foldseek clusters) | 27,105,179 |
| STRING-DB v12 PPI | 15,000,000 |
| ProteinGym DMS (CoSENT target) | 1,000,000 |
| total | 44,335,730 |
Pfam/AFDB pairs are drawn with the disjoint-pair sampler at up to 10 pairs per cluster
(k=10), which spends the pair budget on distinct sequence pairs rather than repeating
members — covers up to 90 distinct sequences per cluster instead of 10. STRING is a fixed
15M-pair subsample (seed 42).
Decontamination
Pfam, AlphaFold DB and STRING were searched against the benchmark test sequences with
MMseqs2 easy-search (40% identity, 80% coverage, --cov-mode 1) and matching sequences
removed before training — the same protsent-data-dc40 corpus used by the ProtSent-V2.5
ESM-2 releases.
Training configuration
| setting | value |
|---|---|
| backbone | ESM-C 300M (Synthyra/ESMplusplus_small, 960 hidden) |
| loss | CachedMultipleNegativesRankingLoss (Pfam/AFDB/STRING) + CoSENT (DMS) |
| contrastive batch | 2048 per device |
| gather across devices | off |
| multi-dataset sampler | proportional, world-size-aligned (see below) |
| GOR weight | 0 (off) |
| Matryoshka | off |
| max sequence length | 512 |
| optimiser | AdamW, LR 1e-4, warmup 300 steps |
| precision / attention | bf16, flash-attention-2 (FA3 unavailable on this hardware) |
| hardware | 3x NVIDIA B300 |
| steps | 7,216 (one epoch) |
| gradient-cache mini-batch | 128 |
Training code: github.com/oriel9p/ProtSent,
train_esmc_300m_v2.sh.
Results
23-task suite, -e test, seed 42, both kNN and linear-probe transfer. Full per-task CSVs:
results/benchmarks/ism/protsent_esmc_300m_v2_{knn,linear}/ in the repo above.
SCOPe-40 structural retrieval, eligible queries only
| method | R@1 | R@10 | MAP |
|---|---|---|---|
| ESM-C 300M (this model's backbone, untrained) | 0.3709 | 0.5794 | 0.2212 |
| ISM-C 300M (structure-distilled ESM-C 300M) | 0.4300 | 0.6592 | 0.2733 |
| HMMER (phmmer, filters off) | 0.7525 | 0.8978 | 0.6067 |
| ProtSent-V2 ESM-C 300M | 0.7974 | 0.9539 | 0.7692 |
23-task suite vs the model's own untrained backbone and vs ISM-C
Win/tie/loss, tie tolerance 0.005:
| probe | vs ESM-C 300M (this backbone, untrained) | vs ISM-C 300M |
|---|---|---|
| kNN | 16W / 0T / 7L | 14W / 1T / 8L |
| linear | 6W / 3T / 14L | 6W / 4T / 13L |
Citation
- Downloads last month
- 11
Model tree for GrimSqueaker/ProtSent-V2-ESMC-300M
Base model
biohub/ESMC-300M