abSFM / README.md
Reddy-BIIE-ETHZ's picture
Model card for public release
cab20dd verified
|
Raw
History Blame Contribute Delete
4.5 kB
---
license: other
tags:
- protein
- antibody
- immunology
- foundation-model
- antibody-antigen
---
# abSFM v1.0 — Antibody–Antigen Specificity Foundation Model
**Reddy Lab — Lab for Systems & Synthetic Immunology — ETH Zurich / Botnar Institute of Immune
Engineering (BIIE), Basel**
abSFM scores **antibody–antigen specificity**. Given an antibody (VH/VL) and an antigen sequence it
returns a binding-compatibility score and localises the paratope and the epitope — from sequence,
without building a structure. Built on the **binderSFM v2.5** protein–protein foundation (two frozen
ESM-2 towers fused by a trainable cross-attention encoder).
Code, configs, evaluation and audit records: **[github.com/Reddy-BIIE-ETHZ/abSFM](https://github.com/Reddy-BIIE-ETHZ/abSFM)**
> **Preprint-stage.** The manuscript is in preparation; the orthogonal AI audit and the wet-lab
> validation of the SFM-display panel are still outstanding. Treat the numbers below accordingly.
## This release
| | |
|---|---|
| File | `abSFM-v1.0.pth` (55.9 MB) |
| **SHA-256** | `8c732efb5c442ccb5349c0587f9d3619d09289c87fa35c057db9dcee7c61b3b7` |
| Trainable parameters | 13.96 M (cross-attention encoder + 3 heads; ESM-2 frozen) |
| Training corpus | 8,499 Ab–Ag pairs, 1,213 antigen clusters (AbSD + SAbDab2 + IEDB), public data only |
| Training cost | ~2 h on a single GPU, 10 epochs |
Configs ship alongside the weights (`config_model.yaml`, `config_data.yaml`, `config_train.yaml`),
so the checkpoint loads without the code repo.
Two other repositories are **not** this model: `SFM-BIIE-ETHZ/abSFM-dev-archive` holds superseded
development checkpoints, and `Reddy-BIIE-ETHZ/abSFM` is an older personal-namespace copy. Verify the
hash.
## Quick start
```bash
hf download SFM-BIIE-ETHZ/abSFM abSFM-v1.0.pth --local-dir ./weights
shasum -a 256 weights/abSFM-v1.0.pth # must match the SHA-256 above
git clone https://github.com/Reddy-BIIE-ETHZ/abSFM.git && cd abSFM && pip install -e .
python scripts/absfm/score_pairs.py \
--antibodies my_abs.tsv --antigens my_antigens.fasta \
--checkpoint weights/abSFM-v1.0.pth --out scores.tsv
```
## Reading the score — read this before ranking anything
- **Not comparable across antigens.** Each antigen carries its own offset; whole-library median
cosine ranged from **−0.04 (IL-2) to +0.23 (HA)**. Rank *within* an antigen, or convert to a
percentile against a fixed background library. An absolute cutoff across targets is meaningless.
- **Don't take only the extreme tip.** Across 7 antigens and ~21,000 genuinely held-out known
binders, no antigen had a validated binder in its **top 10** of 2.59 M. Enrichment peaks around
ranks **50–500** and is flat there — sample a band.
- **No headroom above real binders.** For every antigen tested, the best validated binder scored
*below* the median of our selected candidates. Pushing the score higher is extrapolation.
- **Does not predict structural-verifier confidence** (Spearman ρ ≈ −0.03 vs Boltz-2 iPTM). Use
abSFM and a co-folder as independent filters, never combined into one ranking.
- **Construct-specific.** On full-length dengue E the best known DENV2 binder ranks 864,708 of
2.59 M; on the EDIII domain it ranks **1**. Score the construct you will actually express, and
prefer the epitope-bearing domain over a full ectodomain.
## Performance (held-out)
Whole antibody-clonotype and antigen clusters are held out, then graded by novelty. R@10, antibody →
antigen against a 507-antigen pool (chance 2.0%): **58.0%** Class A (pair seen, memorisation
ceiling), **14.8%** B, **20.6%** C (one partner novel), **8.0%** D (both novel), **15%** overall.
Reverse direction, pool 678 (chance 1.5%): 12% overall. Paratope AUROC 0.950, epitope 0.746.
## Known limitations
- Known-binder scores are inflated by memorisation (10×–2,500×, antigen-specific); any calibration on
known binders must exclude training pairs or it measures recall.
- Large multi-domain antigens score and verify poorly — truncate to the epitope-bearing domain.
Oligomerising the antigen does **not** help and generally hurts.
- The contact head is ~76% antigen-determined, so the predicted epitope is not usable as a binder
filter on its own.
- Public training data only; behaviour on antigen classes absent from SAbDab/AbSD/IEDB is unknown.
## Citation
> Reddy ST. *An Antibody–Antigen Specificity Foundation Model for Discovery and Engineering.*
> In preparation (2026).