--- license: other tags: - protein - antibody - immunology - foundation-model - antibody-antigen --- # abSFM v1.0 — Antibody–Antigen Specificity Foundation Model **Reddy Lab — Lab for Systems & Synthetic Immunology — ETH Zurich / Botnar Institute of Immune Engineering (BIIE), Basel** abSFM scores **antibody–antigen specificity**. Given an antibody (VH/VL) and an antigen sequence it returns a binding-compatibility score and localises the paratope and the epitope — from sequence, without building a structure. Built on the **binderSFM v2.5** protein–protein foundation (two frozen ESM-2 towers fused by a trainable cross-attention encoder). Code, configs, evaluation and audit records: **[github.com/Reddy-BIIE-ETHZ/abSFM](https://github.com/Reddy-BIIE-ETHZ/abSFM)** > **Preprint-stage.** The manuscript is in preparation; the orthogonal AI audit and the wet-lab > validation of the SFM-display panel are still outstanding. Treat the numbers below accordingly. ## This release | | | |---|---| | File | `abSFM-v1.0.pth` (55.9 MB) | | **SHA-256** | `8c732efb5c442ccb5349c0587f9d3619d09289c87fa35c057db9dcee7c61b3b7` | | Trainable parameters | 13.96 M (cross-attention encoder + 3 heads; ESM-2 frozen) | | Training corpus | 8,499 Ab–Ag pairs, 1,213 antigen clusters (AbSD + SAbDab2 + IEDB), public data only | | Training cost | ~2 h on a single GPU, 10 epochs | Configs ship alongside the weights (`config_model.yaml`, `config_data.yaml`, `config_train.yaml`), so the checkpoint loads without the code repo. Two other repositories are **not** this model: `SFM-BIIE-ETHZ/abSFM-dev-archive` holds superseded development checkpoints, and `Reddy-BIIE-ETHZ/abSFM` is an older personal-namespace copy. Verify the hash. ## Quick start ```bash hf download SFM-BIIE-ETHZ/abSFM abSFM-v1.0.pth --local-dir ./weights shasum -a 256 weights/abSFM-v1.0.pth # must match the SHA-256 above git clone https://github.com/Reddy-BIIE-ETHZ/abSFM.git && cd abSFM && pip install -e . python scripts/absfm/score_pairs.py \ --antibodies my_abs.tsv --antigens my_antigens.fasta \ --checkpoint weights/abSFM-v1.0.pth --out scores.tsv ``` ## Reading the score — read this before ranking anything - **Not comparable across antigens.** Each antigen carries its own offset; whole-library median cosine ranged from **−0.04 (IL-2) to +0.23 (HA)**. Rank *within* an antigen, or convert to a percentile against a fixed background library. An absolute cutoff across targets is meaningless. - **Don't take only the extreme tip.** Across 7 antigens and ~21,000 genuinely held-out known binders, no antigen had a validated binder in its **top 10** of 2.59 M. Enrichment peaks around ranks **50–500** and is flat there — sample a band. - **No headroom above real binders.** For every antigen tested, the best validated binder scored *below* the median of our selected candidates. Pushing the score higher is extrapolation. - **Does not predict structural-verifier confidence** (Spearman ρ ≈ −0.03 vs Boltz-2 iPTM). Use abSFM and a co-folder as independent filters, never combined into one ranking. - **Construct-specific.** On full-length dengue E the best known DENV2 binder ranks 864,708 of 2.59 M; on the EDIII domain it ranks **1**. Score the construct you will actually express, and prefer the epitope-bearing domain over a full ectodomain. ## Performance (held-out) Whole antibody-clonotype and antigen clusters are held out, then graded by novelty. R@10, antibody → antigen against a 507-antigen pool (chance 2.0%): **58.0%** Class A (pair seen, memorisation ceiling), **14.8%** B, **20.6%** C (one partner novel), **8.0%** D (both novel), **15%** overall. Reverse direction, pool 678 (chance 1.5%): 12% overall. Paratope AUROC 0.950, epitope 0.746. ## Known limitations - Known-binder scores are inflated by memorisation (10×–2,500×, antigen-specific); any calibration on known binders must exclude training pairs or it measures recall. - Large multi-domain antigens score and verify poorly — truncate to the epitope-bearing domain. Oligomerising the antigen does **not** help and generally hurts. - The contact head is ~76% antigen-determined, so the predicted epitope is not usable as a binder filter on its own. - Public training data only; behaviour on antigen classes absent from SAbDab/AbSD/IEDB is unknown. ## Citation > Reddy ST. *An Antibody–Antigen Specificity Foundation Model for Discovery and Engineering.* > In preparation (2026).