File size: 5,696 Bytes
bf8f5f3 602100a bf8f5f3 602100a bf8f5f3 8161acb bf8f5f3 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 | ---
license: mit
base_model: facebook/esm2_t12_35M_UR50D
tags:
- protein
- structural-biology
- crystallography
- construct-design
- token-classification
pipeline_tag: token-classification
---
# TopPDBLX construct boundary model
**Give it a full-length protein sequence. It tells you, residue by residue, where a
crystallographer would have cut.**
Almost nobody crystallises the full-length gene product. You trim the flexible tails, drop a
disordered linker, and try the folded core. Choosing where to cut is judgement, usually made once
by hand from a disorder plot and an alignment.
The Protein Data Bank already holds **523,018 of those decisions**, each made by someone who then
got a structure. Every deposited chain records which stretch of the full-length protein was
cloned, and SIFTS maps it back residue by residue. This model learns from them.
## What it achieves
On 4,077 held-out proteins, split at 30% sequence identity so no homologue appears on both sides:
| Measure | Value | Why it matters |
|---|---|---|
| **Boundary error** | **9 residues** (median) | How far the predicted cut sits from the real one, on the 1,314 test proteins that were genuinely truncated |
| **MCC** | **0.669** | A model that says "keep everything" scores 0.00 |
| Accuracy | 85.7% | **Do not read alone.** 61.5% of residues really are inside a construct, so "keep everything" already scores 61.5% |
### Read the spread, not just the median
| | boundary error |
|---|---|
| Half the boundaries | within **5 residues** |
| Three quarters | within 56 residues |
| Nine tenths | within 250 residues |
**Excellent on most proteins, badly wrong on a minority.** 60% of boundaries land within 10
residues and 56% of proteins have *both* ends within 25. The mean of 72 residues is the tail.
**It half knows when it is wrong.** Mean predicted probability across the span runs 0.97 on good
predictions and 0.82 on bad ones. Gating at 0.85 covers 68% of proteins and lifts
both-ends-within-25 from 56% to **71%**. Treat a low-confidence span as a hint.
### A worked example
Hen lysozyme (`P00698`) is 147 residues, of which 1 to 18 are the signal peptide and 19 to 147 the
mature chain. Asked cold, the model proposes **19 to 147**. Nobody told it what a signal peptide
is; it learned that crystallographers do not clone them.
For scale: truncated constructs in this corpus trim a median of 62 residues from the N-terminus
and keep under half the chain.
## Usage
```python
import torch, json
from transformers import AutoModel, AutoTokenizer
repo = "Dellboy/toppdblx-construct-boundary"
tok = AutoTokenizer.from_pretrained(repo)
esm = AutoModel.from_pretrained(repo)
head = torch.nn.Sequential(torch.nn.Dropout(0.1),
torch.nn.Linear(esm.config.hidden_size, 1))
head.load_state_dict(torch.load("boundary_head.pt")) # hf_hub_download this file
esm.eval(); head.eval()
sequence = "MKTAYIAKQRQISFVKSHFSRQLEERLGLIEVQ..." # your full-length protein
enc = tok(sequence, return_tensors="pt", truncation=True, max_length=1024)
with torch.no_grad():
prob = torch.sigmoid(head(esm(**enc).last_hidden_state).squeeze(-1))[0, 1:-1]
keep = (prob >= 0.5).nonzero().flatten()
print(f"suggested construct: residues {keep[0].item()+1} to {keep[-1].item()+1}")
```
## Read this before trusting it
- **It is not a disorder predictor.** Residues that were cloned but never appeared in the density
count as *inside*. It predicts what was cloned, not what turned out to be ordered.
- **It only knows successes.** Every label comes from a construct that produced a crystal. It has
never seen one that failed, so it cannot say a boundary is bad, only that it is unlike the ones
that worked.
- **It errs towards keeping residues**, predicting inside for 69% where truth is 61.5%. Proposed
spans run slightly long. Trim rather than extend if choosing between it and your own judgement.
- **Sequences longer than 1,022 residues are truncated** by ESM-2's position limit.
## Why this model and not a bigger one
This checkpoint has been challenged five ways and none of them beat it:
| Attempted | Outcome on the held-out test split |
|---|---|
| Six structural features: coil-avoidance, ESMFold pLDDT, Pfam domain edges, disorder prediction, disorder as a retrained input channel, surface entropy | None improved the boundary |
| Soft targets (train on the fraction of a protein's constructs covering each residue) | No gain, applied everywhere or gated to well-deposited proteins |
| Six epochs instead of three | No gain |
| **ESM-2 t30-150M, four times the parameters** | **MCC 0.673 against 0.669: a dead heat, and worse on boundary error and every coverage@k** |
**The ceiling is in the labels, not the model.** Where a crystallographer cuts carries real signal
and this model extracts most of it, but the decision is also part convention, part whichever vector
was to hand, and part arbitrary. None of that is recoverable from sequence, so a larger model
simply fits the same ceiling more expensively.
## Training
ESM-2 t12-35M fine-tuned end to end, 3 epochs, AdamW with one-cycle LR at 3e-4, batches bucketed
by length into eight fixed widths. Labels are the per-residue consensus across every deposited
construct for that protein, counted as inside where at least half the constructs include the
residue. Validation MCC by epoch: 0.607, 0.681, 0.700, still improving at the end.
The failure condition (MCC below 0.40, or median boundary error worse than 20 residues) was
declared before the run rather than after.
Dataset: [TopPDBLX](https://doi.org/10.5281/zenodo.21807133) ·
Code: [bellcheddar/TopPDBLX](https://github.com/bellcheddar/TopPDBLX)
|