SF-Cluster / bench /configs /cases.yaml
chq1155's picture
Add benchmark reproduction: CPU scoring/eval (evaluate_prediction+batch_eval), reference structures, region manifests, headline prediction sets, reproduce_benchmark.py (reproduces main minority_hit_rate table; 15/15 cells verified by re-scoring 1440 PDBs)
f4e8048 verified
Raw
History Blame Contribute Delete
17.3 kB
cases:
- case_name: KaiB
tier: 1
classification: known-state
query_source_db: UniProt
query_accession: Q79V61
query_organism: Thermosynechococcus vestitus BP-1 (a.k.a. T. elongatus)
canonical_query_sequence: RKTYVLKLYVAGNTPNSVRALKTLNNILEKEFKGVYALKVIDVLKNPQLAEEDKILATPTLAKVLPPPVRRIIGDLSNREKVLIGLDLLYE
canonical_query_length: 91
canonical_query_frame: 2QKE chain B residue numbering, residues 5..95 (matches AF-Cluster
repo 91-aa KaiB_TE query). Protocol §2.1 initially listed 108 the AF-Cluster
repo (data_sep2022/00_KaiB/2qkeE.pdb + 2QKEE_colabfold.a3m) uses this 91-aa trimmed
construct. [V5 resolved; repo pin trumps the 108-aa paper text.]
construct_start: 5
construct_end: 95
states:
- name: state_A_ground
pdb_id: 2QKE
chain_id: B
residue_range: 5..95 (of 1..108 crystal residues)
has_mutations: false
notes: Ground-state βαββααβ KaiBTE. 2QKE has 6 chains (A-F); chain B is the only
complete 1..108 monomer (all others have missing terminal residues). Trimmed
to 5..95 to match AF-Cluster repo working construct.
qc:
method: x-ray diffraction
resolution_A: 2.7
n_models: 1
chains_all:
- A
- B
- C
- D
- E
- F
chain_used: B
chains_dropped:
- A
- C
- D
- E
- F
missing_residues: []
mutations_vs_canonical: []
b_factor_mean: 97.64
heteroatoms_removed:
- HOH
- name: state_B_foldswitch
pdb_id: 5JYT
chain_id: A
residue_range: 5..95 (of 1..106 crystal residues)
has_mutations: true
notes: 'FS-state thioredoxin-like (βαβαββα). 5JYT is a stabilized KaiBTE variant:
point mutations Y8A, N29A, G89A, D91R, Y94A (per RCSB REMARKs & paper p.839)
plus C-terminal tags. Trimmed to 5..95 to exclude the MAPL N-tag and the DYKDDDDK
FLAG tag.'
qc:
method: solution nmr
resolution_A: null
n_models: 20
chains_all:
- A
chain_used: A
chains_dropped: []
missing_residues:
- 107
- 108
mutations_vs_canonical:
- - 8
- Y
- A
- - 29
- N
- A
- - 89
- G
- A
- - 91
- D
- R
- - 94
- Y
- A
- - 100
- Q
- Y
- - 101
- A
- K
- - 102
- E
- D
- - 105
- L
- D
- - 106
- G
- K
b_factor_mean: 0.0
heteroatoms_removed: []
- case_name: GA_GB
tier: 1
classification: known-state
query_source_db: engineered (refs. 49, 50, 51)
query_accession: see_state_notes
query_organism: engineered from Streptococcus protein G GB1/GA domain + HSA-binding
GA domain
canonical_query_sequence: TTYKLILNLKQAKEEAIKELVDAGTAEKYFKLIANAKTVEGVWTLKDEIKTFTVTE
canonical_query_length: 56
canonical_query_frame: 1..56 (representative = GA98 / 2LHC sequence)
construct_start: 1
construct_end: 56
states:
- name: GAWT
pdb_id: null
chain_id: null
residue_range: 1..56
has_mutations: false
notes: 'Sequence-only variant from the AF-Cluster notebook / papers. No deposited
PDB found in RCSB for this exact sequence. Sequence: MEAVDANSLAQAKEAAIKELKQYGIGDYYIKLINNAKTVEGVESLKNEILKALPTE'
qc: null
- name: GA77
pdb_id: null
chain_id: null
residue_range: 1..56
has_mutations: false
notes: 'Sequence-only variant from the AF-Cluster notebook / papers. No deposited
PDB found in RCSB for this exact sequence. Sequence: TTYKLILNLKQAKEEAIKELVDAGIAEKYIKLIANAKTVEGVWTLKDEILKATVTE'
qc: null
- name: GA88
pdb_id: 2JWS
chain_id: A
residue_range: 1..56
has_mutations: false
notes: 'Designed variant in the GA/GB convergence series (refs. 49–51). Sequence:
TTYKLILNLKQAKEEAIKELVDAGIAEKYIKLIANAKTVEGVWTLKDEILTFTVTE'
qc:
method: solution nmr
resolution_A: null
n_models: 20
chains_all:
- A
chain_used: A
chains_dropped: []
missing_residues: []
mutations_vs_canonical: []
b_factor_mean: 0.0
heteroatoms_removed: []
- name: GA91
pdb_id: null
chain_id: null
residue_range: 1..56
has_mutations: false
notes: 'Sequence-only variant from the AF-Cluster notebook / papers. No deposited
PDB found in RCSB for this exact sequence. Sequence: TTYKLILNLKQAKEEAIKELVDAGTAEKYIKLIANAKTVEGVWTLKDEILTFTVTE'
qc: null
- name: GA95
pdb_id: 2KDL
chain_id: A
residue_range: 1..56
has_mutations: false
notes: 'Designed variant in the GA/GB convergence series (refs. 49–51). Sequence:
TTYKLILNLKQAKEEAIKELVDAGTAEKYIKLIANAKTVEGVWTLKDEIKTFTVTE'
qc:
method: solution nmr
resolution_A: null
n_models: 20
chains_all:
- A
chain_used: A
chains_dropped: []
missing_residues: []
mutations_vs_canonical: []
b_factor_mean: 0.0
heteroatoms_removed: []
- name: GA98
pdb_id: 2LHC
chain_id: A
residue_range: 1..56
has_mutations: false
notes: 'Designed variant in the GA/GB convergence series (refs. 49–51). Sequence:
TTYKLILNLKQAKEEAIKELVDAGTAEKYFKLIANAKTVEGVWTLKDEIKTFTVTE'
qc:
method: solution nmr
resolution_A: null
n_models: 20
chains_all:
- A
chain_used: A
chains_dropped: []
missing_residues: []
mutations_vs_canonical: []
b_factor_mean: 0.0
heteroatoms_removed: []
- name: GB98
pdb_id: 2LHD
chain_id: A
residue_range: 1..56
has_mutations: false
notes: 'Designed variant in the GA/GB convergence series (refs. 49–51). Sequence:
TTYKLILNLKQAKEEAIKELVDAGTAEKYFKLIANAKTVEGVWTYKDEIKTFTVTE'
qc:
method: solution nmr
resolution_A: null
n_models: 20
chains_all:
- A
chain_used: A
chains_dropped: []
missing_residues: []
mutations_vs_canonical: []
b_factor_mean: 0.0
heteroatoms_removed: []
- name: GB98_T25I
pdb_id: 2LHG
chain_id: A
residue_range: 1..56
has_mutations: false
notes: 'Designed variant in the GA/GB convergence series (refs. 49–51). Sequence:
TTYKLILNLKQAKEEAIKELVDAGIAEKYFKLIANAKTVEGVWTYKDEIKTFTVTE'
qc:
method: solution nmr
resolution_A: null
n_models: 10
chains_all:
- A
chain_used: A
chains_dropped: []
missing_residues: []
mutations_vs_canonical: []
b_factor_mean: 0.0
heteroatoms_removed: []
- name: GB98_T25I_L20A
pdb_id: 2LHE
chain_id: A
residue_range: 1..56
has_mutations: false
notes: 'Designed variant in the GA/GB convergence series (refs. 49–51). Sequence:
TTYKLILNLKQAKEEAIKEAVDAGIAEKYFKLIANAKTVEGVWTYKDEIKTFTVTE'
qc:
method: solution nmr
resolution_A: null
n_models: 20
chains_all:
- A
chain_used: A
chains_dropped: []
missing_residues: []
mutations_vs_canonical: []
b_factor_mean: 0.0
heteroatoms_removed: []
- name: GB95
pdb_id: 2KDM
chain_id: A
residue_range: 1..56
has_mutations: false
notes: 'Designed variant in the GA/GB convergence series (refs. 49–51). Sequence:
TTYKLILNLKQAKEEAIKEAVDAGTAEKYFKLIANAKTVEGVWTYKDEIKTFTVTE'
qc:
method: solution nmr
resolution_A: null
n_models: 20
chains_all:
- A
chain_used: A
chains_dropped: []
missing_residues: []
mutations_vs_canonical: []
b_factor_mean: 0.0
heteroatoms_removed: []
- name: GB91
pdb_id: null
chain_id: null
residue_range: 1..56
has_mutations: false
notes: 'Sequence-only variant from the AF-Cluster notebook / papers. No deposited
PDB found in RCSB for this exact sequence. Sequence: TTYKLILNLKQAKEEAIKEAVDAGTAEKYFKLYANAKTVEGVWTYKDEIKTFTVTE'
qc: null
- name: GB88
pdb_id: null
chain_id: null
residue_range: 1..56
has_mutations: false
notes: 'Sequence-only variant from the AF-Cluster notebook / papers. No deposited
PDB found in RCSB for this exact sequence. Sequence: TTYKLILNLKQAKEEAITEAVDAGTAEKYFKLYANAKTVEGVWTYKDEIKTFTVTE'
qc: null
- name: GB77
pdb_id: null
chain_id: null
residue_range: 1..56
has_mutations: false
notes: 'Sequence-only variant from the AF-Cluster notebook / papers. No deposited
PDB found in RCSB for this exact sequence. Sequence: TTYKLILNGKQLKEEAITEAVDAATAEKYFKLYANAKTVEGVWTYKDETKTFTVTE'
qc: null
- name: GBWT
pdb_id: null
chain_id: null
residue_range: 1..56
has_mutations: false
notes: 'Sequence-only variant from the AF-Cluster notebook / papers. No deposited
PDB found in RCSB for this exact sequence. Sequence: MTYKLILNGKTLKGETTTEAVDAATAEKVFKQYANDNGVDGEWTYDDATKTFTVTE'
qc: null
note: 'There are 14 sequences in the AF-Cluster notebook (12 engineered mutants
+ GAWT + GBWT). Of these, 7 have deposited RCSB structures: 2LHC (GA98), 2LHD
(GB98), 2LHE (GB98_T25I_L20A), 2LHG (GB98_T25I), 2JWS (GA88), 2KDL (GA95), 2KDM
(GB95). The other 7 are sequence-only variants. 2JWU (a deposited 2008 PNAS precursor)
has a different sequence from the notebook-GB91 (the notebook explicitly flags
`# error in Fig. 2 PNAS 2009`) — it is kept under legacy_2JWU.pdb for traceability
but not used as a primary reference.'
- case_name: Mpt53
tier: 2
classification: discovery
query_source_db: UniProt
query_accession: P9WG65
query_organism: Mycobacterium tuberculosis H37Rv
canonical_query_sequence: ADERLQFTATTLSGAPFDGASLQGKPAVLWFWTPWCPFCNAEAPSLSQVAAANPAVTFVGIATRADVGAMQSFVSKYNLNFTNLNDADGVIWARYNVPWQPAFVFYRADGTSTFVNNPTAAMSQDELSGRVAALTS
canonical_query_length: 136
canonical_query_frame: UniProt 38..173 (mature protein, signal peptide 1..37 cleaved).
This exactly matches the AF-Cluster repo 1LU4A_REF.a3m query (136 aa).
construct_start: 38
construct_end: 173
states:
- name: state_A_reference
pdb_id: 1LU4
chain_id: A
residue_range: 1001..1134 (crystal numbering; = UniProt 38..171). 2 C-term residues
(TS, UniProt 172–173) not in the crystal.
has_mutations: false
notes: Thioredoxin-like reduced state crystal structure of Mpt53. Residue numbering
in 1LU4.pdb starts at 1001; subtract 963 to map to UniProt.
qc:
method: x-ray diffraction
resolution_A: 1.12
n_models: 1
chains_all:
- A
chain_used: A
chains_dropped: []
missing_residues:
- 1135
- 1136
mutations_vs_canonical: []
b_factor_mean: 13.82
heteroatoms_removed:
- HOH
- name: dali_best_info_only
pdb_id: 3EMX
chain_id: A
residue_range: 224..347 (fragment of parent Aeropyrum pernix protein)
has_mutations: false
notes: DALI best hit to the predicted alternative state (per AF-Cluster Fig. 5).
Not a direct evaluation reference; supplied for information only. Discovery-case
§9.2 metrics only use 1LU4.
qc:
method: x-ray diffraction
resolution_A: 2.25
n_models: 1
chains_all:
- A
- B
chain_used: A
chains_dropped:
- B
missing_residues: []
mutations_vs_canonical: []
b_factor_mean: 25.13
heteroatoms_removed:
- HOH
# ── Phase VIII new cases ────────────────────────────────────────────────────
- case_name: RfaH
tier: 1
classification: known-state
query_source_db: UniProt
query_accession: P0AFZ3
query_organism: Escherichia coli K-12
# VERIFY: fetch canonical sequence from UniProt P0AFZ3 before AF2 runs
canonical_query_sequence: SEE_UNIPROT_P0AFZ3
canonical_query_length: 162
canonical_query_frame: UniProt 1..162 (full-length; NTD residues 1-100 + CTD residues 101-162)
construct_start: 1
construct_end: 162
biology: >
RfaH is a transcription elongation factor (NusG paralog). Its C-terminal
domain (CTD, residues ~101-162) undergoes a dramatic fold-switch between a
beta-barrel (free/NusG-like autoinhibited form) and an alpha-helical hairpin
(when engaging the RNA polymerase NTD). The NTD (residues 1-100) is stable
in both states. This is one of the best-characterised natural fold-switching
proteins and a canonical Phase VIII benchmark target.
states:
- name: state_A_NusG_like
# VERIFY: confirm PDB 5OND contains full-length autoinhibited RfaH before use
pdb_id: 5OND
chain_id: A
# VERIFY: check deposited residue range from RCSB before structure cleaning
residue_range: 1..162 (verify)
has_mutations: false
notes: >
Free/NusG-like (autoinhibited) state with CTD in beta-barrel fold.
PDB 5OND is proposed to contain the full-length autoinhibited RfaH
with NTD in ops element-bound form. Chain ID and deposited residue
range MUST be confirmed against RCSB before structure cleaning.
qc:
# VERIFY: fill in method, resolution, chains from RCSB HEADER
method: x-ray (verify)
resolution_A: null
n_models: null
chains_all: []
chain_used: A
chains_dropped: []
missing_residues: []
mutations_vs_canonical: []
b_factor_mean: null
heteroatoms_removed: []
- name: state_B_fold_switched
# VERIFY: 6C6S is proposed fold-switched form; alternatively consider
# 2LCL (isolated CTD beta-barrel, NMR) or 2LCO (isolated CTD alpha, NMR)
# for the CTD-only structures. All IDs must be checked against RCSB.
pdb_id: 6C6S
chain_id: A
# VERIFY: 6C6S may be a full-length NTD-CTD structure; confirm residue range
residue_range: 1..162 (verify)
has_mutations: false
notes: >
Fold-switched state with CTD in alpha-helical hairpin conformation.
If a full-length fold-switched structure is unavailable, consider using
the isolated CTD structures (e.g. 2LCO for alpha-helical CTD, NMR).
PDB ID, chain, method, and residue range MUST be verified against RCSB
before structure cleaning and renumbering.
qc:
# VERIFY: fill in method, resolution, chains from RCSB HEADER
method: x-ray or NMR (verify)
resolution_A: null
n_models: null
chains_all: []
chain_used: A
chains_dropped: []
missing_residues: []
mutations_vs_canonical: []
b_factor_mean: null
heteroatoms_removed: []
- case_name: MAD2
tier: 1
classification: known-state
query_source_db: UniProt
query_accession: O43684
query_organism: Homo sapiens
# VERIFY: fetch canonical sequence from UniProt O43684 before AF2 runs
canonical_query_sequence: SEE_UNIPROT_O43684
canonical_query_length: 205
canonical_query_frame: UniProt 1..205 (full-length MAD2L1)
construct_start: 1
construct_end: 205
biology: >
MAD2 (MAD2L1) is a spindle assembly checkpoint protein that exists in two
conformational states: open (O-MAD2, N1 fold) and closed (C-MAD2, N2 fold).
The switch involves massive topological rearrangement of the C-terminal
"safety belt" region. Closed MAD2 is the active form that sequesters CDC20
to inhibit APC/C. Approximately 205 aa (human MAD2L1).
states:
- name: state_A_open_O_MAD2
# VERIFY: confirm 1DUJ chain A is monomeric open-state MAD2 from RCSB
pdb_id: 1DUJ
chain_id: A
# VERIFY: crystal may be missing terminal residues; confirm exact range from RCSB
residue_range: 1..196 (verify)
has_mutations: false
notes: >
Open state (O-MAD2, N1 fold) monomer. The C-terminal safety belt is in
open topology. Deposited residue range may not cover all 205 residues;
verify from RCSB SEQRES/ATOM records before cleaning.
qc:
# VERIFY: fill in method, resolution, chains from RCSB HEADER
method: x-ray diffraction (verify)
resolution_A: null
n_models: null
chains_all: []
chain_used: A
chains_dropped: []
missing_residues: []
mutations_vs_canonical: []
b_factor_mean: null
heteroatoms_removed: []
- name: state_B_closed_C_MAD2
# VERIFY: 2V64 contains CDC20 peptide; if ligand-free closed state is
# preferred, check 1KLQ or other deposits. Confirm chain ID from RCSB.
pdb_id: 2V64
chain_id: A
# VERIFY: crystal may be missing terminal residues; confirm from RCSB
residue_range: 1..196 (verify)
has_mutations: false
notes: >
Closed state (C-MAD2, N2 fold) with CDC20 peptide bound; the safety belt
wraps around the ligand in a different topology from the open state.
If a ligand-free closed-state structure is available (e.g. 1KLQ), prefer
it. Chain ID and residue range MUST be verified against RCSB before
structure cleaning.
qc:
# VERIFY: fill in method, resolution, chains from RCSB HEADER
method: x-ray diffraction (verify)
resolution_A: null
n_models: null
chains_all: []
chain_used: A
chains_dropped: []
missing_residues: []
mutations_vs_canonical: []
b_factor_mean: null
heteroatoms_removed: []
generated_at: '2026-04-22T18:42:00Z'
updated_at: '2026-04-24T00:00:00Z'