🧬 BioReason-Pro
Advancing Protein Function Prediction with
Multimodal Biological Reasoning
GO-GPT v2
GO-GPT predicts Gene Ontology (GO) terms from a protein sequence. A frozen ESM2-650M protein language model embeds the sequence; an 8-layer GPT-style decoder then generates the GO terms of one aspect at a time (Molecular Function, Biological Process, Cellular Component) as a token sequence, conditioned on the protein and its organism. Its predictions are one of the inputs of BioReason-Pro.
This is version 2. Version 1 is wanglab/gogpt. v2 is trained on the BioReason-Pro v2 dataset (more proteins, tier-1 and tier-2 GO labels) with a smaller protein encoder and a revised decoder; see Changes from v1.
Usage
The code is in gogpt/ of the
BioReason-Pro repository:
git clone https://github.com/NVIDIA-BioNeMo/BioReason-Pro.git
pip install -e BioReason-Pro/gogpt
from gogpt import GOGPTPredictor
predictor = GOGPTPredictor.from_pretrained("wanglab/gogpt-v2")
predictor.predict(
"MERPEPELIRQSWRAVSRSPLEHGTVLFARLFALEPDLLPLFQYNCRQFSSPEDCLSSPEFLDHIRKVMLVIDAAVTNVEDLSSLEEYLASLGRKHRAVGVKLSSFSTVGESLLYMLEKCLGPAFTPATRAAWSQLYGAVVQAMSRGWDGE",
organism=9606, # NCBI taxonomy id, or a name such as "Homo sapiens"
)
# {"MF": ["GO:0003674", "GO:0003824", "GO:0005488", ...], "BP": [...], "CC": [...]}
- Each aspect's terms come in generation order, from general to specific, and include the ancestors of the specific terms (the training labels were propagated the same way).
- The model has embeddings for 200 organisms, the most frequent in its training data (listed in
organism_mapper.json). Other organisms, or none, use a shared "unknown organism" embedding. - Sequences longer than 2,000 residues are truncated.
- Greedy decoding is the default and is what the results below use. Beam search and sampling are available. On GPU the model runs under bf16 autocast, as in training.
The inference notebook and the README cover batch prediction, the command-line tool and training.
Results
Greedy decoding, on the two held-out test sets of the BioReason-Pro v2 dataset (no test protein
was used in training). Each protein is scored on its scored_aspects against tier-1 labels with
the CAFA evaluator (cafaeval): predictions and labels are propagated to their ancestors, and
wF1 is the F1 weighted by information accretion (data/IA.txt in the code repository).
| test set | proteins (protein-aspect pairs) | MF wF1 | BP wF1 | CC wF1 | mean wF1 | mean F1 |
|---|---|---|---|---|---|---|
proteins_extra / test |
6,112 (9,748) | 0.638 | 0.419 | 0.628 | 0.562 | 0.648 |
proteins / test |
3,726 (5,186) | 0.664 | 0.616 | 0.715 | 0.665 | 0.729 |
Against tier-1 ∪ tier-2 labels (what the model is trained on) the mean wF1 is 0.629 and 0.707.
Training
- Data: the
trainsplits of theproteinsandproteins_extraconfigs of wanglab/bioreason-pro-v2 after holding out 8,000 validation proteins: 227,965 training proteins, 392,613 protein-aspect examples. Labels are tier-1 ∪ tier-2 GO annotations propagated to their ancestors, restricted to terms annotated on at least 20 training proteins and ordered by GO depth. - Recipe: 100 epochs on 4 H100 GPUs, effective batch 160, AdamW (lr 3e-4, weight decay 0.2, 5% warmup then cosine decay to 3e-5), label smoothing 0.03, 15% of GO input tokens masked, bf16 mixed precision. The ESM2 encoder is frozen.
- Checkpoint: the epoch with the best validation IA-weighted F1 (Lightning epoch index 89, validation wF1 0.682).
Model
| protein encoder | facebook/esm2_t33_650M_UR50D, frozen; hidden_states[27], no special tokens |
| decoder | 8 layers, 12 heads, 900 dims, RMSNorm, QK-norm, gated attention; 198M parameters |
| attention | protein positions attend to each other; GO tokens attend to all protein positions and to earlier GO tokens |
| GO vocabulary | 13,453 terms (MF 2,501, BP 9,493, CC 1,459); GO release 2023-01-01 |
| organisms | 200 NCBI taxonomy ids + unknown |
| max length | 2,000 residues + 1,000 GO tokens |
Files: model.safetensors (decoder weights, fp32), config.json (architecture and default
generation settings), go_tokenizer.json (GO vocabulary), organism_mapper.json (organism
vocabulary with display names). ESM2 is downloaded from its own repository.
Changes from v1
| v1 | v2 | |
|---|---|---|
| training data | CAFA5-based | BioReason-Pro v2, tier-1 and tier-2 labels |
| protein encoder | ESM2-3B, layer 30 | ESM2-650M, layer 27; up to 2,000 residues |
| decoder | 12 layers, LayerNorm | 8 layers, RMSNorm, QK-norm |
| organism input | species name | NCBI taxonomy id or name |
| default decoding | beam search | greedy |
| weights | pickled Lightning checkpoint | safetensors + JSON |
Limitations
- Predictions are limited to the 13,453 GO terms in the vocabulary; rarer terms are never predicted.
- Proteins from organisms outside the 200 known ones get no organism-specific conditioning.
- Proteins that were in the training data (all of
proteins/trainandproteins_extra/trainexcept the validation proteins) get memorized rather than predicted annotations; evaluate only on held-out proteins. - GO terms follow the 2023-01-01 GO release; terms added or made obsolete since then are not handled.
Citation
@article {Fallahpour2026.03.19.712954,
author = {Fallahpour, Adibvafa and Seyed-Ahmadi, Arman and Idehpour, Parsa and Ibrahim, Omar and Gupta, Purav and Naimer, Jack and Zhu, Kevin and Shah, Arnav and Ma, Shihao and Adduri, Abhinav and G{\"u}loglu, Talu and Liu, Nuo and Cui, Haotian and Jain, Arihant and de Castro, Max and Fallahpour, Amirfaham and Cembellin-Prieto, Antonio and Stiles, John S. and Nem{\v c}ko, Filip and Nevue, Alexander A. and Moon, Hyungseok C. and Sosnick, Lucas and Markham, Olivia and Duan, Haonan and Lee, Michelle Y. Y. and Salvador, Andrea F. M. and Maddison, Chris J. and Thaiss, Christoph A. and Ricci-Tam, Chiara and Plosky, Brian S. and Burke, Dave P. and Hsu, Patrick D. and Goodarzi, Hani and Wang, Bo},
title = {BioReason-Pro: Advancing Protein Function Prediction with Multimodal Biological Reasoning},
elocation-id = {2026.03.19.712954},
year = {2026},
doi = {10.64898/2026.03.19.712954},
publisher = {Cold Spring Harbor Laboratory},
URL = {https://www.biorxiv.org/content/early/2026/03/20/2026.03.19.712954},
eprint = {https://www.biorxiv.org/content/early/2026/03/20/2026.03.19.712954.full.pdf},
journal = {bioRxiv}
}
- Downloads last month
- -