🧬 BioReason-Pro
Advancing Protein Function Prediction with
Multimodal Biological Reasoning

bioRxiv GitHub Website HuggingFace

GO-GPT v2

GO-GPT predicts Gene Ontology (GO) terms from a protein sequence. A frozen ESM2-650M protein language model embeds the sequence; an 8-layer GPT-style decoder then generates the GO terms of one aspect at a time (Molecular Function, Biological Process, Cellular Component) as a token sequence, conditioned on the protein and its organism. Its predictions are one of the inputs of BioReason-Pro.

This is version 2. Version 1 is wanglab/gogpt. v2 is trained on the BioReason-Pro v2 dataset (more proteins, tier-1 and tier-2 GO labels) with a smaller protein encoder and a revised decoder; see Changes from v1.

Usage

The code is in gogpt/ of the BioReason-Pro repository:

git clone https://github.com/NVIDIA-BioNeMo/BioReason-Pro.git
pip install -e BioReason-Pro/gogpt
from gogpt import GOGPTPredictor

predictor = GOGPTPredictor.from_pretrained("wanglab/gogpt-v2")
predictor.predict(
    "MERPEPELIRQSWRAVSRSPLEHGTVLFARLFALEPDLLPLFQYNCRQFSSPEDCLSSPEFLDHIRKVMLVIDAAVTNVEDLSSLEEYLASLGRKHRAVGVKLSSFSTVGESLLYMLEKCLGPAFTPATRAAWSQLYGAVVQAMSRGWDGE",
    organism=9606,  # NCBI taxonomy id, or a name such as "Homo sapiens"
)
# {"MF": ["GO:0003674", "GO:0003824", "GO:0005488", ...], "BP": [...], "CC": [...]}
  • Each aspect's terms come in generation order, from general to specific, and include the ancestors of the specific terms (the training labels were propagated the same way).
  • The model has embeddings for 200 organisms, the most frequent in its training data (listed in organism_mapper.json). Other organisms, or none, use a shared "unknown organism" embedding.
  • Sequences longer than 2,000 residues are truncated.
  • Greedy decoding is the default and is what the results below use. Beam search and sampling are available. On GPU the model runs under bf16 autocast, as in training.

The inference notebook and the README cover batch prediction, the command-line tool and training.

Results

Greedy decoding, on the two held-out test sets of the BioReason-Pro v2 dataset (no test protein was used in training). Each protein is scored on its scored_aspects against tier-1 labels with the CAFA evaluator (cafaeval): predictions and labels are propagated to their ancestors, and wF1 is the F1 weighted by information accretion (data/IA.txt in the code repository).

test set proteins (protein-aspect pairs) MF wF1 BP wF1 CC wF1 mean wF1 mean F1
proteins_extra / test 6,112 (9,748) 0.638 0.419 0.628 0.562 0.648
proteins / test 3,726 (5,186) 0.664 0.616 0.715 0.665 0.729

Against tier-1 ∪ tier-2 labels (what the model is trained on) the mean wF1 is 0.629 and 0.707.

Training

  • Data: the train splits of the proteins and proteins_extra configs of wanglab/bioreason-pro-v2 after holding out 8,000 validation proteins: 227,965 training proteins, 392,613 protein-aspect examples. Labels are tier-1 ∪ tier-2 GO annotations propagated to their ancestors, restricted to terms annotated on at least 20 training proteins and ordered by GO depth.
  • Recipe: 100 epochs on 4 H100 GPUs, effective batch 160, AdamW (lr 3e-4, weight decay 0.2, 5% warmup then cosine decay to 3e-5), label smoothing 0.03, 15% of GO input tokens masked, bf16 mixed precision. The ESM2 encoder is frozen.
  • Checkpoint: the epoch with the best validation IA-weighted F1 (Lightning epoch index 89, validation wF1 0.682).

Model

protein encoder facebook/esm2_t33_650M_UR50D, frozen; hidden_states[27], no special tokens
decoder 8 layers, 12 heads, 900 dims, RMSNorm, QK-norm, gated attention; 198M parameters
attention protein positions attend to each other; GO tokens attend to all protein positions and to earlier GO tokens
GO vocabulary 13,453 terms (MF 2,501, BP 9,493, CC 1,459); GO release 2023-01-01
organisms 200 NCBI taxonomy ids + unknown
max length 2,000 residues + 1,000 GO tokens

Files: model.safetensors (decoder weights, fp32), config.json (architecture and default generation settings), go_tokenizer.json (GO vocabulary), organism_mapper.json (organism vocabulary with display names). ESM2 is downloaded from its own repository.

Changes from v1

v1 v2
training data CAFA5-based BioReason-Pro v2, tier-1 and tier-2 labels
protein encoder ESM2-3B, layer 30 ESM2-650M, layer 27; up to 2,000 residues
decoder 12 layers, LayerNorm 8 layers, RMSNorm, QK-norm
organism input species name NCBI taxonomy id or name
default decoding beam search greedy
weights pickled Lightning checkpoint safetensors + JSON

Limitations

  • Predictions are limited to the 13,453 GO terms in the vocabulary; rarer terms are never predicted.
  • Proteins from organisms outside the 200 known ones get no organism-specific conditioning.
  • Proteins that were in the training data (all of proteins/train and proteins_extra/train except the validation proteins) get memorized rather than predicted annotations; evaluate only on held-out proteins.
  • GO terms follow the 2023-01-01 GO release; terms added or made obsolete since then are not handled.

Citation

@article {Fallahpour2026.03.19.712954,
    author = {Fallahpour, Adibvafa and Seyed-Ahmadi, Arman and Idehpour, Parsa and Ibrahim, Omar and Gupta, Purav and Naimer, Jack and Zhu, Kevin and Shah, Arnav and Ma, Shihao and Adduri, Abhinav and G{\"u}loglu, Talu and Liu, Nuo and Cui, Haotian and Jain, Arihant and de Castro, Max and Fallahpour, Amirfaham and Cembellin-Prieto, Antonio and Stiles, John S. and Nem{\v c}ko, Filip and Nevue, Alexander A. and Moon, Hyungseok C. and Sosnick, Lucas and Markham, Olivia and Duan, Haonan and Lee, Michelle Y. Y. and Salvador, Andrea F. M. and Maddison, Chris J. and Thaiss, Christoph A. and Ricci-Tam, Chiara and Plosky, Brian S. and Burke, Dave P. and Hsu, Patrick D. and Goodarzi, Hani and Wang, Bo},
    title = {BioReason-Pro: Advancing Protein Function Prediction with Multimodal Biological Reasoning},
    elocation-id = {2026.03.19.712954},
    year = {2026},
    doi = {10.64898/2026.03.19.712954},
    publisher = {Cold Spring Harbor Laboratory},
    URL = {https://www.biorxiv.org/content/early/2026/03/20/2026.03.19.712954},
    eprint = {https://www.biorxiv.org/content/early/2026/03/20/2026.03.19.712954.full.pdf},
    journal = {bioRxiv}
}
Downloads last month
-
Safetensors
Model size
0.2B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train wanglab/gogpt-v2