VenusREM / README.md
wuxing0105's picture
Upload folder using huggingface_hub
3c22944 verified
|
Raw
History Blame Contribute Delete
11.6 kB
---
frameworks:
- PyTorch
language:
- en
license: cc-by-nc-nd-4.0
tags:
- OneScience
- bioscience
- protein-mutation-effect-prediction
- protein-engineering
- VenusREM
- ProSST
tasks: []
datasets:
- ICML2022/ProteinGym
---
<p align="center">
<strong>
<span style="font-size: 30px;">VenusREM</span>
</strong>
</p>
# Model Introduction
VenusREM is a zero-shot model for protein mutation effect prediction. It builds on sequence-structure language model logits from ProSST and incorporates evolutionary information derived from homologous sequence alignments. Candidate mutation fitness scores are calculated through retrieval-based logits fusion.
VenusREM can score both single substitutions and mutations containing multiple amino acid substitutions.
The output score represents the model's relative preference for a candidate mutation compared with the wild-type sequence. It should not be interpreted as an experimental measurement with a universal physical unit.
Paper: [From high-throughput evaluation to wet-lab studies: advancing mutation effect prediction with a retrieval-enhanced model](https://academic.oup.com/bioinformatics/article/41/Supplement_1/i401/8199372)
# Model Description
VenusREM uses ProSST-2048 as its base model:
- Represents the primary structure of a protein using its amino acid sequence;
- Quantizes the local three-dimensional environment of each residue into one of 2048 structural tokens;
- Jointly models sequence and structural information to generate amino acid logits at each residue position;
- Estimates evolutionary distributions from homologous sequence alignments in A2M, A3M, or FASTA format;
- Uses the `alpha` parameter to fuse language-model logits with retrieved evolutionary logits;
- Calculates mutation scores from the log-probability difference between the mutant and wild-type amino acids.
# Use Cases
| Use Case | Description |
| :---: | :--- |
| Zero-shot mutation ranking | Rank candidate amino acid substitutions without target-protein-specific supervised labels. |
| ProteinGym evaluation | Generate per-protein mutation scores and summarize Spearman correlations on ProteinGym substitution datasets. |
| MSA-enhanced prediction | Integrate evolutionary information from homologous sequence alignments to improve context-aware mutation scoring. |
| ProSST baseline prediction | Set `alpha=0` to use only sequence-structure language model logits. |
| Structural token generation | Optionally generate ProSST structural tokens when only PDB structures are available. |
# Usage
## 1. OneCode
You can use the OneCode online environment for an intelligent one-click AI4S programming experience:
[Try OneCode for AI4S Programming](https://web-2069360198568017922-iaaj.ksai.scnet.cn:58043/home)
## 2. Manual Installation
**Hardware Requirements**
- Supports both CPU and DCU execution in the OneScience DTK environment;
- DCU is recommended for full ProteinGym inference and PDB-to-structural-token preprocessing;
- CPU execution can be used for basic validation, but full-dataset inference and structural preprocessing are significantly slower.
### Download the Model Package
Install the Hugging Face command-line tool and download the model repository:
```bash
python -m pip install -U huggingface_hub
hf download OneScience-Group/VenusREM --local-dir ./VenusREM
cd VenusREM
```
### Install the Runtime Environment
**DCU Environment**
```bash
conda create -n onescience311 python=3.11 -y
conda activate onescience311
python -m pip install "onescience[bio-dcu]" \
-i http://mirrors.onescience.ai:3141/pypi/simple/ \
--trusted-host mirrors.onescience.ai
```
Install the additional VenusREM dependencies on top of the OneScience base environment:
```bash
python -m pip install --no-deps -r requirements.txt
```
### Weights and Data Preparation
The official ProSST-2048 assets required for inference are located at:
| Asset | Relative Path | Purpose |
| --- | --- | --- |
| Model weights | `weight/ProSST-2048/model.safetensors` | ProSST-2048 parameters |
| Model configuration | `weight/ProSST-2048/config.json` | Network architecture and vocabulary configuration |
| Model implementation | `weight/ProSST-2048/modeling_prosst.py` | Custom Transformers model implementation |
| Configuration implementation | `weight/ProSST-2048/configuration_prosst.py` | Custom Transformers configuration implementation |
| Amino acid vocabulary | `weight/ProSST-2048/vocab.txt` | Sequence tokenization |
| Tokenizer configuration | `weight/ProSST-2048/tokenizer_config.json` | Tokenizer parameters |
Inference datasets use the following relative directory structure:
```text
conf/data/<dataset_name>/
β”œβ”€β”€ aa_seq/
β”‚ └── protein1.fasta
β”œβ”€β”€ aa_seq_aln_a2m/
β”‚ └── protein1.a2m
β”œβ”€β”€ struc_seq/
β”‚ └── 2048/
β”‚ └── protein1.fasta
└── substitutions/
└── protein1.csv
```
The base filenames must match across the corresponding files. For example:
```text
protein1.fasta
protein1.a2m
protein1.csv
```
must all refer to the same protein.
Each file under `substitutions/*.csv` must contain at least:
- `mutant`: mutation notation such as `A10V`; multiple substitutions are separated by colons, for example `A10V:G25D`;
- `DMS_score`: experimental score used to calculate Spearman correlation during evaluation. If no experimental label is available, this field can be set to `0`, but the resulting correlation metric will not have evaluation significance.
Structural tokens are part of the required base-model input.
Even when `alpha=0` is used to disable MSA-based retrieval enhancement, the following structural-token files are still required:
```text
struc_seq/2048/*.fasta
```
### Quick Inference
**Purpose:** Fuse ProSST-2048 sequence-structure logits with evolutionary information from residue-sequence MSAs.
This mode is suitable for full ProteinGym evaluation and formal mutation ranking when precomputed A2M/A3M alignments are available.
Run VenusREM inference on DCU using the ProSST-2048 weights and precomputed residue-sequence MSA:
```bash
export HIP_VISIBLE_DEVICES=0
export CUDA_VISIBLE_DEVICES=0
python scripts/compute_fitness.py \
--model_name weight/ProSST-2048 \
--model_out_name VenusREM_DCU \
--base_dir conf/data/proteingym_v1 \
--out_scores_dir output/proteingym_v1 \
--logit_mode aa_seq_aln \
--alpha 0.8
```
The output structure is:
```text
output/proteingym_v1/
β”œβ”€β”€ scores/
β”‚ └── <protein_name>.csv
└── summary_performance.csv
```
Each per-protein CSV file contains an additional `VenusREM_DCU` score column.
`summary_performance.csv` records the Spearman correlation for each protein dataset.
### ProSST Inference Without MSA
**Purpose:** Run the ProSST-2048 baseline without retrieval enhancement.
This mode can be used to:
- Compare mutation scores before and after MSA fusion;
- Perform ablation analysis;
- Run basic sequence-structure mutation scoring when homologous sequence alignments are unavailable.
Set `alpha=0` to disable residue-sequence and structural-alignment logit fusion.
Amino acid sequences, structural tokens, and substitution files are still required:
```bash
python scripts/compute_fitness.py \
--model_name weight/ProSST-2048 \
--model_out_name ProSST-2048 \
--base_dir conf/data/proteingym_v1 \
--out_scores_dir output/prosst_2048 \
--alpha 0
```
### Structural Sequence Alignment
**Purpose:** Use structural homolog information generated by Foldseek to enhance ProSST logits.
This mode is suitable when structural-alignment results have already been prepared and can be used to evaluate the contribution of structure-based retrieval or compare it with residue-sequence MSA retrieval.
If Foldseek structural-sequence alignments are available, run:
```bash
python scripts/compute_fitness.py \
--model_name weight/ProSST-2048 \
--model_out_name VenusREM_struc \
--base_dir conf/data/proteingym_v1 \
--out_scores_dir output/proteingym_v1_struc \
--logit_mode struc_seq_aln \
--alpha 0.8
```
The corresponding alignment files should be located at:
```text
conf/data/proteingym_v1/struc_seq_aln_foldseek/<protein_name>.fasta
```
### Convert PDB Structures to Structural Tokens
**Purpose:** Quantize the local three-dimensional environment of each residue in a protein PDB structure into structural tokens that can be processed by ProSST-2048.
The generated structural tokens are required inputs for the inference modes described above. This preprocessing step does not directly generate mutation fitness scores.
Run this step only when PDB structures are available but precomputed files under:
```text
struc_seq/2048/*.fasta
```
are not available.
If official precomputed structural tokens are already provided, this step can be skipped.
For a single PDB file:
```bash
python model/data/get_struc_seq.py \
--pdb_file conf/data/proteingym_v1/pdbs/protein1.pdb \
--output_dir conf/data/proteingym_v1/struc_seq \
--vocab_size 2048 \
--overwrite
```
For multiple PDB files:
```bash
python model/data/get_struc_seq.py \
--pdb_dir conf/data/proteingym_v1/pdbs \
--output_dir conf/data/proteingym_v1/struc_seq \
--vocab_size 2048 \
--overwrite
```
### Custom Data Inference
**Purpose:** Perform zero-shot fitness scoring for user-provided proteins and candidate mutations.
This workflow can be used for:
- Candidate prioritization before wet-lab experiments;
- Initial screening in protein engineering;
- Evaluation on custom mutation datasets.
Filenames must remain consistent across the amino acid sequence, MSA, structural-token, and substitution directories.
Create a relative directory structure such as:
```text
conf/data/my_proteins/
β”œβ”€β”€ aa_seq/
β”œβ”€β”€ aa_seq_aln_a2m/
β”œβ”€β”€ struc_seq/2048/
└── substitutions/
```
Then run:
```bash
python scripts/compute_fitness.py \
--model_name weight/ProSST-2048 \
--model_out_name VenusREM \
--base_dir conf/data/my_proteins \
--out_scores_dir output/my_proteins \
--logit_mode aa_seq_aln \
--alpha 0.8
```
If no substitution file is available, the provided helper script can be used to generate all possible single-amino-acid substitutions before further filtering:
```bash
python model/data/get_sav.py \
--fasta_file conf/data/my_proteins/aa_seq/protein1.fasta \
--output_csv conf/data/my_proteins/substitutions/protein1.csv
```
### Training
The official VenusREM repository does not provide a directly executable training entry point, Dataset implementation, optimizer configuration, or complete training loop.
Therefore, this Hugging Face model package does not provide training scripts.
# OneScience Official Resources
| Platform | OneScience Main Repository | Skills Repository |
| --- | --- | --- |
| Gitee | https://gitee.com/onescience-ai/onescience | https://gitee.com/onescience-ai/oneskills |
| GitHub | https://github.com/onescience-ai/OneScience | https://github.com/onescience-ai/oneskills |
# Citation and License
- VenusREM paper: [From high-throughput evaluation to wet-lab studies: advancing mutation effect prediction with a retrieval-enhanced model](https://academic.oup.com/bioinformatics/article/41/Supplement_1/i401/8199372)
- Official implementation: https://github.com/ai4protein/VenusREM
- Base model: https://github.com/ai4protein/ProSST
- This project is provided under the `CC-BY-NC-ND-4.0` license.
- Model weights, source code, datasets, and third-party assets remain subject to their respective original licenses, copyright notices, and terms of use.