File size: 7,296 Bytes
bae5726 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 | ---
frameworks:
- PyTorch
language:
- en
license: mit
tags:
- OneScience
- bioscience
- enzyme-optimum-pH-prediction
- protein-language-model
- EpHod
tasks:
- regression
---
<p align="center">
<strong>
<span style="font-size: 30px;">EpHod</span>
</strong>
</p>
# Model Introduction
EpHod is an ensemble model for predicting the catalytic optimum pH (`pHopt`) of enzymes.
The model first uses ESM-1v to encode amino acid sequences into protein representations and then combines predictions from a Residual Lightweight Attention network (RLATtr) and a Support Vector Regression model (SVR).
Paper: [Machine learning prediction of enzyme optimum pH](https://doi.org/10.1038/s42256-025-01026-6)
# Model Description
The EpHod inference pipeline contains three main prediction components:
- **ESM-1v:** Encodes enzyme sequences into 1280-dimensional residue-level protein representations;
- **RLATtr:** Uses a residual lightweight attention network to predict `pHopt` and can optionally output residue-level attention weights and a 2560-dimensional EpHod protein representation;
- **SVR:** Performs support vector regression using pooled and standardized ESM-1v representations;
- **Ensemble:** Uses the average of the RLATtr and SVR predictions as the final `pHopt` prediction.
The official RLATtr model was first pretrained on approximately 1.9 million proteins labeled with optimum environmental pH (`pHenv`) and was then fine-tuned on 9,855 enzymes labeled with catalytic optimum pH (`pHopt`).
Input sequences longer than 1022 residues are truncated.
To avoid pooling-related bias, the current inference entry point uses a fixed batch size of 1.
# Use Cases
| Use Case | Description |
| --- | --- |
| Enzyme optimum pH prediction | Predict catalytic optimum pH from an enzyme amino acid sequence. |
| Enzyme candidate screening | Compare multiple candidate enzyme sequences based on predicted optimum pH. |
| Attention analysis | Optionally save residue-level RLATtr attention weights. |
| Protein representation extraction | Optionally save the final 2560-dimensional RLATtr protein representation. |
# Usage
## 1. OneCode
You can use the OneCode online environment for an intelligent one-click AI4S programming experience:
[Try OneCode for AI4S Programming](https://web-2069360198568017922-iaaj.ksai.scnet.cn:58043/home)
## 2. Manual Installation
### Hardware Requirements
- Supports CPU and accelerator devices supported by PyTorch;
- GPU or SCNet DCU is recommended for ESM-1v inference;
- CPU execution is supported but is significantly slower;
- ESM-1v contains approximately 650 million parameters;
- Device memory usage depends on sequence length. If device memory is insufficient, reduce the input sequence length or process sequences individually.
### Download the Model Package
Install the Hugging Face command-line tool and download the model repository:
```bash
python -m pip install -U huggingface_hub
hf download OneScience-Group/EpHod --local-dir ./EpHod
cd EpHod
```
### Install the Runtime Environment
**DCU Environment**
```bash
# Activate DTK and Conda first
conda create -n onescience311 python=3.11 -y
conda activate onescience311
python -m pip install "onescience[bio-dcu]" \
-i http://mirrors.onescience.ai:3141/pypi/simple/ \
--trusted-host mirrors.onescience.ai
```
**GPU Environment**
```bash
# Activate Conda first
conda create -n onescience311 python=3.11 -y
conda activate onescience311
python -m pip install "onescience[bio-gpu]" \
-i http://mirrors.onescience.ai:3141/pypi/simple/ \
--trusted-host mirrors.onescience.ai
```
Install the additional dependencies required by EpHod:
```bash
python -m pip install --no-deps -r requirements.txt
```
### Weight Preparation
Inference requires all three of the following assets:
| Asset | Relative Path | Purpose |
| --- | --- | --- |
| ESM-1v 650M weights | `weight/esm1v_t33_650M_UR90S_1.pt` | Generate residue-level protein representations |
| RLATtr weights | `weight/ESM1v-RLATtr.pt` | Neural-network prediction branch |
| SVR model and normalization statistics | `weight/ESM1v-SVR.pkl` | Support Vector Regression prediction branch |
Official sources:
- [ESM-1v main checkpoint](https://dl.fbaipublicfiles.com/fair-esm/models/esm1v_t33_650M_UR90S_1.pt)
- [EpHod RLATtr weights and training data](https://doi.org/10.5281/zenodo.14252615)
- `ESM1v-SVR.pkl` is distributed with the official EpHod repository.
### Quick Inference
The following command uses a validated smoke-test sequence:
```bash
python scripts/inference.py \
--fasta_path conf/data/smoke.fasta \
--output_path output/smoke/prediction.csv \
--verbose 1 \
--save_attention_weights 0 \
--save_embeddings 0
```
A complete example using the provided test sequences:
```bash
python scripts/inference.py \
--fasta_path conf/data/test_sequences.fasta \
--output_path output/inference/prediction.csv \
--verbose 1 \
--save_attention_weights 0 \
--save_embeddings 0
```
The output CSV contains three prediction columns:
```text
RLATtr,SVR,Ensemble
```
Their meanings are:
- `RLATtr`: optimum-pH prediction from the neural-network branch;
- `SVR`: optimum-pH prediction from the support vector regression branch;
- `Ensemble`: arithmetic mean of the RLATtr and SVR predictions and the recommended final EpHod prediction.
The `--output_path` argument directly specifies the complete output CSV path and automatically creates its parent directory when required.
The original `--save_dir` and `--csv_name` options remain available.
If `--output_path` is not specified, the output path is generated from `--save_dir` and `--csv_name`.
### Save Attention Weights and Protein Representations
Set the corresponding options to `1`:
```bash
python scripts/inference.py \
--fasta_path conf/data/smoke.fasta \
--output_path output/features/prediction.csv \
--save_attention_weights 1 \
--save_embeddings 1
```
The output includes:
```text
output/features/
βββ attention_weights/
βββ embeddings.csv
βββ prediction.csv
```
`attention_weights/` stores residue-level RLATtr attention information.
`embeddings.csv` stores the extracted EpHod protein representations.
`prediction.csv` stores the RLATtr, SVR, and ensemble optimum-pH predictions.
# OneScience Official Resources
| Platform | OneScience Main Repository | Skills Repository |
| --- | --- | --- |
| Gitee | https://gitee.com/onescience-ai/onescience | https://gitee.com/onescience-ai/oneskills |
| GitHub | https://github.com/onescience-ai/OneScience | https://github.com/onescience-ai/oneskills |
# Citation and License
- EpHod paper: [Machine learning prediction of enzyme optimum pH](https://doi.org/10.1038/s42256-025-01026-6)
- Official implementation: https://github.com/jafetgado/EpHod
- EpHod model and data: [Machine learning prediction of enzyme optimal pH](https://doi.org/10.5281/zenodo.14252615)
- The upstream EpHod implementation is distributed under the MIT License.
- This model package provides SCNet/DCU runtime adaptation and directory organization based on the official implementation.
- The adaptation does not modify the copyright status, licenses, or terms of use of the original paper, source code, model weights, datasets, ESM-1v assets, or other third-party resources. |