File size: 7,296 Bytes
bae5726
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
---
frameworks:
- PyTorch
language:
- en
license: mit
tags:
- OneScience
- bioscience
- enzyme-optimum-pH-prediction
- protein-language-model
- EpHod
tasks:
- regression
---

<p align="center">
  <strong>
    <span style="font-size: 30px;">EpHod</span>
  </strong>
</p>

# Model Introduction

EpHod is an ensemble model for predicting the catalytic optimum pH (`pHopt`) of enzymes.

The model first uses ESM-1v to encode amino acid sequences into protein representations and then combines predictions from a Residual Lightweight Attention network (RLATtr) and a Support Vector Regression model (SVR).

Paper: [Machine learning prediction of enzyme optimum pH](https://doi.org/10.1038/s42256-025-01026-6)

# Model Description

The EpHod inference pipeline contains three main prediction components:

- **ESM-1v:** Encodes enzyme sequences into 1280-dimensional residue-level protein representations;
- **RLATtr:** Uses a residual lightweight attention network to predict `pHopt` and can optionally output residue-level attention weights and a 2560-dimensional EpHod protein representation;
- **SVR:** Performs support vector regression using pooled and standardized ESM-1v representations;
- **Ensemble:** Uses the average of the RLATtr and SVR predictions as the final `pHopt` prediction.

The official RLATtr model was first pretrained on approximately 1.9 million proteins labeled with optimum environmental pH (`pHenv`) and was then fine-tuned on 9,855 enzymes labeled with catalytic optimum pH (`pHopt`).

Input sequences longer than 1022 residues are truncated.

To avoid pooling-related bias, the current inference entry point uses a fixed batch size of 1.

# Use Cases

| Use Case | Description |
| --- | --- |
| Enzyme optimum pH prediction | Predict catalytic optimum pH from an enzyme amino acid sequence. |
| Enzyme candidate screening | Compare multiple candidate enzyme sequences based on predicted optimum pH. |
| Attention analysis | Optionally save residue-level RLATtr attention weights. |
| Protein representation extraction | Optionally save the final 2560-dimensional RLATtr protein representation. |

# Usage

## 1. OneCode

You can use the OneCode online environment for an intelligent one-click AI4S programming experience:

[Try OneCode for AI4S Programming](https://web-2069360198568017922-iaaj.ksai.scnet.cn:58043/home)

## 2. Manual Installation

### Hardware Requirements

- Supports CPU and accelerator devices supported by PyTorch;
- GPU or SCNet DCU is recommended for ESM-1v inference;
- CPU execution is supported but is significantly slower;
- ESM-1v contains approximately 650 million parameters;
- Device memory usage depends on sequence length. If device memory is insufficient, reduce the input sequence length or process sequences individually.

### Download the Model Package

Install the Hugging Face command-line tool and download the model repository:

```bash

python -m pip install -U huggingface_hub

hf download OneScience-Group/EpHod --local-dir ./EpHod
cd EpHod
```

### Install the Runtime Environment

**DCU Environment**

```bash
# Activate DTK and Conda first
conda create -n onescience311 python=3.11 -y
conda activate onescience311

python -m pip install "onescience[bio-dcu]" \
  -i http://mirrors.onescience.ai:3141/pypi/simple/ \
  --trusted-host mirrors.onescience.ai
```
**GPU Environment**

```bash
# Activate Conda first
conda create -n onescience311 python=3.11 -y
conda activate onescience311

python -m pip install "onescience[bio-gpu]" \
  -i http://mirrors.onescience.ai:3141/pypi/simple/ \
  --trusted-host mirrors.onescience.ai
```

Install the additional dependencies required by EpHod:

```bash
python -m pip install --no-deps -r requirements.txt
```

### Weight Preparation

Inference requires all three of the following assets:

| Asset | Relative Path | Purpose |
| --- | --- | --- |
| ESM-1v 650M weights | `weight/esm1v_t33_650M_UR90S_1.pt` | Generate residue-level protein representations |
| RLATtr weights | `weight/ESM1v-RLATtr.pt` | Neural-network prediction branch |
| SVR model and normalization statistics | `weight/ESM1v-SVR.pkl` | Support Vector Regression prediction branch |

Official sources:

- [ESM-1v main checkpoint](https://dl.fbaipublicfiles.com/fair-esm/models/esm1v_t33_650M_UR90S_1.pt)
- [EpHod RLATtr weights and training data](https://doi.org/10.5281/zenodo.14252615)
- `ESM1v-SVR.pkl` is distributed with the official EpHod repository.

### Quick Inference

The following command uses a validated smoke-test sequence:

```bash
python scripts/inference.py \
  --fasta_path conf/data/smoke.fasta \
  --output_path output/smoke/prediction.csv \
  --verbose 1 \
  --save_attention_weights 0 \
  --save_embeddings 0
```

A complete example using the provided test sequences:

```bash
python scripts/inference.py \
  --fasta_path conf/data/test_sequences.fasta \
  --output_path output/inference/prediction.csv \
  --verbose 1 \
  --save_attention_weights 0 \
  --save_embeddings 0
```

The output CSV contains three prediction columns:

```text
RLATtr,SVR,Ensemble
```

Their meanings are:

- `RLATtr`: optimum-pH prediction from the neural-network branch;
- `SVR`: optimum-pH prediction from the support vector regression branch;
- `Ensemble`: arithmetic mean of the RLATtr and SVR predictions and the recommended final EpHod prediction.

The `--output_path` argument directly specifies the complete output CSV path and automatically creates its parent directory when required.

The original `--save_dir` and `--csv_name` options remain available.

If `--output_path` is not specified, the output path is generated from `--save_dir` and `--csv_name`.

### Save Attention Weights and Protein Representations

Set the corresponding options to `1`:

```bash
python scripts/inference.py \
  --fasta_path conf/data/smoke.fasta \
  --output_path output/features/prediction.csv \
  --save_attention_weights 1 \
  --save_embeddings 1
```

The output includes:

```text
output/features/
β”œβ”€β”€ attention_weights/
β”œβ”€β”€ embeddings.csv
└── prediction.csv
```

`attention_weights/` stores residue-level RLATtr attention information.

`embeddings.csv` stores the extracted EpHod protein representations.

`prediction.csv` stores the RLATtr, SVR, and ensemble optimum-pH predictions.

# OneScience Official Resources

| Platform | OneScience Main Repository | Skills Repository |
| --- | --- | --- |
| Gitee | https://gitee.com/onescience-ai/onescience | https://gitee.com/onescience-ai/oneskills |
| GitHub | https://github.com/onescience-ai/OneScience | https://github.com/onescience-ai/oneskills |

# Citation and License

- EpHod paper: [Machine learning prediction of enzyme optimum pH](https://doi.org/10.1038/s42256-025-01026-6)
- Official implementation: https://github.com/jafetgado/EpHod
- EpHod model and data: [Machine learning prediction of enzyme optimal pH](https://doi.org/10.5281/zenodo.14252615)
- The upstream EpHod implementation is distributed under the MIT License.
- This model package provides SCNet/DCU runtime adaptation and directory organization based on the official implementation.
- The adaptation does not modify the copyright status, licenses, or terms of use of the original paper, source code, model weights, datasets, ESM-1v assets, or other third-party resources.