File size: 10,227 Bytes
34393ef | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 | ---
frameworks:
- TensorFlow
- PyTorch
language:
- en
license: cc-by-nc-sa-2.0
tags:
- OneScience
- bioscience
- 5UTR
- RNA-generation
- translation-efficiency-prediction
- UTRGAN
tasks: []
datasets: []
---
<p align="center">
<strong>
<span style="font-size: 30px;">UTRGAN</span>
</strong>
</p>
# Model Introduction
UTRGAN is a generative and optimization framework for 5β² UTR design. It can generate candidate 5β² UTR sequences and predict and rank them according to gene expression, Mean Ribosome Load (MRL), and Translation Efficiency (TE).
The Hugging Face model package already contains the data and pretrained weights required for the basic workflow, so no additional model files need to be downloaded during inference.
# Model Description
UTRGAN consists of multiple cooperating models:
- **WGAN-GP Generator:** Generates candidate 5β² UTR sequences of up to 128 nt from a 40-dimensional random latent vector;
- **WGAN-GP Critic:** Used during generative-model training;
- **Xpresso:** Predicts gene-expression-related scores;
- **FramePool:** Predicts Mean Ribosome Load (MRL);
- **MTtrans:** Predicts Translation Efficiency (TE);
- **G4Boost:** Performs G4-related classification and regression analyses.
WGAN, Xpresso, and FramePool use TensorFlow/Keras, while MTtrans uses PyTorch.
To maintain compatibility with the legacy H5 model files released by the original projects, Legacy Keras is used when loading the corresponding models.
# Use Cases
| Use Case | Description |
| :---: | :--- |
| 5β² UTR candidate generation | Generate 5β² UTR candidate sequences in batches using the official pretrained WGAN. |
| MRL prediction and ranking | Use FramePool to calculate MRL scores and rank candidate sequences. |
| TE prediction and ranking | Use MTtrans to calculate TE scores and rank candidate sequences. |
| MRL/TE-directed optimization | Freeze the pretrained models and optimize the latent noise to improve the target score. |
| Gene-expression optimization | Combine UTRGAN with Xpresso to evaluate gene-expression-related properties of candidate sequences. |
| WGAN-GP training | Validate or retrain the generative model using the included UTRdb2 dataset. |
# Usage
## 1. OneCode
You can use the OneCode online environment for an intelligent one-click AI4S programming experience:
[Try OneCode for AI4S Programming](https://web-2069360198568017922-iaaj.ksai.scnet.cn:58043/home)
## 2. Manual Installation
**Hardware Requirements**
- CPU inference is supported;
- A OneScience-supported DCU environment is recommended for batch generation, ranking, and training;
- TensorFlow and PyTorch access the same DCU through their compatible accelerator interfaces;
- Full training time depends on the number of candidate sequences, batch size, and device performance.
### Download the Model Package
Install the Hugging Face command-line tool and download the model repository:
```bash
python -m pip install -U huggingface_hub
hf download OneScience-Group/UTRGAN --local-dir ./UTRGAN
cd UTRGAN
```
### Install the Runtime Environment
Create and activate a Python 3.11 environment, and then install the OneScience bioscience base environment:
```bash
conda create -n utrgan python=3.11 -y
conda activate utrgan
pip install onescience[bio-dcu]
```
Then install the additional or replacement dependencies specified in `requirements.txt`:
```bash
python -m pip install --no-deps -r requirements.txt
```
> **Note:** This model adaptation uses **DTK 26.04**. The corresponding TensorFlow runtime has also been upgraded to **TensorFlow 2.18**.
Check the frameworks and available devices:
```bash
python - <<'PY'
import tensorflow as tf
import torch
print("TensorFlow:", tf.__version__)
print("TensorFlow devices:", tf.config.list_physical_devices("GPU"))
print("PyTorch:", torch.__version__)
print("HIP:", torch.version.hip)
print("DCU available:", torch.cuda.is_available())
PY
```
### Weights and Data Preparation
The Hugging Face model package already contains the resources required for basic generation, prediction, ranking, and training validation:
| Resource | Location | Purpose |
| --- | --- | --- |
| UTRdb2 | `conf/data/utrdb2.csv` | WGAN-GP training data |
| Motif data | `conf/data/motifs.csv` | Motif statistics and optimization analysis |
| WGAN Generator | `weight/checkpoint_3000.h5` | Generate candidate 5β² UTR sequences |
| FramePool | `weight/utr_model_combined_residual_new.h5` | MRL prediction |
| Xpresso | `weight/humanMedian_trainepoch.11-0.426.h5` | Gene-expression-related prediction |
| Xpresso | `weight/GM12878_trainepoch.06-0.5062.h5` | GM12878 expression-related prediction |
| Xpresso | `weight/K562_trainepoch.11-0.4917.h5` | K562 expression-related prediction |
| MTtrans | `weight/mttrans/RL_hard_share_MTL/3R/schedule_MTL-model_best_cv1.pth` | TE prediction |
| G4Boost | `weight/G4Boost_classifier.json` | G4 classification |
| G4Boost | `weight/G4Boost_regressor.json` | G4 regression |
After applying the filtering and deduplication logic used by the official training scripts, UTRdb2 contains approximately 33,250 sequences with lengths between 65 and 128 nt.
The basic workflow does not require additional dataset downloads or precomputed features.
### Optional Dependencies
`requirements.txt` includes XGBoost, ViennaRNA, logomaker, ruptures, and cliffs-delta for G4-related and downstream analysis scripts.
NUPACK is used only by optional MFE preprocessing scripts.
Because NUPACK 4 must be obtained and installed separately according to the licensing and installation requirements of its publisher, it is not automatically installed through standard PyPI dependencies in this model package.
NUPACK is not required for:
- Basic inference;
- MRL ranking;
- TE ranking;
- WGAN-GP training.
### Quick Inference
Use the official pretrained WGAN to generate candidate sequences and rank them independently by MRL and TE.
**Purpose:** Generate 5β² UTR candidates in batches on DCU and output separate MRL and TE ranking results.
```bash
python scripts/predict.py \
--device dcu \
--device-id 0 \
--num-candidates 1024 \
--batch-size 128 \
--seed 33 \
--output-dir outputs/pretrained_batch_ranking
```
The results are saved to:
```text
outputs/pretrained_batch_ranking/
βββ all_candidates_scores.csv
βββ ranked_by_mrl.csv
βββ ranked_by_te.csv
βββ generator_probabilities.npy
βββ summary.json
```
The output files contain:
- `all_candidates_scores.csv`: all candidate sequences together with their MRL and TE prediction scores;
- `ranked_by_mrl.csv`: candidates ranked from highest to lowest MRL score;
- `ranked_by_te.csv`: candidates ranked from highest to lowest TE score;
- `is_duplicate`: indicates whether a generated sequence is duplicated;
- MRL and TE are measured on different scales and their raw scores should not be directly added together.
### MRL-Directed Optimization
**Purpose:** Freeze the WGAN and FramePool models and optimize the latent noise to increase the predicted MRL score of generated candidates.
```bash
python scripts/optimize_te_mrl.py \
-gpu 0 \
-task mrl \
-bs 64 \
-s 10 \
--output-dir outputs/optimization_mrl
```
During this process, the WGAN and FramePool model weights remain frozen. Only the latent noise is updated.
### TE-Directed Optimization
**Purpose:** Freeze the WGAN and MTtrans models and optimize the latent noise to increase the predicted TE score of generated candidates.
```bash
python scripts/optimize_te_mrl.py \
-gpu 0 \
-task te \
-bs 64 \
-s 10 \
--output-dir outputs/optimization_te
```
During this process, the WGAN and MTtrans model weights remain frozen.
Therefore, this workflow is latent-space optimization rather than fine-tuning of the pretrained models.
### Training
UTRGAN supports WGAN-GP training.
The trainable components in this workflow are:
- Generator;
- Critic.
Xpresso, FramePool, and MTtrans do not participate in this training entry point.
To run one epoch using the complete data-processing pipeline and the UTRdb2 dataset included in the repository:
```bash
python scripts/train.py \
-gpu 0 \
-bs 64 \
-lr 5 \
-mxl 128 \
-dim 40 \
--epochs 1 \
--output-dir outputs/train_1epoch
```
According to the upstream README, `-lr 5` corresponds to a learning rate of:
```text
1e-5
```
The adapted training entry point adds the `--epochs` argument and corrects the learning-rate expression in the upstream script so that it is consistent with the documented definition.
To run the full upstream training configuration, use the complete UTRdb2 training dataset for 4000 epochs:
```bash
python scripts/train.py \
-gpu 0 \
-bs 64 \
-lr 5 \
-mxl 128 \
-dim 40 \
--epochs 4000 \
--output-dir outputs/train_full
```
This retrains the WGAN-GP Generator and Critic.
Full training can require substantial runtime. Training progress should be evaluated according to available accelerator resources and training logs.
Generated checkpoints are saved in the specified output directory and do not overwrite the official pretrained weights under:
```text
weight/
```
# OneScience Official Resources
| Platform | OneScience Main Repository | Skills Repository |
| --- | --- | --- |
| Gitee | [OneScience](https://gitee.com/onescience-ai/onescience) | [OneSkills](https://gitee.com/onescience-ai/oneskills) |
| GitHub | [OneScience](https://github.com/onescience-ai/OneScience) | [OneSkills](https://github.com/onescience-ai/oneskills) |
# Citation and License
- Upstream implementation: https://github.com/ciceklab/UTRGAN
- Xpresso: https://github.com/vagarwal87/Xpresso
- FramePool: https://github.com/Karollus/5UTR
- MTtrans: https://github.com/holab-hku/MTtrans
- The upstream UTRGAN project is distributed under the [CC BY-NC-SA 2.0](https://creativecommons.org/licenses/by-nc-sa/2.0/) license and is intended for academic and non-commercial use. Commercial use requires permission from the upstream authors.
- Third-party models, datasets, pretrained weights, and software components remain subject to their respective original licenses and terms of use.
- The DCU adaptation in this repository continues to use the same CC BY-NC-SA 2.0 license. |