File size: 5,173 Bytes
7056daa | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 | ---
license: cc-by-4.0
language:
- en
tags:
- biology
- genomics
- gene-regulatory-network
- protein-protein-interaction
- soybean
- plant-genomics
pretty_name: SoyGRN — soybean GRN models, inferred networks, and expression data
---
# SoyGRN: trained models and inferred networks for soybean transcriptome-scale GRN inference
Reusable data and model release for the paper *"Transcriptome-scale gene regulatory network
inference in soybean: GPU-scalable tree ensembles versus a masked-expression foundation model"*
(Genome Biology). This is the deposited resource referenced by the manuscript's *Availability of
data and materials* declaration and Supplementary Tables S1–S2. **Analysis code is on GitHub:
https://github.com/k821209/soygrn**
All identifiers are Wm82.a6 (Phytozome Gmax_880_v6.0) gene loci for soybean and Araport11 for
Arabidopsis. Protein embeddings throughout are mean-pooled ESM-2 `esm2_t12_35M_UR50D` (480-d).
## Trained models
### `grn_fm_W.npz` — SoyGRN-FM (the network *is* the weight tensor)
The trained masked-expression foundation model. NumPy archive:
- `W` — float32, shape **(3419 TF × 48358 target)**: the signed regulator→target weight tensor; this single tensor is the entire GRN (no other reg→target path in the model).
- `tf_idx` — int64 (3419): row → index into `genes` for each transcription factor.
- `genes` — <U15 (48358): gene locus IDs indexing the columns of `W` (and `tf_idx`).
Read out the network as the top-30 regulators per target by `|W[:,j]| · std(ψ(activity))`, sign from
`sign(W)`; or apply the invariance re-ranking (Supplementary Source Code S1, `grn_fm_invar.py`,
γ=3) to reproduce `edges_fm_invar.tsv`.
```python
import numpy as np
z = np.load("grn_fm_W.npz", allow_pickle=True)
W, tf_idx, genes = z["W"], z["tf_idx"], z["genes"]
tf_names = genes[tf_idx] # 3419 TF locus IDs
# top-10 regulators of a target gene j:
j = list(genes).index("Glyma.01G000100")
top = np.argsort(-np.abs(W[:, j]))[:10]
for r in top: print(tf_names[r], W[r, j]) # weight sign = activation(+)/repression(-)
```
### `ppi_fm_intact.pt` — PPI foundation model (PyTorch)
Symmetric Siamese network over ESM-2 embeddings, trained on **all 27,422** Arabidopsis IntAct
experimental-physical interactions (no holdout). This is the model applied **zero-shot to soybean**
in the paper (AUROC 0.953 on the ortholog-transferred gold). `dict` with `state_dict` + `config`.
Order-invariant: `sigmoid(model(emb_a, emb_b))` = interaction probability. Architecture in
`code/grn_ppi_fm_final.py` (class `PPIFM`).
### `binding_promcnn.pt` — TF→promoter binding model (PyTorch)
The best binding recoverer (leave-TF-out AUROC 0.825): promoter 1D-CNN (max+mean pool) + ESM-2 TF
embedding, trained on **all 15,000** Arabidopsis FunTFBS binding pairs. Input = 10-channel 500 bp
promoter (4 one-hot ACGT + 6 z-scored dinucleotide DNA-shape: twist, roll, propeller, slide, rise,
tilt) and the TF ESM-2 embedding. `dict` with `state_dict` + `config`. Architecture in
`code/grn_binding_cnn_final.py` (class `PromCNN`).
## Inferred soybean networks (TF, target, score[, sign]; tab-separated, header row)
| File | Method | Edges | Signed |
|---|---|---|---|
| `edges_clr.tsv` | CLR | 680,381 | no |
| `edges_grnboost_all.tsv` | GRNBoost2 | 4,000,000 | no |
| `edges_fm_invar.tsv` | SoyGRN-FM (invariance re-ranked, top-30/target) | 1,450,740 | yes (+1 activation / −1 repression) |
## Protein embeddings & ortholog map
- `gmax_prot_esm.npz` — soybean proteome ESM-2 embeddings (`emb`, `genes`); inputs to the PPI/binding models for soybean.
- `glyma_to_athal_rbh.tsv` — soybean↔Arabidopsis reciprocal-best-hit ortholog map (12,939 pairs; columns: glyma, athal, bitscore), used for cross-species PPI/binding transfer.
## Large primary inputs (deposited separately as Zenodo files; not in this folder due to size)
- `gmax_gene_tpm.npz` — *G. max* gene-level TPM, 48,358 × 8,177 (1.5 GB) + `gmax_sample_meta.csv` (sample→SRA study).
- `athal_gene_tpm.npz` — *A. thaliana* compendium, 27,572 × 29,089 (3.2 GB) + `athal_sample_meta.csv`.
- `athal_prot_esm.npz` — Arabidopsis proteome ESM-2 embeddings (training inputs for PPI/binding models).
Per-sample SRA run accessions and study (BioProject) IDs are in the `*_sample_meta.csv` tables.
Public input resources (PlantTFDB/PlantRegMap FunTFBS & GO, IntAct, STRING v12, JASPAR, the Wm82.a6
reference) are obtained from their original providers under their own licences and are not
re-deposited here. See Supplementary Table S1 for the full inventory.
## Code
All analysis code (35 Python scripts: inference, evaluation, GO coherence, binding & PPI recovery,
the two model-training scripts above, figures) is at **https://github.com/k821209/soygrn**. See
Supplementary Table S2 for per-model hyperparameters. `grn_fm_W.npz` is produced by `grn_fm.py`;
`ppi_fm_intact.pt` by `grn_ppi_fm_final.py`; `binding_promcnn.pt` by `grn_binding_cnn_final.py`.
## Manifest
SHA-256 checksums for all deposited files are in `SHA256SUMS.txt`.
## Citation
If you use these models or networks, please cite the paper (DOI on publication) and this deposit.
|