File size: 5,173 Bytes
7056daa
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
---
license: cc-by-4.0
language:
- en
tags:
- biology
- genomics
- gene-regulatory-network
- protein-protein-interaction
- soybean
- plant-genomics
pretty_name: SoyGRN  soybean GRN models, inferred networks, and expression data
---

# SoyGRN: trained models and inferred networks for soybean transcriptome-scale GRN inference

Reusable data and model release for the paper *"Transcriptome-scale gene regulatory network
inference in soybean: GPU-scalable tree ensembles versus a masked-expression foundation model"*
(Genome Biology). This is the deposited resource referenced by the manuscript's *Availability of
data and materials* declaration and Supplementary Tables S1–S2. **Analysis code is on GitHub:
https://github.com/k821209/soygrn**

All identifiers are Wm82.a6 (Phytozome Gmax_880_v6.0) gene loci for soybean and Araport11 for
Arabidopsis. Protein embeddings throughout are mean-pooled ESM-2 `esm2_t12_35M_UR50D` (480-d).

## Trained models

### `grn_fm_W.npz` — SoyGRN-FM (the network *is* the weight tensor)
The trained masked-expression foundation model. NumPy archive:
- `W` — float32, shape **(3419 TF × 48358 target)**: the signed regulator→target weight tensor; this single tensor is the entire GRN (no other reg→target path in the model).
- `tf_idx` — int64 (3419): row → index into `genes` for each transcription factor.
- `genes`<U15 (48358): gene locus IDs indexing the columns of `W` (and `tf_idx`).

Read out the network as the top-30 regulators per target by `|W[:,j]| · std(ψ(activity))`, sign from
`sign(W)`; or apply the invariance re-ranking (Supplementary Source Code S1, `grn_fm_invar.py`,
γ=3) to reproduce `edges_fm_invar.tsv`.

```python
import numpy as np
z = np.load("grn_fm_W.npz", allow_pickle=True)
W, tf_idx, genes = z["W"], z["tf_idx"], z["genes"]
tf_names = genes[tf_idx]                      # 3419 TF locus IDs
# top-10 regulators of a target gene j:
j = list(genes).index("Glyma.01G000100")
top = np.argsort(-np.abs(W[:, j]))[:10]
for r in top: print(tf_names[r], W[r, j])     # weight sign = activation(+)/repression(-)
```

### `ppi_fm_intact.pt` — PPI foundation model (PyTorch)
Symmetric Siamese network over ESM-2 embeddings, trained on **all 27,422** Arabidopsis IntAct
experimental-physical interactions (no holdout). This is the model applied **zero-shot to soybean**
in the paper (AUROC 0.953 on the ortholog-transferred gold). `dict` with `state_dict` + `config`.
Order-invariant: `sigmoid(model(emb_a, emb_b))` = interaction probability. Architecture in
`code/grn_ppi_fm_final.py` (class `PPIFM`).

### `binding_promcnn.pt` — TFpromoter binding model (PyTorch)
The best binding recoverer (leave-TF-out AUROC 0.825): promoter 1D-CNN (max+mean pool) + ESM-2 TF
embedding, trained on **all 15,000** Arabidopsis FunTFBS binding pairs. Input = 10-channel 500 bp
promoter (4 one-hot ACGT + 6 z-scored dinucleotide DNA-shape: twist, roll, propeller, slide, rise,
tilt) and the TF ESM-2 embedding. `dict` with `state_dict` + `config`. Architecture in
`code/grn_binding_cnn_final.py` (class `PromCNN`).

## Inferred soybean networks (TF, target, score[, sign]; tab-separated, header row)

| File | Method | Edges | Signed |
|---|---|---|---|
| `edges_clr.tsv` | CLR | 680,381 | no |
| `edges_grnboost_all.tsv` | GRNBoost2 | 4,000,000 | no |
| `edges_fm_invar.tsv` | SoyGRN-FM (invariance re-ranked, top-30/target) | 1,450,740 | yes (+1 activation / −1 repression) |

## Protein embeddings & ortholog map

- `gmax_prot_esm.npz` — soybean proteome ESM-2 embeddings (`emb`, `genes`); inputs to the PPI/binding models for soybean.
- `glyma_to_athal_rbh.tsv` — soybeanArabidopsis reciprocal-best-hit ortholog map (12,939 pairs; columns: glyma, athal, bitscore), used for cross-species PPI/binding transfer.

## Large primary inputs (deposited separately as Zenodo files; not in this folder due to size)

- `gmax_gene_tpm.npz` — *G. max* gene-level TPM, 48,358 × 8,177 (1.5 GB) + `gmax_sample_meta.csv` (sampleSRA study).
- `athal_gene_tpm.npz` — *A. thaliana* compendium, 27,572 × 29,089 (3.2 GB) + `athal_sample_meta.csv`.
- `athal_prot_esm.npz` — Arabidopsis proteome ESM-2 embeddings (training inputs for PPI/binding models).

Per-sample SRA run accessions and study (BioProject) IDs are in the `*_sample_meta.csv` tables.
Public input resources (PlantTFDB/PlantRegMap FunTFBS & GO, IntAct, STRING v12, JASPAR, the Wm82.a6
reference) are obtained from their original providers under their own licences and are not
re-deposited here. See Supplementary Table S1 for the full inventory.

## Code

All analysis code (35 Python scripts: inference, evaluation, GO coherence, binding & PPI recovery,
the two model-training scripts above, figures) is at **https://github.com/k821209/soygrn**. See
Supplementary Table S2 for per-model hyperparameters. `grn_fm_W.npz` is produced by `grn_fm.py`;
`ppi_fm_intact.pt` by `grn_ppi_fm_final.py`; `binding_promcnn.pt` by `grn_binding_cnn_final.py`.

## Manifest

SHA-256 checksums for all deposited files are in `SHA256SUMS.txt`.

## Citation

If you use these models or networks, please cite the paper (DOI on publication) and this deposit.