File size: 4,393 Bytes
240c1e3 8dada60 240c1e3 44d0475 8dada60 ec45e2c 8dada60 8dedbc5 8dada60 44d0475 8dada60 44d0475 8dada60 44d0475 8dada60 44d0475 8dada60 44d0475 8dada60 44d0475 8dada60 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 | ---
license: apache-2.0
tags:
- molecular-toxicity
- drug-discovery
- graph-neural-network
- cheminformatics
- tox21
- interpretability
- gineconv
language:
- en
---
# MolScreen: GNN for Molecular Toxicity Screening
**MolScreen** is a Graph Isomorphism Network with Edge features (GINEConv) trained on the Tox21 dataset for multi-task molecular toxicity prediction. It combines GNN-based predictions with gradient-based atom attribution and LLM-generated triage notes to make toxicity screening interpretable and actionable.
Developed by Veer Goradia.
## Model Description
Drug toxicity is one of the leading causes of late-stage clinical trial failure, costing billions in wasted development. MolScreen addresses two key gaps in existing approaches: most models give a probability but no explanation, and none integrate a human-readable summary for non-expert decision makers.
**Architecture:**
- GINEConv backbone — uses edge features (bond type, aromaticity, conjugation) that plain GCN/GIN discard
- Multi-task classification across 12 Tox21 endpoints simultaneously
- Masked BCE loss for Tox21's missing labels (5,800–7,200 valid labels per task)
- Gradient-based atom attribution: highlights which molecular substructures drove the prediction
- LLM triage layer: generates a 3-sentence actionable summary (SMILES + flagged probabilities + top-3 attributed atoms per task)
**Total parameters:** ~374,000
**Inference time:** ~0.7 seconds end-to-end (excluding LLM API call)
**Training time:** ~6 minutes on laptop CPU (60 epochs)
## Performance
### Random Split (10-seed validation)
| Metric | Value |
|--------|-------|
| Mean ROC-AUC | 0.8442 ± 0.0073 |
### Scaffold Split (10-seed validation)
| Metric | Value |
|--------|-------|
| Mean ROC-AUC | 0.7716 ± 0.0036 |
### Baseline Comparison (Scaffold Split)
| Model | ROC-AUC |
|-------|---------|
| **MolScreen (GINEConv)** | **0.7716** |
| Plain GCN | 0.7481 |
| Random Forest (Morgan FP) | 0.7465 |
| XGBoost (Morgan FP) | 0.7378 |
| Plain GIN | 0.7249 |
| Chen et al. SSL-GCN (2021) | 0.7570 |
MolScreen beats the best non-pretrained baseline (GCN) by 2.4 points. The GINEConv edge-awareness advantage over plain GIN (+4.7 points) confirms that bond-level features matter for generalization.
### Per-Task ROC-AUC (Random Split, representative run)
| Endpoint | AUC |
|----------|-----|
| NR-AR | 0.8518 |
| NR-AR-LBD | 0.9119 |
| NR-AhR | 0.8938 |
| NR-Aromatase | 0.8788 |
| NR-ER | 0.7486 |
| NR-ER-LBD | 0.8400 |
| NR-PPAR-gamma | 0.8918 |
| SR-ARE | 0.8125 |
| SR-ATAD5 | 0.8762 |
| SR-HSE | 0.7374 |
| SR-MMP | 0.9156 |
| SR-p53 | 0.8713 |
## Training Data
- **Dataset:** Tox21 (7,823 valid compounds, 12 toxicity endpoints)
- **Assay categories:** Nuclear receptor panel (NR-*) and stress response panel (SR-*)
- **Class imbalance:** 2.6%–16% positive rate per task
- **Splits:** Random split (80/10/10) and scaffold split (chemotype-based, harder generalization test)
## Intended Use
- Early-stage drug candidate toxicity screening
- Prioritizing which compounds advance to lab testing
- Interpretability analysis via atom-level attribution maps
- Research on GNN-based molecular property prediction
## How to Use
```python
import torch
from huggingface_hub import hf_hub_download
# Load model from HuggingFace
model_path = hf_hub_download(repo_id="vgoradia/MolScreen", filename="molscreen_best.pt")
model = torch.load(model_path, map_location='cpu')
model.eval()
# Or try the live Streamlit app — no code needed
```
## Live Demo
Try the interactive app (draw a molecule, get toxicity predictions + atom attribution):
[MolScreen Streamlit App](https://mol-screen-32jrti5gpuajmd8wse5u3f.streamlit.app)
## Publications & Acceptances
- **NeurIPS 2026 AIDaR Workshop** — ACCEPTED (poster, Paris, December 12 2026)
- **IEEE BIBM 2026 AIPBDA Workshop** — SUBMITTED (Paper ID S49203)
- **Regeneron STS 2027** — In Progress (November 2026)
- **AAAI-27 AISI Track** — Submitted (Submission #598)
## GitHub
[github.com/vgoradia/mol-screen](https://github.com/vgoradia/mol-screen)
## Citation
@misc{goradia2026molscreen,
title={MolScreen: Interpretable Multi-Task Molecular Toxicity Screening
with Graph Neural Networks and LLM Triage},
author={Goradia, Veer},
year={2026},
note={NeurIPS 2026 AIDaR Workshop}
}
## License
Apache 2.0
## Contact
Veer Goradia · vgoradia07@gmail.com |