|
Download README.md from vgoradia/MolScreen: direct link, hf CLI and curl.
- Browser
- Download file 4.39 kB
-
https://huggingface.co/vgoradia/MolScreen/resolve/main/README.md
- Command line
-
hf download hf://vgoradia/MolScreen/README.md
-
curl -L -o README.md https://huggingface.co/vgoradia/MolScreen/resolve/main/README.md
4.39 kB
| license: apache-2.0 | |
| tags: | |
| - molecular-toxicity | |
| - drug-discovery | |
| - graph-neural-network | |
| - cheminformatics | |
| - tox21 | |
| - interpretability | |
| - gineconv | |
| language: | |
| - en | |
| # MolScreen: GNN for Molecular Toxicity Screening | |
| **MolScreen** is a Graph Isomorphism Network with Edge features (GINEConv) trained on the Tox21 dataset for multi-task molecular toxicity prediction. It combines GNN-based predictions with gradient-based atom attribution and LLM-generated triage notes to make toxicity screening interpretable and actionable. | |
| Developed by Veer Goradia. | |
| ## Model Description | |
| Drug toxicity is one of the leading causes of late-stage clinical trial failure, costing billions in wasted development. MolScreen addresses two key gaps in existing approaches: most models give a probability but no explanation, and none integrate a human-readable summary for non-expert decision makers. | |
| **Architecture:** | |
| - GINEConv backbone β uses edge features (bond type, aromaticity, conjugation) that plain GCN/GIN discard | |
| - Multi-task classification across 12 Tox21 endpoints simultaneously | |
| - Masked BCE loss for Tox21's missing labels (5,800β7,200 valid labels per task) | |
| - Gradient-based atom attribution: highlights which molecular substructures drove the prediction | |
| - LLM triage layer: generates a 3-sentence actionable summary (SMILES + flagged probabilities + top-3 attributed atoms per task) | |
| **Total parameters:** ~374,000 | |
| **Inference time:** ~0.7 seconds end-to-end (excluding LLM API call) | |
| **Training time:** ~6 minutes on laptop CPU (60 epochs) | |
| ## Performance | |
| ### Random Split (10-seed validation) | |
| | Metric | Value | | |
| |--------|-------| | |
| | Mean ROC-AUC | 0.8442 Β± 0.0073 | | |
| ### Scaffold Split (10-seed validation) | |
| | Metric | Value | | |
| |--------|-------| | |
| | Mean ROC-AUC | 0.7716 Β± 0.0036 | | |
| ### Baseline Comparison (Scaffold Split) | |
| | Model | ROC-AUC | | |
| |-------|---------| | |
| | **MolScreen (GINEConv)** | **0.7716** | | |
| | Plain GCN | 0.7481 | | |
| | Random Forest (Morgan FP) | 0.7465 | | |
| | XGBoost (Morgan FP) | 0.7378 | | |
| | Plain GIN | 0.7249 | | |
| | Chen et al. SSL-GCN (2021) | 0.7570 | | |
| MolScreen beats the best non-pretrained baseline (GCN) by 2.4 points. The GINEConv edge-awareness advantage over plain GIN (+4.7 points) confirms that bond-level features matter for generalization. | |
| ### Per-Task ROC-AUC (Random Split, representative run) | |
| | Endpoint | AUC | | |
| |----------|-----| | |
| | NR-AR | 0.8518 | | |
| | NR-AR-LBD | 0.9119 | | |
| | NR-AhR | 0.8938 | | |
| | NR-Aromatase | 0.8788 | | |
| | NR-ER | 0.7486 | | |
| | NR-ER-LBD | 0.8400 | | |
| | NR-PPAR-gamma | 0.8918 | | |
| | SR-ARE | 0.8125 | | |
| | SR-ATAD5 | 0.8762 | | |
| | SR-HSE | 0.7374 | | |
| | SR-MMP | 0.9156 | | |
| | SR-p53 | 0.8713 | | |
| ## Training Data | |
| - **Dataset:** Tox21 (7,823 valid compounds, 12 toxicity endpoints) | |
| - **Assay categories:** Nuclear receptor panel (NR-*) and stress response panel (SR-*) | |
| - **Class imbalance:** 2.6%β16% positive rate per task | |
| - **Splits:** Random split (80/10/10) and scaffold split (chemotype-based, harder generalization test) | |
| ## Intended Use | |
| - Early-stage drug candidate toxicity screening | |
| - Prioritizing which compounds advance to lab testing | |
| - Interpretability analysis via atom-level attribution maps | |
| - Research on GNN-based molecular property prediction | |
| ## How to Use | |
| ```python | |
| import torch | |
| from huggingface_hub import hf_hub_download | |
| # Load model from HuggingFace | |
| model_path = hf_hub_download(repo_id="vgoradia/MolScreen", filename="molscreen_best.pt") | |
| model = torch.load(model_path, map_location='cpu') | |
| model.eval() | |
| # Or try the live Streamlit app β no code needed | |
| ``` | |
| ## Live Demo | |
| Try the interactive app (draw a molecule, get toxicity predictions + atom attribution): | |
| [MolScreen Streamlit App](https://mol-screen-32jrti5gpuajmd8wse5u3f.streamlit.app) | |
| ## Publications & Acceptances | |
| - **NeurIPS 2026 AIDaR Workshop** β ACCEPTED (poster, Paris, December 12 2026) | |
| - **IEEE BIBM 2026 AIPBDA Workshop** β SUBMITTED (Paper ID S49203) | |
| - **Regeneron STS 2027** β In Progress (November 2026) | |
| - **AAAI-27 AISI Track** β Submitted (Submission #598) | |
| ## GitHub | |
| [github.com/vgoradia/mol-screen](https://github.com/vgoradia/mol-screen) | |
| ## Citation | |
| @misc{goradia2026molscreen, | |
| title={MolScreen: Interpretable Multi-Task Molecular Toxicity Screening | |
| with Graph Neural Networks and LLM Triage}, | |
| author={Goradia, Veer}, | |
| year={2026}, | |
| note={NeurIPS 2026 AIDaR Workshop} | |
| } | |
| ## License | |
| Apache 2.0 | |
| ## Contact | |
| Veer Goradia Β· vgoradia07@gmail.com |