| --- |
| license: mit |
| library_name: deepbioisostere |
| tags: |
| - chemistry |
| - drug-discovery |
| - molecular-optimization |
| - bioisostere |
| - graph-neural-network |
| --- |
| |
| # DeepBioisostere |
|
|
| Model weights and fragment library for **DeepBioisostere**, a deep generative |
| model that performs end-to-end bioisosteric replacement: it selects which |
| fragment of a molecule to replace, what to replace it with, and how to attach |
| it, so as to satisfy several property targets at once. |
|
|
| - Paper: [Nature Communications (2026)](https://doi.org/10.1038/s41467-026-75512-9) |
| - Code: https://github.com/Hwoo-Kim/DeepBioisostere |
| - Project page: https://hwoo-kim.github.io/DeepBioisostere/ |
|
|
| ## Usage |
|
|
| These files are fetched automatically; you do not normally download them by hand. |
|
|
| ```bash |
| pip install deepbioisostere |
| deepbioisostere generate -s "Cc1ccc2cnc(N(C)CCc3ccccn3)nc2c1" -t mw=0 -t logp=-1 |
| ``` |
|
|
| ```python |
| from deepbioisostere import Conditioner, DeepBioisostere, Generator |
| |
| properties = ["logp", "mw"] |
| model = DeepBioisostere.from_pretrained(properties=properties) |
| generator = Generator( |
| model=model, |
| conditioner=Conditioner(phase="generation", properties=properties), |
| properties=properties, |
| ) |
| df = generator.generate([("Cc1ccc2cnc(N(C)CCc3ccccn3)nc2c1", {"mw": 0, "logp": -1})]) |
| ``` |
|
|
| ## Contents |
|
|
| ### Checkpoints |
|
|
| Each checkpoint is conditioned on a specific set of properties. The filename |
| encodes the property set, sorted alphabetically. |
|
|
| | File | Properties | Variant | |
| |---|---|---| |
| | `DeepBioisostere_logp.pt` | logP | standard | |
| | `DeepBioisostere_mw.pt` | MW | standard | |
| | `DeepBioisostere_qed.pt` | QED | standard | |
| | `DeepBioisostere_sa.pt` | SA | standard | |
| | `DeepBioisostere_logp_mw.pt` | logP + MW | standard | |
| | `DeepBioisostere_mw_qed.pt` | MW + QED | standard | |
| | `DeepBioisostere_qed_sa.pt` | QED + SA | standard | |
| | `DeepBioisostere_logp_mw_ablation.pt` | logP + MW | ablation | |
| | `DeepBioisostere_mw_qed_ablation.pt` | MW + QED | ablation | |
| | `DeepBioisostere_qed_sa_ablation.pt` | QED + SA | ablation | |
|
|
| The remaining pairs (logP+QED, logP+SA, MW+SA) were not trained. Ablation |
| variants set `use_subgraph_AMPN=False`, which changes the architecture, and are |
| selected with `from_pretrained(..., ablation=True)`. |
|
|
| ### Fragment library |
|
|
| | File | Description | |
| |---|---| |
| | `fragment_library.csv` | 145,854 insertion fragments with frequency, BRICS type and train/val/test split | |
| | `frag_features.pkl` | Pre-parsed graph features, keyed by the csv SMILES | |
| | `frag_brics_maskings.pkl` | BRICS-type compatibility masks (training only) | |
|
|
| The two `.pkl` files are derived from the csv and are regenerated automatically |
| if absent, which takes roughly an hour of CPU. They are published so that a |
| first run is a download instead. |
|
|
| > **Note on fragment identity.** The features are keyed on the exact SMILES |
| > string in `fragment_library.csv`, not on a re-canonicalised form. RDKit's |
| > canonical SMILES writer changed after the library was built under 2022.03: |
| > under 2026.03, 1,712 of the fragments are written differently (stereocentre |
| > descriptors flip). The underlying molecules and their parsed features are |
| > unchanged — this is a serialisation difference only — but the csv string is |
| > the authoritative key, since training data references fragments by row index. |
| |
| ## Citation |
| |
| ```bibtex |
| @article{kim2026deepbioisostere, |
| title = {Autonomous bioisosteric replacement for multi-property optimization in drug design}, |
| author = {Kim, Hyeongwoo and Moon, Seokhyun and Zhung, Wonho and |
| Kim, Shinwoo and Lim, Jaechang and Kim, Woo Youn}, |
| journal = {Nature Communications}, |
| year = {2026}, |
| doi = {10.1038/s41467-026-75512-9}, |
| } |
| ``` |
| |
| License: MIT. |
| |