Spaces:
Running
Running
File size: 5,501 Bytes
fc23fc4 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 | ---
title: LEADBOARD - ADMET, Kinase and Toxicity Prediction Benchmark
emoji: π―
colorFrom: blue
colorTo: gray
sdk: docker
app_port: 7860
# hf_oauth λ₯Ό μΌμΌ νλ«νΌμ΄ OAUTH_CLIENT_ID/SECRET λ₯Ό 컨ν
μ΄λμ λ£μ΄μ€λ€.
hf_oauth: true
pinned: false
short_description: Benchmark for drug prediction tools - ADMET, kinase, tox
tags:
- drug-discovery
- admet
- leaderboard
- benchmark
- cheminformatics
- molecular-property-prediction
- qsar
- toxicity
- herg
- solubility
- kinase
- chembl
- cell-painting
- preclinical
---
# LEADBOARD
**A benchmark for drug property prediction tools.** One yardstick, one discipline at a time.
Submit predictions for a held-out set of molecules. We score them against labels
we never release, and place your tool on the board for that discipline.
π **[Open the leaderboard](https://huggingface.co/spaces/FINAL-Bench/leadboard)**
---
## What is measured
Most published benchmarks split their molecules at random. That is the wrong
question. A tool is used to rank compounds nobody has measured yet, so the honest
test is: *train on what was known by a cut-off year, predict what came after.*
We ran both splits on identical data with an identical model. On the hERG
cardiotoxicity board, AUROC was **0.818** under a random split and **0.606** under
a time split. Same molecules, same code β only the question changed. Every board
here uses the harder one.
## What we publish before you submit
For each board, before any entry arrives:
| Published | Why it matters |
|---|---|
| **Untrained baselines** | Constant prediction, nearest neighbour, Morgan+LightGBM. If a tool cannot beat these, the board says so. |
| **Experimental noise floor** | How far apart two labs land when they measure the same compound. Nothing can be measured below it. |
| **Split grade and answer grade** | Exactly how the test set was cut, and whether the labels can be looked up anywhere. |
| **Data and scorer fingerprints** | The SHA of the exact test file and scoring code behind every score. |
A score without its noise floor is a number without a unit. Both are shown.
## Grading notation
Each board carries two letters, for example `[T/P2]`.
**Split grade** β what the board asks of a tool
- `T` time split by first-report year
- `S` scaffold split by Murcko core
- `R` random split (we do not open these)
**Answer grade** β whether an entrant can look the answer up
- `P1` public source, our curation
- `P2` public source, our unit conversion and selection define this revision
- `P3` labels held privately
- `P4` prospective β the answer does not exist yet
## Boards
21 boards across 7 disciplines, 18,382 held-out compounds.
- π« **Absorption** β Solubility
- π₯ **Metabolism** β CYP3A4, CYP2D6, CYP2C9
- β οΈ **Toxicity** β hERG
- π― **Potency** β AChE, MAOB, COX2
- π **Kinase** β EGFR, JAK2, PI3KΞ±, FLT3, VEGFR2, CDK2, HER2, ABL1, BRAF, KIT, ALK
- π¬ **Cell / Phenotype** β JUMP Cell Painting morphology, 115,689 compounds, 16 profile axes
- π₯ **Clinical** β Post-marketing withdrawal, year-matched controls
The withdrawal board is worth a note. Withdrawal rate tracks approval decade
(7.8% in the 1990s, 1.2% in the 2010s), so a predictor that reads nothing but the
approval year scores AUROC 0.636. After year-matching the controls it scores 0.504
β chance. That is the version we opened.
## How to enter
1. Pick a board and download its test set β structures only, no labels.
2. Predict with your own tool. Anything goes: a trained model, a physics engine,
a language model, a rule of thumb.
3. Upload a CSV of `compound_id,prediction`. Sign in with your Hugging Face account.
4. Scoring runs off-platform on hardware that holds the labels. The Space never
sees them.
The Space carries a filled-in prompt and a runnable skeleton for each board, so a
first entry does not require building anything from scratch.
### Repeated submissions
New scores are revealed under the ladder rule (Blum & Hardt, ICML 2015): a fresh
score replaces your best only when it beats it by more than the noise floor.
Otherwise your previous best stands. This is what stops a leaderboard from being
won by whoever submits the most times.
## Data and licences
- **ChEMBL 37** (EMBL-EBI) β CC BY-SA 3.0. Re-curated; attribution travels with every board card.
- **JUMP Cell Painting Consortium** β CC0.
- Withdrawal board assembled from public regulatory records.
Non-commercial research benchmark. Test sets carry structures only; labels are
never distributed.
---
## νκ΅μ΄ μμ½
**μ μ½ μμΈ‘ λꡬλ₯Ό, λΆμΌλ³λ‘, κ°μ μ£λλ‘ μ½λλ€.**
μ λ΅μ λ°°ν¬νμ§ μμ΅λλ€. λμ μ±μ μκ° λ¨Όμ 곡κ°ν©λλ€ β μ무κ²λ νμ΅νμ§ μμ
κΈ°μ€μ , κ·Έ λΆμΌμ μ€ν μ‘μ λ°λ₯, μ μλ§λ€μ λ°μ΄ν°Β·μ±μ κΈ° μ§λ¬Έ.
무μμ λΆν μ νλ¦° μ§λ¬Έμ
λλ€. λꡬλ μμ§ μ무λ μ¬μ§ μμ νν©λ¬Όμ μμλ₯Ό λ§€κΈ°λ
λ° μ°μ΄λ, μ μ§ν μνμ **μ΄λ ν΄κΉμ§ μλ €μ§ κ²μΌλ‘ λ°°μ°κ³ κ·Έ λ€μ λμ¨ κ²μ
λ§νλ κ²**μ
λλ€. κ°μ μλ£Β·κ°μ λͺ¨λΈλ‘ μ¬ λ³΄λ©΄ hERG μ¬μ₯λ
μ± AUROC κ° λ¬΄μμ
λΆν 0.818, μκ° λΆν 0.606 μ΄μμ΅λλ€. μ¬κΈ° λΆλ¬Έμ μ λΆ μ΄λ €μ΄ μͺ½μ μλλ€.
7κ° λΆμΌ 21λΆλ¬Έ Β· ν
μ€νΈ νν©λ¬Ό 18,382. νλ©΄μ νκ΅μ΄Β·μμ΄λ₯Ό μλμΌλ‘ κ°λ¦
λλ€.
---
κ·κ²© v1.1 (8κ° μ‘°ν) Β· **FINAL-Bench / VIDRAFT**
|