Gaia / README.md
deeeech's picture
Align card with ORR-only, AUC-only, and BioBench release
a12b8dc verified
|
Raw
History Blame
3.26 kB
---
license: cc-by-nc-4.0
tags:
- oncology
- spatial-biology
- benchmarks
---
# Gaia
Gaia is a world model for solid-tumor oncology.
## Public benchmark notebooks
The public benchmark release is separated by evaluation scale. Each notebook
is fully executed, explains every reported score, and links its compact
evaluation script and input tables.
- **[Cohort-level benchmarks](https://github.com/alunalabs/gaia/blob/main/01_cohort_level_benchmarks.ipynb):** drug perturbation priors, the complete 44-row trial-ORR source ledger, direct predicted-versus-observed ORR comparisons, and controls.
- **[Patient-level benchmarks](https://github.com/alunalabs/gaia/blob/main/02_patient_level_benchmarks.ipynb):** the current CRC and cSCC patient-response scores, ROC AUC evaluation, exact label controls, and CRC measured on-treatment biology.
- **[Biology-level benchmarks](https://github.com/alunalabs/gaia/blob/main/03_biology_level_benchmarks.ipynb):** the 289-row BioBench protocol, exact KO/OE scoring rules, row-family inventory, and KO gene-direction audit.
The complete benchmark package is in the
**[public Gaia GitHub repository](https://github.com/alunalabs/gaia)**.
## Released results
| Evaluation | Result |
| --- | --- |
| Cohort ORR, 44 disease–drug rows | Pearson 0.650; Spearman 0.594; MAE 10.25 ORR percentage points |
| CRC pretreatment response, 11 patients | ROC AUC 0.800 |
| cSCC pretreatment response, 12 patients | ROC AUC 0.778 |
| CRC measured on-treatment biology, 11 patients | Best descriptive mean cosine 0.304 at rollout step 4 |
| BioBench sender knockout | 257/289 direction passes (88.9%) |
| BioBench sender overexpression | 257/281 scoreable direction passes (91.5%); 8 rows not applicable |
| BioBench KO gene-direction audit | 0/173 audited ligand rows had every observed pathway gene decrease; median fraction decreasing 61.9% |
BioBench contains 183 ligand–receptor pathway rows and 106 regional
tissue-program rows. The separate per-gene audit is available for 173 ligand
rows. It rules out an all-pathway-genes-down explanation on those rows, not all
possible global-shift confounding; 10 ligand rows and the 106 regional rows do
not have the same archived gene-level audit.
## Reproduction boundary
The GitHub release reproduces metrics, controls, tables, and figures from
compact exported benchmark data. A public model-adapter runner owns the cohort,
patient, and BioBench orchestration and score construction; the actual model,
weights, tissue loader, and encoder remain behind the adapter. The included
released-table adapter is explicitly a deterministic smoke test, not fresh
inference. The repository contains no training code.
This Hugging Face repository currently contains the project card only; model
weights are not currently published here.
The released benchmarks are research evaluations, not prospective clinical
validation or a clinical decision system.
## License
Repository-authored release material is available under
[CC BY-NC 4.0](https://github.com/alunalabs/gaia/blob/main/LICENSE). Commercial
use is not permitted under that license. Underlying publications, study
records, databases, and source datasets retain their own terms.
Project overview: https://blog.alunadata.com/