Gaia / README.md
deeeech's picture
Align card with ORR-only, AUC-only, and BioBench release
a12b8dc verified
|
Raw
History Blame
3.26 kB
metadata
license: cc-by-nc-4.0
tags:
  - oncology
  - spatial-biology
  - benchmarks

Gaia

Gaia is a world model for solid-tumor oncology.

Public benchmark notebooks

The public benchmark release is separated by evaluation scale. Each notebook is fully executed, explains every reported score, and links its compact evaluation script and input tables.

  • Cohort-level benchmarks: drug perturbation priors, the complete 44-row trial-ORR source ledger, direct predicted-versus-observed ORR comparisons, and controls.
  • Patient-level benchmarks: the current CRC and cSCC patient-response scores, ROC AUC evaluation, exact label controls, and CRC measured on-treatment biology.
  • Biology-level benchmarks: the 289-row BioBench protocol, exact KO/OE scoring rules, row-family inventory, and KO gene-direction audit.

The complete benchmark package is in the public Gaia GitHub repository.

Released results

Evaluation Result
Cohort ORR, 44 disease–drug rows Pearson 0.650; Spearman 0.594; MAE 10.25 ORR percentage points
CRC pretreatment response, 11 patients ROC AUC 0.800
cSCC pretreatment response, 12 patients ROC AUC 0.778
CRC measured on-treatment biology, 11 patients Best descriptive mean cosine 0.304 at rollout step 4
BioBench sender knockout 257/289 direction passes (88.9%)
BioBench sender overexpression 257/281 scoreable direction passes (91.5%); 8 rows not applicable
BioBench KO gene-direction audit 0/173 audited ligand rows had every observed pathway gene decrease; median fraction decreasing 61.9%

BioBench contains 183 ligand–receptor pathway rows and 106 regional tissue-program rows. The separate per-gene audit is available for 173 ligand rows. It rules out an all-pathway-genes-down explanation on those rows, not all possible global-shift confounding; 10 ligand rows and the 106 regional rows do not have the same archived gene-level audit.

Reproduction boundary

The GitHub release reproduces metrics, controls, tables, and figures from compact exported benchmark data. A public model-adapter runner owns the cohort, patient, and BioBench orchestration and score construction; the actual model, weights, tissue loader, and encoder remain behind the adapter. The included released-table adapter is explicitly a deterministic smoke test, not fresh inference. The repository contains no training code.

This Hugging Face repository currently contains the project card only; model weights are not currently published here.

The released benchmarks are research evaluations, not prospective clinical validation or a clinical decision system.

License

Repository-authored release material is available under CC BY-NC 4.0. Commercial use is not permitted under that license. Underlying publications, study records, databases, and source datasets retain their own terms.

Project overview: https://blog.alunadata.com/