File size: 9,820 Bytes
57efc81
13d0b34
d38e9d6
 
 
57efc81
3e6f48b
72d7def
57efc81
 
d2a6015
383d326
57efc81
 
75e6a8b
13d0b34
75b3042
d38e9d6
75b3042
0f02e4d
 
75e6a8b
 
 
7bfbc7e
c9d0ebc
 
 
 
 
d2058b2
c9d0ebc
88b3ddf
 
75e6a8b
 
 
7bfbc7e
75b3042
7bfbc7e
383d326
7bfbc7e
 
 
0f02e4d
 
 
 
7bfbc7e
75e6a8b
 
 
0f02e4d
69a850d
 
 
d38e9d6
1d983ea
d38e9d6
0f02e4d
d38e9d6
75e6a8b
1d983ea
 
 
75e6a8b
d38e9d6
1d983ea
 
6286da1
 
d38e9d6
0f02e4d
 
 
 
 
 
1d983ea
d38e9d6
6d6d611
 
 
 
 
 
 
 
d38e9d6
69a850d
75e6a8b
7bfbc7e
75e6a8b
 
 
69a850d
 
 
7bfbc7e
cb0c5af
75e6a8b
69a850d
 
d38e9d6
48fcbed
 
75e6a8b
48fcbed
 
75e6a8b
 
 
48fcbed
 
 
d38e9d6
d2058b2
 
0f02e4d
 
 
 
d2058b2
 
 
 
 
0f02e4d
d2058b2
 
d38e9d6
 
 
 
13d0b34
0f02e4d
d38e9d6
 
d2058b2
 
 
2803daa
 
 
41128ce
 
 
d2058b2
d38e9d6
 
d2a6015
 
1d983ea
383d326
 
69a850d
383d326
6286da1
69a850d
cb0c5af
 
7532dab
 
 
 
 
 
 
 
 
 
d894841
7532dab
 
d2a6015
 
75e6a8b
d2058b2
 
 
 
 
 
2803daa
7532dab
 
 
 
 
 
 
 
d2058b2
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
---
title: PRIMO Benchmark
emoji: 🧬
colorFrom: indigo
colorTo: blue
sdk: gradio
sdk_version: 5.50.0
python_version: "3.10"
app_file: app.py
pinned: false
hf_oauth: true
thumbnail: https://huggingface.co/spaces/PRIMOmics/primo-eval/resolve/main/assets/primo-social-preview.png
---

# PRIMO: Patient Representations in Multi-Omics

**A blind benchmark for omics foundation models.**

PRIMO grades how well a model turns a patient's omics data into a useful
**patient embedding**. You select one modality, embed its datasets, and upload
**one** file; a fixed probe scores each hidden task, and the results roll up into a blind,
per-category leaderboard. The datasets are opaque (`d001`, `d002`…) and you never
see the disease, tissue, or target, which leaves you grading the *embedding*
itself with no room for per-task tuning.

A left **rail** navigates six pages: **Boards** (a grid of every board),
**Board** (one leaderboard at a time), **Tasks**, a **Submit** form (sign in with
Hugging Face), **Contribute**, and **Method**. Every board has its own URL,
`?board=rheumatology-bulk-rna`, and every page too, as `?tab=contribute`, so the
rail links and the open cards are shareable deep links.

Boards also shows **open boards**: greyed-out cards for the omics layers and the
therapeutic areas PRIMO does not cover yet, each linking to Contribute. They are
declared in `boards.py` (`OPEN_BOARDS`) and drop out on their own once the
registry covers that slice. Task categories deliberately get none, because a
category is pinned to a single metric, so an open one would advertise a probe
that does not exist.

🌐 Website: http://primomics.org/ ·
📄 Paper: https://openreview.net/forum?id=v2SA8gHwqo ·
📦 Data: https://huggingface.co/datasets/PRIMOmics/primo

## What's in the data

PRIMO benchmarks any omics modality. It includes **bulk RNA** tasks across
immune-mediated inflammatory diseases and a **single-cell RNA** COVID-19 PBMC
severity task, all with real clinical labels or treatment-induced expression
responses from published cohorts:

- **Gastroenterology**: Crohn's disease, ulcerative colitis (anti-TNF response, severity scores)
- **Dermatology**: atopic dermatitis, psoriasis (severity scores)
- **Rheumatology**: rheumatoid arthritis (joint counts, molecular endotype)
- **Infectious diseases**: COVID-19 severity from single-cell PBMC expression
- **Perturbation response**: adalimumab transfer across inflammatory skin
  diseases, rituximab response in Sjögren salivary gland, and mouse intestinal
  anti-TNF response

## Submission format

One file, one row per (`dataset_id`, `sample_id`), spanning one modality:

- **CSV / TSV / Parquet**: a `dataset_id` column, a `sample_id` column, and one
  numeric column per embedding dimension. Embedding dim may differ per dataset
  (pad short datasets with blank columns; blank/NaN padding columns are dropped
  per dataset).
- **NPZ**: `dataset_ids`, `sample_ids`, and a 2-D `embeddings` array.

The valid `dataset_id`s and how to download each dataset's `expression.h5ad` are
listed in the public `datasets.yaml` manifest. Alignment is by join, so row
order does not matter; every labelled sample of a task must be present with no
NaN/inf, or that task is skipped.

Bulk H5AD files have one row per required submission sample. Single-cell H5AD
files have one sparse raw-count row per cell; opaque cell ids are in
`obs_names`, and the only public cell metadata is `obs["sample_id"]`, which maps
each cell to its opaque collection sample. Submissions remain sample-level: emit
exactly one embedding for every unique `sample_id`, not one embedding per cell.

## How it works

Each dataset is embedded once and scored on every hidden task defined for it.
Rows are scaled to unit length, then per task: standardise on the training folds,
fit `LogisticRegressionCV` (classification) or `RidgeCV` (regression) with the
penalty chosen by repeated inner cross-validation, and predict each held-out
fold. We calculate the native metric per fold and average it over frozen,
subject-grouped partitions: **AUROC** (classification) or **Pearson r**
(regression). A few tasks instead use a fixed train/test split: the probe is fit
once on the training portion and scored on the held-out portion.

Tasks are grouped into **four families**, each reported in its own native
metric. Two families are never merged into one column:

- **Treatment outcome**: response to anti-TNF therapy (**AUROC**)
- **Clinical scores**: disease-severity regression (**Pearson r**)
- **Endotype**: molecular-subtype classification (**AUROC**)
- **Perturbation response**: hidden post-minus-pre DEG decoding (**centered
  Spearman**), where 1 is perfect, 0 is uninformative, and -1 is an inverse gene
  ranking.

A leaderboard shows one column per family. A board holding more than one family
also shows **Mean**, the average of those columns. It is what orders the rows,
but it does mix AUROC, Pearson and centered Spearman, so treat it as a tie-break and
compare models on the family columns.

## Boards

Results are shown as **boards**, self-contained leaderboards over a slice of the
registry: the whole modality, one therapeutic area, one task family. A board
ranks only the models that covered **all of its tasks**, so a submission that
skipped Dermatology is still ranked on Rheumatology. Modality is where we draw
the line: an area or category board never spans two modalities, because an AUROC
on bulk RNA and an AUROC on single-cell are not measuring the same thing.

Partial and failed submissions still get feedback, and their scores always appear
in each board's **per-task** table even when they are not ranked.

## Make a submission

`quickstart.py` is the shortest path: it downloads one modality and embeds its
datasets
(log2(CPM+1) → PCA for bulk; per-cell log2(CP10K+1) → sample mean → PCA for
single-cell) and writes the file the Submit tab wants. Swap its `embed`
function for your encoder and nothing else changes. `example_submission.csv`
shows the expected shape in four lines.

```bash
pip install anndata scikit-learn pandas pyyaml huggingface_hub
python quickstart.py --modality bulk-rna --out submission.parquet
```

## Run the scorer locally

```bash
pip install -r requirements.txt
export HF_TOKEN=...   # read access to the PRIMO datasets
python evaluator.py --modality bulk-rna --submission my_embeddings.parquet
```

## Baselines

`task_results.csv` carries an `is_baseline` flag. Reference submissions we
produce ourselves are published with it set, rendered as `name (baseline)`, and
**ranked in place**. A foundation model losing to raw log-CPM expression is
exactly the result worth publishing, so we keep it in the table rather than
tucked underneath. The published `HVG-1200-genes` baseline uses the 1,200
highest-variance genes after log2(CPM+1), without gene scaling. Its selection is
label-blind but fit over every sample of a dataset.

## Space configuration

- **`hf_oauth: true`** (set above) turns on the Submit tab's *Sign in with
  Hugging Face* button; submitting requires a logged-in HF account.
- Set an **`HF_TOKEN`** Space secret (fine-grained) with: **read** on
  `PRIMOmics/primo` (the public `datasets.yaml` manifest) and
  `PRIMOmics/primo-labels` (the private `tasks.yaml` registry +
  `<task_id>/labels.csv` and perturbation `<task_id>/targets.npz`), and **read + write** on
  `PRIMOmics/primo-results` (the persisted leaderboard).
- Results persist as one normalized `task_results.csv` (`model_name, task_id,
  score, repeat_scores, diagnostics, submitted_at, is_baseline, hf_username`) in the results dataset; the
  leaderboard is recomputed from it by joining the registry, so it survives
  Space restarts.
- A board keeps the **latest result for each `(model name, is_baseline, task)`**,
  so submissions accumulate: a later one fills in the tasks it covers and leaves
  the rest standing, and resubmitting a task replaces just that result. Because a
  task belongs to exactly one modality, one name can hold bulk and single-cell
  results at once without either displacing the other.
- A model name is owned by the account that first submitted it: `hf_username`
  locks it, and Submit refuses a name somebody else holds. It is never rendered
  anywhere: it works as a lock rather than as a credit. Names colliding with a
  per-task column (`Task`, `Family`, `Area`, `Modality`, `Metric`, `Best`) are
  refused too, since model names become column headers.
- The boards show a **Submissions count** per model — the most times any one of
  its task results was resubmitted — so a task tuned against by repeated
  resubmission is visible. Nothing is ever blocked.
- Submitter contact metadata (HF username, email, paper / model links, notes)
  persists to a separate `submissions.csv` in the same **private** results
  dataset, and never reaches the public leaderboard.

## Moving to another Hugging Face org

The three dataset repos are derived from one constant, `ORG` in `evaluator.py`.
The rest of the org name is spelled out and has to be changed by hand:

- `SPACE_REPO` in `deploy_space.py`
- `PUBLIC_REPO` in `quickstart.py` (standalone, so it cannot import `ORG`)
- `SOCIAL_PREVIEW_URL` in `app.py`
- the links in this file — including the YAML `thumbnail:` above — and in
  `pages/*.md`, `../public_dataset_README.md` (the dataset card) and
  `../HF_SETUP.md`

That list is exhaustive as of the `ScientaLab``PRIMOmics` move; `app.py` and
the dataset card were the two it missed the first time.

The theme follows the PRIMO charter: Funnel Display for headings, Funnel Sans for
everything else, Scienta Navy `#080F5F` / PRIMO Cyan `#16B3C0` on Paper
`#F3F8F8`. `colorFrom`/`colorTo` above stay `indigo`/`blue` because Hugging Face
only accepts eight named colours and none of them is cyan.