oncodsl / Read docs /PROMPT_module_ranking.md
govindbalki's picture
Upload folder using huggingface_hub
0fff343 verified
|
Raw
History Blame Contribute Delete
5.74 kB
# Prompt β€” ranked list of coordinated gene modules (coherence runs)
When "Prefer coordinated gene modules" (coherence) is on, show a ranked list of the coordinated gene **modules** the run explored β€” ordered by the module's **combined held-out AUROC** (the group's aggregate score vs the label, on unseen patients), with each module's coherence and its individual gene AUROCs/ranks shown. This lets a real co-expressed program (e.g. cell_cycle) surface as a GROUP even when no single gene is the top separator. Read the named code first; don't assume. Airgap stays absolute. `pytest` + `tsc` after.
## 0. Why / honesty
A module's combined AUROC MUST be measured on the **held-out** split (the same train/test split the run used). Larger groups can trivially fit the training patients, so only the held-out group-AUROC is meaningful. The aggregate is parameter-free (mean of the module's genes β†’ AUROC vs label), so there's no model to fit and no leak. Modules were *explored by the GP* so they carry selection bias; the held-out combined AUROC is exactly the honest check on that.
## 1. Harvest the modules from the run
IMPORTANT β€” verified data shape: engine_v2 persists each candidate with a **flat `gene_ids` list** (`engine_v2/gp.py` ~line 136, `list(population[i].feature_ids())`) plus a typed `program_repr` string. It does NOT store separated per-`Select` gene-sets. So:
- **Default (v1 of this feature): a "module" = a candidate's full distinct `gene_ids` set.** Iterate the persisted population across all generations (`GET /runs/{id}/population/{gen}` β†’ `candidates`, which is the FULL population, not the SSE-trimmed top-12; harvest every generation that's persisted), take each candidate's `gene_ids`, dedupe modules as unordered opaque-ID sets, keep size as a column.
- **Optional (only if cleanly splitting multi-Select programs matters): parse `program_repr`** to extract individual `Select` sets β€” the frontend already has a parser (`web/lib/programRepr.ts`); a backend parser would be new work. Don't do this unless the flat-set version proves insufficient.
- Drop singletons if you want modules β‰₯2 genes; exact-set dedupe is fine for v1 (skip Jaccard merging unless needed).
## 2. Score each module (all on the run's train/test split)
First reproduce the run's EXACT held-out split: the module-scoring endpoint must rebuild the same train/test partition the run used β€” `make_split(M.index, y, test_size, random_state=seed, stratify=objective.binary)` with the run's stored `seed` (and matching `test_size`/stratify), over the anonymised matrix from `_prepare_lab_data(target, dataset)`. If the split doesn't match, "held-out" is meaningless. For every distinct module compute, on that split:
- **Combined held-out AUROC** β€” aggregate the module's genes into one per-patient score with the **mean** (matching `Reduce(mean)`), compute orientation-agnostic AUROC vs the label on the **held-out** patients. This is the primary sort key. (For TMB/continuous targets use the run's metric, e.g. |spearman|, not AUROC.)
- **Coherence** β€” mean absolute pairwise correlation among the module's genes, on **train** (the same quantity the coherence prior rewards). Shows *why* it's a module.
- **Size** β€” number of genes.
- **Per-gene** β€” each gene's single-gene AUROC and rank/N from the existing diagnostic (`/diagnostic/...-rank`), so the user sees the individual strengths inside the group.
- Add an endpoint (or extend the result endpoint) that returns this module ranking for a given run, opaque-ID only for the ranking itself.
## 3. UI β€” ranked modules
- Show this section only when coherence was on for the run (gate on the run's coherence flag).
- Render a ranked list (cards or a table), default sorted by **combined held-out AUROC** descending, with columns: combined AUROC Β· coherence Β· size. Let the user re-sort by coherence.
- Each module expands to its genes: revealed **symbol**, single-gene AUROC, rank/N β€” reuse the highlighted-genes styling. Mark genes that are in a reference set (p16 / cell_cycle) with the existing colour scheme, so a cell_cycle-heavy module is visually obvious.
- Header copy + a "?" explaining plainly: "Each row is a group of genes that move together (a coordinated module the engine explored). Combined AUROC = how well the group's average score separates the label on patients it never saw. A group can beat its individual genes β€” that's the point of looking at modules."
- Place it near the result, after the single-gene ranking. It does NOT replace the winner/ranking; it's an additional view.
## 4. Airgap
- The module ranking itself crosses the wire as **opaque IDs + scores**. Reveal **symbols only for the modules actually displayed** (bounded β€” same discipline as revealing a winner; never the whole map). If the list is long, reveal lazily per expanded module.
- Combined AUROC and coherence are computed on the anonymised matrix; no names needed to rank.
## CONSTRAINTS
- Engine name-blind; reveal only the displayed modules' genes. No change to the GP search itself β€” this is a post-run analysis over the persisted population.
- Held-out combined AUROC only; never in-sample. No model fitting in the aggregate (mean only) so there's nothing to leak.
## Checkpoint
- A coherence-on HPV run shows a ranked list of modules by combined held-out AUROC; expanding a module reveals its genes with per-gene AUROC/rank and reference-set marking.
- A cell_cycle-heavy module (if the run explored one) appears with a high combined AUROC even though its genes aren't individually top β€” demonstrating the group view.
- Ranking payload is opaque-only; revealed symbols are bounded to displayed modules; airgap tests green; `tsc`/`pytest` clean.