oncodsl / Read docs /PROMPT_hpv_dataset.md
govindbalki's picture
Upload folder using huggingface_hub
0fff343 verified
|
Raw
History Blame Contribute Delete
7.76 kB

Prompt β€” Add HPV (head & neck) as a second dataset, with a top-level cancer selector

Add the TCGA HNSC (head & neck) cohort and an HPV-detection objective, behind a new top-level "Cancer / problem" selector. The DSL, engine, airgap, parameters, live view, population, program graph, and result/verdict are SHARED and unchanged β€” only the dataset-specific config swaps (objectives, biology panel, reference sets, diagnostics, post-hoc targets). Read the named code first; don't assume. Do this in two stages.

GUIDING PRINCIPLE: the engine is single. Everything that changes per cancer is config around it.

================================================================= ## STAGE 1 β€” Backend: the HNSC dataset + HPV objective

1a. Data pipeline (the main lift)

  • FIRST verify the data source: confirm the cBioPortal TCGA HNSC study id and that its clinical data carries HPV status (HPV+/HPVβˆ’). Mirror the COADREAD provenance approach documented in CLAUDE.md (LFS-resolved media URL). State the study id + HPV field you used.
  • Build a second processed cohort alongside the colorectal one: HNSC expression.parquet + clinical/labels including an hpv_status label (HPV+ / HPVβˆ’), restricted to samples with a called HPV status and complete expression. Put schema constants (an HNSC_PROCESSED_DIR, HNSC filenames, the HPV-status column + how +/βˆ’ is derived) in data_pipeline/schema.py β€” do not hardcode in build. Ensure the build writes hpv_status into the labels so cohort.labels["hpv_status"] exists after Load.
  • LOADING β€” IMPORTANT: dsl.Load (operators.py ~line 62) only schema-resolves the exact string "processed"; ANY other string is treated as a literal directory path. So do NOT write Load("processed_hnsc"). Instead, EITHER (a) extend Load to map named sources ("processed"β†’COADREAD dir, "processed_hnsc"β†’schema.HNSC_PROCESSED_DIR), OR (b) pass the HNSC directory Path straight to Load(hnsc_dir) β€” Load's existing else branch already loads any directory containing clinical.parquet + expression.parquet, so option (b) needs no change to Load.

1b. Objective

  • Add a binary objective for HPV: spec { target: "hpv", metric: "auroc" }, reusing the SAME orientation-agnostic binary-AUROC machinery as MSI. Concretely: extend V2Objective.target (currently Literal["msi","tmb","none"]) to include "hpv" with binary=True; route its score_vector through the AUROC path and worst_scoreβ†’0.5 (identical to MSI); add an HPV_OBJECTIVE and dispatch hpv+auroc in engine_v2.fitness.objective_from_spec. No new operators.

1c. Data loading per (dataset, target)

  • _prepare_lab_data (api/app.py ~line 177) is currently keyed by target and always uses Load("processed") (colorectal). Generalise it to key by (dataset, target): add a dataset argument; for dataset="hnsc", target="hpv", load the HNSC cohort, anonymise its expression, and set y = (hpv_status == HPV+). Keep colorectal (dataset="coadread", msi/tmb/none) exactly as-is.

1d. API

  • RunRequest gains a dataset: str field (default "coadread"). _worker passes it through to _prepare_lab_data. Validate dataset+objective combinations (e.g. hpv only valid for hnsc; msi/tmb only for coadread).
  • REFERENCE_SETS (api/app.py ~line 106) becomes per-dataset: colorectal keeps MMR/immune; HNSC gets HPV reference sets β€” e.g. "HPV_markers" = {CDKN2A, …} and a "cell_cycle" / E2F-target set (define from a small curated list in one place, like MMR_GENES). /evaluate and any diagnostics select the reference sets for the active dataset.
  • Diagnostics: the MSI/TMB-rank diagnostics are colorectal-specific β€” gate them to that dataset; HNSC can have its own (or none) for now.
  • Airgap unchanged: anonymise HNSC expression; HPV status is a label (target), never a gene name. Extend the airgap tests to the HNSC path (no gene-symbol leak).

================================================================= ## STAGE 2 β€” Frontend: cancer selector + per-dataset config + full "?" coverage

2a. Top-level Cancer / problem selector

  • Add a segmented control (same style as the Objective control) ABOVE the Objective row: Colorectal (TCGA COADREAD) | Head & Neck (TCGA HNSC). (No side "what swaps" panel β€” that was mockup annotation only.)
  • Drive everything from a dataset registry: each entry = { label, dataset_id, objectives, biologyPanel, referenceSets, diagnostics, postHocTargets, axisLabels }. The page renders the selected entry's config around the unchanged engine. Selecting a cancer sets dataset on the run request and swaps the objective options, biology panel, reference sets, and diagnostics.

2b. HPV biology panel β€” include the schematic

  • Add a "What is HPV? (the biology)" collapsed-by-default disclosure, mirroring the existing "What is MMR?" panel EXACTLY (same collapsible pattern). Embed the HPV reference SVG from repo root hpv_reference.svg (already created) verbatim, responsive (width 100%, height auto). Show it for the Head & Neck dataset; show the MMR panel for Colorectal.

2c. HPV objective + rich "?" explainer

  • Add the "HPV detection" objective card for the HNSC dataset (objective spec {target:"hpv", metric:"auroc"}), plus "Unsupervised" (same engine).
  • Give it the SAME rich click-to-open ParamHelp modal as the other objectives, in plain English (no assumed knowledge). Content: what it optimises ("a score that separates HPV+ from HPVβˆ’ tumours"); how it's scored (AUROC: 0.5 = coin-flip, 1.0 = perfect; minority ~15–25% HPV+, so AUROC not accuracy); represented as {target: hpv, metric: AUROC}; and the honest note that this is detection/recovery of a known viral signature (p16/CDKN2A, cell-cycle), not new causation. Use the same intro/footer pattern as the other objective modals.

2d. Full "?" coverage β€” every objective, parameter, result metric, and chart axis

This is a hard requirement, including for HPV:

  • Objectives: all have rich "?" explainers (above).
  • Parameters: shared engine β€” the existing ParamHelp "?" already cover them; no change.
  • Result metrics (Held-out, Permutation p, Nodes, Genes) and the verdict / post-hoc "?" and inline text must be objective-aware and read correctly for HPV (held-out = AUROC separation of HPV+/βˆ’; the post-hoc/verdict framing is the supervised one β€” "real and useful" vs "not reliable" β€” NOT the unsupervised two-questions framing).
  • Charts: the fitness-curve y-axis label + its "?" for the HPV objective = "separation (AUROC)" with the AUROC explanation (same as MSI). Ensure the objective-aware label/tooltip records include the hpv case.
  • All "?" text in plain English explaining the concept, consistent with the rest of the Lab.

2e. Reveal & evaluate

  • For HNSC, the reference-set toggle shows the HPV sets (HPV_markers / cell_cycle) instead of MMR/immune; overlap + reveal work the same way. Expected win: revealing the winner shows it recovered CDKN2A/p16 + cell-cycle genes.

CONSTRAINTS

  • Colorectal behaviour, airgap tests, and existing runs unchanged.
  • The engine / DSL / parameters / live view / population / program graph are identical across datasets β€” only config swaps.

Checkpoint

After Stage 1: a CLI/API HPV run on HNSC completes and separates HPV+/βˆ’ (held-out AUROC should be high). After Stage 2: the cancer selector swaps to Head & Neck, the "What is HPV?" schematic shows, an HPV run gives a strong held-out AUROC with a plain verdict, and reveal recovers p16/CDKN2A + cell-cycle genes. Airgap tests green; colorectal unchanged.