oncodsl / Read docs /PROMPT_iterative_discovery.md
govindbalki's picture
Upload folder using huggingface_hub
0fff343 verified
|
Raw
History Blame Contribute Delete
4.37 kB

Prompt β€” Iterative unsupervised discovery ("peel off axes")

Add a "find the next axis" capability to the UNSUPERVISED objective: after a run, residualise the discovered split out of the data and re-run, so the engine surfaces a ranked list of independent axes (Axis 1 = KRT5, Axis 2 = …, each with its own held-out, p, revealed genes, and post-hoc MSI/TMB alignment). Read the named code first; don't assume. Airgap stays absolute (opaque IDs, no labels in the search). Only applies to the unsupervised objective (target "none").

Why this design

To residualise against a prior axis we need that axis's per-patient score over the FULL cohort. Rather than re-parse a program_repr string back into a Node (no such parser exists), we STORE the winner's full-cohort score in the run result and reuse it. So:

PART A β€” engine_v2/pipeline.py

  1. Store full-cohort winner scores. The pipeline already re-executes the winner on ctx_test for held-out scores (run_v2_pipeline_streaming, ~lines 417–431). ALSO execute the winner on a FULL-cohort context (all patients, train+test; labels empty for unsup) and store, in the result's "winning" block (the existing nested dict at line ~462, NOT a new "winner" key), full_scores (per-patient, finite-guarded) and full_sample_ids (str ids aligned to them) β€” alongside the existing test-only holdout_scores/holdout_sample_ids. This is what the next run residualises against.
  2. Residualisation hook. Add an optional param to run_v2_pipeline_streaming (and run_v2_pipeline), e.g. residualize_scores: pd.DataFrame | None (index = sample id, columns = prior axis scores). When provided, IMMEDIATELY after loading M and BEFORE the train/test split, regress every gene column of M on [intercept, *prior scores] (vectorised OLS: resid = M - P @ np.linalg.lstsq(P, M, rcond=None)[0], aligning P to M.index, guarding rank/NaN) and REPLACE M with the residuals (same opaque columns + index). Residualising before the split keeps train and test consistent. (No effect for MSI/TMB β€” only the unsupervised worker passes it.)

PART B β€” api/app.py

  1. RunRequest (~line 294) gains an optional residualize_against: list[str] | None = prior run_ids in this discovery chain.
  2. In _worker (~line 417), for an unsupervised run with residualize_against set: look up each prior run_id in the in-memory run store, pull its stored run.result["winning"]["full_scores"] + ["full_sample_ids"], assemble a DataFrame aligned by sample id, and pass it to the pipeline as residualize_scores. Note _json_finite may have turned non-finite scores into null β€” drop those patients (or treat as NaN) when aligning. Validate the run_ids exist and were unsupervised; ignore the field for MSI/TMB. (In-memory store β‡’ a chain is valid within a server session; note this.)
  3. Keep all payloads opaque-ID-only; extend the airgap test to the residualisation path (scores are opaque-derived numbers β€” no labels, no gene names leak).

PART C β€” Frontend (web/app/Lab.tsx)

  1. After an unsupervised run completes, show a "Find next axis β†’" button. Clicking it POSTs a new run with the SAME params + unsupervised objective and residualize_against = the run_ids of ALL axes discovered so far in this chain.
  2. Accumulate results into a "Discovered axes" stack rendered below the Result: Axis 1, Axis 2, … Each axis is a compact card reusing the existing verdict + post-hoc rendering β€” held-out score, permutation p, revealed gene(s), and "aligns with MSI at AUROC … / TMB …". So the user reads a ranked list of independent axes the engine found blind, each named by its post-hoc alignment.
  3. The chain resets when the user starts a fresh run (not "next axis").

Honesty / scope

  • This is a discovery HEURISTIC: linear residualisation removes the linear component of each prior axis; deeper non-linear structure may remain. State that in the "?" for the axes stack.
  • It does not guarantee MSI ever appears β€” the value is the ranked, blind, honestly-labelled list of axes (and the chance that a deeper axis turns out to be MSI).

Checkpoint

Run unsupervised β†’ Axis 1 should be KRT5. Hit "Find next axis" β†’ Axis 2 should be a DIFFERENT gene/axis (orthogonal to KRT5), with its own held-out/p and a post-hoc MSI/TMB read. Confirm airgap tests stay green and MSI/TMB runs are unaffected.