Spaces:
Sleeping
Sleeping
File size: 10,504 Bytes
0fff343 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 | # Prompt — engine_v2 full DSL grammar + objective explainers (FINAL, run this one)
Two things in one: (A) make the engine compose programs from the ENTIRE DSL, and (B) upgrade the objective "?" help. This is the single source of truth — it supersedes the earlier Effect/Split and full-DSL drafts; do not run those as well. Strongly-typed GP over the complete operator set. Keep the airgap (genes opaque). Keep existing objectives/presets/Streamlit working (simple programs remain a valid region, so prior results still reproduce).
## 0. READ FIRST (don't assume)
Read the engine_v2 grammar, interpreter, GP operators, program_repr, the shared graph parser, the available data (expression matrix, clinical fields, labels), and the existing objective + tooltip code. Verify types/signatures/names and build on what's there.
# PART A — Full DSL grammar
## A1. TYPES (nouns)
Cohort, Matrix, Vector (per-patient score), FeatureSet (opaque gene IDs), Predicate, Scalar, Model, Outcome.
## A2. OPERATORS (verbs) — ALL of them, typed
- `Load` → Cohort: the implicit ROOT — expression Matrix (opaque gene columns) + clinical {stage, age, …} + labels {msi, tmb}. Always present as the starting data; matrix genes stay opaque.
- `Select(Matrix, FeatureSet)` → Matrix. FeatureSet terminal = 1..genes_per_set opaque IDs (full pool, or the prefilter shortlist when prefilter is on).
- `Reduce(Matrix, Agg)` → Vector. Agg ∈ {mean, median, max, min, var}.
- `Combine(Vector, Vector, Op)` → Vector. Op ∈ {add, sub, mul, div(safe), mean}.
- `Split(Cohort|Vector, Predicate)` → two branches → recombine per-patient → Vector. Predicate is NAME-BLIND (a threshold on a Reduce'd score, or a clinical variable like stage). Guards: min subgroup size, ONE level of Split only (no Split-in-Split), deterministic recombine.
- `Associate(Vector, target, kind)` → Scalar. kind ∈ {pearson, spearman}. Observational association with a target column (msi/tmb via Load).
- `Effect(Vector, target, adjust=[stage,age])` → Scalar. Adjusted association = residualize the Vector on the confounders, then associate. Observational backdoor adjustment — "only as good as the measured confounders" (label it so).
- `Fit(Vector|Vectors, labels)` → Model ; `Apply(Model, Cohort)` → Vector (predicted probabilities). Logistic for the MSI label.
- `Search(Matrix, objective, k)` → FeatureSet. SEE A4 — self-referential; must be bounded.
## A3. OUTPUT + FITNESS (handle multiple terminal types)
A valid program outputs Vector, Scalar, or Model. Normalise to a single fitness given the objective's target + direction; everything on TRAIN (k-fold), TEST touched once for the winner:
- Vector → apply the objective metric (AUROC for msi; signed correlation for tmb).
- Model → Apply → predicted Vector → metric.
- Scalar from Associate/Effect against the objective's target → that statistic IS the fitness (re-evaluated held-out, in the objective's direction).
- The objective stays MINIMAL: target + direction + honest held-out evaluation. Everything else (genes, aggregations, adjust?, branch?, fit?) lives in the program.
- Permutation null re-evaluates the winning program (its exact structure) on permuted targets.
## A4. Search inside programs — the self-referential case (bounded)
`Search` is what the GP engine itself does, so a program containing `Search` is a nested/meta-search. To honour "the entire DSL," INCLUDE it but strictly bounded: small k, shallow nested budget, hard time/size caps, memoised, introduced at a LOW rate. Document it as the recursive/meta operator. If nested search exceeds caps → degenerate (worst fitness). If it proves too costly, FLAG it and gate behind a toggle — do not silently drop it.
## A5. TYPED GP (over the full grammar)
- Type-safe init (ramped half-and-half), crossover, and mutation — only combine/replace subtrees whose return types match.
- TERMINATION GUARANTEE: every type must be closable within max-depth. Provide a minimal closer for each type (Vector via Reduce(Matrix-terminal, agg); Scalar via Associate(Vector, target); etc.) so init/mutation never produce an unclosable tree.
- One seed governs ALL randomness (init, selection, crossover, mutation, FeatureSet sampling, Search).
- Bloat control: λ penalises node count; enforce max-depth and max-nodes; Search and Split obey these caps.
- Degenerate guards: empty FeatureSet forbidden; protected division; constant/NaN output → worst fitness; Split min subgroup size.
## A6. VISIBILITY
Initialization and mutation must introduce the richer operators (Effect/adjust, Split, Fit, and—rarely—Search) at a non-trivial rate, so the population genuinely contains programs using them (visible on tiles and in graphs). They survive into the winner only if they improve held-out fitness.
## A7. PARSER / GRAPH / FORMAT HELP
- Extend the ONE shared parser + React Flow renderer to draw every node type: Select, Reduce, Combine, Split (branch→merge), Associate, Effect (adjust wrapper), Fit/Apply, Search (nested). Reuse existing node styling; add distinct but on-theme styles for new ones.
- Update the "Format & examples" panel to document the full grammar with one short annotated example per operator.
## A8. CONSTRAINTS
- Airgap test green: genes opaque everywhere (engine, program_repr, population endpoint); only clinical fields (stage/age) and labels (msi/tmb) are named; no gene symbols leak.
- Existing engine_v2 objectives/presets/Streamlit untouched; simple Select→Reduce programs remain valid so prior results reproduce.
- The full grammar is a vastly larger search space — recommend a bigger default budget and note runs will be messier early; surface the expected runtime increase.
- All new UI pieces are Client Components.
## A9. OBJECTIVE FIX (signed-negative TMB)
If not already done: the mutation-burden (TMB) objective must reward the SIGNED negative correlation between the program's score and TMB (low score ↔ high TMB), NOT |spearman|. Update every UI label that shows "association (|spearman|)" accordingly. Permutation null re-evaluates the winner under this signed objective.
# PART B — Objective "?" explainers (UI, rich click-to-open)
Upgrade the four Objective "?" icons to the SAME rich click-to-open explainer used for the parameters (hover = short summary; click = detailed modal). Reuse that component; presentation-only. ALSO FIX the typo in the Mutation burden tooltip: "brokNA spell-checker" → "broken spell-checker".
Shared intro line (top of each modal): "An objective is the rule that scores every program — the fitness the engine maximises. It's computed only from a program's per-patient output and a target column (never gene names), so it can't smuggle in the answer. It sets what 'good' means; the engine then composes DSL programs to satisfy it."
### MSI separation
SHORT: "Reward programs whose score separates MSI-H from MSS (measured by AUROC). It rewards ANY separator, so it often finds shortcut genes, not the cause."
DETAILED:
- What it optimises: a program whose per-patient score ranks MSI-H patients above MSS.
- Represented as: { target: msi (the subtype label), metric: AUROC }. The engine sees only the program's score + the MSI label — no gene names.
- AUROC: pick a random MSI-H and a random MSS patient; AUROC is the chance the score ranks the MSI-H one higher. 0.5 = coin flip, 1.0 = perfect. We use orientation-agnostic AUROC (a perfectly inverted score isn't discarded) and AUROC rather than accuracy because MSI-H is a ~15% minority.
- Caveat: MSI-H and MSS differ in thousands of genes, so this rewards consequences and coincidences as readily as causes — it tends to grab shortcut genes, not the MMR cause.
### Mutation burden
SHORT: "Reward programs whose score is negatively associated with mutation burden (low score ↔ high TMB) — the broken spell-checker's fingerprint. A sharper proxy for the cause, but still association, not proof."
DETAILED:
- What it optimises: a program whose score goes DOWN as mutations go UP — the directional signature of a broken DNA-repair gene (switch it off → mutations accumulate). The "spell-checker" is the MMR set (MLH1, MSH2, MSH6, PMS2).
- Represented as: { target: tmb (total mutation count), metric: correlation, direction: negative }. Engine sees only the program's score + the TMB numbers.
- Why "negative" (not magnitude): we reward the SIGNED negative correlation. Magnitude (|spearman|) would also reward genes that go UP with TMB — anti-MMR distractors — diluting the target.
- Correlation ≠ causation (honest): this is a mechanism-SHAPED association — a sharper proxy for the cause than predicting the label, but it does not prove causation. Co-regulated or confounded genes can mimic the same low-expression↔high-TMB pattern. Recovering MLH1 would corroborate known biology, not prove it.
- Toward causal: the program itself can choose to ADJUST for confounders via the Effect operator (negative correlation after removing stage/age) — the most causal the observational data honestly allows. The engine decides whether adjusting helps; we don't hardcode it.
### Survival (Stage 2)
SHORT: "Would reward programs whose score relates to how long patients survive. Not built yet."
DETAILED: "Target = survival outcome/time; metric = association with survival (e.g. log-rank / hazard). Would reward scores that track prognosis. Not built yet (Stage 2)."
### Unsupervised (Stage 2)
SHORT: "Would give no target at all — reward programs that capture the strongest natural structure, then you interpret what they found. Not built yet."
DETAILED: "No target column. Reward = how much real structure the score captures (e.g. variance explained / cluster separation), checked for robustness; you interpret the discovery afterward. Not built yet (Stage 2)."
### Shared footer note (small, in each modal)
"The metric and target above can be written in the DSL itself (Associate / Effect / Fit against a Load'ed label); what stays outside the DSL is the compass — what to maximise, in which direction, judged honestly on held-out data."
# CHECKPOINT
Run once and show me: (a) a few program_repr from the population that use the newer operators (Effect, Split, Fit, and any Search); (b) the winner's program_repr — which operators it chose; and (c) confirm the four objective "?" modals open with the new content and the typo is fixed. Confirm the airgap test is green and existing presets/Streamlit still work.
|