Spaces:
Sleeping
Sleeping
| # Prompt — engine_v2 full DSL grammar + objective explainers (FINAL, run this one) | |
| Two things in one: (A) make the engine compose programs from the ENTIRE DSL, and (B) upgrade the objective "?" help. This is the single source of truth — it supersedes the earlier Effect/Split and full-DSL drafts; do not run those as well. Strongly-typed GP over the complete operator set. Keep the airgap (genes opaque). Keep existing objectives/presets/Streamlit working (simple programs remain a valid region, so prior results still reproduce). | |
| ## 0. READ FIRST (don't assume) | |
| Read the engine_v2 grammar, interpreter, GP operators, program_repr, the shared graph parser, the available data (expression matrix, clinical fields, labels), and the existing objective + tooltip code. Verify types/signatures/names and build on what's there. | |
| # PART A — Full DSL grammar | |
| ## A1. TYPES (nouns) | |
| Cohort, Matrix, Vector (per-patient score), FeatureSet (opaque gene IDs), Predicate, Scalar, Model, Outcome. | |
| ## A2. OPERATORS (verbs) — ALL of them, typed | |
| - `Load` → Cohort: the implicit ROOT — expression Matrix (opaque gene columns) + clinical {stage, age, …} + labels {msi, tmb}. Always present as the starting data; matrix genes stay opaque. | |
| - `Select(Matrix, FeatureSet)` → Matrix. FeatureSet terminal = 1..genes_per_set opaque IDs (full pool, or the prefilter shortlist when prefilter is on). | |
| - `Reduce(Matrix, Agg)` → Vector. Agg ∈ {mean, median, max, min, var}. | |
| - `Combine(Vector, Vector, Op)` → Vector. Op ∈ {add, sub, mul, div(safe), mean}. | |
| - `Split(Cohort|Vector, Predicate)` → two branches → recombine per-patient → Vector. Predicate is NAME-BLIND (a threshold on a Reduce'd score, or a clinical variable like stage). Guards: min subgroup size, ONE level of Split only (no Split-in-Split), deterministic recombine. | |
| - `Associate(Vector, target, kind)` → Scalar. kind ∈ {pearson, spearman}. Observational association with a target column (msi/tmb via Load). | |
| - `Effect(Vector, target, adjust=[stage,age])` → Scalar. Adjusted association = residualize the Vector on the confounders, then associate. Observational backdoor adjustment — "only as good as the measured confounders" (label it so). | |
| - `Fit(Vector|Vectors, labels)` → Model ; `Apply(Model, Cohort)` → Vector (predicted probabilities). Logistic for the MSI label. | |
| - `Search(Matrix, objective, k)` → FeatureSet. SEE A4 — self-referential; must be bounded. | |
| ## A3. OUTPUT + FITNESS (handle multiple terminal types) | |
| A valid program outputs Vector, Scalar, or Model. Normalise to a single fitness given the objective's target + direction; everything on TRAIN (k-fold), TEST touched once for the winner: | |
| - Vector → apply the objective metric (AUROC for msi; signed correlation for tmb). | |
| - Model → Apply → predicted Vector → metric. | |
| - Scalar from Associate/Effect against the objective's target → that statistic IS the fitness (re-evaluated held-out, in the objective's direction). | |
| - The objective stays MINIMAL: target + direction + honest held-out evaluation. Everything else (genes, aggregations, adjust?, branch?, fit?) lives in the program. | |
| - Permutation null re-evaluates the winning program (its exact structure) on permuted targets. | |
| ## A4. Search inside programs — the self-referential case (bounded) | |
| `Search` is what the GP engine itself does, so a program containing `Search` is a nested/meta-search. To honour "the entire DSL," INCLUDE it but strictly bounded: small k, shallow nested budget, hard time/size caps, memoised, introduced at a LOW rate. Document it as the recursive/meta operator. If nested search exceeds caps → degenerate (worst fitness). If it proves too costly, FLAG it and gate behind a toggle — do not silently drop it. | |
| ## A5. TYPED GP (over the full grammar) | |
| - Type-safe init (ramped half-and-half), crossover, and mutation — only combine/replace subtrees whose return types match. | |
| - TERMINATION GUARANTEE: every type must be closable within max-depth. Provide a minimal closer for each type (Vector via Reduce(Matrix-terminal, agg); Scalar via Associate(Vector, target); etc.) so init/mutation never produce an unclosable tree. | |
| - One seed governs ALL randomness (init, selection, crossover, mutation, FeatureSet sampling, Search). | |
| - Bloat control: λ penalises node count; enforce max-depth and max-nodes; Search and Split obey these caps. | |
| - Degenerate guards: empty FeatureSet forbidden; protected division; constant/NaN output → worst fitness; Split min subgroup size. | |
| ## A6. VISIBILITY | |
| Initialization and mutation must introduce the richer operators (Effect/adjust, Split, Fit, and—rarely—Search) at a non-trivial rate, so the population genuinely contains programs using them (visible on tiles and in graphs). They survive into the winner only if they improve held-out fitness. | |
| ## A7. PARSER / GRAPH / FORMAT HELP | |
| - Extend the ONE shared parser + React Flow renderer to draw every node type: Select, Reduce, Combine, Split (branch→merge), Associate, Effect (adjust wrapper), Fit/Apply, Search (nested). Reuse existing node styling; add distinct but on-theme styles for new ones. | |
| - Update the "Format & examples" panel to document the full grammar with one short annotated example per operator. | |
| ## A8. CONSTRAINTS | |
| - Airgap test green: genes opaque everywhere (engine, program_repr, population endpoint); only clinical fields (stage/age) and labels (msi/tmb) are named; no gene symbols leak. | |
| - Existing engine_v2 objectives/presets/Streamlit untouched; simple Select→Reduce programs remain valid so prior results reproduce. | |
| - The full grammar is a vastly larger search space — recommend a bigger default budget and note runs will be messier early; surface the expected runtime increase. | |
| - All new UI pieces are Client Components. | |
| ## A9. OBJECTIVE FIX (signed-negative TMB) | |
| If not already done: the mutation-burden (TMB) objective must reward the SIGNED negative correlation between the program's score and TMB (low score ↔ high TMB), NOT |spearman|. Update every UI label that shows "association (|spearman|)" accordingly. Permutation null re-evaluates the winner under this signed objective. | |
| # PART B — Objective "?" explainers (UI, rich click-to-open) | |
| Upgrade the four Objective "?" icons to the SAME rich click-to-open explainer used for the parameters (hover = short summary; click = detailed modal). Reuse that component; presentation-only. ALSO FIX the typo in the Mutation burden tooltip: "brokNA spell-checker" → "broken spell-checker". | |
| Shared intro line (top of each modal): "An objective is the rule that scores every program — the fitness the engine maximises. It's computed only from a program's per-patient output and a target column (never gene names), so it can't smuggle in the answer. It sets what 'good' means; the engine then composes DSL programs to satisfy it." | |
| ### MSI separation | |
| SHORT: "Reward programs whose score separates MSI-H from MSS (measured by AUROC). It rewards ANY separator, so it often finds shortcut genes, not the cause." | |
| DETAILED: | |
| - What it optimises: a program whose per-patient score ranks MSI-H patients above MSS. | |
| - Represented as: { target: msi (the subtype label), metric: AUROC }. The engine sees only the program's score + the MSI label — no gene names. | |
| - AUROC: pick a random MSI-H and a random MSS patient; AUROC is the chance the score ranks the MSI-H one higher. 0.5 = coin flip, 1.0 = perfect. We use orientation-agnostic AUROC (a perfectly inverted score isn't discarded) and AUROC rather than accuracy because MSI-H is a ~15% minority. | |
| - Caveat: MSI-H and MSS differ in thousands of genes, so this rewards consequences and coincidences as readily as causes — it tends to grab shortcut genes, not the MMR cause. | |
| ### Mutation burden | |
| SHORT: "Reward programs whose score is negatively associated with mutation burden (low score ↔ high TMB) — the broken spell-checker's fingerprint. A sharper proxy for the cause, but still association, not proof." | |
| DETAILED: | |
| - What it optimises: a program whose score goes DOWN as mutations go UP — the directional signature of a broken DNA-repair gene (switch it off → mutations accumulate). The "spell-checker" is the MMR set (MLH1, MSH2, MSH6, PMS2). | |
| - Represented as: { target: tmb (total mutation count), metric: correlation, direction: negative }. Engine sees only the program's score + the TMB numbers. | |
| - Why "negative" (not magnitude): we reward the SIGNED negative correlation. Magnitude (|spearman|) would also reward genes that go UP with TMB — anti-MMR distractors — diluting the target. | |
| - Correlation ≠ causation (honest): this is a mechanism-SHAPED association — a sharper proxy for the cause than predicting the label, but it does not prove causation. Co-regulated or confounded genes can mimic the same low-expression↔high-TMB pattern. Recovering MLH1 would corroborate known biology, not prove it. | |
| - Toward causal: the program itself can choose to ADJUST for confounders via the Effect operator (negative correlation after removing stage/age) — the most causal the observational data honestly allows. The engine decides whether adjusting helps; we don't hardcode it. | |
| ### Survival (Stage 2) | |
| SHORT: "Would reward programs whose score relates to how long patients survive. Not built yet." | |
| DETAILED: "Target = survival outcome/time; metric = association with survival (e.g. log-rank / hazard). Would reward scores that track prognosis. Not built yet (Stage 2)." | |
| ### Unsupervised (Stage 2) | |
| SHORT: "Would give no target at all — reward programs that capture the strongest natural structure, then you interpret what they found. Not built yet." | |
| DETAILED: "No target column. Reward = how much real structure the score captures (e.g. variance explained / cluster separation), checked for robustness; you interpret the discovery afterward. Not built yet (Stage 2)." | |
| ### Shared footer note (small, in each modal) | |
| "The metric and target above can be written in the DSL itself (Associate / Effect / Fit against a Load'ed label); what stays outside the DSL is the compass — what to maximise, in which direction, judged honestly on held-out data." | |
| # CHECKPOINT | |
| Run once and show me: (a) a few program_repr from the population that use the newer operators (Effect, Split, Fit, and any Search); (b) the winner's program_repr — which operators it chose; and (c) confirm the four objective "?" modals open with the new content and the typo is fixed. Confirm the airgap test is green and existing presets/Streamlit still work. | |