# SPINE_V2_PLAN.md — Analysis Spine v2 + Visualization Tool **Status:** ✅ **APPROVED 2026-07-13** by Rifqi, with Sofia's sign-off (covers implementing the tool-layer slice — `ToolOutput.kind` + `render_chart` — in this repo). **Delta at approval:** W3 (activate deferred `analyze_*` tools) is **deferred to a later cycle — do not start**. W4 stays gated, W5 stays future-only. FE reminder (plotly rendering) owned by Rifqi, timing TBD. **Delta 2026-07-13 (pr/16, same-day build):** **W2 and W1 are DONE** — see DEV_PLAN §0.6 (V1–V7) for the task rows and evidence. Rifqi ran the §4.4 DDL against dedorch and the **live e2e gate passed in full** (real `POST /api/v2/chat/stream`: viz turn → `message_charts` row keyed by the `done` message_id → `GET /charts` 200 with a valid v1 envelope; chartless turn → 200 empty; injected `render_chart` failure → answer still streams with the data as a table, no row). W1 ✅ (13 local tests, one per CK rule). Planner viz behavior verified in-process EN+ID incl. the viz-infeasible case (Example K was added when the first smoke force-mapped a stand-in column). Two env finds along the way: `eval.chat_sim` was missing from disk+git (restored by Rifqi — its hard-coded user/source ids are still stale) and the local `.env` lagged Go's Supabase-S3 data plane (`storage_provider` — fixed by Rifqi; REPO_STATUS §13 gotcha). **Owner:** Rifqi. Contributors named per workstream (Sofia = tool layer, Harry = Go/dedorch/contract, mentor = FE). **Companions:** [REPO_STATUS.md](REPO_STATUS.md) (built state) · [DEV_PLAN.md](DEV_PLAN.md) (sprint tracker) · [API_CONTRACT_BE_PYTHON.md](API_CONTRACT_BE_PYTHON.md) (live contract). This file is the design + handoff source for the Spine-v2 work; when a workstream lands, DEV_PLAN gets the task rows and this doc gets dated deltas. Status legend (house): ⬜ not started · 🔄 in progress · ✅ done · ⛔ blocked · 🔎 verify · ⏸️ deferred. --- ## 0. Executive summary We evolve the slow path from a **one-shot pipeline** (plan → run → narrate) into a **staged workflow with one bounded self-correction point**, and we grow the tool registry into **four families on one spine**: `check_*`/`retrieve_*` (built) · `analyze_*` (built) · `render_*` (this plan) · `model_*` (future). Inspiration is CoDA (Google, ICLR 2026) and the Julius.ai comparison, with one governing rule: **CoDA's phases, not CoDA's agents.** CoDA pays ~15 LLM calls per output because every phase (viz mapping, design, codegen, debugging, visual evaluation) is its own LLM agent operating on generated code. We keep our declarative engine — the same phases exist here, but only planning is an LLM decision; mapping lives in the planner's existing single call, styling is a fixed preset, "generation" is a deterministic spec builder, and evaluation is deterministic code. Happy path stays **3 LLM calls per turn** (router + planner + assembler); the only growth is an **evidence-gated repair pass (max +1)** that is NOT built until telemetry justifies it. What ships, in order: - **S1a — Quality checkpoint** (0 LLM): after execution, before composition, deterministic code inspects the run (failed branches, empty/degenerate results, later chart-spec sanity). Verdicts: proceed · honest degrade · log-a-repair-candidate. Converts today's silent degrade-and-continue into *explained* degradation and produces the data for the S1b decision. No invariant touched. - **S2 — Visualization tool** (`render_chart`): planner-selected, Pattern A, emits **Plotly JSON** (locked decision, DEV_PLAN §0 deferred row #26 — no matplotlib PNGs, no new dependency). Delivery mirrors traceability: Python-owned `message_charts` store + `GET /api/v1/charts`, SSE stays text-only. Includes formalizing planner **path recipes** (the few-shots become named workflows; viz = any recipe + a `render_chart` tail). - **S1b — Targeted repair (GATED)**: one bounded re-plan of only the failed subgraph, triggered by the checkpoint. Relaxes INV-6 ("no mid-run LLM") → requires team sign-off *and* S1a telemetry showing it pays. - **S3 — `model_*` family (FUTURE, sketch only)**: regression/forecast/segmentation as fixed, audited compute functions on the same spine. Named here so the architecture anticipates it; no build this cycle. What is deliberately **not changing**: no generated code ever executes; INV-4 (LLM never authors numbers); the five query-defense layers; never-throw seams; text-only SSE (`sources` stays `[]`); records-based versioned reports; Go's ownership of dedorch DDL and `analyses_messages`. --- ## 1. Target architecture ``` POST /api/v2/chat/stream (structured_flow turn) │ ├── Input guard ─► Router (1 LLM) ─► analysis-scoped catalog read (metadata only, PII-safe) ▼ PLAN Planner (1 LLM) — classifies the question into a RECIPE, instantiates a staged Task DAG │ ▲ └ validator: 8 existing checks + path-shape checks, re-prompt ≤3 (all PRE-run) │ │ │ └────────────── (S1b, GATED: targeted repair, ≤1 pass, failed subgraph only) ──┐ ▼ │ EXECUTE TaskRunner (0 LLM) — waves, ${t} handoffs, degrade-and-continue │ │ tool families (fixed registry, all never-throw): │ │ check_* / retrieve_* · analyze_* · render_* (S2) · model_* (S3) │ ▼ │ CHECK Quality checkpoint (S1a, 0 LLM, never-throw) │ │ ok → continue · repairable → repair request ────────────────────────────┘ │ unfixable → honest degrade (specific, not "couldn't compute") ▼ COMPOSE Assembler (1 LLM — narrative only, INV-4; structured fields copied by code) ▼ DELIVER SSE: sources[] → status* → chunk → done{message_id} writes: report_inputs · message_charts (S2) · message_traceability FE follow-ups on done: GET /api/v1/traceability · GET /api/v1/charts (S2) later: POST /api/v1/tools/report (records-based, unchanged) ``` **Recipes (the "predetermined workflow" made explicit).** The planner few-shots already are path templates; S2 names them and the validator enforces their shape: | Recipe | Chain | Exists today as | |---|---|---| | R1 descriptive | `retrieve_data → analyze_descriptive` | few-shot A/B | | R2 aggregate / top-N | single grouped-IR `retrieve_data` (± `analyze_aggregate`) | Example G | | R3 trend | `retrieve_data → analyze_trend` | few-shot | | R4 correlation | `retrieve_data → analyze_correlation` | few-shot | | R5 two-metric merge | `retrieve_data ×2 → analyze_merge → …` | Example I | | R6 infeasible | no tasks + `infeasible_reason` → deterministic refusal | Example H | | viz tail (S2) | any of R1–R5 **+ `render_chart`** when the user explicitly asks to plot/visualize | NEW — Example J | CoDA-role mapping (why no new LLM agents): VizMapping → planner args in the existing call · Design Explorer → fixed style preset inside `render_chart` · Code Generator/Debug → deterministic spec builder (nothing to debug) · Visual Evaluator → checkpoint spec checks. **LLM budget:** happy path 3 calls (unchanged) · S1b failure path max 4 · report unchanged at 1. --- ## 2. Workstreams & sequencing | # | Workstream | Owner | Status | Gate | |---|---|---|---|---| | W1 | S1a quality checkpoint (+ telemetry for the S1b decision) | Rifqi | ✅ 2026-07-13 | none — code-only, 0 LLM | | W2 | S2 `render_chart` tool + chart store + `GET /charts` + planner slice (recipes + Example J) | Rifqi + Sofia | ✅ 2026-07-13 (DDL run + live e2e ALL PASS) | Harry: migration handoff + `done.chart_count` still open · FE: Plotly render (Rifqi to remind) | | W3 | Activate the 4 deferred `analyze_*` tools (comparison, contribution, profile, segment) | Sofia + Rifqi | ⏸️ | **Deferred at approval (2026-07-13) — do not start until Rifqi re-opens** | | W4 | S1b targeted repair pass | Rifqi | ⏸️ | **team sign-off on INV-6 relaxation + S1a telemetry evidence** | | W5 | S3 `model_*` family | — | ⏸️ | future design review; sketch in §7 | Sequencing: **W2 leads** (visible product value), **W1 lands alongside or immediately after** (invisible to users, unblocks nothing, improves everything). W3 deferred at approval. W4 only with evidence. Parallel dispatch is safe: W1 and W2 are file-disjoint except `coordinator.py` (W1) — coordinate that one file. --- ## 3. W1 — S1a Quality checkpoint **What it is:** a deterministic, never-throw inspection of `RunState` between `TaskRunner.run` and `Assembler.assemble`. No LLM. No user-visible new surface — only better inputs to the Assembler and structured logs. **Files:** - NEW `src/agents/slow_path/checkpoint.py` — `assess(run_state, task_list) -> RunAssessment`. - `src/agents/slow_path/coordinator.py` — call site between runner and assembler; pass the assessment into `assemble(...)`. - `src/agents/slow_path/prompt.py` — render the assessment as a short "execution assessment" block in the assembler's human content so the narrative names *what specifically* failed/degraded instead of generic caveats. - `src/agents/slow_path/schemas.py` — `RunAssessment` model (pydantic): per-task verdicts + overall verdict (`ok | degraded | failed`) + `repair_candidates: list[RepairCandidate{task_id, reason}]`. **v1 checks (all deterministic):** - CK1 all tasks failed → overall `failed`; coordinator returns a deterministic honest-failure answer (mirrors the existing infeasible path shape; record stays non-substantive so it can't hit the report floor). - CK2 `retrieve_data` returned 0 rows and a dependent consumed it → flag task + downstream. - CK3 table output truncated at the 10k cap → flag (answer must say "based on the first 10,000 rows"). - CK4 `analyze_trend` produced a single bucket → flag (the pr/13 1970-bucket class). - CK5 all-null column consumed by an `analyze_*` task → flag. - CK6 (lands with W2) chart-spec sanity — see §4.6. - Every flag logs `repair_candidate` + reason via structlog → the S1b evidence base. **Explicit non-goals in S1a:** no re-planning, no LLM, no change to degrade-and-continue semantics, no change to never-throw seams. INV-6 untouched. --- ## 4. W2 — S2 Visualization tool (`render_chart`) ### 4.1 Tool (Sofia's slice — tool layer is tool-team-owned) - NEW `src/tools/analytics/visualization.py` — `render_chart(df, chart_type, x, y, series=None, title=None)`. Pure, deterministic, **no LLM, no new dependency**: builds a Plotly-conformant dict by hand (Plotly JSON is a documented schema; the FE renders with plotly.js — Python never imports plotly). v1 chart types: `bar`, `line`, `pie`, `scatter`. `pie` maps `x`→labels, `y`→values. House style preset (colors, axis config) is a module constant — style is not a HAL decision. - `src/tools/contracts.py` — `ToolOutput.kind` Literal gains `"chart"` (one line; tool-team-owned file). - `src/tools/registry.py` — `ToolSpec(name="render_chart", category="analytics.visualization", input_schema={"required": ["data", "chart_type", "x", "y"], "properties": {…, "series", "title"}}, output_kind="chart", description=visualization.DESCRIPTION)` in `ACTIVE_ANALYTICS_TOOLS`. - `src/tools/invoker.py` — `_DISPATCH["render_chart"] = (visualization.render_chart, "chart")`. Never-throw comes free; a failed chart degrades the task, the answer still assembles. ### 4.2 Chart-spec envelope (v1, FE contract) ```json { "schema": "dataeyond.chart.v1", "chart_type": "bar", "title": "Revenue by region", "plotly": { "data": [{ "type": "bar", "x": ["A", "B"], "y": [1, 2], "name": "revenue" }], "layout": { "title": {"text": "Revenue by region"}, "xaxis": {"title": {"text": "region"}}, "yaxis": {"title": {"text": "revenue"}} } } } ``` FE renders with `Plotly.newPlot(el, spec.plotly.data, spec.plotly.layout)`. The envelope (not raw plotly) is what `ToolOutput.value` carries and what the store persists. ### 4.3 Planner slice (Rifqi) - `src/agents/planner/examples.py` — **Example J**: "plot revenue by region" → R2 retrieve → `render_chart` with `data="${t1}"`. - `src/config/prompts/planner.md` — the recipe table (§1) + rule: chart **only when the user explicitly asks** (plot/visualize/chart/grafik/buatkan diagram); never speculative; `render_chart` is always a tail, fed by a table-kind upstream. - `src/agents/planner/validator.py` — path-shape check: `render_chart.data` must reference a task whose tool yields `table` output. - `src/config/prompts/intent_router.md` — only if EN/ID viz phrasings misroute in the smoke; keep few-shots ↔ eval dataset mirrored if touched. - `src/agents/slow_path/prompt.py` — assembler renderer gets a `kind == "chart"` branch: one-line summary (type, title, point count). Without this the full spec (x/y arrays) floods the assembler prompt ([prompt.py:52](src/agents/slow_path/prompt.py:52) currently dumps `output.value` verbatim for non-table kinds). - `src/agents/report/generator.py` — no change needed: `_collect_evidence` copies table-kind only; charts cannot corrupt the markdown report. Report embedding of charts stays **deferred** (standing decision). ### 4.4 Chart store (Rifqi) — mirrors `message_traceability` - NEW `src/charts/store.py` — `ChartStore` Protocol + `PostgresChartStore` (never-throw `save`, `list` read) + `NullChartStore`. ORM row in `src/db/postgres/models.py`. - Write site: `chat_handler._run_slow_path`, after the record persist — scan `record.results_snapshot` outputs for `kind == "chart"`, save each stamped with `pad.message_id`, `analysis_id`, `user_id`, `record_id`. Never-throw; a chart-persist failure must not break the answer. - **Zero DDL from Python** (§2.2). DDL below is run manually against dedorch now and handed to Harry for the migration: ```sql CREATE TABLE IF NOT EXISTS message_charts ( id uuid PRIMARY KEY DEFAULT gen_random_uuid(), message_id text NOT NULL, analysis_id uuid NOT NULL REFERENCES analyses(id), user_id text NOT NULL, record_id text, chart_type text NOT NULL, title text, spec jsonb NOT NULL, -- the dataeyond.chart.v1 envelope (§4.2) created_at timestamptz NOT NULL DEFAULT now() ); CREATE INDEX IF NOT EXISTS idx_message_charts_lookup ON message_charts (analysis_id, message_id); ``` ### 4.5 API surface + contract delta (Rifqi → Harry) - NEW `GET /api/v1/charts?analysis_id=&message_id=` (`src/api/v1/charts.py`, mounted in `main.py`) → `{"count": n, "charts": [{chart_id, chart_type, title, spec, created_at}]}`; empty list is a valid 200 (FE: no charts this turn). Same fetch-on-`done` pattern as traceability; the row is written before `done`, no polling race. - `done` event: **proposed additive field** `chart_count` (or `chart_ids`) so the FE can skip the GET on chartless turns. Harry's call; fallback = FE fetches unconditionally. SSE order and all existing fields unchanged; `sources` stays `[]`. - `API_CONTRACT_BE_PYTHON.md` updated in the same change (new §charts + `done` note). FE ask (via Rifqi): render `spec.plotly` with plotly.js under the assistant message; "chart iteration" v1 = a follow-up chat turn (the planner re-emits `render_chart` with patched args — no new endpoint). ### 4.6 Checkpoint spec checks (CK6, with W1) Empty series · `len(x) != len(y)` · bar/pie with > 30 / > 8 categories → flag (S1a: honest note; S1b later: re-pick args or fall back to table) · non-numeric y for bar/line/scatter. ### 4.7 PII / tracing notes Chart specs contain the same result rows the user already receives in tables — no new exposure class. Langfuse: tool spans stay metadata-only; the traceability scratchpad records the `render_chart` call like any tool (verify the `TraceabilityToolInvoker` summary handles `kind="chart"` sanely — row_count is absent; use point count). --- ## 5. W3 — Activate deferred analytics tools (⏸️ DEFERRED 2026-07-13 — kept for the later cycle) Move `analyze_comparison`, `analyze_contribution`, `analyze_profile`, `analyze_segment` from `DEFERRED_ANALYTICS_TOOLS` to `ACTIVE_ANALYTICS_TOOLS` (`src/tools/registry.py` — compute fns and invoker mappings already exist). Per the registry's own note: each re-activated tool needs planner few-shot coverage (`examples.py`) — that is the real work. Recipes gain R7 comparison / R8 contribution / R9 segment rows. Prompt change → full §7B eval gauntlet. Closes the breadth gap vs Julius at near-zero engineering cost. --- ## 6. W4 — S1b targeted repair (GATED — do not build yet) Trigger: checkpoint returns repair candidates AND the team has approved the INV-6 relaxation AND S1a telemetry shows a meaningful hit-rate (proposal: revisit after 2 weeks of S1a logs). Design (for the future PR): `PlannerService.repair(task_list, run_state, reasons)` → patched subgraph for the failed task ids only; runner re-executes only those tasks + skipped dependents; **max 1 pass per turn**; repair prompt gets the failure evidence (error strings, row counts). Budget: +1 LLM call, failure paths only. If rejected: S1a's honest-degrade stands alone and this section moves to an ADR graveyard note — the checkpoint is still worth it. --- ## 7. W5 — S3 `model_*` family (future sketch, no build) Same spine, fourth family: fixed, audited compute functions (e.g. `model_regression`, `model_forecast`, `model_cluster` — sklearn/statsmodels behind ToolSpecs; **dependency additions need the §6.4 ask**). Recipe: `retrieve → prepare → model_* → evaluate` — CRISP-DM's modeling/evaluation stages become real. Outputs are `stats`/`table`/`series` + optional `render_chart` tails. Reports gain model sections then, not before. Explicitly NOT codegen — if open-ended modeling demand materializes, a sandboxed codegen tool is a separate §6.3 guardrail conversation. --- ## 8. Verification gates (per workstream — house §7) - Every W: `ruff check src/` clean · full local pytest with **exact counts** vs previous · `PYTHONPATH=. ./.venv/Scripts/python.exe -c "import main"` exits 0 · no never-throw seam pierced. - W1: unit tests for every CK rule (fixture RunStates); in-process slow-path run showing the assessment block in the assembler input; no behavior change on a clean run. - W2 prompts: `eval.chat_sim` smoke ≥ last committed score (name both files); `eval.intent` only if router prompt touched; new timestamped results committed, never overwriting. - W2 e2e: live in-process `structured_flow` turn with an explicit viz ask → chart row in `message_charts` → `GET /charts` 200 with a valid v1 envelope; a chartless turn → 200 empty list; a failing `render_chart` → answer still streams, no chart row. - W3: eval smokes for each new few-shot; EN+ID phrasing per eval conventions (§7F). - Contract changes: `API_CONTRACT_BE_PYTHON.md` same-change; REPO_STATUS/DEV_PLAN deltas dated (§7E). ## 9. Decisions log & open questions Locked (approved 2026-07-13): Plotly JSON not PNG (re-affirms DEV_PLAN #26) · no per-phase LLM agents · S1a before/with S2, S1b evidence-gated · ~~charts chat-only for now (report embedding deferred)~~ RESOLVED 2026-07-14: the FE fenced-block hook exists (lead verified), so report embedding LANDED exactly as sketched — ` ```plotly ` fences in `rendered_markdown`'s EDA section, emitted by the generator from `results_snapshot` chart outputs, INV-4 copy-verbatim; `has_successful_analysis` now also counts a successful `render_chart`, so a chart-only session satisfies the report floor · `render_chart` only on explicit user ask · ~~ToolOutput.kind + render_chart ownership split~~ RESOLVED: Sofia signed off, implemented in-repo by the dev session · **W3 deferred to a later cycle**. Open → owner: `done.chart_count` additive field yes/no → Harry · FE chart placement (under message) + plotly.js availability + report fenced-block hook → mentor/FE (Rifqi to remind) · INV-6 relaxation for W4 → team, after S1a telemetry · W3 re-open timing → Rifqi. ## 10. Doc-sync checklist on landing (per workstream) REPO_STATUS §6/§9 (tool count, new checkpoint stage, charts surface) · DEV_PLAN new task rows under a dated sprint section · API_CONTRACT §charts + `done` delta · this file: status flips + dated delta banner.