Rifqi Hafizuddin
[NOTICKET] revert: unwire F-2 service-secret gate β unarmable with a browser-only caller
92a2913 | # Data Eyond β Current Development Plan (post 2026-06-24 β 2026-06-30 checkpoints) | |
| **Purpose:** context file for Claude Code sessions working on the current sprint. | |
| **Branch:** `pr/5` Β· **Snapshot:** 2026-06-30. | |
| **Companion:** [REPO_STATUS.md](REPO_STATUS.md) describes the repo's *current built state*; this file | |
| describes the *in-flight plan* that changes it. The **active sprint is pr/5** ([Β§0](#0-current-sprint--pr5-observability--endpoint-restructure)); sections Β§1βΒ§6 are the prior | |
| 2026-06-24/25 pivot (now largely β ), kept for context. | |
| --- | |
| ## 0. Current sprint β pr/5: Observability + Endpoint Restructure | |
| From the **2026-06-30 checkpoint**. Direction: **Python β generation/AI-only**; Go owns the analysis | |
| lifecycle + data plane. Endpoint contract sent to Harry on 2026-06-30: | |
| [API_ENDPOINTS_RESTRUCTURE.md](API_ENDPOINTS_RESTRUCTURE.md) (chatβv2, tools regroup, observability β | |
| observability marked tentative). REPO_STATUS carries a matching `pr/5` direction banner. | |
| Mentor's task order: **unwire β regroup endpoints β add tools (retrieve-data + observability)**; share | |
| the endpoint contract *before* coding the tools. Status legend: β¬ not started Β· π in progress Β· β done Β· | |
| β blocked Β· π verify Β· βΈοΈ deferred. | |
| | Phase | Task | Owner | Status | Notes | | |
| |---|---|---|---|---| | |
| | **P0 β contract** | Draft + send endpoint contract to Harry (chat v2 Β· tools group Β· observability) | Rifqi + Sofhia | β | `API_ENDPOINTS_RESTRUCTURE.md` sent 2026-06-30 (before-noon deadline met). Observability section flagged tentative. | | |
| | **1 β unwire** | Unwire `users`(login)/`document`/`room`/`db_client`/`data_catalog`/`analysis` from `main` + Swagger | Sofhia | β | **KM-686**, commit `0b2d678`. Commented, not deleted; `chat`/`report`/`tools` kept mounted. Resolves the analysis-CRUD scope Q β whole `analysis` router unwired (Go owns it). | | |
| | **2 β v2 + regroup** | Create `src/api/v2/` and move the chat pilot there | Rifqi | β | New `src/api/v2/__init__.py` + `src/api/v2/chat.py` (`POST /api/v2/chat/stream`), mounted in `main.py`. Only chat in v2; v1 `/chat/stream` kept mounted until FE moves over. Routes import-verified. | | |
| | **2 β v2 + regroup** | Chat: `room_id` β **`analysis_id`** (request field + handler + history) | Rifqi | β | v2 `ChatRequest{user_id, analysis_id, message}`; reuses warm `ChatHandler` + v1 cache/history helpers; `done` returns `{message_id}` (always minted Python-side, server-authoritative β open-Q #1 resolved in pr/6). Persistence kept transitionally β still ties to #25 (`analyses_messages`); ruff-clean. | | |
| | **2 β v2 + regroup** | Move report under tools β `/api/v1/tools/report` (+ version routes) | Rifqi | β | report router re-prefixed `/api/v1` β `/api/v1/tools` (all 3 routes move together), tag β `Tools`; old `/api/v1/report` gone. Same functionality, new home. Import-verified. | | |
| | **2 β v2 + regroup** | Move help under tools β `POST /api/v1/tools/help` (dedicated endpoint) | Sofhia | β | New `src/api/v1/help.py` (SSE: `sources:[]`β`chunk`β`done{message_id}`) + additive `ChatHandler.stream_help()` (reuses HelpAgent+state+readiness, no router). Generative-only (no persist). **Router `help` intent KEPT** β both paths live by design. message_id always minted Python-side, server-authoritative (open-Q #1 resolved, pr/6). Import-verified. | | |
| | **2 β v2 + regroup** | Tools list β `/api/v1/tools/list` | Sofhia | β | Renamed route `GET /api/v1/tools` β `GET /api/v1/tools/list` ([tools.py:133](src/api/v1/tools.py:133)). | | |
| | **2 β v2 + regroup** | FE: slash menu = `/help` only; report = right-side button | Mentor (FE) | β¬ | Coordination note, not Python work. | | |
| | **3 β tools + obs** | Finish `help` so it actually **calls** (not just lists) + test | Sofhia | β¬ | Mentor: help currently only lists tools. Core #2 after chat. | | |
| | **3 β tools + obs** | Traceability **scratchpad** accumulating in the chat agent | Rifqi + Sofhia | β | **KM-691.** `TraceabilityScratchpad` + `TraceabilityToolInvoker` (`src/traceability/`) capture planning / tool I/O / sources during the run; flushed one row before every `done` (all 8 sites; error turns = no row). Renamed observabilityβtraceability (vs. Langfuse). | | |
| | **3 β tools + obs** | Audit `report_inputs` β covers planning + tool I/O + source? add cols / new store | Rifqi | β | **KM-691.** Chose a dedicated store: `message_traceability` = 1 JSONB row per message (Python-owned, like `report_inputs`; DDL run manually against dedorch, handed to Harry). Langfuse kept for engineering. | | |
| | **3 β tools + obs** | Build `GET /api/v1/traceability` (one merged response) | Rifqi | β | **KM-691.** `src/api/v1/traceability.py` β store.get β payload/404. Intent-based source rules (greeting/help/refusals = none; retrieve = required); full planning only on slow path. Contract Β§7 updated. | | |
| | **3 β tools + obs** | Keep stream **text-only**; traceability is a separate parallel call | Rifqi | β | **KM-691.** No trace data in the SSE stream; the FE fetches `/traceability` on `done`. | | |
| | **3 β tools + obs** | Resolve `message_id` correlation (stream β traceability) with Harry | Rifqi β Harry | β | **RESOLVED (pr/6):** Python is the **sole minter** β `message_id` dropped from the `/api/v2/chat/stream` + `/api/v1/tools/help` request bodies; always minted server-side (server-authoritative, FE-security) and returned on `done`. Any caller-sent `message_id` is ignored. Contract open-Q #1 closed. **Updated 2026-07-09 (pr/13):** id format changed from `msg_<hex>` to a canonical UUID string (`str(uuid.uuid4())`, both mint sites) to mirror Go's `analyses_messages.id` shape β the value is still independently Python-minted (not the real row id), only format-compatible for a future swap. | | |
| | **4 β biz questions** | Get Go folder; confirm `business_questions` in create-analysis (max 5); sync Python | Harry/Mentor β Rifqi | β¬ | Go currently missing the field ("lagi difixing"). Python already models objective + business_questions. | | |
| | **deferred** | Report formats: PPT (preferred) / PDF / infographic on download | β | βΈοΈ | MD is fine for the FE preview stage now. | | |
| | **deferred** | Charts (PlotlyβJSON) + images tables | β | βΈοΈ | Carried from Β§4 #26/#27. | | |
| **Next up:** Phase 2 Python work is **done** (chatβv2 `analysis_id`; `help`/`report`/`list` regrouped | |
| under `/api/v1/tools/`). The `message_id` correlation contract is now settled in **pr/6** (Python sole | |
| minter, stream-only). The **Phase 3 traceability build** β scratchpad + `GET /api/v1/traceability` | |
| (contract Β§7) β is **done (KM-691)**. Next is **Phase 4** (business questions, Go-blocked). | |
| --- | |
| ## 0.5. pr/13 sprint β agent-quality fixes (2026-07-08 live-test review) | |
| Findings from the scoped live sessions (mining analysis, 2026-07-07/08 traces): the planner | |
| force-mapped absent measures (`pa` aliased as "revenue"), top-N ranked raw rows (duplicate models), | |
| `analyze_trend` collapsed integer months into a single 1970-01 bucket, an invalid grouped IR reached | |
| Postgres, failed retrievals wrote all-null traceability sources, and numeric catalog samples arrive | |
| base64-mangled from Go. Fix tasks (same status legend as Β§0): | |
| | # | Task | Owner | Status | Note | | |
| |---|---|---|---|---| | |
| | Q1 | IR validator: reject bare selects under `group_by` (planner retry self-corrects) | Rifqi | β | `query/ir/validator.py` | | |
| | Q2 | Planner **infeasible** path: `TaskList.infeasible_reason` + deterministic EN/ID data-gap reply | Rifqi | β | schemas/validator/coordinator/refusals + planner.md "When the catalog cannot answer"; refusal wording β Rifqi to review | | |
| | Q3 | `analyze_trend`: integer year/month handling (epoch-parse bug) | Rifqi | β | `temporal.py` + 5 local tests | | |
| | Q4 | Planner few-shots: top-N (Example G) + infeasible (Example H) + entity-vs-row ranking rule | Rifqi | β | live-tested 2026-07-08: backlog top-3 correct via single-IR group+sum; "best PA performance" correct in-process (avg-per-model, assumption recorded). Stale-server trace was a false alarm | | |
| | Q5 | Catalog numeric `sample_values` base64-decode stopgap (`catalog/sample_decode.py`) | Rifqi | β | self-disabling; **primary fix = Go marshaling β DDL-free handoff to Harry** | | |
| | Q6 | Traceability null-source suppression + `check_data` `-1` row-count hiding | Rifqi | β | `scratchpad.py` / `data_access.py` | | |
| | Q7 | `analyze_merge` two-table combine tool (unblocks "worst A + biggest B" questions) | tool owner | β | tool shipped by Sofia (8abf635, KM-703); planner slice done 2026-07-09: `_validate_data_source` guards `data_right`, two-retrieveβmerge few-shot (Example I), planner.md "Two measures per entity" bullet | | |
| | Q8 | Report v2: business-question answer section, unresolved/excluded sections, evidence tables from `results_snapshot`, caveat dedupe, single language | Rifqi/Sofhia | β | done 2026-07-09: still exactly ONE LLM call (extended to also draft `bq_answers`, index-based record refs, deterministic fallback = v1 behavior); evidence tables from table-kind outputs (β€3/record, β€10 rows, β€8 cols, `check_*` skipped); reply language via `detect_reply_language` on objective+BQs; verified in-process against live analysis 935a091e | | |
| | Q9 | Record-curation endpoint (`GET β¦/records` + `exclude_record_ids`) + readiness GET for the FE delta guard | Rifqi β FE | β | done 2026-07-09: `GET /tools/report/{analysis_id}/records` + `/readiness` (registered before `/{version}` β int-coercion route-order trap), `exclude_record_ids` on POST; contract updated same change; FE wiring pending (Rifqi β FE) | | |
| | Q10 | Traceability `data_used` layer β resolve IR ids β real names for the FE (users couldn't map `c_β¦`/`t_β¦` ids back to their data; aggregate aliases like `total_revenue` looked like real columns) | Rifqi β FE | β | done 2026-07-13 (pr/15): new `src/traceability/resolve.py` builds `data_used[]` (real source/table/column names; joins; plain-language filters; `columns_read` vs `output_columns` with `computed`+`formula`), `tool_calls[].summary`, `sources[]` gains `source_name`+all tables; **ids kept but machine-only (FE must not render)**; deterministic no-LLM, never-throw; catalog threaded to the scratchpad at the composition root. Contract + `TRACEABILITY_FE_HANDOFF.md` updated; FE wiring pending (Rifqi β FE). Additive/non-breaking | | |
| | Q11 | IR wart: `OrderByClause.column_id` may hold a SELECT **alias** (a computed output), not a catalog column_id | Rifqi | π | surfaced by Q10 β the resolver tolerates it (`kind: "computed"` fallback), but the IR field name is misleading. Consider an explicit `by_alias` field or renaming. Low priority; no functional bug | | |
| ## 0.6. pr/16 sprint β Spine v2: W2 charts + W1 checkpoint (SPINE_V2_PLAN, approved 2026-07-13) | |
| Scope approved 2026-07-13 (Rifqi + Sofia): build W2 (`render_chart` + chart store + `GET /charts` | |
| + planner viz slice) and W1 (S1a quality checkpoint). **W3 (activate deferred `analyze_*`) deferred | |
| at approval β do not start until Rifqi re-opens.** W4 (S1b repair) stays gated on INV-6 sign-off + | |
| S1a telemetry. Design + handoff source: `SPINE_V2_PLAN.md`. | |
| | # | Task | Owner | Status | Note | | |
| |---|---|---|---|---| | |
| | V1 | `render_chart` tool slice: `visualization.py` (Plotly-JSON `dataeyond.chart.v1` envelope, no plotly dep) + `ToolOutput.kind` `"chart"` + registry + invoker | Rifqi (Sofia signed off on the tool-layer edit) | β | done 2026-07-13; deterministic spec builder (bar/line/pie/scatter, fixed style preset), traceability scratchpad summarizes chart outputs compactly (point_count, not the raw arrays) | | |
| | V2 | Chart store + API: `MessageChartRow` (`message_charts`) + `src/charts/store.py` (never-throw save) + write site in `_run_slow_path` + `GET /api/v1/charts` + contract Β§charts | Rifqi | β | done 2026-07-13; empty list = valid 200; `done` event unchanged (no `chart_count` β open, Harry); FE fetches unconditionally on `done` | | |
| | V3 | Planner viz slice: recipe table + "Charts only on explicit ask" rule (planner.md), Example J (viz tail) + Example K (viz-infeasible), validator Check 10 (`render_chart.data` must be table-kind), assembler chart one-liner guard | Rifqi | β | done 2026-07-13; behavioral matrix verified in-process (real LLM): explicit ask EN/ID β chart tail; plain question β no chart; absent dimension β infeasible (Example K + prompt guard added after the first smoke force-mapped `status` AS "region") | | |
| | V4 | W1 S1a quality checkpoint: `slow_path/checkpoint.py` (CK1βCK6) + `RunAssessment` schemas + coordinator call site + "Execution assessment" block in the assembler input + `refusals.run_failure_message` | Rifqi | β | done 2026-07-13; 13 local tests (one per CK rule + never-throw + prompt + coordinator); CK1 all-failed β deterministic honest failure, **no assembler call**; clean run renders nothing (zero behavior change); every flag logs `repair_candidate` (S1b evidence) | | |
| | V5 | `message_charts` DDL: run manually against dedorch (block for the live e2e chart test), then hand the schema to Harry for the dedorch migration | Rifqi β Harry | β | Rifqi ran the DDL 2026-07-13; **live e2e ALL PASS** same day (real v2 endpoint: viz turn β chart row keyed by `done` message_id β `GET /charts` valid v1 envelope; chartless β 200 empty; injected `render_chart` failure β answer streams as a table, no row). **Remaining: send Harry the migration handoff** (schema + contract Β§charts) | | |
| | V6 | Restore the `eval.chat_sim` harness β `eval/chat_sim/*.py` is missing from disk AND git (only `__pycache__` remains); Β§7B prompt gate can't run | Rifqi | β | restored by Rifqi 2026-07-13 (accidental delete). β οΈ its hard-coded `DEFAULT_USER_ID`/`TITANIC_SOURCE_ID` are stale (that user has no catalog; the Titanic blobs are gone) β update the constants before the next full run | | |
| | V7 | Local `.env` lagged Go's Supabase-S3 data plane: `storage_provider=azure_blob` + empty `supabase_s3_*` made EVERY local tabular retrieve fail `BlobNotFound` | Rifqi | β | found during the e2e (masked as "data not available" by the honest-degrade path); Rifqi set the six values 2026-07-13. Gotcha documented in REPO_STATUS Β§13 | | |
| | V8 | Lead review of `GET /charts`: lookup by `message_id` alone + tri-state response marker (`status: success \| empty \| not_found` + `message`) instead of a bare list | Rifqi (lead ask) | β | done 2026-07-14; `not_found` vs `empty` decided against the turn's traceability row (PK lookup); always HTTP 200; contract Β§charts updated. β οΈ additive DDL for the new lookup: `CREATE INDEX IF NOT EXISTS idx_message_charts_message ON message_charts (message_id);` (run manually + include in Harry's migration). Traceability GET still takes both params β aligning it is open | | |
| | V9 | Report chart embedding: `AnalysisReport.charts` (verbatim envelopes per record) + `## EDA` section with ` ```plotly ` fences (content = the **full v1 envelope**, pretty-printed β the FE hook's verified shape); `has_successful_analysis` extended so a successful `render_chart` counts (chart-only sessions satisfy the report floor) | Rifqi | β | done 2026-07-14; first cut emitted bare `{data, layout}` β FE test showed the hook parses the full envelope, fixed same day; live reports v3 (wrong fence) β **v4 (correct)** for analysis `7be50846β¦` (3 charts embedded); suite **381 passed, 7 skipped** (+5 chart-embed tests; same 2 pre-existing failures) | | |
| Full-suite evidence for this sprint: **376 passed, 7 skipped** (+13 new checkpoint tests; the 2 | |
| failures β `test_chat_handler::test_structured_flow_runs_slow_path`, | |
| `test_reader::test_structured_read_falls_back_to_user_scope_when_no_analysis_row` β reproduce at | |
| HEAD before this diff, i.e. pre-existing). Ruff clean on all touched paths; `import main` OK. | |
| ## 1. The direction change (locked decisions from 2026-06-24) | |
| 1. **"Problem statement" is replaced by two user-entered fields: `objective` + `business_questions`.** | |
| User fills them at onboarding; **both mandatory to submit; NO agent validation.** | |
| 2. The **gate (`problem_validated`) and the `problem_statement` skill/intent are removed** (comment out, don't delete). | |
| 3. **Report is records-based** (reads persisted `AnalysisRecord`s) β **decided and pushed** (KM-674). | |
| It is formal markdown: title, date, "generated by {user}", objective, business questions, | |
| findings, insights. **NOT gated** on whether business questions were answered. | |
| 4. **`owner_id` β `user_id`** everywhere (Harry mirrors in dedorch/Go). | |
| 5. **State writes go through a request to Go**, not direct Python DB writes. | |
| 6. **FE-callable surface = 4 endpoints:** `call_agent` (chat/stream), `list_skills` (`GET /tools`), | |
| **skill: help**, **skill: report**. `problem_statement` removed; `check_data` not FE-facing from | |
| Python (Go provides it); analysis CRUD not needed from FE (comment, don't delete). | |
| 7. Deliverables for Harry: (a) API endpoint doc (MD); (b) full Python project doc (MD β PDF/Word BRD). | |
| 8. Integration tested via Swagger `/docs` on the HF Python build (simulating FE manually). Target ~Wed. | |
| ## 1.5. 2026-06-25 checkpoint deltas | |
| Confirms the 2026-06-24 direction and adds these concrete changes (folded into Β§4 as tasks 21β28): | |
| 1. **Rename `analysis_records` β `report_inputs`** (DONE #21) β names the table by purpose (the rows | |
| report generation reads); avoids clashing with Go's `analyses_messages` and with Langfuse | |
| observability. **Stays Python-owned**; finalized schema handed to Harry so his dedorch migration | |
| creates it post-`SKIP_INIT_DB` (#22, resolves #16). Write scope = **one row per slow-path analysis | |
| run** (decided β not per-agent-call telemetry; that stays Langfuse). | |
| 2. **`analyses` table (Go) β `status`, `data_bind` + `data_bind_version`, `report_collection`** (id+version). | |
| **Verified 2026-06-25: these + `user_id` are ALREADY present in dedorch `analyses`.** Plus Harry drops | |
| the duplicate/wrong singular `analysis` table. (β #3) | |
| 3. **`analyses_messages` (Go) = the analysis chat room** (user Q + agent A) β replaces the now-**deprecated** | |
| `chat_messages`/`rooms`; Python's chat read/write must migrate here before cutover. (β #25) | |
| 4. **Reports: Go owns ALL writes.** Report stays a **skill** (no router intent): FE β Go β Python; | |
| Python only returns content. Input = the records table (now `agent_observability`); edit-mode may | |
| also need the last report. (β #7/#18/#24) | |
| 5. **Markdown minimum now:** tables, **bold**, *italic*, horizontal separators β optimize that before | |
| anything fancier. (β #23) | |
| 6. **Deferred:** charts (prefer **PlotlyβJSON** in a future `chart` table over matplotlib PNGs) and | |
| images (image table keyed by analysis/message/report + originals in a bucket). (β #26/#27) | |
| 7. **Near-term:** the remove-`problem_statement` work isn't on HF yet β **PR + deploy + test in the | |
| playground** (#13). Harry stabilizes Go ~Fri; FE manual testing ~Mon. **Keep it playground-able.** | |
| 8. **UI research** (no dedicated UI person): new-analysis form (title/objective/business_questions), | |
| knowledge menu (user-level vs analysis-level binding), report artifacts panel + version selector; | |
| interview + old analysis UI removed. (β #28) | |
| ## 2. What is already done (KM-674, pushed on `pr/4`) | |
| Report layer adapted to the new goal shape: | |
| - `report/schemas.py::ProblemStatement` β `objective: str` + `business_questions: list[str]` | |
| (old `target_value`/`scope`/`metric_direction`/`target_metric` dropped). Class name kept for now | |
| (rename to `ReportGoal` once the upstream AnalysisState rename lands). | |
| - `report/generator.py` renders **Objective** + numbered **Business Questions** + a | |
| **"generated by {user}"** line. | |
| - `api/v1/report.py::_problem_statement_from` is **tolerant**: prefers new `objective` / | |
| `business_questions` from state, falls back to legacy `problem_statement` β works before AND after | |
| Harry's migration. | |
| - `config/prompts/report_summary.md` updated to objective + business questions. | |
| - Report stays **records-based**; the floor gate (`problem_validated`) was deliberately left for task #2. | |
| **This tolerant-migration pattern (getattr fallback) is the model for tasks #2 and #4.** | |
| ## 3. Assessment β gaps & contradictions to resolve before building | |
| These came out of reviewing the plan against the actual code. They are folded into the task table (Β§4) as tasks 15β19. | |
| - **G1 (β task 15). Records-based reports need the slow path ON.** `AnalysisRecord`s persist only in | |
| `chat_handler._run_slow_path`, which runs only when `ENABLE_SLOW_PATH=true`. Default is off β no | |
| records β `POST /report` 409s. The Swagger demo can't show a non-empty report unless slow path is | |
| flipped on and a `structured_flow` question is run first. `BusinessContext` is still a stub but the | |
| slow path runs fine on it. | |
| - **G2 (β task 16). `analysis_records` ownership is now required and collides with `SKIP_INIT_DB`.** | |
| It's created today by Python `create_all` (`db/postgres/init_db.py`), is in no dedorch/Go migration, | |
| and after the dedorch cutover (`SKIP_INIT_DB=true`) Python stops running `create_all` β the table | |
| won't exist β reports break. Decide: dedorch migration (Harry) OR a Python carve-out that creates | |
| just this one table even under `SKIP_INIT_DB`. *(Resolved 2026-06-25 β see Β§1.5.1 / #16 / #22.)* | |
| - **G3 (β task 17). `chat_history` in the report contract is vestigial.** Records-based generation | |
| reads records by `analysis_id`; it never uses chat history. Drop `chat_history` from the report | |
| skill contract, or mark it reserved/unused. | |
| - **G4 (β task 4 note). Make #2/#4 tolerant of both state shapes.** If Harry drops | |
| `problem_validated`/`owner_id` from dedorch before Python stops reading them, Python's gate + | |
| state_store break. Use the same `getattr` tolerance KM-674 used. The `owner_id`β`user_id` rename | |
| also touches `api/v1/analysis.py` (`_serialize_state`, `list_analyses`, `get_analysis`), not just | |
| the model + state_store. | |
| - **G5 (β task 18). "State writes via Go" is bigger than `report_id`.** Python still writes state in | |
| `/analysis/create` (state + room + bindings, plus the data-first gate and soon the mandatory-field | |
| check) and in `state_store.ensure` per turn. If creation moves to Go (consistent with commenting | |
| analysis CRUD), then Go owns ALL state writes + both creation gates, and Python's `ensure` must | |
| become a **read-only get** (Go must guarantee the row exists before any chat turn). | |
| - **G6 (smaller).** | |
| - Removing `problem_statement` (task 1) means neutering it in 4 places: the `Intent` literal | |
| (`agents/orchestration.py`), the router prompt (`config/prompts/intent_router.md`), the handler, | |
| and the gate's redirect *target*. Do it with task 2. | |
| - "generated by {user}" currently prints the raw `user_id`; a formal report wants a name β source | |
| from `users.fullname` or have Go pass a display name (task 19). | |
| - The meeting's outline (background / EDA / insights) isn't fully in the renderer; map those | |
| sections onto the record fields deliberately (task 5 follow-up). | |
| - The full project doc (task 11) should reuse [REPO_STATUS.md](REPO_STATUS.md), not restart. | |
| ## 4. Task table | |
| Status legend: β¬ not started Β· π in progress Β· β done Β· β blocked Β· π verify Β· βΈοΈ deferred. | |
| | # | Task | Owner | Status | Note | | |
| |---|---|---|---|---| | |
| | 1 | Comment out `problem_statement` skill **+ `Intent` literal + router prompt + gate redirect target**; remove `/problem-statement` from `list_tools` | Rifqi | β | Done 2026-06-25 (one commit w/ #2). Unwired in `orchestration.py`, `intent_router.md`, `chat_handler.py`, `tools.py`; `problem_statement.py` kept intact | | |
| | 2 | Drop `problem_validated`: gate neutered; `is_report_ready`/`report_floor` β **β₯1 completed analysis** only, no-LLM | Rifqi | β | Done 2026-06-25. `gate.py` no-op, gate call site commented in `chat_handler.py`, `report_floor` drops the goal check. Tests updated (`test_gate`/`test_chat_handler`/`test_readiness`). Suite: **284 passed, 7 skipped**; ruff clean | | |
| | 3 | dedorch `analyses` migration: drop `problem_statement`/`problem_validated`, add `objective` + `business_questions` | Harry | π | **Verified dedorch 2026-06-25:** `analyses` (plural) ALREADY has `user_id` + `status` + `data_bind` + `data_bind_version` + `report_collection` β those parts done. **Remaining:** drop `problem_statement`/`problem_validated` + add `objective`/`business_questions`. Singular `analysis` = deprecated duplicate to drop | | |
| | 4 | Update Python `analyses` model + `state_store` + `analysis.py` to match dedorch; `owner_id`β`user_id` | Rifqi/Sofhia | β | Done 2026-06-26. `owner_id`β`user_id` + added `status`/`data_bind`/`data_bind_version`/`report_collection` (DB-only, not in the `AnalysisState` pydantic) across `models.py`/`gate.py`/`state_store.py`/`analysis.py` + 3 local tests; also `report_inputs` `id`/`analysis_id` β `uuid`. Kept `problem_statement`/`problem_validated`; `objective`/`business_questions` wait on Harry's #3. Suite **284 passed** | | |
| | 5 | Report generator β `objective`+`business_questions`, "generated by {user}", formal outline | Sofhia | β | Goal-shape (KM-674) + author name (#19) + outline (KM-680): Objective β Business Questions β Executive Summary β Key Findings β EDA β Notes & Limitations β How This Was Analyzed | | |
| | 6 | Report skill input contract: `analysis_id` + `user_id` (no `chat_history`) | Sofhia/Rifqi | β | No-op: `POST /report` already takes only analysis_id + user_id (records-based). Documented in API_ENDPOINTS.md Β§5. *(Edit-mode input revisited in #24.)* | | |
| | 7 | `report_id` state update via request to Go, not direct DB | Sofhia + Harry | β¬ | Needs Go endpoint. **Checkpoint:** Go owns ALL `reports` writes; Python stops any direct insert/update and only returns content; report stays a **skill** (no intent). See #18 | | |
| | 8 | Expose/confirm 4 FE endpoints; comment `check_data` + analysis CRUD | Sofhia | β | KM-678: `list_tools` trimmed to `/help` + `/report` (analytics/check/retrieve commented in the **menu**). `help` confirmed as a `call_agent` intent β no own endpoint. Analysis CRUD endpoint left **registered**: "comment the rest" was about the FE slash menu, not killing HTTP routes Go needs | | |
| | 9 | Verify `analysis_id` in `call_agent` contract | Sofhia | β | Verified: no separate field β carried as `room_id` (`analysis_id == room_id`), per REPO_STATUS Β§4/Β§11. Action for Go: send the id as `room_id` | | |
| | 10 | API endpoint doc (MD), 4 endpoints, for Go integration | Rifqi + Sofhia | β | Done 2026-06-25 β `API_ENDPOINTS.md` (repo root). 4 FE surfaces with request/response **examples** (chat SSE transcript, report 201/409 JSON, version list), schemas, Β§9 full 32-route inventory + task-8 reading | | |
| | 11 | Full Python project doc (MD β PDF/Word BRD) | Rifqi | β | Done 2026-06-26 β `PROJECT_BRD.md` (repo root): purpose/context, FR-1..9 capabilities, lifecycle, architecture, data model, API (β API_ENDPOINTS), NFRs, integrations, open items. Reuses REPO_STATUS/API_ENDPOINTS; convert to PDF/Word for distribution | | |
| | 12 | Reconcile/open the `list_tools` PR cleanly (stacked commits) | Rifqi | β | N/A β we develop directly on the single active branch `pr/4` (KM-652 + KM-678 already stacked there); no separate PR to reconcile | | |
| | 13 | Deploy HF Python build (remove-`problem_statement` work) β test 4 endpoints via Swagger / playground | Sofhia + Harry | π | **Unblocked (#15 β ).** Remove-PS work is on `pr/4` but **not on HF `main` yet** β PR + deploy, then manual test. Harry stabilizes Go ~Fri; FE testing ~Mon | | |
| | 14 | `analysis_records` home | Rifqi + Sofhia + lead | β | **Resolved 2026-06-25:** stays Python-owned, **renamed** (β #21); schema handed to Harry so the dedorch migration creates it post-cutover (β #22). Not moved to Go | | |
| | 15 | Flip `ENABLE_SLOW_PATH=true` + verify an `AnalysisRecord` persists from a `structured_flow` question | Rifqi | β | Verified locally 2026-06-25 (in-process). structured_flow on Titanic.csv β 3-task plan `check_dataβretrieve_dataβanalyze_aggregate` (all success) β AnalysisRecord persisted (substantive) β `report_floor` pass β report generates (201). HF env-flip + Swagger run folds into #13 | | |
| | 16 | Decide `analysis_records` creation under `SKIP_INIT_DB` | Rifqi + Harry | β | **Resolved 2026-06-25:** Python defines it; **Harry's dedorch migration creates it** on env-move (Python still creates locally meanwhile) β exists post-cutover. Execution = #22 | | |
| | 17 | Reconcile report contract with records-based: remove/flag `chat_history` | Sofhia/Rifqi | β | Nothing to remove β `chat_history` was never in the report contract/code (only in help.md). Confirmed via grep; API_ENDPOINTS.md Β§5 documents the clean contract | | |
| | 18 | Confirm Go owns ALL analysis-state writes + both creation gates; make Python `state_store.ensure` read-only | Rifqi + Harry | β¬ | **Confirmed by 2026-06-25 checkpoint** (Python read-only; Go owns writes + new tables). Execution pending Go endpoints | | |
| | 19 | Decide report author display-name source (`users.fullname` vs Go-passed name) | Sofhia | β | Done 2026-06-25. `AnalysisReport.user_name`; `generator` renders `user_name or user_id`; `api/v1/report.py::_resolve_user_name` reads `users.fullname` never-throw (fallback `user_id`). Decided: resolve in Python (unblocked); swap to Go-passed name later if preferred | | |
| | 20 | **Help handoff:** update `handlers/help.py` + `help.md` β drop the `problem_validated` tier + `define_problem_statement` action (the skill it points at is gone as of #1) | Sofhia | β | Done 2026-06-25. `help.py`: actions = `ask_analysis_question` (always) + `generate_report` (if ready); renders objective/business_questions (getattr-tolerant). `help.md` v1βv2: 3 tiers, no `/problem_statement`, `/generate report`β`/report`. Local test_help updated β 11 pass | | |
| | 21 | Rename `analysis_records` β **`report_inputs`** (table, ORM `ReportInputRow`, store `*ReportInputStore`) | Rifqi | β | Done 2026-06-26. `sed` rename across 9 files; Pydantic `AnalysisRecord` kept; columns stay String (pure rename β uuid+FK is the #22 Harry schema). Name `report_inputs` (purpose; avoids Langfuse/`analyses_messages` clash). Write scope = one row per slow-path run. Suite **284 passed** | | |
| | 22 | Finalize `report_inputs` schema β hand to Harry for the dedorch migration | Rifqi β Harry | β | **DDL ready** (uuid `id`/`analysis_id` + FKβ`analyses(id)`; `user_id`/`plan_id` text; `data` jsonb = serialized `AnalysisRecord`, shape documented). dedorch has empty `analysis_records` β rename. Resolves #16. ~~**Action: send Harry the DDL + `data` shape**~~ **RESOLVED 2026-07-22:** Harry's `0001_create_core_schema.sql` now carries `report_inputs` with exactly this shape β FK to `analyses(id)` included. Python's `ReportInputRow` diffed against the live Neon table: **column-for-column identical, no drift.** The two remaining un-migrated Python tables (`message_traceability`, `message_charts`) split out as β #32 | | |
| | 23 | Report markdown formatting: tables, **bold**, *italic*, horizontal separators | Sofhia | β | Done 2026-06-25. Added `---` separators between header + each section in `_render_markdown`. Tables (EDA) / bold (method labels) / italic (meta + citations) already emitted. Relaxed `report_summary.md` to allow inline `**bold**`/`*italic*` for emphasis (kept no-headings/no-bullets so it doesn't duplicate the section structure / Key Findings). Compile + ruff clean | | |
| | 24 | Clarify report input contract: records table (+ `last_report` for edit mode?) | Rifqi/Sofhia β Harry | β¬ new | Edit-mode input left open at the checkpoint | | |
| | 25 | Migrate Python chat path to Go `analyses_messages` (+ `analyses`) | Rifqi β Harry | β | Done 2026-07-02. Read path already on `analyses_messages` (commit `0066161`). This change makes Python **read-only**: removed the `save_messages` calls from `/api/v2/chat/stream` so **Go is the sole writer** β fixes the double-write both Go+Python were producing. `load_history` still reads `analyses_messages`. v1 `/chat/stream` is unwired so left untouched | | |
| | 26 | **Charts (DEFERRED):** store Plotly JSON in a future `chart` table (not matplotlib PNG) | β | π | **Landed 2026-07-13 (pr/16, Β§0.6 V1βV3):** `render_chart` + `message_charts` + `GET /api/v1/charts`, Plotly JSON as decided. π pending the dedorch DDL run (V5) + live e2e | | |
| | 27 | **Images (DEFERRED):** image table (id, analysis_id, msg/report ref, order) + originals in a bucket | β | βΈοΈ | Maintenance-heavy; parked | | |
| | 28 | **UI research** (FE): new-analysis form, knowledge menu (user vs analysis level), report artifacts + version selector | Team | β¬ new | No dedicated UI person; interview + old analysis UI removed | | |
| | 29 | **LLM env quad rename `__4o` β `__54m`** + set the four `azureai__*__54m` secrets on the HF Space | Rifqi | π | Code done 2026-07-14 (pr/17): settings quad + 9 call sites + 2 eval runners renamed; deployment is `gpt-5.4-mini`. Suite **381 passed, 2 failed (both pre-existing), 7 skipped** β same as baseline; ruff clean on touched files; `import main` 0. **Hard rename β no `__4o` fallback**, so HF fails loudly on the first LLM call until the four `__54m` secrets are set. Root cause it fixes: HF silently ran GPT-4o while local ran 5.4-mini β same question, same catalog, planner hallucinated catalog ids on HF only. **Action: set HF secrets, redeploy, re-run the failing question.** Local `.env` still carries the now-dead `__4o` keys β safe to delete | | |
| | 30 | **Neon `reports.user_id` NOT NULL** β reconcile the ORM + `ReportStore` write | Rifqi | β | Done 2026-07-22 (pr/18, commit `13f74e1`). `ReportStore.save` never wrote `user_id`; on Neon that column is `text NOT NULL` β **every `POST /api/v1/tools/report` 500'd** with `NotNullViolationError`. Fix: `models.py` adds `user_id = Column(Text)` (**nullable** on purpose β Python matches the loosest deployment shape, Β§7D getattr-tolerance) + `store.py` writes `report.user_id`, which the endpoint already required (`user_id: str = Query(...)`) and threaded via `generator.generate`. Live-tested by Rifqi β **201**. Suite 381 passed / 2 pre-existing failures / 7 skipped; `import main` 0; ruff unchanged (3 pre-existing E501 in `models.py`, same count on HEAD). **Correction 2026-07-22 (after pulling Go `dd37b38`):** the commit message and the first version of this row said the column was *not* in Go migrations `0001β0004`. **That was wrong** β `0001_create_core_schema.sql` has declared `reports.user_id TEXT NOT NULL` since the schema was written. Real root cause β #31 | | |
| | 31 | **Go migration set is not convergent** β fresh vs migrated dedorch DBs get different NOT NULL constraints | Rifqi β Harry | β¬ new | Root cause of #30, verified in the Go source 2026-07-22. `0001_create_core_schema.sql` creates `reports.user_id TEXT NOT NULL` and `analyses_messages.user_id TEXT NOT NULL`; `0002_cleanup_legacy_schema.sql` (L95, L92) and `0004_replace_chat_with_analysis_scope.sql` (L61, L57) retrofit the same two columns onto pre-existing DBs as **nullable** `ALTER TABLE β¦ ADD COLUMN IF NOT EXISTS user_id TEXT`. Because `CREATE TABLE IF NOT EXISTS` no-ops on an existing table, **the same migration set yields two different schemas** β the old dedorch DB got nullable (hiding Python's missing write for months), Neon got NOT NULL. **Ask Harry:** make the retrofit converge (backfill + `ALTER COLUMN β¦ SET NOT NULL`, or add the constraint in a new migration) so every instance matches `0001`. Until then, `information_schema` on the target instance β not the migration files β is the only reliable schema source (REPO_STATUS Β§13). **Action: Rifqi raises with Harry** | | |
| | 32 | **`message_traceability` + `message_charts` are in no Go migration** β hand the DDL to Harry | Rifqi β Harry | β¬ new | Found by the 2026-07-22 drift scan. Both tables exist on Neon **only because they were created by hand** (2026-07-06 / 2026-07-13); neither appears in `0001β0006`. Any newly provisioned dedorch instance will be missing them β traceability flush and chart persist fail (both never-throw, so they degrade **silently** β no 500 like #30 to make it visible). DDL for `message_charts` is in `SPINE_V2_PLAN.md` Β§4.4. This is the surviving half of #22. **Action: Rifqi sends Harry both DDL blocks** | | |
| ## 0.7. pr/19 β code review remediation (2026-07-23) | |
| From the end-to-end review in `CODE_REVIEW_2026-07-23.md` (findings are cited as **F-n** there) | |
| plus the live report bug on analysis `966224d4β¦`. Same status legend as Β§0. | |
| | # | Task | Owner | Status | Note | | |
| |---|---|---|---|---| | |
| | 33 | **Report body vs floor split** β `has_reportable_result` for the body, `has_successful_analysis` stays the floor | Rifqi | β | Shipped 2026-07-23. Root cause: planner R2/R2b make `analyze_*` optional, so a correct analyze-free run was classed non-substantive, dropped from the body, and its business question rendered **"Unanswered"**. Live-verified by Rifqi on `966224d4β¦` | | |
| | 34 | **Report floor extension** β a successful `retrieve_data` **that returned rows** clears the floor | Rifqi | β | Same root cause; fixes a hard **409** for a session where every question is R2/R2b-shaped. Guardrail-adjacent, authorised 2026-07-23. NOT the "Floor Fixer" failure mode: the floor still asks "did we produce a real result" β empty retrievals, `check_*`-only and fully-failed runs all still fail it | | |
| | 35 | **`GET β¦/records` `substantive` flag** repointed to the body predicate + contract Β§records updated | Rifqi | β | The curation list was contradicting the artifact it curates. Behavioral, non-breaking: no field added/removed/retyped | | |
| | 36 | **CK5b** β quality checkpoint covers analyze-free plans | Rifqi | β | CK5 only inspected `analyze_*` tasks, so an all-null aggregate column on an R2/R2b plan reached the answer unflagged. `check_*` excluded (uncounted tables legitimately carry nulls) | | |
| | 37 | **F-2 service-secret gate** β `X-Dataeyond-Service-Secret`, router-level dependency | Rifqi | βΈοΈ | Shipped inert 2026-07-23. **UNWIRED 2026-07-27 (lead decision).** Verified in `E2E-Frontend-Data-Eyond/src/services/agenticApi.ts` that the **browser SPA is the only caller** of Python and we don't own it β it sends only `Content-Type`, and cannot be changed by us to send the header. So the gate could never be armed without a 401 outage: a wired-but-unarmable gate is a footgun (any operator setting `dataeyond__service__secret` breaks prod). Guard dependency removed from all six router mounts in `main.py`; `service_auth.py` kept in-tree but parked (comment-out-don't-delete); the `dataeyond_service_secret` setting commented. A prepared FE proxy (Node `server.js` holding the secret and injecting the header server-side) is the way to arm it later without exposing the secret to the browser β but that needs FE-repo access we don't have. **Net: the live surface is unauthenticated by design; real auth = #43. F-1 tenant predicates (#38) stay** as defensive-in-depth. | | |
| | 38 | **F-1 tenant scoping** β `user_id` predicate on the six analysis-keyed reads | Rifqi | β | `CatalogStore.get_by_analysis` filtered on `analysis_id` alone where Go filters on both; the catalog payload carries the owner's `user_id`, so `DbExecutor`'s ownership check compared the victim's id against itself and passed β cross-tenant **query execution against a customer DB**. Defence-in-depth only until #37 is armed (`user_id` is caller-supplied, and `GET /traceability` leaks it) | | |
| | 39 | **Stale tests resolved** β the two long-standing suite failures | Rifqi | β | Both encoded the pre-2026-07-13 user-scope fallback that `reader.py` deliberately removed. Not product bugs. Suite is now **394 passed / 0 failed / 7 skipped** β fully green for the first time | | |
| | 40 | ~~**F-3** β scope `GET /charts` + `GET /traceability` by `user_id`~~ | Rifqi | β | **RESOLVED 2026-07-23 β DECLINED by Rifqi. Do not re-open without his sign-off.** Both endpoints keep their existing lookup keys: `/traceability` by `(analysis_id, message_id)`, `/charts` by `message_id` alone (the 2026-07-13 lead decision). F-3 proposed adding a `user_id` parameter to both; an optional-param version was implemented on 2026-07-23 and then **reverted in full** (endpoints, charts-store predicates, and the contract notes) once the decision was restated β no FE change is required and none should be requested. **Accepted consequence:** both endpoints remain unauthenticated capability URLs over real customer data (`charts[].spec.plotly.data` is actual table values; the traceability payload carries 5-row previews, the executed SQL, and the owner's `user_id`). **The service-secret gate (#37) is therefore the only control protecting them** β which raises #37 from important to load-bearing. The stores' optional `user_id` parameters are kept as dead capability for a future Go-forwarded identity (#43); `PostgresChartStore` was returned to its message_id-only form. | | |
| | 41 | **F-12 / F-13** β bound the planner catalog render and the tabular blob read | Rifqi | β | Shipped 2026-07-23. **F-13:** neither storage backend could report an object size, so `object_size()` was added to both (S3 `head_object`β`ContentLength`, Azure `get_blob_properties().size`, each returning None rather than raising) and `TabularExecutor` now refuses >500 MB **before** downloading, with a post-download byte check as the fallback when the probe is unavailable. **F-12:** `render()` gained TWO ceilings β `_MAX_TABLES=150` and `_MAX_CATALOG_CHARS=250_000` β because a table-count cap alone leaves wide tables unbounded (20 tables Γ 300 cols is as fatal as 400 Γ 30). Measured after: 400Γ30 went from ~241k tokens to ~63k; 100Γ30 (~45k tokens) still renders **in full**, so no realistic catalog is touched. Truncation emits an explicit "N more tables not shown" line so the planner knows it saw a subset, and logs β that log is the signal to retune. 12 new tests | | |
| | 42 | **F-20 observability** β `degraded_seam=<name>` on every never-throw / silent-drop path | Rifqi | π | The 2026-07-23 report bug was invisible by construction: the record was dropped with zero logging. **Partially shipped 2026-07-23** β the 10 seams where silent degradation is user-visible now emit a stable `degraded_seam` field (+ `repr(e)` instead of `str(e)`, so an empty-`str()` Fernet error is no longer a blank log): `input_guard_fail_open`, `analysis_catalog_read`, `report_floor_record_read`, `traceability_persist`, `traceability_flush`, `chart_persist` (Γ2), `report_input_persist` (Γ2), `analysis_state_ensure`. **Remaining:** the other ~76 `except Exception` sites, most of which are in unwired routers (`db_client`, `data_catalog`, `users`) or non-live paths β deliberately not swept, since a blanket edit across unwired code is exactly the drive-by Β§7A forbids. Control flow unchanged throughout (Β§5.4) | | |
| | 43 | **Go identity contract** β what does Go forward, and when? | Rifqi β Harry | β¬ | **Now the sole path to caller auth** after #37 was unwired 2026-07-27. Python cannot authenticate the caller alone; either Go forwards a verified per-user token, or the FE (once we can change it) forwards the Go bearer token it already holds (`orchestrationApi.ts` shows the FE has one). Until then the live surface is unauthenticated and the #38 predicates are defensive only. Options for arming interim protection when FE access returns: the `server.js` BFF proxy (secret server-side) or route agentic calls through Go | | |
| | 44 | **F-17 / F-18 / F-25 compiler-parity batch** | Rifqi | β new | Shipped 2026-07-23. Three execution-verified review findings that had **no task row** β the tracker jumped from #39 to #40 and lost them. **F-17 (High):** the bare-select check was gated on `if ir.group_by`, so a mixed select with `group_by=[]` passed; Postgres then failed loudly but the pandas path silently DROPPED the column while `output_columns` still advertised it β a real-looking table with a fabricated all-null column. Fixed in `validator.py` (fires whenever any agg is present β no false positives possible) + a presence backstop in `TabularExecutor`. **F-18:** `astype(str)` ran before `na=False`, so `NULL LIKE '%an%'` matched the literal `"nan"`/`"None"`. **F-25:** `SqlCompiler` raised on an empty `in`/`not_in` where pandas and `_column_values`' own docstring implement the empty-set semantics β a two-step plan whose first step legitimately returned zero rows hard-failed instead of answering. 16 new parity tests; one stale test updated (it pinned the old F-25 raise) | | |
| | 45 | **F-9 PII in persisted artifacts** β mask the traceability preview + report evidence tables | Rifqi | β | Shipped 2026-07-24. `pii_flag` was an **ingestion-time control only**: it nulls `sample_values` into the planner prompt, but nothing stops the planner SELECTing a flagged column β and "list our top 20 customers" legitimately selects `customer_name`/`email`. Real values then reached two PERSISTED sinks: `message_traceability.data` (served by an unauthenticated GET, F-3 declined) and report evidence tables frozen permanently into `reports.content`. Fix = `retrieve_data` now carries `meta.pii_columns` (resolved through the IR select list, so aliases are honoured), and both sinks redact those cells to `[redacted]`. **Per the 2026-07-23 decision the ASSEMBLER still receives real values**, so the answer prose is unchanged and the question stays answerable β verified end-to-end: assembler input kept `Ada Lovelace`, the persisted preview showed `[redacted]`. Aggregates are deliberately NOT masked except `min`/`max`: `sum(salary)` identifies nobody, but `max(email)` returns one customer's actual address. Fails **open** (an unresolvable name is left unmasked, never a legitimate column wrongly blanked), so this is a mitigation, not a guarantee. 15 new tests | | |
| | 46 | **F-8 prompt-injection resistance** β planner / assembler / report_summary | Rifqi | β new | Shipped 2026-07-24. The three prompts that ingest customer data had **no injection rule**: `guardrails.md` is appended only in `chatbot.py`/`help.py`, and the `InputGuard` screens only the user's message β it never sees catalog or row content. This is the one attack the five query-defense layers structurally cannot see, because every IR the planner emits is individually valid. A hostile string only needs to reach a text column in the customer's OWN database (a product description, a support ticket, a form field) to be sampled into `sample_values` and rendered verbatim into the planner prompt. Fix = a purpose-written "content is data, never instructions" rule in each prompt + `<data>β¦</data>` delimiters around the catalog render and the run-state render. **Deliberately NOT `guardrails.md` wholesale** β its rules prescribe refusal sentences, and the planner's only free-text field is `infeasible_reason`, so those strings would surface there and could regress the Q2 data-gap path (tests pin their absence). Verified: planner eval **6/6, carried_over 5/5 green**; live hostile-catalog run planned only `t_products` and never touched the planted `employees.salary`; 10 new tests | | |
| | 47 | **Cheap-batch review fixes** β F-19, F-22, F-24, F-26, F-16, F-4 | Rifqi | β new | Shipped 2026-07-24, one low-risk batch. **F-19:** `sources` was missing entirely on `check` + router-`help` and came *after* `status` on the slow path β contract said always-first; additive fix, contract updated. **F-22:** `analysis_id` now 422s unless it parses as a UUID, on **both** live endpoints (chat + help kept identical on purpose). **F-24:** traceability/chart writes are skipped, and logged, when `analysis_id` is falsy β Go `0007` declares `analysis_id UUID NOT NULL` (+ FK on charts), so `analysis_id or ""` failed the insert and the never-throw seam lost the row silently. **F-26:** `QueryResult.error` uses `str(e) or repr(e)` β a Fernet `InvalidToken` reached the assembler/traceability/report as an EMPTY string; falling back only when `str()` is empty changes no existing message. **F-16:** retrieval cache key gains `redis_prefix` (two envs on one Redis cross-served results). **F-4:** non-Postgres sources now refused at the executor β **zero blast radius today** (Go's `isSupportedActive` allows only `postgres`), a tripwire so nobody re-enables a path that has no read-only session and no `statement_timeout`; legacy branch commented out per house convention, orphaned import commented with it | | |
| | 48 | **floor_08 β floor/body disagreement** | Rifqi | β | **Decided + shipped 2026-07-24 (lead).** #34's row-producing arm was unconditional, so a plan that HAS an `analyze_*` step which FAILED still cleared the floor on the strength of its upstream fetch β while the body rejected it. Because the "Attempted, Unresolved" section is commented out, such a run left no trace: as a session's only run it produced an empty report with the business question "Unanswered" (the #33 bug via another door). Fix = the arm now applies **only when the plan has no analysis step**, mirroring `has_reportable_result`; the shared `_plan_has_analysis` helper means the two predicates can no longer drift on that question. `floor_08` flips to `expected_ready: false`. The INTENDED asymmetry is preserved and re-verified: a zero-row retrieval still fails the floor but passes the body (floor_03). Readiness eval **17/17** | | |
| | 49 | **F-5 timeout does not stop the customer's query** | Rifqi | β new | Shipped 2026-07-24. `asyncio.wait_for` cancels the awaiting **coroutine**; the `to_thread` worker is not cancellable and runs to completion, holding a thread and a connection on the **customer's** database after we already answered "timed out". Two fixes. **(a) Dedicated bounded pool** β DB work moves to its own `ThreadPoolExecutor(50, 'dbexec')` via `run_in_executor`; abandoned workers previously accumulated in the shared default pool (`min(32, cpu_count+4)`) alongside every other `to_thread` caller, notably the tabular Parquet loader, so a few slow customer queries could stall unrelated work process-wide. **(b) Session hardening is no longer best-effort** β and a **latent bug** was found while reading it: both SETs shared one `try`, so a `statement_timeout` failure **skipped `default_transaction_read_only` entirely** and the connection served queries in a WRITABLE session behind a `logger.warning`. Now independent: read-only **fails the connection** if it cannot be set (a writable session against a customer DB is not something to degrade into); `statement_timeout` logs at **error** with `degraded_seam` but does not refuse service, since it is their I/O at risk rather than our correctness. β οΈ **Blast radius to watch:** if any deployment currently fails that SET silently, its sources now fail loudly instead. Believed impossible (Neon accepts it as a SET β the existing comment documents this), but it is the one judgement call here. 7 new tests | | |
| **Reading `eval/readiness/results/` (note for future sessions).** Four files are dated | |
| 2026-07-23. `β¦_150632.json` scores **4/15 (26.7%)** β that is **not** a product | |
| regression. It is the run that exposed the eval harness itself being broken by | |
| `fd4865b` (`_FakeRecord` had no `results_snapshot` for #34, `_FakeStore` took no | |
| `user_id` for #38; both errors were swallowed by `report_floor`'s fail-closed seam, so | |
| every case reported "not ready"). `β¦_150859` (15/15) is post-harness-fix, `β¦_150948` | |
| and `β¦_152615` (17/17) add the two cases that exercise #34. **The current baseline is | |
| `β¦_152615.json` β 17/17.** The intermediate files are kept as the audit trail for the | |
| drift; per Β§7F no result file is ever deleted or overwritten. | |
| **Not re-raised:** F-4 (non-Postgres read-only/timeout gap) was **downgraded to latent** β | |
| Go's `database_clients.Service.Create` enforces `isSupportedActive`, and only `postgres` is | |
| `active`, so no such source can be registered today. | |
| **Decided 2026-07-27 (lead) β F-2 gate unwired, not armed.** The service-secret gate (#37) was | |
| removed from the router mounts because the sole caller is a browser SPA we don't own and can't | |
| change to send the header; arming it would 401 the whole app. The live surface is **unauthenticated | |
| by design** until #43 (Go-forwarded identity). Not a gap to re-raise β it is a recorded posture. | |
| CORS was left at `["*"]` on purpose (tightening it needs the FE origin as config, which we chose | |
| not to set for now). | |
| ## 5. Critical path & sequencing | |
| - **Critical path:** ~~#22 (send Harry the `report_inputs` schema)~~ **β resolved 2026-07-22** β now **#32** (`message_traceability` + `message_charts` DDL to Harry) and **#31** (non-convergent migration set). HF deploy (#13) for the playground. (#4 β , #21 β ; Harry's #3 no longer blocks us β Python is getattr-tolerant.) | |
| - **Parallelizable now:** #31 + #32 (both are Harry handoffs). (#4 β , #11 β , #22 β done.) | |
| - **Harry-blocked / coordinated:** #3 (now π, blocks #4), #7 (Go endpoint), #18 (Go state ownership), #24 (contract). **#25 = chat-path migration to `analyses_messages` β a cutover blocker.** | |
| - **Demo gate (playground, #13):** deploy the remove-`problem_statement` work to HF β slow path (#15 β ) | |
| and the report path are verified locally, and #16 is resolved (#22 hands Harry the schema). **Keep it | |
| playground-able.** | |
| ## 6. Decisions still open (need the team / Harry / lead) | |
| - ~~`analysis_records`: dedorch-owned vs Python-owned (#16/#14).~~ RESOLVED: Python-owned + renamed **`report_inputs`** (#21 done); Harry's migration creates it (#22). | |
| - ~~Whether `help` is its own endpoint or via `call_agent` (#8).~~ RESOLVED: `help` is a `call_agent` intent (no own endpoint). | |
| - ~~Author display-name source for the report (#19).~~ RESOLVED: Python resolves `users.fullname` (fallback `user_id`); swap to a Go-passed name later if preferred. | |
| - ~~Keep vs drop `chat_history` in the report contract (#17).~~ RESOLVED: never in the contract; report is records-based (analysis_id + user_id only). | |
| - Confirm Go takes over analysis creation + both creation gates (data-first + mandatory fields) (#18). | |
| - **Report input for edit mode** β does Python need the last report content? (#24) | |
| - ~~`report_inputs` write scope β every agent call vs slow-path-only? (#21)~~ RESOLVED: one row per slow-path run (telemetry stays Langfuse). | |
| - **Python history source** β confirm Go's `analysis_message` (#25). | |
| - **`done.chart_count`** β additive field on the SSE `done` event so the FE can skip `GET /charts` | |
| on chartless turns? (Harry; SPINE_V2_PLAN Β§4.5.) Until decided the FE fetches unconditionally. | |
| - **W3 re-open timing** (deferred `analyze_*` activation) β Rifqi (deferred at the 2026-07-13 approval). | |
| - **INV-6 relaxation for S1b targeted repair** β team, only after S1a `repair_candidate` telemetry | |
| shows a meaningful hit-rate (SPINE_V2_PLAN Β§6). | |