File size: 27,157 Bytes
68f7345
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
c7d0e19
68f7345
c7d0e19
 
68f7345
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
c7d0e19
 
 
 
 
 
 
68f7345
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
c7d0e19
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
68f7345
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
c7d0e19
 
68f7345
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
487
488
489
490
491
492
493
494
495
496
497
498
# Lexsi DS Agent β€” Developer / KT Guide

A code-level companion to [`DSAGENT_OVERVIEW.md`](DSAGENT_OVERVIEW.md) (which is
the what/why). This is the **how**: file map, control flow, the contracts you'll
touch, configuration, and the full benchmark-running mechanics. Read the overview
first, then this.

---

## 1. Mental model in one paragraph

One **question** β†’ one `AgentLoop.run()` β†’ a **ReAct loop**: build a system prompt
(framing + active-dataset summary + tool catalog), ask the LLM for a JSON decision
(`{"thought","tool","args"}` or `{"thought","final_answer"}`), validate args with
the tool's pydantic schema, run the tool against a shared `AgentContext`, append the
observation to history, repeat until a final answer or the step budget. Tools read
and write a shared `ctx.cache` scratchpad. The data substrate is always **DuckDB**;
tabular ML runs on **remote Lexsi pods** via the SDK. The loop is generic β€” it has
no opinion on which tools to call.

---

## 2. Repo layout (code-oriented)

```
lexsi_ds/
  agent/
    loop.py            AgentLoop.run() β€” the ReAct orchestrator (start here)
    context.py         AgentContext, DatasetHandle, TableInfo, ColumnInfo
    prompts.py         PLANNER_SYSTEM, render_tool_catalog, tool-selector (intent routing)
    datasource.py      DuckDB handle loaders + multi-DB materialization + DB cache
    tools/
      base.py          Tool / ToolSpec / ToolResult contract (the interface)
      __init__.py      REGISTRY: dict[str, Tool]  (30 tools)
      <name>.py        one tool per file, each exports `TOOL: Tool`
      _tabular_common.py  shared helpers for the model-lifecycle tools
                          (resolve_tab_project, active_model_name, df_records, …)
    kg_index.py        embedding/lexical index over the PKDD knowledge graph
    table_index.py     embedding index over table schemas (schema linking, big DBs)
    session.py         SessionContext + multi-turn merge
    turn_classifier.py classify a turn: new / refine / abandon
    reference_resolver.py  resolve "it/that/those" against prior turns
    samples.py         curated demo questions (by_id, SAMPLES) β€” UI chips + tests
  llm/client.py        LLMClient protocol, StubLLMClient, LexsiTextClient, cache, timeout
  schema/
    raw.py             PKDD DDL introspection (introspect())
    context_graph.py   ContextGraph β€” the banking knowledge graph
  data/loader.py       PKDD TSVs -> artifacts/financial.duckdb
  paths.py             DUCKDB_PATH, ARTIFACTS_DIR, etc.

app/gradio_app.py      the UI (build_ui(); demo.launch on GRADIO_PORT)

benchmark/
  dab/
    adapter.py         parse a DAB dataset -> DatasetHandle + tasks (+ DB cache key)
    runner.py          run agent per DAB task, score Pass@1 via the task's validate.py
    DataAgentBench/    the (gitignored) upstream checkout + DB files
  pkdd_runner.py       run agent over bench/pkdd/questions.yaml, auto-score
  datagen/             template pipeline that GENERATES bench/pkdd/*.yaml

bench/pkdd/
  questions.yaml       40 single-turn bench items (analytic/predictive/adversarial)
  sessions.yaml        multi-turn sessions
  p2_feature_engineering.yaml   10 feature-eng items (separate `assertions` schema)
benchmark/fixtures/pkdd_seed.duckdb   pinned PKDD snapshot the bench scores against

scripts/
  repro_tabicl/                  TabICL inference + per-case XAI failure repro
  repro_model_lifecycle/         model-lifecycle SDK probe harness (self-contained):
    data/                        the PKDD CSVs (local copy)
    repro_model_lifecycle.py             TabICL-first (clean)
    repro_model_lifecycle_xgboost_first.py  XGBoost-first (#05-013 cascade repro)

tests/                 pytest suite (test_b_tools, test_v1_multiturn, test_dab_*, …)
docs/                  design + results + this guide
```

---

## 3. The run lifecycle ([loop.py](../lexsi_ds/agent/loop.py))

`AgentLoop(ctx, registry=REGISTRY, max_steps=18).run(question) -> AgentRun`

```
run(question):
  1. (multi-turn) classify_turn() -> new/refine/abandon; resolve_reference()
     - "abandon" short-circuits with a polite answer.
  2. ctx.cache["question"] = question
  3. system = _render_system(...)                 # built ONCE per run, re-sent every step
       - _select_tools(question): LLM intent classifier -> render only relevant tool group(s)
       - render_dataset_block(ctx) + render_tool_catalog(tools) + PLANNER_SYSTEM
  4. history = ["USER QUESTION: ..."]
  5. while n < max_steps:
       user_msg = "\n\n".join(history) + "You have {max_steps-n} tool calls left..."
       decision = parse(ctx.llm.complete(system, user_msg))
       if "final_answer": finalize + return
       tool = registry.get(decision["tool"])
       result = _call_tool(tool, decision["args"])   # pydantic-validate then tool.run()
       history.append("TOOL CALL ...\nOBSERVATION: " + result.summary[:1500])
       # anti-loop: same tool fails 2x in a row -> inject "[LOOP GUARD] re-plan" directive
       if result.meta.get("pause_for_user"): return (clarify pause)   # interactive only
  6. budget exhausted -> _force_final(): ask once for a final answer
```

Key methods: `run`, `_render_system`, `_select_tools`, `_call_tool`, `_force_final`.
Dispatch **always uses the full `self.registry`**; the tool selector only narrows
what the planner *sees* in the prompt, never what it can call.

Data classes:
- `AgentRun` β€” `run_id, question, dataset_id, steps[], final_answer, ok, error, pending_clarification, system_prompt`.
- `Step` β€” `n, thought, tool, args, result(ToolResult), raw_llm_text, latency_s, prompt`. The full per-turn user message lives on each Step (so trajectories are reproducible).

---

## 4. Core data structures ([context.py](../lexsi_ds/agent/context.py))

**`AgentContext`** β€” built once per run, passed to every tool:
- `dataset: DatasetHandle`, `run_id`
- `org / text_project / tab_project / tab_projects` β€” Lexsi SDK handles (may be None offline)
- `llm` β€” an `LLMClient`
- `cache: dict` β€” run-scoped scratchpad (see Β§6)
- `session` β€” optional `SessionContext` (multi-turn)
- `interactive: bool` β€” True in the UI (clarify pauses); False in bench/headless (clarify falls back to its `default`)
- `.duck(read_only=None)` β€” open a DuckDB connection on `dataset.duckdb_path` (read-only for bundled, writable otherwise)

**`DatasetHandle`** β€” `id, kind, duckdb_path, tables: list[TableInfo], connector_id, kg`. `kind ∈ {bundled, upload, connector, attached}`. Everything funnels through `ctx.duck()` β†’ `duckdb_path`, so "switch dataset" = build a new handle and assign `ctx.dataset`.

**`TableInfo`** = `name, columns: list[ColumnInfo], n_rows, comment`. **`ColumnInfo`** = `name, dtype, comment, sample` (the `sample` value is the per-column representative value the planner sees in the schema β€” populated at describe time).

---

## 5. The tool contract β€” how to add a tool ([tools/base.py](../lexsi_ds/agent/tools/base.py))

A tool is any object with `.spec: ToolSpec` and `.run(args, ctx) -> ToolResult`.

```python
# lexsi_ds/agent/tools/my_tool.py
from pydantic import BaseModel, Field
from lexsi_ds.agent.context import AgentContext
from lexsi_ds.agent.tools.base import Tool, ToolResult, ToolSpec

class MyToolArgs(BaseModel):
    table_or_label: str = Field(..., description="...")   # match sibling arg names!

def _run(args: MyToolArgs, ctx: AgentContext) -> ToolResult:
    con = ctx.duck()
    ...
    return ToolResult(ok=True, summary="short text the LLM sees (<~200 tok)",
                      payload={"dataframe": df})           # rich data; LLM never sees it

class _MyTool:
    spec = ToolSpec(name="my_tool", description="...(LLM sees verbatim)...",
                    args_schema=MyToolArgs, returns="what the observation looks like")
    def run(self, args, ctx): return _run(args, ctx)

TOOL: Tool = _MyTool()
```

Then register it: add `from .my_tool import TOOL as my_tool_tool` and an entry in
`REGISTRY` in [tools/__init__.py](../lexsi_ds/agent/tools/__init__.py). If it fits an
intent group, add it to `_TOOL_GROUPS` in [prompts.py](../lexsi_ds/agent/prompts.py)
so the tool selector surfaces it (see Β§8).

Contract notes:
- `summary` is the **only** thing the LLM sees; `payload` is for the UI / downstream tools.
- Return `ok=False` with an **actionable** `summary` on expected failures β€” the planner reads it and re-plans. Unexpected exceptions propagate; the loop wraps them.
- **Arg-name convention:** sibling tools that take a source use `table_or_label`. Match it (or alias via pydantic `AliasChoices`) β€” mismatches show up as `validation_error` and waste planner steps (this bit us; see `transform_column`).
- Args are validated in `_call_tool` via `tool.spec.args_schema.model_validate(raw_args)`.

---

## 6. The cache contract (`ctx.cache`)

Run-scoped (session-scoped in multi-turn). Producers β†’ consumers, by key:

| Key | Producer | Consumer |
|---|---|---|
| `question` | loop | tool selector |
| `tool_intents::<q>` | `_select_tools` | (memoize the intent classification) |
| `sql_result:<label>` | `run_sql` | profile_data, sample_rows, transform_column, predict |
| `last_sql_result`, `last_sql` | `run_sql` | PKDD scorer, downstream |
| `profile:<label>` | `profile_data` | train_tabular_model, dq checks |
| `last_train_columns` | `train_tabular_model` | predict (schema alignment) |
| `last_predictions`, `last_predict_tag`, `last_model_id` | `predict` | explain_prediction, evaluate_predictions |
| `last_eval` | `evaluate_predictions` | summarize_result |

When adding a tool that hands data downstream, write a documented key here.

---

## 7. Data layer ([datasource.py](../lexsi_ds/agent/datasource.py))

Handle loaders (each returns a `DatasetHandle`):
- `load_pkdd_handle()` β€” bundled PKDD (`DUCKDB_PATH = artifacts/financial.duckdb`), KG attached.
- `load_upload_handle(spec)` β€” CSV/Parquet β†’ per-session DuckDB.
- `load_s3_connector_handle(spec)` β€” `read_parquet(s3://…)` views.
- `load_attached_handle(specs, *, dataset_id, target_path=None)` β€” fold heterogeneous DBs into **one** DuckDB.
- `load_postgres_connector_handle` / `load_mongo_connector_handle` β€” live DB connectors (reuse the materialize helpers).

**Multi-DB fold-in** (`_build_attached_duckdb`): each source becomes `<alias>_<table>`:
- sqlite/duckdb β†’ `ATTACH` + `CREATE TABLE AS SELECT` (`_materialize_file_db`)
- postgres β†’ restore the `.sql` dump into an embedded **pgserver**, copy via postgres scanner (`_materialize_postgres_dump`); tears down the `pgdata` dir after.
- mongo β†’ read `.bson` server-lessly, flatten nested β†’ JSON text (`_materialize_mongo_dump`).

**Schema sampling:** `_describe_tables` β†’ `_column_samples` runs one `SELECT * LIMIT 5` per table and stores a representative value on `ColumnInfo.sample`. This is what lets the planner see "Date: VARCHAR sample='31 Dec 1986, 00:00'" and parse with `TRY_CAST`/`strptime` instead of naive casts.

**DB cache (build-once)** β€” `load_attached_handle(target_path=...)`: if the master exists at `target_path`, reuse it; else materialize once (atomic `.building` β†’ rename). Each run gets an isolated **working copy** (`shutil.copyfile` β†’ `artifacts/sessions/attach_*.duckdb`) so agent writes never mutate the master. The DAB adapter computes the master path + content-fingerprint (see Β§13).

Session files live in `artifacts/sessions/` via `_new_session_db()`; the DAB runner deletes the per-dataset working copy when done (`_cleanup_working_copy`).

---

## 8. Prompts + the token-lean tool selector ([prompts.py](../lexsi_ds/agent/prompts.py))

- `PLANNER_SYSTEM` β€” the ReAct framing + hard rules + JSON response format.
- `render_tool_catalog(registry)` β€” renders each tool as `## name` + description +
  **one-line arg signature** (`_compact_args`, not full JSON Schema β€” ~64% fewer tokens)
  + Returns. Enums/Literals keep their choices.
- `render_dataset_block(ctx)` / `render_session_context(...)` β€” dataset summary + multi-turn prefix.
- **Tool selector** (gated, LLM-based): `_select_tools` (in loop.py) calls a one-shot
  classifier (`TOOL_INTENT_SYSTEM`) once per run, cached. `parse_tool_intents` β†’
  `subset_for_intents(intents, registry)` renders **core tools + the union of matched
  groups** (`_CORE_TOOLS`, `_TOOL_GROUPS`); ambiguous/multi-intent β†’ full catalog. Toggle
  with `LEXSI_TOOL_SELECTOR=0`. Dispatch is always against the full registry.

---

## 9. LLM client ([llm/client.py](../lexsi_ds/llm/client.py))

- `LLMClient` protocol: `.complete(system, user) -> LLMResult(text, raw)`.
- Implementations: `StubLLMClient` (offline sentinels), `EchoGoldLLMClient` (tests),
  `LexsiTextClient` (real; `LexsiTextClient.from_env()`), built via `factory("lexsi"|"stub")`.
- **Caching:** in-memory LRU (`LEXSI_LLM_CACHE_MAX`, default 256) keyed on
  `(model, provider, max_tokens, system, user)`; optional **disk** layer when
  `LEXSI_LLM_CACHE_DIR` is set (the bench runners set it β†’ replays across runs). Disable
  with `LEXSI_LLM_CACHE=0`.
- **Timeout/retry (important):** every gateway call goes through `_call_with_timeout`
  (daemon-thread guard, `LEXSI_LLM_TIMEOUT_S` default 180, `LEXSI_LLM_RETRIES` default 1).
  Without it a stalled gateway request freezes the whole run β€” this was a real incident.
- **Model-family handling:** GPT-5 / o-series reject `max_tokens`; the client POSTs
  directly with `max_completion_tokens` (`_post_chat_with_completion_tokens`).

---

## 10. Multi-turn

`SessionContext` ([session.py](../lexsi_ds/agent/session.py)) seeds `ctx.cache` from the
session at run start and lifts changes back at the end. `classify_turn`
([turn_classifier.py](../lexsi_ds/agent/turn_classifier.py)) labels each turn
new/refine/abandon; `resolve` ([reference_resolver.py](../lexsi_ds/agent/reference_resolver.py))
resolves anaphora and hands the planner an explicit reference.

---

## 11. Configuration (environment variables)

**Required for the real LLM path:** `SDK_ACCESS_TOKEN`, `LEXSI_ORG_NAME`,
`LEXSI_WORKSPACE_NAME`, `LEXSI_TEXT_PROJECT_NAME`, `LEXSI_TEXT_PROVIDER`,
`LEXSI_TEXT_MODEL`. Optional: `LEXSI_TEXT_PROVIDER_API_KEY`, `LEXSI_TEXT_MAX_TOKENS`.

**Tabular:** `LEXSI_TABULAR_PROJECT_NAME` (per-run projects created on train, so usually unset).

| Var | Default | Purpose |
|---|---|---|
| `LEXSI_TOOL_SELECTOR` | `1` | gate the LLM tool selector (`0` = always full catalog) |
| `LEXSI_LLM_TIMEOUT_S` | `180` | hard timeout per gateway call |
| `LEXSI_LLM_RETRIES` | `1` | retries on timeout/failure |
| `LEXSI_LLM_CACHE` | `1` | in-memory LLM cache on/off |
| `LEXSI_LLM_CACHE_DIR` | β€” | enable on-disk LLM cache (bench runners set this) |
| `LEXSI_LLM_CACHE_MAX` | `256` | LRU size |
| `DUCKDB_PATH` | `artifacts/financial.duckdb` | bundled PKDD DB |
| `DAB_ROOT` | `benchmark/dab/DataAgentBench` | DAB checkout |
| `GRADIO_PORT` | `7860` (7850 in Docker) | UI port |
| `LEXSI_UI_RUN_TIMEOUT_S`, `LEXSI_UI_HEARTBEAT_S`, `LEXSI_UI_SSR` | β€” | UI tuning |
| `LEXSI_TRAIN_TIMEOUT_S`, `LEXSI_PREDICT_TIMEOUT_S`, `LEXSI_XAI_BUDGET_S` | β€” | Lexsi pod waits |
| `LEXSI_ROUNDTRIP_MODEL/_PROVIDER/_DISABLE` | β€” | datagen Stage-4 back-translator |

(Connector creds β€” `PG_*`, `MONGO_URI`, `AWS_*`, `MYSQL_*` β€” are read by the live `connect_datalake` paths.)

---

## 12. Local dev setup

```bash
# deps (extras: lexsi, ui, dev, embeddings, connectors)
uv sync --extra lexsi --extra ui --extra dev --extra embeddings --extra connectors

# bootstrap the bundled PKDD DuckDB (one-time, ~10s)
uv run python -m lexsi_ds.data.loader

# drive the agent from Python (offline stub LLM)
uv run python -c "
from lexsi_ds.agent import AgentContext, AgentLoop
from lexsi_ds.agent.datasource import load_pkdd_handle
from lexsi_ds.llm.client import factory
ctx = AgentContext(dataset=load_pkdd_handle(), run_id='local', llm=factory('stub'))
print(AgentLoop(ctx=ctx).run('Top 5 districts by number of accounts.').final_answer)
"

# UI
uv run python -m app.gradio_app          # http://localhost:7860

# tests (LLM-gated ones skip without the Lexsi env)
uv run pytest -q
```

---

## 13. Benchmark mechanisms (all of them)

There are **three** harnesses. All write full trajectories so failures are debuggable.

### 13a. DataAgentBench (external; Pass@1) β€” `benchmark/dab/`

What it is: 17 datasets across sqlite/duckdb/postgres/mongo; each task ships a
question + `ground_truth.csv` + its own `validate.py`. We score **Pass@1** by calling
that `validate.py` on the agent's final answer.

Code:
- [adapter.py](../benchmark/dab/adapter.py): `load_dataset(dir)` β†’ `DabDataset`
  (parses `db_config.yaml` into `AttachSpec`s via `parse_db_config`, loads each task's
  `validate.py` via `_load_validate`). `build_handle(dataset, cache_dir=...)` β†’
  `load_attached_handle` with a content-keyed master path
  (`<cache_dir>/<dataset>__<_fingerprint>.duckdb`).
- [runner.py](../benchmark/dab/runner.py): `main` β†’ `run_dataset` (builds the handle
  once per dataset, loops tasks) β†’ `run_task` (composes the question with the DAB
  description + table-name mapping, runs `AgentLoop`, calls `validate`). DB working
  copy cleaned per dataset (`_cleanup_working_copy`).

Setup + run:
```bash
git clone https://github.com/ucbepic/DataAgentBench benchmark/dab/DataAgentBench
uv sync --extra connectors        # needed for the postgres/mongo datasets
export SDK_ACCESS_TOKEN=… LEXSI_WORKSPACE_NAME=… LEXSI_TEXT_PROJECT_NAME=…  # + org/model/provider

# one dataset / one task / sweep all
uv run python -m benchmark.dab.runner --dataset DEPS_DEV_V1 --llm lexsi
uv run python -m benchmark.dab.runner --dataset DEPS_DEV_V1 --query query1 --llm lexsi
uv run python -m benchmark.dab.runner --all --llm lexsi

# the sqlite/duckdb "core" (5 datasets / 17 tasks, no connectors extra), real Pass@1
uv run python -m benchmark.dab.runner --all --db-types sqlite,duckdb --llm lexsi \
  --iterations 3 --max-steps 20 --log-dir artifacts/dab_runs/run1
```

Flags: `--dataset/--all`, `--db-types t1,t2` (filter `--all`), `--query`,
`--iterations`, `--max-steps` (default 12), `--use-hints`, `--log-dir`,
`--db-cache-dir`/`--no-db-cache`, `--llm-cache-dir`/`--no-llm-cache`, `-v`.

Caching (two layers, both speed re-runs):
- **DB cache** `artifacts/dab_dbs/` β€” each dataset's fold-in materialized once; reused across runs; per-run working copy auto-deleted.
- **LLM disk cache** `artifacts/llm_cache/` β€” gateway responses keyed on the prompt; a failed sweep resumes cheaply.

Outputs: per-task trace JSON in `--log-dir` (`<dataset>_<query>_iter<N>.json` β€” full
steps, prompts, raw LLM output, payloads, validator verdict) + a console scoreboard
(per-dataset pass counts, wall vs agent time, LLM reuse). Latest results writeup:
[`dab_results_v1.md`](dab_results_v1.md); integration detail: [`dab_integration.md`](dab_integration.md).

### 13b. PKDD-Curated (internal) β€” `benchmark/pkdd_runner.py`

What it is: our own suite over the bundled PKDD dataset, scored against the **pinned
fixture** `benchmark/fixtures/pkdd_seed.duckdb`. Question schema in
[`benchmark/datagen/types.py`](../benchmark/datagen/types.py) (`Row`); the live set is
[`bench/pkdd/questions.yaml`](../bench/pkdd/questions.yaml) (40 items: analytic,
predictive, adversarial).

Scoring ([pkdd_runner.py](../benchmark/pkdd_runner.py), by `scoring_rule`):
- `row_set_equality` (analytic) β†’ `_score_row_set`: agent's `last_sql_result` vs
  `gold_sql` executed on the fixture, as a multiset of sorted-value rows (column-rename
  tolerant; lenient scalar fallback).
- `exact_match` (numeric) β†’ `_score_numeric`: reference value within `tolerance`.
- `judge_rubric` (adversarial) β†’ `_score_behavioral`: heuristic β€” fail if it used a
  `forbidden_tool` or answered when it should clarify/refuse.
- `ndcg@k` (predictive) β†’ captured but **not auto-scored** unless `--include-predictive`
  (those train real models, slow). Plan precision/recall computed for every row.

Run:
```bash
export SDK_ACCESS_TOKEN=… LEXSI_WORKSPACE_NAME=… LEXSI_TEXT_PROJECT_NAME=…  # + org/model/provider
uv run python -m benchmark.pkdd_runner --llm lexsi --log-dir artifacts/pkdd_runs
# scope / debug
uv run python -m benchmark.pkdd_runner --tier analytic --limit 5 --llm lexsi
uv run python -m benchmark.pkdd_runner --include-predictive --llm lexsi
```
Flags: `--tier`, `--include-predictive`, `--limit`, `--llm`, `--log-dir`,
`--llm-cache-dir`/`--no-llm-cache`, `-v`. Outputs: per-row trace JSON + a per-tier
scoreboard (pass/scored + plan p/r). The agent queries the fixture (handle from
`load_pkdd_handle()` with `duckdb_path` repointed to the fixture).

> Note: `bench/pkdd/sessions.yaml` (multi-turn) and `p2_feature_engineering.yaml`
> (separate `assertions` schema) are **not** run by `pkdd_runner` yet.

### 13c. Datagen pipeline (generates the PKDD questions) β€” `benchmark/datagen/`

Template-based; authoring is per-template, so ~10 templates yield 500+ rows. Stages:
column inventory β†’ templates β†’ instantiate β†’ execute gold SQL on the pinned fixture β†’
LLM round-trip verify β†’ predictive enrichment β†’ multi-turn compose β†’ emit YAML.

```bash
uv run python -m benchmark.datagen pin-fixture        # snapshot artifacts/financial.duckdb
uv run python -m benchmark.datagen run --skip-roundtrip   # fast (skip LLM Stage 4)
uv run python -m benchmark.datagen run                    # full (needs Lexsi text env)
```
Every emitted row carries `fixture_sha`; a scoreboard from a different sha isn't comparable.
The grading design (5 axes, tiers, pass^k) lives in [`v1_benchmarking.md`](v1_benchmarking.md).

### 13d. SDK repro / probe scripts β€” `scripts/repro_*`

Standalone, self-contained scripts that hit the live Lexsi SDK against a real
trained model and report OK/FAIL per call. They're how we validate the
SDK-backed tools and file platform bugs. Each catches every call so one run
produces a complete report; each leaves its project on the platform for
Activity-Log inspection.

- **`scripts/repro_tabicl/`** β€” reproduces TabICL `model_inference` + per-case XAI
  failures (empty `Exception` on T4 pods, inference hang/timeout, `case_predict`
  `ValidationError`). Ships its own PKDD CSVs in `data/`.
- **`scripts/repro_model_lifecycle/`** β€” probe harness for the **model-lifecycle
  tools** (`list_available_models`, `select_active_model`, `compare_models`,
  `check_drift`, `monitor_performance`) + the future explanation methods. Trains
  TabICL + XGBoost on the (local) PKDD CSVs and calls each SDK method. Two scripts:
  `repro_model_lifecycle.py` (TabICL-first, the clean path) and
  `repro_model_lifecycle_xgboost_first.py` (XGBoost-first, reproduces the `#05-013`
  cascade). Run: `uv run python scripts/repro_model_lifecycle/repro_model_lifecycle.py`
  (env: `SDK_ACCESS_TOKEN`, `LEXSI_WORKSPACE_NAME`, `LEXSI_ORG_NAME`).

**Live SDK findings from the lifecycle repro (last run 13 OK / 4 FAIL):**
- `#05-013` β€” classic-ML (XGBoost) explainability scans **every stored column,
  including the server-injected `tag` column**, and fails `float()` on a
  non-numeric one. The SDK uploads clean data + an explicit `feature_include`
  (`tabular.py:285,368`); explainability ignores it. Model still **builds**
  (non-fatal), but `upload_data` raises and **rolls back the upload** β†’ cascade
  (`train_model` β†’ "Upload files first", predict upload β†’ "Project Config is
  required"). TabICL's foundation path doesn't trip it. **Fix for the agent:**
  treat `#05-013` as success-with-warning, and train foundation-first / drop or
  encode non-numeric columns.
- `evals_tabular(<foundation model>)` β†’ `Exception: 'model'` (KeyError) β€” works
  for classic ML, breaks for TabICL. Blocks `compare_models` on foundation models.
- `get_model_performance(...)` β†’ `'utf-8' codec can't decode …` β€” response decode
  bug. Blocks `monitor_performance`.
- per-case XAI (`case_predict`, `get_feature_importance`) needs (a) the model
  **active** + an inference run first (else "Inference status inactive"), and
  (b) SHAP **computed at train** (`xai=["shap"]`) β€” which itself trips `#05-013`
  for classic ML. `case_predict(risk_policies=True)` works without SHAP.

The five lifecycle tools degrade gracefully against these (per-model error capture
in `compare_models`; actionable `ok=False` in `monitor_performance`). Record new
findings in [`sdk_issues.md`](sdk_issues.md).

---

## 14. Deploy

- **Docker** ([Dockerfile](../Dockerfile)): `python:3.10-slim`, CPU-only torch, pinned
  deps, **bakes** `artifacts/financial.duckdb` (the PKDD source TSVs are Git-LFS
  pointers), serves Gradio on **7850**. Build for the platform arch:
  `docker buildx build --platform linux/amd64 --provenance=false --sbom=false -t bplexsi/lexsi-ds-agent:prod_vN --push .`
- **HF Space** β€” a Gradio Space (`README.md` frontmatter `sdk: gradio`,
  `app_file: app/gradio_app.py`); the Dockerfile is ignored there.
- **Lexsi platform** β€” [config.yaml](../config.yaml): one container on port 7850; the
  real `SDK_ACCESS_TOKEN` goes in the platform env tab (the repo keeps a placeholder).

---

## 15. Sharp edges (things that bit us)

- **No LLM timeout = frozen run.** Always keep `_call_with_timeout` in the path (Β§9).
- **Tool arg-name drift.** The planner reuses sibling names (`table_or_label`,
  `source_columns`, `prompt`); mismatched schemas β†’ `validation_error` loops. Alias them.
- **DataFrame truthiness.** `x = cache.get(a) or cache.get(b)` raises on a DataFrame; use explicit `is None`.
- **String-typed dates/numbers.** Source columns are often VARCHAR; rely on the
  `ColumnInfo.sample` + the `TRY_CAST`/`try_strptime` guidance in the text_to_sql prompt.
- **DB cache contamination.** The master must stay pristine β€” agents write to the
  working copy, never the master (Β§7).
- **`artifacts/sessions/` growth.** Working copies/pgdata are cleaned now; if you add a
  loader, clean up after it.
- **Single-run bench numbers are noisy** (DAB sits in an 8–9/17 band, 11/17 ceiling).
  Use `--iterations` and read the per-task/failure breakdown, not the headline.

---

## 16. "Where do I look for…"

| Want to… | Go to |
|---|---|
| change the planner behavior / rules | `prompts.py` (`PLANNER_SYSTEM`) |
| add/modify a tool | `tools/<name>.py` + `tools/__init__.py` (+ `_TOOL_GROUPS`) |
| change which tools the planner sees | `prompts.py` selector (`_CORE_TOOLS`, `_TOOL_GROUPS`) |
| connect a new datasource | `datasource.py` (loaders) + `connect_datalake` tool |
| change LLM / caching / timeout | `llm/client.py` |
| run/extend DAB | `benchmark/dab/{adapter,runner}.py` |
| run/extend PKDD | `benchmark/pkdd_runner.py` + `bench/pkdd/questions.yaml` |
| regenerate PKDD questions | `benchmark/datagen/` |
| model-lifecycle tools (list/select/compare/drift/monitor) | `tools/{list_available_models,select_active_model,compare_models,check_drift,monitor_performance}.py` + `tools/_tabular_common.py` |
| probe the live SDK / file an SDK bug | `scripts/repro_model_lifecycle/`, `scripts/repro_tabicl/` β†’ `docs/sdk_issues.md` |
| the UI | `app/gradio_app.py` |
| roadmap / what's next | `v1_tools.md`, `tools_remaining_and_dab_hardening.md` |