File size: 6,842 Bytes
3993fe3
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
# One Pipeline β€” Removing the Legacy Path and the Mode Flag

Goal: there is exactly one way the app builds a demo:

    research -> blueprint -> engine -> validation -> derived DDL -> Snowflake load

No `data_generation_mode` setting, no `DEMOPREP_DATASET_FIRST` env flag, no
"falling back to DDL-first" branch, no LegitData population path. A build
either succeeds with the story the user asked for, or fails with an error
that says exactly why. Nothing in between.

Why this matters beyond tidiness: today the fork is not just leftover β€” it is
**live and defaulted to legacy**. `load_default_settings()` sets
`'data_generation_mode': 'legacy'`, so a fresh user never touches the new
path at all. And even with dataset-first enabled, `build_dataset_first_demo`
returning None triggers a *silent* fallback to LLM-DDL generation β€” a second
router hiding behind the first. Both are exactly the class of "quietly do
something else" behavior that produced the McKesson incident. The new
pipeline never returns None (it raises), so the fallback has nothing to fall
back from; it is dead code that can only cause harm.

---

## 1. chat_interface.py β€” deletions

### 1a. The mode check (delete the whole method)

```python
def _dataset_first_enabled(self) -> bool:
    """Read dataset-first mode from app config/.env or saved settings."""
    setting = str(self.settings.get("data_generation_mode", "")).strip().lower()
    env_flag = os.getenv("DEMOPREP_DATASET_FIRST", "").strip().lower()
    return setting in {...} or env_flag in {...}
```

Delete it. Also delete `'data_generation_mode': 'legacy'` from the defaults
dict in `load_default_settings()`, and any Settings-tab UI bound to it.

### 1b. The schema-creation fork

Current shape:

```python
self._dataset_first_bundle = None
if self._dataset_first_enabled():
    ...build = build_dataset_first_demo(...)
    if build:
        ...success path...
        return response, self.ddl_code
    self.log_feedback("...no scenario generator matched... Falling back to DDL-first generation.")

# Build DDL generation prompt with geo context
geo_scope = ...
...entire legacy LLM-DDL stage: schema_prompt, retries, 'None' checks,
   regex stripping of db prefixes, etc...
```

New shape β€” no condition, no fallback:

```python
from demoprep_app.pipeline.build_demo import build_demo

row_count_guidance = int(self.settings.get("fact_table_size", 5000) or 5000)
company_name = self.demo_builder.extract_company_name()
build = build_demo(
    company_name=company_name,
    company_url=self.demo_builder.company_url,
    use_case=self.demo_builder.use_case,
    vertical=self.vertical,
    function=self.function,
    row_count_guidance=row_count_guidance,
    research_context="\n\n".join(part for part in [
        getattr(self.demo_builder, "combined_research_results", "") or "",
        getattr(self, "generic_use_case_context", "") or "",
    ] if part),
    # verbatim custom-tab text -> binding directives (McKesson fix)
    user_request=getattr(self, "generic_use_case_context", "") or self.demo_builder.use_case,
    llm_model=self.settings.get("model", DEFAULT_LLM_MODEL),
    progress_callback=self.log_feedback,
    prompt_logger=self._prompt_logger,
)
self._demo_bundle = build.dataset
self.demo_builder.schema_generation_results = build.ddl
self.ddl_code = build.ddl
for warning in build.warnings:
    self.log_feedback(f"⚠️ {warning}")
```

`build_demo` raises on failure with the validation report in the message β€”
let the existing per-stage exception handling surface it. Do NOT wrap it in
a try/except that falls through to anything else.

Everything from `# Build DDL generation prompt with geo context` down through
the DDL retry loop, the `'CREATE TABLE' not in ddl_result` checks, the
`ddl == "None"` check, and the db-prefix regex: **delete**. That is the
legacy stage. (Geo scope belongs in the blueprint author prompt if you still
want it β€” one line in PROMPT_TEMPLATE β€” not in a parallel DDL stage.)

### 1c. The population fork

Current shape:

```python
if self._dataset_first_bundle:
    ...populate_dataset_bundle(...)
    return

success, message, results = populate_demo_data(   # LegitData path
    ddl_content=ddl, ..., size=size, ...
)
```

New shape β€” the bundle always exists:

```python
from demoprep_app.integrations.snowflake import populate_dataset_bundle
results = populate_dataset_bundle(deployer.connection, schema_name, self._demo_bundle)
```

Delete the `populate_demo_data(...)` call and, once nothing else imports it,
the LegitData population module itself, plus the 45-minute `POP_TIMEOUT`
machinery sized for it (the bundle load is a bounded INSERT of in-memory
rows β€” it does not need a repair-loop timeout).

### 1d. Cosmetic followers

- `generation_mode: "dataset_first" if self._dataset_first_bundle else "legacy"`
  in `_deploy_meta` β†’ just `"generation_mode": "blueprint"`.
- Rename `self._dataset_first_bundle` β†’ `self._demo_bundle`.
- UI copy "Dataset-First Schema Creation Complete!" β†’ "Schema Creation
  Complete!" β€” there is no other kind. Consider surfacing
  `build.blueprint.insights` headlines and `build.validation.summary()` in
  that message: it tells the SE what stories are planted before they deploy.

## 2. Settings / env cleanup

- Remove `data_generation_mode` from `SETTINGS_SCHEMA`, Supabase-saved
  settings, and any docs mentioning `DEMOPREP_DATASET_FIRST`.
- Architecture doc: delete the "when dataset-first mode is enabled" section β€”
  the flow it describes is now the only flow.

## 3. Module deletions (after the above, nothing imports them)

- Legacy DDL prompt stage inside chat_interface.py (inline, per 1b)
- LegitData population path (`populate_demo_data` and its module)
- `demoprep_app/scenario/selector.py`, `extractor.py`, `families.py`
- `demoprep_app/dataset/generators/` (all)
- Finally: `demoprep_app/pipeline/dataset_first.py` (the compat shim), once
  imports point at `demoprep_app.pipeline.build_demo`.

## 4. What "one process" does NOT remove

- "Use existing model" flows (skip generation entirely) β€” orthogonal, keep.
- The approval checkpoint ("Do you approve this DDL?") β€” keep; it now shows
  derived DDL + planted-insight headlines instead of LLM-guessed DDL.
- The e2e quality harness β€” point it at `build.validation` and assert
  `passed`.

## 5. Order of operations (safe cutover)

1. Land the new package modules (already tested standalone).
2. Apply 1b + 1c so the app runs only the new pipeline.
3. Run one real demo end to end (Snowflake + ThoughtSpot).
4. Delete legacy modules and the mode setting (steps 2-3 above).
5. Delete the `dataset_first.py` shim and the deprecated aliases.

After step 2 the flag is meaningless even before you delete it: both branches
of the old `if` would run the same code. That is the definition of one
process.