Spaces:
Running
A newer version of the Gradio SDK is available: 6.26.0
One Pipeline β Removing the Legacy Path and the Mode Flag
Goal: there is exactly one way the app builds a demo:
research -> blueprint -> engine -> validation -> derived DDL -> Snowflake load
No data_generation_mode setting, no DEMOPREP_DATASET_FIRST env flag, no
"falling back to DDL-first" branch, no LegitData population path. A build
either succeeds with the story the user asked for, or fails with an error
that says exactly why. Nothing in between.
Why this matters beyond tidiness: today the fork is not just leftover β it is
live and defaulted to legacy. load_default_settings() sets
'data_generation_mode': 'legacy', so a fresh user never touches the new
path at all. And even with dataset-first enabled, build_dataset_first_demo
returning None triggers a silent fallback to LLM-DDL generation β a second
router hiding behind the first. Both are exactly the class of "quietly do
something else" behavior that produced the McKesson incident. The new
pipeline never returns None (it raises), so the fallback has nothing to fall
back from; it is dead code that can only cause harm.
1. chat_interface.py β deletions
1a. The mode check (delete the whole method)
def _dataset_first_enabled(self) -> bool:
"""Read dataset-first mode from app config/.env or saved settings."""
setting = str(self.settings.get("data_generation_mode", "")).strip().lower()
env_flag = os.getenv("DEMOPREP_DATASET_FIRST", "").strip().lower()
return setting in {...} or env_flag in {...}
Delete it. Also delete 'data_generation_mode': 'legacy' from the defaults
dict in load_default_settings(), and any Settings-tab UI bound to it.
1b. The schema-creation fork
Current shape:
self._dataset_first_bundle = None
if self._dataset_first_enabled():
...build = build_dataset_first_demo(...)
if build:
...success path...
return response, self.ddl_code
self.log_feedback("...no scenario generator matched... Falling back to DDL-first generation.")
# Build DDL generation prompt with geo context
geo_scope = ...
...entire legacy LLM-DDL stage: schema_prompt, retries, 'None' checks,
regex stripping of db prefixes, etc...
New shape β no condition, no fallback:
from demoprep_app.pipeline.build_demo import build_demo
row_count_guidance = int(self.settings.get("fact_table_size", 5000) or 5000)
company_name = self.demo_builder.extract_company_name()
build = build_demo(
company_name=company_name,
company_url=self.demo_builder.company_url,
use_case=self.demo_builder.use_case,
vertical=self.vertical,
function=self.function,
row_count_guidance=row_count_guidance,
research_context="\n\n".join(part for part in [
getattr(self.demo_builder, "combined_research_results", "") or "",
getattr(self, "generic_use_case_context", "") or "",
] if part),
# verbatim custom-tab text -> binding directives (McKesson fix)
user_request=getattr(self, "generic_use_case_context", "") or self.demo_builder.use_case,
llm_model=self.settings.get("model", DEFAULT_LLM_MODEL),
progress_callback=self.log_feedback,
prompt_logger=self._prompt_logger,
)
self._demo_bundle = build.dataset
self.demo_builder.schema_generation_results = build.ddl
self.ddl_code = build.ddl
for warning in build.warnings:
self.log_feedback(f"β οΈ {warning}")
build_demo raises on failure with the validation report in the message β
let the existing per-stage exception handling surface it. Do NOT wrap it in
a try/except that falls through to anything else.
Everything from # Build DDL generation prompt with geo context down through
the DDL retry loop, the 'CREATE TABLE' not in ddl_result checks, the
ddl == "None" check, and the db-prefix regex: delete. That is the
legacy stage. (Geo scope belongs in the blueprint author prompt if you still
want it β one line in PROMPT_TEMPLATE β not in a parallel DDL stage.)
1c. The population fork
Current shape:
if self._dataset_first_bundle:
...populate_dataset_bundle(...)
return
success, message, results = populate_demo_data( # LegitData path
ddl_content=ddl, ..., size=size, ...
)
New shape β the bundle always exists:
from demoprep_app.integrations.snowflake import populate_dataset_bundle
results = populate_dataset_bundle(deployer.connection, schema_name, self._demo_bundle)
Delete the populate_demo_data(...) call and, once nothing else imports it,
the LegitData population module itself, plus the 45-minute POP_TIMEOUT
machinery sized for it (the bundle load is a bounded INSERT of in-memory
rows β it does not need a repair-loop timeout).
1d. Cosmetic followers
generation_mode: "dataset_first" if self._dataset_first_bundle else "legacy"in_deploy_metaβ just"generation_mode": "blueprint".- Rename
self._dataset_first_bundleβself._demo_bundle. - UI copy "Dataset-First Schema Creation Complete!" β "Schema Creation
Complete!" β there is no other kind. Consider surfacing
build.blueprint.insightsheadlines andbuild.validation.summary()in that message: it tells the SE what stories are planted before they deploy.
2. Settings / env cleanup
- Remove
data_generation_modefromSETTINGS_SCHEMA, Supabase-saved settings, and any docs mentioningDEMOPREP_DATASET_FIRST. - Architecture doc: delete the "when dataset-first mode is enabled" section β the flow it describes is now the only flow.
3. Module deletions (after the above, nothing imports them)
- Legacy DDL prompt stage inside chat_interface.py (inline, per 1b)
- LegitData population path (
populate_demo_dataand its module) demoprep_app/scenario/selector.py,extractor.py,families.pydemoprep_app/dataset/generators/(all)- Finally:
demoprep_app/pipeline/dataset_first.py(the compat shim), once imports point atdemoprep_app.pipeline.build_demo.
4. What "one process" does NOT remove
- "Use existing model" flows (skip generation entirely) β orthogonal, keep.
- The approval checkpoint ("Do you approve this DDL?") β keep; it now shows derived DDL + planted-insight headlines instead of LLM-guessed DDL.
- The e2e quality harness β point it at
build.validationand assertpassed.
5. Order of operations (safe cutover)
- Land the new package modules (already tested standalone).
- Apply 1b + 1c so the app runs only the new pipeline.
- Run one real demo end to end (Snowflake + ThoughtSpot).
- Delete legacy modules and the mode setting (steps 2-3 above).
- Delete the
dataset_first.pyshim and the deprecated aliases.
After step 2 the flag is meaningless even before you delete it: both branches
of the old if would run the same code. That is the definition of one
process.