# MatchDay v3 Demo Script Target length: 2 minutes 35 seconds. Never exceed 3 minutes. ## 0:00-0:10 - Make the value undeniable Show the completed first viewport while saying: > MatchDay remembered that I need seven hours before kickoff. It caught that the lowest-priced live option broke that rule, moved me to a verified earlier flight, and independently verified the repair. Keep the four-state rail and before/after diff visible. Do not begin with slides or architecture. The first viewport should also show the image-led trip artifact so the result reads as a real travel product immediately. ## 0:10-0:32 - Prove cross-session memory Open a fresh judge identity and show that the profile is empty—no demo fact is silently treated as user memory. Briefly add one typed preference to prove the profile is not limited to three fixed fields, then use the prepared two-trip learning state: the first verified choice is `1 of 2` and not yet reusable; the second independent root trip promotes it. Explain that a same-trip refinement cannot satisfy this threshold. Reload the page to show the learned profile and run survive the session boundary. ## 0:32-0:58 - Show the agent, not a canned planner Replay the request. Let the typed timeline show: 1. Qwen selects memory recall; 2. Qwen produces a validated trip intent; 3. SerpApi returns up to three normalized live packages, or the provider failure is labelled before the reproducible fixture fallback runs; 4. Open-Meteo, OpenStreetMap, and Unsplash report separate live/cached/fallback states for the exact trip; 5. deterministic code reports the arrival-buffer breach; 6. Qwen selects the safe candidate; 7. the 14-check independent contract passes, including route, price completeness, component-bound refund coverage, personalized categories, accessibility, entry, and context evidence. Resolve one flight and one hotel action. Show that an exact action preserves the component, dates, travelers, seller, and current price; an unresolved or failed lookup remains an honest date-bound search. Say clearly that the fixture fallback makes judging reproducible and is not presented as live availability. ## 0:58-1:22 - Show learning and user control Start a neutral later-session tour request and show that MatchDay reuses the independently learned package style, museum/history interests, and C$180 activity ceiling; open the result evidence to show all three memory keys were actually applied. Explain that explicit profile actions remain available, grounded candidates resolve to `ADD`, `UPDATE`, `DELETE`, or `NONE`, and repeated facts do not churn memory history. Then add an explicit conflicting instruction and show current intent wins. Briefly point to correction, deletion, learning-signal dismissal, and memory-disable controls. ## 1:22-1:47 - Show real refinement and safe sharing Briefly show the saved Dallas semifinal run to prove that Qwen can search upcoming official fixtures and select a non-demo host from bounded candidates. Then change the active request to "rank the closest hotel to the stadium." Show the selected package, objective, version change, map, one component detail, and match-centered itinerary. Open the return action and show either a seller-bound exact result or the exact route/date fallback; neither is presented as a booking. Create a safe share and open it in a private window. Point out that the public page contains itinerary/source evidence but no fan memory, user identity, Qwen trace, or tool history. ## 1:47-2:07 - Show production hardening Briefly show the verification panel and mention: - HMAC raw-query/body/path/user binding and replay defense; - cross-user memory/run isolation; - provider fallback and Qwen resume behavior; - owner-checked clarification continuation and duplicate-worker prevention; - serialized, database-constrained concurrent memory updates; - unknown-tool rejection and bounded-loop stall handoff; - atomic per-user/global run budgets and bounded model steps. Do not scroll through source code for more than a few seconds. ## 2:07-2:22 - Show architecture and Alibaba deployment Show `docs/architecture.png`. Explain that Vercel serves the Next.js UI and signed facade, while one approved Alibaba ECS host runs Caddy, FastAPI, private PostgreSQL, the server-side travel/context providers, and a six-hour worker that validates all 104 matches from FIFA before updating the cache. Qwen Cloud supplies intent and repair intelligence. Show minimal live `/healthz` readiness through both public origins, then use the credential-free verifier output as proof of the signed operator-only fixture, release, schema, and integration checks; do not expose the operator endpoint or signing material in the demo. ## 2:22-2:35 - Close on measurable impact Show the eval report: - memory required-key recall: 100%, versus 47.37% naive transcript and 0% without memory, for 52.63 points of structured lift; - cross-user leakage: 0%; - the 20-memory cross-city stress case selects eight relevant/critical facts; the observed context used 650 characters under the enforced 1,500-character stress budget; - local retrieval p95 is reported and gated below 250 ms; - memory lifecycle, objective ranking, multi-factor coverage, memory-sensitive ranking, official fixture discovery, and repaired verification: 100%. - accessibility enforcement: 100%, with unsupported trip-wide needs stopped before verification and live hotel-scoped evidence accepted only when it exactly matches; - entry-constraint enforcement: 100%, with self-declared country scope enforced and passport/visa review stopped without a legal eligibility claim. Close with: > MatchDay does not just remember preferences. It turns memory into a verified intervention before a fan loses the match.