# FlowTwin — Demo Guide Everything needed to run the demonstration, in order, with what to say and what to do if something goes wrong. --- ## Before you start ```bash ./run.sh # Windows: run.bat ``` Open **http://127.0.0.1:8000**. Check the top-right connection chip reads **live** once a run starts. Pre-flight checklist (two minutes): - [ ] `python scripts/build_venues.py` — venues and scenarios regenerate cleanly - [ ] `cd backend && python -m pytest -q` — the suite passes - [ ] `python scripts/ui_check.py` — drives the whole acceptance path in a real browser and saves screenshots to `shots/` - [ ] `benchmarks/BENCHMARKS.md` exists and is current - [ ] Optional: `python scripts/fetch_hf_model.py` — with network access, so the Perception panel shows a live model - [ ] Optional: `python scripts/record_fallback.py` — records both scenarios as replay insurance Keyboard shortcuts during the demo: | Key | Action | |---|---| | `Space` | play / pause | | `1`–`6` | speed 1× / 2× / 5× / 10× / 20× / 40× | | `S` | simulate strategies | | `Esc` | close the drawer or a modal | --- ## The 6-minute run ### 0:00 — The hook > "Formula 1 has spent decades learning to turn telemetry into strategy. But the > cars aren't the only thing moving on race day. Hundreds of thousands of people > move through gates, corridors and transport links that have a fixed capacity. > There is telemetry for the car. Where is the telemetry for the crowd?" Point at the header: **FlowTwin · Crowd Race Control**. Four scenarios in the switcher; the map is the product. ### 0:30 — The problem > "Existing crowd monitoring tells an operator where people *are*. The dangerous > question is where the flow is going to *fail* — and what to do about it before > it does." ### 1:00 — Simulation 1: prove the engine **Do:** with *F1 Circuit Stress Test* selected, press **Run simulation**. Set speed to **20×**. > "Forty thousand spectators, a fictional but realistically proportioned Grand > Prix venue. Four exits, six spectator zones. The chequered flag has just > fallen. Those are individual simulated people — each with their own walking > speed, destination and willingness to follow instructions." **Point out:** the venue emptying, the metrics strip filling, the flow arrows. At **T+04:00** a red banner flashes across the map: **Exit B throughput reduced by 50%**. > "That's a scripted infrastructure failure — and it's a real change to the > simulated network, not a caption. Half of Exit B's lanes are out of service." ### 2:00 — SEE and PREDICT Watch the East Concourse corridor turn yellow, then orange. The Alerts panel raises a card. > "FlowTwin isn't reacting to a threshold. The alert carries a *cause* — inflow > at capacity, the queue growing, walking speed collapsing — and, once the > projection crosses the critical line, a *lead time*." **Point out** the alert reading **"Projected critical in ~50 s"** and the Prediction panel bars stepping up across now / +30 / +60 / +90 / +120 s. > "That projection comes from a gradient-boosted model trained on the simulator's > own ground truth and validated on seeds it never saw. Click *Model accuracy* — > it beats the physics baseline by roughly 40 to 60 per cent depending on the > horizon, and if it hadn't, the system would refuse to load it." *(Optional: open **Model accuracy** for five seconds.)* ### 3:00 — SIMULATE **Do:** press **Simulate strategies** (or `S`). Takes a few seconds. > "This is the part that isn't a dashboard. FlowTwin has just cloned the current > crowd state eight times — byte identical, same random seed — applied a > different intervention to each copy, and simulated all of them forward five > minutes." The drawer opens. Walk the table left to right: > "Doing nothing: peak density around 3, the corridor critical for over two > minutes, a queue of nearly 5,000. Redirecting 40%: peak density down about a > quarter, critical time to zero, the queue down by a quarter — for essentially > no change in average journey time. > > These aren't rules of thumb. Every number in this table was measured from a > simulation that actually ran." Then the **Why this strategy?** column: > "Primary bottleneck, predicted critical time, the recommended action, and the > reason — expressed as measured changes against doing nothing. The margin over > the runner-up is there too. No language model is anywhere in this decision." ### 4:00 — ACT **Do:** press **Apply intervention**. > "The operator stays in control. FlowTwin recommends; a human decides." **Point out:** green dashed reroute paths appear on the map; the queue metric starts falling; the alert drops from critical to warning; the intervention is recorded in the event timeline with how many people accepted the instruction. > "Note that not everyone complies — that's modelled per person. About seventy > per cent of those instructed actually change route, which is what makes the > measured improvement believable." ### 4:45 — Simulation 2: prove it matters **Do:** click **02 · Barcelona 2022 Counterfactual**. > "Now the same intelligence against a real Formula 1 crowd-flow failure." **Read the left rail deliberately** — this is the credibility moment: > "The 2022 Spanish Grand Prix. 277,836 reported weekend attendance. Documented > severe road and public-transport congestion, long concession queues. Formula 1 > publicly told the promoter the situation was not acceptable. Those are facts, > each with a source. > > What we do *not* have is the original spectator telemetry — it isn't public. > So we did not recreate Barcelona. We reconstructed the documented conditions > with a transparent model, and every assumption is labelled as one: the mode > split, the corridor capacities, the departure curve, the walking speeds." Scroll to the **Evidence & assumptions** panel and let the two lists be seen. **Do:** run it at 20×. The rail interchange saturates and the north side backs up. > "The failure here isn't inside the circuit. It's the transport interface — and > that matches what was reported." **Do:** press **Simulate strategies**. > "The question is not *what happened*. It's: given the documented conditions, > what would an AI race engineer have recommended?" ### 5:30 — Results > "Across independent random seeds of the full simulation — baseline shortest > path, a static pre-event plan, and the full FlowTwin loop — peak density falls > by more than half, time spent in critical conditions goes to zero, and the > maximum queue falls by around 60 per cent, with average journey time slightly > *better*, not worse." Quote from `benchmarks/BENCHMARKS.md`. Every figure is generated by `scripts/run_benchmarks.py`. ### 5:45 — Close > "Most systems stop at detection. FlowTwin closes the loop: observe, predict, > simulate the alternatives, and recommend the one that measurably works. > > Don't wait for the bottleneck. Simulate the intervention before it happens." --- ## Optional beats **Perception (30 s).** Click **Perception** in the header. Shows the candidate chain, the active Hugging Face model and the observation schema both input modes share. Upload a crowd photo to get a live count. If no model is loaded it says exactly why — which is itself the point: it never invents a number. **What-if (30 s).** Change attendance or Exit B capacity in the left rail and press **Run simulation**. Every run is reproducible from its seed. **Reproducibility (15 s).** Point at the seed in the metrics strip. Same venue, same scenario, same seed, same run — the test suite asserts it. --- ## If something goes wrong **A simulation will not start.** Every scenario has a recorded run. Start it explicitly: ```bash curl -X POST localhost:8000/api/simulation/start \ -H 'content-type: application/json' \ -d '{"venue_id":"circuit_alpha","scenario_id":"circuit_alpha_post_race","use_recording":true}' ``` The dashboard behaves identically — same frames, same strategy comparison. Record them beforehand with `python scripts/record_fallback.py`. **The connection chip says `reconnecting`.** The stream reconnects on its own with backoff. The simulation keeps running; nothing is lost. **Strategy simulation is slow.** It is doing real work — eight full simulations over a five-minute horizon. Lower the horizon, or reduce attendance in the What-If panel before the run. **The browser tab was reloaded.** Just press **Run simulation** again. Orphaned sessions stop themselves and are reaped. **Perception says unavailable.** Expected without `torch`/`transformers` or network access. It is not on the critical path — say so and move on; the panel already explains it on screen. --- ## Questions you should expect **"Isn't this just shortest-path routing?"** No. Shortest path is the *baseline we measure against*. FlowTwin weights routes by live congestion, predicted state, capacity and risk — and then simulates several interventions before choosing one. The benchmark shows shortest path, static routing and FlowTwin side by side. **"Where does the data come from?"** Synthetic agents, because they give exact ground truth: we know precisely what happened under every intervention, which is what makes honest benchmarking possible. For real-world observation there's a Hugging Face crowd model feeding the same schema. For Barcelona, documented facts and labelled assumptions. **"Is this actually AI?"** Machine-learned crowd perception and a trained future-state predictor validated on held-out seeds. The *decision* comes from simulation and optimisation — which is deliberate. A numerical safety-adjacent decision should be measurable and reproducible, and an LLM is neither. **"Did you recreate Barcelona?"** No, and the interface says so. The original crowd telemetry isn't public. We reconstructed documented conditions and separate evidence from assumption on screen throughout. **"Can you guarantee this prevents a crush?"** No. It is decision support. It shows a recommendation, its cause, its lead time and its expected outcome; a trained operator decides. Real deployment would need venue calibration, sensor integration and operational validation. **"Why is average travel time barely different?"** Because that's the honest result over that window, and it's the point: the safety gain doesn't cost mobility. The optimizer weights travel time explicitly, and if a strategy bought density at the price of a much longer walk, the score would say so — some candidates in the table do exactly that.