File size: 10,645 Bytes
e7a9f02
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
# FlowTwin — Demo Guide

Everything needed to run the demonstration, in order, with what to say and what
to do if something goes wrong.

---

## Before you start

```bash
./run.sh                      # Windows: run.bat
```

Open **http://127.0.0.1:8000**. Check the top-right connection chip reads
**live** once a run starts.

Pre-flight checklist (two minutes):

- [ ] `python scripts/build_venues.py` — venues and scenarios regenerate cleanly
- [ ] `cd backend && python -m pytest -q` — the suite passes
- [ ] `python scripts/ui_check.py` — drives the whole acceptance path in a real browser and saves screenshots to `shots/`
- [ ] `benchmarks/BENCHMARKS.md` exists and is current
- [ ] Optional: `python scripts/fetch_hf_model.py` — with network access, so the Perception panel shows a live model
- [ ] Optional: `python scripts/record_fallback.py` — records both scenarios as replay insurance

Keyboard shortcuts during the demo:

| Key | Action |
|---|---|
| `Space` | play / pause |
| `1``6` | speed 1× / 2× / 5× / 10× / 20× / 40× |
| `S` | simulate strategies |
| `Esc` | close the drawer or a modal |

---

## The 6-minute run

### 0:00 — The hook

> "Formula 1 has spent decades learning to turn telemetry into strategy. But the
> cars aren't the only thing moving on race day. Hundreds of thousands of people
> move through gates, corridors and transport links that have a fixed capacity.
> There is telemetry for the car. Where is the telemetry for the crowd?"

Point at the header: **FlowTwin · Crowd Race Control**. Four scenarios in the
switcher; the map is the product.

### 0:30 — The problem

> "Existing crowd monitoring tells an operator where people *are*. The dangerous
> question is where the flow is going to *fail* — and what to do about it before
> it does."

### 1:00 — Simulation 1: prove the engine

**Do:** with *F1 Circuit Stress Test* selected, press **Run simulation**. Set
speed to **20×**.

> "Forty thousand spectators, a fictional but realistically proportioned Grand
> Prix venue. Four exits, six spectator zones. The chequered flag has just
> fallen. Those are individual simulated people — each with their own walking
> speed, destination and willingness to follow instructions."

**Point out:** the venue emptying, the metrics strip filling, the flow arrows.

At **T+04:00** a red banner flashes across the map: **Exit B throughput reduced
by 50%**.

> "That's a scripted infrastructure failure — and it's a real change to the
> simulated network, not a caption. Half of Exit B's lanes are out of service."

### 2:00 — SEE and PREDICT

Watch the East Concourse corridor turn yellow, then orange. The Alerts panel
raises a card.

> "FlowTwin isn't reacting to a threshold. The alert carries a *cause* — inflow
> at capacity, the queue growing, walking speed collapsing — and, once the
> projection crosses the critical line, a *lead time*."

**Point out** the alert reading **"Projected critical in ~50 s"** and the
Prediction panel bars stepping up across now / +30 / +60 / +90 / +120 s.

> "That projection comes from a gradient-boosted model trained on the simulator's
> own ground truth and validated on seeds it never saw. Click *Model accuracy*
> it beats the physics baseline by roughly 40 to 60 per cent depending on the
> horizon, and if it hadn't, the system would refuse to load it."

*(Optional: open **Model accuracy** for five seconds.)*

### 3:00 — SIMULATE

**Do:** press **Simulate strategies** (or `S`). Takes a few seconds.

> "This is the part that isn't a dashboard. FlowTwin has just cloned the current
> crowd state eight times — byte identical, same random seed — applied a
> different intervention to each copy, and simulated all of them forward five
> minutes."

The drawer opens. Walk the table left to right:

> "Doing nothing: peak density around 3, the corridor critical for over two
> minutes, a queue of nearly 5,000. Redirecting 40%: peak density down about a
> quarter, critical time to zero, the queue down by a quarter — for essentially
> no change in average journey time.
>
> These aren't rules of thumb. Every number in this table was measured from a
> simulation that actually ran."

Then the **Why this strategy?** column:

> "Primary bottleneck, predicted critical time, the recommended action, and the
> reason — expressed as measured changes against doing nothing. The margin over
> the runner-up is there too. No language model is anywhere in this decision."

### 4:00 — ACT

**Do:** press **Apply intervention**.

> "The operator stays in control. FlowTwin recommends; a human decides."

**Point out:** green dashed reroute paths appear on the map; the queue metric
starts falling; the alert drops from critical to warning; the intervention is
recorded in the event timeline with how many people accepted the instruction.

> "Note that not everyone complies — that's modelled per person. About seventy
> per cent of those instructed actually change route, which is what makes the
> measured improvement believable."

### 4:45 — Simulation 2: prove it matters

**Do:** click **02 · Barcelona 2022 Counterfactual**.

> "Now the same intelligence against a real Formula 1 crowd-flow failure."

**Read the left rail deliberately** — this is the credibility moment:

> "The 2022 Spanish Grand Prix. 277,836 reported weekend attendance. Documented
> severe road and public-transport congestion, long concession queues. Formula 1
> publicly told the promoter the situation was not acceptable. Those are facts,
> each with a source.
>
> What we do *not* have is the original spectator telemetry — it isn't public.
> So we did not recreate Barcelona. We reconstructed the documented conditions
> with a transparent model, and every assumption is labelled as one: the mode
> split, the corridor capacities, the departure curve, the walking speeds."

Scroll to the **Evidence & assumptions** panel and let the two lists be seen.

**Do:** run it at 20×. The rail interchange saturates and the north side backs
up.

> "The failure here isn't inside the circuit. It's the transport interface — and
> that matches what was reported."

**Do:** press **Simulate strategies**.

> "The question is not *what happened*. It's: given the documented conditions,
> what would an AI race engineer have recommended?"

### 5:30 — Results

> "Across independent random seeds of the full simulation — baseline shortest
> path, a static pre-event plan, and the full FlowTwin loop — peak density falls
> by more than half, time spent in critical conditions goes to zero, and the
> maximum queue falls by around 60 per cent, with average journey time slightly
> *better*, not worse."

Quote from `benchmarks/BENCHMARKS.md`. Every figure is generated by
`scripts/run_benchmarks.py`.

### 5:45 — Close

> "Most systems stop at detection. FlowTwin closes the loop: observe, predict,
> simulate the alternatives, and recommend the one that measurably works.
>
> Don't wait for the bottleneck. Simulate the intervention before it happens."

---

## Optional beats

**Perception (30 s).** Click **Perception** in the header. Shows the candidate
chain, the active Hugging Face model and the observation schema both input modes
share. Upload a crowd photo to get a live count. If no model is loaded it says
exactly why — which is itself the point: it never invents a number.

**What-if (30 s).** Change attendance or Exit B capacity in the left rail and
press **Run simulation**. Every run is reproducible from its seed.

**Reproducibility (15 s).** Point at the seed in the metrics strip. Same venue,
same scenario, same seed, same run — the test suite asserts it.

---

## If something goes wrong

**A simulation will not start.** Every scenario has a recorded run. Start it
explicitly:

```bash
curl -X POST localhost:8000/api/simulation/start \
  -H 'content-type: application/json' \
  -d '{"venue_id":"circuit_alpha","scenario_id":"circuit_alpha_post_race","use_recording":true}'
```

The dashboard behaves identically — same frames, same strategy comparison. Record
them beforehand with `python scripts/record_fallback.py`.

**The connection chip says `reconnecting`.** The stream reconnects on its own with
backoff. The simulation keeps running; nothing is lost.

**Strategy simulation is slow.** It is doing real work — eight full simulations
over a five-minute horizon. Lower the horizon, or reduce attendance in the
What-If panel before the run.

**The browser tab was reloaded.** Just press **Run simulation** again. Orphaned
sessions stop themselves and are reaped.

**Perception says unavailable.** Expected without `torch`/`transformers` or
network access. It is not on the critical path — say so and move on; the panel
already explains it on screen.

---

## Questions you should expect

**"Isn't this just shortest-path routing?"**
No. Shortest path is the *baseline we measure against*. FlowTwin weights routes
by live congestion, predicted state, capacity and risk — and then simulates
several interventions before choosing one. The benchmark shows shortest path,
static routing and FlowTwin side by side.

**"Where does the data come from?"**
Synthetic agents, because they give exact ground truth: we know precisely what
happened under every intervention, which is what makes honest benchmarking
possible. For real-world observation there's a Hugging Face crowd model feeding
the same schema. For Barcelona, documented facts and labelled assumptions.

**"Is this actually AI?"**
Machine-learned crowd perception and a trained future-state predictor validated
on held-out seeds. The *decision* comes from simulation and optimisation — which
is deliberate. A numerical safety-adjacent decision should be measurable and
reproducible, and an LLM is neither.

**"Did you recreate Barcelona?"**
No, and the interface says so. The original crowd telemetry isn't public. We
reconstructed documented conditions and separate evidence from assumption on
screen throughout.

**"Can you guarantee this prevents a crush?"**
No. It is decision support. It shows a recommendation, its cause, its lead time
and its expected outcome; a trained operator decides. Real deployment would need
venue calibration, sensor integration and operational validation.

**"Why is average travel time barely different?"**
Because that's the honest result over that window, and it's the point: the safety
gain doesn't cost mobility. The optimizer weights travel time explicitly, and if
a strategy bought density at the price of a much longer walk, the score would say
so — some candidates in the table do exactly that.