Spaces:
Runtime error
Runtime error
| # FlowTwin — Project Masterfile | |
| *Everything about this project in one place: what it is, why it exists, how every | |
| part works, what was measured, how to pitch it, and how to defend it.* | |
| Written to be read cold. If you have never seen this project before, start at | |
| §1 and keep going — nothing later assumes anything earlier than what you have | |
| already read. | |
| --- | |
| ## Table of contents | |
| **Part I — Understanding the project** | |
| 1. [The one-paragraph version](#1-the-one-paragraph-version) | |
| 2. [The problem, properly explained](#2-the-problem-properly-explained) | |
| 3. [Why existing tools do not solve it](#3-why-existing-tools-do-not-solve-it) | |
| 4. [The core idea: the decision loop](#4-the-core-idea-the-decision-loop) | |
| 5. [A worked example, end to end](#5-a-worked-example-end-to-end) | |
| **Part II — How it actually works** | |
| 6. [The venue model](#6-the-venue-model) | |
| 7. [The simulation engine](#7-the-simulation-engine) | |
| 8. [The Crowd State Engine](#8-the-crowd-state-engine) | |
| 9. [Prediction](#9-prediction) | |
| 10. [The Strategy Engine](#10-the-strategy-engine) | |
| 11. [Counterfactual simulation](#11-counterfactual-simulation) | |
| 12. [Multi-objective optimisation and the decisiveness verdict](#12-multi-objective-optimisation-and-the-decisiveness-verdict) | |
| 13. [Dynamic routing](#13-dynamic-routing) | |
| 14. [Perception — the Hugging Face path](#14-perception--the-hugging-face-path) | |
| **Part III — The system as software** | |
| 15. [Architecture and module map](#15-architecture-and-module-map) | |
| 16. [Data flow and real-time transport](#16-data-flow-and-real-time-transport) | |
| 17. [The frontend](#17-the-frontend) | |
| 18. [Reproducibility and determinism](#18-reproducibility-and-determinism) | |
| 19. [The three venues](#19-the-three-venues) | |
| 20. [Testing and verification](#20-testing-and-verification) | |
| **Part IV — Evidence** | |
| 21. [Measured results](#21-measured-results) | |
| 22. [Every defect found and fixed](#22-every-defect-found-and-fixed) | |
| 23. [What is deliberately not built](#23-what-is-deliberately-not-built) | |
| **Part V — The hackathon** | |
| 24. [Mapping to the evaluation criteria](#24-mapping-to-the-evaluation-criteria) | |
| 25. [The pitch](#25-the-pitch) | |
| 26. [The demo, minute by minute](#26-the-demo-minute-by-minute) | |
| 27. [Q&A defence](#27-qa-defence) | |
| 28. [Failure drills](#28-failure-drills) | |
| --- | |
| --- | |
| # Part I — Understanding the project | |
| ## 1. The one-paragraph version | |
| FlowTwin is a **digital twin of a crowd**. You give it a venue — where the gates, | |
| walkways, concessions, exits and transport links are, and how much each can | |
| handle — plus how many people are coming and when. It then simulates tens of | |
| thousands of individual people walking through that venue, second by second. As | |
| it runs, it continuously asks three questions: *where is flow about to break | |
| down?*, *what could an operator do about it?*, and *which of those options | |
| actually works?* To answer the third question it does something unusual: it takes | |
| a perfect copy of the crowd's current state, applies each candidate action to its | |
| own copy, runs each copy forward four minutes, and **measures** what happened. | |
| Then it recommends the option that measured best, and shows you the arithmetic. | |
| If no option measurably beats doing nothing, it says so instead of inventing a | |
| recommendation. | |
| That last sentence is the project in miniature. Most systems in this space are | |
| dashboards that tell you what is happening. FlowTwin tells you **what to do**, | |
| and it earns the right to say it by simulating the alternatives rather than | |
| applying a rule of thumb. | |
| --- | |
| ## 2. The problem, properly explained | |
| ### 2.1 Crowd disasters are not headcount problems | |
| The intuitive model of crowd danger is "too many people in the building". That | |
| model is wrong, and the wrongness matters. | |
| A venue can sell out completely, admit exactly the number of people it is | |
| licensed for, and still kill someone — because danger is not a property of the | |
| total, it is a property of the **local density and the local flow**. Five people | |
| per square metre in one corridor is dangerous whether the rest of the venue is | |
| empty or full. Crowd crush injuries happen at pinch points: a gate that closed, a | |
| staircase that narrowed, two streams of people trying to cross. | |
| So the quantity that matters is not *how many people are here* but *how many | |
| people are in this twelve metres of corridor, how fast are they moving, and is | |
| that number rising*. | |
| ### 2.2 Flow failures are non-local and delayed | |
| Here is what makes it genuinely hard. Suppose an exit loses half its capacity. | |
| The people at that exit notice immediately. But the *consequence* is not local: | |
| - The queue at that exit grows backwards up the corridor. | |
| - When it reaches the concourse behind it, that concourse starts filling. | |
| - People arriving at the concourse from an entirely different direction — | |
| who have nothing to do with that exit — now find their route blocked. | |
| - The pressure propagates outward, several minutes after the original event, in | |
| places nobody was watching. | |
| This is the same mathematics as a traffic jam. The shockwave travels *backwards* | |
| through the crowd, slower than the people are walking, and it arrives somewhere | |
| unexpected several minutes later. | |
| Two consequences follow, and both shaped this project: | |
| 1. **You cannot reason about it locally.** A camera on the failing exit tells you | |
| about the failing exit. It does not tell you that the west concourse will be | |
| dangerous in six minutes. | |
| 2. **By the time you can see it, it may be too late to fix by rerouting.** Once a | |
| queue of four thousand people exists, it drains at the gate's service rate no | |
| matter where you send new arrivals. The people you would need to move are | |
| already in the queue and physically cannot move. | |
| FlowTwin models both of these explicitly, and — importantly — it *tells you* when | |
| you have hit the second one, rather than pretending it can still help. | |
| ### 2.3 The operator's actual problem | |
| Put yourself in the control room. You have: | |
| - Cameras and counters, so you know roughly where people are. | |
| - A handful of levers: reroute a percentage of people, hold back departures from | |
| a section, open contingency lanes, unlock an emergency gate, change where a | |
| shuttle picks up. | |
| - Minutes, not hours. | |
| - No way to test a decision before making it. | |
| That last one is the gap. Every lever has a cost and a side effect. Rerouting | |
| 40% of a stand relieves one corridor and loads another. Holding back departures | |
| keeps people safe but makes their evening longer, and if you hold too long the | |
| release is worse than the original problem. Opening an emergency gate means | |
| staffing it, breaking a perimeter, and explaining it afterwards. | |
| **An operator has to choose between options whose consequences are separated | |
| from the decision by five minutes and half a venue.** That is exactly the kind of | |
| decision a simulation should make for you, and nobody does it. | |
| ### 2.4 The specific case this project is built around | |
| The 2022 Spanish Grand Prix at the Circuit de Barcelona-Catalunya reported a | |
| weekend attendance of 277,836, with over 120,000 on race day. Contemporary | |
| reporting described severe road and public-transport congestion leaving the | |
| circuit, heavy pressure on the Montmeló rail infrastructure, long concession | |
| queues and water shortages. Formula 1 publicly told the promoter the fan | |
| experience was not acceptable. | |
| Nobody was hurt. That is the point: this is the *ordinary* failure mode, the one | |
| that happens dozens of times a year at venues that are competently run, and the | |
| one that becomes a disaster when the geometry is slightly worse or the crowd is | |
| slightly bigger. | |
| The same shape of failure covers the applications the problem statement names — | |
| railway station design, IPL match egress, airport terminals, Kumbh-scale | |
| gatherings. It is one problem, and it is not a motorsport problem. | |
| --- | |
| ## 3. Why existing tools do not solve it | |
| There are three categories of existing tool, and each stops short in a different | |
| place. | |
| **Crowd monitoring / people counting.** Cameras plus a counting model, feeding a | |
| dashboard with occupancy numbers and threshold alarms. This tells an operator | |
| *where people are*. It is reactive by construction: the alarm fires when the | |
| density is already high, which is after the point at which rerouting could have | |
| helped. It also has no notion of *why*, so it cannot suggest an action. | |
| **Offline crowd simulation.** Professional pedestrian modelling packages are | |
| excellent, and they are used at design time: you model the venue, run scenarios, | |
| and change the architecture or the plan. They are not real-time decision tools — | |
| a run takes minutes to hours, the model is not connected to live conditions, and | |
| the output is a report rather than an instruction. | |
| **Traffic-style routing.** Shortest-path or capacity-aware assignment can tell | |
| people where to go. But a pre-computed plan is blind to what actually happens on | |
| the day, and a purely reactive router chases congestion around the venue, | |
| producing oscillation: send people east, the east fills, send them west, the west | |
| fills. | |
| FlowTwin sits in the hole between these three. It is a **real-time simulation | |
| that is fast enough to run its own hypotheticals while an operator waits**. The | |
| architectural decision that makes this possible is described in §7.1, and it is | |
| the single most important engineering choice in the project. | |
| --- | |
| ## 4. The core idea: the decision loop | |
| ``` | |
| ┌─────────────────────────────────────────────────────────┐ | |
| │ │ | |
| ▼ │ | |
| ┌──────┐ ┌─────────┐ ┌──────────┐ ┌──────┐ │ | |
| │ SEE │ ───► │ PREDICT │ ───► │ SIMULATE │ ───► │ ACT │ ───┘ | |
| └──────┘ └─────────┘ └──────────┘ └──────┘ | |
| where are where will it what would apply the | |
| people, and break down, each option one that | |
| how fast are and when? actually do? measured best | |
| they moving? | |
| ``` | |
| **SEE.** Turn raw positions into the quantities that predict failure: density per | |
| short segment of corridor, walking speed against free speed, inflow and outflow | |
| per minute, queue length, how fast density is *changing*, and whether two streams | |
| are fighting for the same floor. | |
| **PREDICT.** Project each of those forward 30, 60, 90 and 120 seconds, and | |
| convert that into the only number an operator can act on: **how long until this | |
| corridor is critical**. | |
| **SIMULATE.** Generate the candidate actions that this venue's topology actually | |
| permits, then clone the entire crowd state once per candidate, apply the | |
| candidate to its clone, and run each clone forward four simulated minutes. | |
| **ACT.** Score the outcomes on a weighted objective, recommend the best — or | |
| refuse to recommend if nothing beat doing nothing — and show the arithmetic. When | |
| the operator applies it, the intervention enters the live simulation through the | |
| exact same code path that was measured, and the loop starts again. | |
| The loop is what makes this a decision-support system rather than a dashboard. | |
| Each stage exists because the stage after it needs something the stage before | |
| could not provide. | |
| --- | |
| ## 5. A worked example, end to end | |
| Concrete, from the flagship scenario, with real numbers from a real seeded run. | |
| **T+00:15.** The chequered flag. 40,000 spectators begin leaving six seating | |
| areas on an eighteen-minute departure curve. Everyone routes by shortest path | |
| towards one of four destinations: the rail interchange, the coach interchange, | |
| or one of two car parks. | |
| **T+04:00.** A scripted infrastructure failure fires: **Exit B loses half its | |
| throughput**, dropping from 760 people/minute to 380. This is a real change to | |
| the simulated network — the exit's service budget is halved — not a label on a | |
| map. | |
| **T+05:30.** *SEE.* The corridor feeding Exit B (`X_E_EXITB`, 114 m long, 11 m | |
| wide) is now taking more people per minute than it can pass. Measured: inflow 556 | |
| p/min, outflow 380 p/min. Density is rising at 0.14 p/m² per minute. Walking | |
| speed has fallen to 0.13 m/s against a free speed of 1.34. A queue is forming. | |
| **T+05:30.** *PREDICT.* The gradient-boosted model, fed seventeen features from | |
| the Crowd State Engine, projects density at +30/60/90/120 s. Crossing the venue's | |
| critical threshold of 2.8 p/m² happens inside the horizon, so the alert reads | |
| **"critical in 96 seconds"** — and it explains itself: *density rising, velocity | |
| collapsed, queue growing, downstream service constrained*. | |
| **T+07:30.** *SIMULATE.* The operator presses **Simulate strategies**. The engine | |
| inspects the topology around the bottleneck and generates eight candidates, | |
| including: do nothing; redirect 20/30/40% of the affected flow; stagger the | |
| release from the three stands feeding it; open contingency lanes at another exit | |
| and divert 30%; unlock the north-east emergency gate and divert 35%; move 30% of | |
| coach demand to the south apron; and a combined redirect-plus-stagger. | |
| Eight complete copies of the crowd — every agent's position, route, destination, | |
| compliance and the random number generator's internal state — are made. Each | |
| candidate is applied to its own copy. Each copy runs forward 240 simulated | |
| seconds. About nine seconds of wall-clock later, eight measured futures exist. | |
| **T+07:31.** *ACT.* Scored against the do-nothing arm on nine weighted terms. | |
| **Redirect 40%** wins by 17.1%. The panel says why, in measured deltas: peak | |
| density 2.19 → 1.58 (−28%), queue at end of window 1,636 → 1,245 (−24%), critical | |
| duration to zero, average journey time essentially unchanged, 834 people | |
| rerouted. The verdict reads **Decisive**. | |
| **T+07:45.** The operator applies it. 1,700 people are instructed; per-person | |
| compliance means roughly 70% actually change route. Green rerouting paths animate | |
| on the map. Over the next three minutes the queue metric falls and the alert | |
| drops from critical to warning. | |
| **And the counter-example, which is the more interesting demo.** Do nothing until | |
| **T+15:00** and press the button then. All eight candidates now return an | |
| *identical* peak density of 3.31 p/m². The engine does not pick a winner. It | |
| returns: | |
| > **Not decisive.** Every candidate landed within 0.0% of doing nothing. | |
| > `E CONCOURSE → EXIT B` is already discharging at its service limit (380 | |
| > people/min) with 3,275 people held, so it needs about 9 minutes to clear on | |
| > throughput alone. Rerouting only reaches people who have not yet committed to | |
| > this asset, and there are too few of them left for any routing change to | |
| > register. The remaining levers are capacity and staffing, not routing. | |
| Every number in that paragraph is read from the measured state. That is the | |
| system telling you the decision window closed — which is more useful, and far | |
| more credible, than a confident recommendation that would not have worked. | |
| --- | |
| --- | |
| # Part II — How it actually works | |
| ## 6. The venue model | |
| ### 6.1 A venue is a graph | |
| `backend/flowtwin/venue/models.py` | |
| A venue is a **directed, weighted graph**. Nodes are places a person can be; | |
| edges are the walkable links between them. | |
| **Node types**, and what each means to the engine: | |
| | Type | Role | | |
| |---|---| | |
| | `gate` | Entry point with a service rate in people/minute. An origin in arrival scenarios. | | |
| | `grandstand`, `general_admission` | Seating/standing areas. Origins; a route may *end* at one but never pass *through* one. | | |
| | `platform` | Railway platform. Same semantics as a grandstand — you leave from it, you do not walk across it. | | |
| | `concourse`, `junction` | Circulation space. Optionally rate-limited (a foot-over-bridge is a junction with a service rate set by stair width). | | |
| | `concession` | A dwell point. People passing through stop here for a while. | | |
| | `exit` | A perimeter throughput constraint. **Deliberately not a destination** — see §6.3. | | |
| | `emergency_exit` | A route that physically exists but is locked. **Absent from routing until opened** — see §6.4. | | |
| | `transport`, `parking` | Destinations. These absorb people, at a rate. | | |
| **Edges** carry `length_m`, `width_m` and `capacity_ppm` (people per minute that | |
| may *enter*). Capacity follows Fruin-style pedestrian flow: about 70 people per | |
| minute per metre of effective width in one direction. A bidirectional venue edge | |
| compiles into two directed edges that share the same physical floor, which is how | |
| opposing-flow conflict is measured. | |
| Edge lengths are **derived from node geometry** by `scripts/build_venues.py` | |
| rather than hand-written, so the map you see and the physics that runs can never | |
| drift apart. | |
| ### 6.2 Compilation and cells | |
| `CompiledVenue` turns the pydantic model into flat numpy arrays indexed by node | |
| or directed-edge index, so the simulation's inner loop never touches a Python | |
| object. | |
| Then every edge is split into **cells of about 12 metres**. Density and walking | |
| speed are evaluated per cell, not per edge. | |
| This is not a detail. It is the difference between a model that works and one | |
| that does not: | |
| > With edge-average density, a queue at a gate slows down *everyone* on that | |
| > corridor — including a person 200 metres back with completely clear space in | |
| > front of them. Measured effect when this was wrong: network throughput | |
| > collapsed to roughly **one tenth** of its correct value. | |
| Cells on a two-way corridor are mirrored to their opposite-direction twin | |
| (`cell_pair`), so two people walking towards each other in the same twelve metres | |
| are counted as sharing that floor. | |
| ### 6.3 The decision that an exit is not a destination | |
| A perimeter exit is modelled as a **throughput constraint on the way to somewhere | |
| else** — a station, a car park, a coach apron — not as a place journeys end. | |
| If an exit were a sink, everyone reaching it would vanish, and the queue *behind* | |
| it would never form. That queue is the single most important phenomenon this | |
| project exists to predict. Modelling exits as sinks would have made the demo | |
| easier and the model useless. | |
| ### 6.4 The decision that a locked gate is absent, not expensive | |
| An emergency exit is not modelled as an available-but-costly route. It is | |
| **excluded from every routing table for every policy and every destination**. | |
| The reason is precise. If a locked gate were merely expensive, the optimiser | |
| would quietly have access to capacity that nobody has unlocked; under enough | |
| congestion the crowd would start using it on its own, and the recommendation | |
| *"open the north gate"* would never appear, because the crowd would already be | |
| going there. Modelling it as absent makes opening it a real decision with a real | |
| consequence — and it makes `open_emergency_exit` the only candidate in the whole | |
| strategy set that **adds** network capacity rather than redistributing capacity | |
| already in service. | |
| ### 6.5 Concessions as dwell points | |
| A concession node carries `dwell_s` (mean stop time) and `dwell_share` (the | |
| fraction of passers-by who stop). A person who stops: | |
| - still occupies the floor they are standing on, and counts in the queue extent; | |
| - does **not** consume the downstream node's service budget, because they are not | |
| trying to go anywhere. | |
| That is what makes a food court a crowd feature rather than a label. It also | |
| requires the concession to be **on** a route — a dead-end spur is never on | |
| anybody's path, so nobody ever visits it. Both the fan zone at Circuit Alpha and | |
| the food court at Sangam Junction sit on the main circulation route, with a | |
| longer bypass available, which is what gives the strategy engine something to | |
| reroute people *onto*. | |
| The randomness lives in the agent population, sampled once at creation, not in a | |
| live random stream. That is deliberate: it means a counterfactual branch | |
| reproduces the same dwell decisions exactly, so two branches of one state stay | |
| byte-identical. | |
| --- | |
| ## 7. The simulation engine | |
| `backend/flowtwin/simulation/engine.py` | |
| ### 7.1 The critical architectural choice: mesoscopic, not microscopic | |
| A microscopic pedestrian model (social forces, agents in free 2-D space) is more | |
| physically detailed and completely unusable here: it is far too slow to run eight | |
| alternative futures while an operator waits. | |
| FlowTwin is **mesoscopic**. Agents are individuals — each has a personal walking | |
| speed, an origin, a destination, a route, a compliance probability and a position | |
| — but they move **along graph edges**, not across open floor. Agent state is | |
| stored as a **structure of arrays** (numpy), so a step is a handful of vectorised | |
| operations over the whole population rather than a loop over 40,000 objects. | |
| Measured: **2–4 ms per simulated second at 40,000 agents.** | |
| That number is the enabling fact for the entire project. Because a step is | |
| milliseconds, four minutes of simulation is about a second, and eight | |
| counterfactual futures are about nine seconds — short enough that an operator | |
| will actually press the button. Every other capability in this document is | |
| downstream of that choice. | |
| ### 7.2 The walking model | |
| Speed as a function of density uses **Weidmann's (1993) exponential fundamental | |
| diagram**, the standard empirical pedestrian relation: | |
| ``` | |
| v(ρ) = v_free · (1 − exp(−γ · (1/ρ − 1/ρ_jam))) | |
| ``` | |
| with `v_free = 1.34 m/s`, `γ = 1.913`, `ρ_jam = 5.4 p/m²`. Each agent has a | |
| personal multiplier drawn from a clipped normal (σ = 0.16), so a crowd contains | |
| fast and slow walkers. | |
| This reproduces the two behaviours everything else depends on: unimpeded walking | |
| at low density, and speed collapse as density approaches jam. | |
| ### 7.3 The step, in order | |
| Each simulated second: | |
| 1. **Timeline events** fire (capacity changes, phase transitions). | |
| 2. **Cell density and speed** are computed, including the mirrored opposite | |
| direction. | |
| 3. **Queue extent** is derived (§7.4). | |
| 4. **Agents advance** at their cell's speed × personal factor. A walker cannot | |
| step into a cell that is already at 90% of jam density, and stops when it | |
| reaches the back of a standing queue. | |
| 5. **Transition candidates** are gathered: everyone released and waiting, plus | |
| everyone standing at the head of an edge who is not currently dwelling. | |
| 6. **Node service budget** admits people first-come-first-served by how long they | |
| have been queueing. | |
| 7. **Edge admission** is limited by three separate constraints (§7.5). | |
| 8. **Moves and absorptions** apply. | |
| 9. **Measurement** updates the Crowd State Engine. | |
| 10. **Routing tables** refresh on a 5-second cadence. | |
| ### 7.4 Queue extent — a queue is a length, not a point | |
| A queue occupies corridor. If you measure it only at the stop line, the standing | |
| queue has zero physical extent, and the model then makes everyone behind it *walk | |
| through* a near-jammed corridor at a few centimetres per second to reach the back | |
| of it. | |
| Measured consequence when this was wrong: a gate rated at 500 people/minute | |
| discharged at **under 200**. | |
| The fix: queue extent is derived from everyone who has actually stopped — | |
| `queue_len = queued_count / (queue_pack_density × width)` — with a packing | |
| density of 4.6 p/m², lower than jam because a queue that has stopped moving is | |
| not yet a crush. Walkers then join the *back* of the queue where the back | |
| actually is. | |
| ### 7.5 Three admission constraints, and why each is needed | |
| An edge accepts people this second up to the minimum of: | |
| **(a) Nominal capacity.** `capacity_ppm × dt`, with fractional carry so a 70/min | |
| link does not admit zero people every second and then seventy at once. | |
| **(b) The backward-wave receiving function.** As a link fills, the rate at which | |
| it can accept anyone new falls towards zero. Congestion propagates *backwards* at | |
| `backward_wave_mps = 0.36 m/s`: | |
| ``` | |
| receiving_ppm = 0.36 × 60 × free_space / length | |
| ``` | |
| This is the cell-transmission idea from traffic flow. Without it a corridor | |
| silently absorbs an impossible crowd instead of pushing congestion upstream — and | |
| "congestion spills back" is the entire non-local behaviour described in §2.2. | |
| **(c) Entry-cell headroom.** People enter a corridor **at its mouth**, and the | |
| mouth is one cell wide. A 400 m corridor with room for 2,000 people cannot take | |
| 2,000 people this second, because they would all have to stand in the first | |
| twelve metres. | |
| Constraint (c) was added late, after an existing test caught a peak local density | |
| of **8.0 p/m²** against a jam density of 5.4 on a corridor whose mean was 1.3. | |
| Whole-edge headroom had been passing that traffic; the entrance had not. | |
| ### 7.6 Routing rules that had to be added | |
| - **No transit through seating areas or platforms.** A shortest path was | |
| otherwise happy to cut through a grandstand as a shortcut, misrouting the crowd | |
| and deadlocking against the people trying to leave. Barcelona gridlocked with | |
| 18,000 people stranded before this rule existed. | |
| - **No U-turns.** A routing table that has just been re-weighted can briefly make | |
| the corridor an agent is standing in look like the cheapest way onward. After | |
| repeated interventions this left 262 agents bouncing between two nodes forever. | |
| Reversing is refused unless it is genuinely the only option; the residue fell | |
| to 10. | |
| - **Penalty clamping and decay.** Intervention penalties are capped and relax | |
| towards neutral each refresh, so repeated operator action cannot permanently | |
| distort the cost surface. | |
| ### 7.7 Compliance | |
| Rerouting instructs people; it does not teleport them. Each agent carries a | |
| compliance probability sampled per scenario (typically 0.40–0.97). An instruction | |
| to reroute 40% reaches the agents whose route uses the bottleneck, and roughly | |
| 70% of those actually change. The measured improvement is therefore an | |
| improvement *net of people ignoring you*, which is why it is believable. | |
| --- | |
| ## 8. The Crowd State Engine | |
| `backend/flowtwin/crowd/` | |
| Turns raw agent positions into the quantities that predict failure. Per directed | |
| edge and per node, every second: | |
| | Quantity | Why it is measured | | |
| |---|---| | |
| | Occupancy, density | Density, not headcount, is the danger | | |
| | Peak **local** density | The worst 12 m, not the average | | |
| | Velocity, and velocity ratio vs free speed | Speed collapse precedes compression | | |
| | Inflow / outflow (people per minute) | The imbalance *is* the queue growth | | |
| | Capacity utilisation | How close to the design limit | | |
| | Density growth (per minute) | Rate of change is the leading indicator | | |
| | Queue growth (net people/minute) | Same, in people rather than density | | |
| | Opposing-flow conflict | Two streams on one floor is a distinct hazard | | |
| | Composite risk score (0–1) | One number for ranking | | |
| The **risk score** is a weighted sum, not a density threshold, because a single | |
| density number cannot distinguish a busy concourse from a compressing queue: | |
| ``` | |
| risk = 0.30·density + 0.18·utilisation + 0.18·density_growth | |
| + 0.12·queue_growth + 0.12·velocity_drop + 0.10·flow_conflict | |
| ``` | |
| Crucially, `risk_contributions()` exposes the per-term breakdown, so an alert | |
| does not just say "risk 0.81" — it says **why**: *density rising fast, velocity | |
| collapsed, queue growing, opposing flow*. A test asserts the contributions sum | |
| to the score, so the explanation can never drift from the number. | |
| Alerts are raised at 0.42 (watch), 0.58 (warning) and 0.74 (critical), and are | |
| de-duplicated so a two-way corridor produces one alert, not two. | |
| --- | |
| ## 9. Prediction | |
| `backend/flowtwin/prediction/` | |
| ### 9.1 The honest-baseline design | |
| The predictor is a **gradient-boosted regressor** (`HistGradientBoostingRegressor`), | |
| one model per horizon (+30, +60, +90, +120 s), predicting density on each edge. | |
| The important design decision is what it is measured against. There is an | |
| **analytic mass-balance baseline** — project density forward from current inflow, | |
| outflow and free storage — which is genuinely good, because pedestrian flow is | |
| substantially conservation of people. The trained model is used at inference time | |
| **only if it beats that baseline on held-out seeds.** Otherwise the system falls | |
| back to the baseline and says so in the UI. | |
| This is what stops "we used ML" from being decoration. | |
| ### 9.2 Features | |
| Seventeen, all from the Crowd State Engine, all quantities an operator would | |
| recognise: | |
| `density`, `density_growth_per_min`, `velocity_ratio`, `inflow_per_capacity`, | |
| `outflow_per_capacity`, `net_flow_per_capacity`, `occupancy_ratio`, | |
| `queue_ratio`, `flow_conflict`, `risk`, `upstream_density`, `downstream_density`, | |
| `downstream_wait_min`, `downstream_service_ratio`, `free_storage_ratio`, | |
| `length_m`, `width_m`. | |
| Note `upstream_density` and `downstream_density`: the model can see the | |
| neighbourhood, which is how it learns the spill-back behaviour of §2.2. | |
| ### 9.3 Training and validation | |
| The simulator is the data generator, which means **exact ground truth** — the | |
| label for "density here in 60 seconds" is simply what the density was, sixty | |
| seconds later, in a run that actually happened. | |
| Validation is on **disjoint seeds**: five seeds for training, two entirely | |
| different seeds held out, across all four scenarios including the railway | |
| terminus. 421,198 training rows, 169,364 test rows. | |
| Measured on held-out seeds: | |
| | Horizon | Model MAE | Baseline MAE | Improvement | R² | | |
| |---|---|---|---|---| | |
| | +30 s | 0.0097 | 0.0183 | **+47.3%** | 0.999 | | |
| | +60 s | 0.0159 | 0.0352 | **+54.7%** | 0.998 | | |
| | +90 s | 0.0221 | 0.0519 | **+57.4%** | 0.996 | | |
| | +120 s | 0.0278 | 0.0683 | **+59.3%** | 0.993 | | |
| The improvement *grows* with horizon, which is what you would hope: the physics | |
| baseline is nearly right in the short term and degrades as second-order effects | |
| accumulate; the model captures those. | |
| These numbers are visible in the dashboard, not just in a file. | |
| ### 9.4 The output an operator can use | |
| A density number in 90 seconds is not actionable. **"Critical in 96 seconds"** is. | |
| `time_to_threshold` interpolates the projected trajectory against the venue's | |
| critical density and reports lead time, which is what the alert displays and what | |
| the strategy engine uses to decide there is something worth acting on. | |
| ### 9.5 A performance trap worth knowing about | |
| Inference on 66 rows took **1,000 ms**. The same inference on one thread took | |
| **9 ms**. The BLAS/OpenMP thread pools were fighting over a tiny batch. Thread | |
| limits are pinned in `flowtwin/__init__.py` *before* numpy or sklearn are | |
| imported, which is the only place it works. | |
| --- | |
| ## 10. The Strategy Engine | |
| `backend/flowtwin/strategy/interventions.py` | |
| Candidates are **generated from the venue's topology and live state**, not read | |
| from a fixed list. A candidate only exists if the venue can actually support it. | |
| | Candidate | Generated when | What it does | | |
| |---|---|---| | |
| | **No action** | Always | The reference every other option is measured against | | |
| | **Redirect 20 / 30 / 40%** | An alternative path exists | Switches that fraction of the affected agents to adaptive routing with a cost penalty on the bottleneck | | |
| | **Stagger release** | Origin zones still have people to release | Holds 45% of the remaining departures from the top three feeding zones for 150 s | | |
| | **Open contingency lanes** | Another exit has **measured** spare capacity right now | +35% throughput there, and diverts 30% of the flow to it | | |
| | **Open emergency exit** | The venue has one still closed | Unlocks and staffs it — the only option that *adds* capacity — and diverts 35% | | |
| | **Destination split** | Two interchangeable destinations exist | Moves 30% of demand from one to the other: changing *where people are going*, not just how they get there | | |
| | **Combined** | Both a reroute and a stagger are available | Redirect 25% and hold 30% of remaining departures for 120 s | | |
| Two things are worth pointing out to a judge: | |
| - **"Open contingency lanes" quotes measured spare capacity in its own | |
| description.** It is not offered unless the alternative exit genuinely has room | |
| at this instant. | |
| - **Destination split is a different *kind* of lever.** Everything else changes | |
| routes; this changes destinations — operationally, "your coach has been moved to | |
| the south apron". | |
| --- | |
| ## 11. Counterfactual simulation | |
| `backend/flowtwin/strategy/counterfactual.py` | |
| This is the part that makes the recommendation a **measurement** rather than a | |
| rule. | |
| ``` | |
| capture the current state | |
| ├─ clone → apply "no action" → run 240 s → measure | |
| ├─ clone → apply "redirect 20%" → run 240 s → measure | |
| ├─ clone → apply "redirect 30%" → run 240 s → measure | |
| ├─ clone → apply "stagger release" → run 240 s → measure | |
| ├─ clone → apply "open emergency" → run 240 s → measure | |
| └─ … one clone per candidate | |
| compare → score → recommend | |
| ``` | |
| **Every clone starts byte-identical**, including the random number generator's | |
| internal bit-generator state. The only difference between two results is the | |
| intervention. That is the whole scientific claim, and two tests enforce it: one | |
| asserts that two branches of one state produce identical results, another that | |
| evaluating strategies does not advance the live run by a single step or move a | |
| single agent. | |
| Cloning is cheap because of the array layout: copy the agent arrays, three small | |
| integer routing matrices, the capacity budgets and the RNG state. | |
| Each roll-out measures sixteen quantities, including peak density on the watched | |
| asset, density **at the end of the window**, seconds spent critical, network-wide | |
| critical exposure, mean and p95 journey time, throughput, peak and final queue, | |
| aggregate risk, and how many people were rerouted. | |
| Note what is deliberately watched: **peak density on the asset under threat**, | |
| not the network maximum. A network maximum set by some unrelated corridor would | |
| be identical across all candidates and would make every option look the same. | |
| --- | |
| ## 12. Multi-objective optimisation and the decisiveness verdict | |
| `backend/flowtwin/strategy/optimizer.py` | |
| ### 12.1 The score | |
| Nine terms, each normalised against the no-action arm so a strategy's score reads | |
| directly as "fraction of the do-nothing outcome". The recommendation is `argmin J`. | |
| | Term | Weight | Asks | | |
| |---|---|---| | |
| | Peak density | 0.22 | How bad does it get? | | |
| | Critical duration | 0.20 | How long does it stay dangerous? | | |
| | **Density at end of window** | 0.12 | What state am I left in? | | |
| | **Queue at end of window** | 0.10 | What am I still holding? | | |
| | Average travel time | 0.10 | Are we punishing everyone to help a few? | | |
| | Aggregate risk | 0.10 | Integrated exposure, not just the peak | | |
| | Throughput | 0.08 | Are people actually leaving? | | |
| | Maximum queue | 0.04 | Worst single moment of holding | | |
| | Rerouting cost | 0.04 | Moving 20,000 people is heavier than moving 2,000 | | |
| All weights are environment-variable overridable, and the per-term contributions | |
| are exposed per strategy, so the table can be audited row by row. | |
| ### 12.2 Why "end of window" terms exist — the most interesting bug in the project | |
| Originally the score was dominated by peak terms. Intervene early and it worked | |
| beautifully. Intervene late and **every candidate returned an identical peak | |
| density to three decimal places**, and the "winner" was decided by the | |
| reroute-cost tiebreak — whichever option moved fewest people. | |
| The root cause is physical, not a coding error. Once a 4,000-person queue exists | |
| at a service-limited exit, it drains at the gate rate regardless of routing. The | |
| peak over the window is already determined. Peak-only scoring genuinely cannot | |
| tell the candidates apart. | |
| Two things were tried: | |
| 1. **Lengthen the roll-out.** Measured: separation returns only at a **720-second** | |
| horizon, costing 27 seconds of compute — for an answer that is still "this | |
| barely helps". Rejected on evidence. | |
| 2. **Add end-of-window terms.** Peaks ask "how bad does it get"; end-of-window | |
| terms ask "what am I still holding when the window closes". A strategy that | |
| leaves the bottleneck 1,500 people lighter at T+horizon is better even when | |
| both runs touched the same maximum. Adopted. | |
| ### 12.3 The decisiveness verdict | |
| The end-of-window terms sharpened the early case but did not manufacture a | |
| difference where there genuinely was none. So a second mechanism was added: | |
| > A candidate must beat no-action by at least **1.5%** of the do-nothing score | |
| > before it is *recommended*. Below that, the ranking still shows exactly what was | |
| > measured, but the recommendation falls back to no action and the system explains | |
| > why. | |
| The explanation is generated from the measured bottleneck state — queue held, | |
| discharge rate, arrival rate, estimated clearance time — and is quoted in full in | |
| §5. | |
| This turned the weakest moment in the demo into one of the strongest. A system | |
| that knows when it cannot help is more credible than one that always has an | |
| answer, and it removes the landmine of a judge pressing the button at the wrong | |
| moment. | |
| Guarded at both ends by tests: one asserts the early case still separates | |
| decisively, one asserts the late case refuses to pick a winner. The late fix | |
| cannot be obtained by flattening the early case. | |
| ### 12.4 Explainability with no language model anywhere | |
| The "why this strategy" panel is generated from **the same normalised terms that | |
| produced the score**. There is no narrative layer that could drift away from the | |
| arithmetic, and there is no LLM in the decision path. | |
| This is a deliberate, defensible position: every claim on screen is traceable to | |
| a measured number, and the reasoning shown is literally the reasoning used. | |
| --- | |
| ## 13. Dynamic routing | |
| `backend/flowtwin/routing/` | |
| ### 13.1 Next-hop tables | |
| Rather than storing a route per agent, FlowTwin stores, for every **policy** and | |
| every **destination**, the best next edge from each node. 40,000 agents then | |
| route with a single fancy-index lookup, and a change in conditions re-routes | |
| everybody who has not committed, in one Dijkstra per destination. | |
| It is also what makes counterfactuals affordable: cloning the routing state is | |
| cloning three small integer matrices. | |
| ### 13.2 Three policies, which are also the benchmark arms | |
| | Policy | What it is | | |
| |---|---| | |
| | **Shortest path** | Baseline A. Distance only. What people do without guidance. | | |
| | **Static assignment** | Baseline B. A real pre-event plan: method-of-successive-averages traffic assignment with BPR-style congestion costs, computed before the event from expected demand and never revised. | | |
| | **FlowTwin adaptive** | Live cost from distance, travel time, congestion, density, capacity and risk, refreshed every 5 simulated seconds. | | |
| Baseline B matters. It is not a straw man — it is what a competent operations | |
| team actually produces, and beating it is the interesting claim. | |
| ### 13.3 Oscillation control | |
| A naive adaptive router flaps: send people east, the east fills, send them west, | |
| the west fills. Four mechanisms prevent it: | |
| - **Hysteresis** — a node abandons its incumbent next hop only when the | |
| challenger is at least ~22% cheaper. | |
| - **Route commitment** — an agent keeps an adopted route for at least 25 s. | |
| - **Cycle breaking** — asserted acyclic by test. | |
| - **Penalty decay** — intervention penalties relax 2% per refresh towards neutral. | |
| --- | |
| ## 14. Perception — the Hugging Face path | |
| `backend/flowtwin/perception/` | |
| ### 14.1 Where it sits, and why that placement is the point | |
| ``` | |
| camera frame ──► Hugging Face crowd model ──► crowd observation ─┐ | |
| ├─► Crowd State Engine ─► prediction ─► strategy | |
| simulated agents ────────────────────────────────────────────────┘ | |
| ``` | |
| Both input modes converge on **one observation schema**. Density, risk, | |
| prediction, counterfactual and recommendation are then identical code whichever | |
| source is feeding them. A deployment can swap simulated crowds for real cameras | |
| without touching the decision path. | |
| It is deliberately **not** in the decision path itself. Nothing downstream | |
| depends on a neural network's opinion. | |
| ### 14.2 The candidate chain | |
| Tried in order; the first that loads wins; the selection is written to | |
| `models/perception_manifest.json`: | |
| 1. `AbdurRahman011/csrnet-indian-metro-crowd-density` — density-map regression. | |
| Counts by integrating a predicted density map, so it degrades gracefully in | |
| dense crowds where detectors fail. Trained on Indian metro crowds. | |
| 2. `AmineSam/irail-crowd-counting-yolov8n` — head detection fine-tuned on | |
| RPEE-Heads (railway platforms and event entrances). | |
| 3. `hustvl/yolos-tiny` — widely mirrored COCO detector, `person` class. | |
| 4. `facebook/detr-resnet-50` — second fallback. | |
| CSRNet's architecture is defined locally in `perception/csrnet.py` so a bare | |
| `state_dict` checkpoint can be loaded. | |
| ### 14.3 Sample frames with exact ground truth | |
| Three frames ship in `data/perception/`, **rendered from the digital twin** rather | |
| than photographed — a top-down view of a real corridor at a real moment of a real | |
| seeded run, one marker per person actually standing there. | |
| | Frame | People in shot | Area | Density | | |
| |---|---|---|---| | |
| | Exit B approach, free-flowing | 260 | 396 m² | 0.66 p/m² | | |
| | Exit B approach, standing queue | 1,762 | 396 m² | 4.45 p/m² | | |
| | Central foot-over-bridge, surge | 522 | 576 m² | 0.91 p/m² | | |
| Two reasons for renders rather than photographs. Shipping third-party crowd | |
| photographs in a public repository is a licensing problem. And a render has a | |
| property no photograph has: **the count is known exactly**, so the panel reports | |
| the model's *error* and not just its answer. A model that reports 1,300 on a frame | |
| containing 1,762 has undercounted by 26%, and being able to say that is worth more | |
| than a number with nothing to check it against. | |
| The UI labels them as renders. Uploading a real photograph runs the identical path. | |
| ### 14.4 Honest status | |
| **Not yet verified against downloaded weights.** The build environment has no | |
| network route to `huggingface.co` (every attempt returns `403 Tunnel connection | |
| failed`). Implemented and tested: the chain, the loader, the local CSRNet | |
| architecture, the manifest, the image → count → observation path, and the failure | |
| behaviour. Not executed: one real inference against real weights. | |
| One command closes it on any networked machine: | |
| ```bash | |
| pip install -r backend/requirements.txt | |
| python scripts/fetch_hf_model.py | |
| ``` | |
| **If it is never run, the endpoint reports the actual error and returns nothing. | |
| It has never fabricated a count, and a test asserts that.** Full record in | |
| [`HUGGING_FACE.md`](HUGGING_FACE.md). | |
| --- | |
| --- | |
| # Part III — The system as software | |
| ## 15. Architecture and module map | |
| ``` | |
| flowtwin/ | |
| ├── backend/ | |
| │ ├── flowtwin/ | |
| │ │ ├── __init__.py Thread-pool pinning (must precede numpy import) | |
| │ │ ├── config.py Every tuning constant, all env-overridable | |
| │ │ ├── main.py FastAPI app, lifespan, static mount | |
| │ │ ├── venue/ Domain model, compilation, scenario loading | |
| │ │ ├── simulation/ Agents, movement physics, the engine | |
| │ │ ├── crowd/ Density, flow, risk, alerts — the Crowd State Engine | |
| │ │ ├── prediction/ Features, analytic baseline, trained-model inference | |
| │ │ ├── routing/ Cost model, next-hop tables, static assignment | |
| │ │ ├── strategy/ Interventions, counterfactuals, optimiser, explanation | |
| │ │ ├── perception/ Hugging Face chain, CSRNet, observation schema | |
| │ │ ├── benchmarks/ Multi-seed, multi-arm evaluation harness | |
| │ │ ├── runtime/ Session lifecycle, broadcast loop, replay sessions | |
| │ │ └── api/ Routes, request/response schemas, WebSocket | |
| │ └── tests/ 79 tests across simulation, intelligence, API | |
| ├── frontend/ Zero-build ES modules + Canvas 2D | |
| ├── data/ | |
| │ ├── venues/ 3 venue JSON files | |
| │ ├── scenarios/ 4 scenario JSON files | |
| │ ├── perception/ 3 sample frames + ground-truth index | |
| │ └── fallback/ Pre-recorded frames (gitignored, regenerable) | |
| ├── models/ Trained predictor + its validation report | |
| ├── benchmarks/ Generated results, never hand-edited | |
| ├── scripts/ build_venues, train_predictor, run_benchmarks, | |
| │ make_perception_samples, fetch_hf_model, | |
| │ record_fallback, ui_check | |
| └── docs/ This file, ARCHITECTURE, DEMO, PS3_AUDIT, | |
| SPEC_AUDIT, HUGGING_FACE, ROADMAP | |
| ``` | |
| Roughly **6,600 lines of backend Python**, **2,750 lines of frontend**, and | |
| **1,100 lines of tests**. | |
| **Deliberate omissions.** No Redis, no PostgreSQL, no Docker, no build step. A | |
| simulation session is in-memory state on one process by nature; adding a datastore | |
| would mean serialising 40,000 agents per frame to solve a problem that does not | |
| exist at this scale. The rationale is written down in `ARCHITECTURE.md §10` so the | |
| absence reads as a decision rather than an omission. | |
| --- | |
| ## 16. Data flow and real-time transport | |
| ``` | |
| Browser FastAPI Simulator | |
| │ │ │ | |
| ├─ POST /api/simulation/start ───►│─ build venue, population ────►│ | |
| │◄──────── session + first frame ─┤ │ | |
| │ │ │ | |
| ├─ WS /api/simulation/{id}/stream►│ │ | |
| │ │ every 200 ms of wall clock: │ | |
| │ │ step × speed ──────────────►│ | |
| │ │◄──── state ───────────────────┤ | |
| │◄───────────── frame (push) ─────┤ │ | |
| │ │ │ | |
| ├─ POST /strategy/simulate ──────►│─ clone × 8, roll out ────────►│ | |
| │◄──── ranked strategies + why ───┤ │ | |
| ├─ POST /strategy/apply ─────────►│─ apply to the live run ──────►│ | |
| ``` | |
| **No per-frame polling.** The server pushes; the browser renders. Frames carry the | |
| crowd state, a bounded sample of agent positions for drawing (2,600 by default — | |
| a rendering budget, not a simulation limit), alerts, predictions and events. | |
| Sessions with no subscribers idle and are reaped. That was a real bug: a refreshed | |
| browser tab left an orphaned session simulating at 40×, which starved the event | |
| loop and made new runs appear to hang. | |
| --- | |
| ## 17. The frontend | |
| **Zero build step.** Vanilla ES modules served by the same FastAPI process. No | |
| npm, no bundler, no version skew, nothing to break on demo day. The trade-off | |
| against a React/Next.js frontend was made deliberately and is written down. | |
| **Layout.** The map dominates. Panels are subordinate. | |
| - **Left rail** — *Inputs*: expected crowd size, arrival/departure window, reroute | |
| compliance, the scheduled event and its severity, seed, baseline routing policy. | |
| Below it, the *Event schedule* showing what will execute and what has fired. | |
| Below that, on the Barcelona venue only, *Evidence & assumptions*. | |
| - **Centre** — the venue map on Canvas 2D: landmarks, corridors coloured by | |
| measured density, animated agents, predicted congestion drawn distinctly from | |
| current congestion, and rerouting paths when an intervention is applied. | |
| Layer toggles, a density legend and a scale bar. | |
| - **Right rail** — *Alerts* with severity, cause and lead time; *Prediction* with | |
| per-horizon projections and a model-accuracy modal; *Strategy* with the simulate | |
| button and the recommendation card. | |
| - **Drawer** — the strategy simulator: the full comparison table, the "why this | |
| strategy" panel, and the projected-density chart per candidate. | |
| **A rendering bug worth knowing about.** Frames arrive five times a second. | |
| Rebuilding an alert card on every frame restarts its CSS entry animation, which | |
| left the alert panel permanently mid-fade — measured opacity **0.26**, effectively | |
| invisible. Cards are now keyed on structure (`base_id:severity`) and live values | |
| are written in place. This shipped broken once. | |
| --- | |
| ## 18. Reproducibility and determinism | |
| Every run is fully determined by **(venue, scenario, seed, overrides)**. | |
| - The RNG's bit-generator state travels inside the snapshot, so a restored state | |
| produces the identical future. | |
| - Dwell decisions are drawn once at population creation, not from a live stream, | |
| for the same reason. | |
| - Interventions use a separate random stream so that applying a strategy never | |
| perturbs the population's own draws. | |
| - The seed is displayed in the metrics strip during every run. | |
| Tested directly: snapshot/restore is exact; two branches of one state are | |
| identical; branching does not disturb the parent; evaluating strategies does not | |
| advance the live simulation. | |
| This is what makes the benchmark numbers checkable rather than assertable. | |
| --- | |
| ## 19. The three venues | |
| All three are plain JSON against one schema. No venue-specific engine code exists. | |
| ### Circuit Alpha — fictional Grand Prix venue | |
| 30 nodes, 43 edges. Four perimeter exits, six spectator zones, a full concourse | |
| ring, three concession clusters, one emergency egress route, two transport | |
| interfaces and two car parks. **40,000 spectators**, simultaneous egress over an | |
| 18-minute curve, with Exit B losing half its throughput at T+4:00. This is the | |
| controlled stress test — the most instrumented venue, and the one the benchmark | |
| headline comes from. | |
| ### Circuit de Barcelona-Catalunya — documented-condition reconstruction | |
| 22 nodes, 33 edges. **78,000 spectators** at race-day scale, with the Montmeló | |
| rail approach deliberately constrained. | |
| The discipline here is the point. Every documented fact carries a source; every | |
| modelling assumption is labelled as an assumption; **both lists are on screen | |
| throughout**. The disclaimer is in the venue data, the briefing and the UI: | |
| > This is a counterfactual reconstruction using publicly documented event | |
| > conditions and a synthetic crowd model. It is not a replay of original | |
| > spectator telemetry, which is not public. | |
| The question it answers is *"given the documented conditions, what would FlowTwin | |
| have recommended?"* — never *"this is what happened."* | |
| ### Sangam Junction — fictional Indian metropolitan railway terminus | |
| 22 nodes, 34 edges. **26,000 passengers** discharged from six platforms over | |
| sixteen minutes, all of whom must change level through one of three routes: two | |
| foot-over-bridges and a subway. At T+4:30 the west bridge is closed to a quarter | |
| of its capacity on safety orders; at T+10:00 east gate screening slows. | |
| This venue exists as **evidence**, not decoration: | |
| - The **failure mode is different in kind**. A circuit fails at its perimeter; a | |
| terminus fails in the middle, at the level change, and the constraint is stair | |
| width rather than gate count. | |
| - The **food court is on the circulation path**, so about a quarter of the people | |
| crossing it stop for ~95 s and the concourse goes amber before the bridges do. | |
| The north gallery bypasses it at the cost of a longer walk — which is what gives | |
| the strategy engine a real question. | |
| - The **emergency gate is shut** and genuinely absent from routing. | |
| Building it required **one new node type and zero special-case simulation code**. | |
| It is fictional and labelled fictional; no real station is named and no real | |
| incident is reconstructed. | |
| --- | |
| ## 20. Testing and verification | |
| **79 automated tests**, in three files: | |
| - `test_simulation.py` (26) — the walking model's monotonicity, capacity budgets | |
| and fractional carry, queue behaviour, density never exceeding jam, snapshot | |
| exactness, branch independence, diversion and compliance, staggering, the | |
| What-If control genuinely retuning the scheduled event, emergency-exit routing | |
| exclusion and use, concession dwell and its reproducibility. | |
| - `test_intelligence.py` (28) — density and threshold maths, risk contributions | |
| summing to the score, bottleneck detection finding the right asset, alert | |
| de-duplication, feature-matrix sanity, prediction responding to a real change in | |
| state, routing acyclicity under hysteresis, adaptive routing genuinely avoiding | |
| the congested asset, counterfactual determinism, evaluation not advancing the | |
| live run, optimiser separation at an early intervention, optimiser refusal at a | |
| late one. | |
| - `test_api.py` (25) — every endpoint's success and failure modes, validation | |
| rejection, perception failing honestly, the sample route and its path-traversal | |
| guard, replay fallback. | |
| **Beyond unit tests:** | |
| - `scripts/ui_check.py` drives the entire acceptance path in a real Chromium | |
| browser via Playwright — load, run, wait for a critical alert, simulate | |
| strategies, check a recommendation is highlighted, apply it, watch the | |
| redistribution, switch to Barcelona and check the provenance panel, switch to | |
| the terminus and check its schedule, open the perception panel and verify every | |
| sample thumbnail actually loads. **Any console error or failed request fails | |
| the run.** It saves screenshots at each step. | |
| - `scripts/run_benchmarks.py` produces the quantitative results from real | |
| multi-seed runs. No figure in any document is typed by hand. | |
| - `.github/workflows/ci.yml` regenerates the venues and runs the suite on push. | |
| --- | |
| --- | |
| # Part IV — Evidence | |
| ## 21. Measured results | |
| Generated by `scripts/run_benchmarks.py`. Three arms — baseline shortest path, | |
| a static pre-event plan, and the full FlowTwin loop — across **8 independent | |
| seeds** of the complete simulation. Mean ± standard deviation. | |
| ### Circuit Alpha · 40,000 spectators · 8 seeds | |
| | Metric | Shortest path | Static plan | **FlowTwin** | vs baseline | | |
| |---|---|---|---|---| | |
| | Peak density (p/m²) | 3.6 ± 0.0 | 3.4 ± 0.1 | **2.0 ± 0.4** | **−42.6%** | | |
| | Critical exposure (corridor·s) | 1733 ± 130 | 1071 ± 216 | **0 ± 0** | **−100%** | | |
| | Maximum queue (people) | 4327 ± 56 | 4063 ± 111 | **2245 ± 440** | **−48.1%** | | |
| | Average journey (s) | 867 ± 12 | 798 ± 13 | **808 ± 15** | **−6.7%** | | |
| | 95th-percentile journey (s) | 2043 ± 89 | 1772 ± 94 | **1837 ± 106** | **−10.1%** | | |
| | Dispersal time, 95% (s) | 2523 ± 81 | 2241 ± 85 | **2286 ± 154** | **−9.4%** | | |
| | People rerouted | 0 | 2036 | 5814 | — | | |
| This is the headline. Time spent above the critical density goes to **zero on | |
| every seed**, peak density falls by 43%, the worst queue nearly halves — and | |
| average journey time gets *better*, not worse. Crowd-safety interventions usually | |
| trade delay for safety; here the congestion relief more than pays for the detour. | |
| The static plan is a genuine competitor, not a straw man: it beats naive | |
| shortest-path handily. FlowTwin beats it on every safety metric. | |
| The standard deviations are informative too. FlowTwin's peak density varies more | |
| across seeds (±0.4) than the baselines (±0.0–0.1), which is exactly what you | |
| would expect: the baselines always fail the same way, while an adaptive system's | |
| outcome depends on when the bottleneck happened to be caught. | |
| ### Circuit de Barcelona-Catalunya · 78,000 spectators · 6 seeds | |
| | Metric | Shortest path | Static plan | **FlowTwin** | vs baseline | | |
| |---|---|---|---|---| | |
| | Peak density (p/m²) | 3.2 ± 0.1 | 3.2 ± 0.1 | **1.3 ± 0.2** | **−59.0%** | | |
| | Critical exposure (corridor·s) | 1254 ± 165 | 1254 ± 165 | **0 ± 0** | **−100%** | | |
| | Maximum queue (people) | 3677 ± 152 | 3677 ± 152 | **1107 ± 228** | **−69.9%** | | |
| | Average journey (s) | 721 ± 3 | 721 ± 3 | 786 ± 26 | **+9.0%** | | |
| | 95th-percentile journey (s) | 1321 ± 7 | 1321 ± 7 | 1743 ± 144 | **+32.0%** | | |
| | Dispersal time, 95% (s) | 2514 ± 8 | 2514 ± 8 | 2837 ± 40 | **+12.8%** | | |
| | People rerouted | 0 | 0 | 8179 | — | | |
| **This one has a real trade-off and it is reported, not hidden.** Barcelona's | |
| danger sits on a narrow transport interface, and relieving it means sending | |
| thousands of people the long way round. Safety improves dramatically — peak | |
| density down 59%, the worst queue down 70%, critical exposure eliminated on every | |
| seed — and it costs 9% on the average journey and **32% on the 95th percentile**. | |
| That is the honest shape of the decision. One person in twenty gets home | |
| substantially later so that nobody stands in a dangerous crush. An operator | |
| should be told that price rather than sold a free lunch, and the optimiser's | |
| `avg_travel_time` weight is exactly the dial that sets how much of it you are | |
| willing to pay. | |
| **Two baselines, identical results.** On this venue shortest-path and the static | |
| plan produce byte-identical numbers, because most origin–destination pairs in the | |
| reconstructed topology have exactly one sensible route. That is a genuine property | |
| of the topology, not a broken benchmark, and it is documented rather than quietly | |
| dropped. | |
| ### Sangam Junction · railway terminus · 26,000 passengers · 6 seeds | |
| | Metric | Shortest path | Static plan | **FlowTwin** | vs baseline | | |
| |---|---|---|---|---| | |
| | Peak density (p/m²) | 3.0 ± 0.0 | 2.5 ± 0.2 | **2.4 ± 0.3** | **−20.3%** | | |
| | Critical exposure (corridor·s) | 0 | 0 | 0 | — | | |
| | Maximum queue (people) | 3600 ± 56 | 3459 ± 33 | 3883 ± 510 | +7.9% | | |
| | Average journey (s) | 887 ± 4 | 853 ± 40 | 1028 ± 87 | **+15.9%** | | |
| | 95th-percentile journey (s) | 1708 ± 33 | 1849 ± 295 | 2739 ± 613 | **+60.3%** | | |
| | Dispersal time, 95% (s) | 2257 ± 7 | 2431 ± 324 | 3298 ± 581 | **+46.1%** | | |
| | People rerouted | 0 | 2063 | 649 | — | | |
| **This is the worst table in the project and it is here on purpose.** On the | |
| terminus FlowTwin shaves 20% off peak density and pays for it with 16% on the | |
| average journey, 60% on the 95th percentile, and 46% on dispersal. Critical | |
| exposure is zero in *every* arm — at this crowd size the venue never becomes | |
| dangerous. So the system bought a safety improvement nobody needed, with a delay | |
| cost everybody paid. | |
| Do not hide this. Understand it, because the cause is precise and the fix is | |
| known. | |
| **Cause 1 — the venue is capacity-limited, not routing-limited.** Measured at the | |
| peak of the surge, every level-change route is at its service limit at the same | |
| moment: | |
| | Route | Capacity | In use | Spare | | |
| |---|---|---|---| | |
| | West foot-over-bridge (closed to 25%) | 130 /min | 129 | **1** | | |
| | Central foot-over-bridge | 900 /min | 900 | **0** | | |
| | East subway | 780 /min | 729 | 51 | | |
| Rerouting redistributes flow across capacity already in service. When all of it | |
| is saturated there is nothing to redistribute — which is why FlowTwin moves only | |
| 649 people here against 5,814 at Circuit Alpha. The decisiveness verdict is doing | |
| its job: most of the time it declines to act. | |
| **Cause 2 — the benchmark harness acts on a fixed review cycle; a human does | |
| not.** The FlowTwin arm re-evaluates every 180 s and applies whatever clears the | |
| 1.5% decisiveness bar, for the whole run. That makes the benchmark an **upper | |
| bound on intervention frequency**, not a model of the product's behaviour: in the | |
| console an operator presses the button when an alert says something is going | |
| critical, and on this venue nothing ever does. The measured cost above is the | |
| cost of intervening on a venue that did not need intervening on. | |
| **The fix, and it is the top of the roadmap.** The decisiveness threshold guards | |
| against candidates that are *indistinguishable from each other*. It does not yet | |
| guard against acting when *nothing is at risk*. A materiality gate — do not | |
| recommend an intervention if the projected peak stays below the venue's critical | |
| density across the whole window — closes it, and it is the same shape of | |
| judgement as the existing verdict. It is scoped in `ROADMAP.md` and it was found | |
| by this benchmark, which is the benchmark doing exactly what it is for. | |
| **What to say about it in a pitch.** Two true things, in this order: | |
| 1. *"A circuit is routing-limited: one exit failed while others had room, and we | |
| cut critical exposure to zero. A terminus is capacity-limited: all three | |
| staircases saturate at once, so we tell you rerouting won't help. Those are | |
| different problems and the system distinguishes them."* | |
| 2. *"And here's the honest part — on the terminus our benchmark harness keeps | |
| intervening anyway, on a cycle, and it costs journey time for a safety | |
| improvement that venue didn't need. That's a real finding from our own | |
| evaluation, and the gate that fixes it is the next thing we're building."* | |
| Owning that is worth more than a table with no weak column in it. | |
| The generated tables for all three venues, with every seed and every metric, are | |
| in `benchmarks/BENCHMARKS.md`. **No figure in this document was typed by hand.** | |
| --- | |
| ## 22. Every defect found and fixed | |
| This section exists because it is the strongest evidence that the model is right | |
| rather than merely convincing. Each of these was found by testing against physical | |
| reality, not by a linter. | |
| | # | Symptom | Root cause | Fix | | |
| |---|---|---|---| | |
| | 1 | Network throughput collapsed to ~1/10 of correct | Density averaged over a whole corridor, so a queue at a gate slowed people 200 m back with clear space | Density and speed evaluated per ~12 m cell | | |
| | 2 | Corridors absorbed impossible numbers of people | Links accepted at nominal capacity until physically full | Backward-wave receiving function — a link stops accepting *before* it is full, so congestion spills back upstream | | |
| | 3 | A 500/min gate discharged at under 200/min | Queue extent measured only at the stop line, so people had to walk *through* a near-jammed corridor to reach the back of the queue | Queue extent derived from everyone who has actually stopped | | |
| | 4 | Barcelona gridlocked with 18,000 stranded | Shortest paths used seating bowls as shortcuts, deadlocking against people leaving them | A route may start or end at a stand, never transit one | | |
| | 5 | 262 agents bouncing between two nodes forever | Re-weighted routing tables briefly made the corridor an agent was standing in look cheapest | U-turn guard; residue fell to 10 | | |
| | 6 | Repeated operator action permanently distorted the network | Intervention penalties compounded without limit | Penalties capped and decayed towards neutral each refresh | | |
| | 7 | Peak local density of 8.0 p/m² against a jam density of 5.4 | Admission limited by whole-edge headroom but not by space just inside the entrance | Entry-cell headroom limit | | |
| | 8 | The optimiser "recommended" on a rounding difference | Once a queue exists, the peak is already determined, so peak-dominated scoring cannot separate candidates | End-of-window objectives plus a 1.5% decisiveness threshold, with an explained hold verdict | | |
| | 9 | The explainability panel went blank exactly when it mattered | On a hold verdict, the winner *is* the baseline, so the comparison was an arm against itself | Falls back to the best rejected alternative, labelled as such | | |
| | 10 | A control that appeared to work and did nothing | The What-If capacity slider sent an empty override | The slider retunes the scheduled event itself | | |
| | 11 | The alert panel was effectively invisible (opacity 0.26) | Cards rebuilt 5×/second, restarting their entry animation | Cards keyed on structure; live values written in place | | |
| | 12 | New runs appeared to hang | Orphaned sessions from refreshed tabs kept simulating and starved the event loop | Sessions with no subscribers idle and are reaped | | |
| | 13 | A 66-row model inference took 1,000 ms | BLAS/OpenMP thread pools fighting over a tiny batch | Thread limits pinned before numpy is imported (9 ms) | | |
| | 14 | An opened emergency gate attracted nobody | An emergency route is geometrically longer, so the router kept using the old way | Opening a gate applies a routing bonus — unlocking it is also staffing and signing it | | |
| | 15 | The perception sample route 404'd | `samples()` advertised URLs for a route that was never implemented | Route added, with the filename reduced to its basename so a crafted name cannot escape the directory | | |
| --- | |
| ## 23. What is deliberately not built | |
| Recorded rather than hidden. Being able to answer "what's missing?" crisply is | |
| worth more than pretending nothing is. | |
| | Item | Status | Reasoning | | |
| |---|---|---| | |
| | Hugging Face chain verified against live weights | **Open** | No network route from the build environment. One command, one hour, on any networked machine. | | |
| | Personnel dispatch | Not built | The natural next feature — it answers *who should act*, which pairs perfectly with the hold verdict. Scoped in `ROADMAP.md`; about half a day. | | |
| | Ablation study | Not built | Nearly free; the benchmark harness already supports arms. Would answer "which part is doing the work". | | |
| | Venue upload / in-browser editor | Not built | Venues are JSON against a published schema and `build_venues.py` shows how to author one, but there is no upload endpoint. | | |
| | Multi-camera fusion | Not built | Single-frame perception only. | | |
| | Natural-language assistant | **Deliberately excluded** | Keeping every number in the decision path arithmetic is why the explainability story holds. | | |
| | Redis / PostgreSQL / Docker | **Deliberately excluded** | Simulation state is in-memory by nature. Rationale in `ARCHITECTURE.md §10`. | | |
| --- | |
| --- | |
| # Part V — The hackathon | |
| ## 24. Mapping to the evaluation criteria | |
| The rubric is 100 points across eight criteria. Here is what to point at for each. | |
| ### 1. Problem Understanding & Relevance — 15 | |
| Lead with §2.1: **crowd danger is not a headcount problem, it is a local density | |
| and flow problem**, and the failure is non-local and delayed. Then the killer | |
| detail: *by the time you can see it, rerouting may no longer help* — and show | |
| that the system knows this and says so. | |
| Ground it in the documented Barcelona 2022 conditions, then widen to the | |
| applications the problem statement names: railway stations, IPL egress, airport | |
| terminals, mass gatherings. Point at the terminus venue as proof you took | |
| "railway station design" literally rather than rhetorically. | |
| ### 2. Innovation & Originality — 15 | |
| The single strongest claim: **the recommendation is a measurement, not a rule.** | |
| Nobody else in this room will clone their entire simulation state eight times and | |
| race the futures against each other. | |
| Second: **the hold verdict**. A system that refuses to recommend when the | |
| measurement cannot separate the options, and explains why with the real discharge | |
| rate and clearance time, is a genuinely unusual piece of engineering judgement. | |
| Third: **the emergency exit is absent from routing, not expensive** — a small | |
| modelling decision with a large consequence, and easy to explain in ten seconds. | |
| ### 3. Technical Implementation — 20 | |
| The heaviest-weighted criterion, and where the depth lives: | |
| - Mesoscopic architecture chosen *because* counterfactuals must be affordable — | |
| 2–4 ms per step at 40,000 agents. | |
| - Weidmann fundamental diagram, per-cell evaluation, backward-wave receiving | |
| function, entry-cell admission, FIFO capacity budgets with fractional carry. | |
| - Gradient boosting validated on **disjoint seeds** against an analytic baseline, | |
| and used only if it wins. | |
| - Reverse-Dijkstra next-hop tables with hysteresis, commitment and cycle-breaking. | |
| - Byte-identical counterfactual branching including RNG state. | |
| - 79 tests, plus a real-browser acceptance run that fails on any console error. | |
| Have §22 (the defect table) ready. Fifteen real bugs, each with the symptom that | |
| revealed it, is the most persuasive artefact in the project. | |
| ### 4. Impact & Scalability — 15 | |
| Impact: the measured table — **critical exposure to zero, peak density −43%, max | |
| queue −48%, and journey times slightly better** — against a competent static plan, | |
| not a straw man. | |
| Scalability, and be specific rather than hand-wavy: | |
| - **Venue scalability** — three venues, one schema, zero venue-specific code. The | |
| terminus needed one node type. | |
| - **Population scalability** — 40,000 agents at 2–4 ms/step; 78,000 in Barcelona; | |
| hard-capped at 120,000. | |
| - **Deployment scalability** — one process, one command, no datastore, no build. | |
| - **Input scalability** — swap simulated agents for camera observations at the | |
| observation schema; nothing downstream changes. | |
| ### 5. User Experience & Design — 10 | |
| The map dominates; panels are subordinate. Every number on screen is measured; | |
| none are hard-coded. Alerts state their cause and their lead time. The strategy | |
| table is auditable row by row. The "why" panel is generated from the same | |
| arithmetic that produced the score. | |
| Mention the invisible-alert-panel bug (§22 #11) if design comes up — it shows the | |
| polish was verified, not assumed. | |
| ### 6. Completeness & Functionality — 10 | |
| One command, and the whole loop runs end to end without manual intervention. | |
| Three venues, four scenarios, 79 tests, a real-browser acceptance script, | |
| generated benchmarks, and six documents. `scripts/ui_check.py` output is the | |
| proof: it walks the entire acceptance path and fails on any error. | |
| Be honest about the one open item (§23) rather than letting a judge find it. | |
| ### 7. Presentation & Demo — 10 | |
| See §26. The rule: **run it live, and let the numbers on screen be the evidence.** | |
| Never read a figure aloud that is not visible behind you. | |
| ### 8. Q&A & Defense — 5 | |
| See §27. The general strategy: for every question, answer with a measured number | |
| or a named file, and if the answer is "not built", say so immediately and say why. | |
| --- | |
| ## 25. The pitch | |
| ### The 30-second version | |
| > "When a crowd turns dangerous, the problem isn't that there are too many people | |
| > — it's that there are too many people in one corridor, and by the time you can | |
| > see it, the queue that would need to move already can't. | |
| > | |
| > FlowTwin is a digital twin of the crowd. It simulates forty thousand people | |
| > walking through a venue in real time, predicts where flow will break down | |
| > ninety seconds before it does, and then does something no monitoring system | |
| > does: it clones the entire crowd, tries every option an operator has on its own | |
| > copy, and measures which one actually works. | |
| > | |
| > Across eight independent runs, time spent in dangerous density goes to zero — | |
| > and people get home *faster*, not slower." | |
| ### The 90-second version | |
| Add these three beats: | |
| **The mechanism, concretely.** "Eight complete copies of the crowd — every | |
| person's position, route and compliance, and the random number generator's | |
| internal state — one per candidate action. Each runs forward four minutes. Nine | |
| seconds later we have eight measured futures and we pick the best. The | |
| recommendation is a measurement, not a rule, and there is no language model | |
| anywhere in that path." | |
| **The honesty.** "And if you act too late, it tells you. Press the button fifteen | |
| minutes in and every option comes back identical, because a four-thousand-person | |
| queue drains at the gate's rate no matter where you send people. So it says: this | |
| exit is discharging at its limit with 3,275 people held, it needs nine minutes to | |
| clear, rerouting can't reach them, your remaining levers are capacity and | |
| staffing. A system that knows when it can't help is worth more than one that | |
| always has an answer." | |
| **The generality.** "It's not a motorsport product. Same engine, a railway | |
| terminus on a festival night — six platforms emptying through two foot-over-bridges | |
| and a subway. One new node type, zero special-case code. And on that venue it | |
| tells us rerouting *won't* help, because all three staircases are at their limit | |
| at once — which is the difference between a venue with an operations problem and a | |
| venue with a design problem." | |
| ### The one line to leave them with | |
| > **"Don't wait for the bottleneck. Simulate the intervention before it happens."** | |
| --- | |
| ## 26. The demo, minute by minute | |
| **Before you start:** server running, browser at 100% zoom, Simulation 1 | |
| pre-selected but **not** started. Have `benchmarks/BENCHMARKS.md` open in a second | |
| tab. Know your seed. | |
| | Time | What you do | What you say | | |
| |---|---|---| | |
| | **0:00** | Point at the header and the scenario switcher | The hook (§25). Name the three venues in one breath and move on. | | |
| | **0:30** | — | "Monitoring tells you where people *are*. The dangerous question is where flow will *fail*, and what to do before it does." | | |
| | **1:00** | Select **F1 Circuit Stress Test**, press **Run simulation**, set speed **20×** | "Forty thousand spectators, four exits, four destinations. At four minutes, Exit B loses half its throughput — a real change to the network, not an annotation." | | |
| | **1:45** | Point at the map as the east side reddens | "That's measured density per twelve metres of corridor, not a heat blob." | | |
| | **2:00** | Point at the Alerts panel | "Critical in ninety-six seconds. And it says *why*: density rising, velocity collapsed, queue growing. Those are the same six terms that produced the risk score." | | |
| | **2:30** | Point at the Prediction panel, click **Model accuracy** briefly | "Gradient boosting, trained on simulator ground truth, validated on seeds it has never seen — 47 to 59 per cent better than the physics baseline. If it hadn't beaten the baseline we'd be showing you the baseline." | | |
| | **3:00** | Press **Simulate strategies** | "Now the part that matters. Eight copies of the crowd — every person, every route, and the random number generator's internal state. One candidate each. Four minutes forward." | | |
| | **3:20** | Drawer opens; walk the table left to right | "These aren't estimates. Every column is measured from a run that happened. Redirect 40% wins by seventeen per cent." | | |
| | **3:40** | Point at the **why** panel | "Peak density down 28%, queue at end of window down 24%, critical time to zero, journey time unchanged, 834 people rerouted. That's the arithmetic that produced the score — there's no narrative layer that could drift from it." | | |
| | **4:00** | Press **Apply intervention** | "Same code path that was measured." | | |
| | **4:20** | Point at the reroute paths and the falling queue metric | "About seventy per cent comply. That's modelled per person, which is why the improvement is believable." | | |
| | **4:45** | Switch to **Barcelona 2022**, read the left rail | The provenance beat — facts with sources, assumptions labelled, and the disclaimer said out loud: *"we did not recreate Barcelona; we reconstructed the documented conditions."* | | |
| | **5:15** | Switch to **Railway Terminus**, run at 40× | "Same engine, no motorsport. Six platforms, two foot-over-bridges, a subway. The failure happens in the *middle* of the venue, not at the perimeter." Then the food court and the amber emergency gate (§19). | | |
| | **5:30** | Point at the benchmark table | "Eight independent seeds. Critical exposure to zero, peak density down 43%, max queue down 48% — and journeys six per cent *faster*. Every figure generated, none typed." | | |
| | **5:45** | — | The closing line (§25). | | |
| **If you have a spare minute, this is the beat to add:** the hold verdict. | |
| Re-run Simulation 1, jump to T+15:00, press **Simulate strategies**, and read the | |
| verdict aloud. It is the single most memorable thing in the demo. | |
| **Rules for yourself.** Run live. Never read a number that is not on screen. | |
| Press *Simulate strategies* while the prediction still says "critical in N | |
| seconds", not after the alert has been red for five minutes. | |
| --- | |
| ## 27. Q&A defence | |
| **"Is this real or is the simulation faked?"** | |
| Every number on screen is computed. The venue JSON has capacities and areas; the | |
| physics is Weidmann's fundamental diagram; the seed is displayed and the run is | |
| reproducible from it. Change the crowd size in the left rail and re-run — the | |
| outcome changes because the physics changed. | |
| **"Where is the AI?"** | |
| Three places, and be precise about each. A gradient-boosted model predicting | |
| density at four horizons, validated on disjoint seeds and used only because it | |
| beats a strong analytic baseline by 47–59%. A Hugging Face crowd-counting model | |
| on the perception path, converting camera frames into the same observation schema | |
| the simulator produces. And the decision layer — counterfactual search over a | |
| generated candidate set with multi-objective scoring. Deliberately **not** a | |
| language model, because the explainability story depends on the reasoning being | |
| the same arithmetic that produced the score. | |
| **"Isn't this just a shortest-path algorithm?"** | |
| Shortest path is baseline A in the benchmark, and it is the one FlowTwin beats by | |
| 43% on peak density. There is also baseline B — a proper capacity-aware | |
| pre-event plan using method-of-successive-averages assignment — which is what a | |
| competent operations team actually produces. FlowTwin beats that on every safety | |
| metric too. | |
| **"How do you know the recommendation is right?"** | |
| We don't assert it, we measure it. Each candidate is applied to a byte-identical | |
| clone and simulated forward; the numbers in the table come from those runs. And | |
| when the measurement can't separate the options, the system says so rather than | |
| picking one — that threshold is 1.5% and it's in the config. | |
| **"What if a judge presses the button at the wrong moment?"** | |
| Then they see the hold verdict, which is a better demo than the recommendation. | |
| That was a real bug we found and fixed: the optimiser used to pick a winner on a | |
| rounding difference. Now it explains why nothing helps, with the measured | |
| discharge rate and clearance time. | |
| **"Have you verified the Hugging Face model?"** | |
| Not against downloaded weights — the build environment has no route to | |
| huggingface.co, and I'd rather say that than claim otherwise. The chain, the | |
| loader, the local CSRNet architecture, the manifest and the failure behaviour are | |
| all implemented and tested; one command closes it on a networked machine. And | |
| what it does *today* if no model loads is report the actual error — it has never | |
| fabricated a count, and there's a test asserting it. | |
| **"Would this work at my venue?"** | |
| The venue is JSON against a published schema — nodes with positions, areas and | |
| service rates; edges with lengths, widths and capacities. Three venues ship, | |
| including a railway terminus, and none of them required engine changes. What is | |
| *not* built is an upload UI, so today it's a file you author with the script in | |
| `scripts/build_venues.py`. | |
| **"Does it scale to a Kumbh-scale gathering?"** | |
| The simulation is capped at 120,000 agents and runs 78,000 comfortably at 2–4 ms | |
| per step. Beyond that the honest answer is that the mesoscopic model would need | |
| to be partitioned, and that the harder problem at that scale isn't compute — it's | |
| that a single operator can't act on a hundred simultaneous bottlenecks, which is | |
| why personnel dispatch is the next feature. | |
| **"What would you build next?"** | |
| Personnel dispatch. Right now every lever moves the crowd; none of them moves | |
| staff. And it pairs exactly with the hold verdict — when routing can no longer | |
| help, "send four stewards to Exit B" is what the system should be able to say. | |
| **"Your railway venue barely improves. Isn't that a failure?"** | |
| It's the most useful result we have. That venue is *capacity*-limited, not | |
| routing-limited: at the peak of the surge all three level-change routes are at | |
| their service limit simultaneously — one, zero and fifty-one people per minute of | |
| spare capacity, with eight thousand people queued behind them. Rerouting | |
| redistributes capacity that's already in service; when all of it is saturated | |
| there is nothing to redistribute. So the engine says so, instead of claiming a | |
| win. And that answer is actionable in a different way: it says the fix is a | |
| fourth bridge or a phased platform release, not better signage. The problem | |
| statement lists railway station *design* as an application — that is what | |
| designing looks like. | |
| **"What's the weakest part?"** | |
| The Hugging Face path being unverified against live weights, and the absence of | |
| personnel dispatch. Both are in `ROADMAP.md` with the work scoped. The Barcelona | |
| static baseline also produces results identical to shortest path, because that | |
| topology mostly has one sensible route per origin-destination pair — a real | |
| property of the venue, documented rather than hidden. | |
| --- | |
| ## 28. Failure drills | |
| Practise these once. Confidence when something breaks is worth more than the | |
| thing not breaking. | |
| | If this happens | Do this | | |
| |---|---| | |
| | The live run stalls or the connection chip goes red | Re-select the scenario — the app tears the session down and starts cleanly. If it recurs, switch to the recorded fallback: it replays through the identical interface. | | |
| | Strategy simulation takes longer than expected | Say what it is doing: "that's eight full simulations running." It is ~9 s at 40,000 agents; at 78,000 it is longer and that is honest. | | |
| | The perception panel shows no model | This is the designed behaviour and a good beat. "No weights on this machine, so it reports the error instead of guessing. It has never invented a count." | | |
| | A judge asks for a venue you don't have | Show the venue JSON and `scripts/build_venues.py`. The schema is the answer. | | |
| | The projector eats the dark theme | The metrics strip and the strategy table are the highest-contrast elements. Demo from those. | | |
| | Everything fails | `docs/DEMO.md` carries the full narrative and every real number, and `benchmarks/BENCHMARKS.md` is generated evidence you can read from. | | |
| --- | |
| *Last verified against the repository at the commit that introduced the railway | |
| terminus, the emergency-exit routing semantics, the end-of-window optimiser | |
| objectives and the decisiveness verdict.* | |