# FlowTwin — Architecture This document explains how the system is put together and, where a decision was not obvious, why it was made that way. ``` ┌──────────────── OBSERVATION ────────────────┐ │ │ synthetic agents camera frame (simulator, exact ground truth) (Hugging Face crowd model) │ │ └──────────────────┬──────────────────────────┘ ▼ CROWD STATE ENGINE occupancy · density · inflow · outflow · velocity utilisation · density growth · queue growth · risk │ ┌──────────────┴──────────────┐ ▼ ▼ CURRENT STATE PREDICTED STATE (+30 / +60 / +90 / +120 s) └──────────────┬──────────────┘ ▼ BOTTLENECK DETECTION ▼ STRATEGY ENGINE candidates generated from venue topology ▼ COUNTERFACTUAL SIMULATOR each candidate applied to an identical clone of state ▼ OPTIMIZER J = Σ wᵢ · (metricᵢ / no-action metricᵢ) ▼ RECOMMENDATION + EXPLANATION ▼ RACE CONTROL UI ▼ operator applies → NEW STATE ─┐ ▲ │ └────────────────┘ ``` --- ## 1. Venue digital twin A venue is a directed weighted graph (`venue/models.py`). Nodes are places a spectator can be; edges are the pedestrian links between them. An edge carries `length_m`, `width_m` and `capacity_ppm`. A node may carry `area_m2` (so it can hold a crowd) and `service_rate_ppm` (how many people per minute it can process). **A perimeter exit is not a sink.** It is a throughput constraint on the way to somewhere else — a station, a car park. Modelling it as a destination would hide exactly the queue this project exists to predict. Sinks are transport interfaces and car parks; exits are gates in between. `CompiledVenue` is the array-oriented view built once per venue: node and edge attributes as numpy arrays, a CSR-style adjacency, polyline geometry with cumulative arc length, and the pairing between the two directions of a two-way corridor. The hot loop never touches a Python object. Both venues are generated by `scripts/build_venues.py` rather than hand-written JSON, so edge lengths are always derived from the drawn geometry and the map can never disagree with the physics. --- ## 2. Simulation `simulation/engine.py`. A mesoscopic, capacity-constrained pedestrian network model. Agents are individuals — own walking speed, destination, route, reroute compliance — but they travel along graph edges rather than in free 2-D space. **Why not a full social-force model?** A microscopic 2-D simulation of 40,000 agents cannot run five alternative futures while an operator waits. The counterfactual comparison *is* the product, so the movement model was chosen to make it affordable: a step costs ~2–4 ms for 40,000 agents, which makes an eight-strategy sweep over a 300-second horizon about six seconds. Four pieces of physics do the work: **Speed depends on local density.** Weidmann's (1993) exponential fundamental diagram. Free walking at low density, speed collapse approaching jam density. **Density is evaluated per cell, not per edge.** Every corridor is divided into ~12 m cells. This matters more than it sounds: with edge-average density, a queue at a gate slows *everybody* in the corridor, including someone 200 m back with clear space in front of them. The result was a corridor that filled uniformly to jam and delivered a tenth of its real throughput. With cells, the congested region grows upstream one cell at a time, as a queue does. **Throughput is bounded twice, and admission is FIFO.** Moving from one link to the next requires passing a *node* budget (the gate's people-per-minute) and an *edge* budget (what the next corridor accepts). Fractional capacity is carried across steps so a 90/minute gate really passes 90 per minute. Whoever has been waiting longest goes first. **A link stops accepting people before it is physically full.** Receiving capacity falls as a link fills, at the backward wave speed. Without this, a corridor quietly absorbs an impossible number of people instead of pushing the congestion upstream — spillback is what turns one degraded gate into a network-wide event, and it has to be in the model. Agents that reach the head of a queue and cannot pass are marked blocked and spread across the length the queue physically occupies, so the map shows the queue backing up the corridor and approaching walkers meet it where it really is. ### Reproducibility and branching `snapshot()` captures everything: agent arrays, budgets with their fractional carry, cost model, routing tables, crowd-state history, counters, fired events, and the state of both random generators. `branch()` produces a detached copy. This is the foundation of the counterfactual: every candidate strategy starts from a byte-identical state with an identical random stream, so the *only* difference between two results is the intervention. Tests assert it directly. --- ## 3. Crowd State Engine `crowd/state.py`. Converts agent positions into the aggregates everything downstream reasons about, and keeps a rolling history so it can talk about *trajectories*, not just instants. A corridor at 2.1 p/m² filling at 0.4 p/m² per minute is a different operational situation from one sitting at 2.1 p/m² in steady state, and only the first needs an intervention. That distinction is the reason for the history buffers. The composite risk score combines density, capacity utilisation, density growth, queue growth, velocity drop and opposing flow — deliberately not a threshold on raw density, which cannot tell a busy concourse from a compressing queue. Weights are configurable and the per-term contributions are exposed, so an alert can say *why* it fired. Density is reported two ways: the **mean over the corridor** (the headline number, which moves continuously as a queue lengthens) and the **peak in any single cell**. The dashboard labels which is which. --- ## 4. Prediction `prediction/`. Features come straight from the Crowd State Engine — the model sees exactly what the operator sees, with no privileged knowledge of the scenario script. Two predictors exist: - **Analytic mass-balance projection.** `density(t+h) = density + (inflow − outflow)·h/(60·area)`, damped as the corridor approaches jam. Always available. - **Gradient-boosted regressor**, one per horizon, trained by `scripts/train_predictor.py` on data the simulator generates. Because the simulator provides exact ground truth, the model can be validated honestly. Training and test use **disjoint seeds**, and the report records the model's mean absolute error alongside the baseline's. **The trained model is only used if it beat the baseline on held-out seeds**; otherwise `DensityPredictor` refuses to load it. The dashboard shows which predictor is active and its accuracy. Density is a property of the physical corridor, so a projection that differs by direction is an artefact of direction-specific features, not a real disagreement — the carrying direction's projection is mirrored to its pair so the alert list, the prediction panel and the strategy engine cannot quote different futures for the same piece of concrete. Projections are memoised per (simulation, step): one dashboard frame asks for them several times and they must all agree. --- ## 5. Routing `routing/`. FlowTwin stores, for each policy and destination, the best **next edge** from every node, rather than a route per agent. A 40,000-agent population then routes with one fancy-index lookup, and a change in crowd state re-routes everyone who has not committed, in one Dijkstra per destination. Three policies exist so the benchmark can compare like with like: - `shortest_path` — minimise distance. - `static_assignment` — a real method-of-successive-averages traffic assignment with BPR-style congestion costs, computed once before the event from expected demand. Capacity-aware, but blind to what actually happens. - `flowtwin_adaptive` — `C_e = α·distance + β·travel time at current speed + γ·congestion + δ·risk`, plus expected waiting time at each node from its live queue and service rate. That node term is what makes rerouting more than cosmetic: an exit with 2,400 people waiting and a 750/minute service rate is a 192-second delay, and the router has to know it. **Oscillation control.** A node only abandons its incumbent next hop when the challenger is meaningfully cheaper (hysteresis), and agents that adopt the adaptive plan keep it. Hysteresis can in principle retain a hop that closes a loop, so the merged table is checked for termination and any node that fails is reverted to the pure shortest-path hop. A test asserts the tables stay acyclic and that repeated refreshes on an unchanged state change nothing. --- ## 6. Strategy Engine `strategy/`. The candidate set is **not** a fixed list — it is derived from the bottleneck that was detected and what the surrounding network makes possible. A reroute is only offered when an alternative path exists; an alternate exit only when one has measured spare throughput; a destination split only when two interchangeable destinations exist. Families: no action, reroute (20/30/40%), staggered release, open an alternate exit, destination split, and a combined response. Each intervention knows how to apply itself to a simulator. That is the whole contract, and it matters: the counterfactual applies it to a clone, the operator applies it to the live run, and both go through the same code path — **what the operator gets is what was measured.** Compliance is modelled per agent. An instruction reaches everyone selected; only those whose personal compliance clears a random draw act on it. ### Counterfactual and optimizer For each candidate: clone, apply, roll forward, measure. Metrics are scoped to the asset under threat — a network-wide maximum set by some unrelated corridor would make every strategy look identical. ``` J = w₁·peak density + w₂·critical duration + w₃·travel time + w₄·risk + w₅·queue + w₆·(1/throughput) + w₇·reroute cost ``` Every term is normalised against the *no action* counterfactual, so weights express relative importance rather than doing unit conversion, and a score reads directly as "fraction of the do-nothing outcome". The optimum is `argmin J`. The explanation is generated from the same normalised terms that produced the score. There is no separate narrative layer that could drift away from the arithmetic, and no language model anywhere in this path — a numerical safety-adjacent decision should be measurable and reproducible, which an LLM is not. --- ## 7. Runtime `runtime/session.py`. A session owns one simulator, advances it on a wall-clock timer at the requested speed multiplier, and publishes frames to connected dashboards. Stepping and counterfactual sweeps run off the event loop so the WebSocket never stalls; a slow client has its oldest frame dropped rather than slowing the venue down. A session with no subscribers does no work, and idle sessions are reaped. A refreshed browser tab would otherwise leave an orphaned simulation stepping forever, and enough of those starve the event loop. `ReplaySession` implements the same interface from a precomputed recording. It is demo insurance only — the live simulation is always the primary path. --- ## 8. Frontend `frontend/`. A single-page Race Control console served by the backend: no build step, no package install, one process to start. That is a deliberate trade against the framework named in the specification. On demo day, one command that serves both the API and the UI removes an entire class of failure — dependency install, build output, port and CORS configuration — and none of what the dashboard has to do (a canvas map, a WebSocket, a few panels) needs a framework. Everything is vanilla ES modules and hand-written CSS, and it works offline. The map is Canvas 2D, layered: circuit geometry, corridors coloured by measured density, predicted congestion as a dashed overlay, animated flow direction, reroute overlay, agents, nodes with queue rings, labels, and a pulsing halo on the primary bottleneck. Nothing on it is decorative state — if a corridor is orange, its measured density put it there. Panels re-render only when their content would actually differ. Frames arrive five times a second, and rewriting a panel on every one of them restarts its entry animation and leaves it permanently mid-fade. --- ## 9. Configuration Everything a deployment might reasonably want to change lives in `config.py` and is overridable by environment variable: movement physics, risk weights, routing costs, optimizer weights, prediction horizons, server behaviour, perception model. No tuning constant is hard-coded inside an algorithm module. One non-obvious setting is applied at package import: the BLAS/OpenMP thread pools are pinned to one thread. FlowTwin's numeric work is many *small* operations, and on a small container the thread pools spend far longer coordinating than computing — one edge-density inference measured 1,000 ms across two threads and 9 ms on one. ## 10. Deliberate omissions - **Redis** — the state that would live there is owned by a single process, and adding a network hop between a simulation and its own state buys nothing for a single-node demo while adding a thing that can be down. - **PostgreSQL** — nothing in the demo path needs durable storage. Venues and scenarios are JSON; benchmark results and recordings are files. - **A language model in the decision loop** — excluded on purpose. It may be added as an interface layer that reads structured engine output; it must never determine an intervention.