crowd-nav — human-response models for the Foresight controller
Trained weights for Foresight, a convex receding-horizon controller that predicts how each nearby pedestrian will respond to the robot's candidate plan, and then chooses that plan by exact convex optimization. These are the response models the controller calls; the planner is plain code (no learned policy, no RL).
| file | model | role |
|---|---|---|
linres_head.pt |
Linear-response operator (headline) | best closed-loop performance; makes planning an exact convex QP |
residual_predictor_k1.pt |
Residual predictor | most accurate on real data; the safest variant in closed loop |
residual_predictor.pt |
legacy | deprecated — trained with a robot-force prior of 0 but deployed at 1; superseded by residual_predictor_k1.pt |
Which one should I use?
linres_head.pt unless you have a reason not to. It is a few centimetres less accurate as a
pure forecaster, but its prediction is exactly affine in the robot's planned displacement, so the
planner gets an exact constant Jacobian and the safety constraints stay linear — the whole plan
selection becomes a convex QP that solves in ~17 ms with a checkable feasibility certificate.
That trade wins where it counts: it leads our benchmark, and it is the only variant that crossed a
90-pedestrian oncoming corridor with zero contacts.
Measured results
Isaac Sim, 16 scenarios × 10 paired seeds (160 episodes per method), identical hardware:
| model | goal reach | hard collisions / ep | plan latency (mean) |
|---|---|---|---|
Foresight-LinRes (linres_head.pt) |
94.4 % | 0.013 | 16.6 ms |
Foresight-Residual (residual_predictor_k1.pt) |
90.0 % | 0.006 | 19.7 ms |
| Foresight-SFM (physics only, ablation) | 93.8 % | 0.013 | 17.2 ms |
| Foresight-CV (straight-line, ablation) | 90.0 % | 0.019 | 16.4 ms |
| SICNav [T-RO '24] (nonconvex bilevel MPC) | 89.2 % | 0.042 | 1843 ms |
Stress test — a 4.5 m × 105 m corridor with 90 pedestrians all walking against the robot and a continuous inflow so it never empties: LinRes reached the goal in 98.5 s with zero contacts; against pedestrian models it was never tuned on (SRFM, social-force) it reached in 85.0 s and 91.7 s, also contact-free.
Prediction accuracy on the held-out real split (factual ADE, metres — lower is better):
| model | no robot | stationary robot | moving robot |
|---|---|---|---|
| Constant velocity | 0.240 | 0.253 | 0.354 |
| Social force | 0.343 | 0.358 | 0.416 |
| Residual (k1) | 0.192 | 0.217 | 0.285 |
| Linear response | 0.204 | 0.248 | 0.319 |
Architecture
Both models are a physics prior plus a learned correction — small by design, because the real moving-robot data is scarce.
prior (no parameters, per pedestrian per 0.25 s step)
f_goal = (1.3·(g−p)/‖g−p‖ − v)/0.5
f_ped = Σ 2.0·exp(−d/0.4)·d̂ (d < 2.0 m)
f_robot = k·exp(−d/0.5)·d̂ (d < 3.0 m)
v ← clip(v + f·Δt, 1.8); p ← p + v·Δt
residual_predictor_k1.pt ŷ = prior(k=1) + f_θ(x)
f_θ : MLP 22 → 128 → 128 → 16 (21,520 params)
linres_head.pt ŷ = prior(k=0) + f_θ(x) + G_φ(x)·ΔR
f_θ : MLP 22 → 128 → 128 → 16 (21,520 params)
G_φ : MLP 22 → 64 → 96 (7,712 params) ⇒ rank-3 16×16 operator
G = Σ_{r=1..3} a_r b_rᵀ, ΔR = (R − 1⊗r₀)/5.0
⇒ ∂ŷ/∂R = G/5.0 — exact, constant, no differentiation needed
Input x (22-d, pedestrian local frame): 1 s of past positions (4 @ 0.25 s), velocity, goal
direction, robot relative position, a robot-nearby flag, the two nearest neighbours, and a 3-way
condition one-hot. Output: 8 × 2 future displacements — 2.0 s at 0.25 s.
Training: supervised (Adam, lr 1e-3, batch 128, ~25 epochs), synthetic pre-train then fine-tune on the real PeRoI robot–pedestrian recordings. No reinforcement learning anywhere in this stack. Both run in well under 1 ms per call on CPU.
Load
pip install huggingface_hub torch
hf download elmoghany/crowd-nav linres_head.pt residual_predictor_k1.pt --local-dir .
import torch
ck = torch.load("linres_head.pt", map_location="cpu")
ck["config"]
# {'in_dim': 22, 'horizon_dim': 16, 'rank': 3, 'hidden': 64, 'dr_scale': 5.0,
# 'sfm_k_robot': 0.0, 'convention': 'y = sfm(k=0) + f_resid(x_ctx) + head(x_ctx, dR)'}
ck["f_resid"], ck["head"] # two state_dicts
The config block is authoritative: sfm_k_robot must match the robot-force strength used in
your prior at deployment. Mismatching them is a real bug we shipped once — the residual then
corrects an error that is not there. residual_predictor.pt is the artifact of that mistake and is
kept only for reproducibility.
Model classes (ResidualPredictor, LinearResponseHead) and the full controller live in the
github.com/elmoghany/crowd-nav -- a one-file,
simulator-free controller with a quickstart.py that fetches these weights and verifies your
install in one command. The full research stack (convex-QP planner, Isaac benchmarks, paper) is at
github.com/elmoghany/crowd-nav-legacy.
Honest limitations
- Real-data forecasting is a tie. On four real robot–pedestrian datasets (including JRDB), every learned model here ties the analytic baselines on factual accuracy. Our measured explanation: the counterfactual proxy those benchmarks rely on (constant-velocity extrapolation) correlates only r ≈ 0.30 with exact simulated counterfactuals — the target is mostly noise. The gains above are closed-loop gains, not forecasting gains.
- Compliance bias. The training recordings contain cooperative pedestrians. On deliberately stubborn crowds the k1 residual is the weakest learned variant; compliance-randomized training data is the known fix, not yet applied.
- Cross-dataset transfer. Moving
residual_predictor_k1.ptto JRDB over-predicts robot-induced deviation by ~6× (its k=1 prior is calibrated to our robot). It makes a controller conservative, not unsafe — it still threaded real JRDB crowds contact-free — but retunesfm_k_robotfor a different platform.