crowd-nav — human-response models for the Foresight controller

Trained weights for Foresight, a convex receding-horizon controller that predicts how each nearby pedestrian will respond to the robot's candidate plan, and then chooses that plan by exact convex optimization. These are the response models the controller calls; the planner is plain code (no learned policy, no RL).

file model role
linres_head.pt Linear-response operator (headline) best closed-loop performance; makes planning an exact convex QP
residual_predictor_k1.pt Residual predictor most accurate on real data; the safest variant in closed loop
residual_predictor.pt legacy deprecated — trained with a robot-force prior of 0 but deployed at 1; superseded by residual_predictor_k1.pt

Which one should I use?

linres_head.pt unless you have a reason not to. It is a few centimetres less accurate as a pure forecaster, but its prediction is exactly affine in the robot's planned displacement, so the planner gets an exact constant Jacobian and the safety constraints stay linear — the whole plan selection becomes a convex QP that solves in ~17 ms with a checkable feasibility certificate. That trade wins where it counts: it leads our benchmark, and it is the only variant that crossed a 90-pedestrian oncoming corridor with zero contacts.

Measured results

Isaac Sim, 16 scenarios × 10 paired seeds (160 episodes per method), identical hardware:

model goal reach hard collisions / ep plan latency (mean)
Foresight-LinRes (linres_head.pt) 94.4 % 0.013 16.6 ms
Foresight-Residual (residual_predictor_k1.pt) 90.0 % 0.006 19.7 ms
Foresight-SFM (physics only, ablation) 93.8 % 0.013 17.2 ms
Foresight-CV (straight-line, ablation) 90.0 % 0.019 16.4 ms
SICNav [T-RO '24] (nonconvex bilevel MPC) 89.2 % 0.042 1843 ms

Stress test — a 4.5 m × 105 m corridor with 90 pedestrians all walking against the robot and a continuous inflow so it never empties: LinRes reached the goal in 98.5 s with zero contacts; against pedestrian models it was never tuned on (SRFM, social-force) it reached in 85.0 s and 91.7 s, also contact-free.

Prediction accuracy on the held-out real split (factual ADE, metres — lower is better):

model no robot stationary robot moving robot
Constant velocity 0.240 0.253 0.354
Social force 0.343 0.358 0.416
Residual (k1) 0.192 0.217 0.285
Linear response 0.204 0.248 0.319

Architecture

Both models are a physics prior plus a learned correction — small by design, because the real moving-robot data is scarce.

prior (no parameters, per pedestrian per 0.25 s step)
  f_goal  = (1.3·(g−p)/‖g−p‖ − v)/0.5
  f_ped   = Σ 2.0·exp(−d/0.4)·d̂      (d < 2.0 m)
  f_robot = k·exp(−d/0.5)·d̂          (d < 3.0 m)
  v ← clip(v + f·Δt, 1.8);  p ← p + v·Δt

residual_predictor_k1.pt     ŷ = prior(k=1) + f_θ(x)
  f_θ : MLP 22 → 128 → 128 → 16   (21,520 params)

linres_head.pt               ŷ = prior(k=0) + f_θ(x) + G_φ(x)·ΔR
  f_θ : MLP 22 → 128 → 128 → 16   (21,520 params)
  G_φ : MLP 22 → 64 → 96          (7,712 params) ⇒ rank-3 16×16 operator
        G = Σ_{r=1..3} a_r b_rᵀ,  ΔR = (R − 1⊗r₀)/5.0
        ⇒ ∂ŷ/∂R = G/5.0  — exact, constant, no differentiation needed

Input x (22-d, pedestrian local frame): 1 s of past positions (4 @ 0.25 s), velocity, goal direction, robot relative position, a robot-nearby flag, the two nearest neighbours, and a 3-way condition one-hot. Output: 8 × 2 future displacements — 2.0 s at 0.25 s.

Training: supervised (Adam, lr 1e-3, batch 128, ~25 epochs), synthetic pre-train then fine-tune on the real PeRoI robot–pedestrian recordings. No reinforcement learning anywhere in this stack. Both run in well under 1 ms per call on CPU.

Load

pip install huggingface_hub torch
hf download elmoghany/crowd-nav linres_head.pt residual_predictor_k1.pt --local-dir .
import torch
ck = torch.load("linres_head.pt", map_location="cpu")
ck["config"]
# {'in_dim': 22, 'horizon_dim': 16, 'rank': 3, 'hidden': 64, 'dr_scale': 5.0,
#  'sfm_k_robot': 0.0, 'convention': 'y = sfm(k=0) + f_resid(x_ctx) + head(x_ctx, dR)'}
ck["f_resid"], ck["head"]   # two state_dicts

The config block is authoritative: sfm_k_robot must match the robot-force strength used in your prior at deployment. Mismatching them is a real bug we shipped once — the residual then corrects an error that is not there. residual_predictor.pt is the artifact of that mistake and is kept only for reproducibility.

Model classes (ResidualPredictor, LinearResponseHead) and the full controller live in the github.com/elmoghany/crowd-nav -- a one-file, simulator-free controller with a quickstart.py that fetches these weights and verifies your install in one command. The full research stack (convex-QP planner, Isaac benchmarks, paper) is at github.com/elmoghany/crowd-nav-legacy.

Honest limitations

  • Real-data forecasting is a tie. On four real robot–pedestrian datasets (including JRDB), every learned model here ties the analytic baselines on factual accuracy. Our measured explanation: the counterfactual proxy those benchmarks rely on (constant-velocity extrapolation) correlates only r ≈ 0.30 with exact simulated counterfactuals — the target is mostly noise. The gains above are closed-loop gains, not forecasting gains.
  • Compliance bias. The training recordings contain cooperative pedestrians. On deliberately stubborn crowds the k1 residual is the weakest learned variant; compliance-randomized training data is the known fix, not yet applied.
  • Cross-dataset transfer. Moving residual_predictor_k1.pt to JRDB over-predicts robot-induced deviation by ~6× (its k=1 prior is calibrated to our robot). It makes a controller conservative, not unsafe — it still threaded real JRDB crowds contact-free — but retune sfm_k_robot for a different platform.
Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading