--- license: apache-2.0 pipeline_tag: video-classification tags: - speed-estimation - ego-motion - optical-flow - video - regression - pytorch - motorsport - dashcam - imu - gps - dji - gopro - telemetry - karting --- # SpeedNet v5 — ego-speed estimation from onboard video SpeedNet estimates **vehicle speed directly from onboard/POV video** — no GPS, no CAN bus, no sensors. Point it at a dashcam clip, a go-kart helmet-cam recording, or ATV footage, and it returns a per-frame speed curve. It was built to answer a practical question: *can we recover telemetry from footage that has none?* — e.g. indoor karting, where GPS does not work. ## Demo Four unedited minutes across the domains the model has to cope with. The cyan trace is SpeedNet reading the video; the dashed white trace is GPS, shown wherever ground truth exists, with a running MAE. Nothing is cherry-picked — the two car clips are included precisely because open-road driving is the model's weakest regime (see Limitations). **ATV, forest trail** — 9–61 km/h, unseen recording. MAE **3.4 km/h**. **Car, country roads** — 19–64 km/h. MAE **11.1 km/h**. **Car, motorway** — 21–108 km/h. MAE **10.7 km/h**. High speed is a hard case: at 110 km/h the flow field saturates and distant structure carries most of the signal. See *Limitations*. **Indoor kart** — no GPS exists here at all; this is the case the project was built for. No ground-truth trace, because there is nothing to compare against frame by frame; validation is against the venue's timing system (see *Evaluation*). ## Motivation Onboard motorsport footage is everywhere — helmet cams, chest mounts, dashcams — but the telemetry that makes it *interesting* usually is not. Rental karts have no data logger you can access; indoor tracks deny GPS entirely; action cameras often record IMU but no position. We wanted two things: **measure speed** where no sensor can, and **add live telemetry overlays** (speed, g-force, lap timing) to motorsport videos that were recorded without any. Getting there meant overcoming a chain of problems, each documented in this card because they will bite anyone building something similar: 1. **Monocular scale ambiguity** — optical flow alone cannot distinguish "fast and far" from "slow and close". Solved with the RGB context branch, which only generalized once training spanned 7 distinct domains; with fewer, it memorized venues. 2. **No indoor ground truth** — GPS does not exist on an indoor kart track. Solved with weak supervision from the venue's official lap times (mean speed per lap = track length / lap time), aligned to video by detecting recurring track features. 3. **Lap memorization** — with only ~50 supervised laps, the GRU learned to recognize individual laps instead of estimating speed. Solved by augmenting the lap-supervision forward passes (flip, noise, occluders). 4. **Sequence length as a scale cue** — training with two window lengths taught the model to use context length itself as a signal; hence the two documented inference modes. 5. **Compressed dynamics** — regression pulls toward the mean (corners read fast, straights slow). Solved, when IMU data exists, by the physics fusion described below: measured longitudinal g-force drives the acceleration profile, centripetal physics anchors cornering speed, and track length locks the scale. 6. **The published track length is not the driven distance.** Our indoor supervision used `mean speed = track length / lap time` with the venue's advertised 1000 m. That figure is the **centreline**. A racing line cuts every corner by up to half the track width, so on a 13 m wide circuit with 33.1 rad of total turning it saves `6.5 m x 33.1 = 215 m` — the driven lap is ~790 m, and every label was 27% too fast. If you use lap-time weak supervision, measure the turning and subtract; do not trust the marketing number. Corner count is a good sanity check that you have the right layout (we measured 14, matching the venue's spec) while being completely independent of scale. ``` video ──► optical flow (motion field) ──┐ ├──► GRU ──► speed (m/s, per frame) video ──► low-res context frame ────────┘ (tells the model the scene scale) ``` ## Why the two branches? Optical flow encodes how fast pixels move — which depends on **both** speed and scene distance. 130 km/h on an open highway and 57 km/h between the close walls of an indoor kart track can produce similar flow magnitudes. A monocular model therefore cannot know the absolute scale from motion alone. The **context branch** sees one low-resolution RGB frame and learns a scene→scale mapping (indoor track / forest trail / highway), which lets a single checkpoint work across very different environments. This only started working once the training corpus spanned 7 distinct domains — with fewer domains the branch memorized venues instead of generalizing. ## Quickstart ```bash pip install torch opencv-python-headless numpy pandas huggingface_hub ``` Fetch the repo via `huggingface_hub` rather than downloading the weight file directly — this is also what makes your download count toward the model's stats on the Hub: ```python from huggingface_hub import snapshot_download local_dir = snapshot_download("x2q/speednet") ``` ```python import torch from modeling_speednet import SpeedNet, predict_video model = SpeedNet() model.load_state_dict(torch.load("speednet_v5.pt", map_location="cuda")) result = predict_video(model, "onboard_clip.mp4", device="cuda") print(f"mean speed: {result['speed_smooth_mps'].mean() * 3.6:.1f} km/h") ``` Or run the bundled CLI example: ```bash python example.py my_clip.mp4 # -> my_clip_speed.csv + my_clip_speed.png ``` `predict_video` handles everything: decoding, resizing to 320x180, Farneback optical flow, frame-rate normalization, context frames, streaming GRU inference and 1-second median smoothing. ## Critical usage notes 1. **Pick the right inference mode.** The model was trained with two window lengths, so sequence length acts as a scale cue: - *short mode* (default): hidden state reset every 16 frames — for roads, trails and open environments. Reproduces the GPS-domain metrics below. - *long-context mode* (`long_context=True`): hidden state carried across the clip — required for closed-course footage (indoor karting): short windows collapse indoor scale to roughly half the true value. 2. **Any frame rate works** — flow is normalized to a 15 fps convention internally (`raw_flow * fps / 15`). Feeding pre-computed flow without this normalization will mis-scale predictions. 3. **Forward-facing footage.** The model was trained on forward-facing onboard cameras. Rear/side cameras are out of distribution. 4. **Static occluders are fine.** Hoods, helmet edges and mounts were augmented during training; the model ignores zero-flow regions. ## Evaluation Held-out recordings, never trained on. Two inference modes exist (sequence length acts as a scale cue for this architecture); both are reported — pick per the recommendation column. RMSE in m/s against GPS: | Test clip | short windows | streaming | recommended | |---|---|---|---| | ATV, forest trail (unseen day) | **0.99** | 1.60 | short | | Car, coastal Denmark (zero-shot recording) | 1.57 | **1.39** | streaming | | Car, low-speed urban | 4.44 | **4.25** | streaming | | Car, urban/rural | 7.82 | **6.00** | streaming | | Car, motorway 100–135 km/h | 6.50 | **5.87** | streaming | | Indoor go-kart (lap means vs official timing) | — | **3.70 km/h MAE** (bias −2.84) | streaming | The indoor row compares mean predicted speed per lap against `driven lap length / official lap time` on laps excluded from training, using the corrected 790 m driven-lap length (see challenge 6 above — the venue's advertised 1000 m is the centreline). **Scale caveat:** all speeds carry a ±5% scale uncertainty indoors. The 790 m driven-lap figure is bracketed by geometry (a racing line cannot be shorter than ~785 m) and by flow evidence (which integrates to ~753 m); the residual −2.8 km/h indoor bias reflects exactly this uncertainty. ## Training data ~5 hours of labeled video across 7 domains, all with real ground truth: | Domain | Source | Speed signal | |---|---|---| | Indoor go-kart (2 races) | helmet cam, indoor kart circuit | official lap times (weak supervision: mean speed per lap over the 790 m driven racing line) | | ATV / forest | action cam + GPS remote | 10 Hz GPS | | Mini-loader / grass | action cam + GPS remote | 10 Hz GPS | | Car / Denmark | action cam + GPS, dashcam | 10 Hz + 1 Hz GPS | | Car / EU | [L2D](https://huggingface.co/datasets/yaak-ai/L2D) (Apache-2.0) | 10 Hz CAN/GNSS | | Car / US | [comma2k19](https://github.com/commaai/comma2k19) (MIT) | 20 Hz pose velocities | Training tricks that mattered: horizontal-flip and synthetic-occluder augmentation on flow; **augmenting the lap-supervision forward passes** (otherwise the GRU memorizes individual laps); context dropout (15%) and photometric jitter on the context frame; per-recording train/val/test splits. ## Limitations - **Protocol sensitivity**: sequence length is a scale cue (an artifact of two-length training); use the recommended mode per the evaluation table (short for trail/off-road, streaming for cars and closed courses). Length-randomized training removes this — validated on a companion track-specialized model — and is planned for the next release. - **Indoor absolute scale is ±5%**: supervision is lap time x driven lap length, and the driven racing line length is derived, not surveyed. - **Compressed dynamics within a lap/segment**: like most regression models, predictions are pulled toward the mean — cornering speeds read slightly high, straight-line peaks slightly low. If your footage has IMU data, see the physics-fusion recipe below. - Monocular scale is *learned*, not measured: radically new scene types (aircraft, boats, rear-facing cameras) will mis-scale. - GPS ground truth is itself smoothed (~1 s); treat sub-second dynamics as indicative. - Farneback flow degrades with heavy motion blur and very low light. ## Advanced: physics fusion (video + IMU) If the recording also has IMU data, you can fix the compressed dynamics by solving for the speed curve that best satisfies, jointly: - `dv/dt = a_lon` — measured longitudinal acceleration (true accel/braking), - `v = a_lat / ω` — centripetal anchor in corners (|yaw rate| > 30°/s), - `∫ v dt = track_length` per lap (if lap times are known), - the video model as a weak level prior. A few thousand Adam steps over the speed vector suffice. On our indoor race this reproduces official lap averages to **0.07 km/h** with per-lap distance 789/790 m, while restoring realistic braking-into-corner / accelerating-onto-straight dynamics. ## Model details - ~1.0 M parameters (flow CNN 4 conv blocks → 256-d; context CNN 3 blocks → 64-d; 2-layer GRU, 192 hidden; linear head). - Input: flow [T, 2, 72, 128] normalized by 8.0; context RGB [3, 72, 128] in [-0.5, 0.5]. Output × 40.0 = m/s. - Trained ~30 epochs, AdamW, Huber loss, cosine schedule, on a single NVIDIA GB10; checkpoint selected on combined GPS-validation + lap-holdout score. ## License Weights and code: Apache-2.0. Trained only on own recordings and permissively licensed public data (MIT / Apache-2.0). ## Citation ``` @misc{speednet2026, title = {SpeedNet v5: ego-speed estimation from onboard video}, author = {x2q}, year = {2026}, url = {https://huggingface.co/x2q/speednet} } ```