File size: 12,378 Bytes
2162853 046e550 a86caf0 2162853 9285f34 2162853 a86caf0 9285f34 a86caf0 9285f34 a86caf0 9285f34 a86caf0 2162853 fc5224a 2162853 d7cf6e8 2162853 9285f34 2162853 9285f34 2162853 9285f34 2162853 d7cf6e8 2162853 d7cf6e8 2162853 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 | ---
license: apache-2.0
pipeline_tag: video-classification
tags:
- speed-estimation
- ego-motion
- optical-flow
- video
- regression
- pytorch
- motorsport
- dashcam
- imu
- gps
- dji
- gopro
- telemetry
- karting
---
# SpeedNet v5 β ego-speed estimation from onboard video
SpeedNet estimates **vehicle speed directly from onboard/POV video** β no GPS,
no CAN bus, no sensors. Point it at a dashcam clip, a go-kart helmet-cam
recording, or ATV footage, and it returns a per-frame speed curve.
It was built to answer a practical question: *can we recover telemetry from
footage that has none?* β e.g. indoor karting, where GPS does not work.
## Demo
Four unedited minutes across the domains the model has to cope with. The
cyan trace is SpeedNet reading the video; the dashed white trace is GPS,
shown wherever ground truth exists, with a running MAE. Nothing is
cherry-picked β the two car clips are included precisely because open-road
driving is the model's weakest regime (see Limitations).
**ATV, forest trail** β 9β61 km/h, unseen recording. MAE **3.4 km/h**.
<video src="https://huggingface.co/x2q/speednet/resolve/main/assets/demo_atv.mp4" controls muted playsinline width="100%"></video>
**Car, country roads** β 19β64 km/h. MAE **11.1 km/h**.
<video src="https://huggingface.co/x2q/speednet/resolve/main/assets/demo_car_rural.mp4" controls muted playsinline width="100%"></video>
**Car, motorway** β 21β108 km/h. MAE **10.7 km/h**. High speed is a hard
case: at 110 km/h the flow field saturates and distant structure carries
most of the signal. See *Limitations*.
<video src="https://huggingface.co/x2q/speednet/resolve/main/assets/demo_car_motorway.mp4" controls muted playsinline width="100%"></video>
**Indoor kart** β no GPS exists here at all; this is the case the project was
built for. No ground-truth trace, because there is nothing to compare
against frame by frame; validation is against the venue's timing system
(see *Evaluation*).
<video src="https://huggingface.co/x2q/speednet/resolve/main/assets/demo_kart.mp4" controls muted playsinline width="100%"></video>
## Motivation
Onboard motorsport footage is everywhere β helmet cams, chest mounts,
dashcams β but the telemetry that makes it *interesting* usually is not.
Rental karts have no data logger you can access; indoor tracks deny GPS
entirely; action cameras often record IMU but no position. We wanted two
things: **measure speed** where no sensor can, and **add live telemetry
overlays** (speed, g-force, lap timing) to motorsport videos that were
recorded without any.
Getting there meant overcoming a chain of problems, each documented in this
card because they will bite anyone building something similar:
1. **Monocular scale ambiguity** β optical flow alone cannot distinguish
"fast and far" from "slow and close". Solved with the RGB context branch,
which only generalized once training spanned 7 distinct domains; with
fewer, it memorized venues.
2. **No indoor ground truth** β GPS does not exist on an indoor kart track.
Solved with weak supervision from the venue's official lap times
(mean speed per lap = track length / lap time), aligned to video by
detecting recurring track features.
3. **Lap memorization** β with only ~50 supervised laps, the GRU learned to
recognize individual laps instead of estimating speed. Solved by
augmenting the lap-supervision forward passes (flip, noise, occluders).
4. **Sequence length as a scale cue** β training with two window lengths
taught the model to use context length itself as a signal; hence the two
documented inference modes.
5. **Compressed dynamics** β regression pulls toward the mean (corners read
fast, straights slow). Solved, when IMU data exists, by the physics
fusion described below: measured longitudinal g-force drives the
acceleration profile, centripetal physics anchors cornering speed, and
track length locks the scale.
6. **The published track length is not the driven distance.** Our indoor
supervision used `mean speed = track length / lap time` with the venue's
advertised 1000 m. That figure is the **centreline**. A racing line cuts
every corner by up to half the track width, so on a 13 m wide circuit
with 33.1 rad of total turning it saves `6.5 m x 33.1 = 215 m` β the
driven lap is ~790 m, and every label was 27% too fast. If you use
lap-time weak supervision, measure the turning and subtract; do not trust
the marketing number. Corner count is a good sanity check that you have
the right layout (we measured 14, matching the venue's spec) while being
completely independent of scale.
```
video βββΊ optical flow (motion field) βββ
ββββΊ GRU βββΊ speed (m/s, per frame)
video βββΊ low-res context frame βββββββββ
(tells the model the scene scale)
```
## Why the two branches?
Optical flow encodes how fast pixels move β which depends on **both** speed
and scene distance. 130 km/h on an open highway and 57 km/h between the close
walls of an indoor kart track can produce similar flow magnitudes. A
monocular model therefore cannot know the absolute scale from motion alone.
The **context branch** sees one low-resolution RGB frame and learns a
sceneβscale mapping (indoor track / forest trail / highway), which lets a
single checkpoint work across very different environments. This only started
working once the training corpus spanned 7 distinct domains β with fewer
domains the branch memorized venues instead of generalizing.
## Quickstart
```bash
pip install torch opencv-python-headless numpy pandas huggingface_hub
```
Fetch the repo via `huggingface_hub` rather than downloading the weight
file directly β this is also what makes your download count toward the
model's stats on the Hub:
```python
from huggingface_hub import snapshot_download
local_dir = snapshot_download("x2q/speednet")
```
```python
import torch
from modeling_speednet import SpeedNet, predict_video
model = SpeedNet()
model.load_state_dict(torch.load("speednet_v5.pt", map_location="cuda"))
result = predict_video(model, "onboard_clip.mp4", device="cuda")
print(f"mean speed: {result['speed_smooth_mps'].mean() * 3.6:.1f} km/h")
```
Or run the bundled CLI example:
```bash
python example.py my_clip.mp4 # -> my_clip_speed.csv + my_clip_speed.png
```
`predict_video` handles everything: decoding, resizing to 320x180, Farneback
optical flow, frame-rate normalization, context frames, streaming GRU
inference and 1-second median smoothing.
## Critical usage notes
1. **Pick the right inference mode.** The model was trained with two window
lengths, so sequence length acts as a scale cue:
- *short mode* (default): hidden state reset every 16 frames β for roads,
trails and open environments. Reproduces the GPS-domain metrics below.
- *long-context mode* (`long_context=True`): hidden state carried across
the clip β required for closed-course footage (indoor karting): short
windows collapse indoor scale to roughly half the true value.
2. **Any frame rate works** β flow is normalized to a 15 fps convention
internally (`raw_flow * fps / 15`). Feeding pre-computed flow without this
normalization will mis-scale predictions.
3. **Forward-facing footage.** The model was trained on forward-facing
onboard cameras. Rear/side cameras are out of distribution.
4. **Static occluders are fine.** Hoods, helmet edges and mounts were
augmented during training; the model ignores zero-flow regions.
## Evaluation
Held-out recordings, never trained on. Two inference modes exist (sequence
length acts as a scale cue for this architecture); both are reported β
pick per the recommendation column. RMSE in m/s against GPS:
| Test clip | short windows | streaming | recommended |
|---|---|---|---|
| ATV, forest trail (unseen day) | **0.99** | 1.60 | short |
| Car, coastal Denmark (zero-shot recording) | 1.57 | **1.39** | streaming |
| Car, low-speed urban | 4.44 | **4.25** | streaming |
| Car, urban/rural | 7.82 | **6.00** | streaming |
| Car, motorway 100β135 km/h | 6.50 | **5.87** | streaming |
| Indoor go-kart (lap means vs official timing) | β | **3.70 km/h MAE** (bias β2.84) | streaming |
The indoor row compares mean predicted speed per lap against
`driven lap length / official lap time` on laps excluded from training,
using the corrected 790 m driven-lap length (see challenge 6 above β
the venue's advertised 1000 m is the centreline).
**Scale caveat:** all speeds carry a Β±5% scale uncertainty indoors. The
790 m driven-lap figure is bracketed by geometry (a racing line cannot be
shorter than ~785 m) and by flow evidence (which integrates to ~753 m);
the residual β2.8 km/h indoor bias reflects exactly this uncertainty.
## Training data
~5 hours of labeled video across 7 domains, all with real ground truth:
| Domain | Source | Speed signal |
|---|---|---|
| Indoor go-kart (2 races) | helmet cam, indoor kart circuit | official lap times (weak supervision: mean speed per lap over the 790 m driven racing line) |
| ATV / forest | action cam + GPS remote | 10 Hz GPS |
| Mini-loader / grass | action cam + GPS remote | 10 Hz GPS |
| Car / Denmark | action cam + GPS, dashcam | 10 Hz + 1 Hz GPS |
| Car / EU | [L2D](https://huggingface.co/datasets/yaak-ai/L2D) (Apache-2.0) | 10 Hz CAN/GNSS |
| Car / US | [comma2k19](https://github.com/commaai/comma2k19) (MIT) | 20 Hz pose velocities |
Training tricks that mattered: horizontal-flip and synthetic-occluder
augmentation on flow; **augmenting the lap-supervision forward passes**
(otherwise the GRU memorizes individual laps); context dropout (15%) and
photometric jitter on the context frame; per-recording train/val/test splits.
## Limitations
- **Protocol sensitivity**: sequence length is a scale cue (an artifact of
two-length training); use the recommended mode per the evaluation table
(short for trail/off-road, streaming for cars and closed courses).
Length-randomized training removes this β validated on a companion
track-specialized model β and is planned for the next release.
- **Indoor absolute scale is Β±5%**: supervision is lap time x driven lap
length, and the driven racing line length is derived, not surveyed.
- **Compressed dynamics within a lap/segment**: like most regression models,
predictions are pulled toward the mean β cornering speeds read slightly
high, straight-line peaks slightly low. If your footage has IMU data, see
the physics-fusion recipe below.
- Monocular scale is *learned*, not measured: radically new scene types
(aircraft, boats, rear-facing cameras) will mis-scale.
- GPS ground truth is itself smoothed (~1 s); treat sub-second dynamics as
indicative.
- Farneback flow degrades with heavy motion blur and very low light.
## Advanced: physics fusion (video + IMU)
If the recording also has IMU data, you can fix the compressed dynamics by
solving for the speed curve that best satisfies, jointly:
- `dv/dt = a_lon` β measured longitudinal acceleration (true accel/braking),
- `v = a_lat / Ο` β centripetal anchor in corners (|yaw rate| > 30Β°/s),
- `β« v dt = track_length` per lap (if lap times are known),
- the video model as a weak level prior.
A few thousand Adam steps over the speed vector suffice. On our indoor race
this reproduces official lap averages to **0.07 km/h** with per-lap distance
789/790 m, while restoring realistic braking-into-corner /
accelerating-onto-straight dynamics.
## Model details
- ~1.0 M parameters (flow CNN 4 conv blocks β 256-d; context CNN 3 blocks β
64-d; 2-layer GRU, 192 hidden; linear head).
- Input: flow [T, 2, 72, 128] normalized by 8.0; context RGB [3, 72, 128]
in [-0.5, 0.5]. Output Γ 40.0 = m/s.
- Trained ~30 epochs, AdamW, Huber loss, cosine schedule, on a single
NVIDIA GB10; checkpoint selected on combined GPS-validation + lap-holdout
score.
## License
Weights and code: Apache-2.0. Trained only on own recordings and
permissively licensed public data (MIT / Apache-2.0).
## Citation
```
@misc{speednet2026,
title = {SpeedNet v5: ego-speed estimation from onboard video},
author = {x2q},
year = {2026},
url = {https://huggingface.co/x2q/speednet}
}
```
|