File size: 12,378 Bytes
2162853
 
 
 
 
 
 
 
 
 
 
 
046e550
 
 
 
a86caf0
 
2162853
 
9285f34
2162853
 
 
 
 
 
 
 
a86caf0
 
9285f34
 
 
 
 
a86caf0
9285f34
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
a86caf0
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
9285f34
 
 
 
 
 
 
 
 
 
a86caf0
2162853
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
fc5224a
 
 
 
 
 
 
 
 
 
2162853
 
 
 
 
 
 
d7cf6e8
2162853
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
9285f34
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
2162853
 
 
 
 
 
 
9285f34
2162853
 
 
 
 
 
 
 
 
 
 
 
 
 
9285f34
 
 
 
 
 
2162853
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
d7cf6e8
 
2162853
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
d7cf6e8
2162853
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
---
license: apache-2.0
pipeline_tag: video-classification
tags:
  - speed-estimation
  - ego-motion
  - optical-flow
  - video
  - regression
  - pytorch
  - motorsport
  - dashcam
  - imu
  - gps
  - dji
  - gopro
  - telemetry
  - karting
---

# SpeedNet v5 β€” ego-speed estimation from onboard video

SpeedNet estimates **vehicle speed directly from onboard/POV video** β€” no GPS,
no CAN bus, no sensors. Point it at a dashcam clip, a go-kart helmet-cam
recording, or ATV footage, and it returns a per-frame speed curve.

It was built to answer a practical question: *can we recover telemetry from
footage that has none?* β€” e.g. indoor karting, where GPS does not work.

## Demo

Four unedited minutes across the domains the model has to cope with. The
cyan trace is SpeedNet reading the video; the dashed white trace is GPS,
shown wherever ground truth exists, with a running MAE. Nothing is
cherry-picked β€” the two car clips are included precisely because open-road
driving is the model's weakest regime (see Limitations).

**ATV, forest trail** β€” 9–61 km/h, unseen recording. MAE **3.4 km/h**.

<video src="https://huggingface.co/x2q/speednet/resolve/main/assets/demo_atv.mp4" controls muted playsinline width="100%"></video>

**Car, country roads** β€” 19–64 km/h. MAE **11.1 km/h**.

<video src="https://huggingface.co/x2q/speednet/resolve/main/assets/demo_car_rural.mp4" controls muted playsinline width="100%"></video>

**Car, motorway** β€” 21–108 km/h. MAE **10.7 km/h**. High speed is a hard
case: at 110 km/h the flow field saturates and distant structure carries
most of the signal. See *Limitations*.

<video src="https://huggingface.co/x2q/speednet/resolve/main/assets/demo_car_motorway.mp4" controls muted playsinline width="100%"></video>

**Indoor kart** β€” no GPS exists here at all; this is the case the project was
built for. No ground-truth trace, because there is nothing to compare
against frame by frame; validation is against the venue's timing system
(see *Evaluation*).

<video src="https://huggingface.co/x2q/speednet/resolve/main/assets/demo_kart.mp4" controls muted playsinline width="100%"></video>

## Motivation

Onboard motorsport footage is everywhere β€” helmet cams, chest mounts,
dashcams β€” but the telemetry that makes it *interesting* usually is not.
Rental karts have no data logger you can access; indoor tracks deny GPS
entirely; action cameras often record IMU but no position. We wanted two
things: **measure speed** where no sensor can, and **add live telemetry
overlays** (speed, g-force, lap timing) to motorsport videos that were
recorded without any.

Getting there meant overcoming a chain of problems, each documented in this
card because they will bite anyone building something similar:

1. **Monocular scale ambiguity** β€” optical flow alone cannot distinguish
   "fast and far" from "slow and close". Solved with the RGB context branch,
   which only generalized once training spanned 7 distinct domains; with
   fewer, it memorized venues.
2. **No indoor ground truth** β€” GPS does not exist on an indoor kart track.
   Solved with weak supervision from the venue's official lap times
   (mean speed per lap = track length / lap time), aligned to video by
   detecting recurring track features.
3. **Lap memorization** β€” with only ~50 supervised laps, the GRU learned to
   recognize individual laps instead of estimating speed. Solved by
   augmenting the lap-supervision forward passes (flip, noise, occluders).
4. **Sequence length as a scale cue** β€” training with two window lengths
   taught the model to use context length itself as a signal; hence the two
   documented inference modes.
5. **Compressed dynamics** β€” regression pulls toward the mean (corners read
   fast, straights slow). Solved, when IMU data exists, by the physics
   fusion described below: measured longitudinal g-force drives the
   acceleration profile, centripetal physics anchors cornering speed, and
   track length locks the scale.
6. **The published track length is not the driven distance.** Our indoor
   supervision used `mean speed = track length / lap time` with the venue's
   advertised 1000 m. That figure is the **centreline**. A racing line cuts
   every corner by up to half the track width, so on a 13 m wide circuit
   with 33.1 rad of total turning it saves `6.5 m x 33.1 = 215 m` β€” the
   driven lap is ~790 m, and every label was 27% too fast. If you use
   lap-time weak supervision, measure the turning and subtract; do not trust
   the marketing number. Corner count is a good sanity check that you have
   the right layout (we measured 14, matching the venue's spec) while being
   completely independent of scale.

```
video ──► optical flow (motion field) ──┐
                                        β”œβ”€β”€β–Ί GRU ──► speed (m/s, per frame)
video ──► low-res context frame β”€β”€β”€β”€β”€β”€β”€β”€β”˜
          (tells the model the scene scale)
```

## Why the two branches?

Optical flow encodes how fast pixels move β€” which depends on **both** speed
and scene distance. 130 km/h on an open highway and 57 km/h between the close
walls of an indoor kart track can produce similar flow magnitudes. A
monocular model therefore cannot know the absolute scale from motion alone.
The **context branch** sees one low-resolution RGB frame and learns a
scene→scale mapping (indoor track / forest trail / highway), which lets a
single checkpoint work across very different environments. This only started
working once the training corpus spanned 7 distinct domains β€” with fewer
domains the branch memorized venues instead of generalizing.

## Quickstart

```bash
pip install torch opencv-python-headless numpy pandas huggingface_hub
```

Fetch the repo via `huggingface_hub` rather than downloading the weight
file directly β€” this is also what makes your download count toward the
model's stats on the Hub:

```python
from huggingface_hub import snapshot_download
local_dir = snapshot_download("x2q/speednet")
```

```python
import torch
from modeling_speednet import SpeedNet, predict_video

model = SpeedNet()
model.load_state_dict(torch.load("speednet_v5.pt", map_location="cuda"))

result = predict_video(model, "onboard_clip.mp4", device="cuda")
print(f"mean speed: {result['speed_smooth_mps'].mean() * 3.6:.1f} km/h")
```

Or run the bundled CLI example:

```bash
python example.py my_clip.mp4     # -> my_clip_speed.csv + my_clip_speed.png
```

`predict_video` handles everything: decoding, resizing to 320x180, Farneback
optical flow, frame-rate normalization, context frames, streaming GRU
inference and 1-second median smoothing.

## Critical usage notes

1. **Pick the right inference mode.** The model was trained with two window
   lengths, so sequence length acts as a scale cue:
   - *short mode* (default): hidden state reset every 16 frames β€” for roads,
     trails and open environments. Reproduces the GPS-domain metrics below.
   - *long-context mode* (`long_context=True`): hidden state carried across
     the clip β€” required for closed-course footage (indoor karting): short
     windows collapse indoor scale to roughly half the true value.
2. **Any frame rate works** β€” flow is normalized to a 15 fps convention
   internally (`raw_flow * fps / 15`). Feeding pre-computed flow without this
   normalization will mis-scale predictions.
3. **Forward-facing footage.** The model was trained on forward-facing
   onboard cameras. Rear/side cameras are out of distribution.
4. **Static occluders are fine.** Hoods, helmet edges and mounts were
   augmented during training; the model ignores zero-flow regions.

## Evaluation

Held-out recordings, never trained on. Two inference modes exist (sequence
length acts as a scale cue for this architecture); both are reported β€”
pick per the recommendation column. RMSE in m/s against GPS:

| Test clip | short windows | streaming | recommended |
|---|---|---|---|
| ATV, forest trail (unseen day) | **0.99** | 1.60 | short |
| Car, coastal Denmark (zero-shot recording) | 1.57 | **1.39** | streaming |
| Car, low-speed urban | 4.44 | **4.25** | streaming |
| Car, urban/rural | 7.82 | **6.00** | streaming |
| Car, motorway 100–135 km/h | 6.50 | **5.87** | streaming |
| Indoor go-kart (lap means vs official timing) | β€” | **3.70 km/h MAE** (bias βˆ’2.84) | streaming |

The indoor row compares mean predicted speed per lap against
`driven lap length / official lap time` on laps excluded from training,
using the corrected 790 m driven-lap length (see challenge 6 above β€”
the venue's advertised 1000 m is the centreline).

**Scale caveat:** all speeds carry a Β±5% scale uncertainty indoors. The
790 m driven-lap figure is bracketed by geometry (a racing line cannot be
shorter than ~785 m) and by flow evidence (which integrates to ~753 m);
the residual βˆ’2.8 km/h indoor bias reflects exactly this uncertainty.

## Training data

~5 hours of labeled video across 7 domains, all with real ground truth:

| Domain | Source | Speed signal |
|---|---|---|
| Indoor go-kart (2 races) | helmet cam, indoor kart circuit | official lap times (weak supervision: mean speed per lap over the 790 m driven racing line) |
| ATV / forest | action cam + GPS remote | 10 Hz GPS |
| Mini-loader / grass | action cam + GPS remote | 10 Hz GPS |
| Car / Denmark | action cam + GPS, dashcam | 10 Hz + 1 Hz GPS |
| Car / EU | [L2D](https://huggingface.co/datasets/yaak-ai/L2D) (Apache-2.0) | 10 Hz CAN/GNSS |
| Car / US | [comma2k19](https://github.com/commaai/comma2k19) (MIT) | 20 Hz pose velocities |

Training tricks that mattered: horizontal-flip and synthetic-occluder
augmentation on flow; **augmenting the lap-supervision forward passes**
(otherwise the GRU memorizes individual laps); context dropout (15%) and
photometric jitter on the context frame; per-recording train/val/test splits.

## Limitations

- **Protocol sensitivity**: sequence length is a scale cue (an artifact of
  two-length training); use the recommended mode per the evaluation table
  (short for trail/off-road, streaming for cars and closed courses).
  Length-randomized training removes this β€” validated on a companion
  track-specialized model β€” and is planned for the next release.
- **Indoor absolute scale is Β±5%**: supervision is lap time x driven lap
  length, and the driven racing line length is derived, not surveyed.
- **Compressed dynamics within a lap/segment**: like most regression models,
  predictions are pulled toward the mean β€” cornering speeds read slightly
  high, straight-line peaks slightly low. If your footage has IMU data, see
  the physics-fusion recipe below.
- Monocular scale is *learned*, not measured: radically new scene types
  (aircraft, boats, rear-facing cameras) will mis-scale.
- GPS ground truth is itself smoothed (~1 s); treat sub-second dynamics as
  indicative.
- Farneback flow degrades with heavy motion blur and very low light.

## Advanced: physics fusion (video + IMU)

If the recording also has IMU data, you can fix the compressed dynamics by
solving for the speed curve that best satisfies, jointly:

- `dv/dt = a_lon` β€” measured longitudinal acceleration (true accel/braking),
- `v = a_lat / Ο‰` β€” centripetal anchor in corners (|yaw rate| > 30Β°/s),
- `∫ v dt = track_length` per lap (if lap times are known),
- the video model as a weak level prior.

A few thousand Adam steps over the speed vector suffice. On our indoor race
this reproduces official lap averages to **0.07 km/h** with per-lap distance
789/790 m, while restoring realistic braking-into-corner /
accelerating-onto-straight dynamics.

## Model details

- ~1.0 M parameters (flow CNN 4 conv blocks β†’ 256-d; context CNN 3 blocks β†’
  64-d; 2-layer GRU, 192 hidden; linear head).
- Input: flow [T, 2, 72, 128] normalized by 8.0; context RGB [3, 72, 128]
  in [-0.5, 0.5]. Output Γ— 40.0 = m/s.
- Trained ~30 epochs, AdamW, Huber loss, cosine schedule, on a single
  NVIDIA GB10; checkpoint selected on combined GPS-validation + lap-holdout
  score.

## License

Weights and code: Apache-2.0. Trained only on own recordings and
permissively licensed public data (MIT / Apache-2.0).

## Citation

```
@misc{speednet2026,
  title  = {SpeedNet v5: ego-speed estimation from onboard video},
  author = {x2q},
  year   = {2026},
  url    = {https://huggingface.co/x2q/speednet}
}
```