File size: 7,225 Bytes
cb1e428 85a2899 cb1e428 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 | ---
license: apache-2.0
pipeline_tag: video-classification
tags:
- speed-estimation
- ego-motion
- optical-flow
- localization
- video
- regression
- pytorch
- motorsport
- karting
- telemetry
- imu
- dji
- gopro
- track
---
# KartNet v3 β track-specialized telemetry from onboard karting video
KartNet turns onboard video from **one specific indoor kart circuit**
(Racehall Aarhus, Denmark) into telemetry that no sensor can record there β
GPS does not work indoors:
- **speed** (m/s, per frame at 10 Hz)
- **track position** β where on the lap the kart is, as a cyclic phase;
lap times fall out of the wrap points
- **yaw rate**
If the recording carries an IMU (DJI and GoPro action cameras embed one),
an optional input branch uses it and measurably improves accuracy. Without
sensors the model runs on video alone β one checkpoint serves both cases.
It is the specialized sibling of [SpeedNet](https://huggingface.co/x2q/speednet),
our general ego-speed model. Where SpeedNet must work on any road, KartNet
exploits the fact that every clip shows the same 790 m circuit β which buys
per-frame position and much tighter speed.
## Demo
Held-out race (never trained on), same venue, video + IMU. The speed
readout and the lap-position bar are model outputs; there is no GPS in the
loop anywhere:
<video src="https://huggingface.co/x2q/kartnet/resolve/main/assets/demo.mp4" controls muted playsinline width="100%"></video>
## How it was trained β labels without any ground truth sensor
There is no GPS indoors and rental karts expose no data bus, so the
training labels themselves had to be manufactured from video:
1. **Corpus self-localization.** 29 hours of onboard footage from the same
circuit (76 recordings: different drivers, karts, cameras, seasons,
lighting) were all aligned to a single canonical reference lap with a
cyclic Viterbi over frame-signature similarity. Every frame gets a
position `u` around the lap.
2. **Lap times for free.** Laps are the wrap points of `u`. Validated
against lap times published in video titles (median error **0.22 s**)
and against the venue's official SMS-Timing for one race
(27/27 laps, median error **0.19 s**) β data the aligner never saw.
3. **Metres from physics.** The advertised track length (1000 m) is the
centreline; a racing line cuts corners, and three independent estimates
(corner-cut geometry, integrated flow speed, the karts' rated top speed)
agree on **~790 m driven per lap**. Speed is then `790 m x du/dt`.
4. **Result:** 809 laps / 748k frames with per-frame speed, position and
yaw labels β from video alone.
## Evaluation
Held-out **whole videos** (unseen sessions, drivers, karts, lighting):
| protocol | speed MAE | position (median) |
|---|---|---|
| 64-frame windows | 2.2 km/h | 1.3 m |
| whole clip, streaming | 2.5 km/h | 7.1 m |
Length-invariance (the model was trained with randomized sequence lengths,
so window size is a throughput knob, not an accuracy mode):
| window | 16 | 32 | 64 | 128 |
|---|---|---|---|---|
| speed MAE (km/h) | 1.95 | 1.70 | 1.54 | 1.49 |
| position (m) | 1.9 | 1.4 | 1.0 | 0.8 |
**Does it measure, or does it recall?** A model that can localize itself
on a known track could fake speed by recalling what karts usually do at
that corner. Ablation says no:
| inputs | speed MAE |
|---|---|
| flow + appearance (full) | 2.6 km/h |
| flow only | 3.4 km/h |
| appearance only (no motion!) | 8.0 km/h |
Appearance refines; optical flow carries the measurement.
**Cross-camera holdout** β a different camera, mount and month
(a DJI helmet camera never seen in training except via two other sessions
from the same device), judged against official race timing:
| metric | video only | video + IMU |
|---|---|---|
| per-lap mean speed MAE (all 27 laps) | 1.9 km/h | **1.1 km/h** |
| per-lap mean speed bias | +1.9 km/h | +0.8 km/h |
| position (median) | 3.5 m | 3.4 m |
## Limitations β read before use
- **One venue, one layout.** This model is Racehall Aarhus in its standard
direction, by design. Other tracks, or this track in reverse layout,
produce garbage positions (our own reverse-layout footage is correctly
rejected by the label pipeline, but the model has no such guard).
- **Lap times from the position head are unreliable on unfamiliar
cameras.** On the cross-camera holdout the phase runs fast and reads lap
times ~5 s short despite good median position. If you need lap timing on
a new camera setup, derive it from signature alignment against a
reference lap instead β or fine-tune.
- **Absolute speed scale is Β±5%.** The 790 m driven-lap length is derived,
not surveyed.
- **Frame-level dynamics carry label noise.** Within-lap timing of
accel/brake transitions inherits smoothing from the auto-labels
(~0.7 s); per-lap statistics are much tighter than per-frame values.
- The vibration IMU channel needs the full-rate (β₯50 Hz) accelerometer
stream; feeding downsampled IMU zeroes its benefit.
## Model details
- 1.67 M parameters: flow CNN (4 blocks β 256-d) + appearance CNN
(4 blocks β 128-d) + IMU MLP (5 channels + presence flag β 32-d)
β 2-layer GRU (256) β three linear heads.
- Inputs at 10 Hz: Farneback flow 128Γ72 (decoded at 30 fps, every 3rd
frame kept), grayscale 128Γ72, optional IMU
(g_lon, g_lat, yaw/60, |a|β1, vibration RMS).
- Camera-robustness training: random zoom/translation with consistently
scaled flow vectors, resolution degradation, photometric jitter,
appearance dropout (25%), IMU dropout (40%), randomized sequence
lengths (16/32/64/128).
- Training data is **not** redistributed: most of the corpus is publicly
posted third-party onboard footage, used for training only. This
repository contains weights, code and documentation β no footage.
## Quickstart
Fetch the repo via `huggingface_hub` rather than downloading the weight
file directly β this is also what makes your download count toward the
model's stats on the Hub:
```python
from huggingface_hub import snapshot_download
local_dir = snapshot_download("x2q/kartnet")
```
```python
import torch
from modeling_kartnet import KartNet, extract_features, predict
model = KartNet()
model.load_state_dict(torch.load("kartnet_v3.pt", map_location="cuda"))
feats = extract_features("onboard.mp4") # ffmpeg + OpenCV, ~2x realtime
out = predict(model, feats, device="cuda")
out["speed_mps"] # per-frame speed, 10 Hz
out["u"] # lap phase in [0,1)
out["lap_times_s"] # from phase wrap points
```
With IMU (example for DJI embedded telemetry, any source works if
timestamps are on the video clock):
```python
from modeling_kartnet import imu_features
imu = imu_features(len(feats["gray"]), t_imu, accel_xyz_g, yaw_rate_dps,
gforce_lon, gforce_lat)
out = predict(model, feats, imu=imu, device="cuda")
```
Or just: `python example.py my_clip.mp4`
## License
Apache-2.0 (weights and code).
## Citation
```bibtex
@misc{kartnet2026,
title = {KartNet v3: track-specialized telemetry from onboard karting video},
author = {x2q},
year = {2026},
url = {https://huggingface.co/x2q/kartnet}
}
```
|