kartnet / README.md
x2q's picture
quickstart: recommend snapshot_download so downloads register
85a2899 verified
|
Raw
History Blame Contribute Delete
7.23 kB
---
license: apache-2.0
pipeline_tag: video-classification
tags:
- speed-estimation
- ego-motion
- optical-flow
- localization
- video
- regression
- pytorch
- motorsport
- karting
- telemetry
- imu
- dji
- gopro
- track
---
# KartNet v3 β€” track-specialized telemetry from onboard karting video
KartNet turns onboard video from **one specific indoor kart circuit**
(Racehall Aarhus, Denmark) into telemetry that no sensor can record there β€”
GPS does not work indoors:
- **speed** (m/s, per frame at 10 Hz)
- **track position** β€” where on the lap the kart is, as a cyclic phase;
lap times fall out of the wrap points
- **yaw rate**
If the recording carries an IMU (DJI and GoPro action cameras embed one),
an optional input branch uses it and measurably improves accuracy. Without
sensors the model runs on video alone β€” one checkpoint serves both cases.
It is the specialized sibling of [SpeedNet](https://huggingface.co/x2q/speednet),
our general ego-speed model. Where SpeedNet must work on any road, KartNet
exploits the fact that every clip shows the same 790 m circuit β€” which buys
per-frame position and much tighter speed.
## Demo
Held-out race (never trained on), same venue, video + IMU. The speed
readout and the lap-position bar are model outputs; there is no GPS in the
loop anywhere:
<video src="https://huggingface.co/x2q/kartnet/resolve/main/assets/demo.mp4" controls muted playsinline width="100%"></video>
## How it was trained β€” labels without any ground truth sensor
There is no GPS indoors and rental karts expose no data bus, so the
training labels themselves had to be manufactured from video:
1. **Corpus self-localization.** 29 hours of onboard footage from the same
circuit (76 recordings: different drivers, karts, cameras, seasons,
lighting) were all aligned to a single canonical reference lap with a
cyclic Viterbi over frame-signature similarity. Every frame gets a
position `u` around the lap.
2. **Lap times for free.** Laps are the wrap points of `u`. Validated
against lap times published in video titles (median error **0.22 s**)
and against the venue's official SMS-Timing for one race
(27/27 laps, median error **0.19 s**) β€” data the aligner never saw.
3. **Metres from physics.** The advertised track length (1000 m) is the
centreline; a racing line cuts corners, and three independent estimates
(corner-cut geometry, integrated flow speed, the karts' rated top speed)
agree on **~790 m driven per lap**. Speed is then `790 m x du/dt`.
4. **Result:** 809 laps / 748k frames with per-frame speed, position and
yaw labels β€” from video alone.
## Evaluation
Held-out **whole videos** (unseen sessions, drivers, karts, lighting):
| protocol | speed MAE | position (median) |
|---|---|---|
| 64-frame windows | 2.2 km/h | 1.3 m |
| whole clip, streaming | 2.5 km/h | 7.1 m |
Length-invariance (the model was trained with randomized sequence lengths,
so window size is a throughput knob, not an accuracy mode):
| window | 16 | 32 | 64 | 128 |
|---|---|---|---|---|
| speed MAE (km/h) | 1.95 | 1.70 | 1.54 | 1.49 |
| position (m) | 1.9 | 1.4 | 1.0 | 0.8 |
**Does it measure, or does it recall?** A model that can localize itself
on a known track could fake speed by recalling what karts usually do at
that corner. Ablation says no:
| inputs | speed MAE |
|---|---|
| flow + appearance (full) | 2.6 km/h |
| flow only | 3.4 km/h |
| appearance only (no motion!) | 8.0 km/h |
Appearance refines; optical flow carries the measurement.
**Cross-camera holdout** β€” a different camera, mount and month
(a DJI helmet camera never seen in training except via two other sessions
from the same device), judged against official race timing:
| metric | video only | video + IMU |
|---|---|---|
| per-lap mean speed MAE (all 27 laps) | 1.9 km/h | **1.1 km/h** |
| per-lap mean speed bias | +1.9 km/h | +0.8 km/h |
| position (median) | 3.5 m | 3.4 m |
## Limitations β€” read before use
- **One venue, one layout.** This model is Racehall Aarhus in its standard
direction, by design. Other tracks, or this track in reverse layout,
produce garbage positions (our own reverse-layout footage is correctly
rejected by the label pipeline, but the model has no such guard).
- **Lap times from the position head are unreliable on unfamiliar
cameras.** On the cross-camera holdout the phase runs fast and reads lap
times ~5 s short despite good median position. If you need lap timing on
a new camera setup, derive it from signature alignment against a
reference lap instead β€” or fine-tune.
- **Absolute speed scale is Β±5%.** The 790 m driven-lap length is derived,
not surveyed.
- **Frame-level dynamics carry label noise.** Within-lap timing of
accel/brake transitions inherits smoothing from the auto-labels
(~0.7 s); per-lap statistics are much tighter than per-frame values.
- The vibration IMU channel needs the full-rate (β‰₯50 Hz) accelerometer
stream; feeding downsampled IMU zeroes its benefit.
## Model details
- 1.67 M parameters: flow CNN (4 blocks β†’ 256-d) + appearance CNN
(4 blocks β†’ 128-d) + IMU MLP (5 channels + presence flag β†’ 32-d)
β†’ 2-layer GRU (256) β†’ three linear heads.
- Inputs at 10 Hz: Farneback flow 128Γ—72 (decoded at 30 fps, every 3rd
frame kept), grayscale 128Γ—72, optional IMU
(g_lon, g_lat, yaw/60, |a|βˆ’1, vibration RMS).
- Camera-robustness training: random zoom/translation with consistently
scaled flow vectors, resolution degradation, photometric jitter,
appearance dropout (25%), IMU dropout (40%), randomized sequence
lengths (16/32/64/128).
- Training data is **not** redistributed: most of the corpus is publicly
posted third-party onboard footage, used for training only. This
repository contains weights, code and documentation β€” no footage.
## Quickstart
Fetch the repo via `huggingface_hub` rather than downloading the weight
file directly β€” this is also what makes your download count toward the
model's stats on the Hub:
```python
from huggingface_hub import snapshot_download
local_dir = snapshot_download("x2q/kartnet")
```
```python
import torch
from modeling_kartnet import KartNet, extract_features, predict
model = KartNet()
model.load_state_dict(torch.load("kartnet_v3.pt", map_location="cuda"))
feats = extract_features("onboard.mp4") # ffmpeg + OpenCV, ~2x realtime
out = predict(model, feats, device="cuda")
out["speed_mps"] # per-frame speed, 10 Hz
out["u"] # lap phase in [0,1)
out["lap_times_s"] # from phase wrap points
```
With IMU (example for DJI embedded telemetry, any source works if
timestamps are on the video clock):
```python
from modeling_kartnet import imu_features
imu = imu_features(len(feats["gray"]), t_imu, accel_xyz_g, yaw_rate_dps,
gforce_lon, gforce_lat)
out = predict(model, feats, imu=imu, device="cuda")
```
Or just: `python example.py my_clip.mp4`
## License
Apache-2.0 (weights and code).
## Citation
```bibtex
@misc{kartnet2026,
title = {KartNet v3: track-specialized telemetry from onboard karting video},
author = {x2q},
year = {2026},
url = {https://huggingface.co/x2q/kartnet}
}
```