CalfVO
Monocular visual odometry without calibration or test-time optimization.
Project page · Paper · Code
Given uncalibrated video, CalfVO predicts a metric camera trajectory in a single forward pass at 53 FPS, with no camera intrinsics, no bundle adjustment and no loop closure. A transformer regresses relative camera poses over overlapping image windows together with separate rotation and translation confidences, supervised by camera poses alone, and a confidence-weighted module aggregates the overlapping predictions into one trajectory. Scale is recovered from learned priors, accurately enough that the trajectories are evaluated without any alignment to the ground truth.
⚠️ The two checkpoints are under different licences
This repository holds two models that differ only in their frozen image encoder. They are not under the same terms, because their encoders are not.
| File | Encoder | Licence | Use |
|---|---|---|---|
calfvo_croco.pth |
CroCo v2, as distributed with DUSt3R | CC BY-NC-SA 4.0 | Non-commercial, share-alike |
calfvo_dinov2.pth |
DINOv2 ViT-L/14 | Apache-2.0 | Commercial use permitted |
calfvo_croco.pth embeds Naver's CroCo v2 encoder weights, which are CC BY-NC-SA 4.0.
Those terms are not ours to relax, so the file is non-commercial and share-alike
even though the surrounding code is MIT.
calfvo_dinov2.pth embeds only Meta's DINOv2 ViT-L/14, which is Apache-2.0. Combined
with the MIT training code, that file is usable commercially.
If you need a commercially usable model, take calfvo_dinov2.pth. It is the weaker of
the two on most metrics (see below).
Which one to cite
The paper reports calfvo_croco.pth. If you are comparing against CalfVO, that is the
model to use.
Results
Both are step 57000 of the same 57000-step schedule on the same data. rpe_trans and
rpe_rot are evo's relative pose error at a one-frame delta over all pairs.
| Test set | croco rpe_trans | dinov2 rpe_trans | croco rpe_rot | dinov2 rpe_rot |
|---|---|---|---|---|
| ScanNet (97 seq) | 0.00538 | 0.00609 | 0.16340 | 0.20100 |
| TUM RGB-D (9 seq) | 0.00524 | 0.00541 | 0.33554 | 0.35930 |
| TartanAir v1 (18 seq) | 0.10273 | 0.09194 | 0.28150 | 0.39616 |
| KITTI 09-10 (2 seq) | 0.16199 | 0.18672 | 0.12306 | 0.21414 |
| EuRoC MAV (11 seq) | 0.02066 | 0.01772 | 0.31326 | 0.25842 |
CroCo wins rotation on four of five sets, by 7% to 74%. Translation is mixed: DINOv2 is better on TartanAir and EuRoC. CroCo v2 is pretrained on cross-view completion, a geometric correspondence task, which is a plausible reason it recovers inter-frame rotation better than DINOv2's semantic features.
Usage
Clone the code, then:
import torch
from src.model import get_model
from src.utils import io_utils
config = io_utils.read_yaml_file("configs/calfvo_croco.yaml")
config["model"]["num_views"] = config["num_views"]
model = get_model(config["model"])
ckpt = torch.load("calfvo_croco.pth", map_location="cpu")
model.load_state_dict(ckpt["model_state_dict"])
model.eval()
Use configs/calfvo_dinov2.yaml with calfvo_dinov2.pth. The two are not
interchangeable: their positional tables differ in size, because CroCo emits a 14x14
token grid at 224 px and DINOv2 emits 16x16, so loading one into the other's config
raises an error rather than failing quietly.
Each file is self contained, including the frozen encoder, so nothing is fetched at load time. The optimizer state is stripped, so these are for inference and not for resuming training.
To run the model over a video and get a trajectory:
python scripts/demo_run_vo.py \
--config configs/calfvo_croco.yaml \
--checkpoint calfvo_croco.pth \
--video /path/to/clip.mp4 \
--out output/demo/vo
Input
Eight-frame RGB windows at 224x224. The transform centre-crops each frame to its shortest edge before resizing, so aspect ratio is handled for you and no intrinsics are needed. The model was trained on consecutive frames of roughly 30 fps footage and predicts the motion between the frames it is given, so subsample a faster-moving clip rather than feeding it every frame.
Training data
ScanNet, TartanAir v2, KITTI odometry 00-08 and Virtual KITTI 2, mixed 32/36/16/16%. Held out: ScanNet test, TartanAir v1, TUM RGB-D, KITTI 09-10 and EuRoC MAV.
Citation
@misc{yugay2026calfvo,
title={Monocular Visual Odometry without Calibration or Test-time Optimization},
author={Vladimir Yugay and Duy-Kien Nguyen and Theo Gevers and Cees G. M. Snoek and Martin R. Oswald},
year={2026},
eprint={2510.03348},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2510.03348},
}