CalfVO

Monocular visual odometry without calibration or test-time optimization.

Project page · Paper · Code

Given uncalibrated video, CalfVO predicts a metric camera trajectory in a single forward pass at 53 FPS, with no camera intrinsics, no bundle adjustment and no loop closure. A transformer regresses relative camera poses over overlapping image windows together with separate rotation and translation confidences, supervised by camera poses alone, and a confidence-weighted module aggregates the overlapping predictions into one trajectory. Scale is recovered from learned priors, accurately enough that the trajectories are evaluated without any alignment to the ground truth.

⚠️ The two checkpoints are under different licences

This repository holds two models that differ only in their frozen image encoder. They are not under the same terms, because their encoders are not.

File Encoder Licence Use
calfvo_croco.pth CroCo v2, as distributed with DUSt3R CC BY-NC-SA 4.0 Non-commercial, share-alike
calfvo_dinov2.pth DINOv2 ViT-L/14 Apache-2.0 Commercial use permitted

calfvo_croco.pth embeds Naver's CroCo v2 encoder weights, which are CC BY-NC-SA 4.0. Those terms are not ours to relax, so the file is non-commercial and share-alike even though the surrounding code is MIT.

calfvo_dinov2.pth embeds only Meta's DINOv2 ViT-L/14, which is Apache-2.0. Combined with the MIT training code, that file is usable commercially.

If you need a commercially usable model, take calfvo_dinov2.pth. It is the weaker of the two on most metrics (see below).

Which one to cite

The paper reports calfvo_croco.pth. If you are comparing against CalfVO, that is the model to use.

Results

Both are step 57000 of the same 57000-step schedule on the same data. rpe_trans and rpe_rot are evo's relative pose error at a one-frame delta over all pairs.

Test set croco rpe_trans dinov2 rpe_trans croco rpe_rot dinov2 rpe_rot
ScanNet (97 seq) 0.00538 0.00609 0.16340 0.20100
TUM RGB-D (9 seq) 0.00524 0.00541 0.33554 0.35930
TartanAir v1 (18 seq) 0.10273 0.09194 0.28150 0.39616
KITTI 09-10 (2 seq) 0.16199 0.18672 0.12306 0.21414
EuRoC MAV (11 seq) 0.02066 0.01772 0.31326 0.25842

CroCo wins rotation on four of five sets, by 7% to 74%. Translation is mixed: DINOv2 is better on TartanAir and EuRoC. CroCo v2 is pretrained on cross-view completion, a geometric correspondence task, which is a plausible reason it recovers inter-frame rotation better than DINOv2's semantic features.

Usage

Clone the code, then:

import torch
from src.model import get_model
from src.utils import io_utils

config = io_utils.read_yaml_file("configs/calfvo_croco.yaml")
config["model"]["num_views"] = config["num_views"]

model = get_model(config["model"])
ckpt = torch.load("calfvo_croco.pth", map_location="cpu")
model.load_state_dict(ckpt["model_state_dict"])
model.eval()

Use configs/calfvo_dinov2.yaml with calfvo_dinov2.pth. The two are not interchangeable: their positional tables differ in size, because CroCo emits a 14x14 token grid at 224 px and DINOv2 emits 16x16, so loading one into the other's config raises an error rather than failing quietly.

Each file is self contained, including the frozen encoder, so nothing is fetched at load time. The optimizer state is stripped, so these are for inference and not for resuming training.

To run the model over a video and get a trajectory:

python scripts/demo_run_vo.py \
    --config configs/calfvo_croco.yaml \
    --checkpoint calfvo_croco.pth \
    --video /path/to/clip.mp4 \
    --out output/demo/vo

Input

Eight-frame RGB windows at 224x224. The transform centre-crops each frame to its shortest edge before resizing, so aspect ratio is handled for you and no intrinsics are needed. The model was trained on consecutive frames of roughly 30 fps footage and predicts the motion between the frames it is given, so subsample a faster-moving clip rather than feeding it every frame.

Training data

ScanNet, TartanAir v2, KITTI odometry 00-08 and Virtual KITTI 2, mixed 32/36/16/16%. Held out: ScanNet test, TartanAir v1, TUM RGB-D, KITTI 09-10 and EuRoC MAV.

Citation

@misc{yugay2026calfvo,
      title={Monocular Visual Odometry without Calibration or Test-time Optimization},
      author={Vladimir Yugay and Duy-Kien Nguyen and Theo Gevers and Cees G. M. Snoek and Martin R. Oswald},
      year={2026},
      eprint={2510.03348},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2510.03348},
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Paper for voviktyl/CalfVO