TartanMatch: Towards Universal Dense Correspondence Across Modalities
Project page | Code | arXiv (coming soon)
Dense correspondence between any two views captured in RGB, depth, thermal, LiDAR, or event modalities. Given a source view and a target view (same or different modality), TartanMatch predicts, for every source pixel, the matching target pixel (optical flow) and the probability that the pixel is visible in the target (covisibility). One set of weights serves all 25 ordered modality pairs.
Files
| File | Description |
|---|---|
tartanmatch_v1.safetensors |
All-modality model, ViT-L/14 encoder, 420x560 inference resolution, 428.3M params. |
Usage
git clone https://github.com/castacks/tartanmatch.git
cd tartanmatch && pip install -e .
from tartanmatch import TartanMatch
model = TartanMatch.from_pretrained("theairlabcmu/TartanMatch", device="cuda")
out = model.predict(source, "rgb", target, "event", event_resolution=(640, 640))
flow = out.flow[0] # (2, H, W): target pixel = source pixel + flow
covisibility = out.covisibility[0] # (H, W) in [0, 1]
See the repository README for the input format of each modality.
License
The code is released under the BSD-3-Clause license. The model weights inherit the licenses of the training datasets and are released under CC BY-NC-SA 4.0; they may not be used for commercial purposes.
Citation
@article{kwon2026tartanmatch,
title={TartanMatch: Towards Universal Dense Correspondence Across Modalities},
author={Kwon, Hyeokjoon and Cai, Jiting and Li, Ruogu and Bhandari, Kritan and Hemkumar, Geethika and Maheshwari, Parv and Qiu, Yuheng and Zhang, Yuchen and Scherer, Sebastian and Wang, Wenshan},
journal={arXiv preprint},
year={2026}
}
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support