TartanMatch: Towards Universal Dense Correspondence Across Modalities

Project page | Code | arXiv (coming soon)

Dense correspondence between any two views captured in RGB, depth, thermal, LiDAR, or event modalities. Given a source view and a target view (same or different modality), TartanMatch predicts, for every source pixel, the matching target pixel (optical flow) and the probability that the pixel is visible in the target (covisibility). One set of weights serves all 25 ordered modality pairs.

Files

File Description
tartanmatch_v1.safetensors All-modality model, ViT-L/14 encoder, 420x560 inference resolution, 428.3M params.

Usage

git clone https://github.com/castacks/tartanmatch.git
cd tartanmatch && pip install -e .
from tartanmatch import TartanMatch

model = TartanMatch.from_pretrained("theairlabcmu/TartanMatch", device="cuda")
out = model.predict(source, "rgb", target, "event", event_resolution=(640, 640))
flow = out.flow[0]                 # (2, H, W): target pixel = source pixel + flow
covisibility = out.covisibility[0] # (H, W) in [0, 1]

See the repository README for the input format of each modality.

License

The code is released under the BSD-3-Clause license. The model weights inherit the licenses of the training datasets and are released under CC BY-NC-SA 4.0; they may not be used for commercial purposes.

Citation

@article{kwon2026tartanmatch,
 title={TartanMatch: Towards Universal Dense Correspondence Across Modalities},
 author={Kwon, Hyeokjoon and Cai, Jiting and Li, Ruogu and Bhandari, Kritan and Hemkumar, Geethika and Maheshwari, Parv and Qiu, Yuheng and Zhang, Yuchen and Scherer, Sebastian and Wang, Wenshan},
 journal={arXiv preprint},
 year={2026}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support