checkpoint / README.md
nielsr's picture
nielsr HF Staff
Add model card for TrackCraft3R
31b2640 verified
|
Raw
History Blame
2.22 kB
---
pipeline_tag: other
---
# TrackCraft3R: Repurposing Video Diffusion Transformers for Dense 3D Tracking
TrackCraft3R is the first method to repurpose a pre-trained video diffusion transformer (specifically Wan2.1-T2V-1.3B) as a single-pass dense 3D tracker. Given a monocular video along with its predicted depth and camera, TrackCraft3R predicts dense 3D trajectories in a single forward pass.
[**Paper**](https://huggingface.co/papers/2605.12587) | [**Project Page**](https://cvlab-kaist.github.io/TrackCraft3r) | [**GitHub**](https://github.com/cvlab-kaist/TrackCraft3r)
## Abstract
Dense 3D tracking from monocular video is fundamental to dynamic scene understanding. While recent 3D foundation models provide reliable per-frame geometry, recovering object motion in this geometry remains challenging. TrackCraft3R repurposes a pre-trained video diffusion transformer (video DiT) to capture rich spatio-temporal priors from internet-scale videos. By employing a dual-latent representation and temporal RoPE alignment, it converts the generative frame-anchored paradigm into a reference-anchored tracking formulation.
## Sample Usage
To run inference on a video, you first need to follow the environment setup and preprocessing steps (extracting depth and camera information) detailed in the [official repository](https://github.com/cvlab-kaist/TrackCraft3r). Once the input is prepared in `.npz` format, you can run inference as follows:
```bash
MODELSCOPE_CACHE=./checkpoints/wan_models \
python scripts/inference_user_video.py \
--checkpoint_path ./checkpoints/trackcraft3r/model.safetensors \
--input_npz ./breakdance_user.npz \
--output_npz ./breakdance_dense.npz \
--num_frames 12 --frame_stride 5
```
The output contains 3D tracks and visibility information, which can be visualized using the `visualize_dense.py` script provided in the repository.
## Citation
```bibtex
@article{nam2026trackcraft3r,
title={TrackCraft3R: Repurposing Video Diffusion Transformers for Dense 3D Tracking},
author={Nam, Jisu and Koo, Jahyeok and Son, Soowon and Jung, Jaewoo and An, Honggyu and Hur, Junhwa and Kim, Seungryong},
journal={arXiv preprint arXiv:2605.12587},
year={2026}
}
```