pipeline_tag: other
TrackCraft3R: Repurposing Video Diffusion Transformers for Dense 3D Tracking
TrackCraft3R is the first method to repurpose a pre-trained video diffusion transformer (specifically Wan2.1-T2V-1.3B) as a single-pass dense 3D tracker. Given a monocular video along with its predicted depth and camera, TrackCraft3R predicts dense 3D trajectories in a single forward pass.
Paper | Project Page | GitHub
Abstract
Dense 3D tracking from monocular video is fundamental to dynamic scene understanding. While recent 3D foundation models provide reliable per-frame geometry, recovering object motion in this geometry remains challenging. TrackCraft3R repurposes a pre-trained video diffusion transformer (video DiT) to capture rich spatio-temporal priors from internet-scale videos. By employing a dual-latent representation and temporal RoPE alignment, it converts the generative frame-anchored paradigm into a reference-anchored tracking formulation.
Sample Usage
To run inference on a video, you first need to follow the environment setup and preprocessing steps (extracting depth and camera information) detailed in the official repository. Once the input is prepared in .npz format, you can run inference as follows:
MODELSCOPE_CACHE=./checkpoints/wan_models \
python scripts/inference_user_video.py \
--checkpoint_path ./checkpoints/trackcraft3r/model.safetensors \
--input_npz ./breakdance_user.npz \
--output_npz ./breakdance_dense.npz \
--num_frames 12 --frame_stride 5
The output contains 3D tracks and visibility information, which can be visualized using the visualize_dense.py script provided in the repository.
Citation
@article{nam2026trackcraft3r,
title={TrackCraft3R: Repurposing Video Diffusion Transformers for Dense 3D Tracking},
author={Nam, Jisu and Koo, Jahyeok and Son, Soowon and Jung, Jaewoo and An, Honggyu and Hur, Junhwa and Kim, Seungryong},
journal={arXiv preprint arXiv:2605.12587},
year={2026}
}