MoSE3: Learning World-Space SE(3) at Every Pixel

Paper | Project page | Code

MoSE3 is a feed-forward model that predicts world-space SE(3) motion for every pixel of a monocular RGB video, generalizing across rigid, articulated and deformable objects.

Usage

from mose3.models.mose3 import MoSE3

model = MoSE3.from_pretrained("mose3-tracker/MoSE3", strict=True).cuda().eval()

The code repository has the full inference and visualization scripts.

License

The weights are released under CC BY-NC 4.0 (non-commercial). They include the weights of π³, which are released under the same license.

Citation

@inproceedings{cheng2026mose3,
  title     = {{MoSE3}: Learning World-Space {SE(3)} at Every Pixel},
  author    = {Cheng, Jiahuan and Li, Zhiyi and Xia, Tian and
               Cai, Ruojin and Du, Yilun and Wang, Qianqian},
  booktitle = {Advances in Neural Information Processing Systems (NeurIPS)},
  year      = {2026}
}
Downloads last month
126
Safetensors
Model size
2B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Paper for mose3-tracker/MoSE3