--- license: mit library_name: pytorch pipeline_tag: image-to-3d tags: - structure-from-motion - multi-view - pose-estimation - 3d-reconstruction - diffusion - computer-vision --- # DiffusionSfM: Predicting Structure and Motion via Ray Origin and Endpoint Diffusion This repository hosts the pretrained checkpoint for **DiffusionSfM**, a transformer-based denoising diffusion model that directly infers 3D scene geometry and camera poses from multi-view images. The model parameterizes scene geometry and cameras as pixel-wise ray origins and endpoints in a global frame and learns a denoising diffusion model to predict them from multi-view inputs. ## Paper **DiffusionSfM: Predicting Structure and Motion via Ray Origin and Endpoint Diffusion** Qitao Zhao, Amy Lin, Jeff Tan, Jason Y. Zhang, Deva Ramanan, Shubham Tulsiani Accepted to CVPR 2025. For more details, see the paper: https://huggingface.co/papers/2505.05473 - arXiv: https://arxiv.org/abs/2505.05473 - Project page: https://qitaozhao.github.io/DiffusionSfM - Code: https://github.com/QitaoZhao/DiffusionSfM - Interactive demo (Gradio Space): https://huggingface.co/spaces/qitaoz/DiffusionSfM ## Model description DiffusionSfM is a data-driven Structure-from-Motion (SfM) approach that replaces the conventional two-stage pipeline (pairwise reasoning + global optimization) with a single end-to-end multi-view reasoning model. Given a set of multi-view images, it predicts, per pixel, the ray origin and ray endpoint in a global coordinate frame. The endpoints give the 3D scene geometry (point structure), while the ray origins and endpoints together recover the camera poses. A transformer-based denoising diffusion model performs this prediction directly from multi-view image features. Specialized training mechanisms handle missing data and unbounded scene coordinates. The model naturally models uncertainty through the diffusion process and outperforms classical and prior learning-based SfM approaches on both synthetic and real datasets. This repository contains the released checkpoint `ckpt_00800000.pth` (a raw PyTorch checkpoint). It is intended to be used together with the official codebase at https://github.com/QitaoZhao/DiffusionSfM. ## Intended uses - Multi-view Structure-from-Motion (camera pose estimation + 3D structure recovery) from a set of posed or unposed images. - Research on learned SfM, multi-view geometry, and diffusion models for 3D vision. ## How to use The checkpoint is a raw PyTorch `.pth` file and is loaded by the official codebase. It is not directly loadable with `transformers` or `diffusers` `from_pretrained` APIs; use the project's own loading utilities. 1. Clone the official codebase: ```bash git clone https://github.com/QitaoZhao/DiffusionSfM.git cd DiffusionSfM ``` 2. Download the checkpoint from this repo: ```python from huggingface_hub import hf_hub_download ckpt_path = hf_hub_download(repo_id="qitaoz/DiffusionSfM", filename="ckpt_00800000.pth") ``` 3. Follow the installation and inference instructions in the [official repository](https://github.com/QitaoZhao/DiffusionSfM) (environment setup with `torch`, `torchvision`, `pytorch3d`, and the project's `diffusionsfm` package). You can also try the online [Gradio demo](https://huggingface.co/spaces/qitaoz/DiffusionSfM) without any local setup. ## License This model is released under the MIT License (see the [LICENSE](https://github.com/QitaoZhao/DiffusionSfM/blob/main/LICENSE) in the official repository). ## Citation If you use this model, please cite: ```bibtex @inproceedings{zhao2025diffusionsfm, title={DiffusionSfM: Predicting Structure and Motion via Ray Origin and Endpoint Diffusion}, author={Qitao Zhao and Amy Lin and Jeff Tan and Jason Y. Zhang and Deva Ramanan and Shubham Tulsiani}, booktitle={CVPR}, year={2025} } ```