VGGT-Diff: Visual Geometry Meets Diffusion for Sparse-View Novel View Synthesis
Abstract
We present VGGT-Diff, a geometry-routed multi-view diffusion model for sparse-view novel view synthesis. Existing novel view synthesis (NVS) methods face a fundamental trade-off: reconstruction-based approaches preserve observed geometry but struggle to synthesize unseen regions, while diffusion-based methods provide strong generative priors yet rely on implicit source-to-query correspondence. VGGT-Diff bridges these regimes by routing visual geometry latents from VGGT-Ω into a pretrained video diffusion model. Each visual token is associated with a 3D point and confidence, then transformed into query-aligned latent conditions through a confidence-aware Visual Geometry Router (VGR) that preserves front and back surface evidence. These conditions guide joint target-view denoising, while Point-Track Residual Consistency (PTRC) regularizes predicted-clean residuals along reliable 3D tracks, improving multi-view stability. We further introduce robust geometry conditioning, combining training-time regularization with inference-time guidance for improved robustness. Experiments show competitive or state-of-the-art performance across interpolation and extrapolation under different viewpoint difficulties. Our code is available at https://github.com/chenkangjie1123/VGGT-Diff.
Community
VGGT-Diff is a geometry-routed multi-view diffusion model for sparse-view novel view synthesis from six input images. Its first key innovation, the confidence-aware Visual Geometry Router (VGR), transforms VGGT-Ω features into query-aligned geometric conditions while preserving both front- and back-surface evidence. Its second innovation, Point-Track Residual Consistency (PTRC), regularizes denoising residuals along reliable 3D tracks to improve cross-view geometric consistency.
Despite being trained on only 1K scenes with a relatively limited training budget, VGGT-Diff already delivers competitive or state-of-the-art novel-view synthesis performance. It achieves state-of-the-art PSNR / LPIPS in different viewpoint-difficulty settings, with particularly strong results on challenging mid- and far-range interpolation and extrapolation views. Further scaling in both training data and optimization steps could continue improving visual quality and cross-view consistency.
VGGT-Diff supports both pose-aware inference with known camera parameters and pose-free inference directly from six RGB images. The current effective 21K checkpoint—initialized from a 20K half-resolution checkpoint and adapted with only 1K additional full-resolution training steps—already generates high-quality 480p, 80-frame continuous camera-trajectory videos. The model is still being actively trained, and we plan to release multiple fully trained checkpoints to the community later.
Get this paper in your agent:
hf papers read 2609.33253 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
