Spaces:
Running on Zero
Running on Zero
A newer version of the Gradio SDK is available: 6.24.0
metadata
title: VEGA-3D Spatial Reasoning
emoji: 馃Л
colorFrom: red
colorTo: gray
sdk: gradio
sdk_version: 5.49.1
app_file: app.py
short_description: 3D spatial reasoning VLM with video-diffusion 3D priors
python_version: '3.10'
startup_duration_timeout: 1h
pinned: false
VEGA-3D 路 Spatial Reasoning
Interactive demo of VEGA-3D, a spatial-reasoning multimodal LLM from the paper Generation Models Know Space: Unleashing Implicit 3D Priors for Scene Understanding (ECCV 2026).
VEGA-3D augments a Qwen2.5-VL backbone with implicit 3D priors extracted from a frozen Wan2.1-T2V-1.3B video diffusion model. Intermediate spatiotemporal features from the diffusion model are fused into the VLM via token-level gated fusion, giving it stronger geometric and spatial understanding of indoor scenes.
Upload a scene image (or a short scene video) and ask a spatial-reasoning question.