OvDSGG: End-to-End Open-Vocabulary Dynamic Scene Graph Generation
Abstract
OvDSGG is an end-to-end open-vocabulary dynamic scene graph generation framework that uses spatial and temporal backbones with cross-modal alignment to recognize unseen objects and predicates without costly multi-stage training.
Dynamic scene graphs (DSGs) capture spatio-temporal interactions across videos as langlesubject, predicate, objectrangle triplets, and underpin downstream tasks such as video captioning, video question answering, and action analysis. However, end-to-end dynamic scene graph generation (DSGG) methods are closed-set: they recognize only objects and predicates from a fixed training vocabulary and struggle with the long-tailed distribution of rare concepts, severely limiting their real-world applicability. Existing open-vocabulary models typically inherit pretrained large language models, resulting in multi-stage training and inference with substantial cost. We introduce OvDSGG, the first end-to-end framework for open-vocabulary DSGG. OvDSGG builds on top of an open-vocabulary Spatial Backbone and a Temporal Backbone; we further propose a Triplet Feature Extraction Module that bridges them, and a Visual-Language Alignment Module that preserves open-vocabulary recognition by learning an adaptive decision boundary in the joint visual-language feature space, without expensive knowledge distillation in existing methods. We further introduce a rigorous open-vocabulary DSGG benchmark adapted from Action Genome, with disjoint Base/Novel splits for both objects and predicates. OvDSGG significantly outperforms open-vocabulary baselines across all metrics, with zero-shot Recall@K scores 10.0--20.4 percentage point higher than the next-best baseline, while on closed-set DSGG remaining competitive with state-of-the-art models. Code and benchmark are publicly available at https://github.com/jhelsby/OvDSGG/.
Get this paper in your agent:
hf papers read 2608.14835 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 1
Datasets citing this paper 1
jhelsby/ovdsgg-action-genome-split
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper