StoryEngine: A State-Grounded Agentic Framework for Video Storytelling
Abstract
Despite recent progress in agentic multi-shot video generation, producing coherent and consistent long-form stories remains challenging. Existing agentic pipelines typically rely on textual shot plans or previously generated pixels, yet lack an explicit mechanism for propagating the consequences of story events and maintaining the video world state across shots. As a result, missing visual details may be reconstructed inaccurately, while visual drift may propagate across subsequent shots, undermining both narrative coherence and visual consistency. To address these challenges, we propose StoryEngine, a state-grounded agentic framework for video storytelling. StoryEngine establishes a separation between authoritative semantic plans and unreliable visual observations. Specifically, StoryEngine maintains a structured representation of entity placement and story-relevant states, and propagates event-induced changes to define the intended start and end states of each shot. To visually realize these states, StoryEngine constructs canonical references for recurring entities and environments, and compiles state and visual constraints into executable render plans. Meanwhile, to realize these states correctly, a bounded evaluation-guided repair loop further corrects local state inconsistencies. Together, these mechanisms preserve causal story progression and prevent local visual errors from propagating across shots. To comprehensively evaluate long-form storytelling, we construct a benchmark across diverse scenarios and visual styles, with metrics assessing storytelling quality, narrative coherence, and visual consistency. Experimental results demonstrate that StoryEngine consistently outperforms state-of-the-art methods across all evaluation dimensions, validating its effectiveness for coherent and consistent video storytelling.
Community
đŹ StoryEngine: A State-Driven Agentic Framework for Multi-Shot Video Storytelling
AI can generate beautiful short videos. But keeping a story coherent across multiple shots remains a challenge: characters drift, objects become inconsistent, and errors accumulate.
Our approach: give the story an explicit âworld state.â
This state tracks where characters and objects are, and how events change them. It also defines how each shot should begin and endâso generation follows the intended story.
⨠Highlights
đ§ Track the storyâ¨Turn narrative events into explicit state changes, keeping track of what happens and what changes from shot to shot.
đźď¸ Keep visuals consistentâ¨Create shared visual references for characters, objects, and scenes. Select camera viewpoints to suit the action, and translate story states into clear shot-generation requirements.
đ Verify and correctâ¨Use the next shotâs expected state to decide how to use the previous shotâs visuals: reuse them, reference them partially, or discard them.
Check actions and states in every shot, then make targeted corrections within a fixed retry budget. The goal: prevent visual errors from changing what happens next in the story.
đ Evaluate the whole storyânot just individual shotsâ¨We also introduce a video evaluation benchmark with 8 metrics across 3 dimensions:â¨â˘ Storytellingâ¨â˘ Cross-shot coherenceâ¨â˘ Visual consistency
Get this paper in your agent:
hf papers read 2609.33627 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper