Papers
arxiv:2610.12459

WorldGuide: Goal-Directed Video World Model for Procedural Task Execution

Published on Oct 8
· Submitted by
Ankan Deria
on Oct 9
Authors:
,
,
,
,

Abstract

Video generators and video-based world models can synthesize plausible visual trajectories, but long-horizon procedural tasks require generation to adapt to what has actually been produced. A model must determine the next action from its generated state, execute that action, and recognize when the task is complete. Open-loop generation cannot adapt to execution outcomes, while existing closed-loop systems often rely on pretrained executors or indirect verification. This leaves a gap between deciding an action and successfully realizing it. We formulate procedural video generation as closed-loop task execution in visual world space and introduce WorldGuide. Given only an initial image and a task goal, WorldGuide predicts an atomic action, generates its corresponding video clip, and uses the generated result to select the next action or terminate. The Planner and Executor are trained on the same step-level procedural demonstrations: the Planner learns to predict the next atomic action or task completion from visual progress, while the Executor is directly trained to realize the predicted actions. Hierarchical visual memory maintains state across long-horizon execution with bounded history token cost. Due to the lack of step-level action-video supervision for joint planner-executor training, we introduce WorldGuide Bench: approximately 59K step-annotated videos across 245 tasks and 27 procedural categories. WorldGuide achieves a 33.33\% Task Success on WorldGuide-Bench, compared with 29.90\% for the strong recent video model MiniMax-H3, even though MiniMax-H3 receives reference action plans, and achieves 47.69\% on VideoCraft-Bench compared with 32.73\% for MiniMax-H3 under goal-only conditioning. These results demonstrate the importance of coupling planning with learned execution for goal-directed procedural video generation.

Community

Paper submitter

Excited to share WorldGuide, a goal-directed video world model for procedural task execution.

Instead of generating a fixed open-loop rollout, WorldGuide treats procedural video generation as closed-loop task execution: a ContextPlanner predicts the next atomic action (or a completion token) from the generated visual state, an Executor trained on the same step-level demonstrations renders it as a clip, and the result guides the next action or termination. Hierarchical visual memory keeps long rollouts at bounded token cost.

We also introduce WorldGuide Bench (~59K step-annotated videos, 245 tasks, 27 procedural categories) to supply the missing step-level action-video supervision.

WorldGuide reaches 33.33% Task Success on WorldGuide Bench (vs 29.90% for MiniMax-H3, which receives reference action plans) and 47.69% on Video-CraftBench (vs 32.73% under goal-only conditioning).

Project page: https://mbzuai-oryx.github.io/WorldGuide
Code: https://github.com/mbzuai-oryx/WorldGuide

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2610.12459
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 1

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2610.12459 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2610.12459 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.