WorldGuide: Goal-Directed Video World Model for Procedural Task Execution
Paper β’ 2610.12459 β’ Published β’ 4
How to use ankanmbz/WorldGuide-Ckpt with Diffusers:
pip install -U diffusers transformers accelerate
import torch
from diffusers import DiffusionPipeline
from diffusers.utils import load_image, export_to_video
# switch to "mps" for apple devices
pipe = DiffusionPipeline.from_pretrained("ankanmbz/WorldGuide-Ckpt", dtype=torch.bfloat16, device_map="cuda")
pipe.to("cuda")
prompt = "A man with short gray hair plays a red electric guitar."
image = load_image(
"https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/diffusers/guitar-man.png"
)
output = pipe(image=image, prompt=prompt).frames[0]
export_to_video(output, "output.mp4")import torch
from diffusers import DiffusionPipeline
from diffusers.utils import load_image, export_to_video
# switch to "mps" for apple devices
pipe = DiffusionPipeline.from_pretrained("ankanmbz/WorldGuide-Ckpt", dtype=torch.bfloat16, device_map="cuda")
pipe.to("cuda")
prompt = "A man with short gray hair plays a red electric guitar."
image = load_image(
"https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/diffusers/guitar-man.png"
)
output = pipe(image=image, prompt=prompt).frames[0]
export_to_video(output, "output.mp4")
Ankan Deria, Komal Kumar, Hisham Cholakkal, Fahad Shahbaz Khan, Salman Khan
Mohamed bin Zayed University of Artificial Intelligence
Official checkpoints for WorldGuide. Method details, results and training code are on the project page and on GitHub.
WorldGuide-Ckpt/
βββ transformer/ # WorldGuide Video DiT β --action_ckpt
βββ text_encoder/llm/ # WorldGuide ContextPlanner β --planner_model_path
git clone https://github.com/mbzuai-oryx/WorldGuide.git
cd WorldGuide
pip install -r requirements.txt
# Download this repository and all other required components into ./ckpts
python download_models.py --hf_token <your_hf_token>
# Run closed-loop inference
bash scripts/inference/run_v5_qwen_closed_loop_memory.sh
@article{deria2026worldguide,
title={WorldGuide: Goal-Directed Video World Model for Procedural Task Execution},
author={Deria, Ankan and Kumar, Komal and Cholakkal, Hisham and Khan, Fahad Shahbaz and Khan, Salman},
journal={arXiv preprint arXiv:2610.12459},
year={2026}
}