Text-to-Video
Diffusers
Safetensors
world-model
video-generation
image-to-video
interactive
distillation
Instructions to use AlayaLab/Evoke with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use AlayaLab/Evoke with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("AlayaLab/Evoke", dtype=torch.bfloat16, device_map="cuda") prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k" image = pipe(prompt).images[0] - Notebooks
- Google Colab
- Kaggle
| license: apache-2.0 | |
| pipeline_tag: text-to-video | |
| inference: false | |
| tags: | |
| - world-model | |
| - video-generation | |
| - text-to-video | |
| - image-to-video | |
| - interactive | |
| - distillation | |
| <h1 align="center">Alaya-EVOKE: From Linear-Scaling Supervision to Endless World</h1> | |
| <p align="center"> | |
| <a href="https://github.com/AlayaLab/Evoke"><img src="https://img.shields.io/badge/GitHub-AlayaLab/Evoke-181717.svg?logo=github" alt="GitHub"></a> | |
| <a href="https://evoke-world.github.io/Evoke/"><img src="https://img.shields.io/badge/π_Project_Page-evoke--world.github.io-1a73e8.svg" alt="Project Page"></a> | |
| <a href="https://arxiv.org/abs/2608.13546"><img src="https://img.shields.io/badge/arXiv-2608.13546-b31b1b.svg" alt="arXiv"></a> | |
| <a href="https://huggingface.co/papers/2608.13546"><img src="https://img.shields.io/badge/π€_Paper-2608.13546-ffce1c.svg" alt="Paper page"></a> | |
| </p> | |
| Model weights for **EVOKE** ([paper](https://huggingface.co/papers/2608.13546)), a 3-step, CFG-free | |
| interactive world model that generates **384 Γ 640 @ 24 fps** video and stays coherent over 30 s | |
| rollouts. Code, docs and demos live in the GitHub repository β **this repository holds weights only.** | |
| - β‘ **3 steps, zero CFG** β 1.5 s of video every 2.11 s on one H200, one forward per step. | |
| - π **Endless, not windowed** β scene geometry lives in an external camera-indexed world state bank, | |
| so the denoiser context stays bounded however long the session runs. | |
| - ποΈ **Re-promptable mid-flight** β change the prompt while the rollout is running, no cut, no restart. | |
| ## Contents | |
| Every EVOKE directory is the **parent** of a `transformer/`, because it loads as | |
| `from_pretrained(path, subfolder="transformer")`. | |
| ``` | |
| evoke-base/ vae / text_encoder / tokenizer / scheduler only | |
| evoke/ | |
| βββ stage1_camera_control/transformer/ multi-step camera-controllable model | |
| βββ stage2_few_step_training/transformer/ few-step distillation (3-step pyramid) | |
| βββ stage3_long_distillation/transformer/ 30 s long-video distillation (post-distill init) | |
| βββ stage3_post_distillation/transformer/ the shipped model | |
| βββ evoke_teacher/{high,low}_noise/ the two DMD teacher experts -- training only | |
| ``` | |
| ## Usage | |
| ```bash | |
| git clone https://github.com/AlayaLab/Evoke && cd Evoke | |
| pip install -r requirements.txt | |
| hf download AlayaLab/Evoke --local-dir models | |
| hf download pkqbajng/ViGeo --local-dir models/ViGeo1.1 # REQUIRED depth backend | |
| MODE=t2v NUM_CHUNKS=20 bash scripts/inference/infer_post_distill.sh | |
| ``` | |
| `ViGeo` is a separate download and is **required** β every shipped recipe uses it as the depth backend | |
| behind the world state bank. Depth-Anything-3 is optional. Both ship under CC-BY-NC-4.0, which is more | |
| restrictive than this repository's Apache-2.0; check their licences before any commercial use. | |
| Inference modes, the mode Γ model matrix, hour-scale rollouts and training are documented in the | |
| GitHub repository. | |
| ## Notes | |
| The distilled models were trained on v2v conditioning alone, so `MODE=i2v|t2v` on them is **zero-shot**; | |
| only `stage1_camera_control` has all three modes in distribution. | |
| The vae / text encoder / tokenizer / scheduler in `evoke-base/` come from the released | |
| [Helios](https://github.com/PKU-YuanGroup/Helios) base, which traces them to Wan. The EVOKE teacher is | |
| built on [LingBot-World](https://github.com/robbyant/lingbot-world). | |
| ## Citation | |
| ```bibtex | |
| @article{evoke2026, | |
| title = {Alaya-EVOKE: From Linear-Scaling Supervision to Endless World}, | |
| author = {Yin, Yuanyang and Wang, Gongxuan and Zhan, Yifan and | |
| Li, Chuanhao and Zhang, Kaipeng and Zhao, Feng}, | |
| journal = {arXiv preprint arXiv:2608.13546}, | |
| year = {2026}, | |
| } | |
| ``` | |