--- license: apache-2.0 pipeline_tag: text-to-video inference: false tags: - world-model - video-generation - text-to-video - image-to-video - interactive - distillation ---

Alaya-EVOKE: From Linear-Scaling Supervision to Endless World

GitHub Project Page arXiv Paper page

Model weights for **EVOKE** ([paper](https://huggingface.co/papers/2608.13546)), a 3-step, CFG-free interactive world model that generates **384 × 640 @ 24 fps** video and stays coherent over 30 s rollouts. Code, docs and demos live in the GitHub repository — **this repository holds weights only.** - ⚡ **3 steps, zero CFG** — 1.5 s of video every 2.11 s on one H200, one forward per step. - 🌍 **Endless, not windowed** — scene geometry lives in an external camera-indexed world state bank, so the denoiser context stays bounded however long the session runs. - 🎛️ **Re-promptable mid-flight** — change the prompt while the rollout is running, no cut, no restart. ## Contents Every EVOKE directory is the **parent** of a `transformer/`, because it loads as `from_pretrained(path, subfolder="transformer")`. ``` evoke-base/ vae / text_encoder / tokenizer / scheduler only evoke/ ├── stage1_camera_control/transformer/ multi-step camera-controllable model ├── stage2_few_step_training/transformer/ few-step distillation (3-step pyramid) ├── stage3_long_distillation/transformer/ 30 s long-video distillation (post-distill init) ├── stage3_post_distillation/transformer/ the shipped model └── evoke_teacher/{high,low}_noise/ the two DMD teacher experts -- training only ``` ## Usage ```bash git clone https://github.com/AlayaLab/Evoke && cd Evoke pip install -r requirements.txt hf download AlayaLab/Evoke --local-dir models hf download pkqbajng/ViGeo --local-dir models/ViGeo1.1 # REQUIRED depth backend MODE=t2v NUM_CHUNKS=20 bash scripts/inference/infer_post_distill.sh ``` `ViGeo` is a separate download and is **required** — every shipped recipe uses it as the depth backend behind the world state bank. Depth-Anything-3 is optional. Both ship under CC-BY-NC-4.0, which is more restrictive than this repository's Apache-2.0; check their licences before any commercial use. Inference modes, the mode × model matrix, hour-scale rollouts and training are documented in the GitHub repository. ## Notes The distilled models were trained on v2v conditioning alone, so `MODE=i2v|t2v` on them is **zero-shot**; only `stage1_camera_control` has all three modes in distribution. The vae / text encoder / tokenizer / scheduler in `evoke-base/` come from the released [Helios](https://github.com/PKU-YuanGroup/Helios) base, which traces them to Wan. The EVOKE teacher is built on [LingBot-World](https://github.com/robbyant/lingbot-world). ## Citation ```bibtex @article{evoke2026, title = {Alaya-EVOKE: From Linear-Scaling Supervision to Endless World}, author = {Yin, Yuanyang and Wang, Gongxuan and Zhan, Yifan and Li, Chuanhao and Zhang, Kaipeng and Zhao, Feng}, journal = {arXiv preprint arXiv:2608.13546}, year = {2026}, } ```