Instructions to use AlayaLab/AlayaWorld-stage1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use AlayaLab/AlayaWorld-stage1 with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline from diffusers.utils import load_image, export_to_video # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("AlayaLab/AlayaWorld-stage1", dtype=torch.bfloat16, device_map="cuda") pipe.to("cuda") prompt = "A man with short gray hair plays a red electric guitar." image = load_image( "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/diffusers/guitar-man.png" ) output = pipe(image=image, prompt=prompt).frames[0] export_to_video(output, "output.mp4") - Notebooks
- Google Colab
- Kaggle
import torch
from diffusers import DiffusionPipeline
from diffusers.utils import load_image, export_to_video
# switch to "mps" for apple devices
pipe = DiffusionPipeline.from_pretrained("AlayaLab/AlayaWorld-stage1", dtype=torch.bfloat16, device_map="cuda")
pipe.to("cuda")
prompt = "A man with short gray hair plays a red electric guitar."
image = load_image(
"https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/diffusers/guitar-man.png"
)
output = pipe(image=image, prompt=prompt).frames[0]
export_to_video(output, "output.mp4")AlayaWorld Stage1 — shared bidirectional checkpoint
This repository releases the same Stage1 bidirectional pretraining checkpoint shared by the DA3 version of AlayaWorld and the v1.1 ViGeo version. Both versions use this checkpoint before their later autoregressive and spatial-memory training stages.
Stage1 supports image-to-video generation from a first-frame image and a text prompt. This inference path uses neither DA3 nor ViGeo and requires no camera trajectory. The later camera-controlled AR/DMD checkpoints are version-specific; use their corresponding inference configurations and weights.
Code and full setup: AlayaLab/AlayaWorld.
Files
| File | Contents |
|---|---|
diffusion_pytorch_model.safetensors |
Full Stage1 video transformer (26,248,786,528 bytes) |
config.json |
Original transformer configuration |
SHA256SUMS |
SHA-256 checksums of the two checkpoint files |
The release does not bundle the VAE or text encoder. Obtain the LTX-2.3 base
(ltx-2.3-22b-dev.safetensors) and
Gemma
separately, following the code repository's setup instructions.
Inference
From the latest AlayaWorld code repository and its Python environment:
hf download AlayaLab/AlayaWorld-stage1 \
config.json diffusion_pytorch_model.safetensors \
--local-dir weights/AlayaWorld-stage1
VALIDATE_ONLY=1 CONFIG_PATH=configs/infer_i2v_bidir.yaml bash scripts/finetune/train.sh
The config expects the LTX-2.3 base at
weights/ltx-2.3/ltx-2.3-22b-dev.safetensors and Gemma at
weights/ltx-2.3/google/gemma-3-12b-it-qat-q4_0-unquantized.
Adjust paths.base_transformer, paths.vae, paths.gemma and
paths.resume_checkpoint for your installation.
The default input is the bundled playground/case1 image and prompt. For your
own input, edit image_dir and prompt_file under
validation.modes.i2v_bidir.dataset; captions_json supports per-image prompts.
Default generation: 481 frames, 960×544, 24 fps (about 20 seconds), 30 steps,
CFG 3, STG 1 (block 28), rescale 0.7, seed 42. This produces a single
bidirectional clip through the shared RolloutTrainer validation path.
Set LOG_FILTER=all to show loading and sampling progress.
中文说明
本仓库发布 AlayaWorld DA3 版本与 v1.1 ViGeo 版本共享的同一份 Stage1 双向预训练权重。两个版本都从这份 checkpoint 进入后续自回归与空间记忆训练阶段。
Stage1 双向推理只需首帧图片和文本提示词,本身不使用 DA3 或 ViGeo,也不需要 相机轨迹。后续支持相机控制的 AR/DMD 权重仍对应各自版本,应使用相应权重与配置。
发布文件包括完整 transformer 权重、原始 config.json 和文件校验值。
VAE 与文本编码器未打包在本仓库中,需按照
代码仓库
说明单独获取 LTX-2.3 底座和 Gemma。下载及启动命令见上方。
默认配置使用 playground/case1,生成 481 帧、960×544、24fps、约 20 秒的
单段视频,采样 30 步。使用自己的输入时,修改
validation.modes.i2v_bidir.dataset 下的 image_dir 与 prompt_file;
也可用 captions_json 指定每张图片的提示词。推理复用开源代码的
RolloutTrainer 验证流程。通过 LOG_FILTER=all 可查看加载与采样进度。
License / 许可证
These weights are fine-tuned from LTX-2.3 by Alaya Lab and are released under the LTX-2 Community License Agreement, with the AlayaWorld project's academic research and non-commercial-use restriction. See NOTICE for attribution. Third-party dependencies retain their own licenses.
本权重由 Alaya Lab 基于 LTX-2.3 微调,依据 LTX-2 社区许可协议 及 AlayaWorld 项目的学术研究与非商业用途限制发布。署名信息见 NOTICE, 第三方依赖遵循各自许可证。
- Downloads last month
- 10