Instructions to use Overdog/LIFT with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use Overdog/LIFT with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline from diffusers.utils import load_image, export_to_video # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("Overdog/LIFT", dtype=torch.bfloat16, device_map="cuda") pipe.to("cuda") prompt = "A man with short gray hair plays a red electric guitar." image = load_image( "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/diffusers/guitar-man.png" ) output = pipe(image=image, prompt=prompt).frames[0] export_to_video(output, "output.mp4") - Notebooks
- Google Colab
- Kaggle
|
Download README.md from Overdog/LIFT: direct link, hf CLI and curl.
- Browser
- Download file 2.87 kB
-
https://huggingface.co/Overdog/LIFT/resolve/main/README.md
- Command line
-
hf download hf://Overdog/LIFT/README.md
-
curl -L -o README.md https://huggingface.co/Overdog/LIFT/resolve/main/README.md
2.87 kB
metadata
license: apache-2.0
pipeline_tag: image-to-video
tags:
- video-generation
LIFT: Layout-In-Future Video Generation under Large Viewpoint Change via On-Policy Self-Distillation
Given a first frame, users can navigate from the first-frame view along a desired camera path and specify layouts using bounding boxes with local text prompts in the final frame. Then, LIFT generates the intended shot that transitions from the input image to the user-defined last-frame layout following the prescribed camera trajectory.
We introduce LIFT, a unified image-to-video generation framework that complements camera control with Layout-In-FuTure control, enabling users to specify what should appear in a future view and where it should appear.
Model Checkpoints
Our models are built on Wan2.1-Fun-V1.1-1.3B-Control-Camera.
| Model | Description |
|---|---|
LIFT/transformer/ |
Last-frame layout student trained by dual-mode on-policy self-distillation from the dense-layout teacher. |
LIFT_dense_layout_teacher/transformer/ |
Dense-layout teacher fine-tuned with dense per-frame layout |
Download
pip install -U "huggingface_hub[cli]"
# Last-frame layout student (dual-mode OPSD)
hf download Overdog/LIFT --include "LIFT/*" --local-dir models
# Dense-layout teacher
hf download Overdog/LIFT --include "LIFT_dense_layout_teacher/*" --local-dir models
Citation
@article{ji2026lift,
title={LIFT: Layout-In-Future Video Generation under Large Viewpoint Change via On-Policy Self-Distillation},
author={Ji, Shengxiang and Wang, Boyang and Xu, Haiyang and Li, Bingnan and Mao, Yucheng and Chen, Zeyuan and Shan, Xiaojun and Zhang, Xiang and Hua, Gang and Xie, Jianwen and Cheng, Zezhou and Tu, Zhuowen},
journal={arXiv preprint arXiv:2609.38146},
year={2026}
}