Video Prediction Policy 2: Predict Better, Act Better

Official VPP2 weights for zero-shot video prediction, RoboDojo and LIBERO.

Paper · Code · Project page · ModelScope

Model weights

Checkpoint Path Use
VPP2 Stage-1 Video (49 frames) checkpoints_video/vpp2-video-stage1-49f.pth Zero-shot image-to-video prediction
VPP2 Stage-2 Video (17 frames) checkpoints_video/vpp2-video-stage2-17f.pth Zero-shot image-to-video prediction
RoboDojo history-conditioned Video-10k checkpoints/initialization/robodojo_his10k.pt Joint + Action2B training initializer
RoboDojo joint Video-100k checkpoints/joint2b_s100000/video.pt Paired RoboDojo evaluation
RoboDojo Action2B-100k checkpoints/joint2b_s100000/action.pt Paired RoboDojo evaluation
LIBERO Video-10k checkpoints/libero/video_step010000.pt Action training and evaluation
LIBERO Action-30k checkpoints/libero/action_step030000.pt LIBERO, LIBERO-OOD and LIBERO-PRO evaluation

Shared VAE, CLIP, UMT5 and tokenizer assets are under checkpoints/Wan2.1-I2V-14B-480P/. Each policy pair includes its manifest and normalization statistics. Keep each pair together.

Download

Install the Hugging Face CLI:

python -m pip install -U huggingface_hub

Each command downloads the root config.json, which lists the released checkpoints, video frame counts and inference defaults, together with the selected weights and shared Wan encoders.

Stage-1 video

Use --num-frames 49.

hf download Haodong082399/VPP2 --local-dir weights \
  --include 'config.json' \
  --include 'checkpoints_video/vpp2-video-stage1-49f.pth' \
  --include 'checkpoints/Wan2.1-I2V-14B-480P/**'

Stage-2 video

Use --num-frames 17.

hf download Haodong082399/VPP2 --local-dir weights \
  --include 'config.json' \
  --include 'checkpoints_video/vpp2-video-stage2-17f.pth' \
  --include 'checkpoints/Wan2.1-I2V-14B-480P/**'

RoboDojo training

Use Video-10k to initialize joint Video + Action2B training.

hf download Haodong082399/VPP2 --local-dir weights \
  --include 'config.json' \
  --include 'checkpoints/initialization/robodojo_his10k.pt' \
  --include 'checkpoints/Wan2.1-I2V-14B-480P/**'

RoboDojo evaluation

Download the paired Video-100k and Action2B-100k weights with their normalization statistics.

hf download Haodong082399/VPP2 --local-dir weights \
  --include 'config.json' \
  --include 'checkpoints/joint2b_s100000/**' \
  --include 'checkpoints/Wan2.1-I2V-14B-480P/**'

LIBERO

Use the same Video-10k and Action-30k pair for LIBERO, LIBERO-OOD and LIBERO-PRO.

hf download Haodong082399/VPP2 --local-dir weights \
  --include 'config.json' \
  --include 'checkpoints/libero/**' \
  --include 'checkpoints/Wan2.1-I2V-14B-480P/**'

Both standalone video checkpoints use 30 denoising steps, CFG=4 and sigma shift 3. Follow the video prediction guide for image preparation and inference, or the RoboDojo and LIBERO guides for policy training and evaluation.

All weights

To download the complete repository, including config.json:

hf download Haodong082399/VPP2 --local-dir weights

Evaluation settings

Benchmark Checkpoint pair Denoising Action horizon / execution
RoboDojo Video-100k + Action2B-100k 10 Euler steps, shift 1 32 / 24
LIBERO / OOD / PRO Video-10k + Action-30k 10 Euler steps, shift 5 32 / 10

RoboDojo uses its EE16 z-score statistics; LIBERO uses its own 7D action min/max statistics. Use each benchmark's configuration and simulator setup.

Citation

See CITATION.bib for the paper citation.

Licenses

Shared Wan assets retain their upstream license in checkpoints/Wan2.1-I2V-14B-480P/LICENSE.txt. The official implementation has its own code license.

Downloads last month
36
Video Preview
loading

Model tree for Haodong082399/VPP2

Finetuned
(27)
this model

Paper for Haodong082399/VPP2