V-RAE: Rethinking Video Latent Spaces for Generation

Minghui Guo1    Shengqiong Wu2    Hao Fei2

1National University of Singapore    2University of Oxford

Paper Project Page Code

V-RAE is a video representation autoencoder that builds compact generative latents on top of frozen vision foundation model representations. It uses a lightweight temporal pooling module to remove temporal redundancy while preserving semantic structure, together with a video decoder for reconstructing continuous motion.

V-RAE method

V-RAE overview

Released Models

This repository provides V-RAE checkpoints based on four frozen visual encoders:

Checkpoint Encoder
vrae/vrae_dinov3.pt DINOv3 ViT-L/16
vrae/vrae_siglip2.pt SigLIP2 ViT-L/16
vrae/vrae_vjepa2.1.pt V-JEPA2.1 ViT-L/16
vrae/vrae_eupe.pt EUPE ViT-B/16

Class-conditional VideoDiT checkpoints and their matching latent statistics are also provided for UCF101 and Kinetics-600:

  • UCF101: V-JEPA2.1 and EUPE variants
  • Kinetics-600: V-JEPA2.1 variant

Quick Start

Install the official implementation:

git clone https://github.com/V-RAE/V-RAE.git
cd V-RAE

conda create -n vrae python=3.10 -y
conda activate vrae
conda install -c conda-forge ffmpeg -y

pip install uv
uv pip install -e .

Download the released checkpoints:

hf download Guomh0707/V-RAE-Models --local-dir ckpts

Download the matching frozen encoder by following the encoder instructions.

Place three input videos at assets/sample1.mp4, assets/sample2.mp4, and assets/sample3.mp4, then run one of:

python sampling.py dino
python sampling.py siglip
python sampling.py vjepa
python sampling.py eupe

Reconstruction results are saved under outputs/<variant>/.

Results

Evaluation Result
Kinetics-600 reconstruction 2.13 rFVD
UCF101 class-conditional generation 117.86 gFVD
Kinetics-600 class-conditional generation 19.16 gFVD

Notes

  • Each V-RAE checkpoint requires its corresponding frozen encoder weights.
  • These are custom PyTorch checkpoints and require the official V-RAE codebase.
  • This repository is intended primarily for research use.

Citation

@article{guo2026vrae,
  title   = {V-RAE: Rethinking Video Latent Spaces for Generation},
  author  = {Guo, Minghui and Wu, Shengqiong and Fei, Hao},
  journal = {arXiv preprint arXiv:2608.13556},
  year    = {2026},
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train Guomh0707/V-RAE-Models

Paper for Guomh0707/V-RAE-Models