stzhao's picture
|
download
raw
3.31 kB

UCF101 Data Processing & Analysis

Download model weights locally

/mnt/posttrain/zhaoshitian/models/RAE-collections
# https://huggingface.co/nyu-visionx/RAE-collections

/mnt/posttrain/zhaoshitian/models/dinov2-with-registers-base
# https://huggingface.co/facebook/dinov2-with-registers-base

/mnt/posttrain/zhaoshitian/models/vit-mae-base
# https://huggingface.co/facebook/vit-mae-base

/mnt/posttrain/zhaoshitian/models/siglip2-base-patch16-256
# https://huggingface.co/google/siglip2-base-patch16-256

/mnt/posttrain/zhaoshitian/models/FLUX.2-klein-4B/vae/ae_bfl_format.safetensors
# https://huggingface.co/black-forest-labs/FLUX.2-klein-4B

Raw dataset preparation & preprocess with vidaforge

Download data from: https://opendatalab.com/OpenDataLab/UCF101

Expected raw UCF101 dataset root:

/mnt/posttrain/zhaoshitian/datasets/ucf101/OpenDataLab___UCF101/raw/data
uv run bash /home/sz128/projects/video-velocity-model/data_processing/scripts/vidaforge/stage1_ingestion/run_step1_probe.sh
uv run bash /home/sz128/projects/video-velocity-model/data_processing/scripts/vidaforge/stage1_ingestion/run_step2_screen.sh
uv run bash /home/sz128/projects/video-velocity-model/data_processing/scripts/vidaforge/stage1_ingestion/run_step3_transcode.sh

# detect the scene in raw videos with transnetv2
uv run bash /home/sz128/projects/video-velocity-model/data_processing/scripts/vidaforge/stage2_segmentation/run_step1_detect.sh
uv run bash /home/sz128/projects/video-velocity-model/data_processing/scripts/vidaforge/stage2_segmentation/run_step2_clip.sh

UCF101 postprocessing

The preprocessing script reads official UCF101 split files, encodes video frames into RAE latent space, and optionally decodes them back into reconstruction videos.

It now supports three local RAE backbones plus the FLUX.2 autoencoder:

  • dinov2
  • mae
  • siglip2
  • flux2_ae

RAE-collections is used for RAE decoder checkpoints and latent normalization stats, while the three encoder backbones are loaded from the local model directories above. flux2_ae uses the FLUX.2 AE checkpoint directly for both encoding and decoding.

Default inputs:

/mnt/posttrain/zhaoshitian/datasets/ucf101/OpenDataLab___UCF101/raw/data/UCF101
/mnt/posttrain/zhaoshitian/datasets/ucf101/OpenDataLab___UCF101/raw/data/UCF101TrainTestSplits-RecognitionTask.zip

Run the full preprocessing job with the different backbones:

bash scripts/run_data_process_dinov2.sh
bash scripts/run_data_process_mae.sh
bash scripts/run_data_process_siglip2.sh
bash scripts/run_data_process_flux2ae.sh

Backbone defaults:

dinov2  -> encoder_input_size=224
mae     -> encoder_input_size=256
siglip2 -> encoder_input_size=256
flux2_ae -> encoder_input_size=256

If --image-size is not specified, the script uses the selected backbone's default input size automatically.

Default outputs:

Each processed video writes:

  • latent file: <class_name>/<video_name>_patch_tokens.npz
  • reconstruction video: <class_name>/<video_name>.mp4

The latent .npz contains:

  • features: (T, C, H, W); for dinov2/mae/siglip2, usually (T, 768, 16, 16); for flux2_ae at the default 256 crop, (T, 128, 16, 16)
  • timesteps: (T,)

Here T is the number of decoded video frames.

Xet Storage Details

Size:
3.31 kB
·
Xet hash:
7d6ce877c7cbbcf1f14d6b45880ec10aeb5d400d41dbb86ebbf9798c3f7a00c4

Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.