373 GB
289,899 files
Updated 2 days ago
Name
Size
scripts
README.md3.32 kB
xet
download.py678 Bytes
xet
README.md

BridgeData v2 Processing & Analysis

Download model weights locally

/mnt/posttrain/zhaoshitian/models/RAE-collections
# https://huggingface.co/nyu-visionx/RAE-collections

/mnt/posttrain/zhaoshitian/models/dinov2-with-registers-base
# https://huggingface.co/facebook/dinov2-with-registers-base

/mnt/posttrain/zhaoshitian/models/vit-mae-base
# https://huggingface.co/facebook/vit-mae-base

/mnt/posttrain/zhaoshitian/models/siglip2-base-patch16-256
# https://huggingface.co/google/siglip2-base-patch16-256

/mnt/posttrain/zhaoshitian/models/FLUX.2-klein-4B/vae/ae_bfl_format.safetensors
# https://huggingface.co/black-forest-labs/FLUX.2-klein-4B

Raw dataset preparation & preprocess with vidaforge

Download data from: https://opendatalab.com/OpenDataLab/UCF101

Expected raw UCF101 dataset root:

/mnt/posttrain/zhaoshitian/datasets/ucf101/OpenDataLab___UCF101/raw/data
uv run bash /home/sz128/projects/video-velocity-model/data_processing/scripts/vidaforge/stage1_ingestion/run_step1_probe.sh
uv run bash /home/sz128/projects/video-velocity-model/data_processing/scripts/vidaforge/stage1_ingestion/run_step2_screen.sh
uv run bash /home/sz128/projects/video-velocity-model/data_processing/scripts/vidaforge/stage1_ingestion/run_step3_transcode.sh

# detect the scene in raw videos with transnetv2
uv run bash /home/sz128/projects/video-velocity-model/data_processing/scripts/vidaforge/stage2_segmentation/run_step1_detect.sh
uv run bash /home/sz128/projects/video-velocity-model/data_processing/scripts/vidaforge/stage2_segmentation/run_step2_clip.sh

UCF101 postprocessing

The preprocessing script reads official UCF101 split files, encodes video frames into RAE latent space, and optionally decodes them back into reconstruction videos.

It now supports three local RAE backbones plus the FLUX.2 autoencoder:

  • dinov2
  • mae
  • siglip2
  • flux2_ae

RAE-collections is used for RAE decoder checkpoints and latent normalization stats, while the three encoder backbones are loaded from the local model directories above. flux2_ae uses the FLUX.2 AE checkpoint directly for both encoding and decoding.

Default inputs:

/mnt/posttrain/zhaoshitian/datasets/ucf101/OpenDataLab___UCF101/raw/data/UCF101
/mnt/posttrain/zhaoshitian/datasets/ucf101/OpenDataLab___UCF101/raw/data/UCF101TrainTestSplits-RecognitionTask.zip

Run the full preprocessing job with the different backbones:

bash scripts/run_data_process_dinov2.sh
bash scripts/run_data_process_mae.sh
bash scripts/run_data_process_siglip2.sh
bash scripts/run_data_process_flux2ae.sh

Backbone defaults:

dinov2  -> encoder_input_size=224
mae     -> encoder_input_size=256
siglip2 -> encoder_input_size=256
flux2_ae -> encoder_input_size=256

If --image-size is not specified, the script uses the selected backbone's default input size automatically.

Default outputs:

Each processed video writes:

  • latent file: <class_name>/<video_name>_patch_tokens.npz
  • reconstruction video: <class_name>/<video_name>.mp4

The latent .npz contains:

  • features: (T, C, H, W); for dinov2/mae/siglip2, usually (T, 768, 16, 16); for flux2_ae at the default 256 crop, (T, 128, 16, 16)
  • timesteps: (T,)

Here T is the number of decoded video frames.

Total size
373 GB
Files
289,899
Last updated
Aug 4
Pre-warmed CDN
US EU US EU

Contributors