Buckets:
UCF101 Data Processing & Analysis
Download model weights locally
/mnt/posttrain/zhaoshitian/models/RAE-collections
# https://huggingface.co/nyu-visionx/RAE-collections
/mnt/posttrain/zhaoshitian/models/dinov2-with-registers-base
# https://huggingface.co/facebook/dinov2-with-registers-base
/mnt/posttrain/zhaoshitian/models/vit-mae-base
# https://huggingface.co/facebook/vit-mae-base
/mnt/posttrain/zhaoshitian/models/siglip2-base-patch16-256
# https://huggingface.co/google/siglip2-base-patch16-256
/mnt/posttrain/zhaoshitian/models/FLUX.2-klein-4B/vae/ae_bfl_format.safetensors
# https://huggingface.co/black-forest-labs/FLUX.2-klein-4B
Raw dataset preparation & preprocess with vidaforge
Download data from: https://opendatalab.com/OpenDataLab/UCF101
Expected raw UCF101 dataset root:
/mnt/posttrain/zhaoshitian/datasets/ucf101/OpenDataLab___UCF101/raw/data
uv run bash /home/sz128/projects/video-velocity-model/data_processing/scripts/vidaforge/stage1_ingestion/run_step1_probe.sh
uv run bash /home/sz128/projects/video-velocity-model/data_processing/scripts/vidaforge/stage1_ingestion/run_step2_screen.sh
uv run bash /home/sz128/projects/video-velocity-model/data_processing/scripts/vidaforge/stage1_ingestion/run_step3_transcode.sh
# detect the scene in raw videos with transnetv2
uv run bash /home/sz128/projects/video-velocity-model/data_processing/scripts/vidaforge/stage2_segmentation/run_step1_detect.sh
uv run bash /home/sz128/projects/video-velocity-model/data_processing/scripts/vidaforge/stage2_segmentation/run_step2_clip.sh
UCF101 postprocessing
The preprocessing script reads official UCF101 split files, encodes video frames into RAE latent space, and optionally decodes them back into reconstruction videos.
It now supports three local RAE backbones plus the FLUX.2 autoencoder:
dinov2maesiglip2flux2_ae
RAE-collections is used for RAE decoder checkpoints and latent normalization stats, while the three encoder backbones are loaded from the local model directories above. flux2_ae uses the FLUX.2 AE checkpoint directly for both encoding and decoding.
Default inputs:
/mnt/posttrain/zhaoshitian/datasets/ucf101/OpenDataLab___UCF101/raw/data/UCF101
/mnt/posttrain/zhaoshitian/datasets/ucf101/OpenDataLab___UCF101/raw/data/UCF101TrainTestSplits-RecognitionTask.zip
Run the full preprocessing job with the different backbones:
bash scripts/run_data_process_dinov2.sh
bash scripts/run_data_process_mae.sh
bash scripts/run_data_process_siglip2.sh
bash scripts/run_data_process_flux2ae.sh
Backbone defaults:
dinov2 -> encoder_input_size=224
mae -> encoder_input_size=256
siglip2 -> encoder_input_size=256
flux2_ae -> encoder_input_size=256
If --image-size is not specified, the script uses the selected backbone's default input size automatically.
Default outputs:
Each processed video writes:
- latent file:
<class_name>/<video_name>_patch_tokens.npz - reconstruction video:
<class_name>/<video_name>.mp4
The latent .npz contains:
features:(T, C, H, W); fordinov2/mae/siglip2, usually(T, 768, 16, 16); forflux2_aeat the default 256 crop,(T, 128, 16, 16)timesteps:(T,)
Here T is the number of decoded video frames.
Xet Storage Details
- Size:
- 3.31 kB
- Xet hash:
- 7d6ce877c7cbbcf1f14d6b45880ec10aeb5d400d41dbb86ebbf9798c3f7a00c4
Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.