stzhao's picture
|
download
raw
3.54 kB

Data Processing & Analysis

ucf101

bridgedatav2

UMAP visualization

data_processing/umap_alg.py runs UMAP on per-frame features and saves a 2D scatter plot. Points are colored by frame index, which makes it easy to inspect temporal trajectories. It can also mark GOP boundaries by detecting keyframes / I-frames from the original video bitstream.

Supported inputs:

  • .npy: expects a 2D array with shape (T, D)
  • .npz: expects a key named features by default

For the latent files produced by forward.py, features usually has shape (T, 768, 16, 16). umap_alg.py now supports these files directly and will project them to 2D frame features before UMAP:

  • mean_pool (default): spatial average pooling, (T, 768, 16, 16) -> (T, 768)
  • flatten: flatten each frame latent, (T, 768, 16, 16) -> (T, 196608)

PPL smoothness metric

data_processing/cal_metric.py computes a StyleGAN-inspired perceptual path length (PPL) on video latent interpolation paths. It is meant to compare the local smoothness of the dinov2, mae, siglip2, and flux2_ae feature spaces.

The script:

  1. encodes raw video frames into the selected RAE latent space
  2. samples two nearby points on each adjacent-frame latent path
  3. decodes the two nearby latents back to images
  4. computes LPIPS between the decoded images
  5. divides by eps^2 and aggregates the result as PPL

Supported PPL interpolation geometries:

  • linear: standard per-gap linear interpolation
  • slerp: token-wise spherical interpolation
  • smooth_continuous: the same cubic-Hermite path used by stage-2 TrajectoryPlan(sampling_mode="smooth_continuous")
    • frame velocities are estimated from the whole latent trajectory, so each adjacent-frame gap depends on its temporal neighbors
    • this mode is currently implemented in data_processing/cal_metric.py only; data_processing/forward_intrapolate.py still supports linear / slerp

It can also visualize the per-gap PPL distribution for a single video:

  • x-axis: the left frame index of each adjacent-frame gap
  • y-axis: the PPL value of that gap

When --samples-per-gap > 1, the plot uses the mean PPL of each gap, which helps show how local smoothness changes across the whole sequence.

Interpretation:

  • lower PPL -> smoother local geometry / smoother interpolation path
  • higher PPL -> sharper perceptual changes under small latent-space moves

Comparison protocols:

  • default: use each backbone's runtime default interpolation
    • dinov2 -> linear
    • mae -> linear
    • siglip2 -> slerp
    • flux2_ae -> linear
  • default can also be overridden explicitly with --interp-mode smooth_continuous
  • controlled: force all backbones to use linear
  • both: report both sets of numbers in one run

Important outputs per result row:

  • ppl_mean_raw: mean over all collected PPL samples
  • ppl_mean_trimmed: mean after percentile trimming
  • ppl_median: median PPL
  • ppl_std: standard deviation
  • pairs_evaluated: number of adjacent latent-frame pairs
  • samples_evaluated: total sampled (alpha, alpha + eps) comparisons

Trimmed mean uses the percentile window controlled by:

  • --trim-percentile-low (default 1)
  • --trim-percentile-high (default 99)

Weight / dependency notes:

  • LPIPS uses the repo's cached vgg.pth weights under RAE/src/disc/.caches/
  • torchvision also needs a VGG16 backbone checkpoint
  • on this machine, a local VGG16 checkpoint is available at:
/home/i-zhaoshitian/models/vgg/vgg16-397923af.pth

Xet Storage Details

Size:
3.54 kB
·
Xet hash:
45d0e8cfe8d4c48c37f991dd2b2728f27fc08184c2d19c68267a89ba21a9b1f8

Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.