sync run artifacts: inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow
Browse filesThis view is limited to 50 files because it contains too many changes. See raw diff
- inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/best.pt +3 -0
- inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/e116_best_snapshot.pt +3 -0
- inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/e150_snapshot.pt +3 -0
- inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/git-info.txt +2 -0
- inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_25163/25163_0_log.err +102 -0
- inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_25163/25163_0_log.out +0 -0
- inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_25163/25163_1_log.err +42 -0
- inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_25163/25163_1_log.out +212 -0
- inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_25163/25163_2_log.err +36 -0
- inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_25163/25163_2_log.out +212 -0
- inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_25163/25163_3_log.err +48 -0
- inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_25163/25163_3_log.out +212 -0
- inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_25163/25163_4_log.err +50 -0
- inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_25163/25163_4_log.out +212 -0
- inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_25163/25163_5_log.err +45 -0
- inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_25163/25163_5_log.out +212 -0
- inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_25214/25214_0_log.err +124 -0
- inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_25214/25214_0_log.out +0 -0
- inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_25214/25214_1_log.err +41 -0
- inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_25214/25214_1_log.out +218 -0
- inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_25214/25214_2_log.err +61 -0
- inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_25214/25214_2_log.out +218 -0
- inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_25214/25214_3_log.err +19 -0
- inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_25214/25214_3_log.out +218 -0
- inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_25214/25214_4_log.err +33 -0
- inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_25214/25214_4_log.out +218 -0
- inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_25214/25214_5_log.err +47 -0
- inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_25214/25214_5_log.out +218 -0
- inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_26022/26022_0_log.err +18 -0
- inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_26022/26022_0_log.out +241 -0
- inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_26022/26022_1_log.err +6 -0
- inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_26022/26022_1_log.out +217 -0
- inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_26022/26022_2_log.err +6 -0
- inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_26022/26022_2_log.out +217 -0
- inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_26986/26986_0_log.err +148 -0
- inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_26986/26986_0_log.out +0 -0
- inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_26986/26986_1_log.err +82 -0
- inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_26986/26986_1_log.out +218 -0
- inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_26986/26986_2_log.err +79 -0
- inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_26986/26986_2_log.out +218 -0
- inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_27036/27036_0_log.err +161 -0
- inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_27036/27036_0_log.out +0 -0
- inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_27036/27036_1_log.err +102 -0
- inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_27036/27036_1_log.out +218 -0
- inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_27036/27036_2_log.err +103 -0
- inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_27036/27036_2_log.out +218 -0
- inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_27070/27070_0_log.err +167 -0
- inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_27070/27070_0_log.out +0 -0
- inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_27070/27070_1_log.err +91 -0
- inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_27070/27070_1_log.out +218 -0
inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/best.pt
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:8cfe434d017e84085b955e7fce6a4dd28acb5dcda6aa5df2cbe78e27acea9787
|
| 3 |
+
size 5114658721
|
inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/e116_best_snapshot.pt
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:554b01d24c45b26a0071b4d8e7b24b3ffefa65c06c9789435aabfd301649bf28
|
| 3 |
+
size 5114658721
|
inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/e150_snapshot.pt
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:115962937c75e7eb1d496602782262e4784d3c378d047680ca8c9c813f8b3845
|
| 3 |
+
size 5114754241
|
inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/git-info.txt
ADDED
|
@@ -0,0 +1,2 @@
|
|
|
|
|
|
|
|
|
|
| 1 |
+
branch: main
|
| 2 |
+
commit: e82d40bb2e9210dc731666848badd739c694f1ee
|
inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_25163/25163_0_log.err
ADDED
|
@@ -0,0 +1,102 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
/var/scratch/dcanezil/.cache/uv/virtualenvs/vd/lib/python3.12/site-packages/timm/models/layers/__init__.py:49: FutureWarning: Importing from timm.models.layers is deprecated, please import via timm.layers
|
| 2 |
+
warnings.warn(f"Importing from {__name__} is deprecated, please import via timm.layers", FutureWarning)
|
| 3 |
+
wandb: Currently logged in as: dgcnz (uvjepa) to https://api.wandb.ai. Use `wandb login --relogin` to force relogin
|
| 4 |
+
wandb: setting up run vfclercb
|
| 5 |
+
wandb: Tracking run with wandb version 0.23.1
|
| 6 |
+
wandb: Run data is saved locally in /var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/wandb/run-20260611_004330-vfclercb
|
| 7 |
+
wandb: Run `wandb offline` to turn off syncing.
|
| 8 |
+
wandb: Syncing run d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow
|
| 9 |
+
wandb: ⭐️ View project at https://wandb.ai/uvjepa/vjepa_ablation
|
| 10 |
+
wandb: 🚀 View run at https://wandb.ai/uvjepa/vjepa_ablation/runs/vfclercb
|
| 11 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/utils.py:888: FutureWarning: `torch.cuda.amp.GradScaler(args...)` is deprecated. Please use `torch.amp.GradScaler('cuda', args...)` instead.
|
| 12 |
+
scaler = torch.cuda.amp.GradScaler() if mixed_precision else None
|
| 13 |
+
/var/scratch/dcanezil/.cache/uv/virtualenvs/vd/lib/python3.12/site-packages/torch/distributed/c10d_logger.py:83: UserWarning: barrier(): using the device under current context. You can specify `device_id` in `init_process_group` to mute this warning.
|
| 14 |
+
return func(*args, **kwargs)
|
| 15 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 16 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 17 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_flute/co50KUHacYw_000005_000015.avi'
|
| 18 |
+
warnings.warn(f"video path not found {fname=}")
|
| 19 |
+
[01:21:02] /github/workspace/src/video/video_reader.cc:83: ERROR opening: /local/dcanezil/data/kinetics/VideoData/long_jump/jsYiitkbr5w_000004_000014.avi, Invalid data found when processing input
|
| 20 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/frying_vegetables/Nr-zCXgW2Sw_000161_000171.avi'
|
| 21 |
+
warnings.warn(f"video path not found {fname=}")
|
| 22 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/baking_cookies/Q3XGZmqk1Q0_000296_000306.avi'
|
| 23 |
+
warnings.warn(f"video path not found {fname=}")
|
| 24 |
+
/var/scratch/dcanezil/.cache/uv/virtualenvs/vd/lib/python3.12/site-packages/torch/distributed/c10d_logger.py:83: UserWarning: barrier(): using the device under current context. You can specify `device_id` in `init_process_group` to mute this warning.
|
| 25 |
+
return func(*args, **kwargs)
|
| 26 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 27 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 28 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/windsurfing/i-gzh_BPDa8_000154_000164.avi'
|
| 29 |
+
warnings.warn(f"video path not found {fname=}")
|
| 30 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_poker/6-DyF8umej8_000097_000107.avi'
|
| 31 |
+
warnings.warn(f"video path not found {fname=}")
|
| 32 |
+
/var/scratch/dcanezil/.cache/uv/virtualenvs/vd/lib/python3.12/site-packages/torch/distributed/c10d_logger.py:83: UserWarning: barrier(): using the device under current context. You can specify `device_id` in `init_process_group` to mute this warning.
|
| 33 |
+
return func(*args, **kwargs)
|
| 34 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 35 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 36 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/sailing/ZolXG_-RgIs_000970_000980.avi'
|
| 37 |
+
warnings.warn(f"video path not found {fname=}")
|
| 38 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_monopoly/NLL667uPWVA_000066_000076.avi'
|
| 39 |
+
warnings.warn(f"video path not found {fname=}")
|
| 40 |
+
/var/scratch/dcanezil/.cache/uv/virtualenvs/vd/lib/python3.12/site-packages/torch/distributed/c10d_logger.py:83: UserWarning: barrier(): using the device under current context. You can specify `device_id` in `init_process_group` to mute this warning.
|
| 41 |
+
return func(*args, **kwargs)
|
| 42 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 43 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 44 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/robot_dancing/J1x0XlWa6HM_000147_000157.avi'
|
| 45 |
+
warnings.warn(f"video path not found {fname=}")
|
| 46 |
+
/var/scratch/dcanezil/.cache/uv/virtualenvs/vd/lib/python3.12/site-packages/torch/distributed/c10d_logger.py:83: UserWarning: barrier(): using the device under current context. You can specify `device_id` in `init_process_group` to mute this warning.
|
| 47 |
+
return func(*args, **kwargs)
|
| 48 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 49 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 50 |
+
/var/scratch/dcanezil/.cache/uv/virtualenvs/vd/lib/python3.12/site-packages/torch/distributed/c10d_logger.py:83: UserWarning: barrier(): using the device under current context. You can specify `device_id` in `init_process_group` to mute this warning.
|
| 51 |
+
return func(*args, **kwargs)
|
| 52 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 53 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 54 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/waxing_legs/cqUusXBODuw_001046_001056.avi'
|
| 55 |
+
warnings.warn(f"video path not found {fname=}")
|
| 56 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/dancing_macarena/uGQgxFHemyA_000059_000069.avi'
|
| 57 |
+
warnings.warn(f"video path not found {fname=}")
|
| 58 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/robot_dancing/J1x0XlWa6HM_000147_000157.avi'
|
| 59 |
+
warnings.warn(f"video path not found {fname=}")
|
| 60 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/making_tea/mtYFNsRcxY4_000063_000073.avi'
|
| 61 |
+
warnings.warn(f"video path not found {fname=}")
|
| 62 |
+
[12:02:17] /github/workspace/src/video/video_reader.cc:83: ERROR opening: /local/dcanezil/data/kinetics/VideoData/long_jump/jsYiitkbr5w_000004_000014.avi, Invalid data found when processing input
|
| 63 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/windsurfing/i-gzh_BPDa8_000154_000164.avi'
|
| 64 |
+
warnings.warn(f"video path not found {fname=}")
|
| 65 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/robot_dancing/J1x0XlWa6HM_000147_000157.avi'
|
| 66 |
+
warnings.warn(f"video path not found {fname=}")
|
| 67 |
+
/var/scratch/dcanezil/.cache/uv/virtualenvs/vd/lib/python3.12/site-packages/torch/distributed/c10d_logger.py:83: UserWarning: barrier(): using the device under current context. You can specify `device_id` in `init_process_group` to mute this warning.
|
| 68 |
+
return func(*args, **kwargs)
|
| 69 |
+
wandb: uploading output.log; uploading config.yaml
|
| 70 |
+
wandb: uploading history steps 930-930, summary
|
| 71 |
+
wandb:
|
| 72 |
+
wandb: Run history:
|
| 73 |
+
wandb: epoch ▁▁▁▁▁▁▁▁▂▂▂▂▃▃▃▃▄▄▄▄▅▅▅▅▅▅▅▆▆▆▆▆▆▆▆▇▇▇▇█
|
| 74 |
+
wandb: epoch/avg_data_time_ms ▄▅▅▄▄▄▄█▄▆▄▆▆▃▅▅▇▆▄▄▅▁▄▄▅▂▆▃▃▅
|
| 75 |
+
wandb: epoch/avg_gpu_time_ms ▄▄▃▅▅▆▆▅▃▃▃▅▄█▄▃▆▆▅▆█▁▂▂▆▄▄▃▇▇
|
| 76 |
+
wandb: epoch/avg_iter_time_ms ▄▄▄▅▅▆▆▆▃▄▃▅▄█▄▃▇▆▅▆█▁▂▂▆▄▅▃▇▇
|
| 77 |
+
wandb: epoch/avg_loss █▆▅▄▃▃▂▂▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁
|
| 78 |
+
wandb: knn/ssv2_coarse10_temporal_concat_top1 ▁▂▃▄▆▇█
|
| 79 |
+
wandb: knn/ssv2_coarse10_temporal_concat_top5 ▁▃▄▅▅▇█
|
| 80 |
+
wandb: linear/ssv2_coarse10_temporal_concat_top1 ▁▁▄▅▆▆█
|
| 81 |
+
wandb: linear/ssv2_coarse10_temporal_concat_top5 ▁▃▅▆▆██
|
| 82 |
+
wandb: train/ema/cos_sim ▁▃▄▄▄▆▆▇▇███████████████████████████████
|
| 83 |
+
wandb: +34 ...
|
| 84 |
+
wandb:
|
| 85 |
+
wandb: Run summary:
|
| 86 |
+
wandb: epoch 30
|
| 87 |
+
wandb: epoch/avg_data_time_ms 6.85255
|
| 88 |
+
wandb: epoch/avg_gpu_time_ms 4683.10447
|
| 89 |
+
wandb: epoch/avg_iter_time_ms 4696.56027
|
| 90 |
+
wandb: epoch/avg_loss 0.51306
|
| 91 |
+
wandb: knn/ssv2_coarse10_temporal_concat_top1 41.16466
|
| 92 |
+
wandb: knn/ssv2_coarse10_temporal_concat_top5 83.73494
|
| 93 |
+
wandb: linear/ssv2_coarse10_temporal_concat_top1 56.1245
|
| 94 |
+
wandb: linear/ssv2_coarse10_temporal_concat_top5 91.96787
|
| 95 |
+
wandb: train/ema/cos_sim 0.79508
|
| 96 |
+
wandb: +34 ...
|
| 97 |
+
wandb:
|
| 98 |
+
wandb: 🚀 View run d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow at: https://wandb.ai/uvjepa/vjepa_ablation/runs/vfclercb
|
| 99 |
+
wandb: ⭐️ View project at: https://wandb.ai/uvjepa/vjepa_ablation
|
| 100 |
+
wandb: Synced 5 W&B file(s), 0 media file(s), 0 artifact file(s) and 0 other file(s)
|
| 101 |
+
wandb: Find logs at: /var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/wandb/run-20260611_004330-vfclercb/logs
|
| 102 |
+
[rank0]:[W611 12:55:44.290842283 ProcessGroupNCCL.cpp:1553] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())
|
inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_25163/25163_0_log.out
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_25163/25163_1_log.err
ADDED
|
@@ -0,0 +1,42 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
/var/scratch/dcanezil/.cache/uv/virtualenvs/vd/lib/python3.12/site-packages/timm/models/layers/__init__.py:49: FutureWarning: Importing from timm.models.layers is deprecated, please import via timm.layers
|
| 2 |
+
warnings.warn(f"Importing from {__name__} is deprecated, please import via timm.layers", FutureWarning)
|
| 3 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/utils.py:888: FutureWarning: `torch.cuda.amp.GradScaler(args...)` is deprecated. Please use `torch.amp.GradScaler('cuda', args...)` instead.
|
| 4 |
+
scaler = torch.cuda.amp.GradScaler() if mixed_precision else None
|
| 5 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 6 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 7 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/dancing_macarena/uGQgxFHemyA_000059_000069.avi'
|
| 8 |
+
warnings.warn(f"video path not found {fname=}")
|
| 9 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 10 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 11 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/windsurfing/i-gzh_BPDa8_000154_000164.avi'
|
| 12 |
+
warnings.warn(f"video path not found {fname=}")
|
| 13 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/baking_cookies/Q3XGZmqk1Q0_000296_000306.avi'
|
| 14 |
+
warnings.warn(f"video path not found {fname=}")
|
| 15 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/robot_dancing/J1x0XlWa6HM_000147_000157.avi'
|
| 16 |
+
warnings.warn(f"video path not found {fname=}")
|
| 17 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/making_tea/mtYFNsRcxY4_000063_000073.avi'
|
| 18 |
+
warnings.warn(f"video path not found {fname=}")
|
| 19 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_poker/-jGxlNQKkeo_000092_000102.avi'
|
| 20 |
+
warnings.warn(f"video path not found {fname=}")
|
| 21 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 22 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 23 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/waxing_legs/cqUusXBODuw_001046_001056.avi'
|
| 24 |
+
warnings.warn(f"video path not found {fname=}")
|
| 25 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/water_skiing/4bhvIHWADVU_000022_000032.avi'
|
| 26 |
+
warnings.warn(f"video path not found {fname=}")
|
| 27 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 28 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 29 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/eating_carrots/eiZ8Hzc7FPU_000080_000090.avi'
|
| 30 |
+
warnings.warn(f"video path not found {fname=}")
|
| 31 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 32 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 33 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/water_skiing/4bhvIHWADVU_000022_000032.avi'
|
| 34 |
+
warnings.warn(f"video path not found {fname=}")
|
| 35 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/windsurfing/i-gzh_BPDa8_000154_000164.avi'
|
| 36 |
+
warnings.warn(f"video path not found {fname=}")
|
| 37 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 38 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 39 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_saxophone/Zs0_2GMEPXo_000054_000064.avi'
|
| 40 |
+
warnings.warn(f"video path not found {fname=}")
|
| 41 |
+
[12:37:51] /github/workspace/src/video/video_reader.cc:83: ERROR opening: /local/dcanezil/data/kinetics/VideoData/climbing_tree/QX5IzK4iFCM_000002_000012.avi, Invalid data found when processing input
|
| 42 |
+
[rank1]:[W611 12:55:28.432416152 ProcessGroupNCCL.cpp:1553] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())
|
inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_25163/25163_1_log.out
ADDED
|
@@ -0,0 +1,212 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
submitit INFO (2026-06-11 00:43:17,122) - Starting with JobEnvironment(job_id=25163, hostname=node407, local_rank=1(3), node=0(2), global_rank=1(6))
|
| 2 |
+
submitit INFO (2026-06-11 00:43:17,122) - Loading pickle: /var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_25163/25163_submitted.pkl
|
| 3 |
+
INFO:root:loaded pretrain params...
|
| 4 |
+
{ 'app': 'vjepa',
|
| 5 |
+
'cpus_per_task': 16,
|
| 6 |
+
'data': { 'batch_size': 32,
|
| 7 |
+
'crop_size': 224,
|
| 8 |
+
'dataset_fpcs': [8, 8],
|
| 9 |
+
'dataset_type': 'VideoDataset',
|
| 10 |
+
'datasets': [ '/local/dcanezil/data/kinetics/train.csv',
|
| 11 |
+
'/local/dcanezil/data/ssv2/train.csv'],
|
| 12 |
+
'datasets_weights': [0.65, 0.35],
|
| 13 |
+
'fps': 2,
|
| 14 |
+
'num_workers': 4,
|
| 15 |
+
'patch_size': 14,
|
| 16 |
+
'persistent_workers': True,
|
| 17 |
+
'pin_mem': True,
|
| 18 |
+
'tubelet_size': 1},
|
| 19 |
+
'data_aug': { 'auto_augment': False,
|
| 20 |
+
'motion_shift': False,
|
| 21 |
+
'random_resize_aspect_ratio': [0.75, 1.35],
|
| 22 |
+
'random_resize_scale': [0.3, 1.0],
|
| 23 |
+
'reprob': 0.0},
|
| 24 |
+
'folder': '/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow',
|
| 25 |
+
'loss': {'loss_exp': 1.0},
|
| 26 |
+
'mask': [ { 'aspect_ratio': [0.75, 1.5],
|
| 27 |
+
'full_complement': False,
|
| 28 |
+
'max_keep': None,
|
| 29 |
+
'max_temporal_keep': 1.0,
|
| 30 |
+
'num_blocks': 8,
|
| 31 |
+
'spatial_scale': [0.05, 0.05],
|
| 32 |
+
'temporal_scale': [1.0, 1.0]},
|
| 33 |
+
{ 'aspect_ratio': [0.75, 1.5],
|
| 34 |
+
'full_complement': False,
|
| 35 |
+
'max_keep': None,
|
| 36 |
+
'max_temporal_keep': 1.0,
|
| 37 |
+
'num_blocks': 2,
|
| 38 |
+
'spatial_scale': [0.4, 0.4],
|
| 39 |
+
'temporal_scale': [1.0, 1.0]}],
|
| 40 |
+
'mem_per_gpu': '40G',
|
| 41 |
+
'meta': { 'dtype': 'bfloat16',
|
| 42 |
+
'knn_eval_epoch0': True,
|
| 43 |
+
'knn_eval_freq': 5,
|
| 44 |
+
'knn_eval_presets': [ { 'config': { 'batch_size': 8,
|
| 45 |
+
'dataset_train': '/local/dcanezil/data/ssv2/train_coarse10.csv',
|
| 46 |
+
'dataset_val': '/local/dcanezil/data/ssv2/val_coarse10.csv',
|
| 47 |
+
'eval_videos_per_class': 100,
|
| 48 |
+
'linear_probe': True,
|
| 49 |
+
'num_workers': 2,
|
| 50 |
+
'pool_type': 'temporal_concat',
|
| 51 |
+
'train_videos_per_class': 500},
|
| 52 |
+
'preset': 'ssv2_coarse10'}],
|
| 53 |
+
'load_checkpoint': True,
|
| 54 |
+
'read_checkpoint': None,
|
| 55 |
+
'save_every_freq': 10,
|
| 56 |
+
'seed': 239,
|
| 57 |
+
'use_sdpa': True,
|
| 58 |
+
'use_wandb': True,
|
| 59 |
+
'wandb_project': 'vjepa_ablation'},
|
| 60 |
+
'metrics': {'sigreg': {}, 'std': {}},
|
| 61 |
+
'model': { 'class_token': True,
|
| 62 |
+
'freeze_backbone': False,
|
| 63 |
+
'is_causal': True,
|
| 64 |
+
'model_name': 'vit_large_patch14_dinov2_lvd142m',
|
| 65 |
+
'pe_type': '2d',
|
| 66 |
+
'pred_depth': 12,
|
| 67 |
+
'pred_embed_dim': 384,
|
| 68 |
+
'pred_num_heads': 6,
|
| 69 |
+
'predictor': 'v2_cross',
|
| 70 |
+
'stem_type': '2d',
|
| 71 |
+
'target_embed_dim': 1024,
|
| 72 |
+
'target_kind': 'ema',
|
| 73 |
+
'uniform_power': True,
|
| 74 |
+
'use_activation_checkpointing': True,
|
| 75 |
+
'use_alibi': False,
|
| 76 |
+
'use_lora': False,
|
| 77 |
+
'use_mask_tokens': True,
|
| 78 |
+
'use_rope': False,
|
| 79 |
+
'use_sdpa': True,
|
| 80 |
+
'zero_init_mask_tokens': True},
|
| 81 |
+
'nodes': 2,
|
| 82 |
+
'optimization': { 'clip_grad': None,
|
| 83 |
+
'ema': [1.0, 0.99925],
|
| 84 |
+
'ema_hold_epochs': 0,
|
| 85 |
+
'ema_ramp_epochs': 5,
|
| 86 |
+
'encoder_layer_decay': 0.9,
|
| 87 |
+
'epochs': 30,
|
| 88 |
+
'final_lr': 0.0001,
|
| 89 |
+
'final_weight_decay': 0.04,
|
| 90 |
+
'ipe': 300,
|
| 91 |
+
'ipe_scale': 1.0,
|
| 92 |
+
'lr': 0.0001,
|
| 93 |
+
'predictor_lr_scale': 3.0,
|
| 94 |
+
'start_lr': 1e-05,
|
| 95 |
+
'warmup': 2,
|
| 96 |
+
'weight_decay': 0.04},
|
| 97 |
+
'tasks_per_node': 3}
|
| 98 |
+
INFO:root:Running pre-training of app: vjepa
|
| 99 |
+
[INFO ][2026-06-11 00:43:25][app.vjepa.train ][main ] which_dtype='bfloat16'
|
| 100 |
+
[INFO ][2026-06-11 00:43:25][app.vjepa.train ][main ] Disabling persistent_workers (incompatible with KNN eval)
|
| 101 |
+
[INFO ][2026-06-11 00:43:25][app.vjepa.train ][main ] NCCL_SOCKET_IFNAME=None
|
| 102 |
+
[INFO ][2026-06-11 00:43:28][app.vjepa.train ][main ] Initialized (rank/world-size) 1/6, tasks_per_node=3
|
| 103 |
+
[WARNING ][2026-06-11 00:43:34][experiments.stmodels.vision_transformers_v3][_build_dinov2_vit ] DINOv2 is an image model — fwd_chunk is unset so all frames will be processed jointly. Set fwd_chunk=1 for per-frame (standard) processing.
|
| 104 |
+
[INFO ][2026-06-11 00:43:40][experiments.stmodels.vision_transformers_v3][_build_dinov2_vit ] Loading pretrained weights for vit_large_patch14_dinov2_lvd142m from timm (vit_large_patch14_dinov2)
|
| 105 |
+
[INFO ][2026-06-11 00:43:43][timm.models._builder][load_pretrained ] Loading pretrained weights from Hugging Face hub (timm/vit_large_patch14_dinov2.lvd142m)
|
| 106 |
+
[INFO ][2026-06-11 00:43:44][timm.models._hub ][load_state_dict_from_hf ] [timm/vit_large_patch14_dinov2.lvd142m] Safe alternative available for 'pytorch_model.bin' (as 'model.safetensors'). Loading weights using safetensors.
|
| 107 |
+
[INFO ][2026-06-11 00:43:44][timm.layers.pos_embed][resample_abs_pos_embed ] Resized position embedding: (37, 37) to (16, 16).
|
| 108 |
+
[INFO ][2026-06-11 00:43:46][root ][init_video_model ] ViTMultiSeqWrapper(
|
| 109 |
+
(backbone): VisionTransformer(
|
| 110 |
+
(patch_embed): PatchEmbed3D(
|
| 111 |
+
(proj): Conv3d(3, 1024, kernel_size=(1, 14, 14), stride=(1, 14, 14))
|
| 112 |
+
)
|
| 113 |
+
(blocks): ModuleList(
|
| 114 |
+
(0-23): 24 x Block(
|
| 115 |
+
(norm1): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 116 |
+
(norm2): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 117 |
+
(drop_path1): Identity()
|
| 118 |
+
(drop_path2): Identity()
|
| 119 |
+
(attn): Attention(
|
| 120 |
+
(qkv): Linear(in_features=1024, out_features=3072, bias=True)
|
| 121 |
+
(attn_drop): Dropout(p=0.0, inplace=False)
|
| 122 |
+
(proj): Linear(in_features=1024, out_features=1024, bias=True)
|
| 123 |
+
(proj_drop): Dropout(p=0.0, inplace=False)
|
| 124 |
+
)
|
| 125 |
+
(mlp): MLP(
|
| 126 |
+
(fc1): Linear(in_features=1024, out_features=4096, bias=True)
|
| 127 |
+
(act): GELU(approximate='none')
|
| 128 |
+
(fc2): Linear(in_features=4096, out_features=1024, bias=True)
|
| 129 |
+
(drop): Dropout(p=0.0, inplace=False)
|
| 130 |
+
)
|
| 131 |
+
)
|
| 132 |
+
)
|
| 133 |
+
(norm): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 134 |
+
)
|
| 135 |
+
)
|
| 136 |
+
[INFO ][2026-06-11 00:43:46][root ][init_video_model ] VJEPAPredictorMultiSeqWrapper(
|
| 137 |
+
(backbone): PredictorV2(
|
| 138 |
+
(predictor_embed): Linear(in_features=1024, out_features=384, bias=True)
|
| 139 |
+
(mask_tokens): ParameterList(
|
| 140 |
+
(0): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
|
| 141 |
+
(1): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
|
| 142 |
+
(2): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
|
| 143 |
+
(3): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
|
| 144 |
+
)
|
| 145 |
+
(predictor_blocks): ModuleList(
|
| 146 |
+
(0-11): 12 x Block(
|
| 147 |
+
(residual1): EfficientResidual(
|
| 148 |
+
(norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
|
| 149 |
+
(fn): Attention(
|
| 150 |
+
(q_proj): Linear(in_features=384, out_features=384, bias=False)
|
| 151 |
+
(k_proj): Linear(in_features=384, out_features=384, bias=False)
|
| 152 |
+
(v_proj): Linear(in_features=384, out_features=384, bias=False)
|
| 153 |
+
(proj): Linear(in_features=384, out_features=384, bias=False)
|
| 154 |
+
(rope): Rope()
|
| 155 |
+
)
|
| 156 |
+
)
|
| 157 |
+
(residual2): EfficientResidual(
|
| 158 |
+
(norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
|
| 159 |
+
(fn): MLP(
|
| 160 |
+
(fc1): Linear(in_features=384, out_features=1536, bias=False)
|
| 161 |
+
(act): GELU(approximate='none')
|
| 162 |
+
(fc2): Linear(in_features=1536, out_features=384, bias=False)
|
| 163 |
+
(drop): Dropout(p=0.0, inplace=False)
|
| 164 |
+
)
|
| 165 |
+
)
|
| 166 |
+
)
|
| 167 |
+
)
|
| 168 |
+
(predictor_norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
|
| 169 |
+
(predictor_proj): Linear(in_features=384, out_features=1024, bias=True)
|
| 170 |
+
)
|
| 171 |
+
)
|
| 172 |
+
[INFO ][2026-06-11 00:43:46][root ][init_video_model ] ViTMultiSeqWrapper(
|
| 173 |
+
(backbone): VisionTransformer(
|
| 174 |
+
(patch_embed): PatchEmbed3D(
|
| 175 |
+
(proj): Conv3d(3, 1024, kernel_size=(1, 14, 14), stride=(1, 14, 14))
|
| 176 |
+
)
|
| 177 |
+
(blocks): ModuleList(
|
| 178 |
+
(0-23): 24 x Block(
|
| 179 |
+
(norm1): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 180 |
+
(norm2): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 181 |
+
(drop_path1): Identity()
|
| 182 |
+
(drop_path2): Identity()
|
| 183 |
+
(attn): Attention(
|
| 184 |
+
(qkv): Linear(in_features=1024, out_features=3072, bias=True)
|
| 185 |
+
(attn_drop): Dropout(p=0.0, inplace=False)
|
| 186 |
+
(proj): Linear(in_features=1024, out_features=1024, bias=True)
|
| 187 |
+
(proj_drop): Dropout(p=0.0, inplace=False)
|
| 188 |
+
)
|
| 189 |
+
(mlp): MLP(
|
| 190 |
+
(fc1): Linear(in_features=1024, out_features=4096, bias=True)
|
| 191 |
+
(act): GELU(approximate='none')
|
| 192 |
+
(fc2): Linear(in_features=4096, out_features=1024, bias=True)
|
| 193 |
+
(drop): Dropout(p=0.0, inplace=False)
|
| 194 |
+
)
|
| 195 |
+
)
|
| 196 |
+
)
|
| 197 |
+
(norm): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 198 |
+
)
|
| 199 |
+
)
|
| 200 |
+
[INFO ][2026-06-11 00:43:46][root ][init_video_model ] Encoder number of parameters: 302964736
|
| 201 |
+
[INFO ][2026-06-11 00:43:46][root ][init_video_model ] Predictor number of parameters: 22042240
|
| 202 |
+
[INFO ][2026-06-11 00:43:46][root ][init_video_model ] Target encoder number of parameters: 302964736
|
| 203 |
+
[INFO ][2026-06-11 00:43:47][root ][make_videodataset ] VideoDataset dataset created
|
| 204 |
+
[INFO ][2026-06-11 00:43:47][WeightedSampler ][__init__ ] Using DistributedWeightedSampler with rank 1 / 6
|
| 205 |
+
[INFO ][2026-06-11 00:43:47][root ][make_videodataset ] VideoDataset unsupervised data loader created
|
| 206 |
+
[INFO ][2026-06-11 00:43:47][app.vjepa.train ][main ] iterations per epoch/dataset length: 300/2103
|
| 207 |
+
[INFO ][2026-06-11 00:43:47][app.vjepa.train ][main ] Wrapping models in DDP (rank 1)...
|
| 208 |
+
[INFO ][2026-06-11 00:45:06][app.vjepa.train ][main ] Initializing loader...
|
| 209 |
+
submitit INFO (2026-06-11 12:55:27,874) - Job completed successfully
|
| 210 |
+
[INFO ][2026-06-11 12:55:27][submitit ][process_job ] Job completed successfully
|
| 211 |
+
submitit INFO (2026-06-11 12:55:27,880) - Exiting after successful completion
|
| 212 |
+
[INFO ][2026-06-11 12:55:27][submitit ][process_job ] Exiting after successful completion
|
inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_25163/25163_2_log.err
ADDED
|
@@ -0,0 +1,36 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
/var/scratch/dcanezil/.cache/uv/virtualenvs/vd/lib/python3.12/site-packages/timm/models/layers/__init__.py:49: FutureWarning: Importing from timm.models.layers is deprecated, please import via timm.layers
|
| 2 |
+
warnings.warn(f"Importing from {__name__} is deprecated, please import via timm.layers", FutureWarning)
|
| 3 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/utils.py:888: FutureWarning: `torch.cuda.amp.GradScaler(args...)` is deprecated. Please use `torch.amp.GradScaler('cuda', args...)` instead.
|
| 4 |
+
scaler = torch.cuda.amp.GradScaler() if mixed_precision else None
|
| 5 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 6 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 7 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/eating_carrots/eiZ8Hzc7FPU_000080_000090.avi'
|
| 8 |
+
warnings.warn(f"video path not found {fname=}")
|
| 9 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/making_tea/mtYFNsRcxY4_000063_000073.avi'
|
| 10 |
+
warnings.warn(f"video path not found {fname=}")
|
| 11 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 12 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 13 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/sailing/ZolXG_-RgIs_000970_000980.avi'
|
| 14 |
+
warnings.warn(f"video path not found {fname=}")
|
| 15 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/windsurfing/i-gzh_BPDa8_000154_000164.avi'
|
| 16 |
+
warnings.warn(f"video path not found {fname=}")
|
| 17 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 18 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 19 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/picking_fruit/3VvkoFtPCCU_000045_000055.avi'
|
| 20 |
+
warnings.warn(f"video path not found {fname=}")
|
| 21 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 22 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 23 |
+
[07:07:33] /github/workspace/src/video/video_reader.cc:83: ERROR opening: /local/dcanezil/data/kinetics/VideoData/climbing_tree/QX5IzK4iFCM_000002_000012.avi, Invalid data found when processing input
|
| 24 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/waxing_legs/cqUusXBODuw_001046_001056.avi'
|
| 25 |
+
warnings.warn(f"video path not found {fname=}")
|
| 26 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 27 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 28 |
+
[09:22:17] /github/workspace/src/video/video_reader.cc:83: ERROR opening: /local/dcanezil/data/kinetics/VideoData/trimming_trees/T6dqaZioaXs_000004_000014.avi, Invalid data found when processing input
|
| 29 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/dancing_macarena/uGQgxFHemyA_000059_000069.avi'
|
| 30 |
+
warnings.warn(f"video path not found {fname=}")
|
| 31 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/making_tea/mtYFNsRcxY4_000063_000073.avi'
|
| 32 |
+
warnings.warn(f"video path not found {fname=}")
|
| 33 |
+
[10:48:59] /github/workspace/src/video/video_reader.cc:83: ERROR opening: /local/dcanezil/data/kinetics/VideoData/long_jump/jsYiitkbr5w_000004_000014.avi, Invalid data found when processing input
|
| 34 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 35 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 36 |
+
[rank2]:[W611 12:55:28.433278605 ProcessGroupNCCL.cpp:1553] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())
|
inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_25163/25163_2_log.out
ADDED
|
@@ -0,0 +1,212 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
submitit INFO (2026-06-11 00:43:17,122) - Starting with JobEnvironment(job_id=25163, hostname=node407, local_rank=2(3), node=0(2), global_rank=2(6))
|
| 2 |
+
submitit INFO (2026-06-11 00:43:17,122) - Loading pickle: /var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_25163/25163_submitted.pkl
|
| 3 |
+
INFO:root:loaded pretrain params...
|
| 4 |
+
{ 'app': 'vjepa',
|
| 5 |
+
'cpus_per_task': 16,
|
| 6 |
+
'data': { 'batch_size': 32,
|
| 7 |
+
'crop_size': 224,
|
| 8 |
+
'dataset_fpcs': [8, 8],
|
| 9 |
+
'dataset_type': 'VideoDataset',
|
| 10 |
+
'datasets': [ '/local/dcanezil/data/kinetics/train.csv',
|
| 11 |
+
'/local/dcanezil/data/ssv2/train.csv'],
|
| 12 |
+
'datasets_weights': [0.65, 0.35],
|
| 13 |
+
'fps': 2,
|
| 14 |
+
'num_workers': 4,
|
| 15 |
+
'patch_size': 14,
|
| 16 |
+
'persistent_workers': True,
|
| 17 |
+
'pin_mem': True,
|
| 18 |
+
'tubelet_size': 1},
|
| 19 |
+
'data_aug': { 'auto_augment': False,
|
| 20 |
+
'motion_shift': False,
|
| 21 |
+
'random_resize_aspect_ratio': [0.75, 1.35],
|
| 22 |
+
'random_resize_scale': [0.3, 1.0],
|
| 23 |
+
'reprob': 0.0},
|
| 24 |
+
'folder': '/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow',
|
| 25 |
+
'loss': {'loss_exp': 1.0},
|
| 26 |
+
'mask': [ { 'aspect_ratio': [0.75, 1.5],
|
| 27 |
+
'full_complement': False,
|
| 28 |
+
'max_keep': None,
|
| 29 |
+
'max_temporal_keep': 1.0,
|
| 30 |
+
'num_blocks': 8,
|
| 31 |
+
'spatial_scale': [0.05, 0.05],
|
| 32 |
+
'temporal_scale': [1.0, 1.0]},
|
| 33 |
+
{ 'aspect_ratio': [0.75, 1.5],
|
| 34 |
+
'full_complement': False,
|
| 35 |
+
'max_keep': None,
|
| 36 |
+
'max_temporal_keep': 1.0,
|
| 37 |
+
'num_blocks': 2,
|
| 38 |
+
'spatial_scale': [0.4, 0.4],
|
| 39 |
+
'temporal_scale': [1.0, 1.0]}],
|
| 40 |
+
'mem_per_gpu': '40G',
|
| 41 |
+
'meta': { 'dtype': 'bfloat16',
|
| 42 |
+
'knn_eval_epoch0': True,
|
| 43 |
+
'knn_eval_freq': 5,
|
| 44 |
+
'knn_eval_presets': [ { 'config': { 'batch_size': 8,
|
| 45 |
+
'dataset_train': '/local/dcanezil/data/ssv2/train_coarse10.csv',
|
| 46 |
+
'dataset_val': '/local/dcanezil/data/ssv2/val_coarse10.csv',
|
| 47 |
+
'eval_videos_per_class': 100,
|
| 48 |
+
'linear_probe': True,
|
| 49 |
+
'num_workers': 2,
|
| 50 |
+
'pool_type': 'temporal_concat',
|
| 51 |
+
'train_videos_per_class': 500},
|
| 52 |
+
'preset': 'ssv2_coarse10'}],
|
| 53 |
+
'load_checkpoint': True,
|
| 54 |
+
'read_checkpoint': None,
|
| 55 |
+
'save_every_freq': 10,
|
| 56 |
+
'seed': 239,
|
| 57 |
+
'use_sdpa': True,
|
| 58 |
+
'use_wandb': True,
|
| 59 |
+
'wandb_project': 'vjepa_ablation'},
|
| 60 |
+
'metrics': {'sigreg': {}, 'std': {}},
|
| 61 |
+
'model': { 'class_token': True,
|
| 62 |
+
'freeze_backbone': False,
|
| 63 |
+
'is_causal': True,
|
| 64 |
+
'model_name': 'vit_large_patch14_dinov2_lvd142m',
|
| 65 |
+
'pe_type': '2d',
|
| 66 |
+
'pred_depth': 12,
|
| 67 |
+
'pred_embed_dim': 384,
|
| 68 |
+
'pred_num_heads': 6,
|
| 69 |
+
'predictor': 'v2_cross',
|
| 70 |
+
'stem_type': '2d',
|
| 71 |
+
'target_embed_dim': 1024,
|
| 72 |
+
'target_kind': 'ema',
|
| 73 |
+
'uniform_power': True,
|
| 74 |
+
'use_activation_checkpointing': True,
|
| 75 |
+
'use_alibi': False,
|
| 76 |
+
'use_lora': False,
|
| 77 |
+
'use_mask_tokens': True,
|
| 78 |
+
'use_rope': False,
|
| 79 |
+
'use_sdpa': True,
|
| 80 |
+
'zero_init_mask_tokens': True},
|
| 81 |
+
'nodes': 2,
|
| 82 |
+
'optimization': { 'clip_grad': None,
|
| 83 |
+
'ema': [1.0, 0.99925],
|
| 84 |
+
'ema_hold_epochs': 0,
|
| 85 |
+
'ema_ramp_epochs': 5,
|
| 86 |
+
'encoder_layer_decay': 0.9,
|
| 87 |
+
'epochs': 30,
|
| 88 |
+
'final_lr': 0.0001,
|
| 89 |
+
'final_weight_decay': 0.04,
|
| 90 |
+
'ipe': 300,
|
| 91 |
+
'ipe_scale': 1.0,
|
| 92 |
+
'lr': 0.0001,
|
| 93 |
+
'predictor_lr_scale': 3.0,
|
| 94 |
+
'start_lr': 1e-05,
|
| 95 |
+
'warmup': 2,
|
| 96 |
+
'weight_decay': 0.04},
|
| 97 |
+
'tasks_per_node': 3}
|
| 98 |
+
INFO:root:Running pre-training of app: vjepa
|
| 99 |
+
[INFO ][2026-06-11 00:43:25][app.vjepa.train ][main ] which_dtype='bfloat16'
|
| 100 |
+
[INFO ][2026-06-11 00:43:25][app.vjepa.train ][main ] Disabling persistent_workers (incompatible with KNN eval)
|
| 101 |
+
[INFO ][2026-06-11 00:43:25][app.vjepa.train ][main ] NCCL_SOCKET_IFNAME=None
|
| 102 |
+
[INFO ][2026-06-11 00:43:28][app.vjepa.train ][main ] Initialized (rank/world-size) 2/6, tasks_per_node=3
|
| 103 |
+
[WARNING ][2026-06-11 00:43:34][experiments.stmodels.vision_transformers_v3][_build_dinov2_vit ] DINOv2 is an image model — fwd_chunk is unset so all frames will be processed jointly. Set fwd_chunk=1 for per-frame (standard) processing.
|
| 104 |
+
[INFO ][2026-06-11 00:43:40][experiments.stmodels.vision_transformers_v3][_build_dinov2_vit ] Loading pretrained weights for vit_large_patch14_dinov2_lvd142m from timm (vit_large_patch14_dinov2)
|
| 105 |
+
[INFO ][2026-06-11 00:43:44][timm.models._builder][load_pretrained ] Loading pretrained weights from Hugging Face hub (timm/vit_large_patch14_dinov2.lvd142m)
|
| 106 |
+
[INFO ][2026-06-11 00:43:44][timm.models._hub ][load_state_dict_from_hf ] [timm/vit_large_patch14_dinov2.lvd142m] Safe alternative available for 'pytorch_model.bin' (as 'model.safetensors'). Loading weights using safetensors.
|
| 107 |
+
[INFO ][2026-06-11 00:43:44][timm.layers.pos_embed][resample_abs_pos_embed ] Resized position embedding: (37, 37) to (16, 16).
|
| 108 |
+
[INFO ][2026-06-11 00:43:46][root ][init_video_model ] ViTMultiSeqWrapper(
|
| 109 |
+
(backbone): VisionTransformer(
|
| 110 |
+
(patch_embed): PatchEmbed3D(
|
| 111 |
+
(proj): Conv3d(3, 1024, kernel_size=(1, 14, 14), stride=(1, 14, 14))
|
| 112 |
+
)
|
| 113 |
+
(blocks): ModuleList(
|
| 114 |
+
(0-23): 24 x Block(
|
| 115 |
+
(norm1): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 116 |
+
(norm2): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 117 |
+
(drop_path1): Identity()
|
| 118 |
+
(drop_path2): Identity()
|
| 119 |
+
(attn): Attention(
|
| 120 |
+
(qkv): Linear(in_features=1024, out_features=3072, bias=True)
|
| 121 |
+
(attn_drop): Dropout(p=0.0, inplace=False)
|
| 122 |
+
(proj): Linear(in_features=1024, out_features=1024, bias=True)
|
| 123 |
+
(proj_drop): Dropout(p=0.0, inplace=False)
|
| 124 |
+
)
|
| 125 |
+
(mlp): MLP(
|
| 126 |
+
(fc1): Linear(in_features=1024, out_features=4096, bias=True)
|
| 127 |
+
(act): GELU(approximate='none')
|
| 128 |
+
(fc2): Linear(in_features=4096, out_features=1024, bias=True)
|
| 129 |
+
(drop): Dropout(p=0.0, inplace=False)
|
| 130 |
+
)
|
| 131 |
+
)
|
| 132 |
+
)
|
| 133 |
+
(norm): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 134 |
+
)
|
| 135 |
+
)
|
| 136 |
+
[INFO ][2026-06-11 00:43:46][root ][init_video_model ] VJEPAPredictorMultiSeqWrapper(
|
| 137 |
+
(backbone): PredictorV2(
|
| 138 |
+
(predictor_embed): Linear(in_features=1024, out_features=384, bias=True)
|
| 139 |
+
(mask_tokens): ParameterList(
|
| 140 |
+
(0): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
|
| 141 |
+
(1): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
|
| 142 |
+
(2): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
|
| 143 |
+
(3): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
|
| 144 |
+
)
|
| 145 |
+
(predictor_blocks): ModuleList(
|
| 146 |
+
(0-11): 12 x Block(
|
| 147 |
+
(residual1): EfficientResidual(
|
| 148 |
+
(norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
|
| 149 |
+
(fn): Attention(
|
| 150 |
+
(q_proj): Linear(in_features=384, out_features=384, bias=False)
|
| 151 |
+
(k_proj): Linear(in_features=384, out_features=384, bias=False)
|
| 152 |
+
(v_proj): Linear(in_features=384, out_features=384, bias=False)
|
| 153 |
+
(proj): Linear(in_features=384, out_features=384, bias=False)
|
| 154 |
+
(rope): Rope()
|
| 155 |
+
)
|
| 156 |
+
)
|
| 157 |
+
(residual2): EfficientResidual(
|
| 158 |
+
(norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
|
| 159 |
+
(fn): MLP(
|
| 160 |
+
(fc1): Linear(in_features=384, out_features=1536, bias=False)
|
| 161 |
+
(act): GELU(approximate='none')
|
| 162 |
+
(fc2): Linear(in_features=1536, out_features=384, bias=False)
|
| 163 |
+
(drop): Dropout(p=0.0, inplace=False)
|
| 164 |
+
)
|
| 165 |
+
)
|
| 166 |
+
)
|
| 167 |
+
)
|
| 168 |
+
(predictor_norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
|
| 169 |
+
(predictor_proj): Linear(in_features=384, out_features=1024, bias=True)
|
| 170 |
+
)
|
| 171 |
+
)
|
| 172 |
+
[INFO ][2026-06-11 00:43:46][root ][init_video_model ] ViTMultiSeqWrapper(
|
| 173 |
+
(backbone): VisionTransformer(
|
| 174 |
+
(patch_embed): PatchEmbed3D(
|
| 175 |
+
(proj): Conv3d(3, 1024, kernel_size=(1, 14, 14), stride=(1, 14, 14))
|
| 176 |
+
)
|
| 177 |
+
(blocks): ModuleList(
|
| 178 |
+
(0-23): 24 x Block(
|
| 179 |
+
(norm1): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 180 |
+
(norm2): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 181 |
+
(drop_path1): Identity()
|
| 182 |
+
(drop_path2): Identity()
|
| 183 |
+
(attn): Attention(
|
| 184 |
+
(qkv): Linear(in_features=1024, out_features=3072, bias=True)
|
| 185 |
+
(attn_drop): Dropout(p=0.0, inplace=False)
|
| 186 |
+
(proj): Linear(in_features=1024, out_features=1024, bias=True)
|
| 187 |
+
(proj_drop): Dropout(p=0.0, inplace=False)
|
| 188 |
+
)
|
| 189 |
+
(mlp): MLP(
|
| 190 |
+
(fc1): Linear(in_features=1024, out_features=4096, bias=True)
|
| 191 |
+
(act): GELU(approximate='none')
|
| 192 |
+
(fc2): Linear(in_features=4096, out_features=1024, bias=True)
|
| 193 |
+
(drop): Dropout(p=0.0, inplace=False)
|
| 194 |
+
)
|
| 195 |
+
)
|
| 196 |
+
)
|
| 197 |
+
(norm): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 198 |
+
)
|
| 199 |
+
)
|
| 200 |
+
[INFO ][2026-06-11 00:43:46][root ][init_video_model ] Encoder number of parameters: 302964736
|
| 201 |
+
[INFO ][2026-06-11 00:43:46][root ][init_video_model ] Predictor number of parameters: 22042240
|
| 202 |
+
[INFO ][2026-06-11 00:43:46][root ][init_video_model ] Target encoder number of parameters: 302964736
|
| 203 |
+
[INFO ][2026-06-11 00:43:47][root ][make_videodataset ] VideoDataset dataset created
|
| 204 |
+
[INFO ][2026-06-11 00:43:47][WeightedSampler ][__init__ ] Using DistributedWeightedSampler with rank 2 / 6
|
| 205 |
+
[INFO ][2026-06-11 00:43:47][root ][make_videodataset ] VideoDataset unsupervised data loader created
|
| 206 |
+
[INFO ][2026-06-11 00:43:47][app.vjepa.train ][main ] iterations per epoch/dataset length: 300/2103
|
| 207 |
+
[INFO ][2026-06-11 00:43:47][app.vjepa.train ][main ] Wrapping models in DDP (rank 2)...
|
| 208 |
+
[INFO ][2026-06-11 00:45:06][app.vjepa.train ][main ] Initializing loader...
|
| 209 |
+
submitit INFO (2026-06-11 12:55:27,888) - Job completed successfully
|
| 210 |
+
[INFO ][2026-06-11 12:55:27][submitit ][process_job ] Job completed successfully
|
| 211 |
+
submitit INFO (2026-06-11 12:55:27,890) - Exiting after successful completion
|
| 212 |
+
[INFO ][2026-06-11 12:55:27][submitit ][process_job ] Exiting after successful completion
|
inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_25163/25163_3_log.err
ADDED
|
@@ -0,0 +1,48 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
/var/scratch/dcanezil/.cache/uv/virtualenvs/vd/lib/python3.12/site-packages/timm/models/layers/__init__.py:49: FutureWarning: Importing from timm.models.layers is deprecated, please import via timm.layers
|
| 2 |
+
warnings.warn(f"Importing from {__name__} is deprecated, please import via timm.layers", FutureWarning)
|
| 3 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/utils.py:888: FutureWarning: `torch.cuda.amp.GradScaler(args...)` is deprecated. Please use `torch.amp.GradScaler('cuda', args...)` instead.
|
| 4 |
+
scaler = torch.cuda.amp.GradScaler() if mixed_precision else None
|
| 5 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 6 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 7 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_saxophone/Zs0_2GMEPXo_000054_000064.avi'
|
| 8 |
+
warnings.warn(f"video path not found {fname=}")
|
| 9 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_monopoly/NLL667uPWVA_000066_000076.avi'
|
| 10 |
+
warnings.warn(f"video path not found {fname=}")
|
| 11 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/pumping_fist/wKsuyr6Xk30_000094_000104.avi'
|
| 12 |
+
warnings.warn(f"video path not found {fname=}")
|
| 13 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_monopoly/NLL667uPWVA_000066_000076.avi'
|
| 14 |
+
warnings.warn(f"video path not found {fname=}")
|
| 15 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/picking_fruit/3VvkoFtPCCU_000045_000055.avi'
|
| 16 |
+
warnings.warn(f"video path not found {fname=}")
|
| 17 |
+
[04:07:16] /github/workspace/src/video/video_reader.cc:83: ERROR opening: /local/dcanezil/data/kinetics/VideoData/climbing_tree/QX5IzK4iFCM_000002_000012.avi, Invalid data found when processing input
|
| 18 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/waxing_legs/cqUusXBODuw_001046_001056.avi'
|
| 19 |
+
warnings.warn(f"video path not found {fname=}")
|
| 20 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/robot_dancing/J1x0XlWa6HM_000147_000157.avi'
|
| 21 |
+
warnings.warn(f"video path not found {fname=}")
|
| 22 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/water_skiing/4bhvIHWADVU_000022_000032.avi'
|
| 23 |
+
warnings.warn(f"video path not found {fname=}")
|
| 24 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/eating_carrots/eiZ8Hzc7FPU_000080_000090.avi'
|
| 25 |
+
warnings.warn(f"video path not found {fname=}")
|
| 26 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_saxophone/Zs0_2GMEPXo_000054_000064.avi'
|
| 27 |
+
warnings.warn(f"video path not found {fname=}")
|
| 28 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/windsurfing/i-gzh_BPDa8_000154_000164.avi'
|
| 29 |
+
warnings.warn(f"video path not found {fname=}")
|
| 30 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/surfing_water/5l5Pdd96Pao_000161_000171.avi'
|
| 31 |
+
warnings.warn(f"video path not found {fname=}")
|
| 32 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/surfing_water/5l5Pdd96Pao_000161_000171.avi'
|
| 33 |
+
warnings.warn(f"video path not found {fname=}")
|
| 34 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_monopoly/NLL667uPWVA_000066_000076.avi'
|
| 35 |
+
warnings.warn(f"video path not found {fname=}")
|
| 36 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_poker/-jGxlNQKkeo_000092_000102.avi'
|
| 37 |
+
warnings.warn(f"video path not found {fname=}")
|
| 38 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/baking_cookies/Q3XGZmqk1Q0_000296_000306.avi'
|
| 39 |
+
warnings.warn(f"video path not found {fname=}")
|
| 40 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/bowling/f4eb0wOlspM_000053_000063.avi'
|
| 41 |
+
warnings.warn(f"video path not found {fname=}")
|
| 42 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_monopoly/NLL667uPWVA_000066_000076.avi'
|
| 43 |
+
warnings.warn(f"video path not found {fname=}")
|
| 44 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/eating_carrots/eiZ8Hzc7FPU_000080_000090.avi'
|
| 45 |
+
warnings.warn(f"video path not found {fname=}")
|
| 46 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/surfing_water/5l5Pdd96Pao_000161_000171.avi'
|
| 47 |
+
warnings.warn(f"video path not found {fname=}")
|
| 48 |
+
[rank3]:[W611 12:55:28.458244544 ProcessGroupNCCL.cpp:1553] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())
|
inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_25163/25163_3_log.out
ADDED
|
@@ -0,0 +1,212 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
submitit INFO (2026-06-11 00:43:17,099) - Starting with JobEnvironment(job_id=25163, hostname=node411, local_rank=0(3), node=1(2), global_rank=3(6))
|
| 2 |
+
submitit INFO (2026-06-11 00:43:17,099) - Loading pickle: /var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_25163/25163_submitted.pkl
|
| 3 |
+
INFO:root:loaded pretrain params...
|
| 4 |
+
{ 'app': 'vjepa',
|
| 5 |
+
'cpus_per_task': 16,
|
| 6 |
+
'data': { 'batch_size': 32,
|
| 7 |
+
'crop_size': 224,
|
| 8 |
+
'dataset_fpcs': [8, 8],
|
| 9 |
+
'dataset_type': 'VideoDataset',
|
| 10 |
+
'datasets': [ '/local/dcanezil/data/kinetics/train.csv',
|
| 11 |
+
'/local/dcanezil/data/ssv2/train.csv'],
|
| 12 |
+
'datasets_weights': [0.65, 0.35],
|
| 13 |
+
'fps': 2,
|
| 14 |
+
'num_workers': 4,
|
| 15 |
+
'patch_size': 14,
|
| 16 |
+
'persistent_workers': True,
|
| 17 |
+
'pin_mem': True,
|
| 18 |
+
'tubelet_size': 1},
|
| 19 |
+
'data_aug': { 'auto_augment': False,
|
| 20 |
+
'motion_shift': False,
|
| 21 |
+
'random_resize_aspect_ratio': [0.75, 1.35],
|
| 22 |
+
'random_resize_scale': [0.3, 1.0],
|
| 23 |
+
'reprob': 0.0},
|
| 24 |
+
'folder': '/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow',
|
| 25 |
+
'loss': {'loss_exp': 1.0},
|
| 26 |
+
'mask': [ { 'aspect_ratio': [0.75, 1.5],
|
| 27 |
+
'full_complement': False,
|
| 28 |
+
'max_keep': None,
|
| 29 |
+
'max_temporal_keep': 1.0,
|
| 30 |
+
'num_blocks': 8,
|
| 31 |
+
'spatial_scale': [0.05, 0.05],
|
| 32 |
+
'temporal_scale': [1.0, 1.0]},
|
| 33 |
+
{ 'aspect_ratio': [0.75, 1.5],
|
| 34 |
+
'full_complement': False,
|
| 35 |
+
'max_keep': None,
|
| 36 |
+
'max_temporal_keep': 1.0,
|
| 37 |
+
'num_blocks': 2,
|
| 38 |
+
'spatial_scale': [0.4, 0.4],
|
| 39 |
+
'temporal_scale': [1.0, 1.0]}],
|
| 40 |
+
'mem_per_gpu': '40G',
|
| 41 |
+
'meta': { 'dtype': 'bfloat16',
|
| 42 |
+
'knn_eval_epoch0': True,
|
| 43 |
+
'knn_eval_freq': 5,
|
| 44 |
+
'knn_eval_presets': [ { 'config': { 'batch_size': 8,
|
| 45 |
+
'dataset_train': '/local/dcanezil/data/ssv2/train_coarse10.csv',
|
| 46 |
+
'dataset_val': '/local/dcanezil/data/ssv2/val_coarse10.csv',
|
| 47 |
+
'eval_videos_per_class': 100,
|
| 48 |
+
'linear_probe': True,
|
| 49 |
+
'num_workers': 2,
|
| 50 |
+
'pool_type': 'temporal_concat',
|
| 51 |
+
'train_videos_per_class': 500},
|
| 52 |
+
'preset': 'ssv2_coarse10'}],
|
| 53 |
+
'load_checkpoint': True,
|
| 54 |
+
'read_checkpoint': None,
|
| 55 |
+
'save_every_freq': 10,
|
| 56 |
+
'seed': 239,
|
| 57 |
+
'use_sdpa': True,
|
| 58 |
+
'use_wandb': True,
|
| 59 |
+
'wandb_project': 'vjepa_ablation'},
|
| 60 |
+
'metrics': {'sigreg': {}, 'std': {}},
|
| 61 |
+
'model': { 'class_token': True,
|
| 62 |
+
'freeze_backbone': False,
|
| 63 |
+
'is_causal': True,
|
| 64 |
+
'model_name': 'vit_large_patch14_dinov2_lvd142m',
|
| 65 |
+
'pe_type': '2d',
|
| 66 |
+
'pred_depth': 12,
|
| 67 |
+
'pred_embed_dim': 384,
|
| 68 |
+
'pred_num_heads': 6,
|
| 69 |
+
'predictor': 'v2_cross',
|
| 70 |
+
'stem_type': '2d',
|
| 71 |
+
'target_embed_dim': 1024,
|
| 72 |
+
'target_kind': 'ema',
|
| 73 |
+
'uniform_power': True,
|
| 74 |
+
'use_activation_checkpointing': True,
|
| 75 |
+
'use_alibi': False,
|
| 76 |
+
'use_lora': False,
|
| 77 |
+
'use_mask_tokens': True,
|
| 78 |
+
'use_rope': False,
|
| 79 |
+
'use_sdpa': True,
|
| 80 |
+
'zero_init_mask_tokens': True},
|
| 81 |
+
'nodes': 2,
|
| 82 |
+
'optimization': { 'clip_grad': None,
|
| 83 |
+
'ema': [1.0, 0.99925],
|
| 84 |
+
'ema_hold_epochs': 0,
|
| 85 |
+
'ema_ramp_epochs': 5,
|
| 86 |
+
'encoder_layer_decay': 0.9,
|
| 87 |
+
'epochs': 30,
|
| 88 |
+
'final_lr': 0.0001,
|
| 89 |
+
'final_weight_decay': 0.04,
|
| 90 |
+
'ipe': 300,
|
| 91 |
+
'ipe_scale': 1.0,
|
| 92 |
+
'lr': 0.0001,
|
| 93 |
+
'predictor_lr_scale': 3.0,
|
| 94 |
+
'start_lr': 1e-05,
|
| 95 |
+
'warmup': 2,
|
| 96 |
+
'weight_decay': 0.04},
|
| 97 |
+
'tasks_per_node': 3}
|
| 98 |
+
INFO:root:Running pre-training of app: vjepa
|
| 99 |
+
[INFO ][2026-06-11 00:43:25][app.vjepa.train ][main ] which_dtype='bfloat16'
|
| 100 |
+
[INFO ][2026-06-11 00:43:25][app.vjepa.train ][main ] Disabling persistent_workers (incompatible with KNN eval)
|
| 101 |
+
[INFO ][2026-06-11 00:43:25][app.vjepa.train ][main ] NCCL_SOCKET_IFNAME=None
|
| 102 |
+
[INFO ][2026-06-11 00:43:28][app.vjepa.train ][main ] Initialized (rank/world-size) 3/6, tasks_per_node=3
|
| 103 |
+
[WARNING ][2026-06-11 00:43:35][experiments.stmodels.vision_transformers_v3][_build_dinov2_vit ] DINOv2 is an image model — fwd_chunk is unset so all frames will be processed jointly. Set fwd_chunk=1 for per-frame (standard) processing.
|
| 104 |
+
[INFO ][2026-06-11 00:43:40][experiments.stmodels.vision_transformers_v3][_build_dinov2_vit ] Loading pretrained weights for vit_large_patch14_dinov2_lvd142m from timm (vit_large_patch14_dinov2)
|
| 105 |
+
[INFO ][2026-06-11 00:43:43][timm.models._builder][load_pretrained ] Loading pretrained weights from Hugging Face hub (timm/vit_large_patch14_dinov2.lvd142m)
|
| 106 |
+
[INFO ][2026-06-11 00:43:44][timm.models._hub ][load_state_dict_from_hf ] [timm/vit_large_patch14_dinov2.lvd142m] Safe alternative available for 'pytorch_model.bin' (as 'model.safetensors'). Loading weights using safetensors.
|
| 107 |
+
[INFO ][2026-06-11 00:43:44][timm.layers.pos_embed][resample_abs_pos_embed ] Resized position embedding: (37, 37) to (16, 16).
|
| 108 |
+
[INFO ][2026-06-11 00:43:46][root ][init_video_model ] ViTMultiSeqWrapper(
|
| 109 |
+
(backbone): VisionTransformer(
|
| 110 |
+
(patch_embed): PatchEmbed3D(
|
| 111 |
+
(proj): Conv3d(3, 1024, kernel_size=(1, 14, 14), stride=(1, 14, 14))
|
| 112 |
+
)
|
| 113 |
+
(blocks): ModuleList(
|
| 114 |
+
(0-23): 24 x Block(
|
| 115 |
+
(norm1): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 116 |
+
(norm2): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 117 |
+
(drop_path1): Identity()
|
| 118 |
+
(drop_path2): Identity()
|
| 119 |
+
(attn): Attention(
|
| 120 |
+
(qkv): Linear(in_features=1024, out_features=3072, bias=True)
|
| 121 |
+
(attn_drop): Dropout(p=0.0, inplace=False)
|
| 122 |
+
(proj): Linear(in_features=1024, out_features=1024, bias=True)
|
| 123 |
+
(proj_drop): Dropout(p=0.0, inplace=False)
|
| 124 |
+
)
|
| 125 |
+
(mlp): MLP(
|
| 126 |
+
(fc1): Linear(in_features=1024, out_features=4096, bias=True)
|
| 127 |
+
(act): GELU(approximate='none')
|
| 128 |
+
(fc2): Linear(in_features=4096, out_features=1024, bias=True)
|
| 129 |
+
(drop): Dropout(p=0.0, inplace=False)
|
| 130 |
+
)
|
| 131 |
+
)
|
| 132 |
+
)
|
| 133 |
+
(norm): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 134 |
+
)
|
| 135 |
+
)
|
| 136 |
+
[INFO ][2026-06-11 00:43:46][root ][init_video_model ] VJEPAPredictorMultiSeqWrapper(
|
| 137 |
+
(backbone): PredictorV2(
|
| 138 |
+
(predictor_embed): Linear(in_features=1024, out_features=384, bias=True)
|
| 139 |
+
(mask_tokens): ParameterList(
|
| 140 |
+
(0): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
|
| 141 |
+
(1): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
|
| 142 |
+
(2): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
|
| 143 |
+
(3): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
|
| 144 |
+
)
|
| 145 |
+
(predictor_blocks): ModuleList(
|
| 146 |
+
(0-11): 12 x Block(
|
| 147 |
+
(residual1): EfficientResidual(
|
| 148 |
+
(norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
|
| 149 |
+
(fn): Attention(
|
| 150 |
+
(q_proj): Linear(in_features=384, out_features=384, bias=False)
|
| 151 |
+
(k_proj): Linear(in_features=384, out_features=384, bias=False)
|
| 152 |
+
(v_proj): Linear(in_features=384, out_features=384, bias=False)
|
| 153 |
+
(proj): Linear(in_features=384, out_features=384, bias=False)
|
| 154 |
+
(rope): Rope()
|
| 155 |
+
)
|
| 156 |
+
)
|
| 157 |
+
(residual2): EfficientResidual(
|
| 158 |
+
(norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
|
| 159 |
+
(fn): MLP(
|
| 160 |
+
(fc1): Linear(in_features=384, out_features=1536, bias=False)
|
| 161 |
+
(act): GELU(approximate='none')
|
| 162 |
+
(fc2): Linear(in_features=1536, out_features=384, bias=False)
|
| 163 |
+
(drop): Dropout(p=0.0, inplace=False)
|
| 164 |
+
)
|
| 165 |
+
)
|
| 166 |
+
)
|
| 167 |
+
)
|
| 168 |
+
(predictor_norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
|
| 169 |
+
(predictor_proj): Linear(in_features=384, out_features=1024, bias=True)
|
| 170 |
+
)
|
| 171 |
+
)
|
| 172 |
+
[INFO ][2026-06-11 00:43:46][root ][init_video_model ] ViTMultiSeqWrapper(
|
| 173 |
+
(backbone): VisionTransformer(
|
| 174 |
+
(patch_embed): PatchEmbed3D(
|
| 175 |
+
(proj): Conv3d(3, 1024, kernel_size=(1, 14, 14), stride=(1, 14, 14))
|
| 176 |
+
)
|
| 177 |
+
(blocks): ModuleList(
|
| 178 |
+
(0-23): 24 x Block(
|
| 179 |
+
(norm1): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 180 |
+
(norm2): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 181 |
+
(drop_path1): Identity()
|
| 182 |
+
(drop_path2): Identity()
|
| 183 |
+
(attn): Attention(
|
| 184 |
+
(qkv): Linear(in_features=1024, out_features=3072, bias=True)
|
| 185 |
+
(attn_drop): Dropout(p=0.0, inplace=False)
|
| 186 |
+
(proj): Linear(in_features=1024, out_features=1024, bias=True)
|
| 187 |
+
(proj_drop): Dropout(p=0.0, inplace=False)
|
| 188 |
+
)
|
| 189 |
+
(mlp): MLP(
|
| 190 |
+
(fc1): Linear(in_features=1024, out_features=4096, bias=True)
|
| 191 |
+
(act): GELU(approximate='none')
|
| 192 |
+
(fc2): Linear(in_features=4096, out_features=1024, bias=True)
|
| 193 |
+
(drop): Dropout(p=0.0, inplace=False)
|
| 194 |
+
)
|
| 195 |
+
)
|
| 196 |
+
)
|
| 197 |
+
(norm): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 198 |
+
)
|
| 199 |
+
)
|
| 200 |
+
[INFO ][2026-06-11 00:43:46][root ][init_video_model ] Encoder number of parameters: 302964736
|
| 201 |
+
[INFO ][2026-06-11 00:43:46][root ][init_video_model ] Predictor number of parameters: 22042240
|
| 202 |
+
[INFO ][2026-06-11 00:43:46][root ][init_video_model ] Target encoder number of parameters: 302964736
|
| 203 |
+
[INFO ][2026-06-11 00:43:47][root ][make_videodataset ] VideoDataset dataset created
|
| 204 |
+
[INFO ][2026-06-11 00:43:47][WeightedSampler ][__init__ ] Using DistributedWeightedSampler with rank 3 / 6
|
| 205 |
+
[INFO ][2026-06-11 00:43:47][root ][make_videodataset ] VideoDataset unsupervised data loader created
|
| 206 |
+
[INFO ][2026-06-11 00:43:47][app.vjepa.train ][main ] iterations per epoch/dataset length: 300/2103
|
| 207 |
+
[INFO ][2026-06-11 00:43:47][app.vjepa.train ][main ] Wrapping models in DDP (rank 3)...
|
| 208 |
+
[INFO ][2026-06-11 00:45:06][app.vjepa.train ][main ] Initializing loader...
|
| 209 |
+
submitit INFO (2026-06-11 12:55:27,844) - Job completed successfully
|
| 210 |
+
[INFO ][2026-06-11 12:55:27][submitit ][process_job ] Job completed successfully
|
| 211 |
+
submitit INFO (2026-06-11 12:55:27,849) - Exiting after successful completion
|
| 212 |
+
[INFO ][2026-06-11 12:55:27][submitit ][process_job ] Exiting after successful completion
|
inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_25163/25163_4_log.err
ADDED
|
@@ -0,0 +1,50 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
/var/scratch/dcanezil/.cache/uv/virtualenvs/vd/lib/python3.12/site-packages/timm/models/layers/__init__.py:49: FutureWarning: Importing from timm.models.layers is deprecated, please import via timm.layers
|
| 2 |
+
warnings.warn(f"Importing from {__name__} is deprecated, please import via timm.layers", FutureWarning)
|
| 3 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/utils.py:888: FutureWarning: `torch.cuda.amp.GradScaler(args...)` is deprecated. Please use `torch.amp.GradScaler('cuda', args...)` instead.
|
| 4 |
+
scaler = torch.cuda.amp.GradScaler() if mixed_precision else None
|
| 5 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 6 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 7 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/dancing_macarena/uGQgxFHemyA_000059_000069.avi'
|
| 8 |
+
warnings.warn(f"video path not found {fname=}")
|
| 9 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/sailing/ZolXG_-RgIs_000970_000980.avi'
|
| 10 |
+
warnings.warn(f"video path not found {fname=}")
|
| 11 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/surfing_water/5l5Pdd96Pao_000161_000171.avi'
|
| 12 |
+
warnings.warn(f"video path not found {fname=}")
|
| 13 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_poker/6-DyF8umej8_000097_000107.avi'
|
| 14 |
+
warnings.warn(f"video path not found {fname=}")
|
| 15 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/picking_fruit/3VvkoFtPCCU_000045_000055.avi'
|
| 16 |
+
warnings.warn(f"video path not found {fname=}")
|
| 17 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/frying_vegetables/Nr-zCXgW2Sw_000161_000171.avi'
|
| 18 |
+
warnings.warn(f"video path not found {fname=}")
|
| 19 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/water_skiing/4bhvIHWADVU_000022_000032.avi'
|
| 20 |
+
warnings.warn(f"video path not found {fname=}")
|
| 21 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/water_skiing/4bhvIHWADVU_000022_000032.avi'
|
| 22 |
+
warnings.warn(f"video path not found {fname=}")
|
| 23 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_monopoly/NLL667uPWVA_000066_000076.avi'
|
| 24 |
+
warnings.warn(f"video path not found {fname=}")
|
| 25 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/pumping_fist/wKsuyr6Xk30_000094_000104.avi'
|
| 26 |
+
warnings.warn(f"video path not found {fname=}")
|
| 27 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/robot_dancing/J1x0XlWa6HM_000147_000157.avi'
|
| 28 |
+
warnings.warn(f"video path not found {fname=}")
|
| 29 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/dancing_macarena/uGQgxFHemyA_000059_000069.avi'
|
| 30 |
+
warnings.warn(f"video path not found {fname=}")
|
| 31 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/frying_vegetables/Nr-zCXgW2Sw_000161_000171.avi'
|
| 32 |
+
warnings.warn(f"video path not found {fname=}")
|
| 33 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/eating_carrots/eiZ8Hzc7FPU_000080_000090.avi'
|
| 34 |
+
warnings.warn(f"video path not found {fname=}")
|
| 35 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/picking_fruit/3VvkoFtPCCU_000045_000055.avi'
|
| 36 |
+
warnings.warn(f"video path not found {fname=}")
|
| 37 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_saxophone/Zs0_2GMEPXo_000054_000064.avi'
|
| 38 |
+
warnings.warn(f"video path not found {fname=}")
|
| 39 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_poker/-jGxlNQKkeo_000092_000102.avi'
|
| 40 |
+
warnings.warn(f"video path not found {fname=}")
|
| 41 |
+
[11:33:27] /github/workspace/src/video/video_reader.cc:83: ERROR opening: /local/dcanezil/data/kinetics/VideoData/long_jump/jsYiitkbr5w_000004_000014.avi, Invalid data found when processing input
|
| 42 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_monopoly/NLL667uPWVA_000066_000076.avi'
|
| 43 |
+
warnings.warn(f"video path not found {fname=}")
|
| 44 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/making_tea/mtYFNsRcxY4_000063_000073.avi'
|
| 45 |
+
warnings.warn(f"video path not found {fname=}")
|
| 46 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/bowling/f4eb0wOlspM_000053_000063.avi'
|
| 47 |
+
warnings.warn(f"video path not found {fname=}")
|
| 48 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_monopoly/NLL667uPWVA_000066_000076.avi'
|
| 49 |
+
warnings.warn(f"video path not found {fname=}")
|
| 50 |
+
[rank4]:[W611 12:55:28.447750472 ProcessGroupNCCL.cpp:1553] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())
|
inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_25163/25163_4_log.out
ADDED
|
@@ -0,0 +1,212 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
submitit INFO (2026-06-11 00:43:17,099) - Starting with JobEnvironment(job_id=25163, hostname=node411, local_rank=1(3), node=1(2), global_rank=4(6))
|
| 2 |
+
submitit INFO (2026-06-11 00:43:17,099) - Loading pickle: /var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_25163/25163_submitted.pkl
|
| 3 |
+
INFO:root:loaded pretrain params...
|
| 4 |
+
{ 'app': 'vjepa',
|
| 5 |
+
'cpus_per_task': 16,
|
| 6 |
+
'data': { 'batch_size': 32,
|
| 7 |
+
'crop_size': 224,
|
| 8 |
+
'dataset_fpcs': [8, 8],
|
| 9 |
+
'dataset_type': 'VideoDataset',
|
| 10 |
+
'datasets': [ '/local/dcanezil/data/kinetics/train.csv',
|
| 11 |
+
'/local/dcanezil/data/ssv2/train.csv'],
|
| 12 |
+
'datasets_weights': [0.65, 0.35],
|
| 13 |
+
'fps': 2,
|
| 14 |
+
'num_workers': 4,
|
| 15 |
+
'patch_size': 14,
|
| 16 |
+
'persistent_workers': True,
|
| 17 |
+
'pin_mem': True,
|
| 18 |
+
'tubelet_size': 1},
|
| 19 |
+
'data_aug': { 'auto_augment': False,
|
| 20 |
+
'motion_shift': False,
|
| 21 |
+
'random_resize_aspect_ratio': [0.75, 1.35],
|
| 22 |
+
'random_resize_scale': [0.3, 1.0],
|
| 23 |
+
'reprob': 0.0},
|
| 24 |
+
'folder': '/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow',
|
| 25 |
+
'loss': {'loss_exp': 1.0},
|
| 26 |
+
'mask': [ { 'aspect_ratio': [0.75, 1.5],
|
| 27 |
+
'full_complement': False,
|
| 28 |
+
'max_keep': None,
|
| 29 |
+
'max_temporal_keep': 1.0,
|
| 30 |
+
'num_blocks': 8,
|
| 31 |
+
'spatial_scale': [0.05, 0.05],
|
| 32 |
+
'temporal_scale': [1.0, 1.0]},
|
| 33 |
+
{ 'aspect_ratio': [0.75, 1.5],
|
| 34 |
+
'full_complement': False,
|
| 35 |
+
'max_keep': None,
|
| 36 |
+
'max_temporal_keep': 1.0,
|
| 37 |
+
'num_blocks': 2,
|
| 38 |
+
'spatial_scale': [0.4, 0.4],
|
| 39 |
+
'temporal_scale': [1.0, 1.0]}],
|
| 40 |
+
'mem_per_gpu': '40G',
|
| 41 |
+
'meta': { 'dtype': 'bfloat16',
|
| 42 |
+
'knn_eval_epoch0': True,
|
| 43 |
+
'knn_eval_freq': 5,
|
| 44 |
+
'knn_eval_presets': [ { 'config': { 'batch_size': 8,
|
| 45 |
+
'dataset_train': '/local/dcanezil/data/ssv2/train_coarse10.csv',
|
| 46 |
+
'dataset_val': '/local/dcanezil/data/ssv2/val_coarse10.csv',
|
| 47 |
+
'eval_videos_per_class': 100,
|
| 48 |
+
'linear_probe': True,
|
| 49 |
+
'num_workers': 2,
|
| 50 |
+
'pool_type': 'temporal_concat',
|
| 51 |
+
'train_videos_per_class': 500},
|
| 52 |
+
'preset': 'ssv2_coarse10'}],
|
| 53 |
+
'load_checkpoint': True,
|
| 54 |
+
'read_checkpoint': None,
|
| 55 |
+
'save_every_freq': 10,
|
| 56 |
+
'seed': 239,
|
| 57 |
+
'use_sdpa': True,
|
| 58 |
+
'use_wandb': True,
|
| 59 |
+
'wandb_project': 'vjepa_ablation'},
|
| 60 |
+
'metrics': {'sigreg': {}, 'std': {}},
|
| 61 |
+
'model': { 'class_token': True,
|
| 62 |
+
'freeze_backbone': False,
|
| 63 |
+
'is_causal': True,
|
| 64 |
+
'model_name': 'vit_large_patch14_dinov2_lvd142m',
|
| 65 |
+
'pe_type': '2d',
|
| 66 |
+
'pred_depth': 12,
|
| 67 |
+
'pred_embed_dim': 384,
|
| 68 |
+
'pred_num_heads': 6,
|
| 69 |
+
'predictor': 'v2_cross',
|
| 70 |
+
'stem_type': '2d',
|
| 71 |
+
'target_embed_dim': 1024,
|
| 72 |
+
'target_kind': 'ema',
|
| 73 |
+
'uniform_power': True,
|
| 74 |
+
'use_activation_checkpointing': True,
|
| 75 |
+
'use_alibi': False,
|
| 76 |
+
'use_lora': False,
|
| 77 |
+
'use_mask_tokens': True,
|
| 78 |
+
'use_rope': False,
|
| 79 |
+
'use_sdpa': True,
|
| 80 |
+
'zero_init_mask_tokens': True},
|
| 81 |
+
'nodes': 2,
|
| 82 |
+
'optimization': { 'clip_grad': None,
|
| 83 |
+
'ema': [1.0, 0.99925],
|
| 84 |
+
'ema_hold_epochs': 0,
|
| 85 |
+
'ema_ramp_epochs': 5,
|
| 86 |
+
'encoder_layer_decay': 0.9,
|
| 87 |
+
'epochs': 30,
|
| 88 |
+
'final_lr': 0.0001,
|
| 89 |
+
'final_weight_decay': 0.04,
|
| 90 |
+
'ipe': 300,
|
| 91 |
+
'ipe_scale': 1.0,
|
| 92 |
+
'lr': 0.0001,
|
| 93 |
+
'predictor_lr_scale': 3.0,
|
| 94 |
+
'start_lr': 1e-05,
|
| 95 |
+
'warmup': 2,
|
| 96 |
+
'weight_decay': 0.04},
|
| 97 |
+
'tasks_per_node': 3}
|
| 98 |
+
INFO:root:Running pre-training of app: vjepa
|
| 99 |
+
[INFO ][2026-06-11 00:43:25][app.vjepa.train ][main ] which_dtype='bfloat16'
|
| 100 |
+
[INFO ][2026-06-11 00:43:25][app.vjepa.train ][main ] Disabling persistent_workers (incompatible with KNN eval)
|
| 101 |
+
[INFO ][2026-06-11 00:43:25][app.vjepa.train ][main ] NCCL_SOCKET_IFNAME=None
|
| 102 |
+
[INFO ][2026-06-11 00:43:28][app.vjepa.train ][main ] Initialized (rank/world-size) 4/6, tasks_per_node=3
|
| 103 |
+
[WARNING ][2026-06-11 00:43:35][experiments.stmodels.vision_transformers_v3][_build_dinov2_vit ] DINOv2 is an image model — fwd_chunk is unset so all frames will be processed jointly. Set fwd_chunk=1 for per-frame (standard) processing.
|
| 104 |
+
[INFO ][2026-06-11 00:43:40][experiments.stmodels.vision_transformers_v3][_build_dinov2_vit ] Loading pretrained weights for vit_large_patch14_dinov2_lvd142m from timm (vit_large_patch14_dinov2)
|
| 105 |
+
[INFO ][2026-06-11 00:43:43][timm.models._builder][load_pretrained ] Loading pretrained weights from Hugging Face hub (timm/vit_large_patch14_dinov2.lvd142m)
|
| 106 |
+
[INFO ][2026-06-11 00:43:43][timm.models._hub ][load_state_dict_from_hf ] [timm/vit_large_patch14_dinov2.lvd142m] Safe alternative available for 'pytorch_model.bin' (as 'model.safetensors'). Loading weights using safetensors.
|
| 107 |
+
[INFO ][2026-06-11 00:43:44][timm.layers.pos_embed][resample_abs_pos_embed ] Resized position embedding: (37, 37) to (16, 16).
|
| 108 |
+
[INFO ][2026-06-11 00:43:46][root ][init_video_model ] ViTMultiSeqWrapper(
|
| 109 |
+
(backbone): VisionTransformer(
|
| 110 |
+
(patch_embed): PatchEmbed3D(
|
| 111 |
+
(proj): Conv3d(3, 1024, kernel_size=(1, 14, 14), stride=(1, 14, 14))
|
| 112 |
+
)
|
| 113 |
+
(blocks): ModuleList(
|
| 114 |
+
(0-23): 24 x Block(
|
| 115 |
+
(norm1): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 116 |
+
(norm2): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 117 |
+
(drop_path1): Identity()
|
| 118 |
+
(drop_path2): Identity()
|
| 119 |
+
(attn): Attention(
|
| 120 |
+
(qkv): Linear(in_features=1024, out_features=3072, bias=True)
|
| 121 |
+
(attn_drop): Dropout(p=0.0, inplace=False)
|
| 122 |
+
(proj): Linear(in_features=1024, out_features=1024, bias=True)
|
| 123 |
+
(proj_drop): Dropout(p=0.0, inplace=False)
|
| 124 |
+
)
|
| 125 |
+
(mlp): MLP(
|
| 126 |
+
(fc1): Linear(in_features=1024, out_features=4096, bias=True)
|
| 127 |
+
(act): GELU(approximate='none')
|
| 128 |
+
(fc2): Linear(in_features=4096, out_features=1024, bias=True)
|
| 129 |
+
(drop): Dropout(p=0.0, inplace=False)
|
| 130 |
+
)
|
| 131 |
+
)
|
| 132 |
+
)
|
| 133 |
+
(norm): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 134 |
+
)
|
| 135 |
+
)
|
| 136 |
+
[INFO ][2026-06-11 00:43:46][root ][init_video_model ] VJEPAPredictorMultiSeqWrapper(
|
| 137 |
+
(backbone): PredictorV2(
|
| 138 |
+
(predictor_embed): Linear(in_features=1024, out_features=384, bias=True)
|
| 139 |
+
(mask_tokens): ParameterList(
|
| 140 |
+
(0): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
|
| 141 |
+
(1): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
|
| 142 |
+
(2): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
|
| 143 |
+
(3): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
|
| 144 |
+
)
|
| 145 |
+
(predictor_blocks): ModuleList(
|
| 146 |
+
(0-11): 12 x Block(
|
| 147 |
+
(residual1): EfficientResidual(
|
| 148 |
+
(norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
|
| 149 |
+
(fn): Attention(
|
| 150 |
+
(q_proj): Linear(in_features=384, out_features=384, bias=False)
|
| 151 |
+
(k_proj): Linear(in_features=384, out_features=384, bias=False)
|
| 152 |
+
(v_proj): Linear(in_features=384, out_features=384, bias=False)
|
| 153 |
+
(proj): Linear(in_features=384, out_features=384, bias=False)
|
| 154 |
+
(rope): Rope()
|
| 155 |
+
)
|
| 156 |
+
)
|
| 157 |
+
(residual2): EfficientResidual(
|
| 158 |
+
(norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
|
| 159 |
+
(fn): MLP(
|
| 160 |
+
(fc1): Linear(in_features=384, out_features=1536, bias=False)
|
| 161 |
+
(act): GELU(approximate='none')
|
| 162 |
+
(fc2): Linear(in_features=1536, out_features=384, bias=False)
|
| 163 |
+
(drop): Dropout(p=0.0, inplace=False)
|
| 164 |
+
)
|
| 165 |
+
)
|
| 166 |
+
)
|
| 167 |
+
)
|
| 168 |
+
(predictor_norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
|
| 169 |
+
(predictor_proj): Linear(in_features=384, out_features=1024, bias=True)
|
| 170 |
+
)
|
| 171 |
+
)
|
| 172 |
+
[INFO ][2026-06-11 00:43:46][root ][init_video_model ] ViTMultiSeqWrapper(
|
| 173 |
+
(backbone): VisionTransformer(
|
| 174 |
+
(patch_embed): PatchEmbed3D(
|
| 175 |
+
(proj): Conv3d(3, 1024, kernel_size=(1, 14, 14), stride=(1, 14, 14))
|
| 176 |
+
)
|
| 177 |
+
(blocks): ModuleList(
|
| 178 |
+
(0-23): 24 x Block(
|
| 179 |
+
(norm1): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 180 |
+
(norm2): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 181 |
+
(drop_path1): Identity()
|
| 182 |
+
(drop_path2): Identity()
|
| 183 |
+
(attn): Attention(
|
| 184 |
+
(qkv): Linear(in_features=1024, out_features=3072, bias=True)
|
| 185 |
+
(attn_drop): Dropout(p=0.0, inplace=False)
|
| 186 |
+
(proj): Linear(in_features=1024, out_features=1024, bias=True)
|
| 187 |
+
(proj_drop): Dropout(p=0.0, inplace=False)
|
| 188 |
+
)
|
| 189 |
+
(mlp): MLP(
|
| 190 |
+
(fc1): Linear(in_features=1024, out_features=4096, bias=True)
|
| 191 |
+
(act): GELU(approximate='none')
|
| 192 |
+
(fc2): Linear(in_features=4096, out_features=1024, bias=True)
|
| 193 |
+
(drop): Dropout(p=0.0, inplace=False)
|
| 194 |
+
)
|
| 195 |
+
)
|
| 196 |
+
)
|
| 197 |
+
(norm): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 198 |
+
)
|
| 199 |
+
)
|
| 200 |
+
[INFO ][2026-06-11 00:43:46][root ][init_video_model ] Encoder number of parameters: 302964736
|
| 201 |
+
[INFO ][2026-06-11 00:43:46][root ][init_video_model ] Predictor number of parameters: 22042240
|
| 202 |
+
[INFO ][2026-06-11 00:43:46][root ][init_video_model ] Target encoder number of parameters: 302964736
|
| 203 |
+
[INFO ][2026-06-11 00:43:47][root ][make_videodataset ] VideoDataset dataset created
|
| 204 |
+
[INFO ][2026-06-11 00:43:47][WeightedSampler ][__init__ ] Using DistributedWeightedSampler with rank 4 / 6
|
| 205 |
+
[INFO ][2026-06-11 00:43:47][root ][make_videodataset ] VideoDataset unsupervised data loader created
|
| 206 |
+
[INFO ][2026-06-11 00:43:47][app.vjepa.train ][main ] iterations per epoch/dataset length: 300/2103
|
| 207 |
+
[INFO ][2026-06-11 00:43:47][app.vjepa.train ][main ] Wrapping models in DDP (rank 4)...
|
| 208 |
+
[INFO ][2026-06-11 00:45:06][app.vjepa.train ][main ] Initializing loader...
|
| 209 |
+
submitit INFO (2026-06-11 12:55:27,846) - Job completed successfully
|
| 210 |
+
[INFO ][2026-06-11 12:55:27][submitit ][process_job ] Job completed successfully
|
| 211 |
+
submitit INFO (2026-06-11 12:55:27,853) - Exiting after successful completion
|
| 212 |
+
[INFO ][2026-06-11 12:55:27][submitit ][process_job ] Exiting after successful completion
|
inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_25163/25163_5_log.err
ADDED
|
@@ -0,0 +1,45 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
/var/scratch/dcanezil/.cache/uv/virtualenvs/vd/lib/python3.12/site-packages/timm/models/layers/__init__.py:49: FutureWarning: Importing from timm.models.layers is deprecated, please import via timm.layers
|
| 2 |
+
warnings.warn(f"Importing from {__name__} is deprecated, please import via timm.layers", FutureWarning)
|
| 3 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/utils.py:888: FutureWarning: `torch.cuda.amp.GradScaler(args...)` is deprecated. Please use `torch.amp.GradScaler('cuda', args...)` instead.
|
| 4 |
+
scaler = torch.cuda.amp.GradScaler() if mixed_precision else None
|
| 5 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 6 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 7 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/dancing_macarena/uGQgxFHemyA_000059_000069.avi'
|
| 8 |
+
warnings.warn(f"video path not found {fname=}")
|
| 9 |
+
[02:13:55] /github/workspace/src/video/video_reader.cc:83: ERROR opening: /local/dcanezil/data/kinetics/VideoData/climbing_tree/QX5IzK4iFCM_000002_000012.avi, Invalid data found when processing input
|
| 10 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/baking_cookies/Q3XGZmqk1Q0_000296_000306.avi'
|
| 11 |
+
warnings.warn(f"video path not found {fname=}")
|
| 12 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/making_tea/mtYFNsRcxY4_000063_000073.avi'
|
| 13 |
+
warnings.warn(f"video path not found {fname=}")
|
| 14 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/robot_dancing/J1x0XlWa6HM_000147_000157.avi'
|
| 15 |
+
warnings.warn(f"video path not found {fname=}")
|
| 16 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/surfing_water/5l5Pdd96Pao_000161_000171.avi'
|
| 17 |
+
warnings.warn(f"video path not found {fname=}")
|
| 18 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/picking_fruit/3VvkoFtPCCU_000045_000055.avi'
|
| 19 |
+
warnings.warn(f"video path not found {fname=}")
|
| 20 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/eating_carrots/eiZ8Hzc7FPU_000080_000090.avi'
|
| 21 |
+
warnings.warn(f"video path not found {fname=}")
|
| 22 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_poker/-jGxlNQKkeo_000092_000102.avi'
|
| 23 |
+
warnings.warn(f"video path not found {fname=}")
|
| 24 |
+
[04:34:17] /github/workspace/src/video/video_reader.cc:83: ERROR opening: /local/dcanezil/data/kinetics/VideoData/climbing_tree/QX5IzK4iFCM_000002_000012.avi, Invalid data found when processing input
|
| 25 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/frying_vegetables/Nr-zCXgW2Sw_000161_000171.avi'
|
| 26 |
+
warnings.warn(f"video path not found {fname=}")
|
| 27 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/dancing_macarena/uGQgxFHemyA_000059_000069.avi'
|
| 28 |
+
warnings.warn(f"video path not found {fname=}")
|
| 29 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/sailing/ZolXG_-RgIs_000970_000980.avi'
|
| 30 |
+
warnings.warn(f"video path not found {fname=}")
|
| 31 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/baking_cookies/Q3XGZmqk1Q0_000296_000306.avi'
|
| 32 |
+
warnings.warn(f"video path not found {fname=}")
|
| 33 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/water_skiing/4bhvIHWADVU_000022_000032.avi'
|
| 34 |
+
warnings.warn(f"video path not found {fname=}")
|
| 35 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/water_skiing/4bhvIHWADVU_000022_000032.avi'
|
| 36 |
+
warnings.warn(f"video path not found {fname=}")
|
| 37 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_saxophone/Zs0_2GMEPXo_000054_000064.avi'
|
| 38 |
+
warnings.warn(f"video path not found {fname=}")
|
| 39 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_flute/co50KUHacYw_000005_000015.avi'
|
| 40 |
+
warnings.warn(f"video path not found {fname=}")
|
| 41 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/waxing_legs/cqUusXBODuw_001046_001056.avi'
|
| 42 |
+
warnings.warn(f"video path not found {fname=}")
|
| 43 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/eating_carrots/eiZ8Hzc7FPU_000080_000090.avi'
|
| 44 |
+
warnings.warn(f"video path not found {fname=}")
|
| 45 |
+
[rank5]:[W611 12:55:28.438561007 ProcessGroupNCCL.cpp:1553] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())
|
inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_25163/25163_5_log.out
ADDED
|
@@ -0,0 +1,212 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
submitit INFO (2026-06-11 00:43:17,099) - Starting with JobEnvironment(job_id=25163, hostname=node411, local_rank=2(3), node=1(2), global_rank=5(6))
|
| 2 |
+
submitit INFO (2026-06-11 00:43:17,100) - Loading pickle: /var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_25163/25163_submitted.pkl
|
| 3 |
+
INFO:root:loaded pretrain params...
|
| 4 |
+
{ 'app': 'vjepa',
|
| 5 |
+
'cpus_per_task': 16,
|
| 6 |
+
'data': { 'batch_size': 32,
|
| 7 |
+
'crop_size': 224,
|
| 8 |
+
'dataset_fpcs': [8, 8],
|
| 9 |
+
'dataset_type': 'VideoDataset',
|
| 10 |
+
'datasets': [ '/local/dcanezil/data/kinetics/train.csv',
|
| 11 |
+
'/local/dcanezil/data/ssv2/train.csv'],
|
| 12 |
+
'datasets_weights': [0.65, 0.35],
|
| 13 |
+
'fps': 2,
|
| 14 |
+
'num_workers': 4,
|
| 15 |
+
'patch_size': 14,
|
| 16 |
+
'persistent_workers': True,
|
| 17 |
+
'pin_mem': True,
|
| 18 |
+
'tubelet_size': 1},
|
| 19 |
+
'data_aug': { 'auto_augment': False,
|
| 20 |
+
'motion_shift': False,
|
| 21 |
+
'random_resize_aspect_ratio': [0.75, 1.35],
|
| 22 |
+
'random_resize_scale': [0.3, 1.0],
|
| 23 |
+
'reprob': 0.0},
|
| 24 |
+
'folder': '/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow',
|
| 25 |
+
'loss': {'loss_exp': 1.0},
|
| 26 |
+
'mask': [ { 'aspect_ratio': [0.75, 1.5],
|
| 27 |
+
'full_complement': False,
|
| 28 |
+
'max_keep': None,
|
| 29 |
+
'max_temporal_keep': 1.0,
|
| 30 |
+
'num_blocks': 8,
|
| 31 |
+
'spatial_scale': [0.05, 0.05],
|
| 32 |
+
'temporal_scale': [1.0, 1.0]},
|
| 33 |
+
{ 'aspect_ratio': [0.75, 1.5],
|
| 34 |
+
'full_complement': False,
|
| 35 |
+
'max_keep': None,
|
| 36 |
+
'max_temporal_keep': 1.0,
|
| 37 |
+
'num_blocks': 2,
|
| 38 |
+
'spatial_scale': [0.4, 0.4],
|
| 39 |
+
'temporal_scale': [1.0, 1.0]}],
|
| 40 |
+
'mem_per_gpu': '40G',
|
| 41 |
+
'meta': { 'dtype': 'bfloat16',
|
| 42 |
+
'knn_eval_epoch0': True,
|
| 43 |
+
'knn_eval_freq': 5,
|
| 44 |
+
'knn_eval_presets': [ { 'config': { 'batch_size': 8,
|
| 45 |
+
'dataset_train': '/local/dcanezil/data/ssv2/train_coarse10.csv',
|
| 46 |
+
'dataset_val': '/local/dcanezil/data/ssv2/val_coarse10.csv',
|
| 47 |
+
'eval_videos_per_class': 100,
|
| 48 |
+
'linear_probe': True,
|
| 49 |
+
'num_workers': 2,
|
| 50 |
+
'pool_type': 'temporal_concat',
|
| 51 |
+
'train_videos_per_class': 500},
|
| 52 |
+
'preset': 'ssv2_coarse10'}],
|
| 53 |
+
'load_checkpoint': True,
|
| 54 |
+
'read_checkpoint': None,
|
| 55 |
+
'save_every_freq': 10,
|
| 56 |
+
'seed': 239,
|
| 57 |
+
'use_sdpa': True,
|
| 58 |
+
'use_wandb': True,
|
| 59 |
+
'wandb_project': 'vjepa_ablation'},
|
| 60 |
+
'metrics': {'sigreg': {}, 'std': {}},
|
| 61 |
+
'model': { 'class_token': True,
|
| 62 |
+
'freeze_backbone': False,
|
| 63 |
+
'is_causal': True,
|
| 64 |
+
'model_name': 'vit_large_patch14_dinov2_lvd142m',
|
| 65 |
+
'pe_type': '2d',
|
| 66 |
+
'pred_depth': 12,
|
| 67 |
+
'pred_embed_dim': 384,
|
| 68 |
+
'pred_num_heads': 6,
|
| 69 |
+
'predictor': 'v2_cross',
|
| 70 |
+
'stem_type': '2d',
|
| 71 |
+
'target_embed_dim': 1024,
|
| 72 |
+
'target_kind': 'ema',
|
| 73 |
+
'uniform_power': True,
|
| 74 |
+
'use_activation_checkpointing': True,
|
| 75 |
+
'use_alibi': False,
|
| 76 |
+
'use_lora': False,
|
| 77 |
+
'use_mask_tokens': True,
|
| 78 |
+
'use_rope': False,
|
| 79 |
+
'use_sdpa': True,
|
| 80 |
+
'zero_init_mask_tokens': True},
|
| 81 |
+
'nodes': 2,
|
| 82 |
+
'optimization': { 'clip_grad': None,
|
| 83 |
+
'ema': [1.0, 0.99925],
|
| 84 |
+
'ema_hold_epochs': 0,
|
| 85 |
+
'ema_ramp_epochs': 5,
|
| 86 |
+
'encoder_layer_decay': 0.9,
|
| 87 |
+
'epochs': 30,
|
| 88 |
+
'final_lr': 0.0001,
|
| 89 |
+
'final_weight_decay': 0.04,
|
| 90 |
+
'ipe': 300,
|
| 91 |
+
'ipe_scale': 1.0,
|
| 92 |
+
'lr': 0.0001,
|
| 93 |
+
'predictor_lr_scale': 3.0,
|
| 94 |
+
'start_lr': 1e-05,
|
| 95 |
+
'warmup': 2,
|
| 96 |
+
'weight_decay': 0.04},
|
| 97 |
+
'tasks_per_node': 3}
|
| 98 |
+
INFO:root:Running pre-training of app: vjepa
|
| 99 |
+
[INFO ][2026-06-11 00:43:25][app.vjepa.train ][main ] which_dtype='bfloat16'
|
| 100 |
+
[INFO ][2026-06-11 00:43:25][app.vjepa.train ][main ] Disabling persistent_workers (incompatible with KNN eval)
|
| 101 |
+
[INFO ][2026-06-11 00:43:25][app.vjepa.train ][main ] NCCL_SOCKET_IFNAME=None
|
| 102 |
+
[INFO ][2026-06-11 00:43:28][app.vjepa.train ][main ] Initialized (rank/world-size) 5/6, tasks_per_node=3
|
| 103 |
+
[WARNING ][2026-06-11 00:43:35][experiments.stmodels.vision_transformers_v3][_build_dinov2_vit ] DINOv2 is an image model — fwd_chunk is unset so all frames will be processed jointly. Set fwd_chunk=1 for per-frame (standard) processing.
|
| 104 |
+
[INFO ][2026-06-11 00:43:40][experiments.stmodels.vision_transformers_v3][_build_dinov2_vit ] Loading pretrained weights for vit_large_patch14_dinov2_lvd142m from timm (vit_large_patch14_dinov2)
|
| 105 |
+
[INFO ][2026-06-11 00:43:44][timm.models._builder][load_pretrained ] Loading pretrained weights from Hugging Face hub (timm/vit_large_patch14_dinov2.lvd142m)
|
| 106 |
+
[INFO ][2026-06-11 00:43:44][timm.models._hub ][load_state_dict_from_hf ] [timm/vit_large_patch14_dinov2.lvd142m] Safe alternative available for 'pytorch_model.bin' (as 'model.safetensors'). Loading weights using safetensors.
|
| 107 |
+
[INFO ][2026-06-11 00:43:44][timm.layers.pos_embed][resample_abs_pos_embed ] Resized position embedding: (37, 37) to (16, 16).
|
| 108 |
+
[INFO ][2026-06-11 00:43:46][root ][init_video_model ] ViTMultiSeqWrapper(
|
| 109 |
+
(backbone): VisionTransformer(
|
| 110 |
+
(patch_embed): PatchEmbed3D(
|
| 111 |
+
(proj): Conv3d(3, 1024, kernel_size=(1, 14, 14), stride=(1, 14, 14))
|
| 112 |
+
)
|
| 113 |
+
(blocks): ModuleList(
|
| 114 |
+
(0-23): 24 x Block(
|
| 115 |
+
(norm1): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 116 |
+
(norm2): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 117 |
+
(drop_path1): Identity()
|
| 118 |
+
(drop_path2): Identity()
|
| 119 |
+
(attn): Attention(
|
| 120 |
+
(qkv): Linear(in_features=1024, out_features=3072, bias=True)
|
| 121 |
+
(attn_drop): Dropout(p=0.0, inplace=False)
|
| 122 |
+
(proj): Linear(in_features=1024, out_features=1024, bias=True)
|
| 123 |
+
(proj_drop): Dropout(p=0.0, inplace=False)
|
| 124 |
+
)
|
| 125 |
+
(mlp): MLP(
|
| 126 |
+
(fc1): Linear(in_features=1024, out_features=4096, bias=True)
|
| 127 |
+
(act): GELU(approximate='none')
|
| 128 |
+
(fc2): Linear(in_features=4096, out_features=1024, bias=True)
|
| 129 |
+
(drop): Dropout(p=0.0, inplace=False)
|
| 130 |
+
)
|
| 131 |
+
)
|
| 132 |
+
)
|
| 133 |
+
(norm): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 134 |
+
)
|
| 135 |
+
)
|
| 136 |
+
[INFO ][2026-06-11 00:43:46][root ][init_video_model ] VJEPAPredictorMultiSeqWrapper(
|
| 137 |
+
(backbone): PredictorV2(
|
| 138 |
+
(predictor_embed): Linear(in_features=1024, out_features=384, bias=True)
|
| 139 |
+
(mask_tokens): ParameterList(
|
| 140 |
+
(0): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
|
| 141 |
+
(1): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
|
| 142 |
+
(2): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
|
| 143 |
+
(3): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
|
| 144 |
+
)
|
| 145 |
+
(predictor_blocks): ModuleList(
|
| 146 |
+
(0-11): 12 x Block(
|
| 147 |
+
(residual1): EfficientResidual(
|
| 148 |
+
(norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
|
| 149 |
+
(fn): Attention(
|
| 150 |
+
(q_proj): Linear(in_features=384, out_features=384, bias=False)
|
| 151 |
+
(k_proj): Linear(in_features=384, out_features=384, bias=False)
|
| 152 |
+
(v_proj): Linear(in_features=384, out_features=384, bias=False)
|
| 153 |
+
(proj): Linear(in_features=384, out_features=384, bias=False)
|
| 154 |
+
(rope): Rope()
|
| 155 |
+
)
|
| 156 |
+
)
|
| 157 |
+
(residual2): EfficientResidual(
|
| 158 |
+
(norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
|
| 159 |
+
(fn): MLP(
|
| 160 |
+
(fc1): Linear(in_features=384, out_features=1536, bias=False)
|
| 161 |
+
(act): GELU(approximate='none')
|
| 162 |
+
(fc2): Linear(in_features=1536, out_features=384, bias=False)
|
| 163 |
+
(drop): Dropout(p=0.0, inplace=False)
|
| 164 |
+
)
|
| 165 |
+
)
|
| 166 |
+
)
|
| 167 |
+
)
|
| 168 |
+
(predictor_norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
|
| 169 |
+
(predictor_proj): Linear(in_features=384, out_features=1024, bias=True)
|
| 170 |
+
)
|
| 171 |
+
)
|
| 172 |
+
[INFO ][2026-06-11 00:43:46][root ][init_video_model ] ViTMultiSeqWrapper(
|
| 173 |
+
(backbone): VisionTransformer(
|
| 174 |
+
(patch_embed): PatchEmbed3D(
|
| 175 |
+
(proj): Conv3d(3, 1024, kernel_size=(1, 14, 14), stride=(1, 14, 14))
|
| 176 |
+
)
|
| 177 |
+
(blocks): ModuleList(
|
| 178 |
+
(0-23): 24 x Block(
|
| 179 |
+
(norm1): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 180 |
+
(norm2): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 181 |
+
(drop_path1): Identity()
|
| 182 |
+
(drop_path2): Identity()
|
| 183 |
+
(attn): Attention(
|
| 184 |
+
(qkv): Linear(in_features=1024, out_features=3072, bias=True)
|
| 185 |
+
(attn_drop): Dropout(p=0.0, inplace=False)
|
| 186 |
+
(proj): Linear(in_features=1024, out_features=1024, bias=True)
|
| 187 |
+
(proj_drop): Dropout(p=0.0, inplace=False)
|
| 188 |
+
)
|
| 189 |
+
(mlp): MLP(
|
| 190 |
+
(fc1): Linear(in_features=1024, out_features=4096, bias=True)
|
| 191 |
+
(act): GELU(approximate='none')
|
| 192 |
+
(fc2): Linear(in_features=4096, out_features=1024, bias=True)
|
| 193 |
+
(drop): Dropout(p=0.0, inplace=False)
|
| 194 |
+
)
|
| 195 |
+
)
|
| 196 |
+
)
|
| 197 |
+
(norm): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 198 |
+
)
|
| 199 |
+
)
|
| 200 |
+
[INFO ][2026-06-11 00:43:46][root ][init_video_model ] Encoder number of parameters: 302964736
|
| 201 |
+
[INFO ][2026-06-11 00:43:46][root ][init_video_model ] Predictor number of parameters: 22042240
|
| 202 |
+
[INFO ][2026-06-11 00:43:46][root ][init_video_model ] Target encoder number of parameters: 302964736
|
| 203 |
+
[INFO ][2026-06-11 00:43:47][root ][make_videodataset ] VideoDataset dataset created
|
| 204 |
+
[INFO ][2026-06-11 00:43:47][WeightedSampler ][__init__ ] Using DistributedWeightedSampler with rank 5 / 6
|
| 205 |
+
[INFO ][2026-06-11 00:43:47][root ][make_videodataset ] VideoDataset unsupervised data loader created
|
| 206 |
+
[INFO ][2026-06-11 00:43:47][app.vjepa.train ][main ] iterations per epoch/dataset length: 300/2103
|
| 207 |
+
[INFO ][2026-06-11 00:43:47][app.vjepa.train ][main ] Wrapping models in DDP (rank 5)...
|
| 208 |
+
[INFO ][2026-06-11 00:45:06][app.vjepa.train ][main ] Initializing loader...
|
| 209 |
+
submitit INFO (2026-06-11 12:55:27,847) - Job completed successfully
|
| 210 |
+
[INFO ][2026-06-11 12:55:27][submitit ][process_job ] Job completed successfully
|
| 211 |
+
submitit INFO (2026-06-11 12:55:27,853) - Exiting after successful completion
|
| 212 |
+
[INFO ][2026-06-11 12:55:27][submitit ][process_job ] Exiting after successful completion
|
inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_25214/25214_0_log.err
ADDED
|
@@ -0,0 +1,124 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
/var/scratch/dcanezil/.cache/uv/virtualenvs/vd/lib/python3.12/site-packages/timm/models/layers/__init__.py:49: FutureWarning: Importing from timm.models.layers is deprecated, please import via timm.layers
|
| 2 |
+
warnings.warn(f"Importing from {__name__} is deprecated, please import via timm.layers", FutureWarning)
|
| 3 |
+
wandb: Currently logged in as: dgcnz (uvjepa) to https://api.wandb.ai. Use `wandb login --relogin` to force relogin
|
| 4 |
+
wandb: setting up run vfclercb
|
| 5 |
+
wandb: Tracking run with wandb version 0.23.1
|
| 6 |
+
wandb: Run data is saved locally in /var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/wandb/run-20260611_195811-vfclercb
|
| 7 |
+
wandb: Run `wandb offline` to turn off syncing.
|
| 8 |
+
wandb: Resuming run d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow
|
| 9 |
+
wandb: ⭐️ View project at https://wandb.ai/uvjepa/vjepa_ablation
|
| 10 |
+
wandb: 🚀 View run at https://wandb.ai/uvjepa/vjepa_ablation/runs/vfclercb
|
| 11 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/utils.py:888: FutureWarning: `torch.cuda.amp.GradScaler(args...)` is deprecated. Please use `torch.amp.GradScaler('cuda', args...)` instead.
|
| 12 |
+
scaler = torch.cuda.amp.GradScaler() if mixed_precision else None
|
| 13 |
+
/var/scratch/dcanezil/.cache/uv/virtualenvs/vd/lib/python3.12/site-packages/torch/distributed/c10d_logger.py:83: UserWarning: barrier(): using the device under current context. You can specify `device_id` in `init_process_group` to mute this warning.
|
| 14 |
+
return func(*args, **kwargs)
|
| 15 |
+
wandb: WARNING Tried to log to step 0 that is less than the current step 9001. Steps must be monotonically increasing, so this data will be ignored. See https://wandb.me/define-metric to log data out of order.
|
| 16 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 17 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 18 |
+
wandb: WARNING Tried to log to step 9000 that is less than the current step 9001. Steps must be monotonically increasing, so this data will be ignored. See https://wandb.me/define-metric to log data out of order.
|
| 19 |
+
[20:06:18] /github/workspace/src/video/video_reader.cc:83: ERROR opening: /local/dcanezil/data/kinetics/VideoData/long_jump/jsYiitkbr5w_000004_000014.avi, Invalid data found when processing input
|
| 20 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/water_skiing/4bhvIHWADVU_000022_000032.avi'
|
| 21 |
+
warnings.warn(f"video path not found {fname=}")
|
| 22 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/water_skiing/4bhvIHWADVU_000022_000032.avi'
|
| 23 |
+
warnings.warn(f"video path not found {fname=}")
|
| 24 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/baking_cookies/Q3XGZmqk1Q0_000296_000306.avi'
|
| 25 |
+
warnings.warn(f"video path not found {fname=}")
|
| 26 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/robot_dancing/J1x0XlWa6HM_000147_000157.avi'
|
| 27 |
+
warnings.warn(f"video path not found {fname=}")
|
| 28 |
+
/var/scratch/dcanezil/.cache/uv/virtualenvs/vd/lib/python3.12/site-packages/torch/distributed/c10d_logger.py:83: UserWarning: barrier(): using the device under current context. You can specify `device_id` in `init_process_group` to mute this warning.
|
| 29 |
+
return func(*args, **kwargs)
|
| 30 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 31 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 32 |
+
[22:04:49] /github/workspace/src/video/video_reader.cc:83: ERROR opening: /local/dcanezil/data/kinetics/VideoData/trimming_trees/T6dqaZioaXs_000004_000014.avi, Invalid data found when processing input
|
| 33 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/waxing_legs/cqUusXBODuw_001046_001056.avi'
|
| 34 |
+
warnings.warn(f"video path not found {fname=}")
|
| 35 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/water_skiing/4bhvIHWADVU_000022_000032.avi'
|
| 36 |
+
warnings.warn(f"video path not found {fname=}")
|
| 37 |
+
/var/scratch/dcanezil/.cache/uv/virtualenvs/vd/lib/python3.12/site-packages/torch/distributed/c10d_logger.py:83: UserWarning: barrier(): using the device under current context. You can specify `device_id` in `init_process_group` to mute this warning.
|
| 38 |
+
return func(*args, **kwargs)
|
| 39 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 40 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 41 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/surfing_water/5l5Pdd96Pao_000161_000171.avi'
|
| 42 |
+
warnings.warn(f"video path not found {fname=}")
|
| 43 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/frying_vegetables/Nr-zCXgW2Sw_000161_000171.avi'
|
| 44 |
+
warnings.warn(f"video path not found {fname=}")
|
| 45 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/surfing_water/5l5Pdd96Pao_000161_000171.avi'
|
| 46 |
+
warnings.warn(f"video path not found {fname=}")
|
| 47 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/windsurfing/i-gzh_BPDa8_000154_000164.avi'
|
| 48 |
+
warnings.warn(f"video path not found {fname=}")
|
| 49 |
+
/var/scratch/dcanezil/.cache/uv/virtualenvs/vd/lib/python3.12/site-packages/torch/distributed/c10d_logger.py:83: UserWarning: barrier(): using the device under current context. You can specify `device_id` in `init_process_group` to mute this warning.
|
| 50 |
+
return func(*args, **kwargs)
|
| 51 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 52 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 53 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/baking_cookies/Q3XGZmqk1Q0_000296_000306.avi'
|
| 54 |
+
warnings.warn(f"video path not found {fname=}")
|
| 55 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/eating_carrots/eiZ8Hzc7FPU_000080_000090.avi'
|
| 56 |
+
warnings.warn(f"video path not found {fname=}")
|
| 57 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/dancing_macarena/uGQgxFHemyA_000059_000069.avi'
|
| 58 |
+
warnings.warn(f"video path not found {fname=}")
|
| 59 |
+
[03:57:12] /github/workspace/src/video/video_reader.cc:83: ERROR opening: /local/dcanezil/data/kinetics/VideoData/climbing_tree/QX5IzK4iFCM_000002_000012.avi, Invalid data found when processing input
|
| 60 |
+
/var/scratch/dcanezil/.cache/uv/virtualenvs/vd/lib/python3.12/site-packages/torch/distributed/c10d_logger.py:83: UserWarning: barrier(): using the device under current context. You can specify `device_id` in `init_process_group` to mute this warning.
|
| 61 |
+
return func(*args, **kwargs)
|
| 62 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 63 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 64 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/water_skiing/4bhvIHWADVU_000022_000032.avi'
|
| 65 |
+
warnings.warn(f"video path not found {fname=}")
|
| 66 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/water_skiing/4bhvIHWADVU_000022_000032.avi'
|
| 67 |
+
warnings.warn(f"video path not found {fname=}")
|
| 68 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/surfing_water/5l5Pdd96Pao_000161_000171.avi'
|
| 69 |
+
warnings.warn(f"video path not found {fname=}")
|
| 70 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/baking_cookies/Q3XGZmqk1Q0_000296_000306.avi'
|
| 71 |
+
warnings.warn(f"video path not found {fname=}")
|
| 72 |
+
[05:07:48] /github/workspace/src/video/video_reader.cc:83: ERROR opening: /local/dcanezil/data/kinetics/VideoData/long_jump/jsYiitkbr5w_000004_000014.avi, Invalid data found when processing input
|
| 73 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/pumping_fist/wKsuyr6Xk30_000094_000104.avi'
|
| 74 |
+
warnings.warn(f"video path not found {fname=}")
|
| 75 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/windsurfing/i-gzh_BPDa8_000154_000164.avi'
|
| 76 |
+
warnings.warn(f"video path not found {fname=}")
|
| 77 |
+
/var/scratch/dcanezil/.cache/uv/virtualenvs/vd/lib/python3.12/site-packages/torch/distributed/c10d_logger.py:83: UserWarning: barrier(): using the device under current context. You can specify `device_id` in `init_process_group` to mute this warning.
|
| 78 |
+
return func(*args, **kwargs)
|
| 79 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 80 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 81 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/surfing_water/5l5Pdd96Pao_000161_000171.avi'
|
| 82 |
+
warnings.warn(f"video path not found {fname=}")
|
| 83 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/sailing/ZolXG_-RgIs_000970_000980.avi'
|
| 84 |
+
warnings.warn(f"video path not found {fname=}")
|
| 85 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/robot_dancing/J1x0XlWa6HM_000147_000157.avi'
|
| 86 |
+
warnings.warn(f"video path not found {fname=}")
|
| 87 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/dancing_macarena/uGQgxFHemyA_000059_000069.avi'
|
| 88 |
+
warnings.warn(f"video path not found {fname=}")
|
| 89 |
+
/var/scratch/dcanezil/.cache/uv/virtualenvs/vd/lib/python3.12/site-packages/torch/distributed/c10d_logger.py:83: UserWarning: barrier(): using the device under current context. You can specify `device_id` in `init_process_group` to mute this warning.
|
| 90 |
+
return func(*args, **kwargs)
|
| 91 |
+
wandb: uploading output.log; uploading wandb-summary.json
|
| 92 |
+
wandb: uploading history steps 1860-1860, summary, console lines 1357-1366
|
| 93 |
+
wandb:
|
| 94 |
+
wandb: Run history:
|
| 95 |
+
wandb: epoch ▁▁▁▁▁▂▂▂▂▃▃▃▃▄▄▄▄▅▅▅▅▅▅▅▆▆▆▆▆▆▇▇▇███████
|
| 96 |
+
wandb: epoch/avg_data_time_ms ▃▃▃▅▇▁▇▇▇▇▃▃▅▃▂▄▄▅▃▅▅▅█▆▇▄▆██▇
|
| 97 |
+
wandb: epoch/avg_gpu_time_ms ▅▇▃▄▄▃▇▅▁▃▄▅▁▃▁▅▃▂▄▄▅▅▇█▆▇▃▅▄▁
|
| 98 |
+
wandb: epoch/avg_iter_time_ms ▅▇▃▄▄▃▇▅▁▄▄▅▂▃▁▅▃▂▄▄▅▅▇█▆▇▃▆▄▁
|
| 99 |
+
wandb: epoch/avg_loss ▁▁▂▃▃▃▃▄▄▄▅▅▅▅▅▆▆▆▆▆▇▇▇▇▇▇▇▇██
|
| 100 |
+
wandb: knn/ssv2_coarse10_temporal_concat_top1 ▁▂▅▅█▆
|
| 101 |
+
wandb: knn/ssv2_coarse10_temporal_concat_top5 ▁▄▅▃██
|
| 102 |
+
wandb: linear/ssv2_coarse10_temporal_concat_top1 ▁▃█▂▇▅
|
| 103 |
+
wandb: linear/ssv2_coarse10_temporal_concat_top5 ▂▂▁▃▅█
|
| 104 |
+
wandb: train/ema/cos_sim █▅▆█▄▅▄▅▄▅▅▆▂▃▅▄▄▅▃▄▁▃▃▃▅▃▃▄▃▃▂▂▃▂▂▃▁▃▄▁
|
| 105 |
+
wandb: +34 ...
|
| 106 |
+
wandb:
|
| 107 |
+
wandb: Run summary:
|
| 108 |
+
wandb: epoch 60
|
| 109 |
+
wandb: epoch/avg_data_time_ms 7.41441
|
| 110 |
+
wandb: epoch/avg_gpu_time_ms 4642.75277
|
| 111 |
+
wandb: epoch/avg_iter_time_ms 4656.73333
|
| 112 |
+
wandb: epoch/avg_loss 0.53577
|
| 113 |
+
wandb: knn/ssv2_coarse10_temporal_concat_top1 43.4739
|
| 114 |
+
wandb: knn/ssv2_coarse10_temporal_concat_top5 86.44578
|
| 115 |
+
wandb: linear/ssv2_coarse10_temporal_concat_top1 57.12851
|
| 116 |
+
wandb: linear/ssv2_coarse10_temporal_concat_top5 93.07229
|
| 117 |
+
wandb: train/ema/cos_sim 0.76809
|
| 118 |
+
wandb: +34 ...
|
| 119 |
+
wandb:
|
| 120 |
+
wandb: 🚀 View run d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow at: https://wandb.ai/uvjepa/vjepa_ablation/runs/vfclercb
|
| 121 |
+
wandb: ⭐️ View project at: https://wandb.ai/uvjepa/vjepa_ablation
|
| 122 |
+
wandb: Synced 5 W&B file(s), 0 media file(s), 0 artifact file(s) and 0 other file(s)
|
| 123 |
+
wandb: Find logs at: /var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/wandb/run-20260611_195811-vfclercb/logs
|
| 124 |
+
[rank0]:[W612 08:09:43.865534111 ProcessGroupNCCL.cpp:1553] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())
|
inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_25214/25214_0_log.out
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_25214/25214_1_log.err
ADDED
|
@@ -0,0 +1,41 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
/var/scratch/dcanezil/.cache/uv/virtualenvs/vd/lib/python3.12/site-packages/timm/models/layers/__init__.py:49: FutureWarning: Importing from timm.models.layers is deprecated, please import via timm.layers
|
| 2 |
+
warnings.warn(f"Importing from {__name__} is deprecated, please import via timm.layers", FutureWarning)
|
| 3 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/utils.py:888: FutureWarning: `torch.cuda.amp.GradScaler(args...)` is deprecated. Please use `torch.amp.GradScaler('cuda', args...)` instead.
|
| 4 |
+
scaler = torch.cuda.amp.GradScaler() if mixed_precision else None
|
| 5 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 6 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 7 |
+
[20:15:37] /github/workspace/src/video/video_reader.cc:83: ERROR opening: /local/dcanezil/data/kinetics/VideoData/trimming_trees/T6dqaZioaXs_000004_000014.avi, Invalid data found when processing input
|
| 8 |
+
[20:50:50] /github/workspace/src/video/video_reader.cc:83: ERROR opening: /local/dcanezil/data/kinetics/VideoData/trimming_trees/T6dqaZioaXs_000004_000014.avi, Invalid data found when processing input
|
| 9 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_flute/co50KUHacYw_000005_000015.avi'
|
| 10 |
+
warnings.warn(f"video path not found {fname=}")
|
| 11 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/waxing_legs/cqUusXBODuw_001046_001056.avi'
|
| 12 |
+
warnings.warn(f"video path not found {fname=}")
|
| 13 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/windsurfing/i-gzh_BPDa8_000154_000164.avi'
|
| 14 |
+
warnings.warn(f"video path not found {fname=}")
|
| 15 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 16 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 17 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_poker/6-DyF8umej8_000097_000107.avi'
|
| 18 |
+
warnings.warn(f"video path not found {fname=}")
|
| 19 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 20 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 21 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_poker/6-DyF8umej8_000097_000107.avi'
|
| 22 |
+
warnings.warn(f"video path not found {fname=}")
|
| 23 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/windsurfing/i-gzh_BPDa8_000154_000164.avi'
|
| 24 |
+
warnings.warn(f"video path not found {fname=}")
|
| 25 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 26 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 27 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/sailing/ZolXG_-RgIs_000970_000980.avi'
|
| 28 |
+
warnings.warn(f"video path not found {fname=}")
|
| 29 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 30 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 31 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/pumping_fist/wKsuyr6Xk30_000094_000104.avi'
|
| 32 |
+
warnings.warn(f"video path not found {fname=}")
|
| 33 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/water_skiing/4bhvIHWADVU_000022_000032.avi'
|
| 34 |
+
warnings.warn(f"video path not found {fname=}")
|
| 35 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 36 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 37 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_flute/co50KUHacYw_000005_000015.avi'
|
| 38 |
+
warnings.warn(f"video path not found {fname=}")
|
| 39 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/bowling/f4eb0wOlspM_000053_000063.avi'
|
| 40 |
+
warnings.warn(f"video path not found {fname=}")
|
| 41 |
+
[rank1]:[W612 08:09:41.224472889 ProcessGroupNCCL.cpp:1553] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())
|
inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_25214/25214_1_log.out
ADDED
|
@@ -0,0 +1,218 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
submitit INFO (2026-06-11 19:57:57,500) - Starting with JobEnvironment(job_id=25214, hostname=node404, local_rank=1(3), node=0(2), global_rank=1(6))
|
| 2 |
+
submitit INFO (2026-06-11 19:57:57,500) - Loading pickle: /var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_25214/25214_submitted.pkl
|
| 3 |
+
INFO:root:loaded pretrain params...
|
| 4 |
+
{ 'app': 'vjepa',
|
| 5 |
+
'cpus_per_task': 16,
|
| 6 |
+
'data': { 'batch_size': 32,
|
| 7 |
+
'crop_size': 224,
|
| 8 |
+
'dataset_fpcs': [8, 8],
|
| 9 |
+
'dataset_type': 'VideoDataset',
|
| 10 |
+
'datasets': [ '/local/dcanezil/data/kinetics/train.csv',
|
| 11 |
+
'/local/dcanezil/data/ssv2/train.csv'],
|
| 12 |
+
'datasets_weights': [0.65, 0.35],
|
| 13 |
+
'fps': 2,
|
| 14 |
+
'num_workers': 4,
|
| 15 |
+
'patch_size': 14,
|
| 16 |
+
'persistent_workers': True,
|
| 17 |
+
'pin_mem': True,
|
| 18 |
+
'tubelet_size': 1},
|
| 19 |
+
'data_aug': { 'auto_augment': False,
|
| 20 |
+
'motion_shift': False,
|
| 21 |
+
'random_resize_aspect_ratio': [0.75, 1.35],
|
| 22 |
+
'random_resize_scale': [0.3, 1.0],
|
| 23 |
+
'reprob': 0.0},
|
| 24 |
+
'folder': '/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow',
|
| 25 |
+
'loss': {'loss_exp': 1.0},
|
| 26 |
+
'mask': [ { 'aspect_ratio': [0.75, 1.5],
|
| 27 |
+
'full_complement': False,
|
| 28 |
+
'max_keep': None,
|
| 29 |
+
'max_temporal_keep': 1.0,
|
| 30 |
+
'num_blocks': 8,
|
| 31 |
+
'spatial_scale': [0.05, 0.05],
|
| 32 |
+
'temporal_scale': [1.0, 1.0]},
|
| 33 |
+
{ 'aspect_ratio': [0.75, 1.5],
|
| 34 |
+
'full_complement': False,
|
| 35 |
+
'max_keep': None,
|
| 36 |
+
'max_temporal_keep': 1.0,
|
| 37 |
+
'num_blocks': 2,
|
| 38 |
+
'spatial_scale': [0.4, 0.4],
|
| 39 |
+
'temporal_scale': [1.0, 1.0]}],
|
| 40 |
+
'mem_per_gpu': '40G',
|
| 41 |
+
'meta': { 'dtype': 'bfloat16',
|
| 42 |
+
'knn_eval_epoch0': True,
|
| 43 |
+
'knn_eval_freq': 5,
|
| 44 |
+
'knn_eval_presets': [ { 'config': { 'batch_size': 8,
|
| 45 |
+
'dataset_train': '/local/dcanezil/data/ssv2/train_coarse10.csv',
|
| 46 |
+
'dataset_val': '/local/dcanezil/data/ssv2/val_coarse10.csv',
|
| 47 |
+
'eval_videos_per_class': 100,
|
| 48 |
+
'linear_probe': True,
|
| 49 |
+
'num_workers': 2,
|
| 50 |
+
'pool_type': 'temporal_concat',
|
| 51 |
+
'train_videos_per_class': 500},
|
| 52 |
+
'preset': 'ssv2_coarse10'}],
|
| 53 |
+
'load_checkpoint': True,
|
| 54 |
+
'read_checkpoint': None,
|
| 55 |
+
'save_every_freq': 10,
|
| 56 |
+
'seed': 239,
|
| 57 |
+
'use_sdpa': True,
|
| 58 |
+
'use_wandb': True,
|
| 59 |
+
'wandb_project': 'vjepa_ablation'},
|
| 60 |
+
'metrics': {'sigreg': {}, 'std': {}},
|
| 61 |
+
'model': { 'class_token': True,
|
| 62 |
+
'freeze_backbone': False,
|
| 63 |
+
'is_causal': True,
|
| 64 |
+
'model_name': 'vit_large_patch14_dinov2_lvd142m',
|
| 65 |
+
'pe_type': '2d',
|
| 66 |
+
'pred_depth': 12,
|
| 67 |
+
'pred_embed_dim': 384,
|
| 68 |
+
'pred_num_heads': 6,
|
| 69 |
+
'predictor': 'v2_cross',
|
| 70 |
+
'stem_type': '2d',
|
| 71 |
+
'target_embed_dim': 1024,
|
| 72 |
+
'target_kind': 'ema',
|
| 73 |
+
'uniform_power': True,
|
| 74 |
+
'use_activation_checkpointing': True,
|
| 75 |
+
'use_alibi': False,
|
| 76 |
+
'use_lora': False,
|
| 77 |
+
'use_mask_tokens': True,
|
| 78 |
+
'use_rope': False,
|
| 79 |
+
'use_sdpa': True,
|
| 80 |
+
'zero_init_mask_tokens': True},
|
| 81 |
+
'nodes': 2,
|
| 82 |
+
'optimization': { 'clip_grad': None,
|
| 83 |
+
'ema': [1.0, 0.99925],
|
| 84 |
+
'ema_hold_epochs': 0,
|
| 85 |
+
'ema_ramp_epochs': 5,
|
| 86 |
+
'encoder_layer_decay': 0.9,
|
| 87 |
+
'epochs': 60,
|
| 88 |
+
'final_lr': 0.0001,
|
| 89 |
+
'final_weight_decay': 0.04,
|
| 90 |
+
'ipe': 300,
|
| 91 |
+
'ipe_scale': 1.0,
|
| 92 |
+
'lr': 0.0001,
|
| 93 |
+
'predictor_lr_scale': 3.0,
|
| 94 |
+
'start_lr': 1e-05,
|
| 95 |
+
'warmup': 2,
|
| 96 |
+
'weight_decay': 0.04},
|
| 97 |
+
'tasks_per_node': 3}
|
| 98 |
+
INFO:root:Running pre-training of app: vjepa
|
| 99 |
+
[INFO ][2026-06-11 19:58:06][app.vjepa.train ][main ] which_dtype='bfloat16'
|
| 100 |
+
[INFO ][2026-06-11 19:58:06][app.vjepa.train ][main ] Disabling persistent_workers (incompatible with KNN eval)
|
| 101 |
+
[INFO ][2026-06-11 19:58:06][app.vjepa.train ][main ] NCCL_SOCKET_IFNAME=None
|
| 102 |
+
[INFO ][2026-06-11 19:58:10][app.vjepa.train ][main ] Initialized (rank/world-size) 1/6, tasks_per_node=3
|
| 103 |
+
[WARNING ][2026-06-11 19:58:23][experiments.stmodels.vision_transformers_v3][_build_dinov2_vit ] DINOv2 is an image model — fwd_chunk is unset so all frames will be processed jointly. Set fwd_chunk=1 for per-frame (standard) processing.
|
| 104 |
+
[INFO ][2026-06-11 19:58:28][experiments.stmodels.vision_transformers_v3][_build_dinov2_vit ] Loading pretrained weights for vit_large_patch14_dinov2_lvd142m from timm (vit_large_patch14_dinov2)
|
| 105 |
+
[INFO ][2026-06-11 19:58:32][timm.models._builder][load_pretrained ] Loading pretrained weights from Hugging Face hub (timm/vit_large_patch14_dinov2.lvd142m)
|
| 106 |
+
[INFO ][2026-06-11 19:58:32][timm.models._hub ][load_state_dict_from_hf ] [timm/vit_large_patch14_dinov2.lvd142m] Safe alternative available for 'pytorch_model.bin' (as 'model.safetensors'). Loading weights using safetensors.
|
| 107 |
+
[INFO ][2026-06-11 19:58:32][timm.layers.pos_embed][resample_abs_pos_embed ] Resized position embedding: (37, 37) to (16, 16).
|
| 108 |
+
[INFO ][2026-06-11 19:58:36][root ][init_video_model ] ViTMultiSeqWrapper(
|
| 109 |
+
(backbone): VisionTransformer(
|
| 110 |
+
(patch_embed): PatchEmbed3D(
|
| 111 |
+
(proj): Conv3d(3, 1024, kernel_size=(1, 14, 14), stride=(1, 14, 14))
|
| 112 |
+
)
|
| 113 |
+
(blocks): ModuleList(
|
| 114 |
+
(0-23): 24 x Block(
|
| 115 |
+
(norm1): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 116 |
+
(norm2): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 117 |
+
(drop_path1): Identity()
|
| 118 |
+
(drop_path2): Identity()
|
| 119 |
+
(attn): Attention(
|
| 120 |
+
(qkv): Linear(in_features=1024, out_features=3072, bias=True)
|
| 121 |
+
(attn_drop): Dropout(p=0.0, inplace=False)
|
| 122 |
+
(proj): Linear(in_features=1024, out_features=1024, bias=True)
|
| 123 |
+
(proj_drop): Dropout(p=0.0, inplace=False)
|
| 124 |
+
)
|
| 125 |
+
(mlp): MLP(
|
| 126 |
+
(fc1): Linear(in_features=1024, out_features=4096, bias=True)
|
| 127 |
+
(act): GELU(approximate='none')
|
| 128 |
+
(fc2): Linear(in_features=4096, out_features=1024, bias=True)
|
| 129 |
+
(drop): Dropout(p=0.0, inplace=False)
|
| 130 |
+
)
|
| 131 |
+
)
|
| 132 |
+
)
|
| 133 |
+
(norm): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 134 |
+
)
|
| 135 |
+
)
|
| 136 |
+
[INFO ][2026-06-11 19:58:36][root ][init_video_model ] VJEPAPredictorMultiSeqWrapper(
|
| 137 |
+
(backbone): PredictorV2(
|
| 138 |
+
(predictor_embed): Linear(in_features=1024, out_features=384, bias=True)
|
| 139 |
+
(mask_tokens): ParameterList(
|
| 140 |
+
(0): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
|
| 141 |
+
(1): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
|
| 142 |
+
(2): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
|
| 143 |
+
(3): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
|
| 144 |
+
)
|
| 145 |
+
(predictor_blocks): ModuleList(
|
| 146 |
+
(0-11): 12 x Block(
|
| 147 |
+
(residual1): EfficientResidual(
|
| 148 |
+
(norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
|
| 149 |
+
(fn): Attention(
|
| 150 |
+
(q_proj): Linear(in_features=384, out_features=384, bias=False)
|
| 151 |
+
(k_proj): Linear(in_features=384, out_features=384, bias=False)
|
| 152 |
+
(v_proj): Linear(in_features=384, out_features=384, bias=False)
|
| 153 |
+
(proj): Linear(in_features=384, out_features=384, bias=False)
|
| 154 |
+
(rope): Rope()
|
| 155 |
+
)
|
| 156 |
+
)
|
| 157 |
+
(residual2): EfficientResidual(
|
| 158 |
+
(norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
|
| 159 |
+
(fn): MLP(
|
| 160 |
+
(fc1): Linear(in_features=384, out_features=1536, bias=False)
|
| 161 |
+
(act): GELU(approximate='none')
|
| 162 |
+
(fc2): Linear(in_features=1536, out_features=384, bias=False)
|
| 163 |
+
(drop): Dropout(p=0.0, inplace=False)
|
| 164 |
+
)
|
| 165 |
+
)
|
| 166 |
+
)
|
| 167 |
+
)
|
| 168 |
+
(predictor_norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
|
| 169 |
+
(predictor_proj): Linear(in_features=384, out_features=1024, bias=True)
|
| 170 |
+
)
|
| 171 |
+
)
|
| 172 |
+
[INFO ][2026-06-11 19:58:36][root ][init_video_model ] ViTMultiSeqWrapper(
|
| 173 |
+
(backbone): VisionTransformer(
|
| 174 |
+
(patch_embed): PatchEmbed3D(
|
| 175 |
+
(proj): Conv3d(3, 1024, kernel_size=(1, 14, 14), stride=(1, 14, 14))
|
| 176 |
+
)
|
| 177 |
+
(blocks): ModuleList(
|
| 178 |
+
(0-23): 24 x Block(
|
| 179 |
+
(norm1): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 180 |
+
(norm2): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 181 |
+
(drop_path1): Identity()
|
| 182 |
+
(drop_path2): Identity()
|
| 183 |
+
(attn): Attention(
|
| 184 |
+
(qkv): Linear(in_features=1024, out_features=3072, bias=True)
|
| 185 |
+
(attn_drop): Dropout(p=0.0, inplace=False)
|
| 186 |
+
(proj): Linear(in_features=1024, out_features=1024, bias=True)
|
| 187 |
+
(proj_drop): Dropout(p=0.0, inplace=False)
|
| 188 |
+
)
|
| 189 |
+
(mlp): MLP(
|
| 190 |
+
(fc1): Linear(in_features=1024, out_features=4096, bias=True)
|
| 191 |
+
(act): GELU(approximate='none')
|
| 192 |
+
(fc2): Linear(in_features=4096, out_features=1024, bias=True)
|
| 193 |
+
(drop): Dropout(p=0.0, inplace=False)
|
| 194 |
+
)
|
| 195 |
+
)
|
| 196 |
+
)
|
| 197 |
+
(norm): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 198 |
+
)
|
| 199 |
+
)
|
| 200 |
+
[INFO ][2026-06-11 19:58:36][root ][init_video_model ] Encoder number of parameters: 302964736
|
| 201 |
+
[INFO ][2026-06-11 19:58:36][root ][init_video_model ] Predictor number of parameters: 22042240
|
| 202 |
+
[INFO ][2026-06-11 19:58:36][root ][init_video_model ] Target encoder number of parameters: 302964736
|
| 203 |
+
[INFO ][2026-06-11 19:58:37][root ][make_videodataset ] VideoDataset dataset created
|
| 204 |
+
[INFO ][2026-06-11 19:58:37][WeightedSampler ][__init__ ] Using DistributedWeightedSampler with rank 1 / 6
|
| 205 |
+
[INFO ][2026-06-11 19:58:37][root ][make_videodataset ] VideoDataset unsupervised data loader created
|
| 206 |
+
[INFO ][2026-06-11 19:58:37][app.vjepa.train ][main ] iterations per epoch/dataset length: 300/2103
|
| 207 |
+
[INFO ][2026-06-11 19:58:37][app.vjepa.train ][main ] Wrapping models in DDP (rank 1)...
|
| 208 |
+
[INFO ][2026-06-11 19:58:38][root ][load_checkpoint ] Loading checkpoint from /var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/latest.pt
|
| 209 |
+
[INFO ][2026-06-11 19:58:48][root ][load_checkpoint ] loaded pretrained encoder from epoch 30 with msg: <All keys matched successfully>
|
| 210 |
+
[INFO ][2026-06-11 19:58:48][root ][load_checkpoint ] loaded pretrained predictor from epoch 30 with msg: <All keys matched successfully>
|
| 211 |
+
[INFO ][2026-06-11 19:58:48][root ][load_checkpoint ] loaded pretrained target encoder from epoch 30 with msg: <All keys matched successfully>
|
| 212 |
+
[INFO ][2026-06-11 19:58:48][root ][load_checkpoint ] loaded optimizers from epoch 30
|
| 213 |
+
[INFO ][2026-06-11 19:58:48][root ][load_checkpoint ] read-path: /var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/latest.pt
|
| 214 |
+
[INFO ][2026-06-11 20:00:08][app.vjepa.train ][main ] Initializing loader...
|
| 215 |
+
submitit INFO (2026-06-12 08:09:41,007) - Job completed successfully
|
| 216 |
+
[INFO ][2026-06-12 08:09:41][submitit ][process_job ] Job completed successfully
|
| 217 |
+
submitit INFO (2026-06-12 08:09:41,009) - Exiting after successful completion
|
| 218 |
+
[INFO ][2026-06-12 08:09:41][submitit ][process_job ] Exiting after successful completion
|
inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_25214/25214_2_log.err
ADDED
|
@@ -0,0 +1,61 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
/var/scratch/dcanezil/.cache/uv/virtualenvs/vd/lib/python3.12/site-packages/timm/models/layers/__init__.py:49: FutureWarning: Importing from timm.models.layers is deprecated, please import via timm.layers
|
| 2 |
+
warnings.warn(f"Importing from {__name__} is deprecated, please import via timm.layers", FutureWarning)
|
| 3 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/utils.py:888: FutureWarning: `torch.cuda.amp.GradScaler(args...)` is deprecated. Please use `torch.amp.GradScaler('cuda', args...)` instead.
|
| 4 |
+
scaler = torch.cuda.amp.GradScaler() if mixed_precision else None
|
| 5 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 6 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 7 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/sailing/ZolXG_-RgIs_000970_000980.avi'
|
| 8 |
+
warnings.warn(f"video path not found {fname=}")
|
| 9 |
+
[21:23:05] /github/workspace/src/video/video_reader.cc:83: ERROR opening: /local/dcanezil/data/kinetics/VideoData/trimming_trees/T6dqaZioaXs_000004_000014.avi, Invalid data found when processing input
|
| 10 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/making_tea/mtYFNsRcxY4_000063_000073.avi'
|
| 11 |
+
warnings.warn(f"video path not found {fname=}")
|
| 12 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 13 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 14 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/dancing_macarena/uGQgxFHemyA_000059_000069.avi'
|
| 15 |
+
warnings.warn(f"video path not found {fname=}")
|
| 16 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 17 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 18 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/pumping_fist/wKsuyr6Xk30_000094_000104.avi'
|
| 19 |
+
warnings.warn(f"video path not found {fname=}")
|
| 20 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_poker/-jGxlNQKkeo_000092_000102.avi'
|
| 21 |
+
warnings.warn(f"video path not found {fname=}")
|
| 22 |
+
[01:08:05] /github/workspace/src/video/video_reader.cc:83: ERROR opening: /local/dcanezil/data/kinetics/VideoData/climbing_tree/QX5IzK4iFCM_000002_000012.avi, Invalid data found when processing input
|
| 23 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/robot_dancing/J1x0XlWa6HM_000147_000157.avi'
|
| 24 |
+
warnings.warn(f"video path not found {fname=}")
|
| 25 |
+
[01:19:45] /github/workspace/src/video/video_reader.cc:83: ERROR opening: /local/dcanezil/data/kinetics/VideoData/long_jump/jsYiitkbr5w_000004_000014.avi, Invalid data found when processing input
|
| 26 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/frying_vegetables/Nr-zCXgW2Sw_000161_000171.avi'
|
| 27 |
+
warnings.warn(f"video path not found {fname=}")
|
| 28 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 29 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 30 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_saxophone/Zs0_2GMEPXo_000054_000064.avi'
|
| 31 |
+
warnings.warn(f"video path not found {fname=}")
|
| 32 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_flute/co50KUHacYw_000005_000015.avi'
|
| 33 |
+
warnings.warn(f"video path not found {fname=}")
|
| 34 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/baking_cookies/Q3XGZmqk1Q0_000296_000306.avi'
|
| 35 |
+
warnings.warn(f"video path not found {fname=}")
|
| 36 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 37 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 38 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/surfing_water/5l5Pdd96Pao_000161_000171.avi'
|
| 39 |
+
warnings.warn(f"video path not found {fname=}")
|
| 40 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_saxophone/Zs0_2GMEPXo_000054_000064.avi'
|
| 41 |
+
warnings.warn(f"video path not found {fname=}")
|
| 42 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 43 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 44 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_poker/-jGxlNQKkeo_000092_000102.avi'
|
| 45 |
+
warnings.warn(f"video path not found {fname=}")
|
| 46 |
+
[06:15:56] /github/workspace/src/video/video_reader.cc:83: ERROR opening: /local/dcanezil/data/kinetics/VideoData/long_jump/jsYiitkbr5w_000004_000014.avi, Invalid data found when processing input
|
| 47 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_poker/6-DyF8umej8_000097_000107.avi'
|
| 48 |
+
warnings.warn(f"video path not found {fname=}")
|
| 49 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/sailing/ZolXG_-RgIs_000970_000980.avi'
|
| 50 |
+
warnings.warn(f"video path not found {fname=}")
|
| 51 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/dancing_macarena/uGQgxFHemyA_000059_000069.avi'
|
| 52 |
+
warnings.warn(f"video path not found {fname=}")
|
| 53 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_poker/6-DyF8umej8_000097_000107.avi'
|
| 54 |
+
warnings.warn(f"video path not found {fname=}")
|
| 55 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_poker/-jGxlNQKkeo_000092_000102.avi'
|
| 56 |
+
warnings.warn(f"video path not found {fname=}")
|
| 57 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/baking_cookies/Q3XGZmqk1Q0_000296_000306.avi'
|
| 58 |
+
warnings.warn(f"video path not found {fname=}")
|
| 59 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/robot_dancing/J1x0XlWa6HM_000147_000157.avi'
|
| 60 |
+
warnings.warn(f"video path not found {fname=}")
|
| 61 |
+
[rank2]:[W612 08:09:41.221863207 ProcessGroupNCCL.cpp:1553] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())
|
inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_25214/25214_2_log.out
ADDED
|
@@ -0,0 +1,218 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
submitit INFO (2026-06-11 19:57:57,500) - Starting with JobEnvironment(job_id=25214, hostname=node404, local_rank=2(3), node=0(2), global_rank=2(6))
|
| 2 |
+
submitit INFO (2026-06-11 19:57:57,500) - Loading pickle: /var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_25214/25214_submitted.pkl
|
| 3 |
+
INFO:root:loaded pretrain params...
|
| 4 |
+
{ 'app': 'vjepa',
|
| 5 |
+
'cpus_per_task': 16,
|
| 6 |
+
'data': { 'batch_size': 32,
|
| 7 |
+
'crop_size': 224,
|
| 8 |
+
'dataset_fpcs': [8, 8],
|
| 9 |
+
'dataset_type': 'VideoDataset',
|
| 10 |
+
'datasets': [ '/local/dcanezil/data/kinetics/train.csv',
|
| 11 |
+
'/local/dcanezil/data/ssv2/train.csv'],
|
| 12 |
+
'datasets_weights': [0.65, 0.35],
|
| 13 |
+
'fps': 2,
|
| 14 |
+
'num_workers': 4,
|
| 15 |
+
'patch_size': 14,
|
| 16 |
+
'persistent_workers': True,
|
| 17 |
+
'pin_mem': True,
|
| 18 |
+
'tubelet_size': 1},
|
| 19 |
+
'data_aug': { 'auto_augment': False,
|
| 20 |
+
'motion_shift': False,
|
| 21 |
+
'random_resize_aspect_ratio': [0.75, 1.35],
|
| 22 |
+
'random_resize_scale': [0.3, 1.0],
|
| 23 |
+
'reprob': 0.0},
|
| 24 |
+
'folder': '/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow',
|
| 25 |
+
'loss': {'loss_exp': 1.0},
|
| 26 |
+
'mask': [ { 'aspect_ratio': [0.75, 1.5],
|
| 27 |
+
'full_complement': False,
|
| 28 |
+
'max_keep': None,
|
| 29 |
+
'max_temporal_keep': 1.0,
|
| 30 |
+
'num_blocks': 8,
|
| 31 |
+
'spatial_scale': [0.05, 0.05],
|
| 32 |
+
'temporal_scale': [1.0, 1.0]},
|
| 33 |
+
{ 'aspect_ratio': [0.75, 1.5],
|
| 34 |
+
'full_complement': False,
|
| 35 |
+
'max_keep': None,
|
| 36 |
+
'max_temporal_keep': 1.0,
|
| 37 |
+
'num_blocks': 2,
|
| 38 |
+
'spatial_scale': [0.4, 0.4],
|
| 39 |
+
'temporal_scale': [1.0, 1.0]}],
|
| 40 |
+
'mem_per_gpu': '40G',
|
| 41 |
+
'meta': { 'dtype': 'bfloat16',
|
| 42 |
+
'knn_eval_epoch0': True,
|
| 43 |
+
'knn_eval_freq': 5,
|
| 44 |
+
'knn_eval_presets': [ { 'config': { 'batch_size': 8,
|
| 45 |
+
'dataset_train': '/local/dcanezil/data/ssv2/train_coarse10.csv',
|
| 46 |
+
'dataset_val': '/local/dcanezil/data/ssv2/val_coarse10.csv',
|
| 47 |
+
'eval_videos_per_class': 100,
|
| 48 |
+
'linear_probe': True,
|
| 49 |
+
'num_workers': 2,
|
| 50 |
+
'pool_type': 'temporal_concat',
|
| 51 |
+
'train_videos_per_class': 500},
|
| 52 |
+
'preset': 'ssv2_coarse10'}],
|
| 53 |
+
'load_checkpoint': True,
|
| 54 |
+
'read_checkpoint': None,
|
| 55 |
+
'save_every_freq': 10,
|
| 56 |
+
'seed': 239,
|
| 57 |
+
'use_sdpa': True,
|
| 58 |
+
'use_wandb': True,
|
| 59 |
+
'wandb_project': 'vjepa_ablation'},
|
| 60 |
+
'metrics': {'sigreg': {}, 'std': {}},
|
| 61 |
+
'model': { 'class_token': True,
|
| 62 |
+
'freeze_backbone': False,
|
| 63 |
+
'is_causal': True,
|
| 64 |
+
'model_name': 'vit_large_patch14_dinov2_lvd142m',
|
| 65 |
+
'pe_type': '2d',
|
| 66 |
+
'pred_depth': 12,
|
| 67 |
+
'pred_embed_dim': 384,
|
| 68 |
+
'pred_num_heads': 6,
|
| 69 |
+
'predictor': 'v2_cross',
|
| 70 |
+
'stem_type': '2d',
|
| 71 |
+
'target_embed_dim': 1024,
|
| 72 |
+
'target_kind': 'ema',
|
| 73 |
+
'uniform_power': True,
|
| 74 |
+
'use_activation_checkpointing': True,
|
| 75 |
+
'use_alibi': False,
|
| 76 |
+
'use_lora': False,
|
| 77 |
+
'use_mask_tokens': True,
|
| 78 |
+
'use_rope': False,
|
| 79 |
+
'use_sdpa': True,
|
| 80 |
+
'zero_init_mask_tokens': True},
|
| 81 |
+
'nodes': 2,
|
| 82 |
+
'optimization': { 'clip_grad': None,
|
| 83 |
+
'ema': [1.0, 0.99925],
|
| 84 |
+
'ema_hold_epochs': 0,
|
| 85 |
+
'ema_ramp_epochs': 5,
|
| 86 |
+
'encoder_layer_decay': 0.9,
|
| 87 |
+
'epochs': 60,
|
| 88 |
+
'final_lr': 0.0001,
|
| 89 |
+
'final_weight_decay': 0.04,
|
| 90 |
+
'ipe': 300,
|
| 91 |
+
'ipe_scale': 1.0,
|
| 92 |
+
'lr': 0.0001,
|
| 93 |
+
'predictor_lr_scale': 3.0,
|
| 94 |
+
'start_lr': 1e-05,
|
| 95 |
+
'warmup': 2,
|
| 96 |
+
'weight_decay': 0.04},
|
| 97 |
+
'tasks_per_node': 3}
|
| 98 |
+
INFO:root:Running pre-training of app: vjepa
|
| 99 |
+
[INFO ][2026-06-11 19:58:06][app.vjepa.train ][main ] which_dtype='bfloat16'
|
| 100 |
+
[INFO ][2026-06-11 19:58:06][app.vjepa.train ][main ] Disabling persistent_workers (incompatible with KNN eval)
|
| 101 |
+
[INFO ][2026-06-11 19:58:06][app.vjepa.train ][main ] NCCL_SOCKET_IFNAME=None
|
| 102 |
+
[INFO ][2026-06-11 19:58:10][app.vjepa.train ][main ] Initialized (rank/world-size) 2/6, tasks_per_node=3
|
| 103 |
+
[WARNING ][2026-06-11 19:58:23][experiments.stmodels.vision_transformers_v3][_build_dinov2_vit ] DINOv2 is an image model — fwd_chunk is unset so all frames will be processed jointly. Set fwd_chunk=1 for per-frame (standard) processing.
|
| 104 |
+
[INFO ][2026-06-11 19:58:28][experiments.stmodels.vision_transformers_v3][_build_dinov2_vit ] Loading pretrained weights for vit_large_patch14_dinov2_lvd142m from timm (vit_large_patch14_dinov2)
|
| 105 |
+
[INFO ][2026-06-11 19:58:33][timm.models._builder][load_pretrained ] Loading pretrained weights from Hugging Face hub (timm/vit_large_patch14_dinov2.lvd142m)
|
| 106 |
+
[INFO ][2026-06-11 19:58:33][timm.models._hub ][load_state_dict_from_hf ] [timm/vit_large_patch14_dinov2.lvd142m] Safe alternative available for 'pytorch_model.bin' (as 'model.safetensors'). Loading weights using safetensors.
|
| 107 |
+
[INFO ][2026-06-11 19:58:33][timm.layers.pos_embed][resample_abs_pos_embed ] Resized position embedding: (37, 37) to (16, 16).
|
| 108 |
+
[INFO ][2026-06-11 19:58:36][root ][init_video_model ] ViTMultiSeqWrapper(
|
| 109 |
+
(backbone): VisionTransformer(
|
| 110 |
+
(patch_embed): PatchEmbed3D(
|
| 111 |
+
(proj): Conv3d(3, 1024, kernel_size=(1, 14, 14), stride=(1, 14, 14))
|
| 112 |
+
)
|
| 113 |
+
(blocks): ModuleList(
|
| 114 |
+
(0-23): 24 x Block(
|
| 115 |
+
(norm1): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 116 |
+
(norm2): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 117 |
+
(drop_path1): Identity()
|
| 118 |
+
(drop_path2): Identity()
|
| 119 |
+
(attn): Attention(
|
| 120 |
+
(qkv): Linear(in_features=1024, out_features=3072, bias=True)
|
| 121 |
+
(attn_drop): Dropout(p=0.0, inplace=False)
|
| 122 |
+
(proj): Linear(in_features=1024, out_features=1024, bias=True)
|
| 123 |
+
(proj_drop): Dropout(p=0.0, inplace=False)
|
| 124 |
+
)
|
| 125 |
+
(mlp): MLP(
|
| 126 |
+
(fc1): Linear(in_features=1024, out_features=4096, bias=True)
|
| 127 |
+
(act): GELU(approximate='none')
|
| 128 |
+
(fc2): Linear(in_features=4096, out_features=1024, bias=True)
|
| 129 |
+
(drop): Dropout(p=0.0, inplace=False)
|
| 130 |
+
)
|
| 131 |
+
)
|
| 132 |
+
)
|
| 133 |
+
(norm): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 134 |
+
)
|
| 135 |
+
)
|
| 136 |
+
[INFO ][2026-06-11 19:58:36][root ][init_video_model ] VJEPAPredictorMultiSeqWrapper(
|
| 137 |
+
(backbone): PredictorV2(
|
| 138 |
+
(predictor_embed): Linear(in_features=1024, out_features=384, bias=True)
|
| 139 |
+
(mask_tokens): ParameterList(
|
| 140 |
+
(0): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
|
| 141 |
+
(1): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
|
| 142 |
+
(2): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
|
| 143 |
+
(3): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
|
| 144 |
+
)
|
| 145 |
+
(predictor_blocks): ModuleList(
|
| 146 |
+
(0-11): 12 x Block(
|
| 147 |
+
(residual1): EfficientResidual(
|
| 148 |
+
(norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
|
| 149 |
+
(fn): Attention(
|
| 150 |
+
(q_proj): Linear(in_features=384, out_features=384, bias=False)
|
| 151 |
+
(k_proj): Linear(in_features=384, out_features=384, bias=False)
|
| 152 |
+
(v_proj): Linear(in_features=384, out_features=384, bias=False)
|
| 153 |
+
(proj): Linear(in_features=384, out_features=384, bias=False)
|
| 154 |
+
(rope): Rope()
|
| 155 |
+
)
|
| 156 |
+
)
|
| 157 |
+
(residual2): EfficientResidual(
|
| 158 |
+
(norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
|
| 159 |
+
(fn): MLP(
|
| 160 |
+
(fc1): Linear(in_features=384, out_features=1536, bias=False)
|
| 161 |
+
(act): GELU(approximate='none')
|
| 162 |
+
(fc2): Linear(in_features=1536, out_features=384, bias=False)
|
| 163 |
+
(drop): Dropout(p=0.0, inplace=False)
|
| 164 |
+
)
|
| 165 |
+
)
|
| 166 |
+
)
|
| 167 |
+
)
|
| 168 |
+
(predictor_norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
|
| 169 |
+
(predictor_proj): Linear(in_features=384, out_features=1024, bias=True)
|
| 170 |
+
)
|
| 171 |
+
)
|
| 172 |
+
[INFO ][2026-06-11 19:58:36][root ][init_video_model ] ViTMultiSeqWrapper(
|
| 173 |
+
(backbone): VisionTransformer(
|
| 174 |
+
(patch_embed): PatchEmbed3D(
|
| 175 |
+
(proj): Conv3d(3, 1024, kernel_size=(1, 14, 14), stride=(1, 14, 14))
|
| 176 |
+
)
|
| 177 |
+
(blocks): ModuleList(
|
| 178 |
+
(0-23): 24 x Block(
|
| 179 |
+
(norm1): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 180 |
+
(norm2): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 181 |
+
(drop_path1): Identity()
|
| 182 |
+
(drop_path2): Identity()
|
| 183 |
+
(attn): Attention(
|
| 184 |
+
(qkv): Linear(in_features=1024, out_features=3072, bias=True)
|
| 185 |
+
(attn_drop): Dropout(p=0.0, inplace=False)
|
| 186 |
+
(proj): Linear(in_features=1024, out_features=1024, bias=True)
|
| 187 |
+
(proj_drop): Dropout(p=0.0, inplace=False)
|
| 188 |
+
)
|
| 189 |
+
(mlp): MLP(
|
| 190 |
+
(fc1): Linear(in_features=1024, out_features=4096, bias=True)
|
| 191 |
+
(act): GELU(approximate='none')
|
| 192 |
+
(fc2): Linear(in_features=4096, out_features=1024, bias=True)
|
| 193 |
+
(drop): Dropout(p=0.0, inplace=False)
|
| 194 |
+
)
|
| 195 |
+
)
|
| 196 |
+
)
|
| 197 |
+
(norm): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 198 |
+
)
|
| 199 |
+
)
|
| 200 |
+
[INFO ][2026-06-11 19:58:36][root ][init_video_model ] Encoder number of parameters: 302964736
|
| 201 |
+
[INFO ][2026-06-11 19:58:36][root ][init_video_model ] Predictor number of parameters: 22042240
|
| 202 |
+
[INFO ][2026-06-11 19:58:36][root ][init_video_model ] Target encoder number of parameters: 302964736
|
| 203 |
+
[INFO ][2026-06-11 19:58:37][root ][make_videodataset ] VideoDataset dataset created
|
| 204 |
+
[INFO ][2026-06-11 19:58:37][WeightedSampler ][__init__ ] Using DistributedWeightedSampler with rank 2 / 6
|
| 205 |
+
[INFO ][2026-06-11 19:58:37][root ][make_videodataset ] VideoDataset unsupervised data loader created
|
| 206 |
+
[INFO ][2026-06-11 19:58:37][app.vjepa.train ][main ] iterations per epoch/dataset length: 300/2103
|
| 207 |
+
[INFO ][2026-06-11 19:58:37][app.vjepa.train ][main ] Wrapping models in DDP (rank 2)...
|
| 208 |
+
[INFO ][2026-06-11 19:58:38][root ][load_checkpoint ] Loading checkpoint from /var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/latest.pt
|
| 209 |
+
[INFO ][2026-06-11 19:58:48][root ][load_checkpoint ] loaded pretrained encoder from epoch 30 with msg: <All keys matched successfully>
|
| 210 |
+
[INFO ][2026-06-11 19:58:48][root ][load_checkpoint ] loaded pretrained predictor from epoch 30 with msg: <All keys matched successfully>
|
| 211 |
+
[INFO ][2026-06-11 19:58:48][root ][load_checkpoint ] loaded pretrained target encoder from epoch 30 with msg: <All keys matched successfully>
|
| 212 |
+
[INFO ][2026-06-11 19:58:48][root ][load_checkpoint ] loaded optimizers from epoch 30
|
| 213 |
+
[INFO ][2026-06-11 19:58:48][root ][load_checkpoint ] read-path: /var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/latest.pt
|
| 214 |
+
[INFO ][2026-06-11 20:00:08][app.vjepa.train ][main ] Initializing loader...
|
| 215 |
+
submitit INFO (2026-06-12 08:09:41,012) - Job completed successfully
|
| 216 |
+
[INFO ][2026-06-12 08:09:41][submitit ][process_job ] Job completed successfully
|
| 217 |
+
submitit INFO (2026-06-12 08:09:41,014) - Exiting after successful completion
|
| 218 |
+
[INFO ][2026-06-12 08:09:41][submitit ][process_job ] Exiting after successful completion
|
inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_25214/25214_3_log.err
ADDED
|
@@ -0,0 +1,19 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
/var/scratch/dcanezil/.cache/uv/virtualenvs/vd/lib/python3.12/site-packages/timm/models/layers/__init__.py:49: FutureWarning: Importing from timm.models.layers is deprecated, please import via timm.layers
|
| 2 |
+
warnings.warn(f"Importing from {__name__} is deprecated, please import via timm.layers", FutureWarning)
|
| 3 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/utils.py:888: FutureWarning: `torch.cuda.amp.GradScaler(args...)` is deprecated. Please use `torch.amp.GradScaler('cuda', args...)` instead.
|
| 4 |
+
scaler = torch.cuda.amp.GradScaler() if mixed_precision else None
|
| 5 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 6 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 7 |
+
[21:06:09] /github/workspace/src/video/video_reader.cc:83: ERROR opening: /local/dcanezil/data/kinetics/VideoData/long_jump/jsYiitkbr5w_000004_000014.avi, Invalid data found when processing input
|
| 8 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/water_skiing/4bhvIHWADVU_000022_000032.avi'
|
| 9 |
+
warnings.warn(f"video path not found {fname=}")
|
| 10 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/picking_fruit/3VvkoFtPCCU_000045_000055.avi'
|
| 11 |
+
warnings.warn(f"video path not found {fname=}")
|
| 12 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_saxophone/Zs0_2GMEPXo_000054_000064.avi'
|
| 13 |
+
warnings.warn(f"video path not found {fname=}")
|
| 14 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_saxophone/Zs0_2GMEPXo_000054_000064.avi'
|
| 15 |
+
warnings.warn(f"video path not found {fname=}")
|
| 16 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/eating_carrots/eiZ8Hzc7FPU_000080_000090.avi'
|
| 17 |
+
warnings.warn(f"video path not found {fname=}")
|
| 18 |
+
[03:45:16] /github/workspace/src/video/video_reader.cc:83: ERROR opening: /local/dcanezil/data/kinetics/VideoData/climbing_tree/QX5IzK4iFCM_000002_000012.avi, Invalid data found when processing input
|
| 19 |
+
[rank3]:[W612 08:09:42.670453105 ProcessGroupNCCL.cpp:1553] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())
|
inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_25214/25214_3_log.out
ADDED
|
@@ -0,0 +1,218 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
submitit INFO (2026-06-11 19:57:57,500) - Starting with JobEnvironment(job_id=25214, hostname=node406, local_rank=0(3), node=1(2), global_rank=3(6))
|
| 2 |
+
submitit INFO (2026-06-11 19:57:57,500) - Loading pickle: /var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_25214/25214_submitted.pkl
|
| 3 |
+
INFO:root:loaded pretrain params...
|
| 4 |
+
{ 'app': 'vjepa',
|
| 5 |
+
'cpus_per_task': 16,
|
| 6 |
+
'data': { 'batch_size': 32,
|
| 7 |
+
'crop_size': 224,
|
| 8 |
+
'dataset_fpcs': [8, 8],
|
| 9 |
+
'dataset_type': 'VideoDataset',
|
| 10 |
+
'datasets': [ '/local/dcanezil/data/kinetics/train.csv',
|
| 11 |
+
'/local/dcanezil/data/ssv2/train.csv'],
|
| 12 |
+
'datasets_weights': [0.65, 0.35],
|
| 13 |
+
'fps': 2,
|
| 14 |
+
'num_workers': 4,
|
| 15 |
+
'patch_size': 14,
|
| 16 |
+
'persistent_workers': True,
|
| 17 |
+
'pin_mem': True,
|
| 18 |
+
'tubelet_size': 1},
|
| 19 |
+
'data_aug': { 'auto_augment': False,
|
| 20 |
+
'motion_shift': False,
|
| 21 |
+
'random_resize_aspect_ratio': [0.75, 1.35],
|
| 22 |
+
'random_resize_scale': [0.3, 1.0],
|
| 23 |
+
'reprob': 0.0},
|
| 24 |
+
'folder': '/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow',
|
| 25 |
+
'loss': {'loss_exp': 1.0},
|
| 26 |
+
'mask': [ { 'aspect_ratio': [0.75, 1.5],
|
| 27 |
+
'full_complement': False,
|
| 28 |
+
'max_keep': None,
|
| 29 |
+
'max_temporal_keep': 1.0,
|
| 30 |
+
'num_blocks': 8,
|
| 31 |
+
'spatial_scale': [0.05, 0.05],
|
| 32 |
+
'temporal_scale': [1.0, 1.0]},
|
| 33 |
+
{ 'aspect_ratio': [0.75, 1.5],
|
| 34 |
+
'full_complement': False,
|
| 35 |
+
'max_keep': None,
|
| 36 |
+
'max_temporal_keep': 1.0,
|
| 37 |
+
'num_blocks': 2,
|
| 38 |
+
'spatial_scale': [0.4, 0.4],
|
| 39 |
+
'temporal_scale': [1.0, 1.0]}],
|
| 40 |
+
'mem_per_gpu': '40G',
|
| 41 |
+
'meta': { 'dtype': 'bfloat16',
|
| 42 |
+
'knn_eval_epoch0': True,
|
| 43 |
+
'knn_eval_freq': 5,
|
| 44 |
+
'knn_eval_presets': [ { 'config': { 'batch_size': 8,
|
| 45 |
+
'dataset_train': '/local/dcanezil/data/ssv2/train_coarse10.csv',
|
| 46 |
+
'dataset_val': '/local/dcanezil/data/ssv2/val_coarse10.csv',
|
| 47 |
+
'eval_videos_per_class': 100,
|
| 48 |
+
'linear_probe': True,
|
| 49 |
+
'num_workers': 2,
|
| 50 |
+
'pool_type': 'temporal_concat',
|
| 51 |
+
'train_videos_per_class': 500},
|
| 52 |
+
'preset': 'ssv2_coarse10'}],
|
| 53 |
+
'load_checkpoint': True,
|
| 54 |
+
'read_checkpoint': None,
|
| 55 |
+
'save_every_freq': 10,
|
| 56 |
+
'seed': 239,
|
| 57 |
+
'use_sdpa': True,
|
| 58 |
+
'use_wandb': True,
|
| 59 |
+
'wandb_project': 'vjepa_ablation'},
|
| 60 |
+
'metrics': {'sigreg': {}, 'std': {}},
|
| 61 |
+
'model': { 'class_token': True,
|
| 62 |
+
'freeze_backbone': False,
|
| 63 |
+
'is_causal': True,
|
| 64 |
+
'model_name': 'vit_large_patch14_dinov2_lvd142m',
|
| 65 |
+
'pe_type': '2d',
|
| 66 |
+
'pred_depth': 12,
|
| 67 |
+
'pred_embed_dim': 384,
|
| 68 |
+
'pred_num_heads': 6,
|
| 69 |
+
'predictor': 'v2_cross',
|
| 70 |
+
'stem_type': '2d',
|
| 71 |
+
'target_embed_dim': 1024,
|
| 72 |
+
'target_kind': 'ema',
|
| 73 |
+
'uniform_power': True,
|
| 74 |
+
'use_activation_checkpointing': True,
|
| 75 |
+
'use_alibi': False,
|
| 76 |
+
'use_lora': False,
|
| 77 |
+
'use_mask_tokens': True,
|
| 78 |
+
'use_rope': False,
|
| 79 |
+
'use_sdpa': True,
|
| 80 |
+
'zero_init_mask_tokens': True},
|
| 81 |
+
'nodes': 2,
|
| 82 |
+
'optimization': { 'clip_grad': None,
|
| 83 |
+
'ema': [1.0, 0.99925],
|
| 84 |
+
'ema_hold_epochs': 0,
|
| 85 |
+
'ema_ramp_epochs': 5,
|
| 86 |
+
'encoder_layer_decay': 0.9,
|
| 87 |
+
'epochs': 60,
|
| 88 |
+
'final_lr': 0.0001,
|
| 89 |
+
'final_weight_decay': 0.04,
|
| 90 |
+
'ipe': 300,
|
| 91 |
+
'ipe_scale': 1.0,
|
| 92 |
+
'lr': 0.0001,
|
| 93 |
+
'predictor_lr_scale': 3.0,
|
| 94 |
+
'start_lr': 1e-05,
|
| 95 |
+
'warmup': 2,
|
| 96 |
+
'weight_decay': 0.04},
|
| 97 |
+
'tasks_per_node': 3}
|
| 98 |
+
INFO:root:Running pre-training of app: vjepa
|
| 99 |
+
[INFO ][2026-06-11 19:58:06][app.vjepa.train ][main ] which_dtype='bfloat16'
|
| 100 |
+
[INFO ][2026-06-11 19:58:06][app.vjepa.train ][main ] Disabling persistent_workers (incompatible with KNN eval)
|
| 101 |
+
[INFO ][2026-06-11 19:58:06][app.vjepa.train ][main ] NCCL_SOCKET_IFNAME=None
|
| 102 |
+
[INFO ][2026-06-11 19:58:09][app.vjepa.train ][main ] Initialized (rank/world-size) 3/6, tasks_per_node=3
|
| 103 |
+
[WARNING ][2026-06-11 19:58:23][experiments.stmodels.vision_transformers_v3][_build_dinov2_vit ] DINOv2 is an image model — fwd_chunk is unset so all frames will be processed jointly. Set fwd_chunk=1 for per-frame (standard) processing.
|
| 104 |
+
[INFO ][2026-06-11 19:58:28][experiments.stmodels.vision_transformers_v3][_build_dinov2_vit ] Loading pretrained weights for vit_large_patch14_dinov2_lvd142m from timm (vit_large_patch14_dinov2)
|
| 105 |
+
[INFO ][2026-06-11 19:58:32][timm.models._builder][load_pretrained ] Loading pretrained weights from Hugging Face hub (timm/vit_large_patch14_dinov2.lvd142m)
|
| 106 |
+
[INFO ][2026-06-11 19:58:32][timm.models._hub ][load_state_dict_from_hf ] [timm/vit_large_patch14_dinov2.lvd142m] Safe alternative available for 'pytorch_model.bin' (as 'model.safetensors'). Loading weights using safetensors.
|
| 107 |
+
[INFO ][2026-06-11 19:58:32][timm.layers.pos_embed][resample_abs_pos_embed ] Resized position embedding: (37, 37) to (16, 16).
|
| 108 |
+
[INFO ][2026-06-11 19:58:36][root ][init_video_model ] ViTMultiSeqWrapper(
|
| 109 |
+
(backbone): VisionTransformer(
|
| 110 |
+
(patch_embed): PatchEmbed3D(
|
| 111 |
+
(proj): Conv3d(3, 1024, kernel_size=(1, 14, 14), stride=(1, 14, 14))
|
| 112 |
+
)
|
| 113 |
+
(blocks): ModuleList(
|
| 114 |
+
(0-23): 24 x Block(
|
| 115 |
+
(norm1): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 116 |
+
(norm2): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 117 |
+
(drop_path1): Identity()
|
| 118 |
+
(drop_path2): Identity()
|
| 119 |
+
(attn): Attention(
|
| 120 |
+
(qkv): Linear(in_features=1024, out_features=3072, bias=True)
|
| 121 |
+
(attn_drop): Dropout(p=0.0, inplace=False)
|
| 122 |
+
(proj): Linear(in_features=1024, out_features=1024, bias=True)
|
| 123 |
+
(proj_drop): Dropout(p=0.0, inplace=False)
|
| 124 |
+
)
|
| 125 |
+
(mlp): MLP(
|
| 126 |
+
(fc1): Linear(in_features=1024, out_features=4096, bias=True)
|
| 127 |
+
(act): GELU(approximate='none')
|
| 128 |
+
(fc2): Linear(in_features=4096, out_features=1024, bias=True)
|
| 129 |
+
(drop): Dropout(p=0.0, inplace=False)
|
| 130 |
+
)
|
| 131 |
+
)
|
| 132 |
+
)
|
| 133 |
+
(norm): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 134 |
+
)
|
| 135 |
+
)
|
| 136 |
+
[INFO ][2026-06-11 19:58:36][root ][init_video_model ] VJEPAPredictorMultiSeqWrapper(
|
| 137 |
+
(backbone): PredictorV2(
|
| 138 |
+
(predictor_embed): Linear(in_features=1024, out_features=384, bias=True)
|
| 139 |
+
(mask_tokens): ParameterList(
|
| 140 |
+
(0): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
|
| 141 |
+
(1): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
|
| 142 |
+
(2): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
|
| 143 |
+
(3): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
|
| 144 |
+
)
|
| 145 |
+
(predictor_blocks): ModuleList(
|
| 146 |
+
(0-11): 12 x Block(
|
| 147 |
+
(residual1): EfficientResidual(
|
| 148 |
+
(norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
|
| 149 |
+
(fn): Attention(
|
| 150 |
+
(q_proj): Linear(in_features=384, out_features=384, bias=False)
|
| 151 |
+
(k_proj): Linear(in_features=384, out_features=384, bias=False)
|
| 152 |
+
(v_proj): Linear(in_features=384, out_features=384, bias=False)
|
| 153 |
+
(proj): Linear(in_features=384, out_features=384, bias=False)
|
| 154 |
+
(rope): Rope()
|
| 155 |
+
)
|
| 156 |
+
)
|
| 157 |
+
(residual2): EfficientResidual(
|
| 158 |
+
(norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
|
| 159 |
+
(fn): MLP(
|
| 160 |
+
(fc1): Linear(in_features=384, out_features=1536, bias=False)
|
| 161 |
+
(act): GELU(approximate='none')
|
| 162 |
+
(fc2): Linear(in_features=1536, out_features=384, bias=False)
|
| 163 |
+
(drop): Dropout(p=0.0, inplace=False)
|
| 164 |
+
)
|
| 165 |
+
)
|
| 166 |
+
)
|
| 167 |
+
)
|
| 168 |
+
(predictor_norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
|
| 169 |
+
(predictor_proj): Linear(in_features=384, out_features=1024, bias=True)
|
| 170 |
+
)
|
| 171 |
+
)
|
| 172 |
+
[INFO ][2026-06-11 19:58:36][root ][init_video_model ] ViTMultiSeqWrapper(
|
| 173 |
+
(backbone): VisionTransformer(
|
| 174 |
+
(patch_embed): PatchEmbed3D(
|
| 175 |
+
(proj): Conv3d(3, 1024, kernel_size=(1, 14, 14), stride=(1, 14, 14))
|
| 176 |
+
)
|
| 177 |
+
(blocks): ModuleList(
|
| 178 |
+
(0-23): 24 x Block(
|
| 179 |
+
(norm1): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 180 |
+
(norm2): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 181 |
+
(drop_path1): Identity()
|
| 182 |
+
(drop_path2): Identity()
|
| 183 |
+
(attn): Attention(
|
| 184 |
+
(qkv): Linear(in_features=1024, out_features=3072, bias=True)
|
| 185 |
+
(attn_drop): Dropout(p=0.0, inplace=False)
|
| 186 |
+
(proj): Linear(in_features=1024, out_features=1024, bias=True)
|
| 187 |
+
(proj_drop): Dropout(p=0.0, inplace=False)
|
| 188 |
+
)
|
| 189 |
+
(mlp): MLP(
|
| 190 |
+
(fc1): Linear(in_features=1024, out_features=4096, bias=True)
|
| 191 |
+
(act): GELU(approximate='none')
|
| 192 |
+
(fc2): Linear(in_features=4096, out_features=1024, bias=True)
|
| 193 |
+
(drop): Dropout(p=0.0, inplace=False)
|
| 194 |
+
)
|
| 195 |
+
)
|
| 196 |
+
)
|
| 197 |
+
(norm): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 198 |
+
)
|
| 199 |
+
)
|
| 200 |
+
[INFO ][2026-06-11 19:58:36][root ][init_video_model ] Encoder number of parameters: 302964736
|
| 201 |
+
[INFO ][2026-06-11 19:58:36][root ][init_video_model ] Predictor number of parameters: 22042240
|
| 202 |
+
[INFO ][2026-06-11 19:58:36][root ][init_video_model ] Target encoder number of parameters: 302964736
|
| 203 |
+
[INFO ][2026-06-11 19:58:37][root ][make_videodataset ] VideoDataset dataset created
|
| 204 |
+
[INFO ][2026-06-11 19:58:37][WeightedSampler ][__init__ ] Using DistributedWeightedSampler with rank 3 / 6
|
| 205 |
+
[INFO ][2026-06-11 19:58:37][root ][make_videodataset ] VideoDataset unsupervised data loader created
|
| 206 |
+
[INFO ][2026-06-11 19:58:37][app.vjepa.train ][main ] iterations per epoch/dataset length: 300/2103
|
| 207 |
+
[INFO ][2026-06-11 19:58:37][app.vjepa.train ][main ] Wrapping models in DDP (rank 3)...
|
| 208 |
+
[INFO ][2026-06-11 19:58:38][root ][load_checkpoint ] Loading checkpoint from /var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/latest.pt
|
| 209 |
+
[INFO ][2026-06-11 19:58:48][root ][load_checkpoint ] loaded pretrained encoder from epoch 30 with msg: <All keys matched successfully>
|
| 210 |
+
[INFO ][2026-06-11 19:58:48][root ][load_checkpoint ] loaded pretrained predictor from epoch 30 with msg: <All keys matched successfully>
|
| 211 |
+
[INFO ][2026-06-11 19:58:48][root ][load_checkpoint ] loaded pretrained target encoder from epoch 30 with msg: <All keys matched successfully>
|
| 212 |
+
[INFO ][2026-06-11 19:58:48][root ][load_checkpoint ] loaded optimizers from epoch 30
|
| 213 |
+
[INFO ][2026-06-11 19:58:48][root ][load_checkpoint ] read-path: /var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/latest.pt
|
| 214 |
+
[INFO ][2026-06-11 20:00:08][app.vjepa.train ][main ] Initializing loader...
|
| 215 |
+
submitit INFO (2026-06-12 08:09:41,170) - Job completed successfully
|
| 216 |
+
[INFO ][2026-06-12 08:09:41][submitit ][process_job ] Job completed successfully
|
| 217 |
+
submitit INFO (2026-06-12 08:09:41,173) - Exiting after successful completion
|
| 218 |
+
[INFO ][2026-06-12 08:09:41][submitit ][process_job ] Exiting after successful completion
|
inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_25214/25214_4_log.err
ADDED
|
@@ -0,0 +1,33 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
/var/scratch/dcanezil/.cache/uv/virtualenvs/vd/lib/python3.12/site-packages/timm/models/layers/__init__.py:49: FutureWarning: Importing from timm.models.layers is deprecated, please import via timm.layers
|
| 2 |
+
warnings.warn(f"Importing from {__name__} is deprecated, please import via timm.layers", FutureWarning)
|
| 3 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/utils.py:888: FutureWarning: `torch.cuda.amp.GradScaler(args...)` is deprecated. Please use `torch.amp.GradScaler('cuda', args...)` instead.
|
| 4 |
+
scaler = torch.cuda.amp.GradScaler() if mixed_precision else None
|
| 5 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 6 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 7 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/windsurfing/i-gzh_BPDa8_000154_000164.avi'
|
| 8 |
+
warnings.warn(f"video path not found {fname=}")
|
| 9 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_poker/6-DyF8umej8_000097_000107.avi'
|
| 10 |
+
warnings.warn(f"video path not found {fname=}")
|
| 11 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_poker/6-DyF8umej8_000097_000107.avi'
|
| 12 |
+
warnings.warn(f"video path not found {fname=}")
|
| 13 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_flute/co50KUHacYw_000005_000015.avi'
|
| 14 |
+
warnings.warn(f"video path not found {fname=}")
|
| 15 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/baking_cookies/Q3XGZmqk1Q0_000296_000306.avi'
|
| 16 |
+
warnings.warn(f"video path not found {fname=}")
|
| 17 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/surfing_water/5l5Pdd96Pao_000161_000171.avi'
|
| 18 |
+
warnings.warn(f"video path not found {fname=}")
|
| 19 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_poker/-jGxlNQKkeo_000092_000102.avi'
|
| 20 |
+
warnings.warn(f"video path not found {fname=}")
|
| 21 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/picking_fruit/3VvkoFtPCCU_000045_000055.avi'
|
| 22 |
+
warnings.warn(f"video path not found {fname=}")
|
| 23 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/bowling/f4eb0wOlspM_000053_000063.avi'
|
| 24 |
+
warnings.warn(f"video path not found {fname=}")
|
| 25 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_saxophone/Zs0_2GMEPXo_000054_000064.avi'
|
| 26 |
+
warnings.warn(f"video path not found {fname=}")
|
| 27 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/making_tea/mtYFNsRcxY4_000063_000073.avi'
|
| 28 |
+
warnings.warn(f"video path not found {fname=}")
|
| 29 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/pumping_fist/wKsuyr6Xk30_000094_000104.avi'
|
| 30 |
+
warnings.warn(f"video path not found {fname=}")
|
| 31 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/making_tea/mtYFNsRcxY4_000063_000073.avi'
|
| 32 |
+
warnings.warn(f"video path not found {fname=}")
|
| 33 |
+
[rank4]:[W612 08:09:41.289808277 ProcessGroupNCCL.cpp:1553] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())
|
inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_25214/25214_4_log.out
ADDED
|
@@ -0,0 +1,218 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
submitit INFO (2026-06-11 19:57:57,500) - Starting with JobEnvironment(job_id=25214, hostname=node406, local_rank=1(3), node=1(2), global_rank=4(6))
|
| 2 |
+
submitit INFO (2026-06-11 19:57:57,500) - Loading pickle: /var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_25214/25214_submitted.pkl
|
| 3 |
+
INFO:root:loaded pretrain params...
|
| 4 |
+
{ 'app': 'vjepa',
|
| 5 |
+
'cpus_per_task': 16,
|
| 6 |
+
'data': { 'batch_size': 32,
|
| 7 |
+
'crop_size': 224,
|
| 8 |
+
'dataset_fpcs': [8, 8],
|
| 9 |
+
'dataset_type': 'VideoDataset',
|
| 10 |
+
'datasets': [ '/local/dcanezil/data/kinetics/train.csv',
|
| 11 |
+
'/local/dcanezil/data/ssv2/train.csv'],
|
| 12 |
+
'datasets_weights': [0.65, 0.35],
|
| 13 |
+
'fps': 2,
|
| 14 |
+
'num_workers': 4,
|
| 15 |
+
'patch_size': 14,
|
| 16 |
+
'persistent_workers': True,
|
| 17 |
+
'pin_mem': True,
|
| 18 |
+
'tubelet_size': 1},
|
| 19 |
+
'data_aug': { 'auto_augment': False,
|
| 20 |
+
'motion_shift': False,
|
| 21 |
+
'random_resize_aspect_ratio': [0.75, 1.35],
|
| 22 |
+
'random_resize_scale': [0.3, 1.0],
|
| 23 |
+
'reprob': 0.0},
|
| 24 |
+
'folder': '/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow',
|
| 25 |
+
'loss': {'loss_exp': 1.0},
|
| 26 |
+
'mask': [ { 'aspect_ratio': [0.75, 1.5],
|
| 27 |
+
'full_complement': False,
|
| 28 |
+
'max_keep': None,
|
| 29 |
+
'max_temporal_keep': 1.0,
|
| 30 |
+
'num_blocks': 8,
|
| 31 |
+
'spatial_scale': [0.05, 0.05],
|
| 32 |
+
'temporal_scale': [1.0, 1.0]},
|
| 33 |
+
{ 'aspect_ratio': [0.75, 1.5],
|
| 34 |
+
'full_complement': False,
|
| 35 |
+
'max_keep': None,
|
| 36 |
+
'max_temporal_keep': 1.0,
|
| 37 |
+
'num_blocks': 2,
|
| 38 |
+
'spatial_scale': [0.4, 0.4],
|
| 39 |
+
'temporal_scale': [1.0, 1.0]}],
|
| 40 |
+
'mem_per_gpu': '40G',
|
| 41 |
+
'meta': { 'dtype': 'bfloat16',
|
| 42 |
+
'knn_eval_epoch0': True,
|
| 43 |
+
'knn_eval_freq': 5,
|
| 44 |
+
'knn_eval_presets': [ { 'config': { 'batch_size': 8,
|
| 45 |
+
'dataset_train': '/local/dcanezil/data/ssv2/train_coarse10.csv',
|
| 46 |
+
'dataset_val': '/local/dcanezil/data/ssv2/val_coarse10.csv',
|
| 47 |
+
'eval_videos_per_class': 100,
|
| 48 |
+
'linear_probe': True,
|
| 49 |
+
'num_workers': 2,
|
| 50 |
+
'pool_type': 'temporal_concat',
|
| 51 |
+
'train_videos_per_class': 500},
|
| 52 |
+
'preset': 'ssv2_coarse10'}],
|
| 53 |
+
'load_checkpoint': True,
|
| 54 |
+
'read_checkpoint': None,
|
| 55 |
+
'save_every_freq': 10,
|
| 56 |
+
'seed': 239,
|
| 57 |
+
'use_sdpa': True,
|
| 58 |
+
'use_wandb': True,
|
| 59 |
+
'wandb_project': 'vjepa_ablation'},
|
| 60 |
+
'metrics': {'sigreg': {}, 'std': {}},
|
| 61 |
+
'model': { 'class_token': True,
|
| 62 |
+
'freeze_backbone': False,
|
| 63 |
+
'is_causal': True,
|
| 64 |
+
'model_name': 'vit_large_patch14_dinov2_lvd142m',
|
| 65 |
+
'pe_type': '2d',
|
| 66 |
+
'pred_depth': 12,
|
| 67 |
+
'pred_embed_dim': 384,
|
| 68 |
+
'pred_num_heads': 6,
|
| 69 |
+
'predictor': 'v2_cross',
|
| 70 |
+
'stem_type': '2d',
|
| 71 |
+
'target_embed_dim': 1024,
|
| 72 |
+
'target_kind': 'ema',
|
| 73 |
+
'uniform_power': True,
|
| 74 |
+
'use_activation_checkpointing': True,
|
| 75 |
+
'use_alibi': False,
|
| 76 |
+
'use_lora': False,
|
| 77 |
+
'use_mask_tokens': True,
|
| 78 |
+
'use_rope': False,
|
| 79 |
+
'use_sdpa': True,
|
| 80 |
+
'zero_init_mask_tokens': True},
|
| 81 |
+
'nodes': 2,
|
| 82 |
+
'optimization': { 'clip_grad': None,
|
| 83 |
+
'ema': [1.0, 0.99925],
|
| 84 |
+
'ema_hold_epochs': 0,
|
| 85 |
+
'ema_ramp_epochs': 5,
|
| 86 |
+
'encoder_layer_decay': 0.9,
|
| 87 |
+
'epochs': 60,
|
| 88 |
+
'final_lr': 0.0001,
|
| 89 |
+
'final_weight_decay': 0.04,
|
| 90 |
+
'ipe': 300,
|
| 91 |
+
'ipe_scale': 1.0,
|
| 92 |
+
'lr': 0.0001,
|
| 93 |
+
'predictor_lr_scale': 3.0,
|
| 94 |
+
'start_lr': 1e-05,
|
| 95 |
+
'warmup': 2,
|
| 96 |
+
'weight_decay': 0.04},
|
| 97 |
+
'tasks_per_node': 3}
|
| 98 |
+
INFO:root:Running pre-training of app: vjepa
|
| 99 |
+
[INFO ][2026-06-11 19:58:06][app.vjepa.train ][main ] which_dtype='bfloat16'
|
| 100 |
+
[INFO ][2026-06-11 19:58:06][app.vjepa.train ][main ] Disabling persistent_workers (incompatible with KNN eval)
|
| 101 |
+
[INFO ][2026-06-11 19:58:06][app.vjepa.train ][main ] NCCL_SOCKET_IFNAME=None
|
| 102 |
+
[INFO ][2026-06-11 19:58:09][app.vjepa.train ][main ] Initialized (rank/world-size) 4/6, tasks_per_node=3
|
| 103 |
+
[WARNING ][2026-06-11 19:58:23][experiments.stmodels.vision_transformers_v3][_build_dinov2_vit ] DINOv2 is an image model — fwd_chunk is unset so all frames will be processed jointly. Set fwd_chunk=1 for per-frame (standard) processing.
|
| 104 |
+
[INFO ][2026-06-11 19:58:28][experiments.stmodels.vision_transformers_v3][_build_dinov2_vit ] Loading pretrained weights for vit_large_patch14_dinov2_lvd142m from timm (vit_large_patch14_dinov2)
|
| 105 |
+
[INFO ][2026-06-11 19:58:32][timm.models._builder][load_pretrained ] Loading pretrained weights from Hugging Face hub (timm/vit_large_patch14_dinov2.lvd142m)
|
| 106 |
+
[INFO ][2026-06-11 19:58:32][timm.models._hub ][load_state_dict_from_hf ] [timm/vit_large_patch14_dinov2.lvd142m] Safe alternative available for 'pytorch_model.bin' (as 'model.safetensors'). Loading weights using safetensors.
|
| 107 |
+
[INFO ][2026-06-11 19:58:32][timm.layers.pos_embed][resample_abs_pos_embed ] Resized position embedding: (37, 37) to (16, 16).
|
| 108 |
+
[INFO ][2026-06-11 19:58:36][root ][init_video_model ] ViTMultiSeqWrapper(
|
| 109 |
+
(backbone): VisionTransformer(
|
| 110 |
+
(patch_embed): PatchEmbed3D(
|
| 111 |
+
(proj): Conv3d(3, 1024, kernel_size=(1, 14, 14), stride=(1, 14, 14))
|
| 112 |
+
)
|
| 113 |
+
(blocks): ModuleList(
|
| 114 |
+
(0-23): 24 x Block(
|
| 115 |
+
(norm1): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 116 |
+
(norm2): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 117 |
+
(drop_path1): Identity()
|
| 118 |
+
(drop_path2): Identity()
|
| 119 |
+
(attn): Attention(
|
| 120 |
+
(qkv): Linear(in_features=1024, out_features=3072, bias=True)
|
| 121 |
+
(attn_drop): Dropout(p=0.0, inplace=False)
|
| 122 |
+
(proj): Linear(in_features=1024, out_features=1024, bias=True)
|
| 123 |
+
(proj_drop): Dropout(p=0.0, inplace=False)
|
| 124 |
+
)
|
| 125 |
+
(mlp): MLP(
|
| 126 |
+
(fc1): Linear(in_features=1024, out_features=4096, bias=True)
|
| 127 |
+
(act): GELU(approximate='none')
|
| 128 |
+
(fc2): Linear(in_features=4096, out_features=1024, bias=True)
|
| 129 |
+
(drop): Dropout(p=0.0, inplace=False)
|
| 130 |
+
)
|
| 131 |
+
)
|
| 132 |
+
)
|
| 133 |
+
(norm): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 134 |
+
)
|
| 135 |
+
)
|
| 136 |
+
[INFO ][2026-06-11 19:58:36][root ][init_video_model ] VJEPAPredictorMultiSeqWrapper(
|
| 137 |
+
(backbone): PredictorV2(
|
| 138 |
+
(predictor_embed): Linear(in_features=1024, out_features=384, bias=True)
|
| 139 |
+
(mask_tokens): ParameterList(
|
| 140 |
+
(0): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
|
| 141 |
+
(1): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
|
| 142 |
+
(2): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
|
| 143 |
+
(3): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
|
| 144 |
+
)
|
| 145 |
+
(predictor_blocks): ModuleList(
|
| 146 |
+
(0-11): 12 x Block(
|
| 147 |
+
(residual1): EfficientResidual(
|
| 148 |
+
(norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
|
| 149 |
+
(fn): Attention(
|
| 150 |
+
(q_proj): Linear(in_features=384, out_features=384, bias=False)
|
| 151 |
+
(k_proj): Linear(in_features=384, out_features=384, bias=False)
|
| 152 |
+
(v_proj): Linear(in_features=384, out_features=384, bias=False)
|
| 153 |
+
(proj): Linear(in_features=384, out_features=384, bias=False)
|
| 154 |
+
(rope): Rope()
|
| 155 |
+
)
|
| 156 |
+
)
|
| 157 |
+
(residual2): EfficientResidual(
|
| 158 |
+
(norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
|
| 159 |
+
(fn): MLP(
|
| 160 |
+
(fc1): Linear(in_features=384, out_features=1536, bias=False)
|
| 161 |
+
(act): GELU(approximate='none')
|
| 162 |
+
(fc2): Linear(in_features=1536, out_features=384, bias=False)
|
| 163 |
+
(drop): Dropout(p=0.0, inplace=False)
|
| 164 |
+
)
|
| 165 |
+
)
|
| 166 |
+
)
|
| 167 |
+
)
|
| 168 |
+
(predictor_norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
|
| 169 |
+
(predictor_proj): Linear(in_features=384, out_features=1024, bias=True)
|
| 170 |
+
)
|
| 171 |
+
)
|
| 172 |
+
[INFO ][2026-06-11 19:58:36][root ][init_video_model ] ViTMultiSeqWrapper(
|
| 173 |
+
(backbone): VisionTransformer(
|
| 174 |
+
(patch_embed): PatchEmbed3D(
|
| 175 |
+
(proj): Conv3d(3, 1024, kernel_size=(1, 14, 14), stride=(1, 14, 14))
|
| 176 |
+
)
|
| 177 |
+
(blocks): ModuleList(
|
| 178 |
+
(0-23): 24 x Block(
|
| 179 |
+
(norm1): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 180 |
+
(norm2): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 181 |
+
(drop_path1): Identity()
|
| 182 |
+
(drop_path2): Identity()
|
| 183 |
+
(attn): Attention(
|
| 184 |
+
(qkv): Linear(in_features=1024, out_features=3072, bias=True)
|
| 185 |
+
(attn_drop): Dropout(p=0.0, inplace=False)
|
| 186 |
+
(proj): Linear(in_features=1024, out_features=1024, bias=True)
|
| 187 |
+
(proj_drop): Dropout(p=0.0, inplace=False)
|
| 188 |
+
)
|
| 189 |
+
(mlp): MLP(
|
| 190 |
+
(fc1): Linear(in_features=1024, out_features=4096, bias=True)
|
| 191 |
+
(act): GELU(approximate='none')
|
| 192 |
+
(fc2): Linear(in_features=4096, out_features=1024, bias=True)
|
| 193 |
+
(drop): Dropout(p=0.0, inplace=False)
|
| 194 |
+
)
|
| 195 |
+
)
|
| 196 |
+
)
|
| 197 |
+
(norm): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 198 |
+
)
|
| 199 |
+
)
|
| 200 |
+
[INFO ][2026-06-11 19:58:36][root ][init_video_model ] Encoder number of parameters: 302964736
|
| 201 |
+
[INFO ][2026-06-11 19:58:36][root ][init_video_model ] Predictor number of parameters: 22042240
|
| 202 |
+
[INFO ][2026-06-11 19:58:36][root ][init_video_model ] Target encoder number of parameters: 302964736
|
| 203 |
+
[INFO ][2026-06-11 19:58:37][root ][make_videodataset ] VideoDataset dataset created
|
| 204 |
+
[INFO ][2026-06-11 19:58:37][WeightedSampler ][__init__ ] Using DistributedWeightedSampler with rank 4 / 6
|
| 205 |
+
[INFO ][2026-06-11 19:58:37][root ][make_videodataset ] VideoDataset unsupervised data loader created
|
| 206 |
+
[INFO ][2026-06-11 19:58:37][app.vjepa.train ][main ] iterations per epoch/dataset length: 300/2103
|
| 207 |
+
[INFO ][2026-06-11 19:58:37][app.vjepa.train ][main ] Wrapping models in DDP (rank 4)...
|
| 208 |
+
[INFO ][2026-06-11 19:58:38][root ][load_checkpoint ] Loading checkpoint from /var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/latest.pt
|
| 209 |
+
[INFO ][2026-06-11 19:58:48][root ][load_checkpoint ] loaded pretrained encoder from epoch 30 with msg: <All keys matched successfully>
|
| 210 |
+
[INFO ][2026-06-11 19:58:48][root ][load_checkpoint ] loaded pretrained predictor from epoch 30 with msg: <All keys matched successfully>
|
| 211 |
+
[INFO ][2026-06-11 19:58:48][root ][load_checkpoint ] loaded pretrained target encoder from epoch 30 with msg: <All keys matched successfully>
|
| 212 |
+
[INFO ][2026-06-11 19:58:48][root ][load_checkpoint ] loaded optimizers from epoch 30
|
| 213 |
+
[INFO ][2026-06-11 19:58:48][root ][load_checkpoint ] read-path: /var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/latest.pt
|
| 214 |
+
[INFO ][2026-06-11 20:00:08][app.vjepa.train ][main ] Initializing loader...
|
| 215 |
+
submitit INFO (2026-06-12 08:09:41,141) - Job completed successfully
|
| 216 |
+
[INFO ][2026-06-12 08:09:41][submitit ][process_job ] Job completed successfully
|
| 217 |
+
submitit INFO (2026-06-12 08:09:41,144) - Exiting after successful completion
|
| 218 |
+
[INFO ][2026-06-12 08:09:41][submitit ][process_job ] Exiting after successful completion
|
inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_25214/25214_5_log.err
ADDED
|
@@ -0,0 +1,47 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
/var/scratch/dcanezil/.cache/uv/virtualenvs/vd/lib/python3.12/site-packages/timm/models/layers/__init__.py:49: FutureWarning: Importing from timm.models.layers is deprecated, please import via timm.layers
|
| 2 |
+
warnings.warn(f"Importing from {__name__} is deprecated, please import via timm.layers", FutureWarning)
|
| 3 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/utils.py:888: FutureWarning: `torch.cuda.amp.GradScaler(args...)` is deprecated. Please use `torch.amp.GradScaler('cuda', args...)` instead.
|
| 4 |
+
scaler = torch.cuda.amp.GradScaler() if mixed_precision else None
|
| 5 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 6 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 7 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/eating_carrots/eiZ8Hzc7FPU_000080_000090.avi'
|
| 8 |
+
warnings.warn(f"video path not found {fname=}")
|
| 9 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/sailing/ZolXG_-RgIs_000970_000980.avi'
|
| 10 |
+
warnings.warn(f"video path not found {fname=}")
|
| 11 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/surfing_water/5l5Pdd96Pao_000161_000171.avi'
|
| 12 |
+
warnings.warn(f"video path not found {fname=}")
|
| 13 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/sailing/ZolXG_-RgIs_000970_000980.avi'
|
| 14 |
+
warnings.warn(f"video path not found {fname=}")
|
| 15 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/waxing_legs/cqUusXBODuw_001046_001056.avi'
|
| 16 |
+
warnings.warn(f"video path not found {fname=}")
|
| 17 |
+
[23:40:54] /github/workspace/src/video/video_reader.cc:83: ERROR opening: /local/dcanezil/data/kinetics/VideoData/long_jump/jsYiitkbr5w_000004_000014.avi, Invalid data found when processing input
|
| 18 |
+
[00:15:42] /github/workspace/src/video/video_reader.cc:83: ERROR opening: /local/dcanezil/data/kinetics/VideoData/long_jump/jsYiitkbr5w_000004_000014.avi, Invalid data found when processing input
|
| 19 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/pumping_fist/wKsuyr6Xk30_000094_000104.avi'
|
| 20 |
+
warnings.warn(f"video path not found {fname=}")
|
| 21 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_flute/co50KUHacYw_000005_000015.avi'
|
| 22 |
+
warnings.warn(f"video path not found {fname=}")
|
| 23 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/sailing/ZolXG_-RgIs_000970_000980.avi'
|
| 24 |
+
warnings.warn(f"video path not found {fname=}")
|
| 25 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/robot_dancing/J1x0XlWa6HM_000147_000157.avi'
|
| 26 |
+
warnings.warn(f"video path not found {fname=}")
|
| 27 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/waxing_legs/cqUusXBODuw_001046_001056.avi'
|
| 28 |
+
warnings.warn(f"video path not found {fname=}")
|
| 29 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/dancing_macarena/uGQgxFHemyA_000059_000069.avi'
|
| 30 |
+
warnings.warn(f"video path not found {fname=}")
|
| 31 |
+
[02:08:41] /github/workspace/src/video/video_reader.cc:83: ERROR opening: /local/dcanezil/data/kinetics/VideoData/trimming_trees/T6dqaZioaXs_000004_000014.avi, Invalid data found when processing input
|
| 32 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/dancing_macarena/uGQgxFHemyA_000059_000069.avi'
|
| 33 |
+
warnings.warn(f"video path not found {fname=}")
|
| 34 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_poker/6-DyF8umej8_000097_000107.avi'
|
| 35 |
+
warnings.warn(f"video path not found {fname=}")
|
| 36 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_flute/co50KUHacYw_000005_000015.avi'
|
| 37 |
+
warnings.warn(f"video path not found {fname=}")
|
| 38 |
+
[03:06:55] /github/workspace/src/video/video_reader.cc:83: ERROR opening: /local/dcanezil/data/kinetics/VideoData/climbing_tree/QX5IzK4iFCM_000002_000012.avi, Invalid data found when processing input
|
| 39 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/frying_vegetables/Nr-zCXgW2Sw_000161_000171.avi'
|
| 40 |
+
warnings.warn(f"video path not found {fname=}")
|
| 41 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/sailing/ZolXG_-RgIs_000970_000980.avi'
|
| 42 |
+
warnings.warn(f"video path not found {fname=}")
|
| 43 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/picking_fruit/3VvkoFtPCCU_000045_000055.avi'
|
| 44 |
+
warnings.warn(f"video path not found {fname=}")
|
| 45 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/bowling/f4eb0wOlspM_000053_000063.avi'
|
| 46 |
+
warnings.warn(f"video path not found {fname=}")
|
| 47 |
+
[rank5]:[W612 08:09:42.656736785 ProcessGroupNCCL.cpp:1553] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())
|
inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_25214/25214_5_log.out
ADDED
|
@@ -0,0 +1,218 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
submitit INFO (2026-06-11 19:57:57,500) - Starting with JobEnvironment(job_id=25214, hostname=node406, local_rank=2(3), node=1(2), global_rank=5(6))
|
| 2 |
+
submitit INFO (2026-06-11 19:57:57,500) - Loading pickle: /var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_25214/25214_submitted.pkl
|
| 3 |
+
INFO:root:loaded pretrain params...
|
| 4 |
+
{ 'app': 'vjepa',
|
| 5 |
+
'cpus_per_task': 16,
|
| 6 |
+
'data': { 'batch_size': 32,
|
| 7 |
+
'crop_size': 224,
|
| 8 |
+
'dataset_fpcs': [8, 8],
|
| 9 |
+
'dataset_type': 'VideoDataset',
|
| 10 |
+
'datasets': [ '/local/dcanezil/data/kinetics/train.csv',
|
| 11 |
+
'/local/dcanezil/data/ssv2/train.csv'],
|
| 12 |
+
'datasets_weights': [0.65, 0.35],
|
| 13 |
+
'fps': 2,
|
| 14 |
+
'num_workers': 4,
|
| 15 |
+
'patch_size': 14,
|
| 16 |
+
'persistent_workers': True,
|
| 17 |
+
'pin_mem': True,
|
| 18 |
+
'tubelet_size': 1},
|
| 19 |
+
'data_aug': { 'auto_augment': False,
|
| 20 |
+
'motion_shift': False,
|
| 21 |
+
'random_resize_aspect_ratio': [0.75, 1.35],
|
| 22 |
+
'random_resize_scale': [0.3, 1.0],
|
| 23 |
+
'reprob': 0.0},
|
| 24 |
+
'folder': '/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow',
|
| 25 |
+
'loss': {'loss_exp': 1.0},
|
| 26 |
+
'mask': [ { 'aspect_ratio': [0.75, 1.5],
|
| 27 |
+
'full_complement': False,
|
| 28 |
+
'max_keep': None,
|
| 29 |
+
'max_temporal_keep': 1.0,
|
| 30 |
+
'num_blocks': 8,
|
| 31 |
+
'spatial_scale': [0.05, 0.05],
|
| 32 |
+
'temporal_scale': [1.0, 1.0]},
|
| 33 |
+
{ 'aspect_ratio': [0.75, 1.5],
|
| 34 |
+
'full_complement': False,
|
| 35 |
+
'max_keep': None,
|
| 36 |
+
'max_temporal_keep': 1.0,
|
| 37 |
+
'num_blocks': 2,
|
| 38 |
+
'spatial_scale': [0.4, 0.4],
|
| 39 |
+
'temporal_scale': [1.0, 1.0]}],
|
| 40 |
+
'mem_per_gpu': '40G',
|
| 41 |
+
'meta': { 'dtype': 'bfloat16',
|
| 42 |
+
'knn_eval_epoch0': True,
|
| 43 |
+
'knn_eval_freq': 5,
|
| 44 |
+
'knn_eval_presets': [ { 'config': { 'batch_size': 8,
|
| 45 |
+
'dataset_train': '/local/dcanezil/data/ssv2/train_coarse10.csv',
|
| 46 |
+
'dataset_val': '/local/dcanezil/data/ssv2/val_coarse10.csv',
|
| 47 |
+
'eval_videos_per_class': 100,
|
| 48 |
+
'linear_probe': True,
|
| 49 |
+
'num_workers': 2,
|
| 50 |
+
'pool_type': 'temporal_concat',
|
| 51 |
+
'train_videos_per_class': 500},
|
| 52 |
+
'preset': 'ssv2_coarse10'}],
|
| 53 |
+
'load_checkpoint': True,
|
| 54 |
+
'read_checkpoint': None,
|
| 55 |
+
'save_every_freq': 10,
|
| 56 |
+
'seed': 239,
|
| 57 |
+
'use_sdpa': True,
|
| 58 |
+
'use_wandb': True,
|
| 59 |
+
'wandb_project': 'vjepa_ablation'},
|
| 60 |
+
'metrics': {'sigreg': {}, 'std': {}},
|
| 61 |
+
'model': { 'class_token': True,
|
| 62 |
+
'freeze_backbone': False,
|
| 63 |
+
'is_causal': True,
|
| 64 |
+
'model_name': 'vit_large_patch14_dinov2_lvd142m',
|
| 65 |
+
'pe_type': '2d',
|
| 66 |
+
'pred_depth': 12,
|
| 67 |
+
'pred_embed_dim': 384,
|
| 68 |
+
'pred_num_heads': 6,
|
| 69 |
+
'predictor': 'v2_cross',
|
| 70 |
+
'stem_type': '2d',
|
| 71 |
+
'target_embed_dim': 1024,
|
| 72 |
+
'target_kind': 'ema',
|
| 73 |
+
'uniform_power': True,
|
| 74 |
+
'use_activation_checkpointing': True,
|
| 75 |
+
'use_alibi': False,
|
| 76 |
+
'use_lora': False,
|
| 77 |
+
'use_mask_tokens': True,
|
| 78 |
+
'use_rope': False,
|
| 79 |
+
'use_sdpa': True,
|
| 80 |
+
'zero_init_mask_tokens': True},
|
| 81 |
+
'nodes': 2,
|
| 82 |
+
'optimization': { 'clip_grad': None,
|
| 83 |
+
'ema': [1.0, 0.99925],
|
| 84 |
+
'ema_hold_epochs': 0,
|
| 85 |
+
'ema_ramp_epochs': 5,
|
| 86 |
+
'encoder_layer_decay': 0.9,
|
| 87 |
+
'epochs': 60,
|
| 88 |
+
'final_lr': 0.0001,
|
| 89 |
+
'final_weight_decay': 0.04,
|
| 90 |
+
'ipe': 300,
|
| 91 |
+
'ipe_scale': 1.0,
|
| 92 |
+
'lr': 0.0001,
|
| 93 |
+
'predictor_lr_scale': 3.0,
|
| 94 |
+
'start_lr': 1e-05,
|
| 95 |
+
'warmup': 2,
|
| 96 |
+
'weight_decay': 0.04},
|
| 97 |
+
'tasks_per_node': 3}
|
| 98 |
+
INFO:root:Running pre-training of app: vjepa
|
| 99 |
+
[INFO ][2026-06-11 19:58:06][app.vjepa.train ][main ] which_dtype='bfloat16'
|
| 100 |
+
[INFO ][2026-06-11 19:58:06][app.vjepa.train ][main ] Disabling persistent_workers (incompatible with KNN eval)
|
| 101 |
+
[INFO ][2026-06-11 19:58:06][app.vjepa.train ][main ] NCCL_SOCKET_IFNAME=None
|
| 102 |
+
[INFO ][2026-06-11 19:58:09][app.vjepa.train ][main ] Initialized (rank/world-size) 5/6, tasks_per_node=3
|
| 103 |
+
[WARNING ][2026-06-11 19:58:23][experiments.stmodels.vision_transformers_v3][_build_dinov2_vit ] DINOv2 is an image model — fwd_chunk is unset so all frames will be processed jointly. Set fwd_chunk=1 for per-frame (standard) processing.
|
| 104 |
+
[INFO ][2026-06-11 19:58:28][experiments.stmodels.vision_transformers_v3][_build_dinov2_vit ] Loading pretrained weights for vit_large_patch14_dinov2_lvd142m from timm (vit_large_patch14_dinov2)
|
| 105 |
+
[INFO ][2026-06-11 19:58:32][timm.models._builder][load_pretrained ] Loading pretrained weights from Hugging Face hub (timm/vit_large_patch14_dinov2.lvd142m)
|
| 106 |
+
[INFO ][2026-06-11 19:58:32][timm.models._hub ][load_state_dict_from_hf ] [timm/vit_large_patch14_dinov2.lvd142m] Safe alternative available for 'pytorch_model.bin' (as 'model.safetensors'). Loading weights using safetensors.
|
| 107 |
+
[INFO ][2026-06-11 19:58:32][timm.layers.pos_embed][resample_abs_pos_embed ] Resized position embedding: (37, 37) to (16, 16).
|
| 108 |
+
[INFO ][2026-06-11 19:58:36][root ][init_video_model ] ViTMultiSeqWrapper(
|
| 109 |
+
(backbone): VisionTransformer(
|
| 110 |
+
(patch_embed): PatchEmbed3D(
|
| 111 |
+
(proj): Conv3d(3, 1024, kernel_size=(1, 14, 14), stride=(1, 14, 14))
|
| 112 |
+
)
|
| 113 |
+
(blocks): ModuleList(
|
| 114 |
+
(0-23): 24 x Block(
|
| 115 |
+
(norm1): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 116 |
+
(norm2): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 117 |
+
(drop_path1): Identity()
|
| 118 |
+
(drop_path2): Identity()
|
| 119 |
+
(attn): Attention(
|
| 120 |
+
(qkv): Linear(in_features=1024, out_features=3072, bias=True)
|
| 121 |
+
(attn_drop): Dropout(p=0.0, inplace=False)
|
| 122 |
+
(proj): Linear(in_features=1024, out_features=1024, bias=True)
|
| 123 |
+
(proj_drop): Dropout(p=0.0, inplace=False)
|
| 124 |
+
)
|
| 125 |
+
(mlp): MLP(
|
| 126 |
+
(fc1): Linear(in_features=1024, out_features=4096, bias=True)
|
| 127 |
+
(act): GELU(approximate='none')
|
| 128 |
+
(fc2): Linear(in_features=4096, out_features=1024, bias=True)
|
| 129 |
+
(drop): Dropout(p=0.0, inplace=False)
|
| 130 |
+
)
|
| 131 |
+
)
|
| 132 |
+
)
|
| 133 |
+
(norm): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 134 |
+
)
|
| 135 |
+
)
|
| 136 |
+
[INFO ][2026-06-11 19:58:36][root ][init_video_model ] VJEPAPredictorMultiSeqWrapper(
|
| 137 |
+
(backbone): PredictorV2(
|
| 138 |
+
(predictor_embed): Linear(in_features=1024, out_features=384, bias=True)
|
| 139 |
+
(mask_tokens): ParameterList(
|
| 140 |
+
(0): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
|
| 141 |
+
(1): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
|
| 142 |
+
(2): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
|
| 143 |
+
(3): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
|
| 144 |
+
)
|
| 145 |
+
(predictor_blocks): ModuleList(
|
| 146 |
+
(0-11): 12 x Block(
|
| 147 |
+
(residual1): EfficientResidual(
|
| 148 |
+
(norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
|
| 149 |
+
(fn): Attention(
|
| 150 |
+
(q_proj): Linear(in_features=384, out_features=384, bias=False)
|
| 151 |
+
(k_proj): Linear(in_features=384, out_features=384, bias=False)
|
| 152 |
+
(v_proj): Linear(in_features=384, out_features=384, bias=False)
|
| 153 |
+
(proj): Linear(in_features=384, out_features=384, bias=False)
|
| 154 |
+
(rope): Rope()
|
| 155 |
+
)
|
| 156 |
+
)
|
| 157 |
+
(residual2): EfficientResidual(
|
| 158 |
+
(norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
|
| 159 |
+
(fn): MLP(
|
| 160 |
+
(fc1): Linear(in_features=384, out_features=1536, bias=False)
|
| 161 |
+
(act): GELU(approximate='none')
|
| 162 |
+
(fc2): Linear(in_features=1536, out_features=384, bias=False)
|
| 163 |
+
(drop): Dropout(p=0.0, inplace=False)
|
| 164 |
+
)
|
| 165 |
+
)
|
| 166 |
+
)
|
| 167 |
+
)
|
| 168 |
+
(predictor_norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
|
| 169 |
+
(predictor_proj): Linear(in_features=384, out_features=1024, bias=True)
|
| 170 |
+
)
|
| 171 |
+
)
|
| 172 |
+
[INFO ][2026-06-11 19:58:36][root ][init_video_model ] ViTMultiSeqWrapper(
|
| 173 |
+
(backbone): VisionTransformer(
|
| 174 |
+
(patch_embed): PatchEmbed3D(
|
| 175 |
+
(proj): Conv3d(3, 1024, kernel_size=(1, 14, 14), stride=(1, 14, 14))
|
| 176 |
+
)
|
| 177 |
+
(blocks): ModuleList(
|
| 178 |
+
(0-23): 24 x Block(
|
| 179 |
+
(norm1): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 180 |
+
(norm2): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 181 |
+
(drop_path1): Identity()
|
| 182 |
+
(drop_path2): Identity()
|
| 183 |
+
(attn): Attention(
|
| 184 |
+
(qkv): Linear(in_features=1024, out_features=3072, bias=True)
|
| 185 |
+
(attn_drop): Dropout(p=0.0, inplace=False)
|
| 186 |
+
(proj): Linear(in_features=1024, out_features=1024, bias=True)
|
| 187 |
+
(proj_drop): Dropout(p=0.0, inplace=False)
|
| 188 |
+
)
|
| 189 |
+
(mlp): MLP(
|
| 190 |
+
(fc1): Linear(in_features=1024, out_features=4096, bias=True)
|
| 191 |
+
(act): GELU(approximate='none')
|
| 192 |
+
(fc2): Linear(in_features=4096, out_features=1024, bias=True)
|
| 193 |
+
(drop): Dropout(p=0.0, inplace=False)
|
| 194 |
+
)
|
| 195 |
+
)
|
| 196 |
+
)
|
| 197 |
+
(norm): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 198 |
+
)
|
| 199 |
+
)
|
| 200 |
+
[INFO ][2026-06-11 19:58:36][root ][init_video_model ] Encoder number of parameters: 302964736
|
| 201 |
+
[INFO ][2026-06-11 19:58:36][root ][init_video_model ] Predictor number of parameters: 22042240
|
| 202 |
+
[INFO ][2026-06-11 19:58:36][root ][init_video_model ] Target encoder number of parameters: 302964736
|
| 203 |
+
[INFO ][2026-06-11 19:58:37][root ][make_videodataset ] VideoDataset dataset created
|
| 204 |
+
[INFO ][2026-06-11 19:58:37][WeightedSampler ][__init__ ] Using DistributedWeightedSampler with rank 5 / 6
|
| 205 |
+
[INFO ][2026-06-11 19:58:37][root ][make_videodataset ] VideoDataset unsupervised data loader created
|
| 206 |
+
[INFO ][2026-06-11 19:58:37][app.vjepa.train ][main ] iterations per epoch/dataset length: 300/2103
|
| 207 |
+
[INFO ][2026-06-11 19:58:37][app.vjepa.train ][main ] Wrapping models in DDP (rank 5)...
|
| 208 |
+
[INFO ][2026-06-11 19:58:38][root ][load_checkpoint ] Loading checkpoint from /var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/latest.pt
|
| 209 |
+
[INFO ][2026-06-11 19:58:48][root ][load_checkpoint ] loaded pretrained encoder from epoch 30 with msg: <All keys matched successfully>
|
| 210 |
+
[INFO ][2026-06-11 19:58:48][root ][load_checkpoint ] loaded pretrained predictor from epoch 30 with msg: <All keys matched successfully>
|
| 211 |
+
[INFO ][2026-06-11 19:58:48][root ][load_checkpoint ] loaded pretrained target encoder from epoch 30 with msg: <All keys matched successfully>
|
| 212 |
+
[INFO ][2026-06-11 19:58:48][root ][load_checkpoint ] loaded optimizers from epoch 30
|
| 213 |
+
[INFO ][2026-06-11 19:58:48][root ][load_checkpoint ] read-path: /var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/latest.pt
|
| 214 |
+
[INFO ][2026-06-11 20:00:08][app.vjepa.train ][main ] Initializing loader...
|
| 215 |
+
submitit INFO (2026-06-12 08:09:41,163) - Job completed successfully
|
| 216 |
+
[INFO ][2026-06-12 08:09:41][submitit ][process_job ] Job completed successfully
|
| 217 |
+
submitit INFO (2026-06-12 08:09:41,166) - Exiting after successful completion
|
| 218 |
+
[INFO ][2026-06-12 08:09:41][submitit ][process_job ] Exiting after successful completion
|
inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_26022/26022_0_log.err
ADDED
|
@@ -0,0 +1,18 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
/var/scratch/dcanezil/.cache/uv/virtualenvs/vd/lib/python3.12/site-packages/timm/models/layers/__init__.py:49: FutureWarning: Importing from timm.models.layers is deprecated, please import via timm.layers
|
| 2 |
+
warnings.warn(f"Importing from {__name__} is deprecated, please import via timm.layers", FutureWarning)
|
| 3 |
+
wandb: Currently logged in as: dgcnz (uvjepa) to https://api.wandb.ai. Use `wandb login --relogin` to force relogin
|
| 4 |
+
wandb: setting up run vfclercb
|
| 5 |
+
wandb: Tracking run with wandb version 0.23.1
|
| 6 |
+
wandb: Run data is saved locally in /var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/wandb/run-20260623_190228-vfclercb
|
| 7 |
+
wandb: Run `wandb offline` to turn off syncing.
|
| 8 |
+
wandb: Resuming run d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow
|
| 9 |
+
wandb: ⭐️ View project at https://wandb.ai/uvjepa/vjepa_ablation
|
| 10 |
+
wandb: 🚀 View run at https://wandb.ai/uvjepa/vjepa_ablation/runs/vfclercb
|
| 11 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/utils.py:888: FutureWarning: `torch.cuda.amp.GradScaler(args...)` is deprecated. Please use `torch.amp.GradScaler('cuda', args...)` instead.
|
| 12 |
+
scaler = torch.cuda.amp.GradScaler() if mixed_precision else None
|
| 13 |
+
srun: Job step aborted: Waiting up to 32 seconds for job step to finish.
|
| 14 |
+
slurmstepd: error: *** STEP 26022.0 ON node403 CANCELLED AT 2026-06-23T19:04:10 ***
|
| 15 |
+
slurmstepd: error: *** JOB 26022 ON node403 CANCELLED AT 2026-06-23T19:04:10 ***
|
| 16 |
+
submitit WARNING (2026-06-23 19:04:10,889) - Bypassing signal SIGTERM
|
| 17 |
+
submitit WARNING (2026-06-23 19:04:10,890) - Bypassing signal SIGCONT
|
| 18 |
+
slurmstepd: error: Failed to send MESSAGE_TASK_EXIT: Connection refused
|
inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_26022/26022_0_log.out
ADDED
|
@@ -0,0 +1,241 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
submitit INFO (2026-06-23 19:01:59,553) - Starting with JobEnvironment(job_id=26022, hostname=node403, local_rank=0(3), node=0(1), global_rank=0(3))
|
| 2 |
+
submitit INFO (2026-06-23 19:01:59,553) - Loading pickle: /var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_26022/26022_submitted.pkl
|
| 3 |
+
INFO:root:loaded pretrain params...
|
| 4 |
+
{ 'app': 'vjepa',
|
| 5 |
+
'cpus_per_task': 16,
|
| 6 |
+
'data': { 'batch_size': 64,
|
| 7 |
+
'crop_size': 224,
|
| 8 |
+
'dataset_fpcs': [8, 8],
|
| 9 |
+
'dataset_type': 'VideoDataset',
|
| 10 |
+
'datasets': [ '/local/dcanezil/data/kinetics/train.csv',
|
| 11 |
+
'/local/dcanezil/data/ssv2/train.csv'],
|
| 12 |
+
'datasets_weights': [0.65, 0.35],
|
| 13 |
+
'fps': 2,
|
| 14 |
+
'num_workers': 4,
|
| 15 |
+
'patch_size': 14,
|
| 16 |
+
'persistent_workers': True,
|
| 17 |
+
'pin_mem': True,
|
| 18 |
+
'tubelet_size': 1},
|
| 19 |
+
'data_aug': { 'auto_augment': False,
|
| 20 |
+
'motion_shift': False,
|
| 21 |
+
'random_resize_aspect_ratio': [0.75, 1.35],
|
| 22 |
+
'random_resize_scale': [0.3, 1.0],
|
| 23 |
+
'reprob': 0.0},
|
| 24 |
+
'folder': '/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow',
|
| 25 |
+
'loss': {'loss_exp': 1.0},
|
| 26 |
+
'mask': [ { 'aspect_ratio': [0.75, 1.5],
|
| 27 |
+
'full_complement': False,
|
| 28 |
+
'max_keep': None,
|
| 29 |
+
'max_temporal_keep': 1.0,
|
| 30 |
+
'num_blocks': 8,
|
| 31 |
+
'spatial_scale': [0.05, 0.05],
|
| 32 |
+
'temporal_scale': [1.0, 1.0]},
|
| 33 |
+
{ 'aspect_ratio': [0.75, 1.5],
|
| 34 |
+
'full_complement': False,
|
| 35 |
+
'max_keep': None,
|
| 36 |
+
'max_temporal_keep': 1.0,
|
| 37 |
+
'num_blocks': 2,
|
| 38 |
+
'spatial_scale': [0.4, 0.4],
|
| 39 |
+
'temporal_scale': [1.0, 1.0]}],
|
| 40 |
+
'mem_per_gpu': '40G',
|
| 41 |
+
'meta': { 'dtype': 'bfloat16',
|
| 42 |
+
'knn_eval_epoch0': True,
|
| 43 |
+
'knn_eval_freq': 5,
|
| 44 |
+
'knn_eval_presets': [ { 'config': { 'batch_size': 8,
|
| 45 |
+
'dataset_train': '/local/dcanezil/data/ssv2/train_coarse10.csv',
|
| 46 |
+
'dataset_val': '/local/dcanezil/data/ssv2/val_coarse10.csv',
|
| 47 |
+
'eval_videos_per_class': 100,
|
| 48 |
+
'linear_probe': True,
|
| 49 |
+
'num_workers': 2,
|
| 50 |
+
'pool_type': 'temporal_concat',
|
| 51 |
+
'train_videos_per_class': 500},
|
| 52 |
+
'preset': 'ssv2_coarse10'}],
|
| 53 |
+
'load_checkpoint': True,
|
| 54 |
+
'read_checkpoint': None,
|
| 55 |
+
'save_every_freq': 10,
|
| 56 |
+
'seed': 239,
|
| 57 |
+
'use_sdpa': True,
|
| 58 |
+
'use_wandb': True,
|
| 59 |
+
'wandb_project': 'vjepa_ablation'},
|
| 60 |
+
'metrics': {'sigreg': {}, 'std': {}},
|
| 61 |
+
'model': { 'class_token': True,
|
| 62 |
+
'freeze_backbone': False,
|
| 63 |
+
'is_causal': True,
|
| 64 |
+
'model_name': 'vit_large_patch14_dinov2_lvd142m',
|
| 65 |
+
'pe_type': '2d',
|
| 66 |
+
'pred_depth': 12,
|
| 67 |
+
'pred_embed_dim': 384,
|
| 68 |
+
'pred_num_heads': 6,
|
| 69 |
+
'predictor': 'v2_cross',
|
| 70 |
+
'stem_type': '2d',
|
| 71 |
+
'target_embed_dim': 1024,
|
| 72 |
+
'target_kind': 'ema',
|
| 73 |
+
'uniform_power': True,
|
| 74 |
+
'use_activation_checkpointing': True,
|
| 75 |
+
'use_alibi': False,
|
| 76 |
+
'use_lora': False,
|
| 77 |
+
'use_mask_tokens': True,
|
| 78 |
+
'use_rope': False,
|
| 79 |
+
'use_sdpa': True,
|
| 80 |
+
'zero_init_mask_tokens': True},
|
| 81 |
+
'nodes': 1,
|
| 82 |
+
'optimization': { 'clip_grad': None,
|
| 83 |
+
'ema': [1.0, 0.99925],
|
| 84 |
+
'ema_hold_epochs': 0,
|
| 85 |
+
'ema_ramp_epochs': 5,
|
| 86 |
+
'encoder_layer_decay': 0.9,
|
| 87 |
+
'epochs': 60,
|
| 88 |
+
'final_lr': 0.0001,
|
| 89 |
+
'final_weight_decay': 0.04,
|
| 90 |
+
'ipe': 300,
|
| 91 |
+
'ipe_scale': 1.0,
|
| 92 |
+
'lr': 0.0001,
|
| 93 |
+
'predictor_lr_scale': 3.0,
|
| 94 |
+
'start_lr': 1e-05,
|
| 95 |
+
'warmup': 2,
|
| 96 |
+
'weight_decay': 0.04},
|
| 97 |
+
'tasks_per_node': 3}
|
| 98 |
+
INFO:root:Running pre-training of app: vjepa
|
| 99 |
+
[INFO ][2026-06-23 19:02:23][app.vjepa.train ][main ] which_dtype='bfloat16'
|
| 100 |
+
[INFO ][2026-06-23 19:02:23][app.vjepa.train ][main ] Disabling persistent_workers (incompatible with KNN eval)
|
| 101 |
+
[INFO ][2026-06-23 19:02:23][app.vjepa.train ][main ] NCCL_SOCKET_IFNAME=None
|
| 102 |
+
[INFO ][2026-06-23 19:02:26][app.vjepa.train ][main ] Initialized (rank/world-size) 0/3, tasks_per_node=3
|
| 103 |
+
[INFO ][2026-06-23 19:02:27][app.vjepa.train ][main ] Resuming wandb run: vfclercb
|
| 104 |
+
[INFO ][2026-06-23 19:02:34][app.vjepa.train ][main ] Initialized wandb run: d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow (id=vfclercb)
|
| 105 |
+
[WARNING ][2026-06-23 19:03:08][experiments.stmodels.vision_transformers_v3][_build_dinov2_vit ] DINOv2 is an image model — fwd_chunk is unset so all frames will be processed jointly. Set fwd_chunk=1 for per-frame (standard) processing.
|
| 106 |
+
[INFO ][2026-06-23 19:03:19][experiments.stmodels.vision_transformers_v3][_build_dinov2_vit ] Loading pretrained weights for vit_large_patch14_dinov2_lvd142m from timm (vit_large_patch14_dinov2)
|
| 107 |
+
[INFO ][2026-06-23 19:03:21][timm.models._builder][load_pretrained ] Loading pretrained weights from Hugging Face hub (timm/vit_large_patch14_dinov2.lvd142m)
|
| 108 |
+
[INFO ][2026-06-23 19:03:22][timm.models._hub ][load_state_dict_from_hf ] [timm/vit_large_patch14_dinov2.lvd142m] Safe alternative available for 'pytorch_model.bin' (as 'model.safetensors'). Loading weights using safetensors.
|
| 109 |
+
[INFO ][2026-06-23 19:03:24][timm.layers.pos_embed][resample_abs_pos_embed ] Resized position embedding: (37, 37) to (16, 16).
|
| 110 |
+
[INFO ][2026-06-23 19:03:28][root ][init_video_model ] ViTMultiSeqWrapper(
|
| 111 |
+
(backbone): VisionTransformer(
|
| 112 |
+
(patch_embed): PatchEmbed3D(
|
| 113 |
+
(proj): Conv3d(3, 1024, kernel_size=(1, 14, 14), stride=(1, 14, 14))
|
| 114 |
+
)
|
| 115 |
+
(blocks): ModuleList(
|
| 116 |
+
(0-23): 24 x Block(
|
| 117 |
+
(norm1): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 118 |
+
(norm2): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 119 |
+
(drop_path1): Identity()
|
| 120 |
+
(drop_path2): Identity()
|
| 121 |
+
(attn): Attention(
|
| 122 |
+
(qkv): Linear(in_features=1024, out_features=3072, bias=True)
|
| 123 |
+
(attn_drop): Dropout(p=0.0, inplace=False)
|
| 124 |
+
(proj): Linear(in_features=1024, out_features=1024, bias=True)
|
| 125 |
+
(proj_drop): Dropout(p=0.0, inplace=False)
|
| 126 |
+
)
|
| 127 |
+
(mlp): MLP(
|
| 128 |
+
(fc1): Linear(in_features=1024, out_features=4096, bias=True)
|
| 129 |
+
(act): GELU(approximate='none')
|
| 130 |
+
(fc2): Linear(in_features=4096, out_features=1024, bias=True)
|
| 131 |
+
(drop): Dropout(p=0.0, inplace=False)
|
| 132 |
+
)
|
| 133 |
+
)
|
| 134 |
+
)
|
| 135 |
+
(norm): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 136 |
+
)
|
| 137 |
+
)
|
| 138 |
+
[INFO ][2026-06-23 19:03:28][root ][init_video_model ] VJEPAPredictorMultiSeqWrapper(
|
| 139 |
+
(backbone): PredictorV2(
|
| 140 |
+
(predictor_embed): Linear(in_features=1024, out_features=384, bias=True)
|
| 141 |
+
(mask_tokens): ParameterList(
|
| 142 |
+
(0): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
|
| 143 |
+
(1): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
|
| 144 |
+
(2): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
|
| 145 |
+
(3): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
|
| 146 |
+
)
|
| 147 |
+
(predictor_blocks): ModuleList(
|
| 148 |
+
(0-11): 12 x Block(
|
| 149 |
+
(residual1): EfficientResidual(
|
| 150 |
+
(norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
|
| 151 |
+
(fn): Attention(
|
| 152 |
+
(q_proj): Linear(in_features=384, out_features=384, bias=False)
|
| 153 |
+
(k_proj): Linear(in_features=384, out_features=384, bias=False)
|
| 154 |
+
(v_proj): Linear(in_features=384, out_features=384, bias=False)
|
| 155 |
+
(proj): Linear(in_features=384, out_features=384, bias=False)
|
| 156 |
+
(rope): Rope()
|
| 157 |
+
)
|
| 158 |
+
)
|
| 159 |
+
(residual2): EfficientResidual(
|
| 160 |
+
(norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
|
| 161 |
+
(fn): MLP(
|
| 162 |
+
(fc1): Linear(in_features=384, out_features=1536, bias=False)
|
| 163 |
+
(act): GELU(approximate='none')
|
| 164 |
+
(fc2): Linear(in_features=1536, out_features=384, bias=False)
|
| 165 |
+
(drop): Dropout(p=0.0, inplace=False)
|
| 166 |
+
)
|
| 167 |
+
)
|
| 168 |
+
)
|
| 169 |
+
)
|
| 170 |
+
(predictor_norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
|
| 171 |
+
(predictor_proj): Linear(in_features=384, out_features=1024, bias=True)
|
| 172 |
+
)
|
| 173 |
+
)
|
| 174 |
+
[INFO ][2026-06-23 19:03:28][root ][init_video_model ] ViTMultiSeqWrapper(
|
| 175 |
+
(backbone): VisionTransformer(
|
| 176 |
+
(patch_embed): PatchEmbed3D(
|
| 177 |
+
(proj): Conv3d(3, 1024, kernel_size=(1, 14, 14), stride=(1, 14, 14))
|
| 178 |
+
)
|
| 179 |
+
(blocks): ModuleList(
|
| 180 |
+
(0-23): 24 x Block(
|
| 181 |
+
(norm1): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 182 |
+
(norm2): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 183 |
+
(drop_path1): Identity()
|
| 184 |
+
(drop_path2): Identity()
|
| 185 |
+
(attn): Attention(
|
| 186 |
+
(qkv): Linear(in_features=1024, out_features=3072, bias=True)
|
| 187 |
+
(attn_drop): Dropout(p=0.0, inplace=False)
|
| 188 |
+
(proj): Linear(in_features=1024, out_features=1024, bias=True)
|
| 189 |
+
(proj_drop): Dropout(p=0.0, inplace=False)
|
| 190 |
+
)
|
| 191 |
+
(mlp): MLP(
|
| 192 |
+
(fc1): Linear(in_features=1024, out_features=4096, bias=True)
|
| 193 |
+
(act): GELU(approximate='none')
|
| 194 |
+
(fc2): Linear(in_features=4096, out_features=1024, bias=True)
|
| 195 |
+
(drop): Dropout(p=0.0, inplace=False)
|
| 196 |
+
)
|
| 197 |
+
)
|
| 198 |
+
)
|
| 199 |
+
(norm): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 200 |
+
)
|
| 201 |
+
)
|
| 202 |
+
[INFO ][2026-06-23 19:03:28][root ][init_video_model ] Encoder number of parameters: 302964736
|
| 203 |
+
[INFO ][2026-06-23 19:03:28][root ][init_video_model ] Predictor number of parameters: 22042240
|
| 204 |
+
[INFO ][2026-06-23 19:03:28][root ][init_video_model ] Target encoder number of parameters: 302964736
|
| 205 |
+
[INFO ][2026-06-23 19:03:29][root ][make_videodataset ] VideoDataset dataset created
|
| 206 |
+
[INFO ][2026-06-23 19:03:29][WeightedSampler ][__init__ ] Using DistributedWeightedSampler with rank 0 / 3
|
| 207 |
+
[INFO ][2026-06-23 19:03:29][root ][make_videodataset ] VideoDataset unsupervised data loader created
|
| 208 |
+
[INFO ][2026-06-23 19:03:29][app.vjepa.train ][main ] iterations per epoch/dataset length: 300/2103
|
| 209 |
+
[INFO ][2026-06-23 19:03:29][app.vjepa.train ][main ] Wrapping models in DDP (rank 0)...
|
| 210 |
+
[INFO ][2026-06-23 19:03:31][root ][load_checkpoint ] Loading checkpoint from /var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/latest.pt
|
| 211 |
+
[INFO ][2026-06-23 19:03:41][root ][load_checkpoint ] loaded pretrained encoder from epoch 60 with msg: <All keys matched successfully>
|
| 212 |
+
[INFO ][2026-06-23 19:03:41][root ][load_checkpoint ] loaded pretrained predictor from epoch 60 with msg: <All keys matched successfully>
|
| 213 |
+
[INFO ][2026-06-23 19:03:42][root ][load_checkpoint ] loaded pretrained target encoder from epoch 60 with msg: <All keys matched successfully>
|
| 214 |
+
[INFO ][2026-06-23 19:03:42][root ][load_checkpoint ] loaded optimizers from epoch 60
|
| 215 |
+
[INFO ][2026-06-23 19:03:42][root ][load_checkpoint ] read-path: /var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/latest.pt
|
| 216 |
+
[INFO ][2026-06-23 19:03:42][app.vjepa.train ][main ] Running initial KNN evaluation (epoch 0)...
|
| 217 |
+
[INFO ][2026-06-23 19:03:43][src.utils.knn_eval ][run_knn_eval ] [KNN Eval] Starting ssv2_coarse10 KNN evaluation (frames_per_clip=8)...
|
| 218 |
+
[INFO ][2026-06-23 19:03:43][src.utils.knn_eval ][run_knn_eval ] [KNN Eval] Pool type: temporal_concat, CLS token: False
|
| 219 |
+
[INFO ][2026-06-23 19:03:43][src.utils.knn_eval ][run_knn_eval ] [KNN Eval] Using fps=2 (dynamic frame_step per video)
|
| 220 |
+
[INFO ][2026-06-23 19:03:43][src.utils.knn_eval ][run_knn_eval ] [KNN Eval] Presampled train subset with seed=0: 5000 videos across 10 classes
|
| 221 |
+
[INFO ][2026-06-23 19:03:43][src.utils.knn_eval ][run_knn_eval ] [KNN Eval] Presampled val subset with seed=0: 996 videos across 10 classes
|
| 222 |
+
[INFO ][2026-06-23 19:03:43][root ][make_videodataset ] VideoDataset dataset created
|
| 223 |
+
[INFO ][2026-06-23 19:03:43][root ][make_videodataset ] VideoDataset unsupervised data loader created
|
| 224 |
+
[INFO ][2026-06-23 19:03:43][root ][make_videodataset ] VideoDataset dataset created
|
| 225 |
+
[INFO ][2026-06-23 19:03:43][root ][make_videodataset ] VideoDataset unsupervised data loader created
|
| 226 |
+
[INFO ][2026-06-23 19:03:43][src.utils.knn_eval ][run_knn_eval ] [KNN Eval] Train loader: 209 batches, Val loader: 42 batches
|
| 227 |
+
[INFO ][2026-06-23 19:03:43][src.utils.knn_eval ][run_knn_eval ] [KNN Eval] Extracting training features...
|
| 228 |
+
[INFO ][2026-06-23 19:03:56][src.utils.knn_eval ][_extract_features ] [KNN Eval] Extracted 0/209 batches
|
| 229 |
+
[INFO ][2026-06-23 19:04:00][src.utils.knn_eval ][_extract_features ] [KNN Eval] Extracted 20/209 batches
|
| 230 |
+
[INFO ][2026-06-23 19:04:05][src.utils.knn_eval ][_extract_features ] [KNN Eval] Extracted 40/209 batches
|
| 231 |
+
[INFO ][2026-06-23 19:04:10][src.utils.knn_eval ][_extract_features ] [KNN Eval] Extracted 60/209 batches
|
| 232 |
+
submitit WARNING (2026-06-23 19:04:10,889) - Bypassing signal SIGTERM
|
| 233 |
+
[WARNING ][2026-06-23 19:04:10][submitit ][bypass ] Bypassing signal SIGTERM
|
| 234 |
+
submitit WARNING (2026-06-23 19:04:10,890) - Bypassing signal SIGCONT
|
| 235 |
+
[WARNING ][2026-06-23 19:04:10][submitit ][bypass ] Bypassing signal SIGCONT
|
| 236 |
+
[INFO ][2026-06-23 19:04:14][src.utils.knn_eval ][_extract_features ] [KNN Eval] Extracted 80/209 batches
|
| 237 |
+
[INFO ][2026-06-23 19:04:19][src.utils.knn_eval ][_extract_features ] [KNN Eval] Extracted 100/209 batches
|
| 238 |
+
[INFO ][2026-06-23 19:04:24][src.utils.knn_eval ][_extract_features ] [KNN Eval] Extracted 120/209 batches
|
| 239 |
+
[INFO ][2026-06-23 19:04:28][src.utils.knn_eval ][_extract_features ] [KNN Eval] Extracted 140/209 batches
|
| 240 |
+
[INFO ][2026-06-23 19:04:33][src.utils.knn_eval ][_extract_features ] [KNN Eval] Extracted 160/209 batches
|
| 241 |
+
[INFO ][2026-06-23 19:04:37][src.utils.knn_eval ][_extract_features ] [KNN Eval] Extracted 180/209 batches
|
inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_26022/26022_1_log.err
ADDED
|
@@ -0,0 +1,6 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
/var/scratch/dcanezil/.cache/uv/virtualenvs/vd/lib/python3.12/site-packages/timm/models/layers/__init__.py:49: FutureWarning: Importing from timm.models.layers is deprecated, please import via timm.layers
|
| 2 |
+
warnings.warn(f"Importing from {__name__} is deprecated, please import via timm.layers", FutureWarning)
|
| 3 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/utils.py:888: FutureWarning: `torch.cuda.amp.GradScaler(args...)` is deprecated. Please use `torch.amp.GradScaler('cuda', args...)` instead.
|
| 4 |
+
scaler = torch.cuda.amp.GradScaler() if mixed_precision else None
|
| 5 |
+
submitit WARNING (2026-06-23 19:04:10,923) - Bypassing signal SIGTERM
|
| 6 |
+
submitit WARNING (2026-06-23 19:04:10,924) - Bypassing signal SIGCONT
|
inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_26022/26022_1_log.out
ADDED
|
@@ -0,0 +1,217 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
submitit INFO (2026-06-23 19:01:59,553) - Starting with JobEnvironment(job_id=26022, hostname=node403, local_rank=1(3), node=0(1), global_rank=1(3))
|
| 2 |
+
submitit INFO (2026-06-23 19:01:59,553) - Loading pickle: /var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_26022/26022_submitted.pkl
|
| 3 |
+
INFO:root:loaded pretrain params...
|
| 4 |
+
{ 'app': 'vjepa',
|
| 5 |
+
'cpus_per_task': 16,
|
| 6 |
+
'data': { 'batch_size': 64,
|
| 7 |
+
'crop_size': 224,
|
| 8 |
+
'dataset_fpcs': [8, 8],
|
| 9 |
+
'dataset_type': 'VideoDataset',
|
| 10 |
+
'datasets': [ '/local/dcanezil/data/kinetics/train.csv',
|
| 11 |
+
'/local/dcanezil/data/ssv2/train.csv'],
|
| 12 |
+
'datasets_weights': [0.65, 0.35],
|
| 13 |
+
'fps': 2,
|
| 14 |
+
'num_workers': 4,
|
| 15 |
+
'patch_size': 14,
|
| 16 |
+
'persistent_workers': True,
|
| 17 |
+
'pin_mem': True,
|
| 18 |
+
'tubelet_size': 1},
|
| 19 |
+
'data_aug': { 'auto_augment': False,
|
| 20 |
+
'motion_shift': False,
|
| 21 |
+
'random_resize_aspect_ratio': [0.75, 1.35],
|
| 22 |
+
'random_resize_scale': [0.3, 1.0],
|
| 23 |
+
'reprob': 0.0},
|
| 24 |
+
'folder': '/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow',
|
| 25 |
+
'loss': {'loss_exp': 1.0},
|
| 26 |
+
'mask': [ { 'aspect_ratio': [0.75, 1.5],
|
| 27 |
+
'full_complement': False,
|
| 28 |
+
'max_keep': None,
|
| 29 |
+
'max_temporal_keep': 1.0,
|
| 30 |
+
'num_blocks': 8,
|
| 31 |
+
'spatial_scale': [0.05, 0.05],
|
| 32 |
+
'temporal_scale': [1.0, 1.0]},
|
| 33 |
+
{ 'aspect_ratio': [0.75, 1.5],
|
| 34 |
+
'full_complement': False,
|
| 35 |
+
'max_keep': None,
|
| 36 |
+
'max_temporal_keep': 1.0,
|
| 37 |
+
'num_blocks': 2,
|
| 38 |
+
'spatial_scale': [0.4, 0.4],
|
| 39 |
+
'temporal_scale': [1.0, 1.0]}],
|
| 40 |
+
'mem_per_gpu': '40G',
|
| 41 |
+
'meta': { 'dtype': 'bfloat16',
|
| 42 |
+
'knn_eval_epoch0': True,
|
| 43 |
+
'knn_eval_freq': 5,
|
| 44 |
+
'knn_eval_presets': [ { 'config': { 'batch_size': 8,
|
| 45 |
+
'dataset_train': '/local/dcanezil/data/ssv2/train_coarse10.csv',
|
| 46 |
+
'dataset_val': '/local/dcanezil/data/ssv2/val_coarse10.csv',
|
| 47 |
+
'eval_videos_per_class': 100,
|
| 48 |
+
'linear_probe': True,
|
| 49 |
+
'num_workers': 2,
|
| 50 |
+
'pool_type': 'temporal_concat',
|
| 51 |
+
'train_videos_per_class': 500},
|
| 52 |
+
'preset': 'ssv2_coarse10'}],
|
| 53 |
+
'load_checkpoint': True,
|
| 54 |
+
'read_checkpoint': None,
|
| 55 |
+
'save_every_freq': 10,
|
| 56 |
+
'seed': 239,
|
| 57 |
+
'use_sdpa': True,
|
| 58 |
+
'use_wandb': True,
|
| 59 |
+
'wandb_project': 'vjepa_ablation'},
|
| 60 |
+
'metrics': {'sigreg': {}, 'std': {}},
|
| 61 |
+
'model': { 'class_token': True,
|
| 62 |
+
'freeze_backbone': False,
|
| 63 |
+
'is_causal': True,
|
| 64 |
+
'model_name': 'vit_large_patch14_dinov2_lvd142m',
|
| 65 |
+
'pe_type': '2d',
|
| 66 |
+
'pred_depth': 12,
|
| 67 |
+
'pred_embed_dim': 384,
|
| 68 |
+
'pred_num_heads': 6,
|
| 69 |
+
'predictor': 'v2_cross',
|
| 70 |
+
'stem_type': '2d',
|
| 71 |
+
'target_embed_dim': 1024,
|
| 72 |
+
'target_kind': 'ema',
|
| 73 |
+
'uniform_power': True,
|
| 74 |
+
'use_activation_checkpointing': True,
|
| 75 |
+
'use_alibi': False,
|
| 76 |
+
'use_lora': False,
|
| 77 |
+
'use_mask_tokens': True,
|
| 78 |
+
'use_rope': False,
|
| 79 |
+
'use_sdpa': True,
|
| 80 |
+
'zero_init_mask_tokens': True},
|
| 81 |
+
'nodes': 1,
|
| 82 |
+
'optimization': { 'clip_grad': None,
|
| 83 |
+
'ema': [1.0, 0.99925],
|
| 84 |
+
'ema_hold_epochs': 0,
|
| 85 |
+
'ema_ramp_epochs': 5,
|
| 86 |
+
'encoder_layer_decay': 0.9,
|
| 87 |
+
'epochs': 60,
|
| 88 |
+
'final_lr': 0.0001,
|
| 89 |
+
'final_weight_decay': 0.04,
|
| 90 |
+
'ipe': 300,
|
| 91 |
+
'ipe_scale': 1.0,
|
| 92 |
+
'lr': 0.0001,
|
| 93 |
+
'predictor_lr_scale': 3.0,
|
| 94 |
+
'start_lr': 1e-05,
|
| 95 |
+
'warmup': 2,
|
| 96 |
+
'weight_decay': 0.04},
|
| 97 |
+
'tasks_per_node': 3}
|
| 98 |
+
INFO:root:Running pre-training of app: vjepa
|
| 99 |
+
[INFO ][2026-06-23 19:02:23][app.vjepa.train ][main ] which_dtype='bfloat16'
|
| 100 |
+
[INFO ][2026-06-23 19:02:23][app.vjepa.train ][main ] Disabling persistent_workers (incompatible with KNN eval)
|
| 101 |
+
[INFO ][2026-06-23 19:02:23][app.vjepa.train ][main ] NCCL_SOCKET_IFNAME=None
|
| 102 |
+
[INFO ][2026-06-23 19:02:26][app.vjepa.train ][main ] Initialized (rank/world-size) 1/3, tasks_per_node=3
|
| 103 |
+
[WARNING ][2026-06-23 19:03:08][experiments.stmodels.vision_transformers_v3][_build_dinov2_vit ] DINOv2 is an image model — fwd_chunk is unset so all frames will be processed jointly. Set fwd_chunk=1 for per-frame (standard) processing.
|
| 104 |
+
[INFO ][2026-06-23 19:03:19][experiments.stmodels.vision_transformers_v3][_build_dinov2_vit ] Loading pretrained weights for vit_large_patch14_dinov2_lvd142m from timm (vit_large_patch14_dinov2)
|
| 105 |
+
[INFO ][2026-06-23 19:03:22][timm.models._builder][load_pretrained ] Loading pretrained weights from Hugging Face hub (timm/vit_large_patch14_dinov2.lvd142m)
|
| 106 |
+
[INFO ][2026-06-23 19:03:22][timm.models._hub ][load_state_dict_from_hf ] [timm/vit_large_patch14_dinov2.lvd142m] Safe alternative available for 'pytorch_model.bin' (as 'model.safetensors'). Loading weights using safetensors.
|
| 107 |
+
[INFO ][2026-06-23 19:03:24][timm.layers.pos_embed][resample_abs_pos_embed ] Resized position embedding: (37, 37) to (16, 16).
|
| 108 |
+
[INFO ][2026-06-23 19:03:28][root ][init_video_model ] ViTMultiSeqWrapper(
|
| 109 |
+
(backbone): VisionTransformer(
|
| 110 |
+
(patch_embed): PatchEmbed3D(
|
| 111 |
+
(proj): Conv3d(3, 1024, kernel_size=(1, 14, 14), stride=(1, 14, 14))
|
| 112 |
+
)
|
| 113 |
+
(blocks): ModuleList(
|
| 114 |
+
(0-23): 24 x Block(
|
| 115 |
+
(norm1): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 116 |
+
(norm2): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 117 |
+
(drop_path1): Identity()
|
| 118 |
+
(drop_path2): Identity()
|
| 119 |
+
(attn): Attention(
|
| 120 |
+
(qkv): Linear(in_features=1024, out_features=3072, bias=True)
|
| 121 |
+
(attn_drop): Dropout(p=0.0, inplace=False)
|
| 122 |
+
(proj): Linear(in_features=1024, out_features=1024, bias=True)
|
| 123 |
+
(proj_drop): Dropout(p=0.0, inplace=False)
|
| 124 |
+
)
|
| 125 |
+
(mlp): MLP(
|
| 126 |
+
(fc1): Linear(in_features=1024, out_features=4096, bias=True)
|
| 127 |
+
(act): GELU(approximate='none')
|
| 128 |
+
(fc2): Linear(in_features=4096, out_features=1024, bias=True)
|
| 129 |
+
(drop): Dropout(p=0.0, inplace=False)
|
| 130 |
+
)
|
| 131 |
+
)
|
| 132 |
+
)
|
| 133 |
+
(norm): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 134 |
+
)
|
| 135 |
+
)
|
| 136 |
+
[INFO ][2026-06-23 19:03:28][root ][init_video_model ] VJEPAPredictorMultiSeqWrapper(
|
| 137 |
+
(backbone): PredictorV2(
|
| 138 |
+
(predictor_embed): Linear(in_features=1024, out_features=384, bias=True)
|
| 139 |
+
(mask_tokens): ParameterList(
|
| 140 |
+
(0): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
|
| 141 |
+
(1): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
|
| 142 |
+
(2): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
|
| 143 |
+
(3): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
|
| 144 |
+
)
|
| 145 |
+
(predictor_blocks): ModuleList(
|
| 146 |
+
(0-11): 12 x Block(
|
| 147 |
+
(residual1): EfficientResidual(
|
| 148 |
+
(norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
|
| 149 |
+
(fn): Attention(
|
| 150 |
+
(q_proj): Linear(in_features=384, out_features=384, bias=False)
|
| 151 |
+
(k_proj): Linear(in_features=384, out_features=384, bias=False)
|
| 152 |
+
(v_proj): Linear(in_features=384, out_features=384, bias=False)
|
| 153 |
+
(proj): Linear(in_features=384, out_features=384, bias=False)
|
| 154 |
+
(rope): Rope()
|
| 155 |
+
)
|
| 156 |
+
)
|
| 157 |
+
(residual2): EfficientResidual(
|
| 158 |
+
(norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
|
| 159 |
+
(fn): MLP(
|
| 160 |
+
(fc1): Linear(in_features=384, out_features=1536, bias=False)
|
| 161 |
+
(act): GELU(approximate='none')
|
| 162 |
+
(fc2): Linear(in_features=1536, out_features=384, bias=False)
|
| 163 |
+
(drop): Dropout(p=0.0, inplace=False)
|
| 164 |
+
)
|
| 165 |
+
)
|
| 166 |
+
)
|
| 167 |
+
)
|
| 168 |
+
(predictor_norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
|
| 169 |
+
(predictor_proj): Linear(in_features=384, out_features=1024, bias=True)
|
| 170 |
+
)
|
| 171 |
+
)
|
| 172 |
+
[INFO ][2026-06-23 19:03:28][root ][init_video_model ] ViTMultiSeqWrapper(
|
| 173 |
+
(backbone): VisionTransformer(
|
| 174 |
+
(patch_embed): PatchEmbed3D(
|
| 175 |
+
(proj): Conv3d(3, 1024, kernel_size=(1, 14, 14), stride=(1, 14, 14))
|
| 176 |
+
)
|
| 177 |
+
(blocks): ModuleList(
|
| 178 |
+
(0-23): 24 x Block(
|
| 179 |
+
(norm1): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 180 |
+
(norm2): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 181 |
+
(drop_path1): Identity()
|
| 182 |
+
(drop_path2): Identity()
|
| 183 |
+
(attn): Attention(
|
| 184 |
+
(qkv): Linear(in_features=1024, out_features=3072, bias=True)
|
| 185 |
+
(attn_drop): Dropout(p=0.0, inplace=False)
|
| 186 |
+
(proj): Linear(in_features=1024, out_features=1024, bias=True)
|
| 187 |
+
(proj_drop): Dropout(p=0.0, inplace=False)
|
| 188 |
+
)
|
| 189 |
+
(mlp): MLP(
|
| 190 |
+
(fc1): Linear(in_features=1024, out_features=4096, bias=True)
|
| 191 |
+
(act): GELU(approximate='none')
|
| 192 |
+
(fc2): Linear(in_features=4096, out_features=1024, bias=True)
|
| 193 |
+
(drop): Dropout(p=0.0, inplace=False)
|
| 194 |
+
)
|
| 195 |
+
)
|
| 196 |
+
)
|
| 197 |
+
(norm): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 198 |
+
)
|
| 199 |
+
)
|
| 200 |
+
[INFO ][2026-06-23 19:03:28][root ][init_video_model ] Encoder number of parameters: 302964736
|
| 201 |
+
[INFO ][2026-06-23 19:03:28][root ][init_video_model ] Predictor number of parameters: 22042240
|
| 202 |
+
[INFO ][2026-06-23 19:03:28][root ][init_video_model ] Target encoder number of parameters: 302964736
|
| 203 |
+
[INFO ][2026-06-23 19:03:29][root ][make_videodataset ] VideoDataset dataset created
|
| 204 |
+
[INFO ][2026-06-23 19:03:29][WeightedSampler ][__init__ ] Using DistributedWeightedSampler with rank 1 / 3
|
| 205 |
+
[INFO ][2026-06-23 19:03:29][root ][make_videodataset ] VideoDataset unsupervised data loader created
|
| 206 |
+
[INFO ][2026-06-23 19:03:29][app.vjepa.train ][main ] iterations per epoch/dataset length: 300/2103
|
| 207 |
+
[INFO ][2026-06-23 19:03:29][app.vjepa.train ][main ] Wrapping models in DDP (rank 1)...
|
| 208 |
+
[INFO ][2026-06-23 19:03:31][root ][load_checkpoint ] Loading checkpoint from /var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/latest.pt
|
| 209 |
+
[INFO ][2026-06-23 19:03:41][root ][load_checkpoint ] loaded pretrained encoder from epoch 60 with msg: <All keys matched successfully>
|
| 210 |
+
[INFO ][2026-06-23 19:03:41][root ][load_checkpoint ] loaded pretrained predictor from epoch 60 with msg: <All keys matched successfully>
|
| 211 |
+
[INFO ][2026-06-23 19:03:42][root ][load_checkpoint ] loaded pretrained target encoder from epoch 60 with msg: <All keys matched successfully>
|
| 212 |
+
[INFO ][2026-06-23 19:03:42][root ][load_checkpoint ] loaded optimizers from epoch 60
|
| 213 |
+
[INFO ][2026-06-23 19:03:42][root ][load_checkpoint ] read-path: /var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/latest.pt
|
| 214 |
+
submitit WARNING (2026-06-23 19:04:10,923) - Bypassing signal SIGTERM
|
| 215 |
+
[WARNING ][2026-06-23 19:04:10][submitit ][bypass ] Bypassing signal SIGTERM
|
| 216 |
+
submitit WARNING (2026-06-23 19:04:10,924) - Bypassing signal SIGCONT
|
| 217 |
+
[WARNING ][2026-06-23 19:04:10][submitit ][bypass ] Bypassing signal SIGCONT
|
inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_26022/26022_2_log.err
ADDED
|
@@ -0,0 +1,6 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
/var/scratch/dcanezil/.cache/uv/virtualenvs/vd/lib/python3.12/site-packages/timm/models/layers/__init__.py:49: FutureWarning: Importing from timm.models.layers is deprecated, please import via timm.layers
|
| 2 |
+
warnings.warn(f"Importing from {__name__} is deprecated, please import via timm.layers", FutureWarning)
|
| 3 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/utils.py:888: FutureWarning: `torch.cuda.amp.GradScaler(args...)` is deprecated. Please use `torch.amp.GradScaler('cuda', args...)` instead.
|
| 4 |
+
scaler = torch.cuda.amp.GradScaler() if mixed_precision else None
|
| 5 |
+
submitit WARNING (2026-06-23 19:04:10,794) - Bypassing signal SIGTERM
|
| 6 |
+
submitit WARNING (2026-06-23 19:04:10,794) - Bypassing signal SIGCONT
|
inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_26022/26022_2_log.out
ADDED
|
@@ -0,0 +1,217 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
submitit INFO (2026-06-23 19:01:59,553) - Starting with JobEnvironment(job_id=26022, hostname=node403, local_rank=2(3), node=0(1), global_rank=2(3))
|
| 2 |
+
submitit INFO (2026-06-23 19:01:59,553) - Loading pickle: /var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_26022/26022_submitted.pkl
|
| 3 |
+
INFO:root:loaded pretrain params...
|
| 4 |
+
{ 'app': 'vjepa',
|
| 5 |
+
'cpus_per_task': 16,
|
| 6 |
+
'data': { 'batch_size': 64,
|
| 7 |
+
'crop_size': 224,
|
| 8 |
+
'dataset_fpcs': [8, 8],
|
| 9 |
+
'dataset_type': 'VideoDataset',
|
| 10 |
+
'datasets': [ '/local/dcanezil/data/kinetics/train.csv',
|
| 11 |
+
'/local/dcanezil/data/ssv2/train.csv'],
|
| 12 |
+
'datasets_weights': [0.65, 0.35],
|
| 13 |
+
'fps': 2,
|
| 14 |
+
'num_workers': 4,
|
| 15 |
+
'patch_size': 14,
|
| 16 |
+
'persistent_workers': True,
|
| 17 |
+
'pin_mem': True,
|
| 18 |
+
'tubelet_size': 1},
|
| 19 |
+
'data_aug': { 'auto_augment': False,
|
| 20 |
+
'motion_shift': False,
|
| 21 |
+
'random_resize_aspect_ratio': [0.75, 1.35],
|
| 22 |
+
'random_resize_scale': [0.3, 1.0],
|
| 23 |
+
'reprob': 0.0},
|
| 24 |
+
'folder': '/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow',
|
| 25 |
+
'loss': {'loss_exp': 1.0},
|
| 26 |
+
'mask': [ { 'aspect_ratio': [0.75, 1.5],
|
| 27 |
+
'full_complement': False,
|
| 28 |
+
'max_keep': None,
|
| 29 |
+
'max_temporal_keep': 1.0,
|
| 30 |
+
'num_blocks': 8,
|
| 31 |
+
'spatial_scale': [0.05, 0.05],
|
| 32 |
+
'temporal_scale': [1.0, 1.0]},
|
| 33 |
+
{ 'aspect_ratio': [0.75, 1.5],
|
| 34 |
+
'full_complement': False,
|
| 35 |
+
'max_keep': None,
|
| 36 |
+
'max_temporal_keep': 1.0,
|
| 37 |
+
'num_blocks': 2,
|
| 38 |
+
'spatial_scale': [0.4, 0.4],
|
| 39 |
+
'temporal_scale': [1.0, 1.0]}],
|
| 40 |
+
'mem_per_gpu': '40G',
|
| 41 |
+
'meta': { 'dtype': 'bfloat16',
|
| 42 |
+
'knn_eval_epoch0': True,
|
| 43 |
+
'knn_eval_freq': 5,
|
| 44 |
+
'knn_eval_presets': [ { 'config': { 'batch_size': 8,
|
| 45 |
+
'dataset_train': '/local/dcanezil/data/ssv2/train_coarse10.csv',
|
| 46 |
+
'dataset_val': '/local/dcanezil/data/ssv2/val_coarse10.csv',
|
| 47 |
+
'eval_videos_per_class': 100,
|
| 48 |
+
'linear_probe': True,
|
| 49 |
+
'num_workers': 2,
|
| 50 |
+
'pool_type': 'temporal_concat',
|
| 51 |
+
'train_videos_per_class': 500},
|
| 52 |
+
'preset': 'ssv2_coarse10'}],
|
| 53 |
+
'load_checkpoint': True,
|
| 54 |
+
'read_checkpoint': None,
|
| 55 |
+
'save_every_freq': 10,
|
| 56 |
+
'seed': 239,
|
| 57 |
+
'use_sdpa': True,
|
| 58 |
+
'use_wandb': True,
|
| 59 |
+
'wandb_project': 'vjepa_ablation'},
|
| 60 |
+
'metrics': {'sigreg': {}, 'std': {}},
|
| 61 |
+
'model': { 'class_token': True,
|
| 62 |
+
'freeze_backbone': False,
|
| 63 |
+
'is_causal': True,
|
| 64 |
+
'model_name': 'vit_large_patch14_dinov2_lvd142m',
|
| 65 |
+
'pe_type': '2d',
|
| 66 |
+
'pred_depth': 12,
|
| 67 |
+
'pred_embed_dim': 384,
|
| 68 |
+
'pred_num_heads': 6,
|
| 69 |
+
'predictor': 'v2_cross',
|
| 70 |
+
'stem_type': '2d',
|
| 71 |
+
'target_embed_dim': 1024,
|
| 72 |
+
'target_kind': 'ema',
|
| 73 |
+
'uniform_power': True,
|
| 74 |
+
'use_activation_checkpointing': True,
|
| 75 |
+
'use_alibi': False,
|
| 76 |
+
'use_lora': False,
|
| 77 |
+
'use_mask_tokens': True,
|
| 78 |
+
'use_rope': False,
|
| 79 |
+
'use_sdpa': True,
|
| 80 |
+
'zero_init_mask_tokens': True},
|
| 81 |
+
'nodes': 1,
|
| 82 |
+
'optimization': { 'clip_grad': None,
|
| 83 |
+
'ema': [1.0, 0.99925],
|
| 84 |
+
'ema_hold_epochs': 0,
|
| 85 |
+
'ema_ramp_epochs': 5,
|
| 86 |
+
'encoder_layer_decay': 0.9,
|
| 87 |
+
'epochs': 60,
|
| 88 |
+
'final_lr': 0.0001,
|
| 89 |
+
'final_weight_decay': 0.04,
|
| 90 |
+
'ipe': 300,
|
| 91 |
+
'ipe_scale': 1.0,
|
| 92 |
+
'lr': 0.0001,
|
| 93 |
+
'predictor_lr_scale': 3.0,
|
| 94 |
+
'start_lr': 1e-05,
|
| 95 |
+
'warmup': 2,
|
| 96 |
+
'weight_decay': 0.04},
|
| 97 |
+
'tasks_per_node': 3}
|
| 98 |
+
INFO:root:Running pre-training of app: vjepa
|
| 99 |
+
[INFO ][2026-06-23 19:02:23][app.vjepa.train ][main ] which_dtype='bfloat16'
|
| 100 |
+
[INFO ][2026-06-23 19:02:23][app.vjepa.train ][main ] Disabling persistent_workers (incompatible with KNN eval)
|
| 101 |
+
[INFO ][2026-06-23 19:02:23][app.vjepa.train ][main ] NCCL_SOCKET_IFNAME=None
|
| 102 |
+
[INFO ][2026-06-23 19:02:26][app.vjepa.train ][main ] Initialized (rank/world-size) 2/3, tasks_per_node=3
|
| 103 |
+
[WARNING ][2026-06-23 19:03:08][experiments.stmodels.vision_transformers_v3][_build_dinov2_vit ] DINOv2 is an image model — fwd_chunk is unset so all frames will be processed jointly. Set fwd_chunk=1 for per-frame (standard) processing.
|
| 104 |
+
[INFO ][2026-06-23 19:03:19][experiments.stmodels.vision_transformers_v3][_build_dinov2_vit ] Loading pretrained weights for vit_large_patch14_dinov2_lvd142m from timm (vit_large_patch14_dinov2)
|
| 105 |
+
[INFO ][2026-06-23 19:03:21][timm.models._builder][load_pretrained ] Loading pretrained weights from Hugging Face hub (timm/vit_large_patch14_dinov2.lvd142m)
|
| 106 |
+
[INFO ][2026-06-23 19:03:22][timm.models._hub ][load_state_dict_from_hf ] [timm/vit_large_patch14_dinov2.lvd142m] Safe alternative available for 'pytorch_model.bin' (as 'model.safetensors'). Loading weights using safetensors.
|
| 107 |
+
[INFO ][2026-06-23 19:03:24][timm.layers.pos_embed][resample_abs_pos_embed ] Resized position embedding: (37, 37) to (16, 16).
|
| 108 |
+
[INFO ][2026-06-23 19:03:28][root ][init_video_model ] ViTMultiSeqWrapper(
|
| 109 |
+
(backbone): VisionTransformer(
|
| 110 |
+
(patch_embed): PatchEmbed3D(
|
| 111 |
+
(proj): Conv3d(3, 1024, kernel_size=(1, 14, 14), stride=(1, 14, 14))
|
| 112 |
+
)
|
| 113 |
+
(blocks): ModuleList(
|
| 114 |
+
(0-23): 24 x Block(
|
| 115 |
+
(norm1): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 116 |
+
(norm2): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 117 |
+
(drop_path1): Identity()
|
| 118 |
+
(drop_path2): Identity()
|
| 119 |
+
(attn): Attention(
|
| 120 |
+
(qkv): Linear(in_features=1024, out_features=3072, bias=True)
|
| 121 |
+
(attn_drop): Dropout(p=0.0, inplace=False)
|
| 122 |
+
(proj): Linear(in_features=1024, out_features=1024, bias=True)
|
| 123 |
+
(proj_drop): Dropout(p=0.0, inplace=False)
|
| 124 |
+
)
|
| 125 |
+
(mlp): MLP(
|
| 126 |
+
(fc1): Linear(in_features=1024, out_features=4096, bias=True)
|
| 127 |
+
(act): GELU(approximate='none')
|
| 128 |
+
(fc2): Linear(in_features=4096, out_features=1024, bias=True)
|
| 129 |
+
(drop): Dropout(p=0.0, inplace=False)
|
| 130 |
+
)
|
| 131 |
+
)
|
| 132 |
+
)
|
| 133 |
+
(norm): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 134 |
+
)
|
| 135 |
+
)
|
| 136 |
+
[INFO ][2026-06-23 19:03:28][root ][init_video_model ] VJEPAPredictorMultiSeqWrapper(
|
| 137 |
+
(backbone): PredictorV2(
|
| 138 |
+
(predictor_embed): Linear(in_features=1024, out_features=384, bias=True)
|
| 139 |
+
(mask_tokens): ParameterList(
|
| 140 |
+
(0): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
|
| 141 |
+
(1): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
|
| 142 |
+
(2): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
|
| 143 |
+
(3): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
|
| 144 |
+
)
|
| 145 |
+
(predictor_blocks): ModuleList(
|
| 146 |
+
(0-11): 12 x Block(
|
| 147 |
+
(residual1): EfficientResidual(
|
| 148 |
+
(norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
|
| 149 |
+
(fn): Attention(
|
| 150 |
+
(q_proj): Linear(in_features=384, out_features=384, bias=False)
|
| 151 |
+
(k_proj): Linear(in_features=384, out_features=384, bias=False)
|
| 152 |
+
(v_proj): Linear(in_features=384, out_features=384, bias=False)
|
| 153 |
+
(proj): Linear(in_features=384, out_features=384, bias=False)
|
| 154 |
+
(rope): Rope()
|
| 155 |
+
)
|
| 156 |
+
)
|
| 157 |
+
(residual2): EfficientResidual(
|
| 158 |
+
(norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
|
| 159 |
+
(fn): MLP(
|
| 160 |
+
(fc1): Linear(in_features=384, out_features=1536, bias=False)
|
| 161 |
+
(act): GELU(approximate='none')
|
| 162 |
+
(fc2): Linear(in_features=1536, out_features=384, bias=False)
|
| 163 |
+
(drop): Dropout(p=0.0, inplace=False)
|
| 164 |
+
)
|
| 165 |
+
)
|
| 166 |
+
)
|
| 167 |
+
)
|
| 168 |
+
(predictor_norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
|
| 169 |
+
(predictor_proj): Linear(in_features=384, out_features=1024, bias=True)
|
| 170 |
+
)
|
| 171 |
+
)
|
| 172 |
+
[INFO ][2026-06-23 19:03:28][root ][init_video_model ] ViTMultiSeqWrapper(
|
| 173 |
+
(backbone): VisionTransformer(
|
| 174 |
+
(patch_embed): PatchEmbed3D(
|
| 175 |
+
(proj): Conv3d(3, 1024, kernel_size=(1, 14, 14), stride=(1, 14, 14))
|
| 176 |
+
)
|
| 177 |
+
(blocks): ModuleList(
|
| 178 |
+
(0-23): 24 x Block(
|
| 179 |
+
(norm1): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 180 |
+
(norm2): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 181 |
+
(drop_path1): Identity()
|
| 182 |
+
(drop_path2): Identity()
|
| 183 |
+
(attn): Attention(
|
| 184 |
+
(qkv): Linear(in_features=1024, out_features=3072, bias=True)
|
| 185 |
+
(attn_drop): Dropout(p=0.0, inplace=False)
|
| 186 |
+
(proj): Linear(in_features=1024, out_features=1024, bias=True)
|
| 187 |
+
(proj_drop): Dropout(p=0.0, inplace=False)
|
| 188 |
+
)
|
| 189 |
+
(mlp): MLP(
|
| 190 |
+
(fc1): Linear(in_features=1024, out_features=4096, bias=True)
|
| 191 |
+
(act): GELU(approximate='none')
|
| 192 |
+
(fc2): Linear(in_features=4096, out_features=1024, bias=True)
|
| 193 |
+
(drop): Dropout(p=0.0, inplace=False)
|
| 194 |
+
)
|
| 195 |
+
)
|
| 196 |
+
)
|
| 197 |
+
(norm): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 198 |
+
)
|
| 199 |
+
)
|
| 200 |
+
[INFO ][2026-06-23 19:03:28][root ][init_video_model ] Encoder number of parameters: 302964736
|
| 201 |
+
[INFO ][2026-06-23 19:03:28][root ][init_video_model ] Predictor number of parameters: 22042240
|
| 202 |
+
[INFO ][2026-06-23 19:03:28][root ][init_video_model ] Target encoder number of parameters: 302964736
|
| 203 |
+
[INFO ][2026-06-23 19:03:29][root ][make_videodataset ] VideoDataset dataset created
|
| 204 |
+
[INFO ][2026-06-23 19:03:29][WeightedSampler ][__init__ ] Using DistributedWeightedSampler with rank 2 / 3
|
| 205 |
+
[INFO ][2026-06-23 19:03:29][root ][make_videodataset ] VideoDataset unsupervised data loader created
|
| 206 |
+
[INFO ][2026-06-23 19:03:29][app.vjepa.train ][main ] iterations per epoch/dataset length: 300/2103
|
| 207 |
+
[INFO ][2026-06-23 19:03:29][app.vjepa.train ][main ] Wrapping models in DDP (rank 2)...
|
| 208 |
+
[INFO ][2026-06-23 19:03:31][root ][load_checkpoint ] Loading checkpoint from /var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/latest.pt
|
| 209 |
+
[INFO ][2026-06-23 19:03:41][root ][load_checkpoint ] loaded pretrained encoder from epoch 60 with msg: <All keys matched successfully>
|
| 210 |
+
[INFO ][2026-06-23 19:03:41][root ][load_checkpoint ] loaded pretrained predictor from epoch 60 with msg: <All keys matched successfully>
|
| 211 |
+
[INFO ][2026-06-23 19:03:41][root ][load_checkpoint ] loaded pretrained target encoder from epoch 60 with msg: <All keys matched successfully>
|
| 212 |
+
[INFO ][2026-06-23 19:03:42][root ][load_checkpoint ] loaded optimizers from epoch 60
|
| 213 |
+
[INFO ][2026-06-23 19:03:42][root ][load_checkpoint ] read-path: /var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/latest.pt
|
| 214 |
+
submitit WARNING (2026-06-23 19:04:10,794) - Bypassing signal SIGTERM
|
| 215 |
+
[WARNING ][2026-06-23 19:04:10][submitit ][bypass ] Bypassing signal SIGTERM
|
| 216 |
+
submitit WARNING (2026-06-23 19:04:10,794) - Bypassing signal SIGCONT
|
| 217 |
+
[WARNING ][2026-06-23 19:04:10][submitit ][bypass ] Bypassing signal SIGCONT
|
inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_26986/26986_0_log.err
ADDED
|
@@ -0,0 +1,148 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
/var/scratch/dcanezil/.cache/uv/virtualenvs/vd/lib/python3.12/site-packages/timm/models/layers/__init__.py:49: FutureWarning: Importing from timm.models.layers is deprecated, please import via timm.layers
|
| 2 |
+
warnings.warn(f"Importing from {__name__} is deprecated, please import via timm.layers", FutureWarning)
|
| 3 |
+
wandb: Currently logged in as: dgcnz (uvjepa) to https://api.wandb.ai. Use `wandb login --relogin` to force relogin
|
| 4 |
+
wandb: setting up run vfclercb
|
| 5 |
+
wandb: Tracking run with wandb version 0.23.1
|
| 6 |
+
wandb: Run data is saved locally in /var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/wandb/run-20260707_132026-vfclercb
|
| 7 |
+
wandb: Run `wandb offline` to turn off syncing.
|
| 8 |
+
wandb: Resuming run d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow
|
| 9 |
+
wandb: ⭐️ View project at https://wandb.ai/uvjepa/vjepa_ablation
|
| 10 |
+
wandb: 🚀 View run at https://wandb.ai/uvjepa/vjepa_ablation/runs/vfclercb
|
| 11 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/utils.py:888: FutureWarning: `torch.cuda.amp.GradScaler(args...)` is deprecated. Please use `torch.amp.GradScaler('cuda', args...)` instead.
|
| 12 |
+
scaler = torch.cuda.amp.GradScaler() if mixed_precision else None
|
| 13 |
+
/var/scratch/dcanezil/.cache/uv/virtualenvs/vd/lib/python3.12/site-packages/torch/distributed/c10d_logger.py:83: UserWarning: barrier(): using the device under current context. You can specify `device_id` in `init_process_group` to mute this warning.
|
| 14 |
+
return func(*args, **kwargs)
|
| 15 |
+
wandb: WARNING Tried to log to step 0 that is less than the current step 18001. Steps must be monotonically increasing, so this data will be ignored. See https://wandb.me/define-metric to log data out of order.
|
| 16 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 17 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 18 |
+
wandb: WARNING Tried to log to step 18000 that is less than the current step 18001. Steps must be monotonically increasing, so this data will be ignored. See https://wandb.me/define-metric to log data out of order.
|
| 19 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/baking_cookies/Q3XGZmqk1Q0_000296_000306.avi'
|
| 20 |
+
warnings.warn(f"video path not found {fname=}")
|
| 21 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/water_skiing/4bhvIHWADVU_000022_000032.avi'
|
| 22 |
+
warnings.warn(f"video path not found {fname=}")
|
| 23 |
+
[15:57:54] /github/workspace/src/video/video_reader.cc:83: ERROR opening: /local/dcanezil/data/kinetics/VideoData/trimming_trees/T6dqaZioaXs_000004_000014.avi, Invalid data found when processing input
|
| 24 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/bowling/f4eb0wOlspM_000053_000063.avi'
|
| 25 |
+
warnings.warn(f"video path not found {fname=}")
|
| 26 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/making_tea/mtYFNsRcxY4_000063_000073.avi'
|
| 27 |
+
warnings.warn(f"video path not found {fname=}")
|
| 28 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/baking_cookies/Q3XGZmqk1Q0_000296_000306.avi'
|
| 29 |
+
warnings.warn(f"video path not found {fname=}")
|
| 30 |
+
[17:09:51] /github/workspace/src/video/video_reader.cc:83: ERROR opening: /local/dcanezil/data/kinetics/VideoData/long_jump/jsYiitkbr5w_000004_000014.avi, Invalid data found when processing input
|
| 31 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/waxing_legs/cqUusXBODuw_001046_001056.avi'
|
| 32 |
+
warnings.warn(f"video path not found {fname=}")
|
| 33 |
+
/var/scratch/dcanezil/.cache/uv/virtualenvs/vd/lib/python3.12/site-packages/torch/distributed/c10d_logger.py:83: UserWarning: barrier(): using the device under current context. You can specify `device_id` in `init_process_group` to mute this warning.
|
| 34 |
+
return func(*args, **kwargs)
|
| 35 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 36 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 37 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_poker/-jGxlNQKkeo_000092_000102.avi'
|
| 38 |
+
warnings.warn(f"video path not found {fname=}")
|
| 39 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/windsurfing/i-gzh_BPDa8_000154_000164.avi'
|
| 40 |
+
warnings.warn(f"video path not found {fname=}")
|
| 41 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/dancing_macarena/uGQgxFHemyA_000059_000069.avi'
|
| 42 |
+
warnings.warn(f"video path not found {fname=}")
|
| 43 |
+
/var/scratch/dcanezil/.cache/uv/virtualenvs/vd/lib/python3.12/site-packages/torch/distributed/c10d_logger.py:83: UserWarning: barrier(): using the device under current context. You can specify `device_id` in `init_process_group` to mute this warning.
|
| 44 |
+
return func(*args, **kwargs)
|
| 45 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 46 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 47 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/pumping_fist/wKsuyr6Xk30_000094_000104.avi'
|
| 48 |
+
warnings.warn(f"video path not found {fname=}")
|
| 49 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/frying_vegetables/Nr-zCXgW2Sw_000161_000171.avi'
|
| 50 |
+
warnings.warn(f"video path not found {fname=}")
|
| 51 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/pumping_fist/wKsuyr6Xk30_000094_000104.avi'
|
| 52 |
+
warnings.warn(f"video path not found {fname=}")
|
| 53 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_saxophone/Zs0_2GMEPXo_000054_000064.avi'
|
| 54 |
+
warnings.warn(f"video path not found {fname=}")
|
| 55 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/water_skiing/4bhvIHWADVU_000022_000032.avi'
|
| 56 |
+
warnings.warn(f"video path not found {fname=}")
|
| 57 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/windsurfing/i-gzh_BPDa8_000154_000164.avi'
|
| 58 |
+
warnings.warn(f"video path not found {fname=}")
|
| 59 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/water_skiing/4bhvIHWADVU_000022_000032.avi'
|
| 60 |
+
warnings.warn(f"video path not found {fname=}")
|
| 61 |
+
/var/scratch/dcanezil/.cache/uv/virtualenvs/vd/lib/python3.12/site-packages/torch/distributed/c10d_logger.py:83: UserWarning: barrier(): using the device under current context. You can specify `device_id` in `init_process_group` to mute this warning.
|
| 62 |
+
return func(*args, **kwargs)
|
| 63 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 64 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 65 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/surfing_water/5l5Pdd96Pao_000161_000171.avi'
|
| 66 |
+
warnings.warn(f"video path not found {fname=}")
|
| 67 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_poker/6-DyF8umej8_000097_000107.avi'
|
| 68 |
+
warnings.warn(f"video path not found {fname=}")
|
| 69 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/waxing_legs/cqUusXBODuw_001046_001056.avi'
|
| 70 |
+
warnings.warn(f"video path not found {fname=}")
|
| 71 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/waxing_legs/cqUusXBODuw_001046_001056.avi'
|
| 72 |
+
warnings.warn(f"video path not found {fname=}")
|
| 73 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/dancing_macarena/uGQgxFHemyA_000059_000069.avi'
|
| 74 |
+
warnings.warn(f"video path not found {fname=}")
|
| 75 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_poker/-jGxlNQKkeo_000092_000102.avi'
|
| 76 |
+
warnings.warn(f"video path not found {fname=}")
|
| 77 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/frying_vegetables/Nr-zCXgW2Sw_000161_000171.avi'
|
| 78 |
+
warnings.warn(f"video path not found {fname=}")
|
| 79 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_poker/6-DyF8umej8_000097_000107.avi'
|
| 80 |
+
warnings.warn(f"video path not found {fname=}")
|
| 81 |
+
[04:40:06] /github/workspace/src/video/video_reader.cc:83: ERROR opening: /local/dcanezil/data/kinetics/VideoData/long_jump/jsYiitkbr5w_000004_000014.avi, Invalid data found when processing input
|
| 82 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_flute/co50KUHacYw_000005_000015.avi'
|
| 83 |
+
warnings.warn(f"video path not found {fname=}")
|
| 84 |
+
/var/scratch/dcanezil/.cache/uv/virtualenvs/vd/lib/python3.12/site-packages/torch/distributed/c10d_logger.py:83: UserWarning: barrier(): using the device under current context. You can specify `device_id` in `init_process_group` to mute this warning.
|
| 85 |
+
return func(*args, **kwargs)
|
| 86 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 87 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 88 |
+
[05:19:11] /github/workspace/src/video/video_reader.cc:83: ERROR opening: /local/dcanezil/data/kinetics/VideoData/trimming_trees/T6dqaZioaXs_000004_000014.avi, Invalid data found when processing input
|
| 89 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/picking_fruit/3VvkoFtPCCU_000045_000055.avi'
|
| 90 |
+
warnings.warn(f"video path not found {fname=}")
|
| 91 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/windsurfing/i-gzh_BPDa8_000154_000164.avi'
|
| 92 |
+
warnings.warn(f"video path not found {fname=}")
|
| 93 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/pumping_fist/wKsuyr6Xk30_000094_000104.avi'
|
| 94 |
+
warnings.warn(f"video path not found {fname=}")
|
| 95 |
+
/var/scratch/dcanezil/.cache/uv/virtualenvs/vd/lib/python3.12/site-packages/torch/distributed/c10d_logger.py:83: UserWarning: barrier(): using the device under current context. You can specify `device_id` in `init_process_group` to mute this warning.
|
| 96 |
+
return func(*args, **kwargs)
|
| 97 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 98 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 99 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_saxophone/Zs0_2GMEPXo_000054_000064.avi'
|
| 100 |
+
warnings.warn(f"video path not found {fname=}")
|
| 101 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_poker/6-DyF8umej8_000097_000107.avi'
|
| 102 |
+
warnings.warn(f"video path not found {fname=}")
|
| 103 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/frying_vegetables/Nr-zCXgW2Sw_000161_000171.avi'
|
| 104 |
+
warnings.warn(f"video path not found {fname=}")
|
| 105 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_saxophone/Zs0_2GMEPXo_000054_000064.avi'
|
| 106 |
+
warnings.warn(f"video path not found {fname=}")
|
| 107 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_monopoly/NLL667uPWVA_000066_000076.avi'
|
| 108 |
+
warnings.warn(f"video path not found {fname=}")
|
| 109 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/baking_cookies/Q3XGZmqk1Q0_000296_000306.avi'
|
| 110 |
+
warnings.warn(f"video path not found {fname=}")
|
| 111 |
+
[12:42:10] /github/workspace/src/video/video_reader.cc:83: ERROR opening: /local/dcanezil/data/kinetics/VideoData/climbing_tree/QX5IzK4iFCM_000002_000012.avi, Invalid data found when processing input
|
| 112 |
+
/var/scratch/dcanezil/.cache/uv/virtualenvs/vd/lib/python3.12/site-packages/torch/distributed/c10d_logger.py:83: UserWarning: barrier(): using the device under current context. You can specify `device_id` in `init_process_group` to mute this warning.
|
| 113 |
+
return func(*args, **kwargs)
|
| 114 |
+
wandb: updating run metadata
|
| 115 |
+
wandb: uploading output.log; uploading wandb-summary.json
|
| 116 |
+
wandb: uploading history steps 2790-2790, summary, console lines 1360-1367
|
| 117 |
+
wandb:
|
| 118 |
+
wandb: Run history:
|
| 119 |
+
wandb: epoch ▁▁▁▁▁▁▂▂▂▂▂▃▃▃▃▄▅▅▅▅▅▅▅▅▅▅▅▅▆▆▇▇▇▇▇█████
|
| 120 |
+
wandb: epoch/avg_data_time_ms ▆▃▄▄▁▅▃▅▅▅▃▅▄▃▃▃█▃▅▅▂▆▄▆▁▆▅▃▅▄
|
| 121 |
+
wandb: epoch/avg_gpu_time_ms ▃▄▃▂▂▃▂▇▇▅▆▇▄█▅▄▄▁▂▃▇▄▇▃▂▆▆▅▇▆
|
| 122 |
+
wandb: epoch/avg_iter_time_ms ▃▄▃▂▂▃▂▇▇▅▆▇▄█▅▄▄▁▃▄▇▅▇▄▂▆▆▅▇▆
|
| 123 |
+
wandb: epoch/avg_loss ▁▁▁▂▂▂▃▃▃▄▄▄▄▄▅▅▅▆▆▆▆▇▇▇▇▇▇███
|
| 124 |
+
wandb: knn/ssv2_coarse10_temporal_concat_top1 ▁█▇▄▄▂
|
| 125 |
+
wandb: knn/ssv2_coarse10_temporal_concat_top5 █▇▂▃▁▆
|
| 126 |
+
wandb: linear/ssv2_coarse10_temporal_concat_top1 ▁▂▃█▆█
|
| 127 |
+
wandb: linear/ssv2_coarse10_temporal_concat_top5 ▄▄▁▃█▄
|
| 128 |
+
wandb: train/ema/cos_sim ▆▇█▅▅▆▅▃▄▅▅▂▅▄▄▆▅▄▃▅▄▄▆▄▄▃▃█▃▃▄▂▂▂▃▄▄▄▃▁
|
| 129 |
+
wandb: +34 ...
|
| 130 |
+
wandb:
|
| 131 |
+
wandb: Run summary:
|
| 132 |
+
wandb: epoch 90
|
| 133 |
+
wandb: epoch/avg_data_time_ms 15.44802
|
| 134 |
+
wandb: epoch/avg_gpu_time_ms 9210.69954
|
| 135 |
+
wandb: epoch/avg_iter_time_ms 9238.44177
|
| 136 |
+
wandb: epoch/avg_loss 0.54822
|
| 137 |
+
wandb: knn/ssv2_coarse10_temporal_concat_top1 45.08032
|
| 138 |
+
wandb: knn/ssv2_coarse10_temporal_concat_top5 87.249
|
| 139 |
+
wandb: linear/ssv2_coarse10_temporal_concat_top1 58.93574
|
| 140 |
+
wandb: linear/ssv2_coarse10_temporal_concat_top5 92.67068
|
| 141 |
+
wandb: train/ema/cos_sim 0.74331
|
| 142 |
+
wandb: +34 ...
|
| 143 |
+
wandb:
|
| 144 |
+
wandb: 🚀 View run d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow at: https://wandb.ai/uvjepa/vjepa_ablation/runs/vfclercb
|
| 145 |
+
wandb: ⭐️ View project at: https://wandb.ai/uvjepa/vjepa_ablation
|
| 146 |
+
wandb: Synced 5 W&B file(s), 0 media file(s), 0 artifact file(s) and 0 other file(s)
|
| 147 |
+
wandb: Find logs at: /var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/wandb/run-20260707_132026-vfclercb/logs
|
| 148 |
+
[rank0]:[W708 12:54:53.410661669 ProcessGroupNCCL.cpp:1553] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())
|
inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_26986/26986_0_log.out
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_26986/26986_1_log.err
ADDED
|
@@ -0,0 +1,82 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
/var/scratch/dcanezil/.cache/uv/virtualenvs/vd/lib/python3.12/site-packages/timm/models/layers/__init__.py:49: FutureWarning: Importing from timm.models.layers is deprecated, please import via timm.layers
|
| 2 |
+
warnings.warn(f"Importing from {__name__} is deprecated, please import via timm.layers", FutureWarning)
|
| 3 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/utils.py:888: FutureWarning: `torch.cuda.amp.GradScaler(args...)` is deprecated. Please use `torch.amp.GradScaler('cuda', args...)` instead.
|
| 4 |
+
scaler = torch.cuda.amp.GradScaler() if mixed_precision else None
|
| 5 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 6 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 7 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/waxing_legs/cqUusXBODuw_001046_001056.avi'
|
| 8 |
+
warnings.warn(f"video path not found {fname=}")
|
| 9 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/robot_dancing/J1x0XlWa6HM_000147_000157.avi'
|
| 10 |
+
warnings.warn(f"video path not found {fname=}")
|
| 11 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/making_tea/mtYFNsRcxY4_000063_000073.avi'
|
| 12 |
+
warnings.warn(f"video path not found {fname=}")
|
| 13 |
+
[15:05:40] /github/workspace/src/video/video_reader.cc:83: ERROR opening: /local/dcanezil/data/kinetics/VideoData/long_jump/jsYiitkbr5w_000004_000014.avi, Invalid data found when processing input
|
| 14 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/waxing_legs/cqUusXBODuw_001046_001056.avi'
|
| 15 |
+
warnings.warn(f"video path not found {fname=}")
|
| 16 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/windsurfing/i-gzh_BPDa8_000154_000164.avi'
|
| 17 |
+
warnings.warn(f"video path not found {fname=}")
|
| 18 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/baking_cookies/Q3XGZmqk1Q0_000296_000306.avi'
|
| 19 |
+
warnings.warn(f"video path not found {fname=}")
|
| 20 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 21 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 22 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/pumping_fist/wKsuyr6Xk30_000094_000104.avi'
|
| 23 |
+
warnings.warn(f"video path not found {fname=}")
|
| 24 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/water_skiing/4bhvIHWADVU_000022_000032.avi'
|
| 25 |
+
warnings.warn(f"video path not found {fname=}")
|
| 26 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/sailing/ZolXG_-RgIs_000970_000980.avi'
|
| 27 |
+
warnings.warn(f"video path not found {fname=}")
|
| 28 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/dancing_macarena/uGQgxFHemyA_000059_000069.avi'
|
| 29 |
+
warnings.warn(f"video path not found {fname=}")
|
| 30 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_poker/-jGxlNQKkeo_000092_000102.avi'
|
| 31 |
+
warnings.warn(f"video path not found {fname=}")
|
| 32 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/making_tea/mtYFNsRcxY4_000063_000073.avi'
|
| 33 |
+
warnings.warn(f"video path not found {fname=}")
|
| 34 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/dancing_macarena/uGQgxFHemyA_000059_000069.avi'
|
| 35 |
+
warnings.warn(f"video path not found {fname=}")
|
| 36 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 37 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 38 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/water_skiing/4bhvIHWADVU_000022_000032.avi'
|
| 39 |
+
warnings.warn(f"video path not found {fname=}")
|
| 40 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_poker/6-DyF8umej8_000097_000107.avi'
|
| 41 |
+
warnings.warn(f"video path not found {fname=}")
|
| 42 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_poker/-jGxlNQKkeo_000092_000102.avi'
|
| 43 |
+
warnings.warn(f"video path not found {fname=}")
|
| 44 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_monopoly/NLL667uPWVA_000066_000076.avi'
|
| 45 |
+
warnings.warn(f"video path not found {fname=}")
|
| 46 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/sailing/ZolXG_-RgIs_000970_000980.avi'
|
| 47 |
+
warnings.warn(f"video path not found {fname=}")
|
| 48 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/eating_carrots/eiZ8Hzc7FPU_000080_000090.avi'
|
| 49 |
+
warnings.warn(f"video path not found {fname=}")
|
| 50 |
+
[00:44:39] /github/workspace/src/video/video_reader.cc:83: ERROR opening: /local/dcanezil/data/kinetics/VideoData/trimming_trees/T6dqaZioaXs_000004_000014.avi, Invalid data found when processing input
|
| 51 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/water_skiing/4bhvIHWADVU_000022_000032.avi'
|
| 52 |
+
warnings.warn(f"video path not found {fname=}")
|
| 53 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 54 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 55 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/waxing_legs/cqUusXBODuw_001046_001056.avi'
|
| 56 |
+
warnings.warn(f"video path not found {fname=}")
|
| 57 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/sailing/ZolXG_-RgIs_000970_000980.avi'
|
| 58 |
+
warnings.warn(f"video path not found {fname=}")
|
| 59 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/water_skiing/4bhvIHWADVU_000022_000032.avi'
|
| 60 |
+
warnings.warn(f"video path not found {fname=}")
|
| 61 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 62 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 63 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/bowling/f4eb0wOlspM_000053_000063.avi'
|
| 64 |
+
warnings.warn(f"video path not found {fname=}")
|
| 65 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_monopoly/NLL667uPWVA_000066_000076.avi'
|
| 66 |
+
warnings.warn(f"video path not found {fname=}")
|
| 67 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 68 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 69 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/frying_vegetables/Nr-zCXgW2Sw_000161_000171.avi'
|
| 70 |
+
warnings.warn(f"video path not found {fname=}")
|
| 71 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/eating_carrots/eiZ8Hzc7FPU_000080_000090.avi'
|
| 72 |
+
warnings.warn(f"video path not found {fname=}")
|
| 73 |
+
[10:33:29] /github/workspace/src/video/video_reader.cc:83: ERROR opening: /local/dcanezil/data/kinetics/VideoData/climbing_tree/QX5IzK4iFCM_000002_000012.avi, Invalid data found when processing input
|
| 74 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_flute/co50KUHacYw_000005_000015.avi'
|
| 75 |
+
warnings.warn(f"video path not found {fname=}")
|
| 76 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/water_skiing/4bhvIHWADVU_000022_000032.avi'
|
| 77 |
+
warnings.warn(f"video path not found {fname=}")
|
| 78 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_saxophone/Zs0_2GMEPXo_000054_000064.avi'
|
| 79 |
+
warnings.warn(f"video path not found {fname=}")
|
| 80 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/picking_fruit/3VvkoFtPCCU_000045_000055.avi'
|
| 81 |
+
warnings.warn(f"video path not found {fname=}")
|
| 82 |
+
[rank1]:[W708 12:54:52.177463153 ProcessGroupNCCL.cpp:1553] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())
|
inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_26986/26986_1_log.out
ADDED
|
@@ -0,0 +1,218 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
submitit INFO (2026-07-07 13:20:01,953) - Starting with JobEnvironment(job_id=26986, hostname=node404, local_rank=1(3), node=0(1), global_rank=1(3))
|
| 2 |
+
submitit INFO (2026-07-07 13:20:01,953) - Loading pickle: /var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_26986/26986_submitted.pkl
|
| 3 |
+
INFO:root:loaded pretrain params...
|
| 4 |
+
{ 'app': 'vjepa',
|
| 5 |
+
'cpus_per_task': 16,
|
| 6 |
+
'data': { 'batch_size': 64,
|
| 7 |
+
'crop_size': 224,
|
| 8 |
+
'dataset_fpcs': [8, 8],
|
| 9 |
+
'dataset_type': 'VideoDataset',
|
| 10 |
+
'datasets': [ '/local/dcanezil/data/kinetics/train.csv',
|
| 11 |
+
'/local/dcanezil/data/ssv2/train.csv'],
|
| 12 |
+
'datasets_weights': [0.65, 0.35],
|
| 13 |
+
'fps': 2,
|
| 14 |
+
'num_workers': 4,
|
| 15 |
+
'patch_size': 14,
|
| 16 |
+
'persistent_workers': True,
|
| 17 |
+
'pin_mem': True,
|
| 18 |
+
'tubelet_size': 1},
|
| 19 |
+
'data_aug': { 'auto_augment': False,
|
| 20 |
+
'motion_shift': False,
|
| 21 |
+
'random_resize_aspect_ratio': [0.75, 1.35],
|
| 22 |
+
'random_resize_scale': [0.3, 1.0],
|
| 23 |
+
'reprob': 0.0},
|
| 24 |
+
'folder': '/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow',
|
| 25 |
+
'loss': {'loss_exp': 1.0},
|
| 26 |
+
'mask': [ { 'aspect_ratio': [0.75, 1.5],
|
| 27 |
+
'full_complement': False,
|
| 28 |
+
'max_keep': None,
|
| 29 |
+
'max_temporal_keep': 1.0,
|
| 30 |
+
'num_blocks': 8,
|
| 31 |
+
'spatial_scale': [0.05, 0.05],
|
| 32 |
+
'temporal_scale': [1.0, 1.0]},
|
| 33 |
+
{ 'aspect_ratio': [0.75, 1.5],
|
| 34 |
+
'full_complement': False,
|
| 35 |
+
'max_keep': None,
|
| 36 |
+
'max_temporal_keep': 1.0,
|
| 37 |
+
'num_blocks': 2,
|
| 38 |
+
'spatial_scale': [0.4, 0.4],
|
| 39 |
+
'temporal_scale': [1.0, 1.0]}],
|
| 40 |
+
'mem_per_gpu': '40G',
|
| 41 |
+
'meta': { 'dtype': 'bfloat16',
|
| 42 |
+
'knn_eval_epoch0': True,
|
| 43 |
+
'knn_eval_freq': 5,
|
| 44 |
+
'knn_eval_presets': [ { 'config': { 'batch_size': 8,
|
| 45 |
+
'dataset_train': '/local/dcanezil/data/ssv2/train_coarse10.csv',
|
| 46 |
+
'dataset_val': '/local/dcanezil/data/ssv2/val_coarse10.csv',
|
| 47 |
+
'eval_videos_per_class': 100,
|
| 48 |
+
'linear_probe': True,
|
| 49 |
+
'num_workers': 2,
|
| 50 |
+
'pool_type': 'temporal_concat',
|
| 51 |
+
'train_videos_per_class': 500},
|
| 52 |
+
'preset': 'ssv2_coarse10'}],
|
| 53 |
+
'load_checkpoint': True,
|
| 54 |
+
'read_checkpoint': None,
|
| 55 |
+
'save_every_freq': -1,
|
| 56 |
+
'seed': 239,
|
| 57 |
+
'use_sdpa': True,
|
| 58 |
+
'use_wandb': True,
|
| 59 |
+
'wandb_project': 'vjepa_ablation'},
|
| 60 |
+
'metrics': {'sigreg': {}, 'std': {}},
|
| 61 |
+
'model': { 'class_token': True,
|
| 62 |
+
'freeze_backbone': False,
|
| 63 |
+
'is_causal': True,
|
| 64 |
+
'model_name': 'vit_large_patch14_dinov2_lvd142m',
|
| 65 |
+
'pe_type': '2d',
|
| 66 |
+
'pred_depth': 12,
|
| 67 |
+
'pred_embed_dim': 384,
|
| 68 |
+
'pred_num_heads': 6,
|
| 69 |
+
'predictor': 'v2_cross',
|
| 70 |
+
'stem_type': '2d',
|
| 71 |
+
'target_embed_dim': 1024,
|
| 72 |
+
'target_kind': 'ema',
|
| 73 |
+
'uniform_power': True,
|
| 74 |
+
'use_activation_checkpointing': True,
|
| 75 |
+
'use_alibi': False,
|
| 76 |
+
'use_lora': False,
|
| 77 |
+
'use_mask_tokens': True,
|
| 78 |
+
'use_rope': False,
|
| 79 |
+
'use_sdpa': True,
|
| 80 |
+
'zero_init_mask_tokens': True},
|
| 81 |
+
'nodes': 1,
|
| 82 |
+
'optimization': { 'clip_grad': None,
|
| 83 |
+
'ema': [1.0, 0.99925],
|
| 84 |
+
'ema_hold_epochs': 0,
|
| 85 |
+
'ema_ramp_epochs': 5,
|
| 86 |
+
'encoder_layer_decay': 0.9,
|
| 87 |
+
'epochs': 90,
|
| 88 |
+
'final_lr': 0.0001,
|
| 89 |
+
'final_weight_decay': 0.04,
|
| 90 |
+
'ipe': 300,
|
| 91 |
+
'ipe_scale': 1.0,
|
| 92 |
+
'lr': 0.0001,
|
| 93 |
+
'predictor_lr_scale': 3.0,
|
| 94 |
+
'start_lr': 1e-05,
|
| 95 |
+
'warmup': 2,
|
| 96 |
+
'weight_decay': 0.04},
|
| 97 |
+
'tasks_per_node': 3}
|
| 98 |
+
INFO:root:Running pre-training of app: vjepa
|
| 99 |
+
[INFO ][2026-07-07 13:20:22][app.vjepa.train ][main ] which_dtype='bfloat16'
|
| 100 |
+
[INFO ][2026-07-07 13:20:22][app.vjepa.train ][main ] Disabling persistent_workers (incompatible with KNN eval)
|
| 101 |
+
[INFO ][2026-07-07 13:20:22][app.vjepa.train ][main ] NCCL_SOCKET_IFNAME=None
|
| 102 |
+
[INFO ][2026-07-07 13:20:25][app.vjepa.train ][main ] Initialized (rank/world-size) 1/3, tasks_per_node=3
|
| 103 |
+
[WARNING ][2026-07-07 13:21:13][experiments.stmodels.vision_transformers_v3][_build_dinov2_vit ] DINOv2 is an image model — fwd_chunk is unset so all frames will be processed jointly. Set fwd_chunk=1 for per-frame (standard) processing.
|
| 104 |
+
[INFO ][2026-07-07 13:21:23][experiments.stmodels.vision_transformers_v3][_build_dinov2_vit ] Loading pretrained weights for vit_large_patch14_dinov2_lvd142m from timm (vit_large_patch14_dinov2)
|
| 105 |
+
[INFO ][2026-07-07 13:21:26][timm.models._builder][load_pretrained ] Loading pretrained weights from Hugging Face hub (timm/vit_large_patch14_dinov2.lvd142m)
|
| 106 |
+
[INFO ][2026-07-07 13:21:26][timm.models._hub ][load_state_dict_from_hf ] [timm/vit_large_patch14_dinov2.lvd142m] Safe alternative available for 'pytorch_model.bin' (as 'model.safetensors'). Loading weights using safetensors.
|
| 107 |
+
[INFO ][2026-07-07 13:21:27][timm.layers.pos_embed][resample_abs_pos_embed ] Resized position embedding: (37, 37) to (16, 16).
|
| 108 |
+
[INFO ][2026-07-07 13:21:32][root ][init_video_model ] ViTMultiSeqWrapper(
|
| 109 |
+
(backbone): VisionTransformer(
|
| 110 |
+
(patch_embed): PatchEmbed3D(
|
| 111 |
+
(proj): Conv3d(3, 1024, kernel_size=(1, 14, 14), stride=(1, 14, 14))
|
| 112 |
+
)
|
| 113 |
+
(blocks): ModuleList(
|
| 114 |
+
(0-23): 24 x Block(
|
| 115 |
+
(norm1): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 116 |
+
(norm2): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 117 |
+
(drop_path1): Identity()
|
| 118 |
+
(drop_path2): Identity()
|
| 119 |
+
(attn): Attention(
|
| 120 |
+
(qkv): Linear(in_features=1024, out_features=3072, bias=True)
|
| 121 |
+
(attn_drop): Dropout(p=0.0, inplace=False)
|
| 122 |
+
(proj): Linear(in_features=1024, out_features=1024, bias=True)
|
| 123 |
+
(proj_drop): Dropout(p=0.0, inplace=False)
|
| 124 |
+
)
|
| 125 |
+
(mlp): MLP(
|
| 126 |
+
(fc1): Linear(in_features=1024, out_features=4096, bias=True)
|
| 127 |
+
(act): GELU(approximate='none')
|
| 128 |
+
(fc2): Linear(in_features=4096, out_features=1024, bias=True)
|
| 129 |
+
(drop): Dropout(p=0.0, inplace=False)
|
| 130 |
+
)
|
| 131 |
+
)
|
| 132 |
+
)
|
| 133 |
+
(norm): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 134 |
+
)
|
| 135 |
+
)
|
| 136 |
+
[INFO ][2026-07-07 13:21:32][root ][init_video_model ] VJEPAPredictorMultiSeqWrapper(
|
| 137 |
+
(backbone): PredictorV2(
|
| 138 |
+
(predictor_embed): Linear(in_features=1024, out_features=384, bias=True)
|
| 139 |
+
(mask_tokens): ParameterList(
|
| 140 |
+
(0): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
|
| 141 |
+
(1): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
|
| 142 |
+
(2): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
|
| 143 |
+
(3): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
|
| 144 |
+
)
|
| 145 |
+
(predictor_blocks): ModuleList(
|
| 146 |
+
(0-11): 12 x Block(
|
| 147 |
+
(residual1): EfficientResidual(
|
| 148 |
+
(norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
|
| 149 |
+
(fn): Attention(
|
| 150 |
+
(q_proj): Linear(in_features=384, out_features=384, bias=False)
|
| 151 |
+
(k_proj): Linear(in_features=384, out_features=384, bias=False)
|
| 152 |
+
(v_proj): Linear(in_features=384, out_features=384, bias=False)
|
| 153 |
+
(proj): Linear(in_features=384, out_features=384, bias=False)
|
| 154 |
+
(rope): Rope()
|
| 155 |
+
)
|
| 156 |
+
)
|
| 157 |
+
(residual2): EfficientResidual(
|
| 158 |
+
(norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
|
| 159 |
+
(fn): MLP(
|
| 160 |
+
(fc1): Linear(in_features=384, out_features=1536, bias=False)
|
| 161 |
+
(act): GELU(approximate='none')
|
| 162 |
+
(fc2): Linear(in_features=1536, out_features=384, bias=False)
|
| 163 |
+
(drop): Dropout(p=0.0, inplace=False)
|
| 164 |
+
)
|
| 165 |
+
)
|
| 166 |
+
)
|
| 167 |
+
)
|
| 168 |
+
(predictor_norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
|
| 169 |
+
(predictor_proj): Linear(in_features=384, out_features=1024, bias=True)
|
| 170 |
+
)
|
| 171 |
+
)
|
| 172 |
+
[INFO ][2026-07-07 13:21:32][root ][init_video_model ] ViTMultiSeqWrapper(
|
| 173 |
+
(backbone): VisionTransformer(
|
| 174 |
+
(patch_embed): PatchEmbed3D(
|
| 175 |
+
(proj): Conv3d(3, 1024, kernel_size=(1, 14, 14), stride=(1, 14, 14))
|
| 176 |
+
)
|
| 177 |
+
(blocks): ModuleList(
|
| 178 |
+
(0-23): 24 x Block(
|
| 179 |
+
(norm1): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 180 |
+
(norm2): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 181 |
+
(drop_path1): Identity()
|
| 182 |
+
(drop_path2): Identity()
|
| 183 |
+
(attn): Attention(
|
| 184 |
+
(qkv): Linear(in_features=1024, out_features=3072, bias=True)
|
| 185 |
+
(attn_drop): Dropout(p=0.0, inplace=False)
|
| 186 |
+
(proj): Linear(in_features=1024, out_features=1024, bias=True)
|
| 187 |
+
(proj_drop): Dropout(p=0.0, inplace=False)
|
| 188 |
+
)
|
| 189 |
+
(mlp): MLP(
|
| 190 |
+
(fc1): Linear(in_features=1024, out_features=4096, bias=True)
|
| 191 |
+
(act): GELU(approximate='none')
|
| 192 |
+
(fc2): Linear(in_features=4096, out_features=1024, bias=True)
|
| 193 |
+
(drop): Dropout(p=0.0, inplace=False)
|
| 194 |
+
)
|
| 195 |
+
)
|
| 196 |
+
)
|
| 197 |
+
(norm): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 198 |
+
)
|
| 199 |
+
)
|
| 200 |
+
[INFO ][2026-07-07 13:21:32][root ][init_video_model ] Encoder number of parameters: 302964736
|
| 201 |
+
[INFO ][2026-07-07 13:21:32][root ][init_video_model ] Predictor number of parameters: 22042240
|
| 202 |
+
[INFO ][2026-07-07 13:21:32][root ][init_video_model ] Target encoder number of parameters: 302964736
|
| 203 |
+
[INFO ][2026-07-07 13:21:33][root ][make_videodataset ] VideoDataset dataset created
|
| 204 |
+
[INFO ][2026-07-07 13:21:33][WeightedSampler ][__init__ ] Using DistributedWeightedSampler with rank 1 / 3
|
| 205 |
+
[INFO ][2026-07-07 13:21:33][root ][make_videodataset ] VideoDataset unsupervised data loader created
|
| 206 |
+
[INFO ][2026-07-07 13:21:33][app.vjepa.train ][main ] iterations per epoch/dataset length: 300/2103
|
| 207 |
+
[INFO ][2026-07-07 13:21:33][app.vjepa.train ][main ] Wrapping models in DDP (rank 1)...
|
| 208 |
+
[INFO ][2026-07-07 13:21:34][root ][load_checkpoint ] Loading checkpoint from /var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/latest.pt
|
| 209 |
+
[INFO ][2026-07-07 13:21:44][root ][load_checkpoint ] loaded pretrained encoder from epoch 60 with msg: <All keys matched successfully>
|
| 210 |
+
[INFO ][2026-07-07 13:21:44][root ][load_checkpoint ] loaded pretrained predictor from epoch 60 with msg: <All keys matched successfully>
|
| 211 |
+
[INFO ][2026-07-07 13:21:45][root ][load_checkpoint ] loaded pretrained target encoder from epoch 60 with msg: <All keys matched successfully>
|
| 212 |
+
[INFO ][2026-07-07 13:21:45][root ][load_checkpoint ] loaded optimizers from epoch 60
|
| 213 |
+
[INFO ][2026-07-07 13:21:45][root ][load_checkpoint ] read-path: /var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/latest.pt
|
| 214 |
+
[INFO ][2026-07-07 13:23:07][app.vjepa.train ][main ] Initializing loader...
|
| 215 |
+
submitit INFO (2026-07-08 12:54:51,517) - Job completed successfully
|
| 216 |
+
[INFO ][2026-07-08 12:54:51][submitit ][process_job ] Job completed successfully
|
| 217 |
+
submitit INFO (2026-07-08 12:54:51,520) - Exiting after successful completion
|
| 218 |
+
[INFO ][2026-07-08 12:54:51][submitit ][process_job ] Exiting after successful completion
|
inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_26986/26986_2_log.err
ADDED
|
@@ -0,0 +1,79 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
/var/scratch/dcanezil/.cache/uv/virtualenvs/vd/lib/python3.12/site-packages/timm/models/layers/__init__.py:49: FutureWarning: Importing from timm.models.layers is deprecated, please import via timm.layers
|
| 2 |
+
warnings.warn(f"Importing from {__name__} is deprecated, please import via timm.layers", FutureWarning)
|
| 3 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/utils.py:888: FutureWarning: `torch.cuda.amp.GradScaler(args...)` is deprecated. Please use `torch.amp.GradScaler('cuda', args...)` instead.
|
| 4 |
+
scaler = torch.cuda.amp.GradScaler() if mixed_precision else None
|
| 5 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 6 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 7 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_poker/6-DyF8umej8_000097_000107.avi'
|
| 8 |
+
warnings.warn(f"video path not found {fname=}")
|
| 9 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/surfing_water/5l5Pdd96Pao_000161_000171.avi'
|
| 10 |
+
warnings.warn(f"video path not found {fname=}")
|
| 11 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/robot_dancing/J1x0XlWa6HM_000147_000157.avi'
|
| 12 |
+
warnings.warn(f"video path not found {fname=}")
|
| 13 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/water_skiing/4bhvIHWADVU_000022_000032.avi'
|
| 14 |
+
warnings.warn(f"video path not found {fname=}")
|
| 15 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/water_skiing/4bhvIHWADVU_000022_000032.avi'
|
| 16 |
+
warnings.warn(f"video path not found {fname=}")
|
| 17 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/surfing_water/5l5Pdd96Pao_000161_000171.avi'
|
| 18 |
+
warnings.warn(f"video path not found {fname=}")
|
| 19 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_monopoly/NLL667uPWVA_000066_000076.avi'
|
| 20 |
+
warnings.warn(f"video path not found {fname=}")
|
| 21 |
+
[17:04:08] /github/workspace/src/video/video_reader.cc:83: ERROR opening: /local/dcanezil/data/kinetics/VideoData/long_jump/jsYiitkbr5w_000004_000014.avi, Invalid data found when processing input
|
| 22 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/frying_vegetables/Nr-zCXgW2Sw_000161_000171.avi'
|
| 23 |
+
warnings.warn(f"video path not found {fname=}")
|
| 24 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 25 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 26 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/windsurfing/i-gzh_BPDa8_000154_000164.avi'
|
| 27 |
+
warnings.warn(f"video path not found {fname=}")
|
| 28 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_saxophone/Zs0_2GMEPXo_000054_000064.avi'
|
| 29 |
+
warnings.warn(f"video path not found {fname=}")
|
| 30 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_flute/co50KUHacYw_000005_000015.avi'
|
| 31 |
+
warnings.warn(f"video path not found {fname=}")
|
| 32 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/making_tea/mtYFNsRcxY4_000063_000073.avi'
|
| 33 |
+
warnings.warn(f"video path not found {fname=}")
|
| 34 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/dancing_macarena/uGQgxFHemyA_000059_000069.avi'
|
| 35 |
+
warnings.warn(f"video path not found {fname=}")
|
| 36 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_poker/-jGxlNQKkeo_000092_000102.avi'
|
| 37 |
+
warnings.warn(f"video path not found {fname=}")
|
| 38 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/water_skiing/4bhvIHWADVU_000022_000032.avi'
|
| 39 |
+
warnings.warn(f"video path not found {fname=}")
|
| 40 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/frying_vegetables/Nr-zCXgW2Sw_000161_000171.avi'
|
| 41 |
+
warnings.warn(f"video path not found {fname=}")
|
| 42 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/sailing/ZolXG_-RgIs_000970_000980.avi'
|
| 43 |
+
warnings.warn(f"video path not found {fname=}")
|
| 44 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 45 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 46 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_poker/6-DyF8umej8_000097_000107.avi'
|
| 47 |
+
warnings.warn(f"video path not found {fname=}")
|
| 48 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/bowling/f4eb0wOlspM_000053_000063.avi'
|
| 49 |
+
warnings.warn(f"video path not found {fname=}")
|
| 50 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/eating_carrots/eiZ8Hzc7FPU_000080_000090.avi'
|
| 51 |
+
warnings.warn(f"video path not found {fname=}")
|
| 52 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_poker/6-DyF8umej8_000097_000107.avi'
|
| 53 |
+
warnings.warn(f"video path not found {fname=}")
|
| 54 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/waxing_legs/cqUusXBODuw_001046_001056.avi'
|
| 55 |
+
warnings.warn(f"video path not found {fname=}")
|
| 56 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/waxing_legs/cqUusXBODuw_001046_001056.avi'
|
| 57 |
+
warnings.warn(f"video path not found {fname=}")
|
| 58 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 59 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 60 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/picking_fruit/3VvkoFtPCCU_000045_000055.avi'
|
| 61 |
+
warnings.warn(f"video path not found {fname=}")
|
| 62 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/making_tea/mtYFNsRcxY4_000063_000073.avi'
|
| 63 |
+
warnings.warn(f"video path not found {fname=}")
|
| 64 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/windsurfing/i-gzh_BPDa8_000154_000164.avi'
|
| 65 |
+
warnings.warn(f"video path not found {fname=}")
|
| 66 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/picking_fruit/3VvkoFtPCCU_000045_000055.avi'
|
| 67 |
+
warnings.warn(f"video path not found {fname=}")
|
| 68 |
+
[04:14:21] /github/workspace/src/video/video_reader.cc:83: ERROR opening: /local/dcanezil/data/kinetics/VideoData/climbing_tree/QX5IzK4iFCM_000002_000012.avi, Invalid data found when processing input
|
| 69 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 70 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 71 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 72 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 73 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/dancing_macarena/uGQgxFHemyA_000059_000069.avi'
|
| 74 |
+
warnings.warn(f"video path not found {fname=}")
|
| 75 |
+
[09:40:33] /github/workspace/src/video/video_reader.cc:83: ERROR opening: /local/dcanezil/data/kinetics/VideoData/climbing_tree/QX5IzK4iFCM_000002_000012.avi, Invalid data found when processing input
|
| 76 |
+
[11:51:11] /github/workspace/src/video/video_reader.cc:83: ERROR opening: /local/dcanezil/data/kinetics/VideoData/trimming_trees/T6dqaZioaXs_000004_000014.avi, Invalid data found when processing input
|
| 77 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/making_tea/mtYFNsRcxY4_000063_000073.avi'
|
| 78 |
+
warnings.warn(f"video path not found {fname=}")
|
| 79 |
+
[rank2]:[W708 12:54:52.180945773 ProcessGroupNCCL.cpp:1553] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())
|
inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_26986/26986_2_log.out
ADDED
|
@@ -0,0 +1,218 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
submitit INFO (2026-07-07 13:20:01,953) - Starting with JobEnvironment(job_id=26986, hostname=node404, local_rank=2(3), node=0(1), global_rank=2(3))
|
| 2 |
+
submitit INFO (2026-07-07 13:20:01,953) - Loading pickle: /var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_26986/26986_submitted.pkl
|
| 3 |
+
INFO:root:loaded pretrain params...
|
| 4 |
+
{ 'app': 'vjepa',
|
| 5 |
+
'cpus_per_task': 16,
|
| 6 |
+
'data': { 'batch_size': 64,
|
| 7 |
+
'crop_size': 224,
|
| 8 |
+
'dataset_fpcs': [8, 8],
|
| 9 |
+
'dataset_type': 'VideoDataset',
|
| 10 |
+
'datasets': [ '/local/dcanezil/data/kinetics/train.csv',
|
| 11 |
+
'/local/dcanezil/data/ssv2/train.csv'],
|
| 12 |
+
'datasets_weights': [0.65, 0.35],
|
| 13 |
+
'fps': 2,
|
| 14 |
+
'num_workers': 4,
|
| 15 |
+
'patch_size': 14,
|
| 16 |
+
'persistent_workers': True,
|
| 17 |
+
'pin_mem': True,
|
| 18 |
+
'tubelet_size': 1},
|
| 19 |
+
'data_aug': { 'auto_augment': False,
|
| 20 |
+
'motion_shift': False,
|
| 21 |
+
'random_resize_aspect_ratio': [0.75, 1.35],
|
| 22 |
+
'random_resize_scale': [0.3, 1.0],
|
| 23 |
+
'reprob': 0.0},
|
| 24 |
+
'folder': '/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow',
|
| 25 |
+
'loss': {'loss_exp': 1.0},
|
| 26 |
+
'mask': [ { 'aspect_ratio': [0.75, 1.5],
|
| 27 |
+
'full_complement': False,
|
| 28 |
+
'max_keep': None,
|
| 29 |
+
'max_temporal_keep': 1.0,
|
| 30 |
+
'num_blocks': 8,
|
| 31 |
+
'spatial_scale': [0.05, 0.05],
|
| 32 |
+
'temporal_scale': [1.0, 1.0]},
|
| 33 |
+
{ 'aspect_ratio': [0.75, 1.5],
|
| 34 |
+
'full_complement': False,
|
| 35 |
+
'max_keep': None,
|
| 36 |
+
'max_temporal_keep': 1.0,
|
| 37 |
+
'num_blocks': 2,
|
| 38 |
+
'spatial_scale': [0.4, 0.4],
|
| 39 |
+
'temporal_scale': [1.0, 1.0]}],
|
| 40 |
+
'mem_per_gpu': '40G',
|
| 41 |
+
'meta': { 'dtype': 'bfloat16',
|
| 42 |
+
'knn_eval_epoch0': True,
|
| 43 |
+
'knn_eval_freq': 5,
|
| 44 |
+
'knn_eval_presets': [ { 'config': { 'batch_size': 8,
|
| 45 |
+
'dataset_train': '/local/dcanezil/data/ssv2/train_coarse10.csv',
|
| 46 |
+
'dataset_val': '/local/dcanezil/data/ssv2/val_coarse10.csv',
|
| 47 |
+
'eval_videos_per_class': 100,
|
| 48 |
+
'linear_probe': True,
|
| 49 |
+
'num_workers': 2,
|
| 50 |
+
'pool_type': 'temporal_concat',
|
| 51 |
+
'train_videos_per_class': 500},
|
| 52 |
+
'preset': 'ssv2_coarse10'}],
|
| 53 |
+
'load_checkpoint': True,
|
| 54 |
+
'read_checkpoint': None,
|
| 55 |
+
'save_every_freq': -1,
|
| 56 |
+
'seed': 239,
|
| 57 |
+
'use_sdpa': True,
|
| 58 |
+
'use_wandb': True,
|
| 59 |
+
'wandb_project': 'vjepa_ablation'},
|
| 60 |
+
'metrics': {'sigreg': {}, 'std': {}},
|
| 61 |
+
'model': { 'class_token': True,
|
| 62 |
+
'freeze_backbone': False,
|
| 63 |
+
'is_causal': True,
|
| 64 |
+
'model_name': 'vit_large_patch14_dinov2_lvd142m',
|
| 65 |
+
'pe_type': '2d',
|
| 66 |
+
'pred_depth': 12,
|
| 67 |
+
'pred_embed_dim': 384,
|
| 68 |
+
'pred_num_heads': 6,
|
| 69 |
+
'predictor': 'v2_cross',
|
| 70 |
+
'stem_type': '2d',
|
| 71 |
+
'target_embed_dim': 1024,
|
| 72 |
+
'target_kind': 'ema',
|
| 73 |
+
'uniform_power': True,
|
| 74 |
+
'use_activation_checkpointing': True,
|
| 75 |
+
'use_alibi': False,
|
| 76 |
+
'use_lora': False,
|
| 77 |
+
'use_mask_tokens': True,
|
| 78 |
+
'use_rope': False,
|
| 79 |
+
'use_sdpa': True,
|
| 80 |
+
'zero_init_mask_tokens': True},
|
| 81 |
+
'nodes': 1,
|
| 82 |
+
'optimization': { 'clip_grad': None,
|
| 83 |
+
'ema': [1.0, 0.99925],
|
| 84 |
+
'ema_hold_epochs': 0,
|
| 85 |
+
'ema_ramp_epochs': 5,
|
| 86 |
+
'encoder_layer_decay': 0.9,
|
| 87 |
+
'epochs': 90,
|
| 88 |
+
'final_lr': 0.0001,
|
| 89 |
+
'final_weight_decay': 0.04,
|
| 90 |
+
'ipe': 300,
|
| 91 |
+
'ipe_scale': 1.0,
|
| 92 |
+
'lr': 0.0001,
|
| 93 |
+
'predictor_lr_scale': 3.0,
|
| 94 |
+
'start_lr': 1e-05,
|
| 95 |
+
'warmup': 2,
|
| 96 |
+
'weight_decay': 0.04},
|
| 97 |
+
'tasks_per_node': 3}
|
| 98 |
+
INFO:root:Running pre-training of app: vjepa
|
| 99 |
+
[INFO ][2026-07-07 13:20:22][app.vjepa.train ][main ] which_dtype='bfloat16'
|
| 100 |
+
[INFO ][2026-07-07 13:20:22][app.vjepa.train ][main ] Disabling persistent_workers (incompatible with KNN eval)
|
| 101 |
+
[INFO ][2026-07-07 13:20:22][app.vjepa.train ][main ] NCCL_SOCKET_IFNAME=None
|
| 102 |
+
[INFO ][2026-07-07 13:20:25][app.vjepa.train ][main ] Initialized (rank/world-size) 2/3, tasks_per_node=3
|
| 103 |
+
[WARNING ][2026-07-07 13:21:13][experiments.stmodels.vision_transformers_v3][_build_dinov2_vit ] DINOv2 is an image model — fwd_chunk is unset so all frames will be processed jointly. Set fwd_chunk=1 for per-frame (standard) processing.
|
| 104 |
+
[INFO ][2026-07-07 13:21:22][experiments.stmodels.vision_transformers_v3][_build_dinov2_vit ] Loading pretrained weights for vit_large_patch14_dinov2_lvd142m from timm (vit_large_patch14_dinov2)
|
| 105 |
+
[INFO ][2026-07-07 13:21:25][timm.models._builder][load_pretrained ] Loading pretrained weights from Hugging Face hub (timm/vit_large_patch14_dinov2.lvd142m)
|
| 106 |
+
[INFO ][2026-07-07 13:21:25][timm.models._hub ][load_state_dict_from_hf ] [timm/vit_large_patch14_dinov2.lvd142m] Safe alternative available for 'pytorch_model.bin' (as 'model.safetensors'). Loading weights using safetensors.
|
| 107 |
+
[INFO ][2026-07-07 13:21:27][timm.layers.pos_embed][resample_abs_pos_embed ] Resized position embedding: (37, 37) to (16, 16).
|
| 108 |
+
[INFO ][2026-07-07 13:21:32][root ][init_video_model ] ViTMultiSeqWrapper(
|
| 109 |
+
(backbone): VisionTransformer(
|
| 110 |
+
(patch_embed): PatchEmbed3D(
|
| 111 |
+
(proj): Conv3d(3, 1024, kernel_size=(1, 14, 14), stride=(1, 14, 14))
|
| 112 |
+
)
|
| 113 |
+
(blocks): ModuleList(
|
| 114 |
+
(0-23): 24 x Block(
|
| 115 |
+
(norm1): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 116 |
+
(norm2): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 117 |
+
(drop_path1): Identity()
|
| 118 |
+
(drop_path2): Identity()
|
| 119 |
+
(attn): Attention(
|
| 120 |
+
(qkv): Linear(in_features=1024, out_features=3072, bias=True)
|
| 121 |
+
(attn_drop): Dropout(p=0.0, inplace=False)
|
| 122 |
+
(proj): Linear(in_features=1024, out_features=1024, bias=True)
|
| 123 |
+
(proj_drop): Dropout(p=0.0, inplace=False)
|
| 124 |
+
)
|
| 125 |
+
(mlp): MLP(
|
| 126 |
+
(fc1): Linear(in_features=1024, out_features=4096, bias=True)
|
| 127 |
+
(act): GELU(approximate='none')
|
| 128 |
+
(fc2): Linear(in_features=4096, out_features=1024, bias=True)
|
| 129 |
+
(drop): Dropout(p=0.0, inplace=False)
|
| 130 |
+
)
|
| 131 |
+
)
|
| 132 |
+
)
|
| 133 |
+
(norm): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 134 |
+
)
|
| 135 |
+
)
|
| 136 |
+
[INFO ][2026-07-07 13:21:32][root ][init_video_model ] VJEPAPredictorMultiSeqWrapper(
|
| 137 |
+
(backbone): PredictorV2(
|
| 138 |
+
(predictor_embed): Linear(in_features=1024, out_features=384, bias=True)
|
| 139 |
+
(mask_tokens): ParameterList(
|
| 140 |
+
(0): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
|
| 141 |
+
(1): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
|
| 142 |
+
(2): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
|
| 143 |
+
(3): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
|
| 144 |
+
)
|
| 145 |
+
(predictor_blocks): ModuleList(
|
| 146 |
+
(0-11): 12 x Block(
|
| 147 |
+
(residual1): EfficientResidual(
|
| 148 |
+
(norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
|
| 149 |
+
(fn): Attention(
|
| 150 |
+
(q_proj): Linear(in_features=384, out_features=384, bias=False)
|
| 151 |
+
(k_proj): Linear(in_features=384, out_features=384, bias=False)
|
| 152 |
+
(v_proj): Linear(in_features=384, out_features=384, bias=False)
|
| 153 |
+
(proj): Linear(in_features=384, out_features=384, bias=False)
|
| 154 |
+
(rope): Rope()
|
| 155 |
+
)
|
| 156 |
+
)
|
| 157 |
+
(residual2): EfficientResidual(
|
| 158 |
+
(norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
|
| 159 |
+
(fn): MLP(
|
| 160 |
+
(fc1): Linear(in_features=384, out_features=1536, bias=False)
|
| 161 |
+
(act): GELU(approximate='none')
|
| 162 |
+
(fc2): Linear(in_features=1536, out_features=384, bias=False)
|
| 163 |
+
(drop): Dropout(p=0.0, inplace=False)
|
| 164 |
+
)
|
| 165 |
+
)
|
| 166 |
+
)
|
| 167 |
+
)
|
| 168 |
+
(predictor_norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
|
| 169 |
+
(predictor_proj): Linear(in_features=384, out_features=1024, bias=True)
|
| 170 |
+
)
|
| 171 |
+
)
|
| 172 |
+
[INFO ][2026-07-07 13:21:32][root ][init_video_model ] ViTMultiSeqWrapper(
|
| 173 |
+
(backbone): VisionTransformer(
|
| 174 |
+
(patch_embed): PatchEmbed3D(
|
| 175 |
+
(proj): Conv3d(3, 1024, kernel_size=(1, 14, 14), stride=(1, 14, 14))
|
| 176 |
+
)
|
| 177 |
+
(blocks): ModuleList(
|
| 178 |
+
(0-23): 24 x Block(
|
| 179 |
+
(norm1): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 180 |
+
(norm2): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 181 |
+
(drop_path1): Identity()
|
| 182 |
+
(drop_path2): Identity()
|
| 183 |
+
(attn): Attention(
|
| 184 |
+
(qkv): Linear(in_features=1024, out_features=3072, bias=True)
|
| 185 |
+
(attn_drop): Dropout(p=0.0, inplace=False)
|
| 186 |
+
(proj): Linear(in_features=1024, out_features=1024, bias=True)
|
| 187 |
+
(proj_drop): Dropout(p=0.0, inplace=False)
|
| 188 |
+
)
|
| 189 |
+
(mlp): MLP(
|
| 190 |
+
(fc1): Linear(in_features=1024, out_features=4096, bias=True)
|
| 191 |
+
(act): GELU(approximate='none')
|
| 192 |
+
(fc2): Linear(in_features=4096, out_features=1024, bias=True)
|
| 193 |
+
(drop): Dropout(p=0.0, inplace=False)
|
| 194 |
+
)
|
| 195 |
+
)
|
| 196 |
+
)
|
| 197 |
+
(norm): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 198 |
+
)
|
| 199 |
+
)
|
| 200 |
+
[INFO ][2026-07-07 13:21:32][root ][init_video_model ] Encoder number of parameters: 302964736
|
| 201 |
+
[INFO ][2026-07-07 13:21:32][root ][init_video_model ] Predictor number of parameters: 22042240
|
| 202 |
+
[INFO ][2026-07-07 13:21:32][root ][init_video_model ] Target encoder number of parameters: 302964736
|
| 203 |
+
[INFO ][2026-07-07 13:21:33][root ][make_videodataset ] VideoDataset dataset created
|
| 204 |
+
[INFO ][2026-07-07 13:21:33][WeightedSampler ][__init__ ] Using DistributedWeightedSampler with rank 2 / 3
|
| 205 |
+
[INFO ][2026-07-07 13:21:33][root ][make_videodataset ] VideoDataset unsupervised data loader created
|
| 206 |
+
[INFO ][2026-07-07 13:21:33][app.vjepa.train ][main ] iterations per epoch/dataset length: 300/2103
|
| 207 |
+
[INFO ][2026-07-07 13:21:33][app.vjepa.train ][main ] Wrapping models in DDP (rank 2)...
|
| 208 |
+
[INFO ][2026-07-07 13:21:34][root ][load_checkpoint ] Loading checkpoint from /var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/latest.pt
|
| 209 |
+
[INFO ][2026-07-07 13:21:44][root ][load_checkpoint ] loaded pretrained encoder from epoch 60 with msg: <All keys matched successfully>
|
| 210 |
+
[INFO ][2026-07-07 13:21:44][root ][load_checkpoint ] loaded pretrained predictor from epoch 60 with msg: <All keys matched successfully>
|
| 211 |
+
[INFO ][2026-07-07 13:21:45][root ][load_checkpoint ] loaded pretrained target encoder from epoch 60 with msg: <All keys matched successfully>
|
| 212 |
+
[INFO ][2026-07-07 13:21:45][root ][load_checkpoint ] loaded optimizers from epoch 60
|
| 213 |
+
[INFO ][2026-07-07 13:21:45][root ][load_checkpoint ] read-path: /var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/latest.pt
|
| 214 |
+
[INFO ][2026-07-07 13:23:07][app.vjepa.train ][main ] Initializing loader...
|
| 215 |
+
submitit INFO (2026-07-08 12:54:51,527) - Job completed successfully
|
| 216 |
+
[INFO ][2026-07-08 12:54:51][submitit ][process_job ] Job completed successfully
|
| 217 |
+
submitit INFO (2026-07-08 12:54:51,529) - Exiting after successful completion
|
| 218 |
+
[INFO ][2026-07-08 12:54:51][submitit ][process_job ] Exiting after successful completion
|
inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_27036/27036_0_log.err
ADDED
|
@@ -0,0 +1,161 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
/var/scratch/dcanezil/.cache/uv/virtualenvs/vd/lib/python3.12/site-packages/timm/models/layers/__init__.py:49: FutureWarning: Importing from timm.models.layers is deprecated, please import via timm.layers
|
| 2 |
+
warnings.warn(f"Importing from {__name__} is deprecated, please import via timm.layers", FutureWarning)
|
| 3 |
+
wandb: Currently logged in as: dgcnz (uvjepa) to https://api.wandb.ai. Use `wandb login --relogin` to force relogin
|
| 4 |
+
wandb: setting up run vfclercb
|
| 5 |
+
wandb: Tracking run with wandb version 0.23.1
|
| 6 |
+
wandb: Run data is saved locally in /var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/wandb/run-20260708_144006-vfclercb
|
| 7 |
+
wandb: Run `wandb offline` to turn off syncing.
|
| 8 |
+
wandb: Resuming run d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow
|
| 9 |
+
wandb: ⭐️ View project at https://wandb.ai/uvjepa/vjepa_ablation
|
| 10 |
+
wandb: 🚀 View run at https://wandb.ai/uvjepa/vjepa_ablation/runs/vfclercb
|
| 11 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/utils.py:888: FutureWarning: `torch.cuda.amp.GradScaler(args...)` is deprecated. Please use `torch.amp.GradScaler('cuda', args...)` instead.
|
| 12 |
+
scaler = torch.cuda.amp.GradScaler() if mixed_precision else None
|
| 13 |
+
/var/scratch/dcanezil/.cache/uv/virtualenvs/vd/lib/python3.12/site-packages/torch/distributed/c10d_logger.py:83: UserWarning: barrier(): using the device under current context. You can specify `device_id` in `init_process_group` to mute this warning.
|
| 14 |
+
return func(*args, **kwargs)
|
| 15 |
+
wandb: WARNING Tried to log to step 0 that is less than the current step 27001. Steps must be monotonically increasing, so this data will be ignored. See https://wandb.me/define-metric to log data out of order.
|
| 16 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 17 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 18 |
+
wandb: WARNING Tried to log to step 27000 that is less than the current step 27001. Steps must be monotonically increasing, so this data will be ignored. See https://wandb.me/define-metric to log data out of order.
|
| 19 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_poker/6-DyF8umej8_000097_000107.avi'
|
| 20 |
+
warnings.warn(f"video path not found {fname=}")
|
| 21 |
+
[15:19:04] /github/workspace/src/video/video_reader.cc:83: ERROR opening: /local/dcanezil/data/kinetics/VideoData/climbing_tree/QX5IzK4iFCM_000002_000012.avi, Invalid data found when processing input
|
| 22 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_flute/co50KUHacYw_000005_000015.avi'
|
| 23 |
+
warnings.warn(f"video path not found {fname=}")
|
| 24 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_poker/-jGxlNQKkeo_000092_000102.avi'
|
| 25 |
+
warnings.warn(f"video path not found {fname=}")
|
| 26 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/waxing_legs/cqUusXBODuw_001046_001056.avi'
|
| 27 |
+
warnings.warn(f"video path not found {fname=}")
|
| 28 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/surfing_water/5l5Pdd96Pao_000161_000171.avi'
|
| 29 |
+
warnings.warn(f"video path not found {fname=}")
|
| 30 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/water_skiing/4bhvIHWADVU_000022_000032.avi'
|
| 31 |
+
warnings.warn(f"video path not found {fname=}")
|
| 32 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/eating_carrots/eiZ8Hzc7FPU_000080_000090.avi'
|
| 33 |
+
warnings.warn(f"video path not found {fname=}")
|
| 34 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/picking_fruit/3VvkoFtPCCU_000045_000055.avi'
|
| 35 |
+
warnings.warn(f"video path not found {fname=}")
|
| 36 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/making_tea/mtYFNsRcxY4_000063_000073.avi'
|
| 37 |
+
warnings.warn(f"video path not found {fname=}")
|
| 38 |
+
/var/scratch/dcanezil/.cache/uv/virtualenvs/vd/lib/python3.12/site-packages/torch/distributed/c10d_logger.py:83: UserWarning: barrier(): using the device under current context. You can specify `device_id` in `init_process_group` to mute this warning.
|
| 39 |
+
return func(*args, **kwargs)
|
| 40 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 41 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 42 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/pumping_fist/wKsuyr6Xk30_000094_000104.avi'
|
| 43 |
+
warnings.warn(f"video path not found {fname=}")
|
| 44 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/dancing_macarena/uGQgxFHemyA_000059_000069.avi'
|
| 45 |
+
warnings.warn(f"video path not found {fname=}")
|
| 46 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/water_skiing/4bhvIHWADVU_000022_000032.avi'
|
| 47 |
+
warnings.warn(f"video path not found {fname=}")
|
| 48 |
+
[21:50:35] /github/workspace/src/video/video_reader.cc:83: ERROR opening: /local/dcanezil/data/kinetics/VideoData/climbing_tree/QX5IzK4iFCM_000002_000012.avi, Invalid data found when processing input
|
| 49 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/waxing_legs/cqUusXBODuw_001046_001056.avi'
|
| 50 |
+
warnings.warn(f"video path not found {fname=}")
|
| 51 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_flute/co50KUHacYw_000005_000015.avi'
|
| 52 |
+
warnings.warn(f"video path not found {fname=}")
|
| 53 |
+
/var/scratch/dcanezil/.cache/uv/virtualenvs/vd/lib/python3.12/site-packages/torch/distributed/c10d_logger.py:83: UserWarning: barrier(): using the device under current context. You can specify `device_id` in `init_process_group` to mute this warning.
|
| 54 |
+
return func(*args, **kwargs)
|
| 55 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 56 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 57 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/sailing/ZolXG_-RgIs_000970_000980.avi'
|
| 58 |
+
warnings.warn(f"video path not found {fname=}")
|
| 59 |
+
[22:45:51] /github/workspace/src/video/video_reader.cc:83: ERROR opening: /local/dcanezil/data/kinetics/VideoData/long_jump/jsYiitkbr5w_000004_000014.avi, Invalid data found when processing input
|
| 60 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/eating_carrots/eiZ8Hzc7FPU_000080_000090.avi'
|
| 61 |
+
warnings.warn(f"video path not found {fname=}")
|
| 62 |
+
[00:28:09] /github/workspace/src/video/video_reader.cc:83: ERROR opening: /local/dcanezil/data/kinetics/VideoData/climbing_tree/QX5IzK4iFCM_000002_000012.avi, Invalid data found when processing input
|
| 63 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_poker/-jGxlNQKkeo_000092_000102.avi'
|
| 64 |
+
warnings.warn(f"video path not found {fname=}")
|
| 65 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/waxing_legs/cqUusXBODuw_001046_001056.avi'
|
| 66 |
+
warnings.warn(f"video path not found {fname=}")
|
| 67 |
+
[01:44:13] /github/workspace/src/video/video_reader.cc:83: ERROR opening: /local/dcanezil/data/kinetics/VideoData/long_jump/jsYiitkbr5w_000004_000014.avi, Invalid data found when processing input
|
| 68 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_saxophone/Zs0_2GMEPXo_000054_000064.avi'
|
| 69 |
+
warnings.warn(f"video path not found {fname=}")
|
| 70 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/surfing_water/5l5Pdd96Pao_000161_000171.avi'
|
| 71 |
+
warnings.warn(f"video path not found {fname=}")
|
| 72 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/frying_vegetables/Nr-zCXgW2Sw_000161_000171.avi'
|
| 73 |
+
warnings.warn(f"video path not found {fname=}")
|
| 74 |
+
/var/scratch/dcanezil/.cache/uv/virtualenvs/vd/lib/python3.12/site-packages/torch/distributed/c10d_logger.py:83: UserWarning: barrier(): using the device under current context. You can specify `device_id` in `init_process_group` to mute this warning.
|
| 75 |
+
return func(*args, **kwargs)
|
| 76 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 77 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 78 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/waxing_legs/cqUusXBODuw_001046_001056.avi'
|
| 79 |
+
warnings.warn(f"video path not found {fname=}")
|
| 80 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/waxing_legs/cqUusXBODuw_001046_001056.avi'
|
| 81 |
+
warnings.warn(f"video path not found {fname=}")
|
| 82 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/making_tea/mtYFNsRcxY4_000063_000073.avi'
|
| 83 |
+
warnings.warn(f"video path not found {fname=}")
|
| 84 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/picking_fruit/3VvkoFtPCCU_000045_000055.avi'
|
| 85 |
+
warnings.warn(f"video path not found {fname=}")
|
| 86 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/surfing_water/5l5Pdd96Pao_000161_000171.avi'
|
| 87 |
+
warnings.warn(f"video path not found {fname=}")
|
| 88 |
+
/var/scratch/dcanezil/.cache/uv/virtualenvs/vd/lib/python3.12/site-packages/torch/distributed/c10d_logger.py:83: UserWarning: barrier(): using the device under current context. You can specify `device_id` in `init_process_group` to mute this warning.
|
| 89 |
+
return func(*args, **kwargs)
|
| 90 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 91 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 92 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/eating_carrots/eiZ8Hzc7FPU_000080_000090.avi'
|
| 93 |
+
warnings.warn(f"video path not found {fname=}")
|
| 94 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_monopoly/NLL667uPWVA_000066_000076.avi'
|
| 95 |
+
warnings.warn(f"video path not found {fname=}")
|
| 96 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/windsurfing/i-gzh_BPDa8_000154_000164.avi'
|
| 97 |
+
warnings.warn(f"video path not found {fname=}")
|
| 98 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/baking_cookies/Q3XGZmqk1Q0_000296_000306.avi'
|
| 99 |
+
warnings.warn(f"video path not found {fname=}")
|
| 100 |
+
[08:00:09] /github/workspace/src/video/video_reader.cc:83: ERROR opening: /local/dcanezil/data/kinetics/VideoData/climbing_tree/QX5IzK4iFCM_000002_000012.avi, Invalid data found when processing input
|
| 101 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_poker/-jGxlNQKkeo_000092_000102.avi'
|
| 102 |
+
warnings.warn(f"video path not found {fname=}")
|
| 103 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/surfing_water/5l5Pdd96Pao_000161_000171.avi'
|
| 104 |
+
warnings.warn(f"video path not found {fname=}")
|
| 105 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_poker/6-DyF8umej8_000097_000107.avi'
|
| 106 |
+
warnings.warn(f"video path not found {fname=}")
|
| 107 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/bowling/f4eb0wOlspM_000053_000063.avi'
|
| 108 |
+
warnings.warn(f"video path not found {fname=}")
|
| 109 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_monopoly/NLL667uPWVA_000066_000076.avi'
|
| 110 |
+
warnings.warn(f"video path not found {fname=}")
|
| 111 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/frying_vegetables/Nr-zCXgW2Sw_000161_000171.avi'
|
| 112 |
+
warnings.warn(f"video path not found {fname=}")
|
| 113 |
+
/var/scratch/dcanezil/.cache/uv/virtualenvs/vd/lib/python3.12/site-packages/torch/distributed/c10d_logger.py:83: UserWarning: barrier(): using the device under current context. You can specify `device_id` in `init_process_group` to mute this warning.
|
| 114 |
+
return func(*args, **kwargs)
|
| 115 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 116 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 117 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/dancing_macarena/uGQgxFHemyA_000059_000069.avi'
|
| 118 |
+
warnings.warn(f"video path not found {fname=}")
|
| 119 |
+
[11:51:53] /github/workspace/src/video/video_reader.cc:83: ERROR opening: /local/dcanezil/data/kinetics/VideoData/trimming_trees/T6dqaZioaXs_000004_000014.avi, Invalid data found when processing input
|
| 120 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_saxophone/Zs0_2GMEPXo_000054_000064.avi'
|
| 121 |
+
warnings.warn(f"video path not found {fname=}")
|
| 122 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/windsurfing/i-gzh_BPDa8_000154_000164.avi'
|
| 123 |
+
warnings.warn(f"video path not found {fname=}")
|
| 124 |
+
[14:05:19] /github/workspace/src/video/video_reader.cc:83: ERROR opening: /local/dcanezil/data/kinetics/VideoData/trimming_trees/T6dqaZioaXs_000004_000014.avi, Invalid data found when processing input
|
| 125 |
+
/var/scratch/dcanezil/.cache/uv/virtualenvs/vd/lib/python3.12/site-packages/torch/distributed/c10d_logger.py:83: UserWarning: barrier(): using the device under current context. You can specify `device_id` in `init_process_group` to mute this warning.
|
| 126 |
+
return func(*args, **kwargs)
|
| 127 |
+
wandb: updating run metadata
|
| 128 |
+
wandb: updating run metadata; uploading output.log; uploading wandb-summary.json
|
| 129 |
+
wandb: uploading output.log; uploading config.yaml
|
| 130 |
+
wandb:
|
| 131 |
+
wandb: Run history:
|
| 132 |
+
wandb: epoch ▁▁▁▁▂▂▂▂▂▂▂▂▃▃▃▃▄▄▄▄▄▄▄▄▅▅▅▅▅▆▆▆▆▇▇▇▇▇▇█
|
| 133 |
+
wandb: epoch/avg_data_time_ms ▆█▇▂▇▆▄█▇▄▄▅▇▅▆█▆▆▇▆▄▅▂▅▇▃▁▅▂▆
|
| 134 |
+
wandb: epoch/avg_gpu_time_ms █▆▃▃▇▃█▆▅▃▃▂▃▆▇▂▂▄▆▆▇▅▂█▆▆▂▁▅▄
|
| 135 |
+
wandb: epoch/avg_iter_time_ms █▆▃▃▇▃█▆▅▃▃▂▃▆▇▂▂▄▆▆▇▅▂█▆▆▂▁▅▄
|
| 136 |
+
wandb: epoch/avg_loss ▁▂▄▅▄▅▄▅▆▇▇█▇▆▆█▇▅▄▅▅▅▅▂▃▄▄▂▂▂
|
| 137 |
+
wandb: knn/ssv2_coarse10_temporal_concat_top1 ▃▁▄█▆▅
|
| 138 |
+
wandb: knn/ssv2_coarse10_temporal_concat_top5 ▇█▇▄▅▁
|
| 139 |
+
wandb: linear/ssv2_coarse10_temporal_concat_top1 ▄▁▂▆█▅
|
| 140 |
+
wandb: linear/ssv2_coarse10_temporal_concat_top5 ▁▆▆▇█▇
|
| 141 |
+
wandb: train/ema/cos_sim ▄▇▅▆▄▃▂▂▅▅▇▄▄▆▃▄▃▂▄▃▄▃▃▅█▆▂▄▃▄▆▅▆▂▅▄▁▄▆█
|
| 142 |
+
wandb: +34 ...
|
| 143 |
+
wandb:
|
| 144 |
+
wandb: Run summary:
|
| 145 |
+
wandb: epoch 120
|
| 146 |
+
wandb: epoch/avg_data_time_ms 15.90538
|
| 147 |
+
wandb: epoch/avg_gpu_time_ms 9191.97061
|
| 148 |
+
wandb: epoch/avg_iter_time_ms 9220.20652
|
| 149 |
+
wandb: epoch/avg_loss 0.54899
|
| 150 |
+
wandb: knn/ssv2_coarse10_temporal_concat_top1 46.98795
|
| 151 |
+
wandb: knn/ssv2_coarse10_temporal_concat_top5 86.14458
|
| 152 |
+
wandb: linear/ssv2_coarse10_temporal_concat_top1 59.13655
|
| 153 |
+
wandb: linear/ssv2_coarse10_temporal_concat_top5 93.4739
|
| 154 |
+
wandb: train/ema/cos_sim 0.75001
|
| 155 |
+
wandb: +34 ...
|
| 156 |
+
wandb:
|
| 157 |
+
wandb: 🚀 View run d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow at: https://wandb.ai/uvjepa/vjepa_ablation/runs/vfclercb
|
| 158 |
+
wandb: ⭐️ View project at: https://wandb.ai/uvjepa/vjepa_ablation
|
| 159 |
+
wandb: Synced 5 W&B file(s), 0 media file(s), 0 artifact file(s) and 0 other file(s)
|
| 160 |
+
wandb: Find logs at: /var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/wandb/run-20260708_144006-vfclercb/logs
|
| 161 |
+
[rank0]:[W709 14:15:10.545157731 ProcessGroupNCCL.cpp:1553] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())
|
inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_27036/27036_0_log.out
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_27036/27036_1_log.err
ADDED
|
@@ -0,0 +1,102 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
/var/scratch/dcanezil/.cache/uv/virtualenvs/vd/lib/python3.12/site-packages/timm/models/layers/__init__.py:49: FutureWarning: Importing from timm.models.layers is deprecated, please import via timm.layers
|
| 2 |
+
warnings.warn(f"Importing from {__name__} is deprecated, please import via timm.layers", FutureWarning)
|
| 3 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/utils.py:888: FutureWarning: `torch.cuda.amp.GradScaler(args...)` is deprecated. Please use `torch.amp.GradScaler('cuda', args...)` instead.
|
| 4 |
+
scaler = torch.cuda.amp.GradScaler() if mixed_precision else None
|
| 5 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 6 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 7 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/waxing_legs/cqUusXBODuw_001046_001056.avi'
|
| 8 |
+
warnings.warn(f"video path not found {fname=}")
|
| 9 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/water_skiing/4bhvIHWADVU_000022_000032.avi'
|
| 10 |
+
warnings.warn(f"video path not found {fname=}")
|
| 11 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/water_skiing/4bhvIHWADVU_000022_000032.avi'
|
| 12 |
+
warnings.warn(f"video path not found {fname=}")
|
| 13 |
+
[16:14:19] /github/workspace/src/video/video_reader.cc:83: ERROR opening: /local/dcanezil/data/kinetics/VideoData/climbing_tree/QX5IzK4iFCM_000002_000012.avi, Invalid data found when processing input
|
| 14 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/making_tea/mtYFNsRcxY4_000063_000073.avi'
|
| 15 |
+
warnings.warn(f"video path not found {fname=}")
|
| 16 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_saxophone/Zs0_2GMEPXo_000054_000064.avi'
|
| 17 |
+
warnings.warn(f"video path not found {fname=}")
|
| 18 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/baking_cookies/Q3XGZmqk1Q0_000296_000306.avi'
|
| 19 |
+
warnings.warn(f"video path not found {fname=}")
|
| 20 |
+
[18:11:00] /github/workspace/src/video/video_reader.cc:83: ERROR opening: /local/dcanezil/data/kinetics/VideoData/long_jump/jsYiitkbr5w_000004_000014.avi, Invalid data found when processing input
|
| 21 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/bowling/f4eb0wOlspM_000053_000063.avi'
|
| 22 |
+
warnings.warn(f"video path not found {fname=}")
|
| 23 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 24 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 25 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/windsurfing/i-gzh_BPDa8_000154_000164.avi'
|
| 26 |
+
warnings.warn(f"video path not found {fname=}")
|
| 27 |
+
[20:03:38] /github/workspace/src/video/video_reader.cc:83: ERROR opening: /local/dcanezil/data/kinetics/VideoData/climbing_tree/QX5IzK4iFCM_000002_000012.avi, Invalid data found when processing input
|
| 28 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/bowling/f4eb0wOlspM_000053_000063.avi'
|
| 29 |
+
warnings.warn(f"video path not found {fname=}")
|
| 30 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/eating_carrots/eiZ8Hzc7FPU_000080_000090.avi'
|
| 31 |
+
warnings.warn(f"video path not found {fname=}")
|
| 32 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/frying_vegetables/Nr-zCXgW2Sw_000161_000171.avi'
|
| 33 |
+
warnings.warn(f"video path not found {fname=}")
|
| 34 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_saxophone/Zs0_2GMEPXo_000054_000064.avi'
|
| 35 |
+
warnings.warn(f"video path not found {fname=}")
|
| 36 |
+
[22:11:43] /github/workspace/src/video/video_reader.cc:83: ERROR opening: /local/dcanezil/data/kinetics/VideoData/climbing_tree/QX5IzK4iFCM_000002_000012.avi, Invalid data found when processing input
|
| 37 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 38 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 39 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_poker/6-DyF8umej8_000097_000107.avi'
|
| 40 |
+
warnings.warn(f"video path not found {fname=}")
|
| 41 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/water_skiing/4bhvIHWADVU_000022_000032.avi'
|
| 42 |
+
warnings.warn(f"video path not found {fname=}")
|
| 43 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/making_tea/mtYFNsRcxY4_000063_000073.avi'
|
| 44 |
+
warnings.warn(f"video path not found {fname=}")
|
| 45 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/windsurfing/i-gzh_BPDa8_000154_000164.avi'
|
| 46 |
+
warnings.warn(f"video path not found {fname=}")
|
| 47 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_poker/6-DyF8umej8_000097_000107.avi'
|
| 48 |
+
warnings.warn(f"video path not found {fname=}")
|
| 49 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/waxing_legs/cqUusXBODuw_001046_001056.avi'
|
| 50 |
+
warnings.warn(f"video path not found {fname=}")
|
| 51 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/robot_dancing/J1x0XlWa6HM_000147_000157.avi'
|
| 52 |
+
warnings.warn(f"video path not found {fname=}")
|
| 53 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/water_skiing/4bhvIHWADVU_000022_000032.avi'
|
| 54 |
+
warnings.warn(f"video path not found {fname=}")
|
| 55 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 56 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 57 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/surfing_water/5l5Pdd96Pao_000161_000171.avi'
|
| 58 |
+
warnings.warn(f"video path not found {fname=}")
|
| 59 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/pumping_fist/wKsuyr6Xk30_000094_000104.avi'
|
| 60 |
+
warnings.warn(f"video path not found {fname=}")
|
| 61 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/waxing_legs/cqUusXBODuw_001046_001056.avi'
|
| 62 |
+
warnings.warn(f"video path not found {fname=}")
|
| 63 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_poker/-jGxlNQKkeo_000092_000102.avi'
|
| 64 |
+
warnings.warn(f"video path not found {fname=}")
|
| 65 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/baking_cookies/Q3XGZmqk1Q0_000296_000306.avi'
|
| 66 |
+
warnings.warn(f"video path not found {fname=}")
|
| 67 |
+
[03:54:03] /github/workspace/src/video/video_reader.cc:83: ERROR opening: /local/dcanezil/data/kinetics/VideoData/climbing_tree/QX5IzK4iFCM_000002_000012.avi, Invalid data found when processing input
|
| 68 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/sailing/ZolXG_-RgIs_000970_000980.avi'
|
| 69 |
+
warnings.warn(f"video path not found {fname=}")
|
| 70 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/pumping_fist/wKsuyr6Xk30_000094_000104.avi'
|
| 71 |
+
warnings.warn(f"video path not found {fname=}")
|
| 72 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/dancing_macarena/uGQgxFHemyA_000059_000069.avi'
|
| 73 |
+
warnings.warn(f"video path not found {fname=}")
|
| 74 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 75 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 76 |
+
[06:54:57] /github/workspace/src/video/video_reader.cc:83: ERROR opening: /local/dcanezil/data/kinetics/VideoData/trimming_trees/T6dqaZioaXs_000004_000014.avi, Invalid data found when processing input
|
| 77 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/water_skiing/4bhvIHWADVU_000022_000032.avi'
|
| 78 |
+
warnings.warn(f"video path not found {fname=}")
|
| 79 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/eating_carrots/eiZ8Hzc7FPU_000080_000090.avi'
|
| 80 |
+
warnings.warn(f"video path not found {fname=}")
|
| 81 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/making_tea/mtYFNsRcxY4_000063_000073.avi'
|
| 82 |
+
warnings.warn(f"video path not found {fname=}")
|
| 83 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/surfing_water/5l5Pdd96Pao_000161_000171.avi'
|
| 84 |
+
warnings.warn(f"video path not found {fname=}")
|
| 85 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 86 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 87 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/robot_dancing/J1x0XlWa6HM_000147_000157.avi'
|
| 88 |
+
warnings.warn(f"video path not found {fname=}")
|
| 89 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/picking_fruit/3VvkoFtPCCU_000045_000055.avi'
|
| 90 |
+
warnings.warn(f"video path not found {fname=}")
|
| 91 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/water_skiing/4bhvIHWADVU_000022_000032.avi'
|
| 92 |
+
warnings.warn(f"video path not found {fname=}")
|
| 93 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/picking_fruit/3VvkoFtPCCU_000045_000055.avi'
|
| 94 |
+
warnings.warn(f"video path not found {fname=}")
|
| 95 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/making_tea/mtYFNsRcxY4_000063_000073.avi'
|
| 96 |
+
warnings.warn(f"video path not found {fname=}")
|
| 97 |
+
[11:37:52] /github/workspace/src/video/video_reader.cc:83: ERROR opening: /local/dcanezil/data/kinetics/VideoData/trimming_trees/T6dqaZioaXs_000004_000014.avi, Invalid data found when processing input
|
| 98 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_monopoly/NLL667uPWVA_000066_000076.avi'
|
| 99 |
+
warnings.warn(f"video path not found {fname=}")
|
| 100 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/pumping_fist/wKsuyr6Xk30_000094_000104.avi'
|
| 101 |
+
warnings.warn(f"video path not found {fname=}")
|
| 102 |
+
[rank1]:[W709 14:15:08.752473589 ProcessGroupNCCL.cpp:1553] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())
|
inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_27036/27036_1_log.out
ADDED
|
@@ -0,0 +1,218 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
submitit INFO (2026-07-08 14:39:49,900) - Starting with JobEnvironment(job_id=27036, hostname=node404, local_rank=1(3), node=0(1), global_rank=1(3))
|
| 2 |
+
submitit INFO (2026-07-08 14:39:49,900) - Loading pickle: /var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_27036/27036_submitted.pkl
|
| 3 |
+
INFO:root:loaded pretrain params...
|
| 4 |
+
{ 'app': 'vjepa',
|
| 5 |
+
'cpus_per_task': 16,
|
| 6 |
+
'data': { 'batch_size': 64,
|
| 7 |
+
'crop_size': 224,
|
| 8 |
+
'dataset_fpcs': [8, 8],
|
| 9 |
+
'dataset_type': 'VideoDataset',
|
| 10 |
+
'datasets': [ '/local/dcanezil/data/kinetics/train.csv',
|
| 11 |
+
'/local/dcanezil/data/ssv2/train.csv'],
|
| 12 |
+
'datasets_weights': [0.65, 0.35],
|
| 13 |
+
'fps': 2,
|
| 14 |
+
'num_workers': 4,
|
| 15 |
+
'patch_size': 14,
|
| 16 |
+
'persistent_workers': True,
|
| 17 |
+
'pin_mem': True,
|
| 18 |
+
'tubelet_size': 1},
|
| 19 |
+
'data_aug': { 'auto_augment': False,
|
| 20 |
+
'motion_shift': False,
|
| 21 |
+
'random_resize_aspect_ratio': [0.75, 1.35],
|
| 22 |
+
'random_resize_scale': [0.3, 1.0],
|
| 23 |
+
'reprob': 0.0},
|
| 24 |
+
'folder': '/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow',
|
| 25 |
+
'loss': {'loss_exp': 1.0},
|
| 26 |
+
'mask': [ { 'aspect_ratio': [0.75, 1.5],
|
| 27 |
+
'full_complement': False,
|
| 28 |
+
'max_keep': None,
|
| 29 |
+
'max_temporal_keep': 1.0,
|
| 30 |
+
'num_blocks': 8,
|
| 31 |
+
'spatial_scale': [0.05, 0.05],
|
| 32 |
+
'temporal_scale': [1.0, 1.0]},
|
| 33 |
+
{ 'aspect_ratio': [0.75, 1.5],
|
| 34 |
+
'full_complement': False,
|
| 35 |
+
'max_keep': None,
|
| 36 |
+
'max_temporal_keep': 1.0,
|
| 37 |
+
'num_blocks': 2,
|
| 38 |
+
'spatial_scale': [0.4, 0.4],
|
| 39 |
+
'temporal_scale': [1.0, 1.0]}],
|
| 40 |
+
'mem_per_gpu': '40G',
|
| 41 |
+
'meta': { 'dtype': 'bfloat16',
|
| 42 |
+
'knn_eval_epoch0': True,
|
| 43 |
+
'knn_eval_freq': 5,
|
| 44 |
+
'knn_eval_presets': [ { 'config': { 'batch_size': 8,
|
| 45 |
+
'dataset_train': '/local/dcanezil/data/ssv2/train_coarse10.csv',
|
| 46 |
+
'dataset_val': '/local/dcanezil/data/ssv2/val_coarse10.csv',
|
| 47 |
+
'eval_videos_per_class': 100,
|
| 48 |
+
'linear_probe': True,
|
| 49 |
+
'num_workers': 2,
|
| 50 |
+
'pool_type': 'temporal_concat',
|
| 51 |
+
'train_videos_per_class': 500},
|
| 52 |
+
'preset': 'ssv2_coarse10'}],
|
| 53 |
+
'load_checkpoint': True,
|
| 54 |
+
'read_checkpoint': None,
|
| 55 |
+
'save_every_freq': -1,
|
| 56 |
+
'seed': 239,
|
| 57 |
+
'use_sdpa': True,
|
| 58 |
+
'use_wandb': True,
|
| 59 |
+
'wandb_project': 'vjepa_ablation'},
|
| 60 |
+
'metrics': {'sigreg': {}, 'std': {}},
|
| 61 |
+
'model': { 'class_token': True,
|
| 62 |
+
'freeze_backbone': False,
|
| 63 |
+
'is_causal': True,
|
| 64 |
+
'model_name': 'vit_large_patch14_dinov2_lvd142m',
|
| 65 |
+
'pe_type': '2d',
|
| 66 |
+
'pred_depth': 12,
|
| 67 |
+
'pred_embed_dim': 384,
|
| 68 |
+
'pred_num_heads': 6,
|
| 69 |
+
'predictor': 'v2_cross',
|
| 70 |
+
'stem_type': '2d',
|
| 71 |
+
'target_embed_dim': 1024,
|
| 72 |
+
'target_kind': 'ema',
|
| 73 |
+
'uniform_power': True,
|
| 74 |
+
'use_activation_checkpointing': True,
|
| 75 |
+
'use_alibi': False,
|
| 76 |
+
'use_lora': False,
|
| 77 |
+
'use_mask_tokens': True,
|
| 78 |
+
'use_rope': False,
|
| 79 |
+
'use_sdpa': True,
|
| 80 |
+
'zero_init_mask_tokens': True},
|
| 81 |
+
'nodes': 1,
|
| 82 |
+
'optimization': { 'clip_grad': None,
|
| 83 |
+
'ema': [1.0, 0.99925],
|
| 84 |
+
'ema_hold_epochs': 0,
|
| 85 |
+
'ema_ramp_epochs': 5,
|
| 86 |
+
'encoder_layer_decay': 0.9,
|
| 87 |
+
'epochs': 120,
|
| 88 |
+
'final_lr': 0.0001,
|
| 89 |
+
'final_weight_decay': 0.04,
|
| 90 |
+
'ipe': 300,
|
| 91 |
+
'ipe_scale': 1.0,
|
| 92 |
+
'lr': 0.0001,
|
| 93 |
+
'predictor_lr_scale': 3.0,
|
| 94 |
+
'start_lr': 1e-05,
|
| 95 |
+
'warmup': 2,
|
| 96 |
+
'weight_decay': 0.04},
|
| 97 |
+
'tasks_per_node': 3}
|
| 98 |
+
INFO:root:Running pre-training of app: vjepa
|
| 99 |
+
[INFO ][2026-07-08 14:40:02][app.vjepa.train ][main ] which_dtype='bfloat16'
|
| 100 |
+
[INFO ][2026-07-08 14:40:02][app.vjepa.train ][main ] Disabling persistent_workers (incompatible with KNN eval)
|
| 101 |
+
[INFO ][2026-07-08 14:40:02][app.vjepa.train ][main ] NCCL_SOCKET_IFNAME=None
|
| 102 |
+
[INFO ][2026-07-08 14:40:05][app.vjepa.train ][main ] Initialized (rank/world-size) 1/3, tasks_per_node=3
|
| 103 |
+
[WARNING ][2026-07-08 14:40:37][experiments.stmodels.vision_transformers_v3][_build_dinov2_vit ] DINOv2 is an image model — fwd_chunk is unset so all frames will be processed jointly. Set fwd_chunk=1 for per-frame (standard) processing.
|
| 104 |
+
[INFO ][2026-07-08 14:40:44][experiments.stmodels.vision_transformers_v3][_build_dinov2_vit ] Loading pretrained weights for vit_large_patch14_dinov2_lvd142m from timm (vit_large_patch14_dinov2)
|
| 105 |
+
[INFO ][2026-07-08 14:40:47][timm.models._builder][load_pretrained ] Loading pretrained weights from Hugging Face hub (timm/vit_large_patch14_dinov2.lvd142m)
|
| 106 |
+
[INFO ][2026-07-08 14:40:48][timm.models._hub ][load_state_dict_from_hf ] [timm/vit_large_patch14_dinov2.lvd142m] Safe alternative available for 'pytorch_model.bin' (as 'model.safetensors'). Loading weights using safetensors.
|
| 107 |
+
[INFO ][2026-07-08 14:40:48][timm.layers.pos_embed][resample_abs_pos_embed ] Resized position embedding: (37, 37) to (16, 16).
|
| 108 |
+
[INFO ][2026-07-08 14:40:50][root ][init_video_model ] ViTMultiSeqWrapper(
|
| 109 |
+
(backbone): VisionTransformer(
|
| 110 |
+
(patch_embed): PatchEmbed3D(
|
| 111 |
+
(proj): Conv3d(3, 1024, kernel_size=(1, 14, 14), stride=(1, 14, 14))
|
| 112 |
+
)
|
| 113 |
+
(blocks): ModuleList(
|
| 114 |
+
(0-23): 24 x Block(
|
| 115 |
+
(norm1): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 116 |
+
(norm2): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 117 |
+
(drop_path1): Identity()
|
| 118 |
+
(drop_path2): Identity()
|
| 119 |
+
(attn): Attention(
|
| 120 |
+
(qkv): Linear(in_features=1024, out_features=3072, bias=True)
|
| 121 |
+
(attn_drop): Dropout(p=0.0, inplace=False)
|
| 122 |
+
(proj): Linear(in_features=1024, out_features=1024, bias=True)
|
| 123 |
+
(proj_drop): Dropout(p=0.0, inplace=False)
|
| 124 |
+
)
|
| 125 |
+
(mlp): MLP(
|
| 126 |
+
(fc1): Linear(in_features=1024, out_features=4096, bias=True)
|
| 127 |
+
(act): GELU(approximate='none')
|
| 128 |
+
(fc2): Linear(in_features=4096, out_features=1024, bias=True)
|
| 129 |
+
(drop): Dropout(p=0.0, inplace=False)
|
| 130 |
+
)
|
| 131 |
+
)
|
| 132 |
+
)
|
| 133 |
+
(norm): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 134 |
+
)
|
| 135 |
+
)
|
| 136 |
+
[INFO ][2026-07-08 14:40:50][root ][init_video_model ] VJEPAPredictorMultiSeqWrapper(
|
| 137 |
+
(backbone): PredictorV2(
|
| 138 |
+
(predictor_embed): Linear(in_features=1024, out_features=384, bias=True)
|
| 139 |
+
(mask_tokens): ParameterList(
|
| 140 |
+
(0): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
|
| 141 |
+
(1): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
|
| 142 |
+
(2): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
|
| 143 |
+
(3): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
|
| 144 |
+
)
|
| 145 |
+
(predictor_blocks): ModuleList(
|
| 146 |
+
(0-11): 12 x Block(
|
| 147 |
+
(residual1): EfficientResidual(
|
| 148 |
+
(norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
|
| 149 |
+
(fn): Attention(
|
| 150 |
+
(q_proj): Linear(in_features=384, out_features=384, bias=False)
|
| 151 |
+
(k_proj): Linear(in_features=384, out_features=384, bias=False)
|
| 152 |
+
(v_proj): Linear(in_features=384, out_features=384, bias=False)
|
| 153 |
+
(proj): Linear(in_features=384, out_features=384, bias=False)
|
| 154 |
+
(rope): Rope()
|
| 155 |
+
)
|
| 156 |
+
)
|
| 157 |
+
(residual2): EfficientResidual(
|
| 158 |
+
(norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
|
| 159 |
+
(fn): MLP(
|
| 160 |
+
(fc1): Linear(in_features=384, out_features=1536, bias=False)
|
| 161 |
+
(act): GELU(approximate='none')
|
| 162 |
+
(fc2): Linear(in_features=1536, out_features=384, bias=False)
|
| 163 |
+
(drop): Dropout(p=0.0, inplace=False)
|
| 164 |
+
)
|
| 165 |
+
)
|
| 166 |
+
)
|
| 167 |
+
)
|
| 168 |
+
(predictor_norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
|
| 169 |
+
(predictor_proj): Linear(in_features=384, out_features=1024, bias=True)
|
| 170 |
+
)
|
| 171 |
+
)
|
| 172 |
+
[INFO ][2026-07-08 14:40:50][root ][init_video_model ] ViTMultiSeqWrapper(
|
| 173 |
+
(backbone): VisionTransformer(
|
| 174 |
+
(patch_embed): PatchEmbed3D(
|
| 175 |
+
(proj): Conv3d(3, 1024, kernel_size=(1, 14, 14), stride=(1, 14, 14))
|
| 176 |
+
)
|
| 177 |
+
(blocks): ModuleList(
|
| 178 |
+
(0-23): 24 x Block(
|
| 179 |
+
(norm1): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 180 |
+
(norm2): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 181 |
+
(drop_path1): Identity()
|
| 182 |
+
(drop_path2): Identity()
|
| 183 |
+
(attn): Attention(
|
| 184 |
+
(qkv): Linear(in_features=1024, out_features=3072, bias=True)
|
| 185 |
+
(attn_drop): Dropout(p=0.0, inplace=False)
|
| 186 |
+
(proj): Linear(in_features=1024, out_features=1024, bias=True)
|
| 187 |
+
(proj_drop): Dropout(p=0.0, inplace=False)
|
| 188 |
+
)
|
| 189 |
+
(mlp): MLP(
|
| 190 |
+
(fc1): Linear(in_features=1024, out_features=4096, bias=True)
|
| 191 |
+
(act): GELU(approximate='none')
|
| 192 |
+
(fc2): Linear(in_features=4096, out_features=1024, bias=True)
|
| 193 |
+
(drop): Dropout(p=0.0, inplace=False)
|
| 194 |
+
)
|
| 195 |
+
)
|
| 196 |
+
)
|
| 197 |
+
(norm): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 198 |
+
)
|
| 199 |
+
)
|
| 200 |
+
[INFO ][2026-07-08 14:40:50][root ][init_video_model ] Encoder number of parameters: 302964736
|
| 201 |
+
[INFO ][2026-07-08 14:40:50][root ][init_video_model ] Predictor number of parameters: 22042240
|
| 202 |
+
[INFO ][2026-07-08 14:40:50][root ][init_video_model ] Target encoder number of parameters: 302964736
|
| 203 |
+
[INFO ][2026-07-08 14:40:51][root ][make_videodataset ] VideoDataset dataset created
|
| 204 |
+
[INFO ][2026-07-08 14:40:51][WeightedSampler ][__init__ ] Using DistributedWeightedSampler with rank 1 / 3
|
| 205 |
+
[INFO ][2026-07-08 14:40:51][root ][make_videodataset ] VideoDataset unsupervised data loader created
|
| 206 |
+
[INFO ][2026-07-08 14:40:51][app.vjepa.train ][main ] iterations per epoch/dataset length: 300/2103
|
| 207 |
+
[INFO ][2026-07-08 14:40:51][app.vjepa.train ][main ] Wrapping models in DDP (rank 1)...
|
| 208 |
+
[INFO ][2026-07-08 14:40:52][root ][load_checkpoint ] Loading checkpoint from /var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/latest.pt
|
| 209 |
+
[INFO ][2026-07-08 14:41:00][root ][load_checkpoint ] loaded pretrained encoder from epoch 90 with msg: <All keys matched successfully>
|
| 210 |
+
[INFO ][2026-07-08 14:41:00][root ][load_checkpoint ] loaded pretrained predictor from epoch 90 with msg: <All keys matched successfully>
|
| 211 |
+
[INFO ][2026-07-08 14:41:00][root ][load_checkpoint ] loaded pretrained target encoder from epoch 90 with msg: <All keys matched successfully>
|
| 212 |
+
[INFO ][2026-07-08 14:41:01][root ][load_checkpoint ] loaded optimizers from epoch 90
|
| 213 |
+
[INFO ][2026-07-08 14:41:01][root ][load_checkpoint ] read-path: /var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/latest.pt
|
| 214 |
+
[INFO ][2026-07-08 14:42:20][app.vjepa.train ][main ] Initializing loader...
|
| 215 |
+
submitit INFO (2026-07-09 14:15:07,327) - Job completed successfully
|
| 216 |
+
[INFO ][2026-07-09 14:15:07][submitit ][process_job ] Job completed successfully
|
| 217 |
+
submitit INFO (2026-07-09 14:15:07,329) - Exiting after successful completion
|
| 218 |
+
[INFO ][2026-07-09 14:15:07][submitit ][process_job ] Exiting after successful completion
|
inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_27036/27036_2_log.err
ADDED
|
@@ -0,0 +1,103 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
/var/scratch/dcanezil/.cache/uv/virtualenvs/vd/lib/python3.12/site-packages/timm/models/layers/__init__.py:49: FutureWarning: Importing from timm.models.layers is deprecated, please import via timm.layers
|
| 2 |
+
warnings.warn(f"Importing from {__name__} is deprecated, please import via timm.layers", FutureWarning)
|
| 3 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/utils.py:888: FutureWarning: `torch.cuda.amp.GradScaler(args...)` is deprecated. Please use `torch.amp.GradScaler('cuda', args...)` instead.
|
| 4 |
+
scaler = torch.cuda.amp.GradScaler() if mixed_precision else None
|
| 5 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 6 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 7 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/making_tea/mtYFNsRcxY4_000063_000073.avi'
|
| 8 |
+
warnings.warn(f"video path not found {fname=}")
|
| 9 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/water_skiing/4bhvIHWADVU_000022_000032.avi'
|
| 10 |
+
warnings.warn(f"video path not found {fname=}")
|
| 11 |
+
[15:09:37] /github/workspace/src/video/video_reader.cc:83: ERROR opening: /local/dcanezil/data/kinetics/VideoData/trimming_trees/T6dqaZioaXs_000004_000014.avi, Invalid data found when processing input
|
| 12 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/eating_carrots/eiZ8Hzc7FPU_000080_000090.avi'
|
| 13 |
+
warnings.warn(f"video path not found {fname=}")
|
| 14 |
+
[16:15:57] /github/workspace/src/video/video_reader.cc:83: ERROR opening: /local/dcanezil/data/kinetics/VideoData/climbing_tree/QX5IzK4iFCM_000002_000012.avi, Invalid data found when processing input
|
| 15 |
+
[17:10:23] /github/workspace/src/video/video_reader.cc:83: ERROR opening: /local/dcanezil/data/kinetics/VideoData/trimming_trees/T6dqaZioaXs_000004_000014.avi, Invalid data found when processing input
|
| 16 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_poker/6-DyF8umej8_000097_000107.avi'
|
| 17 |
+
warnings.warn(f"video path not found {fname=}")
|
| 18 |
+
[18:14:21] /github/workspace/src/video/video_reader.cc:83: ERROR opening: /local/dcanezil/data/kinetics/VideoData/trimming_trees/T6dqaZioaXs_000004_000014.avi, Invalid data found when processing input
|
| 19 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 20 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 21 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/making_tea/mtYFNsRcxY4_000063_000073.avi'
|
| 22 |
+
warnings.warn(f"video path not found {fname=}")
|
| 23 |
+
[19:11:49] /github/workspace/src/video/video_reader.cc:83: ERROR opening: /local/dcanezil/data/kinetics/VideoData/long_jump/jsYiitkbr5w_000004_000014.avi, Invalid data found when processing input
|
| 24 |
+
[19:24:06] /github/workspace/src/video/video_reader.cc:83: ERROR opening: /local/dcanezil/data/kinetics/VideoData/climbing_tree/QX5IzK4iFCM_000002_000012.avi, Invalid data found when processing input
|
| 25 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/pumping_fist/wKsuyr6Xk30_000094_000104.avi'
|
| 26 |
+
warnings.warn(f"video path not found {fname=}")
|
| 27 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/dancing_macarena/uGQgxFHemyA_000059_000069.avi'
|
| 28 |
+
warnings.warn(f"video path not found {fname=}")
|
| 29 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/waxing_legs/cqUusXBODuw_001046_001056.avi'
|
| 30 |
+
warnings.warn(f"video path not found {fname=}")
|
| 31 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/robot_dancing/J1x0XlWa6HM_000147_000157.avi'
|
| 32 |
+
warnings.warn(f"video path not found {fname=}")
|
| 33 |
+
[22:28:03] /github/workspace/src/video/video_reader.cc:83: ERROR opening: /local/dcanezil/data/kinetics/VideoData/trimming_trees/T6dqaZioaXs_000004_000014.avi, Invalid data found when processing input
|
| 34 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 35 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 36 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/windsurfing/i-gzh_BPDa8_000154_000164.avi'
|
| 37 |
+
warnings.warn(f"video path not found {fname=}")
|
| 38 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/waxing_legs/cqUusXBODuw_001046_001056.avi'
|
| 39 |
+
warnings.warn(f"video path not found {fname=}")
|
| 40 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/surfing_water/5l5Pdd96Pao_000161_000171.avi'
|
| 41 |
+
warnings.warn(f"video path not found {fname=}")
|
| 42 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/eating_carrots/eiZ8Hzc7FPU_000080_000090.avi'
|
| 43 |
+
warnings.warn(f"video path not found {fname=}")
|
| 44 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/waxing_legs/cqUusXBODuw_001046_001056.avi'
|
| 45 |
+
warnings.warn(f"video path not found {fname=}")
|
| 46 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_poker/-jGxlNQKkeo_000092_000102.avi'
|
| 47 |
+
warnings.warn(f"video path not found {fname=}")
|
| 48 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_poker/-jGxlNQKkeo_000092_000102.avi'
|
| 49 |
+
warnings.warn(f"video path not found {fname=}")
|
| 50 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 51 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 52 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/sailing/ZolXG_-RgIs_000970_000980.avi'
|
| 53 |
+
warnings.warn(f"video path not found {fname=}")
|
| 54 |
+
[03:32:39] /github/workspace/src/video/video_reader.cc:83: ERROR opening: /local/dcanezil/data/kinetics/VideoData/long_jump/jsYiitkbr5w_000004_000014.avi, Invalid data found when processing input
|
| 55 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/eating_carrots/eiZ8Hzc7FPU_000080_000090.avi'
|
| 56 |
+
warnings.warn(f"video path not found {fname=}")
|
| 57 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/eating_carrots/eiZ8Hzc7FPU_000080_000090.avi'
|
| 58 |
+
warnings.warn(f"video path not found {fname=}")
|
| 59 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/making_tea/mtYFNsRcxY4_000063_000073.avi'
|
| 60 |
+
warnings.warn(f"video path not found {fname=}")
|
| 61 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/frying_vegetables/Nr-zCXgW2Sw_000161_000171.avi'
|
| 62 |
+
warnings.warn(f"video path not found {fname=}")
|
| 63 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/waxing_legs/cqUusXBODuw_001046_001056.avi'
|
| 64 |
+
warnings.warn(f"video path not found {fname=}")
|
| 65 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_saxophone/Zs0_2GMEPXo_000054_000064.avi'
|
| 66 |
+
warnings.warn(f"video path not found {fname=}")
|
| 67 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 68 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 69 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/picking_fruit/3VvkoFtPCCU_000045_000055.avi'
|
| 70 |
+
warnings.warn(f"video path not found {fname=}")
|
| 71 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_saxophone/Zs0_2GMEPXo_000054_000064.avi'
|
| 72 |
+
warnings.warn(f"video path not found {fname=}")
|
| 73 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/making_tea/mtYFNsRcxY4_000063_000073.avi'
|
| 74 |
+
warnings.warn(f"video path not found {fname=}")
|
| 75 |
+
[07:47:35] /github/workspace/src/video/video_reader.cc:83: ERROR opening: /local/dcanezil/data/kinetics/VideoData/trimming_trees/T6dqaZioaXs_000004_000014.avi, Invalid data found when processing input
|
| 76 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/baking_cookies/Q3XGZmqk1Q0_000296_000306.avi'
|
| 77 |
+
warnings.warn(f"video path not found {fname=}")
|
| 78 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/robot_dancing/J1x0XlWa6HM_000147_000157.avi'
|
| 79 |
+
warnings.warn(f"video path not found {fname=}")
|
| 80 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_poker/6-DyF8umej8_000097_000107.avi'
|
| 81 |
+
warnings.warn(f"video path not found {fname=}")
|
| 82 |
+
[09:49:39] /github/workspace/src/video/video_reader.cc:83: ERROR opening: /local/dcanezil/data/kinetics/VideoData/long_jump/jsYiitkbr5w_000004_000014.avi, Invalid data found when processing input
|
| 83 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/baking_cookies/Q3XGZmqk1Q0_000296_000306.avi'
|
| 84 |
+
warnings.warn(f"video path not found {fname=}")
|
| 85 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 86 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 87 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/surfing_water/5l5Pdd96Pao_000161_000171.avi'
|
| 88 |
+
warnings.warn(f"video path not found {fname=}")
|
| 89 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_monopoly/NLL667uPWVA_000066_000076.avi'
|
| 90 |
+
warnings.warn(f"video path not found {fname=}")
|
| 91 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/dancing_macarena/uGQgxFHemyA_000059_000069.avi'
|
| 92 |
+
warnings.warn(f"video path not found {fname=}")
|
| 93 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_saxophone/Zs0_2GMEPXo_000054_000064.avi'
|
| 94 |
+
warnings.warn(f"video path not found {fname=}")
|
| 95 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/eating_carrots/eiZ8Hzc7FPU_000080_000090.avi'
|
| 96 |
+
warnings.warn(f"video path not found {fname=}")
|
| 97 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/making_tea/mtYFNsRcxY4_000063_000073.avi'
|
| 98 |
+
warnings.warn(f"video path not found {fname=}")
|
| 99 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/sailing/ZolXG_-RgIs_000970_000980.avi'
|
| 100 |
+
warnings.warn(f"video path not found {fname=}")
|
| 101 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_poker/6-DyF8umej8_000097_000107.avi'
|
| 102 |
+
warnings.warn(f"video path not found {fname=}")
|
| 103 |
+
[rank2]:[W709 14:15:08.751823810 ProcessGroupNCCL.cpp:1553] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())
|
inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_27036/27036_2_log.out
ADDED
|
@@ -0,0 +1,218 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
submitit INFO (2026-07-08 14:39:49,900) - Starting with JobEnvironment(job_id=27036, hostname=node404, local_rank=2(3), node=0(1), global_rank=2(3))
|
| 2 |
+
submitit INFO (2026-07-08 14:39:49,900) - Loading pickle: /var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_27036/27036_submitted.pkl
|
| 3 |
+
INFO:root:loaded pretrain params...
|
| 4 |
+
{ 'app': 'vjepa',
|
| 5 |
+
'cpus_per_task': 16,
|
| 6 |
+
'data': { 'batch_size': 64,
|
| 7 |
+
'crop_size': 224,
|
| 8 |
+
'dataset_fpcs': [8, 8],
|
| 9 |
+
'dataset_type': 'VideoDataset',
|
| 10 |
+
'datasets': [ '/local/dcanezil/data/kinetics/train.csv',
|
| 11 |
+
'/local/dcanezil/data/ssv2/train.csv'],
|
| 12 |
+
'datasets_weights': [0.65, 0.35],
|
| 13 |
+
'fps': 2,
|
| 14 |
+
'num_workers': 4,
|
| 15 |
+
'patch_size': 14,
|
| 16 |
+
'persistent_workers': True,
|
| 17 |
+
'pin_mem': True,
|
| 18 |
+
'tubelet_size': 1},
|
| 19 |
+
'data_aug': { 'auto_augment': False,
|
| 20 |
+
'motion_shift': False,
|
| 21 |
+
'random_resize_aspect_ratio': [0.75, 1.35],
|
| 22 |
+
'random_resize_scale': [0.3, 1.0],
|
| 23 |
+
'reprob': 0.0},
|
| 24 |
+
'folder': '/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow',
|
| 25 |
+
'loss': {'loss_exp': 1.0},
|
| 26 |
+
'mask': [ { 'aspect_ratio': [0.75, 1.5],
|
| 27 |
+
'full_complement': False,
|
| 28 |
+
'max_keep': None,
|
| 29 |
+
'max_temporal_keep': 1.0,
|
| 30 |
+
'num_blocks': 8,
|
| 31 |
+
'spatial_scale': [0.05, 0.05],
|
| 32 |
+
'temporal_scale': [1.0, 1.0]},
|
| 33 |
+
{ 'aspect_ratio': [0.75, 1.5],
|
| 34 |
+
'full_complement': False,
|
| 35 |
+
'max_keep': None,
|
| 36 |
+
'max_temporal_keep': 1.0,
|
| 37 |
+
'num_blocks': 2,
|
| 38 |
+
'spatial_scale': [0.4, 0.4],
|
| 39 |
+
'temporal_scale': [1.0, 1.0]}],
|
| 40 |
+
'mem_per_gpu': '40G',
|
| 41 |
+
'meta': { 'dtype': 'bfloat16',
|
| 42 |
+
'knn_eval_epoch0': True,
|
| 43 |
+
'knn_eval_freq': 5,
|
| 44 |
+
'knn_eval_presets': [ { 'config': { 'batch_size': 8,
|
| 45 |
+
'dataset_train': '/local/dcanezil/data/ssv2/train_coarse10.csv',
|
| 46 |
+
'dataset_val': '/local/dcanezil/data/ssv2/val_coarse10.csv',
|
| 47 |
+
'eval_videos_per_class': 100,
|
| 48 |
+
'linear_probe': True,
|
| 49 |
+
'num_workers': 2,
|
| 50 |
+
'pool_type': 'temporal_concat',
|
| 51 |
+
'train_videos_per_class': 500},
|
| 52 |
+
'preset': 'ssv2_coarse10'}],
|
| 53 |
+
'load_checkpoint': True,
|
| 54 |
+
'read_checkpoint': None,
|
| 55 |
+
'save_every_freq': -1,
|
| 56 |
+
'seed': 239,
|
| 57 |
+
'use_sdpa': True,
|
| 58 |
+
'use_wandb': True,
|
| 59 |
+
'wandb_project': 'vjepa_ablation'},
|
| 60 |
+
'metrics': {'sigreg': {}, 'std': {}},
|
| 61 |
+
'model': { 'class_token': True,
|
| 62 |
+
'freeze_backbone': False,
|
| 63 |
+
'is_causal': True,
|
| 64 |
+
'model_name': 'vit_large_patch14_dinov2_lvd142m',
|
| 65 |
+
'pe_type': '2d',
|
| 66 |
+
'pred_depth': 12,
|
| 67 |
+
'pred_embed_dim': 384,
|
| 68 |
+
'pred_num_heads': 6,
|
| 69 |
+
'predictor': 'v2_cross',
|
| 70 |
+
'stem_type': '2d',
|
| 71 |
+
'target_embed_dim': 1024,
|
| 72 |
+
'target_kind': 'ema',
|
| 73 |
+
'uniform_power': True,
|
| 74 |
+
'use_activation_checkpointing': True,
|
| 75 |
+
'use_alibi': False,
|
| 76 |
+
'use_lora': False,
|
| 77 |
+
'use_mask_tokens': True,
|
| 78 |
+
'use_rope': False,
|
| 79 |
+
'use_sdpa': True,
|
| 80 |
+
'zero_init_mask_tokens': True},
|
| 81 |
+
'nodes': 1,
|
| 82 |
+
'optimization': { 'clip_grad': None,
|
| 83 |
+
'ema': [1.0, 0.99925],
|
| 84 |
+
'ema_hold_epochs': 0,
|
| 85 |
+
'ema_ramp_epochs': 5,
|
| 86 |
+
'encoder_layer_decay': 0.9,
|
| 87 |
+
'epochs': 120,
|
| 88 |
+
'final_lr': 0.0001,
|
| 89 |
+
'final_weight_decay': 0.04,
|
| 90 |
+
'ipe': 300,
|
| 91 |
+
'ipe_scale': 1.0,
|
| 92 |
+
'lr': 0.0001,
|
| 93 |
+
'predictor_lr_scale': 3.0,
|
| 94 |
+
'start_lr': 1e-05,
|
| 95 |
+
'warmup': 2,
|
| 96 |
+
'weight_decay': 0.04},
|
| 97 |
+
'tasks_per_node': 3}
|
| 98 |
+
INFO:root:Running pre-training of app: vjepa
|
| 99 |
+
[INFO ][2026-07-08 14:40:02][app.vjepa.train ][main ] which_dtype='bfloat16'
|
| 100 |
+
[INFO ][2026-07-08 14:40:02][app.vjepa.train ][main ] Disabling persistent_workers (incompatible with KNN eval)
|
| 101 |
+
[INFO ][2026-07-08 14:40:02][app.vjepa.train ][main ] NCCL_SOCKET_IFNAME=None
|
| 102 |
+
[INFO ][2026-07-08 14:40:05][app.vjepa.train ][main ] Initialized (rank/world-size) 2/3, tasks_per_node=3
|
| 103 |
+
[WARNING ][2026-07-08 14:40:37][experiments.stmodels.vision_transformers_v3][_build_dinov2_vit ] DINOv2 is an image model — fwd_chunk is unset so all frames will be processed jointly. Set fwd_chunk=1 for per-frame (standard) processing.
|
| 104 |
+
[INFO ][2026-07-08 14:40:43][experiments.stmodels.vision_transformers_v3][_build_dinov2_vit ] Loading pretrained weights for vit_large_patch14_dinov2_lvd142m from timm (vit_large_patch14_dinov2)
|
| 105 |
+
[INFO ][2026-07-08 14:40:47][timm.models._builder][load_pretrained ] Loading pretrained weights from Hugging Face hub (timm/vit_large_patch14_dinov2.lvd142m)
|
| 106 |
+
[INFO ][2026-07-08 14:40:47][timm.models._hub ][load_state_dict_from_hf ] [timm/vit_large_patch14_dinov2.lvd142m] Safe alternative available for 'pytorch_model.bin' (as 'model.safetensors'). Loading weights using safetensors.
|
| 107 |
+
[INFO ][2026-07-08 14:40:47][timm.layers.pos_embed][resample_abs_pos_embed ] Resized position embedding: (37, 37) to (16, 16).
|
| 108 |
+
[INFO ][2026-07-08 14:40:50][root ][init_video_model ] ViTMultiSeqWrapper(
|
| 109 |
+
(backbone): VisionTransformer(
|
| 110 |
+
(patch_embed): PatchEmbed3D(
|
| 111 |
+
(proj): Conv3d(3, 1024, kernel_size=(1, 14, 14), stride=(1, 14, 14))
|
| 112 |
+
)
|
| 113 |
+
(blocks): ModuleList(
|
| 114 |
+
(0-23): 24 x Block(
|
| 115 |
+
(norm1): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 116 |
+
(norm2): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 117 |
+
(drop_path1): Identity()
|
| 118 |
+
(drop_path2): Identity()
|
| 119 |
+
(attn): Attention(
|
| 120 |
+
(qkv): Linear(in_features=1024, out_features=3072, bias=True)
|
| 121 |
+
(attn_drop): Dropout(p=0.0, inplace=False)
|
| 122 |
+
(proj): Linear(in_features=1024, out_features=1024, bias=True)
|
| 123 |
+
(proj_drop): Dropout(p=0.0, inplace=False)
|
| 124 |
+
)
|
| 125 |
+
(mlp): MLP(
|
| 126 |
+
(fc1): Linear(in_features=1024, out_features=4096, bias=True)
|
| 127 |
+
(act): GELU(approximate='none')
|
| 128 |
+
(fc2): Linear(in_features=4096, out_features=1024, bias=True)
|
| 129 |
+
(drop): Dropout(p=0.0, inplace=False)
|
| 130 |
+
)
|
| 131 |
+
)
|
| 132 |
+
)
|
| 133 |
+
(norm): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 134 |
+
)
|
| 135 |
+
)
|
| 136 |
+
[INFO ][2026-07-08 14:40:50][root ][init_video_model ] VJEPAPredictorMultiSeqWrapper(
|
| 137 |
+
(backbone): PredictorV2(
|
| 138 |
+
(predictor_embed): Linear(in_features=1024, out_features=384, bias=True)
|
| 139 |
+
(mask_tokens): ParameterList(
|
| 140 |
+
(0): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
|
| 141 |
+
(1): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
|
| 142 |
+
(2): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
|
| 143 |
+
(3): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
|
| 144 |
+
)
|
| 145 |
+
(predictor_blocks): ModuleList(
|
| 146 |
+
(0-11): 12 x Block(
|
| 147 |
+
(residual1): EfficientResidual(
|
| 148 |
+
(norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
|
| 149 |
+
(fn): Attention(
|
| 150 |
+
(q_proj): Linear(in_features=384, out_features=384, bias=False)
|
| 151 |
+
(k_proj): Linear(in_features=384, out_features=384, bias=False)
|
| 152 |
+
(v_proj): Linear(in_features=384, out_features=384, bias=False)
|
| 153 |
+
(proj): Linear(in_features=384, out_features=384, bias=False)
|
| 154 |
+
(rope): Rope()
|
| 155 |
+
)
|
| 156 |
+
)
|
| 157 |
+
(residual2): EfficientResidual(
|
| 158 |
+
(norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
|
| 159 |
+
(fn): MLP(
|
| 160 |
+
(fc1): Linear(in_features=384, out_features=1536, bias=False)
|
| 161 |
+
(act): GELU(approximate='none')
|
| 162 |
+
(fc2): Linear(in_features=1536, out_features=384, bias=False)
|
| 163 |
+
(drop): Dropout(p=0.0, inplace=False)
|
| 164 |
+
)
|
| 165 |
+
)
|
| 166 |
+
)
|
| 167 |
+
)
|
| 168 |
+
(predictor_norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
|
| 169 |
+
(predictor_proj): Linear(in_features=384, out_features=1024, bias=True)
|
| 170 |
+
)
|
| 171 |
+
)
|
| 172 |
+
[INFO ][2026-07-08 14:40:50][root ][init_video_model ] ViTMultiSeqWrapper(
|
| 173 |
+
(backbone): VisionTransformer(
|
| 174 |
+
(patch_embed): PatchEmbed3D(
|
| 175 |
+
(proj): Conv3d(3, 1024, kernel_size=(1, 14, 14), stride=(1, 14, 14))
|
| 176 |
+
)
|
| 177 |
+
(blocks): ModuleList(
|
| 178 |
+
(0-23): 24 x Block(
|
| 179 |
+
(norm1): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 180 |
+
(norm2): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 181 |
+
(drop_path1): Identity()
|
| 182 |
+
(drop_path2): Identity()
|
| 183 |
+
(attn): Attention(
|
| 184 |
+
(qkv): Linear(in_features=1024, out_features=3072, bias=True)
|
| 185 |
+
(attn_drop): Dropout(p=0.0, inplace=False)
|
| 186 |
+
(proj): Linear(in_features=1024, out_features=1024, bias=True)
|
| 187 |
+
(proj_drop): Dropout(p=0.0, inplace=False)
|
| 188 |
+
)
|
| 189 |
+
(mlp): MLP(
|
| 190 |
+
(fc1): Linear(in_features=1024, out_features=4096, bias=True)
|
| 191 |
+
(act): GELU(approximate='none')
|
| 192 |
+
(fc2): Linear(in_features=4096, out_features=1024, bias=True)
|
| 193 |
+
(drop): Dropout(p=0.0, inplace=False)
|
| 194 |
+
)
|
| 195 |
+
)
|
| 196 |
+
)
|
| 197 |
+
(norm): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 198 |
+
)
|
| 199 |
+
)
|
| 200 |
+
[INFO ][2026-07-08 14:40:50][root ][init_video_model ] Encoder number of parameters: 302964736
|
| 201 |
+
[INFO ][2026-07-08 14:40:50][root ][init_video_model ] Predictor number of parameters: 22042240
|
| 202 |
+
[INFO ][2026-07-08 14:40:50][root ][init_video_model ] Target encoder number of parameters: 302964736
|
| 203 |
+
[INFO ][2026-07-08 14:40:51][root ][make_videodataset ] VideoDataset dataset created
|
| 204 |
+
[INFO ][2026-07-08 14:40:51][WeightedSampler ][__init__ ] Using DistributedWeightedSampler with rank 2 / 3
|
| 205 |
+
[INFO ][2026-07-08 14:40:51][root ][make_videodataset ] VideoDataset unsupervised data loader created
|
| 206 |
+
[INFO ][2026-07-08 14:40:51][app.vjepa.train ][main ] iterations per epoch/dataset length: 300/2103
|
| 207 |
+
[INFO ][2026-07-08 14:40:51][app.vjepa.train ][main ] Wrapping models in DDP (rank 2)...
|
| 208 |
+
[INFO ][2026-07-08 14:40:52][root ][load_checkpoint ] Loading checkpoint from /var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/latest.pt
|
| 209 |
+
[INFO ][2026-07-08 14:41:00][root ][load_checkpoint ] loaded pretrained encoder from epoch 90 with msg: <All keys matched successfully>
|
| 210 |
+
[INFO ][2026-07-08 14:41:00][root ][load_checkpoint ] loaded pretrained predictor from epoch 90 with msg: <All keys matched successfully>
|
| 211 |
+
[INFO ][2026-07-08 14:41:00][root ][load_checkpoint ] loaded pretrained target encoder from epoch 90 with msg: <All keys matched successfully>
|
| 212 |
+
[INFO ][2026-07-08 14:41:01][root ][load_checkpoint ] loaded optimizers from epoch 90
|
| 213 |
+
[INFO ][2026-07-08 14:41:01][root ][load_checkpoint ] read-path: /var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/latest.pt
|
| 214 |
+
[INFO ][2026-07-08 14:42:20][app.vjepa.train ][main ] Initializing loader...
|
| 215 |
+
submitit INFO (2026-07-09 14:15:07,350) - Job completed successfully
|
| 216 |
+
[INFO ][2026-07-09 14:15:07][submitit ][process_job ] Job completed successfully
|
| 217 |
+
submitit INFO (2026-07-09 14:15:07,355) - Exiting after successful completion
|
| 218 |
+
[INFO ][2026-07-09 14:15:07][submitit ][process_job ] Exiting after successful completion
|
inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_27070/27070_0_log.err
ADDED
|
@@ -0,0 +1,167 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
/var/scratch/dcanezil/.cache/uv/virtualenvs/vd/lib/python3.12/site-packages/timm/models/layers/__init__.py:49: FutureWarning: Importing from timm.models.layers is deprecated, please import via timm.layers
|
| 2 |
+
warnings.warn(f"Importing from {__name__} is deprecated, please import via timm.layers", FutureWarning)
|
| 3 |
+
wandb: Currently logged in as: dgcnz (uvjepa) to https://api.wandb.ai. Use `wandb login --relogin` to force relogin
|
| 4 |
+
wandb: setting up run vfclercb
|
| 5 |
+
wandb: Tracking run with wandb version 0.23.1
|
| 6 |
+
wandb: Run data is saved locally in /var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/wandb/run-20260709_143819-vfclercb
|
| 7 |
+
wandb: Run `wandb offline` to turn off syncing.
|
| 8 |
+
wandb: Resuming run d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow
|
| 9 |
+
wandb: ⭐️ View project at https://wandb.ai/uvjepa/vjepa_ablation
|
| 10 |
+
wandb: 🚀 View run at https://wandb.ai/uvjepa/vjepa_ablation/runs/vfclercb
|
| 11 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/utils.py:888: FutureWarning: `torch.cuda.amp.GradScaler(args...)` is deprecated. Please use `torch.amp.GradScaler('cuda', args...)` instead.
|
| 12 |
+
scaler = torch.cuda.amp.GradScaler() if mixed_precision else None
|
| 13 |
+
/var/scratch/dcanezil/.cache/uv/virtualenvs/vd/lib/python3.12/site-packages/torch/distributed/c10d_logger.py:83: UserWarning: barrier(): using the device under current context. You can specify `device_id` in `init_process_group` to mute this warning.
|
| 14 |
+
return func(*args, **kwargs)
|
| 15 |
+
wandb: WARNING Tried to log to step 0 that is less than the current step 36001. Steps must be monotonically increasing, so this data will be ignored. See https://wandb.me/define-metric to log data out of order.
|
| 16 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 17 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 18 |
+
wandb: WARNING Tried to log to step 36000 that is less than the current step 36001. Steps must be monotonically increasing, so this data will be ignored. See https://wandb.me/define-metric to log data out of order.
|
| 19 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_saxophone/Zs0_2GMEPXo_000054_000064.avi'
|
| 20 |
+
warnings.warn(f"video path not found {fname=}")
|
| 21 |
+
[15:20:37] /github/workspace/src/video/video_reader.cc:83: ERROR opening: /local/dcanezil/data/kinetics/VideoData/long_jump/jsYiitkbr5w_000004_000014.avi, Invalid data found when processing input
|
| 22 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/pumping_fist/wKsuyr6Xk30_000094_000104.avi'
|
| 23 |
+
warnings.warn(f"video path not found {fname=}")
|
| 24 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/baking_cookies/Q3XGZmqk1Q0_000296_000306.avi'
|
| 25 |
+
warnings.warn(f"video path not found {fname=}")
|
| 26 |
+
[17:06:20] /github/workspace/src/video/video_reader.cc:83: ERROR opening: /local/dcanezil/data/kinetics/VideoData/climbing_tree/QX5IzK4iFCM_000002_000012.avi, Invalid data found when processing input
|
| 27 |
+
[17:09:20] /github/workspace/src/video/video_reader.cc:83: ERROR opening: /local/dcanezil/data/kinetics/VideoData/trimming_trees/T6dqaZioaXs_000004_000014.avi, Invalid data found when processing input
|
| 28 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/sailing/ZolXG_-RgIs_000970_000980.avi'
|
| 29 |
+
warnings.warn(f"video path not found {fname=}")
|
| 30 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_poker/-jGxlNQKkeo_000092_000102.avi'
|
| 31 |
+
warnings.warn(f"video path not found {fname=}")
|
| 32 |
+
/var/scratch/dcanezil/.cache/uv/virtualenvs/vd/lib/python3.12/site-packages/torch/distributed/c10d_logger.py:83: UserWarning: barrier(): using the device under current context. You can specify `device_id` in `init_process_group` to mute this warning.
|
| 33 |
+
return func(*args, **kwargs)
|
| 34 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 35 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 36 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/robot_dancing/J1x0XlWa6HM_000147_000157.avi'
|
| 37 |
+
warnings.warn(f"video path not found {fname=}")
|
| 38 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/windsurfing/i-gzh_BPDa8_000154_000164.avi'
|
| 39 |
+
warnings.warn(f"video path not found {fname=}")
|
| 40 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/sailing/ZolXG_-RgIs_000970_000980.avi'
|
| 41 |
+
warnings.warn(f"video path not found {fname=}")
|
| 42 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_saxophone/Zs0_2GMEPXo_000054_000064.avi'
|
| 43 |
+
warnings.warn(f"video path not found {fname=}")
|
| 44 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/surfing_water/5l5Pdd96Pao_000161_000171.avi'
|
| 45 |
+
warnings.warn(f"video path not found {fname=}")
|
| 46 |
+
[19:44:41] /github/workspace/src/video/video_reader.cc:83: ERROR opening: /local/dcanezil/data/kinetics/VideoData/climbing_tree/QX5IzK4iFCM_000002_000012.avi, Invalid data found when processing input
|
| 47 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_poker/6-DyF8umej8_000097_000107.avi'
|
| 48 |
+
warnings.warn(f"video path not found {fname=}")
|
| 49 |
+
/var/scratch/dcanezil/.cache/uv/virtualenvs/vd/lib/python3.12/site-packages/torch/distributed/c10d_logger.py:83: UserWarning: barrier(): using the device under current context. You can specify `device_id` in `init_process_group` to mute this warning.
|
| 50 |
+
return func(*args, **kwargs)
|
| 51 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 52 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 53 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/waxing_legs/cqUusXBODuw_001046_001056.avi'
|
| 54 |
+
warnings.warn(f"video path not found {fname=}")
|
| 55 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/water_skiing/4bhvIHWADVU_000022_000032.avi'
|
| 56 |
+
warnings.warn(f"video path not found {fname=}")
|
| 57 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/robot_dancing/J1x0XlWa6HM_000147_000157.avi'
|
| 58 |
+
warnings.warn(f"video path not found {fname=}")
|
| 59 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_saxophone/Zs0_2GMEPXo_000054_000064.avi'
|
| 60 |
+
warnings.warn(f"video path not found {fname=}")
|
| 61 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/picking_fruit/3VvkoFtPCCU_000045_000055.avi'
|
| 62 |
+
warnings.warn(f"video path not found {fname=}")
|
| 63 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/frying_vegetables/Nr-zCXgW2Sw_000161_000171.avi'
|
| 64 |
+
warnings.warn(f"video path not found {fname=}")
|
| 65 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/pumping_fist/wKsuyr6Xk30_000094_000104.avi'
|
| 66 |
+
warnings.warn(f"video path not found {fname=}")
|
| 67 |
+
[00:47:16] /github/workspace/src/video/video_reader.cc:83: ERROR opening: /local/dcanezil/data/kinetics/VideoData/trimming_trees/T6dqaZioaXs_000004_000014.avi, Invalid data found when processing input
|
| 68 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/frying_vegetables/Nr-zCXgW2Sw_000161_000171.avi'
|
| 69 |
+
warnings.warn(f"video path not found {fname=}")
|
| 70 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/bowling/f4eb0wOlspM_000053_000063.avi'
|
| 71 |
+
warnings.warn(f"video path not found {fname=}")
|
| 72 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/making_tea/mtYFNsRcxY4_000063_000073.avi'
|
| 73 |
+
warnings.warn(f"video path not found {fname=}")
|
| 74 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/waxing_legs/cqUusXBODuw_001046_001056.avi'
|
| 75 |
+
warnings.warn(f"video path not found {fname=}")
|
| 76 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/waxing_legs/cqUusXBODuw_001046_001056.avi'
|
| 77 |
+
warnings.warn(f"video path not found {fname=}")
|
| 78 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/robot_dancing/J1x0XlWa6HM_000147_000157.avi'
|
| 79 |
+
warnings.warn(f"video path not found {fname=}")
|
| 80 |
+
/var/scratch/dcanezil/.cache/uv/virtualenvs/vd/lib/python3.12/site-packages/torch/distributed/c10d_logger.py:83: UserWarning: barrier(): using the device under current context. You can specify `device_id` in `init_process_group` to mute this warning.
|
| 81 |
+
return func(*args, **kwargs)
|
| 82 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 83 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 84 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/baking_cookies/Q3XGZmqk1Q0_000296_000306.avi'
|
| 85 |
+
warnings.warn(f"video path not found {fname=}")
|
| 86 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/sailing/ZolXG_-RgIs_000970_000980.avi'
|
| 87 |
+
warnings.warn(f"video path not found {fname=}")
|
| 88 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_flute/co50KUHacYw_000005_000015.avi'
|
| 89 |
+
warnings.warn(f"video path not found {fname=}")
|
| 90 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/making_tea/mtYFNsRcxY4_000063_000073.avi'
|
| 91 |
+
warnings.warn(f"video path not found {fname=}")
|
| 92 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_saxophone/Zs0_2GMEPXo_000054_000064.avi'
|
| 93 |
+
warnings.warn(f"video path not found {fname=}")
|
| 94 |
+
[06:08:27] /github/workspace/src/video/video_reader.cc:83: ERROR opening: /local/dcanezil/data/kinetics/VideoData/climbing_tree/QX5IzK4iFCM_000002_000012.avi, Invalid data found when processing input
|
| 95 |
+
/var/scratch/dcanezil/.cache/uv/virtualenvs/vd/lib/python3.12/site-packages/torch/distributed/c10d_logger.py:83: UserWarning: barrier(): using the device under current context. You can specify `device_id` in `init_process_group` to mute this warning.
|
| 96 |
+
return func(*args, **kwargs)
|
| 97 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 98 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 99 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/making_tea/mtYFNsRcxY4_000063_000073.avi'
|
| 100 |
+
warnings.warn(f"video path not found {fname=}")
|
| 101 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/waxing_legs/cqUusXBODuw_001046_001056.avi'
|
| 102 |
+
warnings.warn(f"video path not found {fname=}")
|
| 103 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/pumping_fist/wKsuyr6Xk30_000094_000104.avi'
|
| 104 |
+
warnings.warn(f"video path not found {fname=}")
|
| 105 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/sailing/ZolXG_-RgIs_000970_000980.avi'
|
| 106 |
+
warnings.warn(f"video path not found {fname=}")
|
| 107 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/pumping_fist/wKsuyr6Xk30_000094_000104.avi'
|
| 108 |
+
warnings.warn(f"video path not found {fname=}")
|
| 109 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/baking_cookies/Q3XGZmqk1Q0_000296_000306.avi'
|
| 110 |
+
warnings.warn(f"video path not found {fname=}")
|
| 111 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/dancing_macarena/uGQgxFHemyA_000059_000069.avi'
|
| 112 |
+
warnings.warn(f"video path not found {fname=}")
|
| 113 |
+
/var/scratch/dcanezil/.cache/uv/virtualenvs/vd/lib/python3.12/site-packages/torch/distributed/c10d_logger.py:83: UserWarning: barrier(): using the device under current context. You can specify `device_id` in `init_process_group` to mute this warning.
|
| 114 |
+
return func(*args, **kwargs)
|
| 115 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 116 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 117 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/baking_cookies/Q3XGZmqk1Q0_000296_000306.avi'
|
| 118 |
+
warnings.warn(f"video path not found {fname=}")
|
| 119 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/frying_vegetables/Nr-zCXgW2Sw_000161_000171.avi'
|
| 120 |
+
warnings.warn(f"video path not found {fname=}")
|
| 121 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/pumping_fist/wKsuyr6Xk30_000094_000104.avi'
|
| 122 |
+
warnings.warn(f"video path not found {fname=}")
|
| 123 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_poker/6-DyF8umej8_000097_000107.avi'
|
| 124 |
+
warnings.warn(f"video path not found {fname=}")
|
| 125 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/frying_vegetables/Nr-zCXgW2Sw_000161_000171.avi'
|
| 126 |
+
warnings.warn(f"video path not found {fname=}")
|
| 127 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/picking_fruit/3VvkoFtPCCU_000045_000055.avi'
|
| 128 |
+
warnings.warn(f"video path not found {fname=}")
|
| 129 |
+
/var/scratch/dcanezil/.cache/uv/virtualenvs/vd/lib/python3.12/site-packages/torch/distributed/c10d_logger.py:83: UserWarning: barrier(): using the device under current context. You can specify `device_id` in `init_process_group` to mute this warning.
|
| 130 |
+
return func(*args, **kwargs)
|
| 131 |
+
wandb: updating run metadata; uploading wandb-summary.json; uploading output.log
|
| 132 |
+
wandb: updating run metadata; uploading output.log
|
| 133 |
+
wandb: updating run metadata
|
| 134 |
+
wandb: uploading config.yaml
|
| 135 |
+
wandb: uploading history steps 4650-4650, summary, console lines 1356-1364
|
| 136 |
+
wandb:
|
| 137 |
+
wandb: Run history:
|
| 138 |
+
wandb: epoch ▁▁▁▁▂▂▃▃▃▃▃▃▃▃▄▄▄▄▅▅▅▅▅▅▆▆▆▆▆▆▇▇▇▇▇▇████
|
| 139 |
+
wandb: epoch/avg_data_time_ms ▄▄▄▂▃▃▂▂▂▄█▄▂▁▃▂▆▄▅▃▃▄▂▄▃▂▆▇▁▅
|
| 140 |
+
wandb: epoch/avg_gpu_time_ms ▃▂▅▄▂▃▅▂▆▃▃▃▆▄▃▂▄▁▅▅▃▂▃▂▄▁▁▁█▂
|
| 141 |
+
wandb: epoch/avg_iter_time_ms ▃▂▅▄▂▃▅▂▆▃▃▃▆▄▃▂▄▁▅▅▃▂▃▂▄▁▁▁█▂
|
| 142 |
+
wandb: epoch/avg_loss ██▇▇██▇▇▆▇▇▇▆▆▆▆▆▇▅▆▅▅▅▆▄▃▄▃▁▂
|
| 143 |
+
wandb: knn/ssv2_coarse10_temporal_concat_top1 █▄▆▇▁▁
|
| 144 |
+
wandb: knn/ssv2_coarse10_temporal_concat_top5 ██▇█▁▅
|
| 145 |
+
wandb: linear/ssv2_coarse10_temporal_concat_top1 █▆▆▅▁▂
|
| 146 |
+
wandb: linear/ssv2_coarse10_temporal_concat_top5 ▂█▇▄▁▃
|
| 147 |
+
wandb: train/ema/cos_sim ▅▅▃▄▅▄▅▁▄▇▅▆▃▅█▄▄▄▆█▃▃▅▂▅▄▃▄▄▄▆▆▄▅▅▅▃▇▅▅
|
| 148 |
+
wandb: +34 ...
|
| 149 |
+
wandb:
|
| 150 |
+
wandb: Run summary:
|
| 151 |
+
wandb: epoch 150
|
| 152 |
+
wandb: epoch/avg_data_time_ms 15.81663
|
| 153 |
+
wandb: epoch/avg_gpu_time_ms 9190.28191
|
| 154 |
+
wandb: epoch/avg_iter_time_ms 9218.5123
|
| 155 |
+
wandb: epoch/avg_loss 0.54449
|
| 156 |
+
wandb: knn/ssv2_coarse10_temporal_concat_top1 44.57831
|
| 157 |
+
wandb: knn/ssv2_coarse10_temporal_concat_top5 86.84739
|
| 158 |
+
wandb: linear/ssv2_coarse10_temporal_concat_top1 58.33333
|
| 159 |
+
wandb: linear/ssv2_coarse10_temporal_concat_top5 92.87149
|
| 160 |
+
wandb: train/ema/cos_sim 0.74967
|
| 161 |
+
wandb: +34 ...
|
| 162 |
+
wandb:
|
| 163 |
+
wandb: 🚀 View run d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow at: https://wandb.ai/uvjepa/vjepa_ablation/runs/vfclercb
|
| 164 |
+
wandb: ⭐️ View project at: https://wandb.ai/uvjepa/vjepa_ablation
|
| 165 |
+
wandb: Synced 5 W&B file(s), 0 media file(s), 0 artifact file(s) and 0 other file(s)
|
| 166 |
+
wandb: Find logs at: /var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/wandb/run-20260709_143819-vfclercb/logs
|
| 167 |
+
[rank0]:[W710 14:14:36.404469477 ProcessGroupNCCL.cpp:1553] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())
|
inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_27070/27070_0_log.out
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_27070/27070_1_log.err
ADDED
|
@@ -0,0 +1,91 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
/var/scratch/dcanezil/.cache/uv/virtualenvs/vd/lib/python3.12/site-packages/timm/models/layers/__init__.py:49: FutureWarning: Importing from timm.models.layers is deprecated, please import via timm.layers
|
| 2 |
+
warnings.warn(f"Importing from {__name__} is deprecated, please import via timm.layers", FutureWarning)
|
| 3 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/utils.py:888: FutureWarning: `torch.cuda.amp.GradScaler(args...)` is deprecated. Please use `torch.amp.GradScaler('cuda', args...)` instead.
|
| 4 |
+
scaler = torch.cuda.amp.GradScaler() if mixed_precision else None
|
| 5 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 6 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 7 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/eating_carrots/eiZ8Hzc7FPU_000080_000090.avi'
|
| 8 |
+
warnings.warn(f"video path not found {fname=}")
|
| 9 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_poker/-jGxlNQKkeo_000092_000102.avi'
|
| 10 |
+
warnings.warn(f"video path not found {fname=}")
|
| 11 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/bowling/f4eb0wOlspM_000053_000063.avi'
|
| 12 |
+
warnings.warn(f"video path not found {fname=}")
|
| 13 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/pumping_fist/wKsuyr6Xk30_000094_000104.avi'
|
| 14 |
+
warnings.warn(f"video path not found {fname=}")
|
| 15 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/robot_dancing/J1x0XlWa6HM_000147_000157.avi'
|
| 16 |
+
warnings.warn(f"video path not found {fname=}")
|
| 17 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 18 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 19 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/baking_cookies/Q3XGZmqk1Q0_000296_000306.avi'
|
| 20 |
+
warnings.warn(f"video path not found {fname=}")
|
| 21 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 22 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 23 |
+
[22:34:34] /github/workspace/src/video/video_reader.cc:83: ERROR opening: /local/dcanezil/data/kinetics/VideoData/climbing_tree/QX5IzK4iFCM_000002_000012.avi, Invalid data found when processing input
|
| 24 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/robot_dancing/J1x0XlWa6HM_000147_000157.avi'
|
| 25 |
+
warnings.warn(f"video path not found {fname=}")
|
| 26 |
+
[22:58:07] /github/workspace/src/video/video_reader.cc:83: ERROR opening: /local/dcanezil/data/kinetics/VideoData/trimming_trees/T6dqaZioaXs_000004_000014.avi, Invalid data found when processing input
|
| 27 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/picking_fruit/3VvkoFtPCCU_000045_000055.avi'
|
| 28 |
+
warnings.warn(f"video path not found {fname=}")
|
| 29 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/windsurfing/i-gzh_BPDa8_000154_000164.avi'
|
| 30 |
+
warnings.warn(f"video path not found {fname=}")
|
| 31 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_monopoly/NLL667uPWVA_000066_000076.avi'
|
| 32 |
+
warnings.warn(f"video path not found {fname=}")
|
| 33 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/making_tea/mtYFNsRcxY4_000063_000073.avi'
|
| 34 |
+
warnings.warn(f"video path not found {fname=}")
|
| 35 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_poker/6-DyF8umej8_000097_000107.avi'
|
| 36 |
+
warnings.warn(f"video path not found {fname=}")
|
| 37 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/surfing_water/5l5Pdd96Pao_000161_000171.avi'
|
| 38 |
+
warnings.warn(f"video path not found {fname=}")
|
| 39 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_poker/-jGxlNQKkeo_000092_000102.avi'
|
| 40 |
+
warnings.warn(f"video path not found {fname=}")
|
| 41 |
+
[02:07:20] /github/workspace/src/video/video_reader.cc:83: ERROR opening: /local/dcanezil/data/kinetics/VideoData/trimming_trees/T6dqaZioaXs_000004_000014.avi, Invalid data found when processing input
|
| 42 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 43 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 44 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_poker/-jGxlNQKkeo_000092_000102.avi'
|
| 45 |
+
warnings.warn(f"video path not found {fname=}")
|
| 46 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_poker/6-DyF8umej8_000097_000107.avi'
|
| 47 |
+
warnings.warn(f"video path not found {fname=}")
|
| 48 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_poker/6-DyF8umej8_000097_000107.avi'
|
| 49 |
+
warnings.warn(f"video path not found {fname=}")
|
| 50 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/windsurfing/i-gzh_BPDa8_000154_000164.avi'
|
| 51 |
+
warnings.warn(f"video path not found {fname=}")
|
| 52 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_poker/6-DyF8umej8_000097_000107.avi'
|
| 53 |
+
warnings.warn(f"video path not found {fname=}")
|
| 54 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/waxing_legs/cqUusXBODuw_001046_001056.avi'
|
| 55 |
+
warnings.warn(f"video path not found {fname=}")
|
| 56 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 57 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 58 |
+
[06:39:37] /github/workspace/src/video/video_reader.cc:83: ERROR opening: /local/dcanezil/data/kinetics/VideoData/long_jump/jsYiitkbr5w_000004_000014.avi, Invalid data found when processing input
|
| 59 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/picking_fruit/3VvkoFtPCCU_000045_000055.avi'
|
| 60 |
+
warnings.warn(f"video path not found {fname=}")
|
| 61 |
+
[07:56:42] /github/workspace/src/video/video_reader.cc:83: ERROR opening: /local/dcanezil/data/kinetics/VideoData/trimming_trees/T6dqaZioaXs_000004_000014.avi, Invalid data found when processing input
|
| 62 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_monopoly/NLL667uPWVA_000066_000076.avi'
|
| 63 |
+
warnings.warn(f"video path not found {fname=}")
|
| 64 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/eating_carrots/eiZ8Hzc7FPU_000080_000090.avi'
|
| 65 |
+
warnings.warn(f"video path not found {fname=}")
|
| 66 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/windsurfing/i-gzh_BPDa8_000154_000164.avi'
|
| 67 |
+
warnings.warn(f"video path not found {fname=}")
|
| 68 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_flute/co50KUHacYw_000005_000015.avi'
|
| 69 |
+
warnings.warn(f"video path not found {fname=}")
|
| 70 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/waxing_legs/cqUusXBODuw_001046_001056.avi'
|
| 71 |
+
warnings.warn(f"video path not found {fname=}")
|
| 72 |
+
[10:02:45] /github/workspace/src/video/video_reader.cc:83: ERROR opening: /local/dcanezil/data/kinetics/VideoData/trimming_trees/T6dqaZioaXs_000004_000014.avi, Invalid data found when processing input
|
| 73 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/frying_vegetables/Nr-zCXgW2Sw_000161_000171.avi'
|
| 74 |
+
warnings.warn(f"video path not found {fname=}")
|
| 75 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/app/vjepa/train.py:983: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
|
| 76 |
+
with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
|
| 77 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_saxophone/Zs0_2GMEPXo_000054_000064.avi'
|
| 78 |
+
warnings.warn(f"video path not found {fname=}")
|
| 79 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/surfing_water/5l5Pdd96Pao_000161_000171.avi'
|
| 80 |
+
warnings.warn(f"video path not found {fname=}")
|
| 81 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/baking_cookies/Q3XGZmqk1Q0_000296_000306.avi'
|
| 82 |
+
warnings.warn(f"video path not found {fname=}")
|
| 83 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/pumping_fist/wKsuyr6Xk30_000094_000104.avi'
|
| 84 |
+
warnings.warn(f"video path not found {fname=}")
|
| 85 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/picking_fruit/3VvkoFtPCCU_000045_000055.avi'
|
| 86 |
+
warnings.warn(f"video path not found {fname=}")
|
| 87 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/playing_flute/co50KUHacYw_000005_000015.avi'
|
| 88 |
+
warnings.warn(f"video path not found {fname=}")
|
| 89 |
+
/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/code/src/datasets/video_dataset.py:418: UserWarning: video path not found fname='/local/dcanezil/data/kinetics/VideoData/frying_vegetables/Nr-zCXgW2Sw_000161_000171.avi'
|
| 90 |
+
warnings.warn(f"video path not found {fname=}")
|
| 91 |
+
[rank1]:[W710 14:14:33.737608791 ProcessGroupNCCL.cpp:1553] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())
|
inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_27070/27070_1_log.out
ADDED
|
@@ -0,0 +1,218 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
submitit INFO (2026-07-09 14:38:05,873) - Starting with JobEnvironment(job_id=27070, hostname=node404, local_rank=1(3), node=0(1), global_rank=1(3))
|
| 2 |
+
submitit INFO (2026-07-09 14:38:05,873) - Loading pickle: /var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/job_27070/27070_submitted.pkl
|
| 3 |
+
INFO:root:loaded pretrain params...
|
| 4 |
+
{ 'app': 'vjepa',
|
| 5 |
+
'cpus_per_task': 16,
|
| 6 |
+
'data': { 'batch_size': 64,
|
| 7 |
+
'crop_size': 224,
|
| 8 |
+
'dataset_fpcs': [8, 8],
|
| 9 |
+
'dataset_type': 'VideoDataset',
|
| 10 |
+
'datasets': [ '/local/dcanezil/data/kinetics/train.csv',
|
| 11 |
+
'/local/dcanezil/data/ssv2/train.csv'],
|
| 12 |
+
'datasets_weights': [0.65, 0.35],
|
| 13 |
+
'fps': 2,
|
| 14 |
+
'num_workers': 4,
|
| 15 |
+
'patch_size': 14,
|
| 16 |
+
'persistent_workers': True,
|
| 17 |
+
'pin_mem': True,
|
| 18 |
+
'tubelet_size': 1},
|
| 19 |
+
'data_aug': { 'auto_augment': False,
|
| 20 |
+
'motion_shift': False,
|
| 21 |
+
'random_resize_aspect_ratio': [0.75, 1.35],
|
| 22 |
+
'random_resize_scale': [0.3, 1.0],
|
| 23 |
+
'reprob': 0.0},
|
| 24 |
+
'folder': '/var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow',
|
| 25 |
+
'loss': {'loss_exp': 1.0},
|
| 26 |
+
'mask': [ { 'aspect_ratio': [0.75, 1.5],
|
| 27 |
+
'full_complement': False,
|
| 28 |
+
'max_keep': None,
|
| 29 |
+
'max_temporal_keep': 1.0,
|
| 30 |
+
'num_blocks': 8,
|
| 31 |
+
'spatial_scale': [0.05, 0.05],
|
| 32 |
+
'temporal_scale': [1.0, 1.0]},
|
| 33 |
+
{ 'aspect_ratio': [0.75, 1.5],
|
| 34 |
+
'full_complement': False,
|
| 35 |
+
'max_keep': None,
|
| 36 |
+
'max_temporal_keep': 1.0,
|
| 37 |
+
'num_blocks': 2,
|
| 38 |
+
'spatial_scale': [0.4, 0.4],
|
| 39 |
+
'temporal_scale': [1.0, 1.0]}],
|
| 40 |
+
'mem_per_gpu': '40G',
|
| 41 |
+
'meta': { 'dtype': 'bfloat16',
|
| 42 |
+
'knn_eval_epoch0': True,
|
| 43 |
+
'knn_eval_freq': 5,
|
| 44 |
+
'knn_eval_presets': [ { 'config': { 'batch_size': 8,
|
| 45 |
+
'dataset_train': '/local/dcanezil/data/ssv2/train_coarse10.csv',
|
| 46 |
+
'dataset_val': '/local/dcanezil/data/ssv2/val_coarse10.csv',
|
| 47 |
+
'eval_videos_per_class': 100,
|
| 48 |
+
'linear_probe': True,
|
| 49 |
+
'num_workers': 2,
|
| 50 |
+
'pool_type': 'temporal_concat',
|
| 51 |
+
'train_videos_per_class': 500},
|
| 52 |
+
'preset': 'ssv2_coarse10'}],
|
| 53 |
+
'load_checkpoint': True,
|
| 54 |
+
'read_checkpoint': None,
|
| 55 |
+
'save_every_freq': 10,
|
| 56 |
+
'seed': 239,
|
| 57 |
+
'use_sdpa': True,
|
| 58 |
+
'use_wandb': True,
|
| 59 |
+
'wandb_project': 'vjepa_ablation'},
|
| 60 |
+
'metrics': {'sigreg': {}, 'std': {}},
|
| 61 |
+
'model': { 'class_token': True,
|
| 62 |
+
'freeze_backbone': False,
|
| 63 |
+
'is_causal': True,
|
| 64 |
+
'model_name': 'vit_large_patch14_dinov2_lvd142m',
|
| 65 |
+
'pe_type': '2d',
|
| 66 |
+
'pred_depth': 12,
|
| 67 |
+
'pred_embed_dim': 384,
|
| 68 |
+
'pred_num_heads': 6,
|
| 69 |
+
'predictor': 'v2_cross',
|
| 70 |
+
'stem_type': '2d',
|
| 71 |
+
'target_embed_dim': 1024,
|
| 72 |
+
'target_kind': 'ema',
|
| 73 |
+
'uniform_power': True,
|
| 74 |
+
'use_activation_checkpointing': True,
|
| 75 |
+
'use_alibi': False,
|
| 76 |
+
'use_lora': False,
|
| 77 |
+
'use_mask_tokens': True,
|
| 78 |
+
'use_rope': False,
|
| 79 |
+
'use_sdpa': True,
|
| 80 |
+
'zero_init_mask_tokens': True},
|
| 81 |
+
'nodes': 1,
|
| 82 |
+
'optimization': { 'clip_grad': None,
|
| 83 |
+
'ema': [1.0, 0.99925],
|
| 84 |
+
'ema_hold_epochs': 0,
|
| 85 |
+
'ema_ramp_epochs': 5,
|
| 86 |
+
'encoder_layer_decay': 0.9,
|
| 87 |
+
'epochs': 150,
|
| 88 |
+
'final_lr': 0.0001,
|
| 89 |
+
'final_weight_decay': 0.04,
|
| 90 |
+
'ipe': 300,
|
| 91 |
+
'ipe_scale': 1.0,
|
| 92 |
+
'lr': 0.0001,
|
| 93 |
+
'predictor_lr_scale': 3.0,
|
| 94 |
+
'start_lr': 1e-05,
|
| 95 |
+
'warmup': 2,
|
| 96 |
+
'weight_decay': 0.04},
|
| 97 |
+
'tasks_per_node': 3}
|
| 98 |
+
INFO:root:Running pre-training of app: vjepa
|
| 99 |
+
[INFO ][2026-07-09 14:38:14][app.vjepa.train ][main ] which_dtype='bfloat16'
|
| 100 |
+
[INFO ][2026-07-09 14:38:14][app.vjepa.train ][main ] Disabling persistent_workers (incompatible with KNN eval)
|
| 101 |
+
[INFO ][2026-07-09 14:38:14][app.vjepa.train ][main ] NCCL_SOCKET_IFNAME=None
|
| 102 |
+
[INFO ][2026-07-09 14:38:17][app.vjepa.train ][main ] Initialized (rank/world-size) 1/3, tasks_per_node=3
|
| 103 |
+
[WARNING ][2026-07-09 14:38:56][experiments.stmodels.vision_transformers_v3][_build_dinov2_vit ] DINOv2 is an image model — fwd_chunk is unset so all frames will be processed jointly. Set fwd_chunk=1 for per-frame (standard) processing.
|
| 104 |
+
[INFO ][2026-07-09 14:39:01][experiments.stmodels.vision_transformers_v3][_build_dinov2_vit ] Loading pretrained weights for vit_large_patch14_dinov2_lvd142m from timm (vit_large_patch14_dinov2)
|
| 105 |
+
[INFO ][2026-07-09 14:39:04][timm.models._builder][load_pretrained ] Loading pretrained weights from Hugging Face hub (timm/vit_large_patch14_dinov2.lvd142m)
|
| 106 |
+
[INFO ][2026-07-09 14:39:04][timm.models._hub ][load_state_dict_from_hf ] [timm/vit_large_patch14_dinov2.lvd142m] Safe alternative available for 'pytorch_model.bin' (as 'model.safetensors'). Loading weights using safetensors.
|
| 107 |
+
[INFO ][2026-07-09 14:39:04][timm.layers.pos_embed][resample_abs_pos_embed ] Resized position embedding: (37, 37) to (16, 16).
|
| 108 |
+
[INFO ][2026-07-09 14:39:07][root ][init_video_model ] ViTMultiSeqWrapper(
|
| 109 |
+
(backbone): VisionTransformer(
|
| 110 |
+
(patch_embed): PatchEmbed3D(
|
| 111 |
+
(proj): Conv3d(3, 1024, kernel_size=(1, 14, 14), stride=(1, 14, 14))
|
| 112 |
+
)
|
| 113 |
+
(blocks): ModuleList(
|
| 114 |
+
(0-23): 24 x Block(
|
| 115 |
+
(norm1): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 116 |
+
(norm2): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 117 |
+
(drop_path1): Identity()
|
| 118 |
+
(drop_path2): Identity()
|
| 119 |
+
(attn): Attention(
|
| 120 |
+
(qkv): Linear(in_features=1024, out_features=3072, bias=True)
|
| 121 |
+
(attn_drop): Dropout(p=0.0, inplace=False)
|
| 122 |
+
(proj): Linear(in_features=1024, out_features=1024, bias=True)
|
| 123 |
+
(proj_drop): Dropout(p=0.0, inplace=False)
|
| 124 |
+
)
|
| 125 |
+
(mlp): MLP(
|
| 126 |
+
(fc1): Linear(in_features=1024, out_features=4096, bias=True)
|
| 127 |
+
(act): GELU(approximate='none')
|
| 128 |
+
(fc2): Linear(in_features=4096, out_features=1024, bias=True)
|
| 129 |
+
(drop): Dropout(p=0.0, inplace=False)
|
| 130 |
+
)
|
| 131 |
+
)
|
| 132 |
+
)
|
| 133 |
+
(norm): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 134 |
+
)
|
| 135 |
+
)
|
| 136 |
+
[INFO ][2026-07-09 14:39:07][root ][init_video_model ] VJEPAPredictorMultiSeqWrapper(
|
| 137 |
+
(backbone): PredictorV2(
|
| 138 |
+
(predictor_embed): Linear(in_features=1024, out_features=384, bias=True)
|
| 139 |
+
(mask_tokens): ParameterList(
|
| 140 |
+
(0): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
|
| 141 |
+
(1): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
|
| 142 |
+
(2): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
|
| 143 |
+
(3): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
|
| 144 |
+
)
|
| 145 |
+
(predictor_blocks): ModuleList(
|
| 146 |
+
(0-11): 12 x Block(
|
| 147 |
+
(residual1): EfficientResidual(
|
| 148 |
+
(norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
|
| 149 |
+
(fn): Attention(
|
| 150 |
+
(q_proj): Linear(in_features=384, out_features=384, bias=False)
|
| 151 |
+
(k_proj): Linear(in_features=384, out_features=384, bias=False)
|
| 152 |
+
(v_proj): Linear(in_features=384, out_features=384, bias=False)
|
| 153 |
+
(proj): Linear(in_features=384, out_features=384, bias=False)
|
| 154 |
+
(rope): Rope()
|
| 155 |
+
)
|
| 156 |
+
)
|
| 157 |
+
(residual2): EfficientResidual(
|
| 158 |
+
(norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
|
| 159 |
+
(fn): MLP(
|
| 160 |
+
(fc1): Linear(in_features=384, out_features=1536, bias=False)
|
| 161 |
+
(act): GELU(approximate='none')
|
| 162 |
+
(fc2): Linear(in_features=1536, out_features=384, bias=False)
|
| 163 |
+
(drop): Dropout(p=0.0, inplace=False)
|
| 164 |
+
)
|
| 165 |
+
)
|
| 166 |
+
)
|
| 167 |
+
)
|
| 168 |
+
(predictor_norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
|
| 169 |
+
(predictor_proj): Linear(in_features=384, out_features=1024, bias=True)
|
| 170 |
+
)
|
| 171 |
+
)
|
| 172 |
+
[INFO ][2026-07-09 14:39:07][root ][init_video_model ] ViTMultiSeqWrapper(
|
| 173 |
+
(backbone): VisionTransformer(
|
| 174 |
+
(patch_embed): PatchEmbed3D(
|
| 175 |
+
(proj): Conv3d(3, 1024, kernel_size=(1, 14, 14), stride=(1, 14, 14))
|
| 176 |
+
)
|
| 177 |
+
(blocks): ModuleList(
|
| 178 |
+
(0-23): 24 x Block(
|
| 179 |
+
(norm1): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 180 |
+
(norm2): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 181 |
+
(drop_path1): Identity()
|
| 182 |
+
(drop_path2): Identity()
|
| 183 |
+
(attn): Attention(
|
| 184 |
+
(qkv): Linear(in_features=1024, out_features=3072, bias=True)
|
| 185 |
+
(attn_drop): Dropout(p=0.0, inplace=False)
|
| 186 |
+
(proj): Linear(in_features=1024, out_features=1024, bias=True)
|
| 187 |
+
(proj_drop): Dropout(p=0.0, inplace=False)
|
| 188 |
+
)
|
| 189 |
+
(mlp): MLP(
|
| 190 |
+
(fc1): Linear(in_features=1024, out_features=4096, bias=True)
|
| 191 |
+
(act): GELU(approximate='none')
|
| 192 |
+
(fc2): Linear(in_features=4096, out_features=1024, bias=True)
|
| 193 |
+
(drop): Dropout(p=0.0, inplace=False)
|
| 194 |
+
)
|
| 195 |
+
)
|
| 196 |
+
)
|
| 197 |
+
(norm): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
|
| 198 |
+
)
|
| 199 |
+
)
|
| 200 |
+
[INFO ][2026-07-09 14:39:07][root ][init_video_model ] Encoder number of parameters: 302964736
|
| 201 |
+
[INFO ][2026-07-09 14:39:07][root ][init_video_model ] Predictor number of parameters: 22042240
|
| 202 |
+
[INFO ][2026-07-09 14:39:07][root ][init_video_model ] Target encoder number of parameters: 302964736
|
| 203 |
+
[INFO ][2026-07-09 14:39:08][root ][make_videodataset ] VideoDataset dataset created
|
| 204 |
+
[INFO ][2026-07-09 14:39:08][WeightedSampler ][__init__ ] Using DistributedWeightedSampler with rank 1 / 3
|
| 205 |
+
[INFO ][2026-07-09 14:39:08][root ][make_videodataset ] VideoDataset unsupervised data loader created
|
| 206 |
+
[INFO ][2026-07-09 14:39:08][app.vjepa.train ][main ] iterations per epoch/dataset length: 300/2103
|
| 207 |
+
[INFO ][2026-07-09 14:39:08][app.vjepa.train ][main ] Wrapping models in DDP (rank 1)...
|
| 208 |
+
[INFO ][2026-07-09 14:39:09][root ][load_checkpoint ] Loading checkpoint from /var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/latest.pt
|
| 209 |
+
[INFO ][2026-07-09 14:39:17][root ][load_checkpoint ] loaded pretrained encoder from epoch 120 with msg: <All keys matched successfully>
|
| 210 |
+
[INFO ][2026-07-09 14:39:17][root ][load_checkpoint ] loaded pretrained predictor from epoch 120 with msg: <All keys matched successfully>
|
| 211 |
+
[INFO ][2026-07-09 14:39:17][root ][load_checkpoint ] loaded pretrained target encoder from epoch 120 with msg: <All keys matched successfully>
|
| 212 |
+
[INFO ][2026-07-09 14:39:18][root ][load_checkpoint ] loaded optimizers from epoch 120
|
| 213 |
+
[INFO ][2026-07-09 14:39:18][root ][load_checkpoint ] read-path: /var/scratch/dcanezil/runs/vjepa/inflated/d073_full_ema_dinov2l_ssv2k400_8f_emainv_ramp5_predlr3_llrd0p9_masklow/latest.pt
|
| 214 |
+
[INFO ][2026-07-09 14:40:37][app.vjepa.train ][main ] Initializing loader...
|
| 215 |
+
submitit INFO (2026-07-10 14:14:32,240) - Job completed successfully
|
| 216 |
+
[INFO ][2026-07-10 14:14:32][submitit ][process_job ] Job completed successfully
|
| 217 |
+
submitit INFO (2026-07-10 14:14:32,242) - Exiting after successful completion
|
| 218 |
+
[INFO ][2026-07-10 14:14:32][submitit ][process_job ] Exiting after successful completion
|