dgcnz commited on
Commit
49afa9b
·
verified ·
1 Parent(s): 4d35c15

sync run artifacts: vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable

Browse files
This view is limited to 50 files because it contains too many changes.   See raw diff
Files changed (50) hide show
  1. vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/best.pt +3 -0
  2. vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/crash_rank0.log +29 -0
  3. vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/crash_rank1.log +29 -0
  4. vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/crash_rank10.log +29 -0
  5. vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/crash_rank11.log +29 -0
  6. vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/crash_rank12.log +29 -0
  7. vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/crash_rank13.log +29 -0
  8. vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/crash_rank14.log +29 -0
  9. vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/crash_rank15.log +29 -0
  10. vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/crash_rank2.log +29 -0
  11. vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/crash_rank3.log +29 -0
  12. vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/crash_rank4.log +29 -0
  13. vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/crash_rank5.log +29 -0
  14. vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/crash_rank6.log +29 -0
  15. vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/crash_rank7.log +29 -0
  16. vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/crash_rank8.log +29 -0
  17. vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/crash_rank9.log +29 -0
  18. vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/git-info.txt +2 -0
  19. vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/job_22614766/22614766_0_log.err +77 -0
  20. vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/job_22614766/22614766_0_log.out +0 -0
  21. vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/job_22614766/22614766_10_log.err +68 -0
  22. vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/job_22614766/22614766_10_log.out +276 -0
  23. vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/job_22614766/22614766_11_log.err +158 -0
  24. vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/job_22614766/22614766_11_log.out +280 -0
  25. vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/job_22614766/22614766_12_log.err +158 -0
  26. vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/job_22614766/22614766_12_log.out +284 -0
  27. vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/job_22614766/22614766_13_log.err +180 -0
  28. vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/job_22614766/22614766_13_log.out +280 -0
  29. vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/job_22614766/22614766_14_log.err +79 -0
  30. vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/job_22614766/22614766_14_log.out +276 -0
  31. vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/job_22614766/22614766_15_log.err +68 -0
  32. vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/job_22614766/22614766_15_log.out +276 -0
  33. vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/job_22614766/22614766_1_log.err +41 -0
  34. vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/job_22614766/22614766_1_log.out +276 -0
  35. vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/job_22614766/22614766_2_log.err +41 -0
  36. vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/job_22614766/22614766_2_log.out +276 -0
  37. vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/job_22614766/22614766_3_log.err +41 -0
  38. vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/job_22614766/22614766_3_log.out +276 -0
  39. vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/job_22614766/22614766_4_log.err +180 -0
  40. vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/job_22614766/22614766_4_log.out +284 -0
  41. vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/job_22614766/22614766_5_log.err +158 -0
  42. vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/job_22614766/22614766_5_log.out +280 -0
  43. vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/job_22614766/22614766_6_log.err +68 -0
  44. vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/job_22614766/22614766_6_log.out +276 -0
  45. vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/job_22614766/22614766_7_log.err +68 -0
  46. vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/job_22614766/22614766_7_log.out +276 -0
  47. vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/job_22614766/22614766_8_log.err +68 -0
  48. vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/job_22614766/22614766_8_log.out +280 -0
  49. vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/job_22614766/22614766_9_log.err +191 -0
  50. vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/job_22614766/22614766_9_log.out +280 -0
vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/best.pt ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:e173a3e93f86a39c26ad135b2c2127d6a1435f88a18d51f9e41a456b94c6f3e8
3
+ size 6188235489
vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/crash_rank0.log ADDED
@@ -0,0 +1,29 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ Rank 0 crashed at epoch 15, itr 250
2
+ Error: loss is nan
3
+
4
+ Traceback (most recent call last):
5
+ File "<frozen runpy>", line 198, in _run_module_as_main
6
+ File "<frozen runpy>", line 88, in _run_code
7
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/_submit.py", line 11, in <module>
8
+ submitit_main()
9
+ ~~~~~~~~~~~~~^^
10
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 76, in submitit_main
11
+ process_job(args.folder)
12
+ ~~~~~~~~~~~^^^^^^^^^^^^^
13
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 69, in process_job
14
+ raise error
15
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 55, in process_job
16
+ result = delayed.result()
17
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/utils.py", line 137, in result
18
+ self._result = self.function(*self.args, **self.kwargs)
19
+ ~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^
20
+ File "/gpfs/home2/dcanez/vd/app/main_distributed.py", line 102, in __call__
21
+ app_main(app, args=params, resume_preempt=resume_preempt)
22
+ ~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
23
+ File "/gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/scaffold.py", line 17, in main
24
+ return importlib.import_module(f"app.{app}.train").main(args=args, resume_preempt=resume_preempt)
25
+ ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
26
+ File "/gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/vjepa/train.py", line 1189, in main
27
+ assert not np.isnan(loss), "loss is nan"
28
+ ^^^^^^^^^^^^^^^^^^
29
+ AssertionError: loss is nan
vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/crash_rank1.log ADDED
@@ -0,0 +1,29 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ Rank 1 crashed at epoch 15, itr 250
2
+ Error: loss is nan
3
+
4
+ Traceback (most recent call last):
5
+ File "<frozen runpy>", line 198, in _run_module_as_main
6
+ File "<frozen runpy>", line 88, in _run_code
7
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/_submit.py", line 11, in <module>
8
+ submitit_main()
9
+ ~~~~~~~~~~~~~^^
10
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 76, in submitit_main
11
+ process_job(args.folder)
12
+ ~~~~~~~~~~~^^^^^^^^^^^^^
13
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 69, in process_job
14
+ raise error
15
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 55, in process_job
16
+ result = delayed.result()
17
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/utils.py", line 137, in result
18
+ self._result = self.function(*self.args, **self.kwargs)
19
+ ~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^
20
+ File "/gpfs/home2/dcanez/vd/app/main_distributed.py", line 102, in __call__
21
+ app_main(app, args=params, resume_preempt=resume_preempt)
22
+ ~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
23
+ File "/gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/scaffold.py", line 17, in main
24
+ return importlib.import_module(f"app.{app}.train").main(args=args, resume_preempt=resume_preempt)
25
+ ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
26
+ File "/gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/vjepa/train.py", line 1189, in main
27
+ assert not np.isnan(loss), "loss is nan"
28
+ ^^^^^^^^^^^^^^^^^^
29
+ AssertionError: loss is nan
vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/crash_rank10.log ADDED
@@ -0,0 +1,29 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ Rank 10 crashed at epoch 15, itr 250
2
+ Error: loss is nan
3
+
4
+ Traceback (most recent call last):
5
+ File "<frozen runpy>", line 198, in _run_module_as_main
6
+ File "<frozen runpy>", line 88, in _run_code
7
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/_submit.py", line 11, in <module>
8
+ submitit_main()
9
+ ~~~~~~~~~~~~~^^
10
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 76, in submitit_main
11
+ process_job(args.folder)
12
+ ~~~~~~~~~~~^^^^^^^^^^^^^
13
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 69, in process_job
14
+ raise error
15
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 55, in process_job
16
+ result = delayed.result()
17
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/utils.py", line 137, in result
18
+ self._result = self.function(*self.args, **self.kwargs)
19
+ ~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^
20
+ File "/gpfs/home2/dcanez/vd/app/main_distributed.py", line 102, in __call__
21
+ app_main(app, args=params, resume_preempt=resume_preempt)
22
+ ~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
23
+ File "/gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/scaffold.py", line 17, in main
24
+ return importlib.import_module(f"app.{app}.train").main(args=args, resume_preempt=resume_preempt)
25
+ ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
26
+ File "/gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/vjepa/train.py", line 1189, in main
27
+ assert not np.isnan(loss), "loss is nan"
28
+ ^^^^^^^^^^^^^^^^^^
29
+ AssertionError: loss is nan
vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/crash_rank11.log ADDED
@@ -0,0 +1,29 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ Rank 11 crashed at epoch 15, itr 250
2
+ Error: loss is nan
3
+
4
+ Traceback (most recent call last):
5
+ File "<frozen runpy>", line 198, in _run_module_as_main
6
+ File "<frozen runpy>", line 88, in _run_code
7
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/_submit.py", line 11, in <module>
8
+ submitit_main()
9
+ ~~~~~~~~~~~~~^^
10
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 76, in submitit_main
11
+ process_job(args.folder)
12
+ ~~~~~~~~~~~^^^^^^^^^^^^^
13
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 69, in process_job
14
+ raise error
15
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 55, in process_job
16
+ result = delayed.result()
17
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/utils.py", line 137, in result
18
+ self._result = self.function(*self.args, **self.kwargs)
19
+ ~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^
20
+ File "/gpfs/home2/dcanez/vd/app/main_distributed.py", line 102, in __call__
21
+ app_main(app, args=params, resume_preempt=resume_preempt)
22
+ ~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
23
+ File "/gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/scaffold.py", line 17, in main
24
+ return importlib.import_module(f"app.{app}.train").main(args=args, resume_preempt=resume_preempt)
25
+ ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
26
+ File "/gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/vjepa/train.py", line 1189, in main
27
+ assert not np.isnan(loss), "loss is nan"
28
+ ^^^^^^^^^^^^^^^^^^
29
+ AssertionError: loss is nan
vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/crash_rank12.log ADDED
@@ -0,0 +1,29 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ Rank 12 crashed at epoch 15, itr 250
2
+ Error: loss is nan
3
+
4
+ Traceback (most recent call last):
5
+ File "<frozen runpy>", line 198, in _run_module_as_main
6
+ File "<frozen runpy>", line 88, in _run_code
7
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/_submit.py", line 11, in <module>
8
+ submitit_main()
9
+ ~~~~~~~~~~~~~^^
10
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 76, in submitit_main
11
+ process_job(args.folder)
12
+ ~~~~~~~~~~~^^^^^^^^^^^^^
13
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 69, in process_job
14
+ raise error
15
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 55, in process_job
16
+ result = delayed.result()
17
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/utils.py", line 137, in result
18
+ self._result = self.function(*self.args, **self.kwargs)
19
+ ~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^
20
+ File "/gpfs/home2/dcanez/vd/app/main_distributed.py", line 102, in __call__
21
+ app_main(app, args=params, resume_preempt=resume_preempt)
22
+ ~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
23
+ File "/gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/scaffold.py", line 17, in main
24
+ return importlib.import_module(f"app.{app}.train").main(args=args, resume_preempt=resume_preempt)
25
+ ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
26
+ File "/gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/vjepa/train.py", line 1189, in main
27
+ assert not np.isnan(loss), "loss is nan"
28
+ ^^^^^^^^^^^^^^^^^^
29
+ AssertionError: loss is nan
vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/crash_rank13.log ADDED
@@ -0,0 +1,29 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ Rank 13 crashed at epoch 15, itr 250
2
+ Error: loss is nan
3
+
4
+ Traceback (most recent call last):
5
+ File "<frozen runpy>", line 198, in _run_module_as_main
6
+ File "<frozen runpy>", line 88, in _run_code
7
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/_submit.py", line 11, in <module>
8
+ submitit_main()
9
+ ~~~~~~~~~~~~~^^
10
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 76, in submitit_main
11
+ process_job(args.folder)
12
+ ~~~~~~~~~~~^^^^^^^^^^^^^
13
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 69, in process_job
14
+ raise error
15
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 55, in process_job
16
+ result = delayed.result()
17
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/utils.py", line 137, in result
18
+ self._result = self.function(*self.args, **self.kwargs)
19
+ ~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^
20
+ File "/gpfs/home2/dcanez/vd/app/main_distributed.py", line 102, in __call__
21
+ app_main(app, args=params, resume_preempt=resume_preempt)
22
+ ~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
23
+ File "/gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/scaffold.py", line 17, in main
24
+ return importlib.import_module(f"app.{app}.train").main(args=args, resume_preempt=resume_preempt)
25
+ ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
26
+ File "/gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/vjepa/train.py", line 1189, in main
27
+ assert not np.isnan(loss), "loss is nan"
28
+ ^^^^^^^^^^^^^^^^^^
29
+ AssertionError: loss is nan
vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/crash_rank14.log ADDED
@@ -0,0 +1,29 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ Rank 14 crashed at epoch 15, itr 250
2
+ Error: loss is nan
3
+
4
+ Traceback (most recent call last):
5
+ File "<frozen runpy>", line 198, in _run_module_as_main
6
+ File "<frozen runpy>", line 88, in _run_code
7
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/_submit.py", line 11, in <module>
8
+ submitit_main()
9
+ ~~~~~~~~~~~~~^^
10
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 76, in submitit_main
11
+ process_job(args.folder)
12
+ ~~~~~~~~~~~^^^^^^^^^^^^^
13
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 69, in process_job
14
+ raise error
15
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 55, in process_job
16
+ result = delayed.result()
17
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/utils.py", line 137, in result
18
+ self._result = self.function(*self.args, **self.kwargs)
19
+ ~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^
20
+ File "/gpfs/home2/dcanez/vd/app/main_distributed.py", line 102, in __call__
21
+ app_main(app, args=params, resume_preempt=resume_preempt)
22
+ ~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
23
+ File "/gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/scaffold.py", line 17, in main
24
+ return importlib.import_module(f"app.{app}.train").main(args=args, resume_preempt=resume_preempt)
25
+ ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
26
+ File "/gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/vjepa/train.py", line 1189, in main
27
+ assert not np.isnan(loss), "loss is nan"
28
+ ^^^^^^^^^^^^^^^^^^
29
+ AssertionError: loss is nan
vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/crash_rank15.log ADDED
@@ -0,0 +1,29 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ Rank 15 crashed at epoch 15, itr 250
2
+ Error: loss is nan
3
+
4
+ Traceback (most recent call last):
5
+ File "<frozen runpy>", line 198, in _run_module_as_main
6
+ File "<frozen runpy>", line 88, in _run_code
7
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/_submit.py", line 11, in <module>
8
+ submitit_main()
9
+ ~~~~~~~~~~~~~^^
10
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 76, in submitit_main
11
+ process_job(args.folder)
12
+ ~~~~~~~~~~~^^^^^^^^^^^^^
13
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 69, in process_job
14
+ raise error
15
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 55, in process_job
16
+ result = delayed.result()
17
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/utils.py", line 137, in result
18
+ self._result = self.function(*self.args, **self.kwargs)
19
+ ~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^
20
+ File "/gpfs/home2/dcanez/vd/app/main_distributed.py", line 102, in __call__
21
+ app_main(app, args=params, resume_preempt=resume_preempt)
22
+ ~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
23
+ File "/gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/scaffold.py", line 17, in main
24
+ return importlib.import_module(f"app.{app}.train").main(args=args, resume_preempt=resume_preempt)
25
+ ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
26
+ File "/gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/vjepa/train.py", line 1189, in main
27
+ assert not np.isnan(loss), "loss is nan"
28
+ ^^^^^^^^^^^^^^^^^^
29
+ AssertionError: loss is nan
vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/crash_rank2.log ADDED
@@ -0,0 +1,29 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ Rank 2 crashed at epoch 15, itr 250
2
+ Error: loss is nan
3
+
4
+ Traceback (most recent call last):
5
+ File "<frozen runpy>", line 198, in _run_module_as_main
6
+ File "<frozen runpy>", line 88, in _run_code
7
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/_submit.py", line 11, in <module>
8
+ submitit_main()
9
+ ~~~~~~~~~~~~~^^
10
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 76, in submitit_main
11
+ process_job(args.folder)
12
+ ~~~~~~~~~~~^^^^^^^^^^^^^
13
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 69, in process_job
14
+ raise error
15
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 55, in process_job
16
+ result = delayed.result()
17
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/utils.py", line 137, in result
18
+ self._result = self.function(*self.args, **self.kwargs)
19
+ ~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^
20
+ File "/gpfs/home2/dcanez/vd/app/main_distributed.py", line 102, in __call__
21
+ app_main(app, args=params, resume_preempt=resume_preempt)
22
+ ~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
23
+ File "/gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/scaffold.py", line 17, in main
24
+ return importlib.import_module(f"app.{app}.train").main(args=args, resume_preempt=resume_preempt)
25
+ ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
26
+ File "/gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/vjepa/train.py", line 1189, in main
27
+ assert not np.isnan(loss), "loss is nan"
28
+ ^^^^^^^^^^^^^^^^^^
29
+ AssertionError: loss is nan
vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/crash_rank3.log ADDED
@@ -0,0 +1,29 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ Rank 3 crashed at epoch 15, itr 250
2
+ Error: loss is nan
3
+
4
+ Traceback (most recent call last):
5
+ File "<frozen runpy>", line 198, in _run_module_as_main
6
+ File "<frozen runpy>", line 88, in _run_code
7
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/_submit.py", line 11, in <module>
8
+ submitit_main()
9
+ ~~~~~~~~~~~~~^^
10
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 76, in submitit_main
11
+ process_job(args.folder)
12
+ ~~~~~~~~~~~^^^^^^^^^^^^^
13
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 69, in process_job
14
+ raise error
15
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 55, in process_job
16
+ result = delayed.result()
17
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/utils.py", line 137, in result
18
+ self._result = self.function(*self.args, **self.kwargs)
19
+ ~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^
20
+ File "/gpfs/home2/dcanez/vd/app/main_distributed.py", line 102, in __call__
21
+ app_main(app, args=params, resume_preempt=resume_preempt)
22
+ ~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
23
+ File "/gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/scaffold.py", line 17, in main
24
+ return importlib.import_module(f"app.{app}.train").main(args=args, resume_preempt=resume_preempt)
25
+ ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
26
+ File "/gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/vjepa/train.py", line 1189, in main
27
+ assert not np.isnan(loss), "loss is nan"
28
+ ^^^^^^^^^^^^^^^^^^
29
+ AssertionError: loss is nan
vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/crash_rank4.log ADDED
@@ -0,0 +1,29 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ Rank 4 crashed at epoch 15, itr 250
2
+ Error: loss is nan
3
+
4
+ Traceback (most recent call last):
5
+ File "<frozen runpy>", line 198, in _run_module_as_main
6
+ File "<frozen runpy>", line 88, in _run_code
7
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/_submit.py", line 11, in <module>
8
+ submitit_main()
9
+ ~~~~~~~~~~~~~^^
10
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 76, in submitit_main
11
+ process_job(args.folder)
12
+ ~~~~~~~~~~~^^^^^^^^^^^^^
13
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 69, in process_job
14
+ raise error
15
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 55, in process_job
16
+ result = delayed.result()
17
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/utils.py", line 137, in result
18
+ self._result = self.function(*self.args, **self.kwargs)
19
+ ~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^
20
+ File "/gpfs/home2/dcanez/vd/app/main_distributed.py", line 102, in __call__
21
+ app_main(app, args=params, resume_preempt=resume_preempt)
22
+ ~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
23
+ File "/gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/scaffold.py", line 17, in main
24
+ return importlib.import_module(f"app.{app}.train").main(args=args, resume_preempt=resume_preempt)
25
+ ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
26
+ File "/gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/vjepa/train.py", line 1189, in main
27
+ assert not np.isnan(loss), "loss is nan"
28
+ ^^^^^^^^^^^^^^^^^^
29
+ AssertionError: loss is nan
vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/crash_rank5.log ADDED
@@ -0,0 +1,29 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ Rank 5 crashed at epoch 15, itr 250
2
+ Error: loss is nan
3
+
4
+ Traceback (most recent call last):
5
+ File "<frozen runpy>", line 198, in _run_module_as_main
6
+ File "<frozen runpy>", line 88, in _run_code
7
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/_submit.py", line 11, in <module>
8
+ submitit_main()
9
+ ~~~~~~~~~~~~~^^
10
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 76, in submitit_main
11
+ process_job(args.folder)
12
+ ~~~~~~~~~~~^^^^^^^^^^^^^
13
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 69, in process_job
14
+ raise error
15
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 55, in process_job
16
+ result = delayed.result()
17
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/utils.py", line 137, in result
18
+ self._result = self.function(*self.args, **self.kwargs)
19
+ ~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^
20
+ File "/gpfs/home2/dcanez/vd/app/main_distributed.py", line 102, in __call__
21
+ app_main(app, args=params, resume_preempt=resume_preempt)
22
+ ~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
23
+ File "/gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/scaffold.py", line 17, in main
24
+ return importlib.import_module(f"app.{app}.train").main(args=args, resume_preempt=resume_preempt)
25
+ ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
26
+ File "/gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/vjepa/train.py", line 1189, in main
27
+ assert not np.isnan(loss), "loss is nan"
28
+ ^^^^^^^^^^^^^^^^^^
29
+ AssertionError: loss is nan
vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/crash_rank6.log ADDED
@@ -0,0 +1,29 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ Rank 6 crashed at epoch 15, itr 250
2
+ Error: loss is nan
3
+
4
+ Traceback (most recent call last):
5
+ File "<frozen runpy>", line 198, in _run_module_as_main
6
+ File "<frozen runpy>", line 88, in _run_code
7
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/_submit.py", line 11, in <module>
8
+ submitit_main()
9
+ ~~~~~~~~~~~~~^^
10
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 76, in submitit_main
11
+ process_job(args.folder)
12
+ ~~~~~~~~~~~^^^^^^^^^^^^^
13
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 69, in process_job
14
+ raise error
15
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 55, in process_job
16
+ result = delayed.result()
17
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/utils.py", line 137, in result
18
+ self._result = self.function(*self.args, **self.kwargs)
19
+ ~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^
20
+ File "/gpfs/home2/dcanez/vd/app/main_distributed.py", line 102, in __call__
21
+ app_main(app, args=params, resume_preempt=resume_preempt)
22
+ ~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
23
+ File "/gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/scaffold.py", line 17, in main
24
+ return importlib.import_module(f"app.{app}.train").main(args=args, resume_preempt=resume_preempt)
25
+ ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
26
+ File "/gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/vjepa/train.py", line 1189, in main
27
+ assert not np.isnan(loss), "loss is nan"
28
+ ^^^^^^^^^^^^^^^^^^
29
+ AssertionError: loss is nan
vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/crash_rank7.log ADDED
@@ -0,0 +1,29 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ Rank 7 crashed at epoch 15, itr 250
2
+ Error: loss is nan
3
+
4
+ Traceback (most recent call last):
5
+ File "<frozen runpy>", line 198, in _run_module_as_main
6
+ File "<frozen runpy>", line 88, in _run_code
7
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/_submit.py", line 11, in <module>
8
+ submitit_main()
9
+ ~~~~~~~~~~~~~^^
10
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 76, in submitit_main
11
+ process_job(args.folder)
12
+ ~~~~~~~~~~~^^^^^^^^^^^^^
13
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 69, in process_job
14
+ raise error
15
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 55, in process_job
16
+ result = delayed.result()
17
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/utils.py", line 137, in result
18
+ self._result = self.function(*self.args, **self.kwargs)
19
+ ~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^
20
+ File "/gpfs/home2/dcanez/vd/app/main_distributed.py", line 102, in __call__
21
+ app_main(app, args=params, resume_preempt=resume_preempt)
22
+ ~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
23
+ File "/gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/scaffold.py", line 17, in main
24
+ return importlib.import_module(f"app.{app}.train").main(args=args, resume_preempt=resume_preempt)
25
+ ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
26
+ File "/gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/vjepa/train.py", line 1189, in main
27
+ assert not np.isnan(loss), "loss is nan"
28
+ ^^^^^^^^^^^^^^^^^^
29
+ AssertionError: loss is nan
vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/crash_rank8.log ADDED
@@ -0,0 +1,29 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ Rank 8 crashed at epoch 15, itr 250
2
+ Error: loss is nan
3
+
4
+ Traceback (most recent call last):
5
+ File "<frozen runpy>", line 198, in _run_module_as_main
6
+ File "<frozen runpy>", line 88, in _run_code
7
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/_submit.py", line 11, in <module>
8
+ submitit_main()
9
+ ~~~~~~~~~~~~~^^
10
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 76, in submitit_main
11
+ process_job(args.folder)
12
+ ~~~~~~~~~~~^^^^^^^^^^^^^
13
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 69, in process_job
14
+ raise error
15
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 55, in process_job
16
+ result = delayed.result()
17
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/utils.py", line 137, in result
18
+ self._result = self.function(*self.args, **self.kwargs)
19
+ ~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^
20
+ File "/gpfs/home2/dcanez/vd/app/main_distributed.py", line 102, in __call__
21
+ app_main(app, args=params, resume_preempt=resume_preempt)
22
+ ~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
23
+ File "/gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/scaffold.py", line 17, in main
24
+ return importlib.import_module(f"app.{app}.train").main(args=args, resume_preempt=resume_preempt)
25
+ ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
26
+ File "/gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/vjepa/train.py", line 1189, in main
27
+ assert not np.isnan(loss), "loss is nan"
28
+ ^^^^^^^^^^^^^^^^^^
29
+ AssertionError: loss is nan
vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/crash_rank9.log ADDED
@@ -0,0 +1,29 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ Rank 9 crashed at epoch 15, itr 250
2
+ Error: loss is nan
3
+
4
+ Traceback (most recent call last):
5
+ File "<frozen runpy>", line 198, in _run_module_as_main
6
+ File "<frozen runpy>", line 88, in _run_code
7
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/_submit.py", line 11, in <module>
8
+ submitit_main()
9
+ ~~~~~~~~~~~~~^^
10
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 76, in submitit_main
11
+ process_job(args.folder)
12
+ ~~~~~~~~~~~^^^^^^^^^^^^^
13
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 69, in process_job
14
+ raise error
15
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 55, in process_job
16
+ result = delayed.result()
17
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/utils.py", line 137, in result
18
+ self._result = self.function(*self.args, **self.kwargs)
19
+ ~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^
20
+ File "/gpfs/home2/dcanez/vd/app/main_distributed.py", line 102, in __call__
21
+ app_main(app, args=params, resume_preempt=resume_preempt)
22
+ ~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
23
+ File "/gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/scaffold.py", line 17, in main
24
+ return importlib.import_module(f"app.{app}.train").main(args=args, resume_preempt=resume_preempt)
25
+ ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
26
+ File "/gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/vjepa/train.py", line 1189, in main
27
+ assert not np.isnan(loss), "loss is nan"
28
+ ^^^^^^^^^^^^^^^^^^
29
+ AssertionError: loss is nan
vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/git-info.txt ADDED
@@ -0,0 +1,2 @@
 
 
 
1
+ branch: main
2
+ commit: 4f75caae343937a59636f9600dd815d5d6b06a6f
vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/job_22614766/22614766_0_log.err ADDED
@@ -0,0 +1,77 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/timm/models/layers/__init__.py:49: FutureWarning: Importing from timm.models.layers is deprecated, please import via timm.layers
2
+ warnings.warn(f"Importing from {__name__} is deprecated, please import via timm.layers", FutureWarning)
3
+ /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/distributed/c10d_logger.py:83: UserWarning: barrier(): using the device under current context. You can specify `device_id` in `init_process_group` to mute this warning.
4
+ return func(*args, **kwargs)
5
+ [rank0]:[W510 12:20:58.732173736 ProcessGroupNCCL.cpp:5138] Guessing device ID based on global rank. This can cause a hang if rank to GPU mapping is heterogeneous. You can specify device_id in init_process_group()
6
+ wandb: Currently logged in as: dgcnz (uvjepa) to https://api.wandb.ai. Use `wandb login --relogin` to force relogin
7
+ wandb: setting up run 4zt31gid
8
+ wandb: Tracking run with wandb version 0.23.1
9
+ wandb: Run data is saved locally in /scratch-shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/wandb/run-20260510_122107-4zt31gid
10
+ wandb: Run `wandb offline` to turn off syncing.
11
+ wandb: Syncing run c002_vitl_k16_simple_cross_stable
12
+ wandb: ⭐️ View project at https://wandb.ai/uvjepa/vjepa_ujepaside
13
+ wandb: 🚀 View run at https://wandb.ai/uvjepa/vjepa_ujepaside/runs/4zt31gid
14
+ /gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/vjepa/utils.py:796: FutureWarning: `torch.cuda.amp.GradScaler(args...)` is deprecated. Please use `torch.amp.GradScaler('cuda', args...)` instead.
15
+ scaler = torch.cuda.amp.GradScaler() if mixed_precision else None
16
+ /gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/vjepa/train.py:955: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
17
+ with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
18
+ /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/distributed/c10d_logger.py:83: UserWarning: barrier(): using the device under current context. You can specify `device_id` in `init_process_group` to mute this warning.
19
+ return func(*args, **kwargs)
20
+ /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/distributed/c10d_logger.py:83: UserWarning: barrier(): using the device under current context. You can specify `device_id` in `init_process_group` to mute this warning.
21
+ return func(*args, **kwargs)
22
+ /gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/vjepa/train.py:955: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
23
+ with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
24
+ /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/distributed/c10d_logger.py:83: UserWarning: barrier(): using the device under current context. You can specify `device_id` in `init_process_group` to mute this warning.
25
+ return func(*args, **kwargs)
26
+ /gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/vjepa/train.py:955: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
27
+ with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
28
+ /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/distributed/c10d_logger.py:83: UserWarning: barrier(): using the device under current context. You can specify `device_id` in `init_process_group` to mute this warning.
29
+ return func(*args, **kwargs)
30
+ /gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/vjepa/train.py:955: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
31
+ with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
32
+ submitit ERROR (2026-05-10 16:19:44,648) - Submitted job triggered an exception
33
+ Traceback (most recent call last):
34
+ File "<frozen runpy>", line 198, in _run_module_as_main
35
+ File "<frozen runpy>", line 88, in _run_code
36
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/_submit.py", line 11, in <module>
37
+ submitit_main()
38
+ ~~~~~~~~~~~~~^^
39
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 76, in submitit_main
40
+ process_job(args.folder)
41
+ ~~~~~~~~~~~^^^^^^^^^^^^^
42
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 69, in process_job
43
+ raise error
44
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 55, in process_job
45
+ result = delayed.result()
46
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/utils.py", line 137, in result
47
+ self._result = self.function(*self.args, **self.kwargs)
48
+ ~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^
49
+ File "/gpfs/home2/dcanez/vd/app/main_distributed.py", line 102, in __call__
50
+ app_main(app, args=params, resume_preempt=resume_preempt)
51
+ ~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
52
+ File "/gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/scaffold.py", line 17, in main
53
+ return importlib.import_module(f"app.{app}.train").main(args=args, resume_preempt=resume_preempt)
54
+ ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
55
+ File "/gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/vjepa/train.py", line 1189, in main
56
+ assert not np.isnan(loss), "loss is nan"
57
+ ^^^^^^^^^^^^^^^^^^
58
+ AssertionError: loss is nan
59
+ [rank0]:[W510 16:19:47.642195666 ProcessGroupNCCL.cpp:1553] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())
60
+ srun: error: gcn87: task 15: Exited with exit code 1
61
+ srun: Terminating StepId=22614766.0
62
+ [2026-05-10T16:19:56.527] error: *** STEP 22614766.0 ON gcn80 CANCELLED AT 2026-05-10T16:19:56 DUE TO TASK FAILURE ***
63
+ srun: error: gcn80: task 1: Exited with exit code 1
64
+ srun: error: gcn85: task 10: Exited with exit code 1
65
+ srun: error: gcn82: task 6: Exited with exit code 1
66
+ srun: error: gcn87: task 14: Exited with exit code 1
67
+ srun: error: gcn85: task 8: Exited with exit code 1
68
+ srun: error: gcn80: task 3: Exited with exit code 1
69
+ srun: error: gcn82: task 7: Exited with exit code 1
70
+ srun: error: gcn80: task 2: Exited with exit code 1
71
+ srun: error: gcn80: task 0: Exited with exit code 1
72
+ srun: error: gcn82: task 5: Exited with exit code 1
73
+ srun: error: gcn85: task 11: Exited with exit code 1
74
+ srun: error: gcn87: task 12: Exited with exit code 1
75
+ srun: error: gcn82: task 4: Exited with exit code 1
76
+ srun: error: gcn87: task 13: Exited with exit code 1
77
+ srun: error: gcn85: task 9: Exited with exit code 1
vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/job_22614766/22614766_0_log.out ADDED
The diff for this file is too large to render. See raw diff
 
vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/job_22614766/22614766_10_log.err ADDED
@@ -0,0 +1,68 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/timm/models/layers/__init__.py:49: FutureWarning: Importing from timm.models.layers is deprecated, please import via timm.layers
2
+ warnings.warn(f"Importing from {__name__} is deprecated, please import via timm.layers", FutureWarning)
3
+ [rank10]:[W510 12:20:01.953233658 ProcessGroupNCCL.cpp:5138] Guessing device ID based on global rank. This can cause a hang if rank to GPU mapping is heterogeneous. You can specify device_id in init_process_group()
4
+ /gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/vjepa/utils.py:796: FutureWarning: `torch.cuda.amp.GradScaler(args...)` is deprecated. Please use `torch.amp.GradScaler('cuda', args...)` instead.
5
+ scaler = torch.cuda.amp.GradScaler() if mixed_precision else None
6
+ /gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/vjepa/train.py:955: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
7
+ with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
8
+ [rank10]:[W510 16:19:49.713092222 TCPStore.cpp:125] [c10d] recvValue failed on SocketImpl(fd=3, addr=[gcn85.local.snellius.surf.nl]:48690, remote=[gcn80.local.snellius.surf.nl]:37129): Failed to recv, got 0 bytes. Connection was likely closed. Did the remote server shutdown or crash?
9
+ Exception raised from recvBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:682 (most recent call first):
10
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x150c7f339fdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
11
+ frame #1: <unknown function> + 0x6a3325d (0x150cc34c825d in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
12
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x273 (0x150cc34c61f3 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
13
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x150c805a8eec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
14
+ frame #4: <unknown function> + 0xed164 (0x150dbfa12164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
15
+ frame #5: <unknown function> + 0x8a19a (0x150dd048419a in /lib64/libc.so.6)
16
+ frame #6: <unknown function> + 0x10f100 (0x150dd0509100 in /lib64/libc.so.6)
17
+
18
+ [rank10]:[W510 16:19:49.756854865 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 10] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Failed to recv, got 0 bytes. Connection was likely closed. Did the remote server shutdown or crash?
19
+ submitit ERROR (2026-05-10 16:19:50,186) - Submitted job triggered an exception
20
+ [rank10]:[W510 16:19:50.757056654 TCPStore.cpp:106] [c10d] sendBytes failed on SocketImpl(fd=3, addr=[gcn85.local.snellius.surf.nl]:48690, remote=[gcn80.local.snellius.surf.nl]:37129): Broken pipe
21
+ Exception raised from sendBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:653 (most recent call first):
22
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x150c7f339fdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
23
+ frame #1: <unknown function> + 0x6a326d1 (0x150cc34c76d1 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
24
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x24d (0x150cc34c61cd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
25
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x150c805a8eec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
26
+ frame #4: <unknown function> + 0xed164 (0x150dbfa12164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
27
+ frame #5: <unknown function> + 0x8a19a (0x150dd048419a in /lib64/libc.so.6)
28
+ frame #6: <unknown function> + 0x10f100 (0x150dd0509100 in /lib64/libc.so.6)
29
+
30
+ [rank10]:[W510 16:19:50.760593632 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 10] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Broken pipe
31
+ Traceback (most recent call last):
32
+ File "<frozen runpy>", line 198, in _run_module_as_main
33
+ File "<frozen runpy>", line 88, in _run_code
34
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/_submit.py", line 11, in <module>
35
+ submitit_main()
36
+ ~~~~~~~~~~~~~^^
37
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 76, in submitit_main
38
+ process_job(args.folder)
39
+ ~~~~~~~~~~~^^^^^^^^^^^^^
40
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 69, in process_job
41
+ raise error
42
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 55, in process_job
43
+ result = delayed.result()
44
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/utils.py", line 137, in result
45
+ self._result = self.function(*self.args, **self.kwargs)
46
+ ~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^
47
+ File "/gpfs/home2/dcanez/vd/app/main_distributed.py", line 102, in __call__
48
+ app_main(app, args=params, resume_preempt=resume_preempt)
49
+ ~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
50
+ File "/gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/scaffold.py", line 17, in main
51
+ return importlib.import_module(f"app.{app}.train").main(args=args, resume_preempt=resume_preempt)
52
+ ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
53
+ File "/gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/vjepa/train.py", line 1189, in main
54
+ assert not np.isnan(loss), "loss is nan"
55
+ ^^^^^^^^^^^^^^^^^^
56
+ AssertionError: loss is nan
57
+ [rank10]:[W510 16:19:51.760775392 TCPStore.cpp:106] [c10d] sendBytes failed on SocketImpl(fd=3, addr=[gcn85.local.snellius.surf.nl]:48690, remote=[gcn80.local.snellius.surf.nl]:37129): Broken pipe
58
+ Exception raised from sendBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:653 (most recent call first):
59
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x150c7f339fdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
60
+ frame #1: <unknown function> + 0x6a326d1 (0x150cc34c76d1 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
61
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x24d (0x150cc34c61cd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
62
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x150c805a8eec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
63
+ frame #4: <unknown function> + 0xed164 (0x150dbfa12164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
64
+ frame #5: <unknown function> + 0x8a19a (0x150dd048419a in /lib64/libc.so.6)
65
+ frame #6: <unknown function> + 0x10f100 (0x150dd0509100 in /lib64/libc.so.6)
66
+
67
+ [rank10]:[W510 16:19:51.763251976 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 10] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Broken pipe
68
+ [rank10]:[W510 16:19:51.202554786 ProcessGroupNCCL.cpp:1553] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())
vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/job_22614766/22614766_10_log.out ADDED
@@ -0,0 +1,276 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ submitit INFO (2026-05-10 12:14:34,049) - Starting with JobEnvironment(job_id=22614766, hostname=gcn85.local.snellius.surf.nl, local_rank=2(4), node=2(4), global_rank=10(16))
2
+ submitit INFO (2026-05-10 12:14:34,049) - Loading pickle: /scratch-shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/job_22614766/22614766_submitted.pkl
3
+ INFO:root:loaded pretrain params...
4
+ { 'app': 'vjepa',
5
+ 'cpus_per_task': 16,
6
+ 'data': { 'batch_size': 64,
7
+ 'crop_size': 224,
8
+ 'dataset_fpcs': [16, 16],
9
+ 'dataset_type': 'VideoDataset',
10
+ 'datasets': [ '/scratch-shared/dcanez/data/kinetics/k400/train.csv',
11
+ '/scratch-shared/dcanez/data/ssv2/train.csv'],
12
+ 'datasets_weights': [0.65, 0.35],
13
+ 'fps': 4,
14
+ 'num_workers': 10,
15
+ 'patch_size': 14,
16
+ 'persistent_workers': True,
17
+ 'pin_mem': True,
18
+ 'stage': [ { 'dest': 'kinetics_240',
19
+ 'format': 'targz_parts',
20
+ 'src': '/scratch-shared/dcanez/data/kinetics/k400/tars_240/'},
21
+ { 'dest': 'ssv2',
22
+ 'format': 'multipart_tar',
23
+ 'src': '/scratch-nvme/ml-datasets/something-something-v2/'}],
24
+ 'tubelet_size': 1},
25
+ 'data_aug': { 'auto_augment': False,
26
+ 'motion_shift': False,
27
+ 'random_resize_aspect_ratio': [0.75, 1.35],
28
+ 'random_resize_scale': [0.3, 1.0],
29
+ 'reprob': 0.0},
30
+ 'folder': '/scratch-shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable',
31
+ 'loss': {'loss_exp': 1.0},
32
+ 'mask': [ { 'aspect_ratio': [0.75, 1.5],
33
+ 'full_complement': False,
34
+ 'max_keep': None,
35
+ 'max_temporal_keep': 1.0,
36
+ 'num_blocks': 8,
37
+ 'spatial_scale': [0.15, 0.15],
38
+ 'temporal_scale': [1.0, 1.0]},
39
+ { 'aspect_ratio': [0.75, 1.5],
40
+ 'full_complement': False,
41
+ 'max_keep': None,
42
+ 'max_temporal_keep': 1.0,
43
+ 'num_blocks': 2,
44
+ 'spatial_scale': [0.7, 0.7],
45
+ 'temporal_scale': [1.0, 1.0]}],
46
+ 'mem_per_gpu': '180G',
47
+ 'meta': { 'dtype': 'bfloat16',
48
+ 'knn_eval_epoch0': False,
49
+ 'knn_eval_freq': 5,
50
+ 'knn_eval_presets': [ { 'config': { 'batch_size': 64,
51
+ 'dataset_train': '/scratch-shared/mdorkenw/ucf101/ucfTrainTestlist/train.csv',
52
+ 'dataset_val': '/scratch-shared/mdorkenw/ucf101/ucfTrainTestlist/val.csv',
53
+ 'eval_videos_per_class': 25,
54
+ 'num_workers': 8,
55
+ 'pool_type': 'slot_temporal_concat',
56
+ 'train_videos_per_class': 100},
57
+ 'preset': 'ucf101'},
58
+ { 'config': { 'batch_size': 64,
59
+ 'dataset_train': '/scratch-shared/mdorkenw/20bn-something-something-v2/something-something-v2-annotations/train_coarse10.csv',
60
+ 'dataset_val': '/scratch-shared/mdorkenw/20bn-something-something-v2/something-something-v2-annotations/val_coarse10.csv',
61
+ 'eval_videos_per_class': 100,
62
+ 'linear_probe': True,
63
+ 'num_workers': 8,
64
+ 'pool_type': 'slot_temporal_concat',
65
+ 'train_videos_per_class': 500},
66
+ 'preset': 'ssv2_coarse10'}],
67
+ 'load_checkpoint': True,
68
+ 'read_checkpoint': None,
69
+ 'save_every_freq': 5,
70
+ 'seed': 239,
71
+ 'use_sdpa': True,
72
+ 'use_wandb': True,
73
+ 'wandb_project': 'vjepa_ujepaside'},
74
+ 'metrics': {'sigreg': {}, 'std': {}},
75
+ 'model': { 'model_name': 'ujepaside_large_patch14_capi_lvd1689m',
76
+ 'pred_depth': 6,
77
+ 'pred_embed_dim': 384,
78
+ 'pred_num_heads': 12,
79
+ 'predictor': 'v2_cross',
80
+ 'st_causal': False,
81
+ 'st_drop_path': 0.2,
82
+ 'st_flex_enable': False,
83
+ 'st_layer_scale_init': 1e-05,
84
+ 'st_num_slots': 16,
85
+ 'st_slots_causal_within_frame': False,
86
+ 'target_kind': 'frozen_2d',
87
+ 'target_type': 'vit_large_patch14_capi.lvd1689m',
88
+ 'temporal_spacing': 1.0,
89
+ 'uniform_power': True,
90
+ 'use_activation_checkpointing': True,
91
+ 'use_mask_tokens': True,
92
+ 'use_rope': True,
93
+ 'use_sdpa': True,
94
+ 'zero_init_mask_tokens': True},
95
+ 'nodes': 4,
96
+ 'optimization': { 'clip_grad': 3.0,
97
+ 'ema': [0.99925, 0.99925],
98
+ 'epochs': 100,
99
+ 'final_lr': 0.0001,
100
+ 'final_weight_decay': 0.04,
101
+ 'ipe': 300,
102
+ 'ipe_scale': 1.0,
103
+ 'lr': 0.0005,
104
+ 'start_lr': 0.0001,
105
+ 'warmup': 10,
106
+ 'weight_decay': 0.04},
107
+ 'tasks_per_node': 4}
108
+ INFO:root:Running pre-training of app: vjepa
109
+ [INFO ][2026-05-10 12:16:54][app.vjepa.train ][main ] which_dtype='bfloat16'
110
+ [INFO ][2026-05-10 12:16:54][app.vjepa.train ][main ] Disabling persistent_workers (incompatible with KNN eval)
111
+ [INFO ][2026-05-10 12:16:54][app.vjepa.train ][main ] NCCL_SOCKET_IFNAME=eno
112
+ [INFO ][2026-05-10 12:16:57][app.vjepa.train ][main ] Initialized (rank/world-size) 10/16, tasks_per_node=4
113
+ [INFO ][2026-05-10 12:16:57][root ][stage_datasets ] [local_rank 2/4] Staging kinetics_240 (targz_parts)
114
+ [INFO ][2026-05-10 12:16:57][root ][_stage_targz_parts ] [rank 2] Extracting 26/103 tar.gz parts to /scratch-node/dcanez.22614766/kinetics_240
115
+ [INFO ][2026-05-10 12:17:11][root ][_stage_targz_parts ] [local_rank 2] Extracted 2/26 parts
116
+ [INFO ][2026-05-10 12:17:25][root ][_stage_targz_parts ] [local_rank 2] Extracted 4/26 parts
117
+ [INFO ][2026-05-10 12:17:39][root ][_stage_targz_parts ] [local_rank 2] Extracted 6/26 parts
118
+ [INFO ][2026-05-10 12:17:53][root ][_stage_targz_parts ] [local_rank 2] Extracted 8/26 parts
119
+ [INFO ][2026-05-10 12:18:07][root ][_stage_targz_parts ] [local_rank 2] Extracted 10/26 parts
120
+ [INFO ][2026-05-10 12:18:21][root ][_stage_targz_parts ] [local_rank 2] Extracted 12/26 parts
121
+ [INFO ][2026-05-10 12:18:35][root ][_stage_targz_parts ] [local_rank 2] Extracted 14/26 parts
122
+ [INFO ][2026-05-10 12:18:49][root ][_stage_targz_parts ] [local_rank 2] Extracted 16/26 parts
123
+ [INFO ][2026-05-10 12:19:04][root ][_stage_targz_parts ] [local_rank 2] Extracted 18/26 parts
124
+ [INFO ][2026-05-10 12:19:18][root ][_stage_targz_parts ] [local_rank 2] Extracted 20/26 parts
125
+ [INFO ][2026-05-10 12:19:32][root ][_stage_targz_parts ] [local_rank 2] Extracted 22/26 parts
126
+ [INFO ][2026-05-10 12:19:46][root ][_stage_targz_parts ] [local_rank 2] Extracted 24/26 parts
127
+ [INFO ][2026-05-10 12:20:01][root ][_stage_targz_parts ] [local_rank 2] Extracted 26/26 parts
128
+ [INFO ][2026-05-10 12:20:01][root ][stage_datasets ] [local_rank 2/4] Staging ssv2 (multipart_tar)
129
+ [INFO ][2026-05-10 12:21:03][app.vjepa.train ][main ] Staged datasets: ['/scratch-node/dcanez.22614766/kinetics_240/train.csv', '/scratch-node/dcanez.22614766/ssv2/train.csv']
130
+ [INFO ][2026-05-10 12:23:31][experiments.stmodels.vision_transformers_v3][vit_large_patch14_capi_lvd1689m] Loading pretrained weights for vit_large_patch14_capi_lvd1689m from https://dl.fbaipublicfiles.com/capi/capi_vitl14_lvd.pth
131
+ [INFO ][2026-05-10 12:23:51][root ][_build_frozen_2d_target ] Loaded frozen_2d target from timm: vit_large_patch14_capi.lvd1689m
132
+ [INFO ][2026-05-10 12:23:52][root ][init_video_model ] ViTMultiSeqWrapper(
133
+ (backbone): UJEPAside(
134
+ (patch_embed): PatchEmbed3D(
135
+ (proj): Conv3d(3, 1024, kernel_size=(1, 14, 14), stride=(1, 14, 14))
136
+ )
137
+ (rope): CAPI2DRoPE()
138
+ (blocks): ModuleList(
139
+ (0-23): 24 x Block(
140
+ (norm1): RMSNorm((1024,), eps=1e-05, elementwise_affine=True)
141
+ (rope_impl): CAPI2DRoPE()
142
+ (norm2): RMSNorm((1024,), eps=1e-05, elementwise_affine=True)
143
+ (drop_path1): Identity()
144
+ (drop_path2): Identity()
145
+ (attn): Attention(
146
+ (qkv): Linear(in_features=1024, out_features=3072, bias=False)
147
+ (attn_drop): Dropout(p=0.0, inplace=False)
148
+ (proj): Linear(in_features=1024, out_features=1024, bias=False)
149
+ (proj_drop): Dropout(p=0.0, inplace=False)
150
+ (rope_impl): CAPI2DRoPE()
151
+ )
152
+ (mlp): MLP(
153
+ (fc1): Linear(in_features=1024, out_features=4096, bias=False)
154
+ (act): GELU(approximate='none')
155
+ (fc2): Linear(in_features=4096, out_features=1024, bias=False)
156
+ (drop): Dropout(p=0.0, inplace=False)
157
+ )
158
+ )
159
+ )
160
+ (norm): RMSNorm((1024,), eps=1e-05, elementwise_affine=False)
161
+ (st_blocks): ModuleList(
162
+ (0-23): 24 x Block(
163
+ (residual1): EfficientResidual(
164
+ (norm): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
165
+ (fn): Attention(
166
+ (q_proj): Linear(in_features=1024, out_features=1024, bias=False)
167
+ (k_proj): Linear(in_features=1024, out_features=1024, bias=False)
168
+ (v_proj): Linear(in_features=1024, out_features=1024, bias=False)
169
+ (proj): Linear(in_features=1024, out_features=1024, bias=False)
170
+ (rope_impl): CAPI3DRoPE()
171
+ )
172
+ )
173
+ (residual2): EfficientResidual(
174
+ (norm): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
175
+ (fn): MLP(
176
+ (fc1): Linear(in_features=1024, out_features=4096, bias=False)
177
+ (act): GELU(approximate='none')
178
+ (fc2): Linear(in_features=4096, out_features=1024, bias=False)
179
+ (drop): Dropout(p=0.0, inplace=False)
180
+ )
181
+ )
182
+ )
183
+ )
184
+ (st_rope): CAPI3DRoPE()
185
+ )
186
+ )
187
+ [INFO ][2026-05-10 12:23:52][root ][init_video_model ] UJEPAsidePredictorMultiSeqWrapper(
188
+ (backbone): PredictorV2(
189
+ (predictor_embed): Linear(in_features=1024, out_features=384, bias=True)
190
+ (mask_tokens): ParameterList(
191
+ (0): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
192
+ (1): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
193
+ (2): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
194
+ (3): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
195
+ )
196
+ (predictor_blocks): ModuleList(
197
+ (0-5): 6 x Block(
198
+ (residual1): EfficientResidual(
199
+ (norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
200
+ (fn): Attention(
201
+ (q_proj): Linear(in_features=384, out_features=384, bias=False)
202
+ (k_proj): Linear(in_features=384, out_features=384, bias=False)
203
+ (v_proj): Linear(in_features=384, out_features=384, bias=False)
204
+ (proj): Linear(in_features=384, out_features=384, bias=False)
205
+ (rope): Rope()
206
+ )
207
+ )
208
+ (residual2): EfficientResidual(
209
+ (norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
210
+ (fn): MLP(
211
+ (fc1): Linear(in_features=384, out_features=1536, bias=False)
212
+ (act): GELU(approximate='none')
213
+ (fc2): Linear(in_features=1536, out_features=384, bias=False)
214
+ (drop): Dropout(p=0.0, inplace=False)
215
+ )
216
+ )
217
+ )
218
+ )
219
+ (predictor_norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
220
+ (predictor_proj): Linear(in_features=384, out_features=1024, bias=True)
221
+ )
222
+ )
223
+ [INFO ][2026-05-10 12:23:52][root ][init_video_model ] MultiSeqWrapper(
224
+ (backbone): Frozen2DTargetWrapper(
225
+ (backbone): Eva(
226
+ (patch_embed): PatchEmbed(
227
+ (proj): Conv2d(3, 1024, kernel_size=(14, 14), stride=(14, 14))
228
+ (norm): Identity()
229
+ )
230
+ (pos_drop): Dropout(p=0.0, inplace=False)
231
+ (norm_pre): Identity()
232
+ (blocks): ModuleList(
233
+ (0-23): 24 x EvaBlock(
234
+ (norm1): RMSNorm((1024,), eps=1e-05, elementwise_affine=True)
235
+ (attn): EvaAttention(
236
+ (qkv): Linear(in_features=1024, out_features=3072, bias=False)
237
+ (q_norm): Identity()
238
+ (k_norm): Identity()
239
+ (attn_drop): Dropout(p=0.0, inplace=False)
240
+ (norm): Identity()
241
+ (proj): Linear(in_features=1024, out_features=1024, bias=True)
242
+ (proj_drop): Dropout(p=0.0, inplace=False)
243
+ )
244
+ (drop_path1): Identity()
245
+ (norm2): RMSNorm((1024,), eps=1e-05, elementwise_affine=True)
246
+ (mlp): Mlp(
247
+ (fc1): Linear(in_features=1024, out_features=4096, bias=True)
248
+ (act): GELU(approximate='none')
249
+ (drop1): Dropout(p=0.0, inplace=False)
250
+ (norm): Identity()
251
+ (fc2): Linear(in_features=4096, out_features=1024, bias=True)
252
+ (drop2): Dropout(p=0.0, inplace=False)
253
+ )
254
+ (drop_path2): Identity()
255
+ )
256
+ )
257
+ (norm): RMSNorm((1024,), eps=1e-05, elementwise_affine=True)
258
+ (fc_norm): Identity()
259
+ (head_drop): Dropout(p=0.0, inplace=False)
260
+ (head): Identity()
261
+ (rope): _CapiPatchRoPE()
262
+ )
263
+ )
264
+ )
265
+ [INFO ][2026-05-10 12:23:52][root ][init_video_model ] Encoder number of parameters: 302399488
266
+ [INFO ][2026-05-10 12:23:52][root ][init_video_model ] Predictor number of parameters: 11416192
267
+ [INFO ][2026-05-10 12:23:52][root ][init_video_model ] Target encoder number of parameters: 0
268
+ [INFO ][2026-05-10 12:23:52][root ][make_videodataset ] VideoDataset dataset created
269
+ [INFO ][2026-05-10 12:23:52][WeightedSampler ][__init__ ] Using DistributedWeightedSampler with rank 10 / 16
270
+ [INFO ][2026-05-10 12:23:52][root ][make_videodataset ] VideoDataset unsupervised data loader created
271
+ [INFO ][2026-05-10 12:23:52][app.vjepa.train ][main ] iterations per epoch/dataset length: 300/399
272
+ [INFO ][2026-05-10 12:23:52][app.vjepa.train ][main ] Wrapping models in DDP (rank 10)...
273
+ [INFO ][2026-05-10 12:24:08][app.vjepa.train ][main ] Initializing loader...
274
+ submitit ERROR (2026-05-10 16:19:50,186) - Submitted job triggered an exception
275
+ [ERROR ][2026-05-10 16:19:50][submitit ][process_job ] Submitted job triggered an exception
276
+ [ERROR ][2026-05-10 16:19:51][app.vjepa.train ][_save_crash ] Saved crash log to /scratch-shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/crash_rank10.log
vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/job_22614766/22614766_11_log.err ADDED
@@ -0,0 +1,158 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/timm/models/layers/__init__.py:49: FutureWarning: Importing from timm.models.layers is deprecated, please import via timm.layers
2
+ warnings.warn(f"Importing from {__name__} is deprecated, please import via timm.layers", FutureWarning)
3
+ [rank11]:[W510 12:19:56.602814346 ProcessGroupNCCL.cpp:5138] Guessing device ID based on global rank. This can cause a hang if rank to GPU mapping is heterogeneous. You can specify device_id in init_process_group()
4
+ /gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/vjepa/utils.py:796: FutureWarning: `torch.cuda.amp.GradScaler(args...)` is deprecated. Please use `torch.amp.GradScaler('cuda', args...)` instead.
5
+ scaler = torch.cuda.amp.GradScaler() if mixed_precision else None
6
+ /gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/vjepa/train.py:955: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
7
+ with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
8
+ [rank11]:[W510 16:19:49.713098182 TCPStore.cpp:125] [c10d] recvValue failed on SocketImpl(fd=3, addr=[gcn85.local.snellius.surf.nl]:48702, remote=[gcn80.local.snellius.surf.nl]:37129): Failed to recv, got 0 bytes. Connection was likely closed. Did the remote server shutdown or crash?
9
+ Exception raised from recvBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:682 (most recent call first):
10
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x147d17e52fdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
11
+ frame #1: <unknown function> + 0x6a3325d (0x147d5bfe125d in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
12
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x273 (0x147d5bfdf1f3 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
13
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x147d190c1eec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
14
+ frame #4: <unknown function> + 0xed164 (0x147e5852b164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
15
+ frame #5: <unknown function> + 0x8a19a (0x147e68f9d19a in /lib64/libc.so.6)
16
+ frame #6: <unknown function> + 0x10f100 (0x147e69022100 in /lib64/libc.so.6)
17
+
18
+ [rank11]:[W510 16:19:49.756860915 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 11] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Failed to recv, got 0 bytes. Connection was likely closed. Did the remote server shutdown or crash?
19
+ [rank11]:[W510 16:19:50.757031945 TCPStore.cpp:106] [c10d] sendBytes failed on SocketImpl(fd=3, addr=[gcn85.local.snellius.surf.nl]:48702, remote=[gcn80.local.snellius.surf.nl]:37129): Broken pipe
20
+ Exception raised from sendBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:653 (most recent call first):
21
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x147d17e52fdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
22
+ frame #1: <unknown function> + 0x6a326d1 (0x147d5bfe06d1 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
23
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x24d (0x147d5bfdf1cd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
24
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x147d190c1eec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
25
+ frame #4: <unknown function> + 0xed164 (0x147e5852b164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
26
+ frame #5: <unknown function> + 0x8a19a (0x147e68f9d19a in /lib64/libc.so.6)
27
+ frame #6: <unknown function> + 0x10f100 (0x147e69022100 in /lib64/libc.so.6)
28
+
29
+ [rank11]:[W510 16:19:50.760781021 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 11] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Broken pipe
30
+ [rank11]:[W510 16:19:51.760872672 TCPStore.cpp:106] [c10d] sendBytes failed on SocketImpl(fd=3, addr=[gcn85.local.snellius.surf.nl]:48702, remote=[gcn80.local.snellius.surf.nl]:37129): Broken pipe
31
+ Exception raised from sendBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:653 (most recent call first):
32
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x147d17e52fdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
33
+ frame #1: <unknown function> + 0x6a326d1 (0x147d5bfe06d1 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
34
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x24d (0x147d5bfdf1cd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
35
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x147d190c1eec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
36
+ frame #4: <unknown function> + 0xed164 (0x147e5852b164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
37
+ frame #5: <unknown function> + 0x8a19a (0x147e68f9d19a in /lib64/libc.so.6)
38
+ frame #6: <unknown function> + 0x10f100 (0x147e69022100 in /lib64/libc.so.6)
39
+
40
+ [rank11]:[W510 16:19:51.763308266 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 11] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Broken pipe
41
+ [rank11]:[W510 16:19:52.763400316 TCPStore.cpp:106] [c10d] sendBytes failed on SocketImpl(fd=3, addr=[gcn85.local.snellius.surf.nl]:48702, remote=[gcn80.local.snellius.surf.nl]:37129): Broken pipe
42
+ Exception raised from sendBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:653 (most recent call first):
43
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x147d17e52fdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
44
+ frame #1: <unknown function> + 0x6a326d1 (0x147d5bfe06d1 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
45
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x24d (0x147d5bfdf1cd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
46
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x147d190c1eec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
47
+ frame #4: <unknown function> + 0xed164 (0x147e5852b164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
48
+ frame #5: <unknown function> + 0x8a19a (0x147e68f9d19a in /lib64/libc.so.6)
49
+ frame #6: <unknown function> + 0x10f100 (0x147e69022100 in /lib64/libc.so.6)
50
+
51
+ [rank11]:[W510 16:19:52.765625912 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 11] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Broken pipe
52
+ [rank11]:[W510 16:19:53.765744111 TCPStore.cpp:106] [c10d] sendBytes failed on SocketImpl(fd=3, addr=[gcn85.local.snellius.surf.nl]:48702, remote=[gcn80.local.snellius.surf.nl]:37129): Broken pipe
53
+ Exception raised from sendBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:653 (most recent call first):
54
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x147d17e52fdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
55
+ frame #1: <unknown function> + 0x6a326d1 (0x147d5bfe06d1 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
56
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x24d (0x147d5bfdf1cd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
57
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x147d190c1eec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
58
+ frame #4: <unknown function> + 0xed164 (0x147e5852b164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
59
+ frame #5: <unknown function> + 0x8a19a (0x147e68f9d19a in /lib64/libc.so.6)
60
+ frame #6: <unknown function> + 0x10f100 (0x147e69022100 in /lib64/libc.so.6)
61
+
62
+ [rank11]:[W510 16:19:53.768030807 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 11] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Broken pipe
63
+ [rank11]:[W510 16:19:54.768144006 TCPStore.cpp:106] [c10d] sendBytes failed on SocketImpl(fd=3, addr=[gcn85.local.snellius.surf.nl]:48702, remote=[gcn80.local.snellius.surf.nl]:37129): Broken pipe
64
+ Exception raised from sendBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:653 (most recent call first):
65
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x147d17e52fdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
66
+ frame #1: <unknown function> + 0x6a326d1 (0x147d5bfe06d1 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
67
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x24d (0x147d5bfdf1cd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
68
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x147d190c1eec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
69
+ frame #4: <unknown function> + 0xed164 (0x147e5852b164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
70
+ frame #5: <unknown function> + 0x8a19a (0x147e68f9d19a in /lib64/libc.so.6)
71
+ frame #6: <unknown function> + 0x10f100 (0x147e69022100 in /lib64/libc.so.6)
72
+
73
+ [rank11]:[W510 16:19:54.770513821 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 11] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Broken pipe
74
+ [rank11]:[W510 16:19:55.770599811 TCPStore.cpp:106] [c10d] sendBytes failed on SocketImpl(fd=3, addr=[gcn85.local.snellius.surf.nl]:48702, remote=[gcn80.local.snellius.surf.nl]:37129): Broken pipe
75
+ Exception raised from sendBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:653 (most recent call first):
76
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x147d17e52fdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
77
+ frame #1: <unknown function> + 0x6a326d1 (0x147d5bfe06d1 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
78
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x24d (0x147d5bfdf1cd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
79
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x147d190c1eec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
80
+ frame #4: <unknown function> + 0xed164 (0x147e5852b164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
81
+ frame #5: <unknown function> + 0x8a19a (0x147e68f9d19a in /lib64/libc.so.6)
82
+ frame #6: <unknown function> + 0x10f100 (0x147e69022100 in /lib64/libc.so.6)
83
+
84
+ [rank11]:[W510 16:19:55.772834467 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 11] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Broken pipe
85
+ [rank11]:[W510 16:19:56.772941577 TCPStore.cpp:106] [c10d] sendBytes failed on SocketImpl(fd=3, addr=[gcn85.local.snellius.surf.nl]:48702, remote=[gcn80.local.snellius.surf.nl]:37129): Broken pipe
86
+ Exception raised from sendBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:653 (most recent call first):
87
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x147d17e52fdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
88
+ frame #1: <unknown function> + 0x6a326d1 (0x147d5bfe06d1 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
89
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x24d (0x147d5bfdf1cd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
90
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x147d190c1eec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
91
+ frame #4: <unknown function> + 0xed164 (0x147e5852b164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
92
+ frame #5: <unknown function> + 0x8a19a (0x147e68f9d19a in /lib64/libc.so.6)
93
+ frame #6: <unknown function> + 0x10f100 (0x147e69022100 in /lib64/libc.so.6)
94
+
95
+ [rank11]:[W510 16:19:56.775308622 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 11] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Broken pipe
96
+ [rank11]:[W510 16:19:57.775405062 TCPStore.cpp:106] [c10d] sendBytes failed on SocketImpl(fd=3, addr=[gcn85.local.snellius.surf.nl]:48702, remote=[gcn80.local.snellius.surf.nl]:37129): Broken pipe
97
+ Exception raised from sendBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:653 (most recent call first):
98
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x147d17e52fdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
99
+ frame #1: <unknown function> + 0x6a326d1 (0x147d5bfe06d1 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
100
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x24d (0x147d5bfdf1cd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
101
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x147d190c1eec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
102
+ frame #4: <unknown function> + 0xed164 (0x147e5852b164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
103
+ frame #5: <unknown function> + 0x8a19a (0x147e68f9d19a in /lib64/libc.so.6)
104
+ frame #6: <unknown function> + 0x10f100 (0x147e69022100 in /lib64/libc.so.6)
105
+
106
+ [rank11]:[W510 16:19:57.777864156 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 11] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Broken pipe
107
+ submitit WARNING (2026-05-10 16:19:58,222) - Bypassing signal SIGTERM
108
+ submitit WARNING (2026-05-10 16:19:58,234) - Bypassing signal SIGCONT
109
+ submitit ERROR (2026-05-10 16:19:58,234) - Submitted job triggered an exception
110
+ Traceback (most recent call last):
111
+ File "<frozen runpy>", line 198, in _run_module_as_main
112
+ File "<frozen runpy>", line 88, in _run_code
113
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/_submit.py", line 11, in <module>
114
+ submitit_main()
115
+ ~~~~~~~~~~~~~^^
116
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 76, in submitit_main
117
+ process_job(args.folder)
118
+ ~~~~~~~~~~~^^^^^^^^^^^^^
119
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 69, in process_job
120
+ raise error
121
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 55, in process_job
122
+ result = delayed.result()
123
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/utils.py", line 137, in result
124
+ self._result = self.function(*self.args, **self.kwargs)
125
+ ~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^
126
+ File "/gpfs/home2/dcanez/vd/app/main_distributed.py", line 102, in __call__
127
+ app_main(app, args=params, resume_preempt=resume_preempt)
128
+ ~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
129
+ File "/gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/scaffold.py", line 17, in main
130
+ return importlib.import_module(f"app.{app}.train").main(args=args, resume_preempt=resume_preempt)
131
+ ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
132
+ File "/gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/vjepa/train.py", line 1189, in main
133
+ assert not np.isnan(loss), "loss is nan"
134
+ ^^^^^^^^^^^^^^^^^^
135
+ AssertionError: loss is nan
136
+ [rank11]:[W510 16:19:58.778051455 TCPStore.cpp:106] [c10d] sendBytes failed on SocketImpl(fd=3, addr=[gcn85.local.snellius.surf.nl]:48702, remote=[gcn80.local.snellius.surf.nl]:37129): Broken pipe
137
+ Exception raised from sendBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:653 (most recent call first):
138
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x147d17e52fdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
139
+ frame #1: <unknown function> + 0x6a326d1 (0x147d5bfe06d1 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
140
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x24d (0x147d5bfdf1cd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
141
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x147d190c1eec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
142
+ frame #4: <unknown function> + 0xed164 (0x147e5852b164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
143
+ frame #5: <unknown function> + 0x8a19a (0x147e68f9d19a in /lib64/libc.so.6)
144
+ frame #6: <unknown function> + 0x10f100 (0x147e69022100 in /lib64/libc.so.6)
145
+
146
+ [rank11]:[W510 16:19:58.780439720 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 11] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Broken pipe
147
+ [rank11]:[W510 16:19:58.283183608 ProcessGroupNCCL.cpp:1553] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())
148
+ [rank11]:[W510 16:19:59.780555439 TCPStore.cpp:106] [c10d] sendBytes failed on SocketImpl(fd=3, addr=[gcn85.local.snellius.surf.nl]:48702, remote=[gcn80.local.snellius.surf.nl]:37129): Broken pipe
149
+ Exception raised from sendBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:653 (most recent call first):
150
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x147d17e52fdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
151
+ frame #1: <unknown function> + 0x6a326d1 (0x147d5bfe06d1 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
152
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x24d (0x147d5bfdf1cd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
153
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x147d190c1eec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
154
+ frame #4: <unknown function> + 0xed164 (0x147e5852b164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
155
+ frame #5: <unknown function> + 0x8a19a (0x147e68f9d19a in /lib64/libc.so.6)
156
+ frame #6: <unknown function> + 0x10f100 (0x147e69022100 in /lib64/libc.so.6)
157
+
158
+ [rank11]:[W510 16:19:59.782896125 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 11] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Broken pipe
vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/job_22614766/22614766_11_log.out ADDED
@@ -0,0 +1,280 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ submitit INFO (2026-05-10 12:14:34,049) - Starting with JobEnvironment(job_id=22614766, hostname=gcn85.local.snellius.surf.nl, local_rank=3(4), node=2(4), global_rank=11(16))
2
+ submitit INFO (2026-05-10 12:14:34,049) - Loading pickle: /scratch-shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/job_22614766/22614766_submitted.pkl
3
+ INFO:root:loaded pretrain params...
4
+ { 'app': 'vjepa',
5
+ 'cpus_per_task': 16,
6
+ 'data': { 'batch_size': 64,
7
+ 'crop_size': 224,
8
+ 'dataset_fpcs': [16, 16],
9
+ 'dataset_type': 'VideoDataset',
10
+ 'datasets': [ '/scratch-shared/dcanez/data/kinetics/k400/train.csv',
11
+ '/scratch-shared/dcanez/data/ssv2/train.csv'],
12
+ 'datasets_weights': [0.65, 0.35],
13
+ 'fps': 4,
14
+ 'num_workers': 10,
15
+ 'patch_size': 14,
16
+ 'persistent_workers': True,
17
+ 'pin_mem': True,
18
+ 'stage': [ { 'dest': 'kinetics_240',
19
+ 'format': 'targz_parts',
20
+ 'src': '/scratch-shared/dcanez/data/kinetics/k400/tars_240/'},
21
+ { 'dest': 'ssv2',
22
+ 'format': 'multipart_tar',
23
+ 'src': '/scratch-nvme/ml-datasets/something-something-v2/'}],
24
+ 'tubelet_size': 1},
25
+ 'data_aug': { 'auto_augment': False,
26
+ 'motion_shift': False,
27
+ 'random_resize_aspect_ratio': [0.75, 1.35],
28
+ 'random_resize_scale': [0.3, 1.0],
29
+ 'reprob': 0.0},
30
+ 'folder': '/scratch-shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable',
31
+ 'loss': {'loss_exp': 1.0},
32
+ 'mask': [ { 'aspect_ratio': [0.75, 1.5],
33
+ 'full_complement': False,
34
+ 'max_keep': None,
35
+ 'max_temporal_keep': 1.0,
36
+ 'num_blocks': 8,
37
+ 'spatial_scale': [0.15, 0.15],
38
+ 'temporal_scale': [1.0, 1.0]},
39
+ { 'aspect_ratio': [0.75, 1.5],
40
+ 'full_complement': False,
41
+ 'max_keep': None,
42
+ 'max_temporal_keep': 1.0,
43
+ 'num_blocks': 2,
44
+ 'spatial_scale': [0.7, 0.7],
45
+ 'temporal_scale': [1.0, 1.0]}],
46
+ 'mem_per_gpu': '180G',
47
+ 'meta': { 'dtype': 'bfloat16',
48
+ 'knn_eval_epoch0': False,
49
+ 'knn_eval_freq': 5,
50
+ 'knn_eval_presets': [ { 'config': { 'batch_size': 64,
51
+ 'dataset_train': '/scratch-shared/mdorkenw/ucf101/ucfTrainTestlist/train.csv',
52
+ 'dataset_val': '/scratch-shared/mdorkenw/ucf101/ucfTrainTestlist/val.csv',
53
+ 'eval_videos_per_class': 25,
54
+ 'num_workers': 8,
55
+ 'pool_type': 'slot_temporal_concat',
56
+ 'train_videos_per_class': 100},
57
+ 'preset': 'ucf101'},
58
+ { 'config': { 'batch_size': 64,
59
+ 'dataset_train': '/scratch-shared/mdorkenw/20bn-something-something-v2/something-something-v2-annotations/train_coarse10.csv',
60
+ 'dataset_val': '/scratch-shared/mdorkenw/20bn-something-something-v2/something-something-v2-annotations/val_coarse10.csv',
61
+ 'eval_videos_per_class': 100,
62
+ 'linear_probe': True,
63
+ 'num_workers': 8,
64
+ 'pool_type': 'slot_temporal_concat',
65
+ 'train_videos_per_class': 500},
66
+ 'preset': 'ssv2_coarse10'}],
67
+ 'load_checkpoint': True,
68
+ 'read_checkpoint': None,
69
+ 'save_every_freq': 5,
70
+ 'seed': 239,
71
+ 'use_sdpa': True,
72
+ 'use_wandb': True,
73
+ 'wandb_project': 'vjepa_ujepaside'},
74
+ 'metrics': {'sigreg': {}, 'std': {}},
75
+ 'model': { 'model_name': 'ujepaside_large_patch14_capi_lvd1689m',
76
+ 'pred_depth': 6,
77
+ 'pred_embed_dim': 384,
78
+ 'pred_num_heads': 12,
79
+ 'predictor': 'v2_cross',
80
+ 'st_causal': False,
81
+ 'st_drop_path': 0.2,
82
+ 'st_flex_enable': False,
83
+ 'st_layer_scale_init': 1e-05,
84
+ 'st_num_slots': 16,
85
+ 'st_slots_causal_within_frame': False,
86
+ 'target_kind': 'frozen_2d',
87
+ 'target_type': 'vit_large_patch14_capi.lvd1689m',
88
+ 'temporal_spacing': 1.0,
89
+ 'uniform_power': True,
90
+ 'use_activation_checkpointing': True,
91
+ 'use_mask_tokens': True,
92
+ 'use_rope': True,
93
+ 'use_sdpa': True,
94
+ 'zero_init_mask_tokens': True},
95
+ 'nodes': 4,
96
+ 'optimization': { 'clip_grad': 3.0,
97
+ 'ema': [0.99925, 0.99925],
98
+ 'epochs': 100,
99
+ 'final_lr': 0.0001,
100
+ 'final_weight_decay': 0.04,
101
+ 'ipe': 300,
102
+ 'ipe_scale': 1.0,
103
+ 'lr': 0.0005,
104
+ 'start_lr': 0.0001,
105
+ 'warmup': 10,
106
+ 'weight_decay': 0.04},
107
+ 'tasks_per_node': 4}
108
+ INFO:root:Running pre-training of app: vjepa
109
+ [INFO ][2026-05-10 12:16:54][app.vjepa.train ][main ] which_dtype='bfloat16'
110
+ [INFO ][2026-05-10 12:16:54][app.vjepa.train ][main ] Disabling persistent_workers (incompatible with KNN eval)
111
+ [INFO ][2026-05-10 12:16:54][app.vjepa.train ][main ] NCCL_SOCKET_IFNAME=eno
112
+ [INFO ][2026-05-10 12:16:57][app.vjepa.train ][main ] Initialized (rank/world-size) 11/16, tasks_per_node=4
113
+ [INFO ][2026-05-10 12:16:57][root ][stage_datasets ] [local_rank 3/4] Staging kinetics_240 (targz_parts)
114
+ [INFO ][2026-05-10 12:16:57][root ][_stage_targz_parts ] [rank 3] Extracting 25/103 tar.gz parts to /scratch-node/dcanez.22614766/kinetics_240
115
+ [INFO ][2026-05-10 12:17:11][root ][_stage_targz_parts ] [local_rank 3] Extracted 2/25 parts
116
+ [INFO ][2026-05-10 12:17:25][root ][_stage_targz_parts ] [local_rank 3] Extracted 4/25 parts
117
+ [INFO ][2026-05-10 12:17:39][root ][_stage_targz_parts ] [local_rank 3] Extracted 6/25 parts
118
+ [INFO ][2026-05-10 12:17:53][root ][_stage_targz_parts ] [local_rank 3] Extracted 8/25 parts
119
+ [INFO ][2026-05-10 12:18:07][root ][_stage_targz_parts ] [local_rank 3] Extracted 10/25 parts
120
+ [INFO ][2026-05-10 12:18:22][root ][_stage_targz_parts ] [local_rank 3] Extracted 12/25 parts
121
+ [INFO ][2026-05-10 12:18:36][root ][_stage_targz_parts ] [local_rank 3] Extracted 14/25 parts
122
+ [INFO ][2026-05-10 12:18:51][root ][_stage_targz_parts ] [local_rank 3] Extracted 16/25 parts
123
+ [INFO ][2026-05-10 12:19:04][root ][_stage_targz_parts ] [local_rank 3] Extracted 18/25 parts
124
+ [INFO ][2026-05-10 12:19:19][root ][_stage_targz_parts ] [local_rank 3] Extracted 20/25 parts
125
+ [INFO ][2026-05-10 12:19:33][root ][_stage_targz_parts ] [local_rank 3] Extracted 22/25 parts
126
+ [INFO ][2026-05-10 12:19:47][root ][_stage_targz_parts ] [local_rank 3] Extracted 24/25 parts
127
+ [INFO ][2026-05-10 12:19:55][root ][_stage_targz_parts ] [local_rank 3] Extracted 25/25 parts
128
+ [INFO ][2026-05-10 12:19:55][root ][stage_datasets ] [local_rank 3/4] Staging ssv2 (multipart_tar)
129
+ [INFO ][2026-05-10 12:21:03][app.vjepa.train ][main ] Staged datasets: ['/scratch-node/dcanez.22614766/kinetics_240/train.csv', '/scratch-node/dcanez.22614766/ssv2/train.csv']
130
+ [INFO ][2026-05-10 12:23:31][experiments.stmodels.vision_transformers_v3][vit_large_patch14_capi_lvd1689m] Loading pretrained weights for vit_large_patch14_capi_lvd1689m from https://dl.fbaipublicfiles.com/capi/capi_vitl14_lvd.pth
131
+ [INFO ][2026-05-10 12:23:55][root ][_build_frozen_2d_target ] Loaded frozen_2d target from timm: vit_large_patch14_capi.lvd1689m
132
+ [INFO ][2026-05-10 12:23:55][root ][init_video_model ] ViTMultiSeqWrapper(
133
+ (backbone): UJEPAside(
134
+ (patch_embed): PatchEmbed3D(
135
+ (proj): Conv3d(3, 1024, kernel_size=(1, 14, 14), stride=(1, 14, 14))
136
+ )
137
+ (rope): CAPI2DRoPE()
138
+ (blocks): ModuleList(
139
+ (0-23): 24 x Block(
140
+ (norm1): RMSNorm((1024,), eps=1e-05, elementwise_affine=True)
141
+ (rope_impl): CAPI2DRoPE()
142
+ (norm2): RMSNorm((1024,), eps=1e-05, elementwise_affine=True)
143
+ (drop_path1): Identity()
144
+ (drop_path2): Identity()
145
+ (attn): Attention(
146
+ (qkv): Linear(in_features=1024, out_features=3072, bias=False)
147
+ (attn_drop): Dropout(p=0.0, inplace=False)
148
+ (proj): Linear(in_features=1024, out_features=1024, bias=False)
149
+ (proj_drop): Dropout(p=0.0, inplace=False)
150
+ (rope_impl): CAPI2DRoPE()
151
+ )
152
+ (mlp): MLP(
153
+ (fc1): Linear(in_features=1024, out_features=4096, bias=False)
154
+ (act): GELU(approximate='none')
155
+ (fc2): Linear(in_features=4096, out_features=1024, bias=False)
156
+ (drop): Dropout(p=0.0, inplace=False)
157
+ )
158
+ )
159
+ )
160
+ (norm): RMSNorm((1024,), eps=1e-05, elementwise_affine=False)
161
+ (st_blocks): ModuleList(
162
+ (0-23): 24 x Block(
163
+ (residual1): EfficientResidual(
164
+ (norm): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
165
+ (fn): Attention(
166
+ (q_proj): Linear(in_features=1024, out_features=1024, bias=False)
167
+ (k_proj): Linear(in_features=1024, out_features=1024, bias=False)
168
+ (v_proj): Linear(in_features=1024, out_features=1024, bias=False)
169
+ (proj): Linear(in_features=1024, out_features=1024, bias=False)
170
+ (rope_impl): CAPI3DRoPE()
171
+ )
172
+ )
173
+ (residual2): EfficientResidual(
174
+ (norm): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
175
+ (fn): MLP(
176
+ (fc1): Linear(in_features=1024, out_features=4096, bias=False)
177
+ (act): GELU(approximate='none')
178
+ (fc2): Linear(in_features=4096, out_features=1024, bias=False)
179
+ (drop): Dropout(p=0.0, inplace=False)
180
+ )
181
+ )
182
+ )
183
+ )
184
+ (st_rope): CAPI3DRoPE()
185
+ )
186
+ )
187
+ [INFO ][2026-05-10 12:23:55][root ][init_video_model ] UJEPAsidePredictorMultiSeqWrapper(
188
+ (backbone): PredictorV2(
189
+ (predictor_embed): Linear(in_features=1024, out_features=384, bias=True)
190
+ (mask_tokens): ParameterList(
191
+ (0): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
192
+ (1): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
193
+ (2): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
194
+ (3): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
195
+ )
196
+ (predictor_blocks): ModuleList(
197
+ (0-5): 6 x Block(
198
+ (residual1): EfficientResidual(
199
+ (norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
200
+ (fn): Attention(
201
+ (q_proj): Linear(in_features=384, out_features=384, bias=False)
202
+ (k_proj): Linear(in_features=384, out_features=384, bias=False)
203
+ (v_proj): Linear(in_features=384, out_features=384, bias=False)
204
+ (proj): Linear(in_features=384, out_features=384, bias=False)
205
+ (rope): Rope()
206
+ )
207
+ )
208
+ (residual2): EfficientResidual(
209
+ (norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
210
+ (fn): MLP(
211
+ (fc1): Linear(in_features=384, out_features=1536, bias=False)
212
+ (act): GELU(approximate='none')
213
+ (fc2): Linear(in_features=1536, out_features=384, bias=False)
214
+ (drop): Dropout(p=0.0, inplace=False)
215
+ )
216
+ )
217
+ )
218
+ )
219
+ (predictor_norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
220
+ (predictor_proj): Linear(in_features=384, out_features=1024, bias=True)
221
+ )
222
+ )
223
+ [INFO ][2026-05-10 12:23:55][root ][init_video_model ] MultiSeqWrapper(
224
+ (backbone): Frozen2DTargetWrapper(
225
+ (backbone): Eva(
226
+ (patch_embed): PatchEmbed(
227
+ (proj): Conv2d(3, 1024, kernel_size=(14, 14), stride=(14, 14))
228
+ (norm): Identity()
229
+ )
230
+ (pos_drop): Dropout(p=0.0, inplace=False)
231
+ (norm_pre): Identity()
232
+ (blocks): ModuleList(
233
+ (0-23): 24 x EvaBlock(
234
+ (norm1): RMSNorm((1024,), eps=1e-05, elementwise_affine=True)
235
+ (attn): EvaAttention(
236
+ (qkv): Linear(in_features=1024, out_features=3072, bias=False)
237
+ (q_norm): Identity()
238
+ (k_norm): Identity()
239
+ (attn_drop): Dropout(p=0.0, inplace=False)
240
+ (norm): Identity()
241
+ (proj): Linear(in_features=1024, out_features=1024, bias=True)
242
+ (proj_drop): Dropout(p=0.0, inplace=False)
243
+ )
244
+ (drop_path1): Identity()
245
+ (norm2): RMSNorm((1024,), eps=1e-05, elementwise_affine=True)
246
+ (mlp): Mlp(
247
+ (fc1): Linear(in_features=1024, out_features=4096, bias=True)
248
+ (act): GELU(approximate='none')
249
+ (drop1): Dropout(p=0.0, inplace=False)
250
+ (norm): Identity()
251
+ (fc2): Linear(in_features=4096, out_features=1024, bias=True)
252
+ (drop2): Dropout(p=0.0, inplace=False)
253
+ )
254
+ (drop_path2): Identity()
255
+ )
256
+ )
257
+ (norm): RMSNorm((1024,), eps=1e-05, elementwise_affine=True)
258
+ (fc_norm): Identity()
259
+ (head_drop): Dropout(p=0.0, inplace=False)
260
+ (head): Identity()
261
+ (rope): _CapiPatchRoPE()
262
+ )
263
+ )
264
+ )
265
+ [INFO ][2026-05-10 12:23:55][root ][init_video_model ] Encoder number of parameters: 302399488
266
+ [INFO ][2026-05-10 12:23:55][root ][init_video_model ] Predictor number of parameters: 11416192
267
+ [INFO ][2026-05-10 12:23:55][root ][init_video_model ] Target encoder number of parameters: 0
268
+ [INFO ][2026-05-10 12:23:56][root ][make_videodataset ] VideoDataset dataset created
269
+ [INFO ][2026-05-10 12:23:56][WeightedSampler ][__init__ ] Using DistributedWeightedSampler with rank 11 / 16
270
+ [INFO ][2026-05-10 12:23:56][root ][make_videodataset ] VideoDataset unsupervised data loader created
271
+ [INFO ][2026-05-10 12:23:56][app.vjepa.train ][main ] iterations per epoch/dataset length: 300/399
272
+ [INFO ][2026-05-10 12:23:56][app.vjepa.train ][main ] Wrapping models in DDP (rank 11)...
273
+ [INFO ][2026-05-10 12:24:08][app.vjepa.train ][main ] Initializing loader...
274
+ submitit WARNING (2026-05-10 16:19:58,222) - Bypassing signal SIGTERM
275
+ [WARNING ][2026-05-10 16:19:58][submitit ][bypass ] Bypassing signal SIGTERM
276
+ submitit WARNING (2026-05-10 16:19:58,234) - Bypassing signal SIGCONT
277
+ [WARNING ][2026-05-10 16:19:58][submitit ][bypass ] Bypassing signal SIGCONT
278
+ submitit ERROR (2026-05-10 16:19:58,234) - Submitted job triggered an exception
279
+ [ERROR ][2026-05-10 16:19:58][submitit ][process_job ] Submitted job triggered an exception
280
+ [ERROR ][2026-05-10 16:19:58][app.vjepa.train ][_save_crash ] Saved crash log to /scratch-shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/crash_rank11.log
vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/job_22614766/22614766_12_log.err ADDED
@@ -0,0 +1,158 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/timm/models/layers/__init__.py:49: FutureWarning: Importing from timm.models.layers is deprecated, please import via timm.layers
2
+ warnings.warn(f"Importing from {__name__} is deprecated, please import via timm.layers", FutureWarning)
3
+ [rank12]:[W510 12:20:52.172098656 ProcessGroupNCCL.cpp:5138] Guessing device ID based on global rank. This can cause a hang if rank to GPU mapping is heterogeneous. You can specify device_id in init_process_group()
4
+ /gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/vjepa/utils.py:796: FutureWarning: `torch.cuda.amp.GradScaler(args...)` is deprecated. Please use `torch.amp.GradScaler('cuda', args...)` instead.
5
+ scaler = torch.cuda.amp.GradScaler() if mixed_precision else None
6
+ /gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/vjepa/train.py:955: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
7
+ with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
8
+ [rank12]:[W510 16:19:49.920378344 TCPStore.cpp:125] [c10d] recvValue failed on SocketImpl(fd=4, addr=[gcn87.local.snellius.surf.nl]:50944, remote=[gcn80.local.snellius.surf.nl]:37129): Failed to recv, got 0 bytes. Connection was likely closed. Did the remote server shutdown or crash?
9
+ Exception raised from recvBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:682 (most recent call first):
10
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x14a51dc5efdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
11
+ frame #1: <unknown function> + 0x6a3325d (0x14a561ded25d in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
12
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x273 (0x14a561deb1f3 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
13
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x14a51eecdeec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
14
+ frame #4: <unknown function> + 0xed164 (0x14a65e337164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
15
+ frame #5: <unknown function> + 0x8a19a (0x14a66eda919a in /lib64/libc.so.6)
16
+ frame #6: <unknown function> + 0x10f100 (0x14a66ee2e100 in /lib64/libc.so.6)
17
+
18
+ [rank12]:[W510 16:19:49.926036927 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 12] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Failed to recv, got 0 bytes. Connection was likely closed. Did the remote server shutdown or crash?
19
+ [rank12]:[W510 16:19:50.926190511 TCPStore.cpp:106] [c10d] sendBytes failed on SocketImpl(fd=4, addr=[gcn87.local.snellius.surf.nl]:50944, remote=[gcn80.local.snellius.surf.nl]:37129): Broken pipe
20
+ Exception raised from sendBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:653 (most recent call first):
21
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x14a51dc5efdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
22
+ frame #1: <unknown function> + 0x6a326d1 (0x14a561dec6d1 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
23
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x24d (0x14a561deb1cd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
24
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x14a51eecdeec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
25
+ frame #4: <unknown function> + 0xed164 (0x14a65e337164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
26
+ frame #5: <unknown function> + 0x8a19a (0x14a66eda919a in /lib64/libc.so.6)
27
+ frame #6: <unknown function> + 0x10f100 (0x14a66ee2e100 in /lib64/libc.so.6)
28
+
29
+ [rank12]:[W510 16:19:50.929892860 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 12] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Broken pipe
30
+ [rank12]:[W510 16:19:51.930033624 TCPStore.cpp:106] [c10d] sendBytes failed on SocketImpl(fd=4, addr=[gcn87.local.snellius.surf.nl]:50944, remote=[gcn80.local.snellius.surf.nl]:37129): Broken pipe
31
+ Exception raised from sendBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:653 (most recent call first):
32
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x14a51dc5efdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
33
+ frame #1: <unknown function> + 0x6a326d1 (0x14a561dec6d1 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
34
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x24d (0x14a561deb1cd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
35
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x14a51eecdeec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
36
+ frame #4: <unknown function> + 0xed164 (0x14a65e337164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
37
+ frame #5: <unknown function> + 0x8a19a (0x14a66eda919a in /lib64/libc.so.6)
38
+ frame #6: <unknown function> + 0x10f100 (0x14a66ee2e100 in /lib64/libc.so.6)
39
+
40
+ [rank12]:[W510 16:19:51.932343947 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 12] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Broken pipe
41
+ [rank12]:[W510 16:19:52.932472232 TCPStore.cpp:106] [c10d] sendBytes failed on SocketImpl(fd=4, addr=[gcn87.local.snellius.surf.nl]:50944, remote=[gcn80.local.snellius.surf.nl]:37129): Broken pipe
42
+ Exception raised from sendBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:653 (most recent call first):
43
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x14a51dc5efdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
44
+ frame #1: <unknown function> + 0x6a326d1 (0x14a561dec6d1 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
45
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x24d (0x14a561deb1cd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
46
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x14a51eecdeec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
47
+ frame #4: <unknown function> + 0xed164 (0x14a65e337164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
48
+ frame #5: <unknown function> + 0x8a19a (0x14a66eda919a in /lib64/libc.so.6)
49
+ frame #6: <unknown function> + 0x10f100 (0x14a66ee2e100 in /lib64/libc.so.6)
50
+
51
+ [rank12]:[W510 16:19:52.934789205 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 12] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Broken pipe
52
+ [rank12]:[W510 16:19:53.934942109 TCPStore.cpp:106] [c10d] sendBytes failed on SocketImpl(fd=4, addr=[gcn87.local.snellius.surf.nl]:50944, remote=[gcn80.local.snellius.surf.nl]:37129): Broken pipe
53
+ Exception raised from sendBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:653 (most recent call first):
54
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x14a51dc5efdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
55
+ frame #1: <unknown function> + 0x6a326d1 (0x14a561dec6d1 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
56
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x24d (0x14a561deb1cd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
57
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x14a51eecdeec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
58
+ frame #4: <unknown function> + 0xed164 (0x14a65e337164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
59
+ frame #5: <unknown function> + 0x8a19a (0x14a66eda919a in /lib64/libc.so.6)
60
+ frame #6: <unknown function> + 0x10f100 (0x14a66ee2e100 in /lib64/libc.so.6)
61
+
62
+ [rank12]:[W510 16:19:53.937294752 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 12] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Broken pipe
63
+ [rank12]:[W510 16:19:54.937423426 TCPStore.cpp:106] [c10d] sendBytes failed on SocketImpl(fd=4, addr=[gcn87.local.snellius.surf.nl]:50944, remote=[gcn80.local.snellius.surf.nl]:37129): Broken pipe
64
+ Exception raised from sendBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:653 (most recent call first):
65
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x14a51dc5efdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
66
+ frame #1: <unknown function> + 0x6a326d1 (0x14a561dec6d1 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
67
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x24d (0x14a561deb1cd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
68
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x14a51eecdeec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
69
+ frame #4: <unknown function> + 0xed164 (0x14a65e337164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
70
+ frame #5: <unknown function> + 0x8a19a (0x14a66eda919a in /lib64/libc.so.6)
71
+ frame #6: <unknown function> + 0x10f100 (0x14a66ee2e100 in /lib64/libc.so.6)
72
+
73
+ [rank12]:[W510 16:19:54.939693100 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 12] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Broken pipe
74
+ [rank12]:[W510 16:19:55.939816784 TCPStore.cpp:106] [c10d] sendBytes failed on SocketImpl(fd=4, addr=[gcn87.local.snellius.surf.nl]:50944, remote=[gcn80.local.snellius.surf.nl]:37129): Broken pipe
75
+ Exception raised from sendBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:653 (most recent call first):
76
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x14a51dc5efdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
77
+ frame #1: <unknown function> + 0x6a326d1 (0x14a561dec6d1 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
78
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x24d (0x14a561deb1cd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
79
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x14a51eecdeec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
80
+ frame #4: <unknown function> + 0xed164 (0x14a65e337164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
81
+ frame #5: <unknown function> + 0x8a19a (0x14a66eda919a in /lib64/libc.so.6)
82
+ frame #6: <unknown function> + 0x10f100 (0x14a66ee2e100 in /lib64/libc.so.6)
83
+
84
+ [rank12]:[W510 16:19:55.942163637 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 12] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Broken pipe
85
+ [rank12]:[W510 16:19:56.942310751 TCPStore.cpp:106] [c10d] sendBytes failed on SocketImpl(fd=4, addr=[gcn87.local.snellius.surf.nl]:50944, remote=[gcn80.local.snellius.surf.nl]:37129): Broken pipe
86
+ Exception raised from sendBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:653 (most recent call first):
87
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x14a51dc5efdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
88
+ frame #1: <unknown function> + 0x6a326d1 (0x14a561dec6d1 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
89
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x24d (0x14a561deb1cd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
90
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x14a51eecdeec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
91
+ frame #4: <unknown function> + 0xed164 (0x14a65e337164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
92
+ frame #5: <unknown function> + 0x8a19a (0x14a66eda919a in /lib64/libc.so.6)
93
+ frame #6: <unknown function> + 0x10f100 (0x14a66ee2e100 in /lib64/libc.so.6)
94
+
95
+ [rank12]:[W510 16:19:56.944697424 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 12] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Broken pipe
96
+ [rank12]:[W510 16:19:57.944832968 TCPStore.cpp:106] [c10d] sendBytes failed on SocketImpl(fd=4, addr=[gcn87.local.snellius.surf.nl]:50944, remote=[gcn80.local.snellius.surf.nl]:37129): Broken pipe
97
+ Exception raised from sendBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:653 (most recent call first):
98
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x14a51dc5efdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
99
+ frame #1: <unknown function> + 0x6a326d1 (0x14a561dec6d1 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
100
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x24d (0x14a561deb1cd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
101
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x14a51eecdeec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
102
+ frame #4: <unknown function> + 0xed164 (0x14a65e337164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
103
+ frame #5: <unknown function> + 0x8a19a (0x14a66eda919a in /lib64/libc.so.6)
104
+ frame #6: <unknown function> + 0x10f100 (0x14a66ee2e100 in /lib64/libc.so.6)
105
+
106
+ [rank12]:[W510 16:19:57.947139571 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 12] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Broken pipe
107
+ submitit WARNING (2026-05-10 16:19:58,222) - Bypassing signal SIGTERM
108
+ [rank12]:[W510 16:19:58.947284656 TCPStore.cpp:106] [c10d] sendBytes failed on SocketImpl(fd=4, addr=[gcn87.local.snellius.surf.nl]:50944, remote=[gcn80.local.snellius.surf.nl]:37129): Broken pipe
109
+ Exception raised from sendBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:653 (most recent call first):
110
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x14a51dc5efdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
111
+ frame #1: <unknown function> + 0x6a326d1 (0x14a561dec6d1 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
112
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x24d (0x14a561deb1cd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
113
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x14a51eecdeec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
114
+ frame #4: <unknown function> + 0xed164 (0x14a65e337164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
115
+ frame #5: <unknown function> + 0x8a19a (0x14a66eda919a in /lib64/libc.so.6)
116
+ frame #6: <unknown function> + 0x10f100 (0x14a66ee2e100 in /lib64/libc.so.6)
117
+
118
+ [rank12]:[W510 16:19:58.949610339 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 12] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Broken pipe
119
+ submitit WARNING (2026-05-10 16:19:58,948) - Bypassing signal SIGCONT
120
+ submitit ERROR (2026-05-10 16:19:58,949) - Submitted job triggered an exception
121
+ Traceback (most recent call last):
122
+ File "<frozen runpy>", line 198, in _run_module_as_main
123
+ File "<frozen runpy>", line 88, in _run_code
124
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/_submit.py", line 11, in <module>
125
+ submitit_main()
126
+ ~~~~~~~~~~~~~^^
127
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 76, in submitit_main
128
+ process_job(args.folder)
129
+ ~~~~~~~~~~~^^^^^^^^^^^^^
130
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 69, in process_job
131
+ raise error
132
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 55, in process_job
133
+ result = delayed.result()
134
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/utils.py", line 137, in result
135
+ self._result = self.function(*self.args, **self.kwargs)
136
+ ~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^
137
+ File "/gpfs/home2/dcanez/vd/app/main_distributed.py", line 102, in __call__
138
+ app_main(app, args=params, resume_preempt=resume_preempt)
139
+ ~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
140
+ File "/gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/scaffold.py", line 17, in main
141
+ return importlib.import_module(f"app.{app}.train").main(args=args, resume_preempt=resume_preempt)
142
+ ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
143
+ File "/gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/vjepa/train.py", line 1189, in main
144
+ assert not np.isnan(loss), "loss is nan"
145
+ ^^^^^^^^^^^^^^^^^^
146
+ AssertionError: loss is nan
147
+ [rank12]:[W510 16:19:59.949764342 TCPStore.cpp:106] [c10d] sendBytes failed on SocketImpl(fd=4, addr=[gcn87.local.snellius.surf.nl]:50944, remote=[gcn80.local.snellius.surf.nl]:37129): Broken pipe
148
+ Exception raised from sendBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:653 (most recent call first):
149
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x14a51dc5efdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
150
+ frame #1: <unknown function> + 0x6a326d1 (0x14a561dec6d1 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
151
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x24d (0x14a561deb1cd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
152
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x14a51eecdeec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
153
+ frame #4: <unknown function> + 0xed164 (0x14a65e337164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
154
+ frame #5: <unknown function> + 0x8a19a (0x14a66eda919a in /lib64/libc.so.6)
155
+ frame #6: <unknown function> + 0x10f100 (0x14a66ee2e100 in /lib64/libc.so.6)
156
+
157
+ [rank12]:[W510 16:19:59.952134326 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 12] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Broken pipe
158
+ [rank12]:[W510 16:19:59.180205490 ProcessGroupNCCL.cpp:1553] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())
vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/job_22614766/22614766_12_log.out ADDED
@@ -0,0 +1,284 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ submitit INFO (2026-05-10 12:14:34,055) - Starting with JobEnvironment(job_id=22614766, hostname=gcn87.local.snellius.surf.nl, local_rank=0(4), node=3(4), global_rank=12(16))
2
+ submitit INFO (2026-05-10 12:14:34,056) - Loading pickle: /scratch-shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/job_22614766/22614766_submitted.pkl
3
+ INFO:root:loaded pretrain params...
4
+ { 'app': 'vjepa',
5
+ 'cpus_per_task': 16,
6
+ 'data': { 'batch_size': 64,
7
+ 'crop_size': 224,
8
+ 'dataset_fpcs': [16, 16],
9
+ 'dataset_type': 'VideoDataset',
10
+ 'datasets': [ '/scratch-shared/dcanez/data/kinetics/k400/train.csv',
11
+ '/scratch-shared/dcanez/data/ssv2/train.csv'],
12
+ 'datasets_weights': [0.65, 0.35],
13
+ 'fps': 4,
14
+ 'num_workers': 10,
15
+ 'patch_size': 14,
16
+ 'persistent_workers': True,
17
+ 'pin_mem': True,
18
+ 'stage': [ { 'dest': 'kinetics_240',
19
+ 'format': 'targz_parts',
20
+ 'src': '/scratch-shared/dcanez/data/kinetics/k400/tars_240/'},
21
+ { 'dest': 'ssv2',
22
+ 'format': 'multipart_tar',
23
+ 'src': '/scratch-nvme/ml-datasets/something-something-v2/'}],
24
+ 'tubelet_size': 1},
25
+ 'data_aug': { 'auto_augment': False,
26
+ 'motion_shift': False,
27
+ 'random_resize_aspect_ratio': [0.75, 1.35],
28
+ 'random_resize_scale': [0.3, 1.0],
29
+ 'reprob': 0.0},
30
+ 'folder': '/scratch-shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable',
31
+ 'loss': {'loss_exp': 1.0},
32
+ 'mask': [ { 'aspect_ratio': [0.75, 1.5],
33
+ 'full_complement': False,
34
+ 'max_keep': None,
35
+ 'max_temporal_keep': 1.0,
36
+ 'num_blocks': 8,
37
+ 'spatial_scale': [0.15, 0.15],
38
+ 'temporal_scale': [1.0, 1.0]},
39
+ { 'aspect_ratio': [0.75, 1.5],
40
+ 'full_complement': False,
41
+ 'max_keep': None,
42
+ 'max_temporal_keep': 1.0,
43
+ 'num_blocks': 2,
44
+ 'spatial_scale': [0.7, 0.7],
45
+ 'temporal_scale': [1.0, 1.0]}],
46
+ 'mem_per_gpu': '180G',
47
+ 'meta': { 'dtype': 'bfloat16',
48
+ 'knn_eval_epoch0': False,
49
+ 'knn_eval_freq': 5,
50
+ 'knn_eval_presets': [ { 'config': { 'batch_size': 64,
51
+ 'dataset_train': '/scratch-shared/mdorkenw/ucf101/ucfTrainTestlist/train.csv',
52
+ 'dataset_val': '/scratch-shared/mdorkenw/ucf101/ucfTrainTestlist/val.csv',
53
+ 'eval_videos_per_class': 25,
54
+ 'num_workers': 8,
55
+ 'pool_type': 'slot_temporal_concat',
56
+ 'train_videos_per_class': 100},
57
+ 'preset': 'ucf101'},
58
+ { 'config': { 'batch_size': 64,
59
+ 'dataset_train': '/scratch-shared/mdorkenw/20bn-something-something-v2/something-something-v2-annotations/train_coarse10.csv',
60
+ 'dataset_val': '/scratch-shared/mdorkenw/20bn-something-something-v2/something-something-v2-annotations/val_coarse10.csv',
61
+ 'eval_videos_per_class': 100,
62
+ 'linear_probe': True,
63
+ 'num_workers': 8,
64
+ 'pool_type': 'slot_temporal_concat',
65
+ 'train_videos_per_class': 500},
66
+ 'preset': 'ssv2_coarse10'}],
67
+ 'load_checkpoint': True,
68
+ 'read_checkpoint': None,
69
+ 'save_every_freq': 5,
70
+ 'seed': 239,
71
+ 'use_sdpa': True,
72
+ 'use_wandb': True,
73
+ 'wandb_project': 'vjepa_ujepaside'},
74
+ 'metrics': {'sigreg': {}, 'std': {}},
75
+ 'model': { 'model_name': 'ujepaside_large_patch14_capi_lvd1689m',
76
+ 'pred_depth': 6,
77
+ 'pred_embed_dim': 384,
78
+ 'pred_num_heads': 12,
79
+ 'predictor': 'v2_cross',
80
+ 'st_causal': False,
81
+ 'st_drop_path': 0.2,
82
+ 'st_flex_enable': False,
83
+ 'st_layer_scale_init': 1e-05,
84
+ 'st_num_slots': 16,
85
+ 'st_slots_causal_within_frame': False,
86
+ 'target_kind': 'frozen_2d',
87
+ 'target_type': 'vit_large_patch14_capi.lvd1689m',
88
+ 'temporal_spacing': 1.0,
89
+ 'uniform_power': True,
90
+ 'use_activation_checkpointing': True,
91
+ 'use_mask_tokens': True,
92
+ 'use_rope': True,
93
+ 'use_sdpa': True,
94
+ 'zero_init_mask_tokens': True},
95
+ 'nodes': 4,
96
+ 'optimization': { 'clip_grad': 3.0,
97
+ 'ema': [0.99925, 0.99925],
98
+ 'epochs': 100,
99
+ 'final_lr': 0.0001,
100
+ 'final_weight_decay': 0.04,
101
+ 'ipe': 300,
102
+ 'ipe_scale': 1.0,
103
+ 'lr': 0.0005,
104
+ 'start_lr': 0.0001,
105
+ 'warmup': 10,
106
+ 'weight_decay': 0.04},
107
+ 'tasks_per_node': 4}
108
+ INFO:root:Running pre-training of app: vjepa
109
+ [INFO ][2026-05-10 12:16:33][app.vjepa.train ][main ] which_dtype='bfloat16'
110
+ [INFO ][2026-05-10 12:16:33][app.vjepa.train ][main ] Disabling persistent_workers (incompatible with KNN eval)
111
+ [INFO ][2026-05-10 12:16:33][app.vjepa.train ][main ] NCCL_SOCKET_IFNAME=eno
112
+ [INFO ][2026-05-10 12:16:54][app.vjepa.train ][main ] Initialized (rank/world-size) 12/16, tasks_per_node=4
113
+ [INFO ][2026-05-10 12:16:54][root ][stage_datasets ] [local_rank 0/4] Staging kinetics_240 (targz_parts)
114
+ [INFO ][2026-05-10 12:16:54][root ][_stage_targz_parts ] [rank 0] Extracting 26/103 tar.gz parts to /scratch-node/dcanez.22614766/kinetics_240
115
+ [INFO ][2026-05-10 12:17:08][root ][_stage_targz_parts ] [local_rank 0] Extracted 2/26 parts
116
+ [INFO ][2026-05-10 12:17:22][root ][_stage_targz_parts ] [local_rank 0] Extracted 4/26 parts
117
+ [INFO ][2026-05-10 12:17:36][root ][_stage_targz_parts ] [local_rank 0] Extracted 6/26 parts
118
+ [INFO ][2026-05-10 12:17:50][root ][_stage_targz_parts ] [local_rank 0] Extracted 8/26 parts
119
+ [INFO ][2026-05-10 12:18:04][root ][_stage_targz_parts ] [local_rank 0] Extracted 10/26 parts
120
+ [INFO ][2026-05-10 12:18:19][root ][_stage_targz_parts ] [local_rank 0] Extracted 12/26 parts
121
+ [INFO ][2026-05-10 12:18:33][root ][_stage_targz_parts ] [local_rank 0] Extracted 14/26 parts
122
+ [INFO ][2026-05-10 12:18:47][root ][_stage_targz_parts ] [local_rank 0] Extracted 16/26 parts
123
+ [INFO ][2026-05-10 12:19:01][root ][_stage_targz_parts ] [local_rank 0] Extracted 18/26 parts
124
+ [INFO ][2026-05-10 12:19:16][root ][_stage_targz_parts ] [local_rank 0] Extracted 20/26 parts
125
+ [INFO ][2026-05-10 12:19:30][root ][_stage_targz_parts ] [local_rank 0] Extracted 22/26 parts
126
+ [INFO ][2026-05-10 12:19:44][root ][_stage_targz_parts ] [local_rank 0] Extracted 24/26 parts
127
+ [INFO ][2026-05-10 12:19:58][root ][_stage_targz_parts ] [local_rank 0] Extracted 26/26 parts
128
+ [INFO ][2026-05-10 12:19:58][root ][stage_datasets ] [local_rank 0/4] Staging ssv2 (multipart_tar)
129
+ [INFO ][2026-05-10 12:19:58][root ][_stage_multipart_tar ] [rank 0] Extracting multipart tar (2 files) to /scratch-node/dcanez.22614766/ssv2
130
+ [INFO ][2026-05-10 12:21:02][root ][stage_datasets ] Data staging completed in 248.1s (4.1min)
131
+ [INFO ][2026-05-10 12:21:02][root ][_rewrite_csv ] Wrote local CSV: /scratch-node/dcanez.22614766/kinetics_240/train.csv (239789 entries)
132
+ [INFO ][2026-05-10 12:21:03][root ][_rewrite_csv ] Wrote local CSV: /scratch-node/dcanez.22614766/ssv2/train.csv (168913 entries)
133
+ [INFO ][2026-05-10 12:21:03][app.vjepa.train ][main ] Staged datasets: ['/scratch-node/dcanez.22614766/kinetics_240/train.csv', '/scratch-node/dcanez.22614766/ssv2/train.csv']
134
+ [INFO ][2026-05-10 12:23:31][experiments.stmodels.vision_transformers_v3][vit_large_patch14_capi_lvd1689m] Loading pretrained weights for vit_large_patch14_capi_lvd1689m from https://dl.fbaipublicfiles.com/capi/capi_vitl14_lvd.pth
135
+ [INFO ][2026-05-10 12:23:57][root ][_build_frozen_2d_target ] Loaded frozen_2d target from timm: vit_large_patch14_capi.lvd1689m
136
+ [INFO ][2026-05-10 12:23:57][root ][init_video_model ] ViTMultiSeqWrapper(
137
+ (backbone): UJEPAside(
138
+ (patch_embed): PatchEmbed3D(
139
+ (proj): Conv3d(3, 1024, kernel_size=(1, 14, 14), stride=(1, 14, 14))
140
+ )
141
+ (rope): CAPI2DRoPE()
142
+ (blocks): ModuleList(
143
+ (0-23): 24 x Block(
144
+ (norm1): RMSNorm((1024,), eps=1e-05, elementwise_affine=True)
145
+ (rope_impl): CAPI2DRoPE()
146
+ (norm2): RMSNorm((1024,), eps=1e-05, elementwise_affine=True)
147
+ (drop_path1): Identity()
148
+ (drop_path2): Identity()
149
+ (attn): Attention(
150
+ (qkv): Linear(in_features=1024, out_features=3072, bias=False)
151
+ (attn_drop): Dropout(p=0.0, inplace=False)
152
+ (proj): Linear(in_features=1024, out_features=1024, bias=False)
153
+ (proj_drop): Dropout(p=0.0, inplace=False)
154
+ (rope_impl): CAPI2DRoPE()
155
+ )
156
+ (mlp): MLP(
157
+ (fc1): Linear(in_features=1024, out_features=4096, bias=False)
158
+ (act): GELU(approximate='none')
159
+ (fc2): Linear(in_features=4096, out_features=1024, bias=False)
160
+ (drop): Dropout(p=0.0, inplace=False)
161
+ )
162
+ )
163
+ )
164
+ (norm): RMSNorm((1024,), eps=1e-05, elementwise_affine=False)
165
+ (st_blocks): ModuleList(
166
+ (0-23): 24 x Block(
167
+ (residual1): EfficientResidual(
168
+ (norm): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
169
+ (fn): Attention(
170
+ (q_proj): Linear(in_features=1024, out_features=1024, bias=False)
171
+ (k_proj): Linear(in_features=1024, out_features=1024, bias=False)
172
+ (v_proj): Linear(in_features=1024, out_features=1024, bias=False)
173
+ (proj): Linear(in_features=1024, out_features=1024, bias=False)
174
+ (rope_impl): CAPI3DRoPE()
175
+ )
176
+ )
177
+ (residual2): EfficientResidual(
178
+ (norm): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
179
+ (fn): MLP(
180
+ (fc1): Linear(in_features=1024, out_features=4096, bias=False)
181
+ (act): GELU(approximate='none')
182
+ (fc2): Linear(in_features=4096, out_features=1024, bias=False)
183
+ (drop): Dropout(p=0.0, inplace=False)
184
+ )
185
+ )
186
+ )
187
+ )
188
+ (st_rope): CAPI3DRoPE()
189
+ )
190
+ )
191
+ [INFO ][2026-05-10 12:23:57][root ][init_video_model ] UJEPAsidePredictorMultiSeqWrapper(
192
+ (backbone): PredictorV2(
193
+ (predictor_embed): Linear(in_features=1024, out_features=384, bias=True)
194
+ (mask_tokens): ParameterList(
195
+ (0): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
196
+ (1): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
197
+ (2): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
198
+ (3): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
199
+ )
200
+ (predictor_blocks): ModuleList(
201
+ (0-5): 6 x Block(
202
+ (residual1): EfficientResidual(
203
+ (norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
204
+ (fn): Attention(
205
+ (q_proj): Linear(in_features=384, out_features=384, bias=False)
206
+ (k_proj): Linear(in_features=384, out_features=384, bias=False)
207
+ (v_proj): Linear(in_features=384, out_features=384, bias=False)
208
+ (proj): Linear(in_features=384, out_features=384, bias=False)
209
+ (rope): Rope()
210
+ )
211
+ )
212
+ (residual2): EfficientResidual(
213
+ (norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
214
+ (fn): MLP(
215
+ (fc1): Linear(in_features=384, out_features=1536, bias=False)
216
+ (act): GELU(approximate='none')
217
+ (fc2): Linear(in_features=1536, out_features=384, bias=False)
218
+ (drop): Dropout(p=0.0, inplace=False)
219
+ )
220
+ )
221
+ )
222
+ )
223
+ (predictor_norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
224
+ (predictor_proj): Linear(in_features=384, out_features=1024, bias=True)
225
+ )
226
+ )
227
+ [INFO ][2026-05-10 12:23:57][root ][init_video_model ] MultiSeqWrapper(
228
+ (backbone): Frozen2DTargetWrapper(
229
+ (backbone): Eva(
230
+ (patch_embed): PatchEmbed(
231
+ (proj): Conv2d(3, 1024, kernel_size=(14, 14), stride=(14, 14))
232
+ (norm): Identity()
233
+ )
234
+ (pos_drop): Dropout(p=0.0, inplace=False)
235
+ (norm_pre): Identity()
236
+ (blocks): ModuleList(
237
+ (0-23): 24 x EvaBlock(
238
+ (norm1): RMSNorm((1024,), eps=1e-05, elementwise_affine=True)
239
+ (attn): EvaAttention(
240
+ (qkv): Linear(in_features=1024, out_features=3072, bias=False)
241
+ (q_norm): Identity()
242
+ (k_norm): Identity()
243
+ (attn_drop): Dropout(p=0.0, inplace=False)
244
+ (norm): Identity()
245
+ (proj): Linear(in_features=1024, out_features=1024, bias=True)
246
+ (proj_drop): Dropout(p=0.0, inplace=False)
247
+ )
248
+ (drop_path1): Identity()
249
+ (norm2): RMSNorm((1024,), eps=1e-05, elementwise_affine=True)
250
+ (mlp): Mlp(
251
+ (fc1): Linear(in_features=1024, out_features=4096, bias=True)
252
+ (act): GELU(approximate='none')
253
+ (drop1): Dropout(p=0.0, inplace=False)
254
+ (norm): Identity()
255
+ (fc2): Linear(in_features=4096, out_features=1024, bias=True)
256
+ (drop2): Dropout(p=0.0, inplace=False)
257
+ )
258
+ (drop_path2): Identity()
259
+ )
260
+ )
261
+ (norm): RMSNorm((1024,), eps=1e-05, elementwise_affine=True)
262
+ (fc_norm): Identity()
263
+ (head_drop): Dropout(p=0.0, inplace=False)
264
+ (head): Identity()
265
+ (rope): _CapiPatchRoPE()
266
+ )
267
+ )
268
+ )
269
+ [INFO ][2026-05-10 12:23:57][root ][init_video_model ] Encoder number of parameters: 302399488
270
+ [INFO ][2026-05-10 12:23:57][root ][init_video_model ] Predictor number of parameters: 11416192
271
+ [INFO ][2026-05-10 12:23:57][root ][init_video_model ] Target encoder number of parameters: 0
272
+ [INFO ][2026-05-10 12:24:08][root ][make_videodataset ] VideoDataset dataset created
273
+ [INFO ][2026-05-10 12:24:08][WeightedSampler ][__init__ ] Using DistributedWeightedSampler with rank 12 / 16
274
+ [INFO ][2026-05-10 12:24:08][root ][make_videodataset ] VideoDataset unsupervised data loader created
275
+ [INFO ][2026-05-10 12:24:08][app.vjepa.train ][main ] iterations per epoch/dataset length: 300/399
276
+ [INFO ][2026-05-10 12:24:08][app.vjepa.train ][main ] Wrapping models in DDP (rank 12)...
277
+ [INFO ][2026-05-10 12:24:08][app.vjepa.train ][main ] Initializing loader...
278
+ submitit WARNING (2026-05-10 16:19:58,222) - Bypassing signal SIGTERM
279
+ [WARNING ][2026-05-10 16:19:58][submitit ][bypass ] Bypassing signal SIGTERM
280
+ submitit WARNING (2026-05-10 16:19:58,948) - Bypassing signal SIGCONT
281
+ [WARNING ][2026-05-10 16:19:58][submitit ][bypass ] Bypassing signal SIGCONT
282
+ submitit ERROR (2026-05-10 16:19:58,949) - Submitted job triggered an exception
283
+ [ERROR ][2026-05-10 16:19:58][submitit ][process_job ] Submitted job triggered an exception
284
+ [ERROR ][2026-05-10 16:19:58][app.vjepa.train ][_save_crash ] Saved crash log to /scratch-shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/crash_rank12.log
vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/job_22614766/22614766_13_log.err ADDED
@@ -0,0 +1,180 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/timm/models/layers/__init__.py:49: FutureWarning: Importing from timm.models.layers is deprecated, please import via timm.layers
2
+ warnings.warn(f"Importing from {__name__} is deprecated, please import via timm.layers", FutureWarning)
3
+ [rank13]:[W510 12:19:58.813234569 ProcessGroupNCCL.cpp:5138] Guessing device ID based on global rank. This can cause a hang if rank to GPU mapping is heterogeneous. You can specify device_id in init_process_group()
4
+ /gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/vjepa/utils.py:796: FutureWarning: `torch.cuda.amp.GradScaler(args...)` is deprecated. Please use `torch.amp.GradScaler('cuda', args...)` instead.
5
+ scaler = torch.cuda.amp.GradScaler() if mixed_precision else None
6
+ /gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/vjepa/train.py:955: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
7
+ with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
8
+ [rank13]:[W510 16:19:49.920365474 TCPStore.cpp:125] [c10d] recvValue failed on SocketImpl(fd=3, addr=[gcn87.local.snellius.surf.nl]:50950, remote=[gcn80.local.snellius.surf.nl]:37129): Failed to recv, got 0 bytes. Connection was likely closed. Did the remote server shutdown or crash?
9
+ Exception raised from recvBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:682 (most recent call first):
10
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x147788ed1fdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
11
+ frame #1: <unknown function> + 0x6a3325d (0x1477cd06025d in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
12
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x273 (0x1477cd05e1f3 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
13
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x14778a140eec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
14
+ frame #4: <unknown function> + 0xed164 (0x1478c95aa164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
15
+ frame #5: <unknown function> + 0x8a19a (0x1478da01c19a in /lib64/libc.so.6)
16
+ frame #6: <unknown function> + 0x10f100 (0x1478da0a1100 in /lib64/libc.so.6)
17
+
18
+ [rank13]:[W510 16:19:49.926025477 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 13] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Failed to recv, got 0 bytes. Connection was likely closed. Did the remote server shutdown or crash?
19
+ [rank13]:[W510 16:19:50.926173311 TCPStore.cpp:106] [c10d] sendBytes failed on SocketImpl(fd=3, addr=[gcn87.local.snellius.surf.nl]:50950, remote=[gcn80.local.snellius.surf.nl]:37129): Broken pipe
20
+ Exception raised from sendBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:653 (most recent call first):
21
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x147788ed1fdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
22
+ frame #1: <unknown function> + 0x6a326d1 (0x1477cd05f6d1 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
23
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x24d (0x1477cd05e1cd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
24
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x14778a140eec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
25
+ frame #4: <unknown function> + 0xed164 (0x1478c95aa164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
26
+ frame #5: <unknown function> + 0x8a19a (0x1478da01c19a in /lib64/libc.so.6)
27
+ frame #6: <unknown function> + 0x10f100 (0x1478da0a1100 in /lib64/libc.so.6)
28
+
29
+ [rank13]:[W510 16:19:50.929575721 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 13] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Broken pipe
30
+ [rank13]:[W510 16:19:51.929706165 TCPStore.cpp:106] [c10d] sendBytes failed on SocketImpl(fd=3, addr=[gcn87.local.snellius.surf.nl]:50950, remote=[gcn80.local.snellius.surf.nl]:37129): Broken pipe
31
+ Exception raised from sendBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:653 (most recent call first):
32
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x147788ed1fdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
33
+ frame #1: <unknown function> + 0x6a326d1 (0x1477cd05f6d1 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
34
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x24d (0x1477cd05e1cd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
35
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x14778a140eec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
36
+ frame #4: <unknown function> + 0xed164 (0x1478c95aa164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
37
+ frame #5: <unknown function> + 0x8a19a (0x1478da01c19a in /lib64/libc.so.6)
38
+ frame #6: <unknown function> + 0x10f100 (0x1478da0a1100 in /lib64/libc.so.6)
39
+
40
+ [rank13]:[W510 16:19:51.932054678 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 13] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Broken pipe
41
+ [rank13]:[W510 16:19:52.932179413 TCPStore.cpp:106] [c10d] sendBytes failed on SocketImpl(fd=3, addr=[gcn87.local.snellius.surf.nl]:50950, remote=[gcn80.local.snellius.surf.nl]:37129): Broken pipe
42
+ Exception raised from sendBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:653 (most recent call first):
43
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x147788ed1fdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
44
+ frame #1: <unknown function> + 0x6a326d1 (0x1477cd05f6d1 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
45
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x24d (0x1477cd05e1cd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
46
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x14778a140eec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
47
+ frame #4: <unknown function> + 0xed164 (0x1478c95aa164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
48
+ frame #5: <unknown function> + 0x8a19a (0x1478da01c19a in /lib64/libc.so.6)
49
+ frame #6: <unknown function> + 0x10f100 (0x1478da0a1100 in /lib64/libc.so.6)
50
+
51
+ [rank13]:[W510 16:19:52.934418546 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 13] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Broken pipe
52
+ [rank13]:[W510 16:19:53.934561220 TCPStore.cpp:106] [c10d] sendBytes failed on SocketImpl(fd=3, addr=[gcn87.local.snellius.surf.nl]:50950, remote=[gcn80.local.snellius.surf.nl]:37129): Broken pipe
53
+ Exception raised from sendBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:653 (most recent call first):
54
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x147788ed1fdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
55
+ frame #1: <unknown function> + 0x6a326d1 (0x1477cd05f6d1 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
56
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x24d (0x1477cd05e1cd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
57
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x14778a140eec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
58
+ frame #4: <unknown function> + 0xed164 (0x1478c95aa164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
59
+ frame #5: <unknown function> + 0x8a19a (0x1478da01c19a in /lib64/libc.so.6)
60
+ frame #6: <unknown function> + 0x10f100 (0x1478da0a1100 in /lib64/libc.so.6)
61
+
62
+ [rank13]:[W510 16:19:53.936917383 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 13] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Broken pipe
63
+ [rank13]:[W510 16:19:54.937063907 TCPStore.cpp:106] [c10d] sendBytes failed on SocketImpl(fd=3, addr=[gcn87.local.snellius.surf.nl]:50950, remote=[gcn80.local.snellius.surf.nl]:37129): Broken pipe
64
+ Exception raised from sendBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:653 (most recent call first):
65
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x147788ed1fdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
66
+ frame #1: <unknown function> + 0x6a326d1 (0x1477cd05f6d1 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
67
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x24d (0x1477cd05e1cd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
68
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x14778a140eec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
69
+ frame #4: <unknown function> + 0xed164 (0x1478c95aa164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
70
+ frame #5: <unknown function> + 0x8a19a (0x1478da01c19a in /lib64/libc.so.6)
71
+ frame #6: <unknown function> + 0x10f100 (0x1478da0a1100 in /lib64/libc.so.6)
72
+
73
+ [rank13]:[W510 16:19:54.939398820 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 13] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Broken pipe
74
+ [rank13]:[W510 16:19:55.939534535 TCPStore.cpp:106] [c10d] sendBytes failed on SocketImpl(fd=3, addr=[gcn87.local.snellius.surf.nl]:50950, remote=[gcn80.local.snellius.surf.nl]:37129): Broken pipe
75
+ Exception raised from sendBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:653 (most recent call first):
76
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x147788ed1fdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
77
+ frame #1: <unknown function> + 0x6a326d1 (0x1477cd05f6d1 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
78
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x24d (0x1477cd05e1cd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
79
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x14778a140eec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
80
+ frame #4: <unknown function> + 0xed164 (0x1478c95aa164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
81
+ frame #5: <unknown function> + 0x8a19a (0x1478da01c19a in /lib64/libc.so.6)
82
+ frame #6: <unknown function> + 0x10f100 (0x1478da0a1100 in /lib64/libc.so.6)
83
+
84
+ [rank13]:[W510 16:19:55.941842188 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 13] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Broken pipe
85
+ [rank13]:[W510 16:19:56.941987682 TCPStore.cpp:106] [c10d] sendBytes failed on SocketImpl(fd=3, addr=[gcn87.local.snellius.surf.nl]:50950, remote=[gcn80.local.snellius.surf.nl]:37129): Broken pipe
86
+ Exception raised from sendBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:653 (most recent call first):
87
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x147788ed1fdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
88
+ frame #1: <unknown function> + 0x6a326d1 (0x1477cd05f6d1 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
89
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x24d (0x1477cd05e1cd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
90
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x14778a140eec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
91
+ frame #4: <unknown function> + 0xed164 (0x1478c95aa164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
92
+ frame #5: <unknown function> + 0x8a19a (0x1478da01c19a in /lib64/libc.so.6)
93
+ frame #6: <unknown function> + 0x10f100 (0x1478da0a1100 in /lib64/libc.so.6)
94
+
95
+ [rank13]:[W510 16:19:56.944320255 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 13] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Broken pipe
96
+ [rank13]:[W510 16:19:57.944470919 TCPStore.cpp:106] [c10d] sendBytes failed on SocketImpl(fd=3, addr=[gcn87.local.snellius.surf.nl]:50950, remote=[gcn80.local.snellius.surf.nl]:37129): Broken pipe
97
+ Exception raised from sendBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:653 (most recent call first):
98
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x147788ed1fdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
99
+ frame #1: <unknown function> + 0x6a326d1 (0x1477cd05f6d1 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
100
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x24d (0x1477cd05e1cd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
101
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x14778a140eec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
102
+ frame #4: <unknown function> + 0xed164 (0x1478c95aa164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
103
+ frame #5: <unknown function> + 0x8a19a (0x1478da01c19a in /lib64/libc.so.6)
104
+ frame #6: <unknown function> + 0x10f100 (0x1478da0a1100 in /lib64/libc.so.6)
105
+
106
+ [rank13]:[W510 16:19:57.946740642 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 13] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Broken pipe
107
+ [rank13]:[W510 16:19:58.946891777 TCPStore.cpp:106] [c10d] sendBytes failed on SocketImpl(fd=3, addr=[gcn87.local.snellius.surf.nl]:50950, remote=[gcn80.local.snellius.surf.nl]:37129): Broken pipe
108
+ Exception raised from sendBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:653 (most recent call first):
109
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x147788ed1fdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
110
+ frame #1: <unknown function> + 0x6a326d1 (0x1477cd05f6d1 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
111
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x24d (0x1477cd05e1cd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
112
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x14778a140eec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
113
+ frame #4: <unknown function> + 0xed164 (0x1478c95aa164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
114
+ frame #5: <unknown function> + 0x8a19a (0x1478da01c19a in /lib64/libc.so.6)
115
+ frame #6: <unknown function> + 0x10f100 (0x1478da0a1100 in /lib64/libc.so.6)
116
+
117
+ [rank13]:[W510 16:19:58.949185960 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 13] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Broken pipe
118
+ [rank13]:[W510 16:19:59.949324054 TCPStore.cpp:106] [c10d] sendBytes failed on SocketImpl(fd=3, addr=[gcn87.local.snellius.surf.nl]:50950, remote=[gcn80.local.snellius.surf.nl]:37129): Broken pipe
119
+ Exception raised from sendBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:653 (most recent call first):
120
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x147788ed1fdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
121
+ frame #1: <unknown function> + 0x6a326d1 (0x1477cd05f6d1 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
122
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x24d (0x1477cd05e1cd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
123
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x14778a140eec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
124
+ frame #4: <unknown function> + 0xed164 (0x1478c95aa164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
125
+ frame #5: <unknown function> + 0x8a19a (0x1478da01c19a in /lib64/libc.so.6)
126
+ frame #6: <unknown function> + 0x10f100 (0x1478da0a1100 in /lib64/libc.so.6)
127
+
128
+ [rank13]:[W510 16:19:59.951693287 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 13] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Broken pipe
129
+ submitit WARNING (2026-05-10 16:20:00,106) - Bypassing signal SIGTERM
130
+ submitit WARNING (2026-05-10 16:20:00,196) - Bypassing signal SIGCONT
131
+ submitit ERROR (2026-05-10 16:20:00,196) - Submitted job triggered an exception
132
+ [rank13]:[W510 16:20:00.951837531 TCPStore.cpp:106] [c10d] sendBytes failed on SocketImpl(fd=3, addr=[gcn87.local.snellius.surf.nl]:50950, remote=[gcn80.local.snellius.surf.nl]:37129): Broken pipe
133
+ Exception raised from sendBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:653 (most recent call first):
134
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x147788ed1fdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
135
+ frame #1: <unknown function> + 0x6a326d1 (0x1477cd05f6d1 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
136
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x24d (0x1477cd05e1cd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
137
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x14778a140eec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
138
+ frame #4: <unknown function> + 0xed164 (0x1478c95aa164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
139
+ frame #5: <unknown function> + 0x8a19a (0x1478da01c19a in /lib64/libc.so.6)
140
+ frame #6: <unknown function> + 0x10f100 (0x1478da0a1100 in /lib64/libc.so.6)
141
+
142
+ [rank13]:[W510 16:20:00.954199564 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 13] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Broken pipe
143
+ Traceback (most recent call last):
144
+ File "<frozen runpy>", line 198, in _run_module_as_main
145
+ File "<frozen runpy>", line 88, in _run_code
146
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/_submit.py", line 11, in <module>
147
+ submitit_main()
148
+ ~~~~~~~~~~~~~^^
149
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 76, in submitit_main
150
+ process_job(args.folder)
151
+ ~~~~~~~~~~~^^^^^^^^^^^^^
152
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 69, in process_job
153
+ raise error
154
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 55, in process_job
155
+ result = delayed.result()
156
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/utils.py", line 137, in result
157
+ self._result = self.function(*self.args, **self.kwargs)
158
+ ~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^
159
+ File "/gpfs/home2/dcanez/vd/app/main_distributed.py", line 102, in __call__
160
+ app_main(app, args=params, resume_preempt=resume_preempt)
161
+ ~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
162
+ File "/gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/scaffold.py", line 17, in main
163
+ return importlib.import_module(f"app.{app}.train").main(args=args, resume_preempt=resume_preempt)
164
+ ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
165
+ File "/gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/vjepa/train.py", line 1189, in main
166
+ assert not np.isnan(loss), "loss is nan"
167
+ ^^^^^^^^^^^^^^^^^^
168
+ AssertionError: loss is nan
169
+ [rank13]:[W510 16:20:01.632183756 ProcessGroupNCCL.cpp:1553] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())
170
+ [rank13]:[W510 16:20:01.954349608 TCPStore.cpp:106] [c10d] sendBytes failed on SocketImpl(fd=3, addr=[gcn87.local.snellius.surf.nl]:50950, remote=[gcn80.local.snellius.surf.nl]:37129): Broken pipe
171
+ Exception raised from sendBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:653 (most recent call first):
172
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x147788ed1fdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
173
+ frame #1: <unknown function> + 0x6a326d1 (0x1477cd05f6d1 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
174
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x24d (0x1477cd05e1cd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
175
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x14778a140eec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
176
+ frame #4: <unknown function> + 0xed164 (0x1478c95aa164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
177
+ frame #5: <unknown function> + 0x8a19a (0x1478da01c19a in /lib64/libc.so.6)
178
+ frame #6: <unknown function> + 0x10f100 (0x1478da0a1100 in /lib64/libc.so.6)
179
+
180
+ [rank13]:[W510 16:20:01.956608451 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 13] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Broken pipe
vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/job_22614766/22614766_13_log.out ADDED
@@ -0,0 +1,280 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ submitit INFO (2026-05-10 12:14:34,055) - Starting with JobEnvironment(job_id=22614766, hostname=gcn87.local.snellius.surf.nl, local_rank=1(4), node=3(4), global_rank=13(16))
2
+ submitit INFO (2026-05-10 12:14:34,056) - Loading pickle: /scratch-shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/job_22614766/22614766_submitted.pkl
3
+ INFO:root:loaded pretrain params...
4
+ { 'app': 'vjepa',
5
+ 'cpus_per_task': 16,
6
+ 'data': { 'batch_size': 64,
7
+ 'crop_size': 224,
8
+ 'dataset_fpcs': [16, 16],
9
+ 'dataset_type': 'VideoDataset',
10
+ 'datasets': [ '/scratch-shared/dcanez/data/kinetics/k400/train.csv',
11
+ '/scratch-shared/dcanez/data/ssv2/train.csv'],
12
+ 'datasets_weights': [0.65, 0.35],
13
+ 'fps': 4,
14
+ 'num_workers': 10,
15
+ 'patch_size': 14,
16
+ 'persistent_workers': True,
17
+ 'pin_mem': True,
18
+ 'stage': [ { 'dest': 'kinetics_240',
19
+ 'format': 'targz_parts',
20
+ 'src': '/scratch-shared/dcanez/data/kinetics/k400/tars_240/'},
21
+ { 'dest': 'ssv2',
22
+ 'format': 'multipart_tar',
23
+ 'src': '/scratch-nvme/ml-datasets/something-something-v2/'}],
24
+ 'tubelet_size': 1},
25
+ 'data_aug': { 'auto_augment': False,
26
+ 'motion_shift': False,
27
+ 'random_resize_aspect_ratio': [0.75, 1.35],
28
+ 'random_resize_scale': [0.3, 1.0],
29
+ 'reprob': 0.0},
30
+ 'folder': '/scratch-shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable',
31
+ 'loss': {'loss_exp': 1.0},
32
+ 'mask': [ { 'aspect_ratio': [0.75, 1.5],
33
+ 'full_complement': False,
34
+ 'max_keep': None,
35
+ 'max_temporal_keep': 1.0,
36
+ 'num_blocks': 8,
37
+ 'spatial_scale': [0.15, 0.15],
38
+ 'temporal_scale': [1.0, 1.0]},
39
+ { 'aspect_ratio': [0.75, 1.5],
40
+ 'full_complement': False,
41
+ 'max_keep': None,
42
+ 'max_temporal_keep': 1.0,
43
+ 'num_blocks': 2,
44
+ 'spatial_scale': [0.7, 0.7],
45
+ 'temporal_scale': [1.0, 1.0]}],
46
+ 'mem_per_gpu': '180G',
47
+ 'meta': { 'dtype': 'bfloat16',
48
+ 'knn_eval_epoch0': False,
49
+ 'knn_eval_freq': 5,
50
+ 'knn_eval_presets': [ { 'config': { 'batch_size': 64,
51
+ 'dataset_train': '/scratch-shared/mdorkenw/ucf101/ucfTrainTestlist/train.csv',
52
+ 'dataset_val': '/scratch-shared/mdorkenw/ucf101/ucfTrainTestlist/val.csv',
53
+ 'eval_videos_per_class': 25,
54
+ 'num_workers': 8,
55
+ 'pool_type': 'slot_temporal_concat',
56
+ 'train_videos_per_class': 100},
57
+ 'preset': 'ucf101'},
58
+ { 'config': { 'batch_size': 64,
59
+ 'dataset_train': '/scratch-shared/mdorkenw/20bn-something-something-v2/something-something-v2-annotations/train_coarse10.csv',
60
+ 'dataset_val': '/scratch-shared/mdorkenw/20bn-something-something-v2/something-something-v2-annotations/val_coarse10.csv',
61
+ 'eval_videos_per_class': 100,
62
+ 'linear_probe': True,
63
+ 'num_workers': 8,
64
+ 'pool_type': 'slot_temporal_concat',
65
+ 'train_videos_per_class': 500},
66
+ 'preset': 'ssv2_coarse10'}],
67
+ 'load_checkpoint': True,
68
+ 'read_checkpoint': None,
69
+ 'save_every_freq': 5,
70
+ 'seed': 239,
71
+ 'use_sdpa': True,
72
+ 'use_wandb': True,
73
+ 'wandb_project': 'vjepa_ujepaside'},
74
+ 'metrics': {'sigreg': {}, 'std': {}},
75
+ 'model': { 'model_name': 'ujepaside_large_patch14_capi_lvd1689m',
76
+ 'pred_depth': 6,
77
+ 'pred_embed_dim': 384,
78
+ 'pred_num_heads': 12,
79
+ 'predictor': 'v2_cross',
80
+ 'st_causal': False,
81
+ 'st_drop_path': 0.2,
82
+ 'st_flex_enable': False,
83
+ 'st_layer_scale_init': 1e-05,
84
+ 'st_num_slots': 16,
85
+ 'st_slots_causal_within_frame': False,
86
+ 'target_kind': 'frozen_2d',
87
+ 'target_type': 'vit_large_patch14_capi.lvd1689m',
88
+ 'temporal_spacing': 1.0,
89
+ 'uniform_power': True,
90
+ 'use_activation_checkpointing': True,
91
+ 'use_mask_tokens': True,
92
+ 'use_rope': True,
93
+ 'use_sdpa': True,
94
+ 'zero_init_mask_tokens': True},
95
+ 'nodes': 4,
96
+ 'optimization': { 'clip_grad': 3.0,
97
+ 'ema': [0.99925, 0.99925],
98
+ 'epochs': 100,
99
+ 'final_lr': 0.0001,
100
+ 'final_weight_decay': 0.04,
101
+ 'ipe': 300,
102
+ 'ipe_scale': 1.0,
103
+ 'lr': 0.0005,
104
+ 'start_lr': 0.0001,
105
+ 'warmup': 10,
106
+ 'weight_decay': 0.04},
107
+ 'tasks_per_node': 4}
108
+ INFO:root:Running pre-training of app: vjepa
109
+ [INFO ][2026-05-10 12:16:34][app.vjepa.train ][main ] which_dtype='bfloat16'
110
+ [INFO ][2026-05-10 12:16:34][app.vjepa.train ][main ] Disabling persistent_workers (incompatible with KNN eval)
111
+ [INFO ][2026-05-10 12:16:34][app.vjepa.train ][main ] NCCL_SOCKET_IFNAME=eno
112
+ [INFO ][2026-05-10 12:16:56][app.vjepa.train ][main ] Initialized (rank/world-size) 13/16, tasks_per_node=4
113
+ [INFO ][2026-05-10 12:16:56][root ][stage_datasets ] [local_rank 1/4] Staging kinetics_240 (targz_parts)
114
+ [INFO ][2026-05-10 12:16:56][root ][_stage_targz_parts ] [rank 1] Extracting 26/103 tar.gz parts to /scratch-node/dcanez.22614766/kinetics_240
115
+ [INFO ][2026-05-10 12:17:08][root ][_stage_targz_parts ] [local_rank 1] Extracted 2/26 parts
116
+ [INFO ][2026-05-10 12:17:22][root ][_stage_targz_parts ] [local_rank 1] Extracted 4/26 parts
117
+ [INFO ][2026-05-10 12:17:36][root ][_stage_targz_parts ] [local_rank 1] Extracted 6/26 parts
118
+ [INFO ][2026-05-10 12:17:50][root ][_stage_targz_parts ] [local_rank 1] Extracted 8/26 parts
119
+ [INFO ][2026-05-10 12:18:04][root ][_stage_targz_parts ] [local_rank 1] Extracted 10/26 parts
120
+ [INFO ][2026-05-10 12:18:19][root ][_stage_targz_parts ] [local_rank 1] Extracted 12/26 parts
121
+ [INFO ][2026-05-10 12:18:32][root ][_stage_targz_parts ] [local_rank 1] Extracted 14/26 parts
122
+ [INFO ][2026-05-10 12:18:46][root ][_stage_targz_parts ] [local_rank 1] Extracted 16/26 parts
123
+ [INFO ][2026-05-10 12:19:00][root ][_stage_targz_parts ] [local_rank 1] Extracted 18/26 parts
124
+ [INFO ][2026-05-10 12:19:14][root ][_stage_targz_parts ] [local_rank 1] Extracted 20/26 parts
125
+ [INFO ][2026-05-10 12:19:29][root ][_stage_targz_parts ] [local_rank 1] Extracted 22/26 parts
126
+ [INFO ][2026-05-10 12:19:43][root ][_stage_targz_parts ] [local_rank 1] Extracted 24/26 parts
127
+ [INFO ][2026-05-10 12:19:58][root ][_stage_targz_parts ] [local_rank 1] Extracted 26/26 parts
128
+ [INFO ][2026-05-10 12:19:58][root ][stage_datasets ] [local_rank 1/4] Staging ssv2 (multipart_tar)
129
+ [INFO ][2026-05-10 12:21:03][app.vjepa.train ][main ] Staged datasets: ['/scratch-node/dcanez.22614766/kinetics_240/train.csv', '/scratch-node/dcanez.22614766/ssv2/train.csv']
130
+ [INFO ][2026-05-10 12:23:31][experiments.stmodels.vision_transformers_v3][vit_large_patch14_capi_lvd1689m] Loading pretrained weights for vit_large_patch14_capi_lvd1689m from https://dl.fbaipublicfiles.com/capi/capi_vitl14_lvd.pth
131
+ [INFO ][2026-05-10 12:23:57][root ][_build_frozen_2d_target ] Loaded frozen_2d target from timm: vit_large_patch14_capi.lvd1689m
132
+ [INFO ][2026-05-10 12:23:57][root ][init_video_model ] ViTMultiSeqWrapper(
133
+ (backbone): UJEPAside(
134
+ (patch_embed): PatchEmbed3D(
135
+ (proj): Conv3d(3, 1024, kernel_size=(1, 14, 14), stride=(1, 14, 14))
136
+ )
137
+ (rope): CAPI2DRoPE()
138
+ (blocks): ModuleList(
139
+ (0-23): 24 x Block(
140
+ (norm1): RMSNorm((1024,), eps=1e-05, elementwise_affine=True)
141
+ (rope_impl): CAPI2DRoPE()
142
+ (norm2): RMSNorm((1024,), eps=1e-05, elementwise_affine=True)
143
+ (drop_path1): Identity()
144
+ (drop_path2): Identity()
145
+ (attn): Attention(
146
+ (qkv): Linear(in_features=1024, out_features=3072, bias=False)
147
+ (attn_drop): Dropout(p=0.0, inplace=False)
148
+ (proj): Linear(in_features=1024, out_features=1024, bias=False)
149
+ (proj_drop): Dropout(p=0.0, inplace=False)
150
+ (rope_impl): CAPI2DRoPE()
151
+ )
152
+ (mlp): MLP(
153
+ (fc1): Linear(in_features=1024, out_features=4096, bias=False)
154
+ (act): GELU(approximate='none')
155
+ (fc2): Linear(in_features=4096, out_features=1024, bias=False)
156
+ (drop): Dropout(p=0.0, inplace=False)
157
+ )
158
+ )
159
+ )
160
+ (norm): RMSNorm((1024,), eps=1e-05, elementwise_affine=False)
161
+ (st_blocks): ModuleList(
162
+ (0-23): 24 x Block(
163
+ (residual1): EfficientResidual(
164
+ (norm): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
165
+ (fn): Attention(
166
+ (q_proj): Linear(in_features=1024, out_features=1024, bias=False)
167
+ (k_proj): Linear(in_features=1024, out_features=1024, bias=False)
168
+ (v_proj): Linear(in_features=1024, out_features=1024, bias=False)
169
+ (proj): Linear(in_features=1024, out_features=1024, bias=False)
170
+ (rope_impl): CAPI3DRoPE()
171
+ )
172
+ )
173
+ (residual2): EfficientResidual(
174
+ (norm): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
175
+ (fn): MLP(
176
+ (fc1): Linear(in_features=1024, out_features=4096, bias=False)
177
+ (act): GELU(approximate='none')
178
+ (fc2): Linear(in_features=4096, out_features=1024, bias=False)
179
+ (drop): Dropout(p=0.0, inplace=False)
180
+ )
181
+ )
182
+ )
183
+ )
184
+ (st_rope): CAPI3DRoPE()
185
+ )
186
+ )
187
+ [INFO ][2026-05-10 12:23:57][root ][init_video_model ] UJEPAsidePredictorMultiSeqWrapper(
188
+ (backbone): PredictorV2(
189
+ (predictor_embed): Linear(in_features=1024, out_features=384, bias=True)
190
+ (mask_tokens): ParameterList(
191
+ (0): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
192
+ (1): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
193
+ (2): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
194
+ (3): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
195
+ )
196
+ (predictor_blocks): ModuleList(
197
+ (0-5): 6 x Block(
198
+ (residual1): EfficientResidual(
199
+ (norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
200
+ (fn): Attention(
201
+ (q_proj): Linear(in_features=384, out_features=384, bias=False)
202
+ (k_proj): Linear(in_features=384, out_features=384, bias=False)
203
+ (v_proj): Linear(in_features=384, out_features=384, bias=False)
204
+ (proj): Linear(in_features=384, out_features=384, bias=False)
205
+ (rope): Rope()
206
+ )
207
+ )
208
+ (residual2): EfficientResidual(
209
+ (norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
210
+ (fn): MLP(
211
+ (fc1): Linear(in_features=384, out_features=1536, bias=False)
212
+ (act): GELU(approximate='none')
213
+ (fc2): Linear(in_features=1536, out_features=384, bias=False)
214
+ (drop): Dropout(p=0.0, inplace=False)
215
+ )
216
+ )
217
+ )
218
+ )
219
+ (predictor_norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
220
+ (predictor_proj): Linear(in_features=384, out_features=1024, bias=True)
221
+ )
222
+ )
223
+ [INFO ][2026-05-10 12:23:57][root ][init_video_model ] MultiSeqWrapper(
224
+ (backbone): Frozen2DTargetWrapper(
225
+ (backbone): Eva(
226
+ (patch_embed): PatchEmbed(
227
+ (proj): Conv2d(3, 1024, kernel_size=(14, 14), stride=(14, 14))
228
+ (norm): Identity()
229
+ )
230
+ (pos_drop): Dropout(p=0.0, inplace=False)
231
+ (norm_pre): Identity()
232
+ (blocks): ModuleList(
233
+ (0-23): 24 x EvaBlock(
234
+ (norm1): RMSNorm((1024,), eps=1e-05, elementwise_affine=True)
235
+ (attn): EvaAttention(
236
+ (qkv): Linear(in_features=1024, out_features=3072, bias=False)
237
+ (q_norm): Identity()
238
+ (k_norm): Identity()
239
+ (attn_drop): Dropout(p=0.0, inplace=False)
240
+ (norm): Identity()
241
+ (proj): Linear(in_features=1024, out_features=1024, bias=True)
242
+ (proj_drop): Dropout(p=0.0, inplace=False)
243
+ )
244
+ (drop_path1): Identity()
245
+ (norm2): RMSNorm((1024,), eps=1e-05, elementwise_affine=True)
246
+ (mlp): Mlp(
247
+ (fc1): Linear(in_features=1024, out_features=4096, bias=True)
248
+ (act): GELU(approximate='none')
249
+ (drop1): Dropout(p=0.0, inplace=False)
250
+ (norm): Identity()
251
+ (fc2): Linear(in_features=4096, out_features=1024, bias=True)
252
+ (drop2): Dropout(p=0.0, inplace=False)
253
+ )
254
+ (drop_path2): Identity()
255
+ )
256
+ )
257
+ (norm): RMSNorm((1024,), eps=1e-05, elementwise_affine=True)
258
+ (fc_norm): Identity()
259
+ (head_drop): Dropout(p=0.0, inplace=False)
260
+ (head): Identity()
261
+ (rope): _CapiPatchRoPE()
262
+ )
263
+ )
264
+ )
265
+ [INFO ][2026-05-10 12:23:57][root ][init_video_model ] Encoder number of parameters: 302399488
266
+ [INFO ][2026-05-10 12:23:57][root ][init_video_model ] Predictor number of parameters: 11416192
267
+ [INFO ][2026-05-10 12:23:57][root ][init_video_model ] Target encoder number of parameters: 0
268
+ [INFO ][2026-05-10 12:24:08][root ][make_videodataset ] VideoDataset dataset created
269
+ [INFO ][2026-05-10 12:24:08][WeightedSampler ][__init__ ] Using DistributedWeightedSampler with rank 13 / 16
270
+ [INFO ][2026-05-10 12:24:08][root ][make_videodataset ] VideoDataset unsupervised data loader created
271
+ [INFO ][2026-05-10 12:24:08][app.vjepa.train ][main ] iterations per epoch/dataset length: 300/399
272
+ [INFO ][2026-05-10 12:24:08][app.vjepa.train ][main ] Wrapping models in DDP (rank 13)...
273
+ [INFO ][2026-05-10 12:24:08][app.vjepa.train ][main ] Initializing loader...
274
+ submitit WARNING (2026-05-10 16:20:00,106) - Bypassing signal SIGTERM
275
+ [WARNING ][2026-05-10 16:20:00][submitit ][bypass ] Bypassing signal SIGTERM
276
+ submitit WARNING (2026-05-10 16:20:00,196) - Bypassing signal SIGCONT
277
+ [WARNING ][2026-05-10 16:20:00][submitit ][bypass ] Bypassing signal SIGCONT
278
+ submitit ERROR (2026-05-10 16:20:00,196) - Submitted job triggered an exception
279
+ [ERROR ][2026-05-10 16:20:00][submitit ][process_job ] Submitted job triggered an exception
280
+ [ERROR ][2026-05-10 16:20:00][app.vjepa.train ][_save_crash ] Saved crash log to /scratch-shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/crash_rank13.log
vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/job_22614766/22614766_14_log.err ADDED
@@ -0,0 +1,79 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/timm/models/layers/__init__.py:49: FutureWarning: Importing from timm.models.layers is deprecated, please import via timm.layers
2
+ warnings.warn(f"Importing from {__name__} is deprecated, please import via timm.layers", FutureWarning)
3
+ [rank14]:[W510 12:20:03.054704675 ProcessGroupNCCL.cpp:5138] Guessing device ID based on global rank. This can cause a hang if rank to GPU mapping is heterogeneous. You can specify device_id in init_process_group()
4
+ /gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/vjepa/utils.py:796: FutureWarning: `torch.cuda.amp.GradScaler(args...)` is deprecated. Please use `torch.amp.GradScaler('cuda', args...)` instead.
5
+ scaler = torch.cuda.amp.GradScaler() if mixed_precision else None
6
+ /gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/vjepa/train.py:955: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
7
+ with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
8
+ [rank14]:[W510 16:19:49.920375084 TCPStore.cpp:125] [c10d] recvValue failed on SocketImpl(fd=3, addr=[gcn87.local.snellius.surf.nl]:34352, remote=[gcn80.local.snellius.surf.nl]:37129): Failed to recv, got 0 bytes. Connection was likely closed. Did the remote server shutdown or crash?
9
+ Exception raised from recvBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:682 (most recent call first):
10
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x1496358b8fdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
11
+ frame #1: <unknown function> + 0x6a3325d (0x149679a4725d in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
12
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x273 (0x149679a451f3 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
13
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x149636b27eec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
14
+ frame #4: <unknown function> + 0xed164 (0x149775f91164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
15
+ frame #5: <unknown function> + 0x8a19a (0x149786a0319a in /lib64/libc.so.6)
16
+ frame #6: <unknown function> + 0x10f100 (0x149786a88100 in /lib64/libc.so.6)
17
+
18
+ [rank14]:[W510 16:19:49.926028947 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 14] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Failed to recv, got 0 bytes. Connection was likely closed. Did the remote server shutdown or crash?
19
+ submitit ERROR (2026-05-10 16:19:50,186) - Submitted job triggered an exception
20
+ [rank14]:[W510 16:19:50.926154361 TCPStore.cpp:106] [c10d] sendBytes failed on SocketImpl(fd=3, addr=[gcn87.local.snellius.surf.nl]:34352, remote=[gcn80.local.snellius.surf.nl]:37129): Broken pipe
21
+ Exception raised from sendBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:653 (most recent call first):
22
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x1496358b8fdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
23
+ frame #1: <unknown function> + 0x6a326d1 (0x149679a466d1 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
24
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x24d (0x149679a451cd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
25
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x149636b27eec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
26
+ frame #4: <unknown function> + 0xed164 (0x149775f91164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
27
+ frame #5: <unknown function> + 0x8a19a (0x149786a0319a in /lib64/libc.so.6)
28
+ frame #6: <unknown function> + 0x10f100 (0x149786a88100 in /lib64/libc.so.6)
29
+
30
+ [rank14]:[W510 16:19:50.929831610 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 14] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Broken pipe
31
+ Traceback (most recent call last):
32
+ File "<frozen runpy>", line 198, in _run_module_as_main
33
+ File "<frozen runpy>", line 88, in _run_code
34
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/_submit.py", line 11, in <module>
35
+ submitit_main()
36
+ ~~~~~~~~~~~~~^^
37
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 76, in submitit_main
38
+ process_job(args.folder)
39
+ ~~~~~~~~~~~^^^^^^^^^^^^^
40
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 69, in process_job
41
+ raise error
42
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 55, in process_job
43
+ result = delayed.result()
44
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/utils.py", line 137, in result
45
+ self._result = self.function(*self.args, **self.kwargs)
46
+ ~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^
47
+ File "/gpfs/home2/dcanez/vd/app/main_distributed.py", line 102, in __call__
48
+ app_main(app, args=params, resume_preempt=resume_preempt)
49
+ ~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
50
+ File "/gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/scaffold.py", line 17, in main
51
+ return importlib.import_module(f"app.{app}.train").main(args=args, resume_preempt=resume_preempt)
52
+ ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
53
+ File "/gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/vjepa/train.py", line 1189, in main
54
+ assert not np.isnan(loss), "loss is nan"
55
+ ^^^^^^^^^^^^^^^^^^
56
+ AssertionError: loss is nan
57
+ [rank14]:[W510 16:19:51.929981635 TCPStore.cpp:106] [c10d] sendBytes failed on SocketImpl(fd=3, addr=[gcn87.local.snellius.surf.nl]:34352, remote=[gcn80.local.snellius.surf.nl]:37129): Broken pipe
58
+ Exception raised from sendBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:653 (most recent call first):
59
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x1496358b8fdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
60
+ frame #1: <unknown function> + 0x6a326d1 (0x149679a466d1 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
61
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x24d (0x149679a451cd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
62
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x149636b27eec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
63
+ frame #4: <unknown function> + 0xed164 (0x149775f91164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
64
+ frame #5: <unknown function> + 0x8a19a (0x149786a0319a in /lib64/libc.so.6)
65
+ frame #6: <unknown function> + 0x10f100 (0x149786a88100 in /lib64/libc.so.6)
66
+
67
+ [rank14]:[W510 16:19:51.932269128 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 14] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Broken pipe
68
+ [rank14]:[W510 16:19:51.445846354 ProcessGroupNCCL.cpp:1553] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())
69
+ [rank14]:[W510 16:19:52.932408362 TCPStore.cpp:106] [c10d] sendBytes failed on SocketImpl(fd=3, addr=[gcn87.local.snellius.surf.nl]:34352, remote=[gcn80.local.snellius.surf.nl]:37129): Broken pipe
70
+ Exception raised from sendBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:653 (most recent call first):
71
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x1496358b8fdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
72
+ frame #1: <unknown function> + 0x6a326d1 (0x149679a466d1 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
73
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x24d (0x149679a451cd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
74
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x149636b27eec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
75
+ frame #4: <unknown function> + 0xed164 (0x149775f91164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
76
+ frame #5: <unknown function> + 0x8a19a (0x149786a0319a in /lib64/libc.so.6)
77
+ frame #6: <unknown function> + 0x10f100 (0x149786a88100 in /lib64/libc.so.6)
78
+
79
+ [rank14]:[W510 16:19:52.934668395 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 14] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Broken pipe
vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/job_22614766/22614766_14_log.out ADDED
@@ -0,0 +1,276 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ submitit INFO (2026-05-10 12:14:34,056) - Starting with JobEnvironment(job_id=22614766, hostname=gcn87.local.snellius.surf.nl, local_rank=2(4), node=3(4), global_rank=14(16))
2
+ submitit INFO (2026-05-10 12:14:34,056) - Loading pickle: /scratch-shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/job_22614766/22614766_submitted.pkl
3
+ INFO:root:loaded pretrain params...
4
+ { 'app': 'vjepa',
5
+ 'cpus_per_task': 16,
6
+ 'data': { 'batch_size': 64,
7
+ 'crop_size': 224,
8
+ 'dataset_fpcs': [16, 16],
9
+ 'dataset_type': 'VideoDataset',
10
+ 'datasets': [ '/scratch-shared/dcanez/data/kinetics/k400/train.csv',
11
+ '/scratch-shared/dcanez/data/ssv2/train.csv'],
12
+ 'datasets_weights': [0.65, 0.35],
13
+ 'fps': 4,
14
+ 'num_workers': 10,
15
+ 'patch_size': 14,
16
+ 'persistent_workers': True,
17
+ 'pin_mem': True,
18
+ 'stage': [ { 'dest': 'kinetics_240',
19
+ 'format': 'targz_parts',
20
+ 'src': '/scratch-shared/dcanez/data/kinetics/k400/tars_240/'},
21
+ { 'dest': 'ssv2',
22
+ 'format': 'multipart_tar',
23
+ 'src': '/scratch-nvme/ml-datasets/something-something-v2/'}],
24
+ 'tubelet_size': 1},
25
+ 'data_aug': { 'auto_augment': False,
26
+ 'motion_shift': False,
27
+ 'random_resize_aspect_ratio': [0.75, 1.35],
28
+ 'random_resize_scale': [0.3, 1.0],
29
+ 'reprob': 0.0},
30
+ 'folder': '/scratch-shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable',
31
+ 'loss': {'loss_exp': 1.0},
32
+ 'mask': [ { 'aspect_ratio': [0.75, 1.5],
33
+ 'full_complement': False,
34
+ 'max_keep': None,
35
+ 'max_temporal_keep': 1.0,
36
+ 'num_blocks': 8,
37
+ 'spatial_scale': [0.15, 0.15],
38
+ 'temporal_scale': [1.0, 1.0]},
39
+ { 'aspect_ratio': [0.75, 1.5],
40
+ 'full_complement': False,
41
+ 'max_keep': None,
42
+ 'max_temporal_keep': 1.0,
43
+ 'num_blocks': 2,
44
+ 'spatial_scale': [0.7, 0.7],
45
+ 'temporal_scale': [1.0, 1.0]}],
46
+ 'mem_per_gpu': '180G',
47
+ 'meta': { 'dtype': 'bfloat16',
48
+ 'knn_eval_epoch0': False,
49
+ 'knn_eval_freq': 5,
50
+ 'knn_eval_presets': [ { 'config': { 'batch_size': 64,
51
+ 'dataset_train': '/scratch-shared/mdorkenw/ucf101/ucfTrainTestlist/train.csv',
52
+ 'dataset_val': '/scratch-shared/mdorkenw/ucf101/ucfTrainTestlist/val.csv',
53
+ 'eval_videos_per_class': 25,
54
+ 'num_workers': 8,
55
+ 'pool_type': 'slot_temporal_concat',
56
+ 'train_videos_per_class': 100},
57
+ 'preset': 'ucf101'},
58
+ { 'config': { 'batch_size': 64,
59
+ 'dataset_train': '/scratch-shared/mdorkenw/20bn-something-something-v2/something-something-v2-annotations/train_coarse10.csv',
60
+ 'dataset_val': '/scratch-shared/mdorkenw/20bn-something-something-v2/something-something-v2-annotations/val_coarse10.csv',
61
+ 'eval_videos_per_class': 100,
62
+ 'linear_probe': True,
63
+ 'num_workers': 8,
64
+ 'pool_type': 'slot_temporal_concat',
65
+ 'train_videos_per_class': 500},
66
+ 'preset': 'ssv2_coarse10'}],
67
+ 'load_checkpoint': True,
68
+ 'read_checkpoint': None,
69
+ 'save_every_freq': 5,
70
+ 'seed': 239,
71
+ 'use_sdpa': True,
72
+ 'use_wandb': True,
73
+ 'wandb_project': 'vjepa_ujepaside'},
74
+ 'metrics': {'sigreg': {}, 'std': {}},
75
+ 'model': { 'model_name': 'ujepaside_large_patch14_capi_lvd1689m',
76
+ 'pred_depth': 6,
77
+ 'pred_embed_dim': 384,
78
+ 'pred_num_heads': 12,
79
+ 'predictor': 'v2_cross',
80
+ 'st_causal': False,
81
+ 'st_drop_path': 0.2,
82
+ 'st_flex_enable': False,
83
+ 'st_layer_scale_init': 1e-05,
84
+ 'st_num_slots': 16,
85
+ 'st_slots_causal_within_frame': False,
86
+ 'target_kind': 'frozen_2d',
87
+ 'target_type': 'vit_large_patch14_capi.lvd1689m',
88
+ 'temporal_spacing': 1.0,
89
+ 'uniform_power': True,
90
+ 'use_activation_checkpointing': True,
91
+ 'use_mask_tokens': True,
92
+ 'use_rope': True,
93
+ 'use_sdpa': True,
94
+ 'zero_init_mask_tokens': True},
95
+ 'nodes': 4,
96
+ 'optimization': { 'clip_grad': 3.0,
97
+ 'ema': [0.99925, 0.99925],
98
+ 'epochs': 100,
99
+ 'final_lr': 0.0001,
100
+ 'final_weight_decay': 0.04,
101
+ 'ipe': 300,
102
+ 'ipe_scale': 1.0,
103
+ 'lr': 0.0005,
104
+ 'start_lr': 0.0001,
105
+ 'warmup': 10,
106
+ 'weight_decay': 0.04},
107
+ 'tasks_per_node': 4}
108
+ INFO:root:Running pre-training of app: vjepa
109
+ [INFO ][2026-05-10 12:16:33][app.vjepa.train ][main ] which_dtype='bfloat16'
110
+ [INFO ][2026-05-10 12:16:33][app.vjepa.train ][main ] Disabling persistent_workers (incompatible with KNN eval)
111
+ [INFO ][2026-05-10 12:16:33][app.vjepa.train ][main ] NCCL_SOCKET_IFNAME=eno
112
+ [INFO ][2026-05-10 12:17:00][app.vjepa.train ][main ] Initialized (rank/world-size) 14/16, tasks_per_node=4
113
+ [INFO ][2026-05-10 12:17:00][root ][stage_datasets ] [local_rank 2/4] Staging kinetics_240 (targz_parts)
114
+ [INFO ][2026-05-10 12:17:00][root ][_stage_targz_parts ] [rank 2] Extracting 26/103 tar.gz parts to /scratch-node/dcanez.22614766/kinetics_240
115
+ [INFO ][2026-05-10 12:17:14][root ][_stage_targz_parts ] [local_rank 2] Extracted 2/26 parts
116
+ [INFO ][2026-05-10 12:17:27][root ][_stage_targz_parts ] [local_rank 2] Extracted 4/26 parts
117
+ [INFO ][2026-05-10 12:17:42][root ][_stage_targz_parts ] [local_rank 2] Extracted 6/26 parts
118
+ [INFO ][2026-05-10 12:17:56][root ][_stage_targz_parts ] [local_rank 2] Extracted 8/26 parts
119
+ [INFO ][2026-05-10 12:18:10][root ][_stage_targz_parts ] [local_rank 2] Extracted 10/26 parts
120
+ [INFO ][2026-05-10 12:18:24][root ][_stage_targz_parts ] [local_rank 2] Extracted 12/26 parts
121
+ [INFO ][2026-05-10 12:18:38][root ][_stage_targz_parts ] [local_rank 2] Extracted 14/26 parts
122
+ [INFO ][2026-05-10 12:18:52][root ][_stage_targz_parts ] [local_rank 2] Extracted 16/26 parts
123
+ [INFO ][2026-05-10 12:19:06][root ][_stage_targz_parts ] [local_rank 2] Extracted 18/26 parts
124
+ [INFO ][2026-05-10 12:19:20][root ][_stage_targz_parts ] [local_rank 2] Extracted 20/26 parts
125
+ [INFO ][2026-05-10 12:19:35][root ][_stage_targz_parts ] [local_rank 2] Extracted 22/26 parts
126
+ [INFO ][2026-05-10 12:19:49][root ][_stage_targz_parts ] [local_rank 2] Extracted 24/26 parts
127
+ [INFO ][2026-05-10 12:20:03][root ][_stage_targz_parts ] [local_rank 2] Extracted 26/26 parts
128
+ [INFO ][2026-05-10 12:20:03][root ][stage_datasets ] [local_rank 2/4] Staging ssv2 (multipart_tar)
129
+ [INFO ][2026-05-10 12:21:03][app.vjepa.train ][main ] Staged datasets: ['/scratch-node/dcanez.22614766/kinetics_240/train.csv', '/scratch-node/dcanez.22614766/ssv2/train.csv']
130
+ [INFO ][2026-05-10 12:23:31][experiments.stmodels.vision_transformers_v3][vit_large_patch14_capi_lvd1689m] Loading pretrained weights for vit_large_patch14_capi_lvd1689m from https://dl.fbaipublicfiles.com/capi/capi_vitl14_lvd.pth
131
+ [INFO ][2026-05-10 12:23:57][root ][_build_frozen_2d_target ] Loaded frozen_2d target from timm: vit_large_patch14_capi.lvd1689m
132
+ [INFO ][2026-05-10 12:23:57][root ][init_video_model ] ViTMultiSeqWrapper(
133
+ (backbone): UJEPAside(
134
+ (patch_embed): PatchEmbed3D(
135
+ (proj): Conv3d(3, 1024, kernel_size=(1, 14, 14), stride=(1, 14, 14))
136
+ )
137
+ (rope): CAPI2DRoPE()
138
+ (blocks): ModuleList(
139
+ (0-23): 24 x Block(
140
+ (norm1): RMSNorm((1024,), eps=1e-05, elementwise_affine=True)
141
+ (rope_impl): CAPI2DRoPE()
142
+ (norm2): RMSNorm((1024,), eps=1e-05, elementwise_affine=True)
143
+ (drop_path1): Identity()
144
+ (drop_path2): Identity()
145
+ (attn): Attention(
146
+ (qkv): Linear(in_features=1024, out_features=3072, bias=False)
147
+ (attn_drop): Dropout(p=0.0, inplace=False)
148
+ (proj): Linear(in_features=1024, out_features=1024, bias=False)
149
+ (proj_drop): Dropout(p=0.0, inplace=False)
150
+ (rope_impl): CAPI2DRoPE()
151
+ )
152
+ (mlp): MLP(
153
+ (fc1): Linear(in_features=1024, out_features=4096, bias=False)
154
+ (act): GELU(approximate='none')
155
+ (fc2): Linear(in_features=4096, out_features=1024, bias=False)
156
+ (drop): Dropout(p=0.0, inplace=False)
157
+ )
158
+ )
159
+ )
160
+ (norm): RMSNorm((1024,), eps=1e-05, elementwise_affine=False)
161
+ (st_blocks): ModuleList(
162
+ (0-23): 24 x Block(
163
+ (residual1): EfficientResidual(
164
+ (norm): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
165
+ (fn): Attention(
166
+ (q_proj): Linear(in_features=1024, out_features=1024, bias=False)
167
+ (k_proj): Linear(in_features=1024, out_features=1024, bias=False)
168
+ (v_proj): Linear(in_features=1024, out_features=1024, bias=False)
169
+ (proj): Linear(in_features=1024, out_features=1024, bias=False)
170
+ (rope_impl): CAPI3DRoPE()
171
+ )
172
+ )
173
+ (residual2): EfficientResidual(
174
+ (norm): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
175
+ (fn): MLP(
176
+ (fc1): Linear(in_features=1024, out_features=4096, bias=False)
177
+ (act): GELU(approximate='none')
178
+ (fc2): Linear(in_features=4096, out_features=1024, bias=False)
179
+ (drop): Dropout(p=0.0, inplace=False)
180
+ )
181
+ )
182
+ )
183
+ )
184
+ (st_rope): CAPI3DRoPE()
185
+ )
186
+ )
187
+ [INFO ][2026-05-10 12:23:57][root ][init_video_model ] UJEPAsidePredictorMultiSeqWrapper(
188
+ (backbone): PredictorV2(
189
+ (predictor_embed): Linear(in_features=1024, out_features=384, bias=True)
190
+ (mask_tokens): ParameterList(
191
+ (0): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
192
+ (1): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
193
+ (2): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
194
+ (3): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
195
+ )
196
+ (predictor_blocks): ModuleList(
197
+ (0-5): 6 x Block(
198
+ (residual1): EfficientResidual(
199
+ (norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
200
+ (fn): Attention(
201
+ (q_proj): Linear(in_features=384, out_features=384, bias=False)
202
+ (k_proj): Linear(in_features=384, out_features=384, bias=False)
203
+ (v_proj): Linear(in_features=384, out_features=384, bias=False)
204
+ (proj): Linear(in_features=384, out_features=384, bias=False)
205
+ (rope): Rope()
206
+ )
207
+ )
208
+ (residual2): EfficientResidual(
209
+ (norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
210
+ (fn): MLP(
211
+ (fc1): Linear(in_features=384, out_features=1536, bias=False)
212
+ (act): GELU(approximate='none')
213
+ (fc2): Linear(in_features=1536, out_features=384, bias=False)
214
+ (drop): Dropout(p=0.0, inplace=False)
215
+ )
216
+ )
217
+ )
218
+ )
219
+ (predictor_norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
220
+ (predictor_proj): Linear(in_features=384, out_features=1024, bias=True)
221
+ )
222
+ )
223
+ [INFO ][2026-05-10 12:23:57][root ][init_video_model ] MultiSeqWrapper(
224
+ (backbone): Frozen2DTargetWrapper(
225
+ (backbone): Eva(
226
+ (patch_embed): PatchEmbed(
227
+ (proj): Conv2d(3, 1024, kernel_size=(14, 14), stride=(14, 14))
228
+ (norm): Identity()
229
+ )
230
+ (pos_drop): Dropout(p=0.0, inplace=False)
231
+ (norm_pre): Identity()
232
+ (blocks): ModuleList(
233
+ (0-23): 24 x EvaBlock(
234
+ (norm1): RMSNorm((1024,), eps=1e-05, elementwise_affine=True)
235
+ (attn): EvaAttention(
236
+ (qkv): Linear(in_features=1024, out_features=3072, bias=False)
237
+ (q_norm): Identity()
238
+ (k_norm): Identity()
239
+ (attn_drop): Dropout(p=0.0, inplace=False)
240
+ (norm): Identity()
241
+ (proj): Linear(in_features=1024, out_features=1024, bias=True)
242
+ (proj_drop): Dropout(p=0.0, inplace=False)
243
+ )
244
+ (drop_path1): Identity()
245
+ (norm2): RMSNorm((1024,), eps=1e-05, elementwise_affine=True)
246
+ (mlp): Mlp(
247
+ (fc1): Linear(in_features=1024, out_features=4096, bias=True)
248
+ (act): GELU(approximate='none')
249
+ (drop1): Dropout(p=0.0, inplace=False)
250
+ (norm): Identity()
251
+ (fc2): Linear(in_features=4096, out_features=1024, bias=True)
252
+ (drop2): Dropout(p=0.0, inplace=False)
253
+ )
254
+ (drop_path2): Identity()
255
+ )
256
+ )
257
+ (norm): RMSNorm((1024,), eps=1e-05, elementwise_affine=True)
258
+ (fc_norm): Identity()
259
+ (head_drop): Dropout(p=0.0, inplace=False)
260
+ (head): Identity()
261
+ (rope): _CapiPatchRoPE()
262
+ )
263
+ )
264
+ )
265
+ [INFO ][2026-05-10 12:23:57][root ][init_video_model ] Encoder number of parameters: 302399488
266
+ [INFO ][2026-05-10 12:23:57][root ][init_video_model ] Predictor number of parameters: 11416192
267
+ [INFO ][2026-05-10 12:23:57][root ][init_video_model ] Target encoder number of parameters: 0
268
+ [INFO ][2026-05-10 12:24:08][root ][make_videodataset ] VideoDataset dataset created
269
+ [INFO ][2026-05-10 12:24:08][WeightedSampler ][__init__ ] Using DistributedWeightedSampler with rank 14 / 16
270
+ [INFO ][2026-05-10 12:24:08][root ][make_videodataset ] VideoDataset unsupervised data loader created
271
+ [INFO ][2026-05-10 12:24:08][app.vjepa.train ][main ] iterations per epoch/dataset length: 300/399
272
+ [INFO ][2026-05-10 12:24:08][app.vjepa.train ][main ] Wrapping models in DDP (rank 14)...
273
+ [INFO ][2026-05-10 12:24:08][app.vjepa.train ][main ] Initializing loader...
274
+ submitit ERROR (2026-05-10 16:19:50,186) - Submitted job triggered an exception
275
+ [ERROR ][2026-05-10 16:19:50][submitit ][process_job ] Submitted job triggered an exception
276
+ [ERROR ][2026-05-10 16:19:51][app.vjepa.train ][_save_crash ] Saved crash log to /scratch-shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/crash_rank14.log
vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/job_22614766/22614766_15_log.err ADDED
@@ -0,0 +1,68 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/timm/models/layers/__init__.py:49: FutureWarning: Importing from timm.models.layers is deprecated, please import via timm.layers
2
+ warnings.warn(f"Importing from {__name__} is deprecated, please import via timm.layers", FutureWarning)
3
+ [rank15]:[W510 12:19:56.808669295 ProcessGroupNCCL.cpp:5138] Guessing device ID based on global rank. This can cause a hang if rank to GPU mapping is heterogeneous. You can specify device_id in init_process_group()
4
+ /gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/vjepa/utils.py:796: FutureWarning: `torch.cuda.amp.GradScaler(args...)` is deprecated. Please use `torch.amp.GradScaler('cuda', args...)` instead.
5
+ scaler = torch.cuda.amp.GradScaler() if mixed_precision else None
6
+ /gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/vjepa/train.py:955: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
7
+ with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
8
+ submitit ERROR (2026-05-10 16:19:48,628) - Submitted job triggered an exception
9
+ [rank15]:[W510 16:19:49.920372704 TCPStore.cpp:125] [c10d] recvValue failed on SocketImpl(fd=3, addr=[gcn87.local.snellius.surf.nl]:50954, remote=[gcn80.local.snellius.surf.nl]:37129): Failed to recv, got 0 bytes. Connection was likely closed. Did the remote server shutdown or crash?
10
+ Exception raised from recvBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:682 (most recent call first):
11
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x1523c078ffdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
12
+ frame #1: <unknown function> + 0x6a3325d (0x15240491e25d in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
13
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x273 (0x15240491c1f3 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
14
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x1523c19feeec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
15
+ frame #4: <unknown function> + 0xed164 (0x152500e68164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
16
+ frame #5: <unknown function> + 0x8a19a (0x1525118da19a in /lib64/libc.so.6)
17
+ frame #6: <unknown function> + 0x10f100 (0x15251195f100 in /lib64/libc.so.6)
18
+
19
+ [rank15]:[W510 16:19:49.926026027 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 15] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Failed to recv, got 0 bytes. Connection was likely closed. Did the remote server shutdown or crash?
20
+ [rank15]:[W510 16:19:50.926168251 TCPStore.cpp:106] [c10d] sendBytes failed on SocketImpl(fd=3, addr=[gcn87.local.snellius.surf.nl]:50954, remote=[gcn80.local.snellius.surf.nl]:37129): Broken pipe
21
+ Exception raised from sendBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:653 (most recent call first):
22
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x1523c078ffdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
23
+ frame #1: <unknown function> + 0x6a326d1 (0x15240491d6d1 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
24
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x24d (0x15240491c1cd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
25
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x1523c19feeec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
26
+ frame #4: <unknown function> + 0xed164 (0x152500e68164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
27
+ frame #5: <unknown function> + 0x8a19a (0x1525118da19a in /lib64/libc.so.6)
28
+ frame #6: <unknown function> + 0x10f100 (0x15251195f100 in /lib64/libc.so.6)
29
+
30
+ [rank15]:[W510 16:19:50.929624471 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 15] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Broken pipe
31
+ Traceback (most recent call last):
32
+ File "<frozen runpy>", line 198, in _run_module_as_main
33
+ File "<frozen runpy>", line 88, in _run_code
34
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/_submit.py", line 11, in <module>
35
+ submitit_main()
36
+ ~~~~~~~~~~~~~^^
37
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 76, in submitit_main
38
+ process_job(args.folder)
39
+ ~~~~~~~~~~~^^^^^^^^^^^^^
40
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 69, in process_job
41
+ raise error
42
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 55, in process_job
43
+ result = delayed.result()
44
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/utils.py", line 137, in result
45
+ self._result = self.function(*self.args, **self.kwargs)
46
+ ~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^
47
+ File "/gpfs/home2/dcanez/vd/app/main_distributed.py", line 102, in __call__
48
+ app_main(app, args=params, resume_preempt=resume_preempt)
49
+ ~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
50
+ File "/gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/scaffold.py", line 17, in main
51
+ return importlib.import_module(f"app.{app}.train").main(args=args, resume_preempt=resume_preempt)
52
+ ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
53
+ File "/gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/vjepa/train.py", line 1189, in main
54
+ assert not np.isnan(loss), "loss is nan"
55
+ ^^^^^^^^^^^^^^^^^^
56
+ AssertionError: loss is nan
57
+ [rank15]:[W510 16:19:51.929784985 TCPStore.cpp:106] [c10d] sendBytes failed on SocketImpl(fd=3, addr=[gcn87.local.snellius.surf.nl]:50954, remote=[gcn80.local.snellius.surf.nl]:37129): Broken pipe
58
+ Exception raised from sendBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:653 (most recent call first):
59
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x1523c078ffdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
60
+ frame #1: <unknown function> + 0x6a326d1 (0x15240491d6d1 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
61
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x24d (0x15240491c1cd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
62
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x1523c19feeec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
63
+ frame #4: <unknown function> + 0xed164 (0x152500e68164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
64
+ frame #5: <unknown function> + 0x8a19a (0x1525118da19a in /lib64/libc.so.6)
65
+ frame #6: <unknown function> + 0x10f100 (0x15251195f100 in /lib64/libc.so.6)
66
+
67
+ [rank15]:[W510 16:19:51.932123188 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 15] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Broken pipe
68
+ [rank15]:[W510 16:19:51.445334286 ProcessGroupNCCL.cpp:1553] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())
vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/job_22614766/22614766_15_log.out ADDED
@@ -0,0 +1,276 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ submitit INFO (2026-05-10 12:14:34,055) - Starting with JobEnvironment(job_id=22614766, hostname=gcn87.local.snellius.surf.nl, local_rank=3(4), node=3(4), global_rank=15(16))
2
+ submitit INFO (2026-05-10 12:14:34,056) - Loading pickle: /scratch-shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/job_22614766/22614766_submitted.pkl
3
+ INFO:root:loaded pretrain params...
4
+ { 'app': 'vjepa',
5
+ 'cpus_per_task': 16,
6
+ 'data': { 'batch_size': 64,
7
+ 'crop_size': 224,
8
+ 'dataset_fpcs': [16, 16],
9
+ 'dataset_type': 'VideoDataset',
10
+ 'datasets': [ '/scratch-shared/dcanez/data/kinetics/k400/train.csv',
11
+ '/scratch-shared/dcanez/data/ssv2/train.csv'],
12
+ 'datasets_weights': [0.65, 0.35],
13
+ 'fps': 4,
14
+ 'num_workers': 10,
15
+ 'patch_size': 14,
16
+ 'persistent_workers': True,
17
+ 'pin_mem': True,
18
+ 'stage': [ { 'dest': 'kinetics_240',
19
+ 'format': 'targz_parts',
20
+ 'src': '/scratch-shared/dcanez/data/kinetics/k400/tars_240/'},
21
+ { 'dest': 'ssv2',
22
+ 'format': 'multipart_tar',
23
+ 'src': '/scratch-nvme/ml-datasets/something-something-v2/'}],
24
+ 'tubelet_size': 1},
25
+ 'data_aug': { 'auto_augment': False,
26
+ 'motion_shift': False,
27
+ 'random_resize_aspect_ratio': [0.75, 1.35],
28
+ 'random_resize_scale': [0.3, 1.0],
29
+ 'reprob': 0.0},
30
+ 'folder': '/scratch-shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable',
31
+ 'loss': {'loss_exp': 1.0},
32
+ 'mask': [ { 'aspect_ratio': [0.75, 1.5],
33
+ 'full_complement': False,
34
+ 'max_keep': None,
35
+ 'max_temporal_keep': 1.0,
36
+ 'num_blocks': 8,
37
+ 'spatial_scale': [0.15, 0.15],
38
+ 'temporal_scale': [1.0, 1.0]},
39
+ { 'aspect_ratio': [0.75, 1.5],
40
+ 'full_complement': False,
41
+ 'max_keep': None,
42
+ 'max_temporal_keep': 1.0,
43
+ 'num_blocks': 2,
44
+ 'spatial_scale': [0.7, 0.7],
45
+ 'temporal_scale': [1.0, 1.0]}],
46
+ 'mem_per_gpu': '180G',
47
+ 'meta': { 'dtype': 'bfloat16',
48
+ 'knn_eval_epoch0': False,
49
+ 'knn_eval_freq': 5,
50
+ 'knn_eval_presets': [ { 'config': { 'batch_size': 64,
51
+ 'dataset_train': '/scratch-shared/mdorkenw/ucf101/ucfTrainTestlist/train.csv',
52
+ 'dataset_val': '/scratch-shared/mdorkenw/ucf101/ucfTrainTestlist/val.csv',
53
+ 'eval_videos_per_class': 25,
54
+ 'num_workers': 8,
55
+ 'pool_type': 'slot_temporal_concat',
56
+ 'train_videos_per_class': 100},
57
+ 'preset': 'ucf101'},
58
+ { 'config': { 'batch_size': 64,
59
+ 'dataset_train': '/scratch-shared/mdorkenw/20bn-something-something-v2/something-something-v2-annotations/train_coarse10.csv',
60
+ 'dataset_val': '/scratch-shared/mdorkenw/20bn-something-something-v2/something-something-v2-annotations/val_coarse10.csv',
61
+ 'eval_videos_per_class': 100,
62
+ 'linear_probe': True,
63
+ 'num_workers': 8,
64
+ 'pool_type': 'slot_temporal_concat',
65
+ 'train_videos_per_class': 500},
66
+ 'preset': 'ssv2_coarse10'}],
67
+ 'load_checkpoint': True,
68
+ 'read_checkpoint': None,
69
+ 'save_every_freq': 5,
70
+ 'seed': 239,
71
+ 'use_sdpa': True,
72
+ 'use_wandb': True,
73
+ 'wandb_project': 'vjepa_ujepaside'},
74
+ 'metrics': {'sigreg': {}, 'std': {}},
75
+ 'model': { 'model_name': 'ujepaside_large_patch14_capi_lvd1689m',
76
+ 'pred_depth': 6,
77
+ 'pred_embed_dim': 384,
78
+ 'pred_num_heads': 12,
79
+ 'predictor': 'v2_cross',
80
+ 'st_causal': False,
81
+ 'st_drop_path': 0.2,
82
+ 'st_flex_enable': False,
83
+ 'st_layer_scale_init': 1e-05,
84
+ 'st_num_slots': 16,
85
+ 'st_slots_causal_within_frame': False,
86
+ 'target_kind': 'frozen_2d',
87
+ 'target_type': 'vit_large_patch14_capi.lvd1689m',
88
+ 'temporal_spacing': 1.0,
89
+ 'uniform_power': True,
90
+ 'use_activation_checkpointing': True,
91
+ 'use_mask_tokens': True,
92
+ 'use_rope': True,
93
+ 'use_sdpa': True,
94
+ 'zero_init_mask_tokens': True},
95
+ 'nodes': 4,
96
+ 'optimization': { 'clip_grad': 3.0,
97
+ 'ema': [0.99925, 0.99925],
98
+ 'epochs': 100,
99
+ 'final_lr': 0.0001,
100
+ 'final_weight_decay': 0.04,
101
+ 'ipe': 300,
102
+ 'ipe_scale': 1.0,
103
+ 'lr': 0.0005,
104
+ 'start_lr': 0.0001,
105
+ 'warmup': 10,
106
+ 'weight_decay': 0.04},
107
+ 'tasks_per_node': 4}
108
+ INFO:root:Running pre-training of app: vjepa
109
+ [INFO ][2026-05-10 12:16:33][app.vjepa.train ][main ] which_dtype='bfloat16'
110
+ [INFO ][2026-05-10 12:16:33][app.vjepa.train ][main ] Disabling persistent_workers (incompatible with KNN eval)
111
+ [INFO ][2026-05-10 12:16:33][app.vjepa.train ][main ] NCCL_SOCKET_IFNAME=eno
112
+ [INFO ][2026-05-10 12:16:57][app.vjepa.train ][main ] Initialized (rank/world-size) 15/16, tasks_per_node=4
113
+ [INFO ][2026-05-10 12:16:57][root ][stage_datasets ] [local_rank 3/4] Staging kinetics_240 (targz_parts)
114
+ [INFO ][2026-05-10 12:16:57][root ][_stage_targz_parts ] [rank 3] Extracting 25/103 tar.gz parts to /scratch-node/dcanez.22614766/kinetics_240
115
+ [INFO ][2026-05-10 12:17:11][root ][_stage_targz_parts ] [local_rank 3] Extracted 2/25 parts
116
+ [INFO ][2026-05-10 12:17:25][root ][_stage_targz_parts ] [local_rank 3] Extracted 4/25 parts
117
+ [INFO ][2026-05-10 12:17:39][root ][_stage_targz_parts ] [local_rank 3] Extracted 6/25 parts
118
+ [INFO ][2026-05-10 12:17:53][root ][_stage_targz_parts ] [local_rank 3] Extracted 8/25 parts
119
+ [INFO ][2026-05-10 12:18:08][root ][_stage_targz_parts ] [local_rank 3] Extracted 10/25 parts
120
+ [INFO ][2026-05-10 12:18:22][root ][_stage_targz_parts ] [local_rank 3] Extracted 12/25 parts
121
+ [INFO ][2026-05-10 12:18:36][root ][_stage_targz_parts ] [local_rank 3] Extracted 14/25 parts
122
+ [INFO ][2026-05-10 12:18:51][root ][_stage_targz_parts ] [local_rank 3] Extracted 16/25 parts
123
+ [INFO ][2026-05-10 12:19:04][root ][_stage_targz_parts ] [local_rank 3] Extracted 18/25 parts
124
+ [INFO ][2026-05-10 12:19:19][root ][_stage_targz_parts ] [local_rank 3] Extracted 20/25 parts
125
+ [INFO ][2026-05-10 12:19:33][root ][_stage_targz_parts ] [local_rank 3] Extracted 22/25 parts
126
+ [INFO ][2026-05-10 12:19:47][root ][_stage_targz_parts ] [local_rank 3] Extracted 24/25 parts
127
+ [INFO ][2026-05-10 12:19:54][root ][_stage_targz_parts ] [local_rank 3] Extracted 25/25 parts
128
+ [INFO ][2026-05-10 12:19:54][root ][stage_datasets ] [local_rank 3/4] Staging ssv2 (multipart_tar)
129
+ [INFO ][2026-05-10 12:21:03][app.vjepa.train ][main ] Staged datasets: ['/scratch-node/dcanez.22614766/kinetics_240/train.csv', '/scratch-node/dcanez.22614766/ssv2/train.csv']
130
+ [INFO ][2026-05-10 12:23:31][experiments.stmodels.vision_transformers_v3][vit_large_patch14_capi_lvd1689m] Loading pretrained weights for vit_large_patch14_capi_lvd1689m from https://dl.fbaipublicfiles.com/capi/capi_vitl14_lvd.pth
131
+ [INFO ][2026-05-10 12:23:57][root ][_build_frozen_2d_target ] Loaded frozen_2d target from timm: vit_large_patch14_capi.lvd1689m
132
+ [INFO ][2026-05-10 12:23:57][root ][init_video_model ] ViTMultiSeqWrapper(
133
+ (backbone): UJEPAside(
134
+ (patch_embed): PatchEmbed3D(
135
+ (proj): Conv3d(3, 1024, kernel_size=(1, 14, 14), stride=(1, 14, 14))
136
+ )
137
+ (rope): CAPI2DRoPE()
138
+ (blocks): ModuleList(
139
+ (0-23): 24 x Block(
140
+ (norm1): RMSNorm((1024,), eps=1e-05, elementwise_affine=True)
141
+ (rope_impl): CAPI2DRoPE()
142
+ (norm2): RMSNorm((1024,), eps=1e-05, elementwise_affine=True)
143
+ (drop_path1): Identity()
144
+ (drop_path2): Identity()
145
+ (attn): Attention(
146
+ (qkv): Linear(in_features=1024, out_features=3072, bias=False)
147
+ (attn_drop): Dropout(p=0.0, inplace=False)
148
+ (proj): Linear(in_features=1024, out_features=1024, bias=False)
149
+ (proj_drop): Dropout(p=0.0, inplace=False)
150
+ (rope_impl): CAPI2DRoPE()
151
+ )
152
+ (mlp): MLP(
153
+ (fc1): Linear(in_features=1024, out_features=4096, bias=False)
154
+ (act): GELU(approximate='none')
155
+ (fc2): Linear(in_features=4096, out_features=1024, bias=False)
156
+ (drop): Dropout(p=0.0, inplace=False)
157
+ )
158
+ )
159
+ )
160
+ (norm): RMSNorm((1024,), eps=1e-05, elementwise_affine=False)
161
+ (st_blocks): ModuleList(
162
+ (0-23): 24 x Block(
163
+ (residual1): EfficientResidual(
164
+ (norm): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
165
+ (fn): Attention(
166
+ (q_proj): Linear(in_features=1024, out_features=1024, bias=False)
167
+ (k_proj): Linear(in_features=1024, out_features=1024, bias=False)
168
+ (v_proj): Linear(in_features=1024, out_features=1024, bias=False)
169
+ (proj): Linear(in_features=1024, out_features=1024, bias=False)
170
+ (rope_impl): CAPI3DRoPE()
171
+ )
172
+ )
173
+ (residual2): EfficientResidual(
174
+ (norm): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
175
+ (fn): MLP(
176
+ (fc1): Linear(in_features=1024, out_features=4096, bias=False)
177
+ (act): GELU(approximate='none')
178
+ (fc2): Linear(in_features=4096, out_features=1024, bias=False)
179
+ (drop): Dropout(p=0.0, inplace=False)
180
+ )
181
+ )
182
+ )
183
+ )
184
+ (st_rope): CAPI3DRoPE()
185
+ )
186
+ )
187
+ [INFO ][2026-05-10 12:23:57][root ][init_video_model ] UJEPAsidePredictorMultiSeqWrapper(
188
+ (backbone): PredictorV2(
189
+ (predictor_embed): Linear(in_features=1024, out_features=384, bias=True)
190
+ (mask_tokens): ParameterList(
191
+ (0): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
192
+ (1): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
193
+ (2): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
194
+ (3): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
195
+ )
196
+ (predictor_blocks): ModuleList(
197
+ (0-5): 6 x Block(
198
+ (residual1): EfficientResidual(
199
+ (norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
200
+ (fn): Attention(
201
+ (q_proj): Linear(in_features=384, out_features=384, bias=False)
202
+ (k_proj): Linear(in_features=384, out_features=384, bias=False)
203
+ (v_proj): Linear(in_features=384, out_features=384, bias=False)
204
+ (proj): Linear(in_features=384, out_features=384, bias=False)
205
+ (rope): Rope()
206
+ )
207
+ )
208
+ (residual2): EfficientResidual(
209
+ (norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
210
+ (fn): MLP(
211
+ (fc1): Linear(in_features=384, out_features=1536, bias=False)
212
+ (act): GELU(approximate='none')
213
+ (fc2): Linear(in_features=1536, out_features=384, bias=False)
214
+ (drop): Dropout(p=0.0, inplace=False)
215
+ )
216
+ )
217
+ )
218
+ )
219
+ (predictor_norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
220
+ (predictor_proj): Linear(in_features=384, out_features=1024, bias=True)
221
+ )
222
+ )
223
+ [INFO ][2026-05-10 12:23:57][root ][init_video_model ] MultiSeqWrapper(
224
+ (backbone): Frozen2DTargetWrapper(
225
+ (backbone): Eva(
226
+ (patch_embed): PatchEmbed(
227
+ (proj): Conv2d(3, 1024, kernel_size=(14, 14), stride=(14, 14))
228
+ (norm): Identity()
229
+ )
230
+ (pos_drop): Dropout(p=0.0, inplace=False)
231
+ (norm_pre): Identity()
232
+ (blocks): ModuleList(
233
+ (0-23): 24 x EvaBlock(
234
+ (norm1): RMSNorm((1024,), eps=1e-05, elementwise_affine=True)
235
+ (attn): EvaAttention(
236
+ (qkv): Linear(in_features=1024, out_features=3072, bias=False)
237
+ (q_norm): Identity()
238
+ (k_norm): Identity()
239
+ (attn_drop): Dropout(p=0.0, inplace=False)
240
+ (norm): Identity()
241
+ (proj): Linear(in_features=1024, out_features=1024, bias=True)
242
+ (proj_drop): Dropout(p=0.0, inplace=False)
243
+ )
244
+ (drop_path1): Identity()
245
+ (norm2): RMSNorm((1024,), eps=1e-05, elementwise_affine=True)
246
+ (mlp): Mlp(
247
+ (fc1): Linear(in_features=1024, out_features=4096, bias=True)
248
+ (act): GELU(approximate='none')
249
+ (drop1): Dropout(p=0.0, inplace=False)
250
+ (norm): Identity()
251
+ (fc2): Linear(in_features=4096, out_features=1024, bias=True)
252
+ (drop2): Dropout(p=0.0, inplace=False)
253
+ )
254
+ (drop_path2): Identity()
255
+ )
256
+ )
257
+ (norm): RMSNorm((1024,), eps=1e-05, elementwise_affine=True)
258
+ (fc_norm): Identity()
259
+ (head_drop): Dropout(p=0.0, inplace=False)
260
+ (head): Identity()
261
+ (rope): _CapiPatchRoPE()
262
+ )
263
+ )
264
+ )
265
+ [INFO ][2026-05-10 12:23:57][root ][init_video_model ] Encoder number of parameters: 302399488
266
+ [INFO ][2026-05-10 12:23:57][root ][init_video_model ] Predictor number of parameters: 11416192
267
+ [INFO ][2026-05-10 12:23:57][root ][init_video_model ] Target encoder number of parameters: 0
268
+ [INFO ][2026-05-10 12:24:08][root ][make_videodataset ] VideoDataset dataset created
269
+ [INFO ][2026-05-10 12:24:08][WeightedSampler ][__init__ ] Using DistributedWeightedSampler with rank 15 / 16
270
+ [INFO ][2026-05-10 12:24:08][root ][make_videodataset ] VideoDataset unsupervised data loader created
271
+ [INFO ][2026-05-10 12:24:08][app.vjepa.train ][main ] iterations per epoch/dataset length: 300/399
272
+ [INFO ][2026-05-10 12:24:08][app.vjepa.train ][main ] Wrapping models in DDP (rank 15)...
273
+ [INFO ][2026-05-10 12:24:08][app.vjepa.train ][main ] Initializing loader...
274
+ submitit ERROR (2026-05-10 16:19:48,628) - Submitted job triggered an exception
275
+ [ERROR ][2026-05-10 16:19:48][submitit ][process_job ] Submitted job triggered an exception
276
+ [ERROR ][2026-05-10 16:19:51][app.vjepa.train ][_save_crash ] Saved crash log to /scratch-shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/crash_rank15.log
vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/job_22614766/22614766_1_log.err ADDED
@@ -0,0 +1,41 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/timm/models/layers/__init__.py:49: FutureWarning: Importing from timm.models.layers is deprecated, please import via timm.layers
2
+ warnings.warn(f"Importing from {__name__} is deprecated, please import via timm.layers", FutureWarning)
3
+ [rank1]:[W510 12:19:58.977747785 ProcessGroupNCCL.cpp:5138] Guessing device ID based on global rank. This can cause a hang if rank to GPU mapping is heterogeneous. You can specify device_id in init_process_group()
4
+ /gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/vjepa/utils.py:796: FutureWarning: `torch.cuda.amp.GradScaler(args...)` is deprecated. Please use `torch.amp.GradScaler('cuda', args...)` instead.
5
+ scaler = torch.cuda.amp.GradScaler() if mixed_precision else None
6
+ /gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/vjepa/train.py:955: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
7
+ with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
8
+ /gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/vjepa/train.py:955: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
9
+ with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
10
+ /gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/vjepa/train.py:955: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
11
+ with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
12
+ /gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/vjepa/train.py:955: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
13
+ with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
14
+ submitit ERROR (2026-05-10 16:19:44,648) - Submitted job triggered an exception
15
+ Traceback (most recent call last):
16
+ File "<frozen runpy>", line 198, in _run_module_as_main
17
+ File "<frozen runpy>", line 88, in _run_code
18
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/_submit.py", line 11, in <module>
19
+ submitit_main()
20
+ ~~~~~~~~~~~~~^^
21
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 76, in submitit_main
22
+ process_job(args.folder)
23
+ ~~~~~~~~~~~^^^^^^^^^^^^^
24
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 69, in process_job
25
+ raise error
26
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 55, in process_job
27
+ result = delayed.result()
28
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/utils.py", line 137, in result
29
+ self._result = self.function(*self.args, **self.kwargs)
30
+ ~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^
31
+ File "/gpfs/home2/dcanez/vd/app/main_distributed.py", line 102, in __call__
32
+ app_main(app, args=params, resume_preempt=resume_preempt)
33
+ ~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
34
+ File "/gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/scaffold.py", line 17, in main
35
+ return importlib.import_module(f"app.{app}.train").main(args=args, resume_preempt=resume_preempt)
36
+ ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
37
+ File "/gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/vjepa/train.py", line 1189, in main
38
+ assert not np.isnan(loss), "loss is nan"
39
+ ^^^^^^^^^^^^^^^^^^
40
+ AssertionError: loss is nan
41
+ [rank1]:[W510 16:19:45.445536181 ProcessGroupNCCL.cpp:1553] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())
vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/job_22614766/22614766_1_log.out ADDED
@@ -0,0 +1,276 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ submitit INFO (2026-05-10 12:14:34,051) - Starting with JobEnvironment(job_id=22614766, hostname=gcn80.local.snellius.surf.nl, local_rank=1(4), node=0(4), global_rank=1(16))
2
+ submitit INFO (2026-05-10 12:14:34,051) - Loading pickle: /scratch-shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/job_22614766/22614766_submitted.pkl
3
+ INFO:root:loaded pretrain params...
4
+ { 'app': 'vjepa',
5
+ 'cpus_per_task': 16,
6
+ 'data': { 'batch_size': 64,
7
+ 'crop_size': 224,
8
+ 'dataset_fpcs': [16, 16],
9
+ 'dataset_type': 'VideoDataset',
10
+ 'datasets': [ '/scratch-shared/dcanez/data/kinetics/k400/train.csv',
11
+ '/scratch-shared/dcanez/data/ssv2/train.csv'],
12
+ 'datasets_weights': [0.65, 0.35],
13
+ 'fps': 4,
14
+ 'num_workers': 10,
15
+ 'patch_size': 14,
16
+ 'persistent_workers': True,
17
+ 'pin_mem': True,
18
+ 'stage': [ { 'dest': 'kinetics_240',
19
+ 'format': 'targz_parts',
20
+ 'src': '/scratch-shared/dcanez/data/kinetics/k400/tars_240/'},
21
+ { 'dest': 'ssv2',
22
+ 'format': 'multipart_tar',
23
+ 'src': '/scratch-nvme/ml-datasets/something-something-v2/'}],
24
+ 'tubelet_size': 1},
25
+ 'data_aug': { 'auto_augment': False,
26
+ 'motion_shift': False,
27
+ 'random_resize_aspect_ratio': [0.75, 1.35],
28
+ 'random_resize_scale': [0.3, 1.0],
29
+ 'reprob': 0.0},
30
+ 'folder': '/scratch-shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable',
31
+ 'loss': {'loss_exp': 1.0},
32
+ 'mask': [ { 'aspect_ratio': [0.75, 1.5],
33
+ 'full_complement': False,
34
+ 'max_keep': None,
35
+ 'max_temporal_keep': 1.0,
36
+ 'num_blocks': 8,
37
+ 'spatial_scale': [0.15, 0.15],
38
+ 'temporal_scale': [1.0, 1.0]},
39
+ { 'aspect_ratio': [0.75, 1.5],
40
+ 'full_complement': False,
41
+ 'max_keep': None,
42
+ 'max_temporal_keep': 1.0,
43
+ 'num_blocks': 2,
44
+ 'spatial_scale': [0.7, 0.7],
45
+ 'temporal_scale': [1.0, 1.0]}],
46
+ 'mem_per_gpu': '180G',
47
+ 'meta': { 'dtype': 'bfloat16',
48
+ 'knn_eval_epoch0': False,
49
+ 'knn_eval_freq': 5,
50
+ 'knn_eval_presets': [ { 'config': { 'batch_size': 64,
51
+ 'dataset_train': '/scratch-shared/mdorkenw/ucf101/ucfTrainTestlist/train.csv',
52
+ 'dataset_val': '/scratch-shared/mdorkenw/ucf101/ucfTrainTestlist/val.csv',
53
+ 'eval_videos_per_class': 25,
54
+ 'num_workers': 8,
55
+ 'pool_type': 'slot_temporal_concat',
56
+ 'train_videos_per_class': 100},
57
+ 'preset': 'ucf101'},
58
+ { 'config': { 'batch_size': 64,
59
+ 'dataset_train': '/scratch-shared/mdorkenw/20bn-something-something-v2/something-something-v2-annotations/train_coarse10.csv',
60
+ 'dataset_val': '/scratch-shared/mdorkenw/20bn-something-something-v2/something-something-v2-annotations/val_coarse10.csv',
61
+ 'eval_videos_per_class': 100,
62
+ 'linear_probe': True,
63
+ 'num_workers': 8,
64
+ 'pool_type': 'slot_temporal_concat',
65
+ 'train_videos_per_class': 500},
66
+ 'preset': 'ssv2_coarse10'}],
67
+ 'load_checkpoint': True,
68
+ 'read_checkpoint': None,
69
+ 'save_every_freq': 5,
70
+ 'seed': 239,
71
+ 'use_sdpa': True,
72
+ 'use_wandb': True,
73
+ 'wandb_project': 'vjepa_ujepaside'},
74
+ 'metrics': {'sigreg': {}, 'std': {}},
75
+ 'model': { 'model_name': 'ujepaside_large_patch14_capi_lvd1689m',
76
+ 'pred_depth': 6,
77
+ 'pred_embed_dim': 384,
78
+ 'pred_num_heads': 12,
79
+ 'predictor': 'v2_cross',
80
+ 'st_causal': False,
81
+ 'st_drop_path': 0.2,
82
+ 'st_flex_enable': False,
83
+ 'st_layer_scale_init': 1e-05,
84
+ 'st_num_slots': 16,
85
+ 'st_slots_causal_within_frame': False,
86
+ 'target_kind': 'frozen_2d',
87
+ 'target_type': 'vit_large_patch14_capi.lvd1689m',
88
+ 'temporal_spacing': 1.0,
89
+ 'uniform_power': True,
90
+ 'use_activation_checkpointing': True,
91
+ 'use_mask_tokens': True,
92
+ 'use_rope': True,
93
+ 'use_sdpa': True,
94
+ 'zero_init_mask_tokens': True},
95
+ 'nodes': 4,
96
+ 'optimization': { 'clip_grad': 3.0,
97
+ 'ema': [0.99925, 0.99925],
98
+ 'epochs': 100,
99
+ 'final_lr': 0.0001,
100
+ 'final_weight_decay': 0.04,
101
+ 'ipe': 300,
102
+ 'ipe_scale': 1.0,
103
+ 'lr': 0.0005,
104
+ 'start_lr': 0.0001,
105
+ 'warmup': 10,
106
+ 'weight_decay': 0.04},
107
+ 'tasks_per_node': 4}
108
+ INFO:root:Running pre-training of app: vjepa
109
+ [INFO ][2026-05-10 12:16:53][app.vjepa.train ][main ] which_dtype='bfloat16'
110
+ [INFO ][2026-05-10 12:16:53][app.vjepa.train ][main ] Disabling persistent_workers (incompatible with KNN eval)
111
+ [INFO ][2026-05-10 12:16:53][app.vjepa.train ][main ] NCCL_SOCKET_IFNAME=eno
112
+ [INFO ][2026-05-10 12:16:55][app.vjepa.train ][main ] Initialized (rank/world-size) 1/16, tasks_per_node=4
113
+ [INFO ][2026-05-10 12:16:55][root ][stage_datasets ] [local_rank 1/4] Staging kinetics_240 (targz_parts)
114
+ [INFO ][2026-05-10 12:16:55][root ][_stage_targz_parts ] [rank 1] Extracting 26/103 tar.gz parts to /scratch-node/dcanez.22614766/kinetics_240
115
+ [INFO ][2026-05-10 12:17:08][root ][_stage_targz_parts ] [local_rank 1] Extracted 2/26 parts
116
+ [INFO ][2026-05-10 12:17:22][root ][_stage_targz_parts ] [local_rank 1] Extracted 4/26 parts
117
+ [INFO ][2026-05-10 12:17:36][root ][_stage_targz_parts ] [local_rank 1] Extracted 6/26 parts
118
+ [INFO ][2026-05-10 12:17:50][root ][_stage_targz_parts ] [local_rank 1] Extracted 8/26 parts
119
+ [INFO ][2026-05-10 12:18:04][root ][_stage_targz_parts ] [local_rank 1] Extracted 10/26 parts
120
+ [INFO ][2026-05-10 12:18:19][root ][_stage_targz_parts ] [local_rank 1] Extracted 12/26 parts
121
+ [INFO ][2026-05-10 12:18:32][root ][_stage_targz_parts ] [local_rank 1] Extracted 14/26 parts
122
+ [INFO ][2026-05-10 12:18:46][root ][_stage_targz_parts ] [local_rank 1] Extracted 16/26 parts
123
+ [INFO ][2026-05-10 12:19:00][root ][_stage_targz_parts ] [local_rank 1] Extracted 18/26 parts
124
+ [INFO ][2026-05-10 12:19:15][root ][_stage_targz_parts ] [local_rank 1] Extracted 20/26 parts
125
+ [INFO ][2026-05-10 12:19:29][root ][_stage_targz_parts ] [local_rank 1] Extracted 22/26 parts
126
+ [INFO ][2026-05-10 12:19:43][root ][_stage_targz_parts ] [local_rank 1] Extracted 24/26 parts
127
+ [INFO ][2026-05-10 12:19:57][root ][_stage_targz_parts ] [local_rank 1] Extracted 26/26 parts
128
+ [INFO ][2026-05-10 12:19:57][root ][stage_datasets ] [local_rank 1/4] Staging ssv2 (multipart_tar)
129
+ [INFO ][2026-05-10 12:21:03][app.vjepa.train ][main ] Staged datasets: ['/scratch-node/dcanez.22614766/kinetics_240/train.csv', '/scratch-node/dcanez.22614766/ssv2/train.csv']
130
+ [INFO ][2026-05-10 12:23:31][experiments.stmodels.vision_transformers_v3][vit_large_patch14_capi_lvd1689m] Loading pretrained weights for vit_large_patch14_capi_lvd1689m from https://dl.fbaipublicfiles.com/capi/capi_vitl14_lvd.pth
131
+ [INFO ][2026-05-10 12:23:57][root ][_build_frozen_2d_target ] Loaded frozen_2d target from timm: vit_large_patch14_capi.lvd1689m
132
+ [INFO ][2026-05-10 12:23:57][root ][init_video_model ] ViTMultiSeqWrapper(
133
+ (backbone): UJEPAside(
134
+ (patch_embed): PatchEmbed3D(
135
+ (proj): Conv3d(3, 1024, kernel_size=(1, 14, 14), stride=(1, 14, 14))
136
+ )
137
+ (rope): CAPI2DRoPE()
138
+ (blocks): ModuleList(
139
+ (0-23): 24 x Block(
140
+ (norm1): RMSNorm((1024,), eps=1e-05, elementwise_affine=True)
141
+ (rope_impl): CAPI2DRoPE()
142
+ (norm2): RMSNorm((1024,), eps=1e-05, elementwise_affine=True)
143
+ (drop_path1): Identity()
144
+ (drop_path2): Identity()
145
+ (attn): Attention(
146
+ (qkv): Linear(in_features=1024, out_features=3072, bias=False)
147
+ (attn_drop): Dropout(p=0.0, inplace=False)
148
+ (proj): Linear(in_features=1024, out_features=1024, bias=False)
149
+ (proj_drop): Dropout(p=0.0, inplace=False)
150
+ (rope_impl): CAPI2DRoPE()
151
+ )
152
+ (mlp): MLP(
153
+ (fc1): Linear(in_features=1024, out_features=4096, bias=False)
154
+ (act): GELU(approximate='none')
155
+ (fc2): Linear(in_features=4096, out_features=1024, bias=False)
156
+ (drop): Dropout(p=0.0, inplace=False)
157
+ )
158
+ )
159
+ )
160
+ (norm): RMSNorm((1024,), eps=1e-05, elementwise_affine=False)
161
+ (st_blocks): ModuleList(
162
+ (0-23): 24 x Block(
163
+ (residual1): EfficientResidual(
164
+ (norm): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
165
+ (fn): Attention(
166
+ (q_proj): Linear(in_features=1024, out_features=1024, bias=False)
167
+ (k_proj): Linear(in_features=1024, out_features=1024, bias=False)
168
+ (v_proj): Linear(in_features=1024, out_features=1024, bias=False)
169
+ (proj): Linear(in_features=1024, out_features=1024, bias=False)
170
+ (rope_impl): CAPI3DRoPE()
171
+ )
172
+ )
173
+ (residual2): EfficientResidual(
174
+ (norm): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
175
+ (fn): MLP(
176
+ (fc1): Linear(in_features=1024, out_features=4096, bias=False)
177
+ (act): GELU(approximate='none')
178
+ (fc2): Linear(in_features=4096, out_features=1024, bias=False)
179
+ (drop): Dropout(p=0.0, inplace=False)
180
+ )
181
+ )
182
+ )
183
+ )
184
+ (st_rope): CAPI3DRoPE()
185
+ )
186
+ )
187
+ [INFO ][2026-05-10 12:23:57][root ][init_video_model ] UJEPAsidePredictorMultiSeqWrapper(
188
+ (backbone): PredictorV2(
189
+ (predictor_embed): Linear(in_features=1024, out_features=384, bias=True)
190
+ (mask_tokens): ParameterList(
191
+ (0): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
192
+ (1): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
193
+ (2): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
194
+ (3): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
195
+ )
196
+ (predictor_blocks): ModuleList(
197
+ (0-5): 6 x Block(
198
+ (residual1): EfficientResidual(
199
+ (norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
200
+ (fn): Attention(
201
+ (q_proj): Linear(in_features=384, out_features=384, bias=False)
202
+ (k_proj): Linear(in_features=384, out_features=384, bias=False)
203
+ (v_proj): Linear(in_features=384, out_features=384, bias=False)
204
+ (proj): Linear(in_features=384, out_features=384, bias=False)
205
+ (rope): Rope()
206
+ )
207
+ )
208
+ (residual2): EfficientResidual(
209
+ (norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
210
+ (fn): MLP(
211
+ (fc1): Linear(in_features=384, out_features=1536, bias=False)
212
+ (act): GELU(approximate='none')
213
+ (fc2): Linear(in_features=1536, out_features=384, bias=False)
214
+ (drop): Dropout(p=0.0, inplace=False)
215
+ )
216
+ )
217
+ )
218
+ )
219
+ (predictor_norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
220
+ (predictor_proj): Linear(in_features=384, out_features=1024, bias=True)
221
+ )
222
+ )
223
+ [INFO ][2026-05-10 12:23:57][root ][init_video_model ] MultiSeqWrapper(
224
+ (backbone): Frozen2DTargetWrapper(
225
+ (backbone): Eva(
226
+ (patch_embed): PatchEmbed(
227
+ (proj): Conv2d(3, 1024, kernel_size=(14, 14), stride=(14, 14))
228
+ (norm): Identity()
229
+ )
230
+ (pos_drop): Dropout(p=0.0, inplace=False)
231
+ (norm_pre): Identity()
232
+ (blocks): ModuleList(
233
+ (0-23): 24 x EvaBlock(
234
+ (norm1): RMSNorm((1024,), eps=1e-05, elementwise_affine=True)
235
+ (attn): EvaAttention(
236
+ (qkv): Linear(in_features=1024, out_features=3072, bias=False)
237
+ (q_norm): Identity()
238
+ (k_norm): Identity()
239
+ (attn_drop): Dropout(p=0.0, inplace=False)
240
+ (norm): Identity()
241
+ (proj): Linear(in_features=1024, out_features=1024, bias=True)
242
+ (proj_drop): Dropout(p=0.0, inplace=False)
243
+ )
244
+ (drop_path1): Identity()
245
+ (norm2): RMSNorm((1024,), eps=1e-05, elementwise_affine=True)
246
+ (mlp): Mlp(
247
+ (fc1): Linear(in_features=1024, out_features=4096, bias=True)
248
+ (act): GELU(approximate='none')
249
+ (drop1): Dropout(p=0.0, inplace=False)
250
+ (norm): Identity()
251
+ (fc2): Linear(in_features=4096, out_features=1024, bias=True)
252
+ (drop2): Dropout(p=0.0, inplace=False)
253
+ )
254
+ (drop_path2): Identity()
255
+ )
256
+ )
257
+ (norm): RMSNorm((1024,), eps=1e-05, elementwise_affine=True)
258
+ (fc_norm): Identity()
259
+ (head_drop): Dropout(p=0.0, inplace=False)
260
+ (head): Identity()
261
+ (rope): _CapiPatchRoPE()
262
+ )
263
+ )
264
+ )
265
+ [INFO ][2026-05-10 12:23:57][root ][init_video_model ] Encoder number of parameters: 302399488
266
+ [INFO ][2026-05-10 12:23:57][root ][init_video_model ] Predictor number of parameters: 11416192
267
+ [INFO ][2026-05-10 12:23:57][root ][init_video_model ] Target encoder number of parameters: 0
268
+ [INFO ][2026-05-10 12:24:07][root ][make_videodataset ] VideoDataset dataset created
269
+ [INFO ][2026-05-10 12:24:07][WeightedSampler ][__init__ ] Using DistributedWeightedSampler with rank 1 / 16
270
+ [INFO ][2026-05-10 12:24:07][root ][make_videodataset ] VideoDataset unsupervised data loader created
271
+ [INFO ][2026-05-10 12:24:07][app.vjepa.train ][main ] iterations per epoch/dataset length: 300/399
272
+ [INFO ][2026-05-10 12:24:07][app.vjepa.train ][main ] Wrapping models in DDP (rank 1)...
273
+ [INFO ][2026-05-10 12:24:08][app.vjepa.train ][main ] Initializing loader...
274
+ submitit ERROR (2026-05-10 16:19:44,648) - Submitted job triggered an exception
275
+ [ERROR ][2026-05-10 16:19:44][submitit ][process_job ] Submitted job triggered an exception
276
+ [ERROR ][2026-05-10 16:19:44][app.vjepa.train ][_save_crash ] Saved crash log to /scratch-shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/crash_rank1.log
vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/job_22614766/22614766_2_log.err ADDED
@@ -0,0 +1,41 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/timm/models/layers/__init__.py:49: FutureWarning: Importing from timm.models.layers is deprecated, please import via timm.layers
2
+ warnings.warn(f"Importing from {__name__} is deprecated, please import via timm.layers", FutureWarning)
3
+ [rank2]:[W510 12:20:01.138964796 ProcessGroupNCCL.cpp:5138] Guessing device ID based on global rank. This can cause a hang if rank to GPU mapping is heterogeneous. You can specify device_id in init_process_group()
4
+ /gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/vjepa/utils.py:796: FutureWarning: `torch.cuda.amp.GradScaler(args...)` is deprecated. Please use `torch.amp.GradScaler('cuda', args...)` instead.
5
+ scaler = torch.cuda.amp.GradScaler() if mixed_precision else None
6
+ /gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/vjepa/train.py:955: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
7
+ with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
8
+ /gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/vjepa/train.py:955: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
9
+ with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
10
+ /gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/vjepa/train.py:955: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
11
+ with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
12
+ /gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/vjepa/train.py:955: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
13
+ with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
14
+ submitit ERROR (2026-05-10 16:19:44,648) - Submitted job triggered an exception
15
+ Traceback (most recent call last):
16
+ File "<frozen runpy>", line 198, in _run_module_as_main
17
+ File "<frozen runpy>", line 88, in _run_code
18
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/_submit.py", line 11, in <module>
19
+ submitit_main()
20
+ ~~~~~~~~~~~~~^^
21
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 76, in submitit_main
22
+ process_job(args.folder)
23
+ ~~~~~~~~~~~^^^^^^^^^^^^^
24
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 69, in process_job
25
+ raise error
26
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 55, in process_job
27
+ result = delayed.result()
28
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/utils.py", line 137, in result
29
+ self._result = self.function(*self.args, **self.kwargs)
30
+ ~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^
31
+ File "/gpfs/home2/dcanez/vd/app/main_distributed.py", line 102, in __call__
32
+ app_main(app, args=params, resume_preempt=resume_preempt)
33
+ ~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
34
+ File "/gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/scaffold.py", line 17, in main
35
+ return importlib.import_module(f"app.{app}.train").main(args=args, resume_preempt=resume_preempt)
36
+ ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
37
+ File "/gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/vjepa/train.py", line 1189, in main
38
+ assert not np.isnan(loss), "loss is nan"
39
+ ^^^^^^^^^^^^^^^^^^
40
+ AssertionError: loss is nan
41
+ [rank2]:[W510 16:19:45.446521046 ProcessGroupNCCL.cpp:1553] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())
vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/job_22614766/22614766_2_log.out ADDED
@@ -0,0 +1,276 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ submitit INFO (2026-05-10 12:14:34,051) - Starting with JobEnvironment(job_id=22614766, hostname=gcn80.local.snellius.surf.nl, local_rank=2(4), node=0(4), global_rank=2(16))
2
+ submitit INFO (2026-05-10 12:14:34,051) - Loading pickle: /scratch-shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/job_22614766/22614766_submitted.pkl
3
+ INFO:root:loaded pretrain params...
4
+ { 'app': 'vjepa',
5
+ 'cpus_per_task': 16,
6
+ 'data': { 'batch_size': 64,
7
+ 'crop_size': 224,
8
+ 'dataset_fpcs': [16, 16],
9
+ 'dataset_type': 'VideoDataset',
10
+ 'datasets': [ '/scratch-shared/dcanez/data/kinetics/k400/train.csv',
11
+ '/scratch-shared/dcanez/data/ssv2/train.csv'],
12
+ 'datasets_weights': [0.65, 0.35],
13
+ 'fps': 4,
14
+ 'num_workers': 10,
15
+ 'patch_size': 14,
16
+ 'persistent_workers': True,
17
+ 'pin_mem': True,
18
+ 'stage': [ { 'dest': 'kinetics_240',
19
+ 'format': 'targz_parts',
20
+ 'src': '/scratch-shared/dcanez/data/kinetics/k400/tars_240/'},
21
+ { 'dest': 'ssv2',
22
+ 'format': 'multipart_tar',
23
+ 'src': '/scratch-nvme/ml-datasets/something-something-v2/'}],
24
+ 'tubelet_size': 1},
25
+ 'data_aug': { 'auto_augment': False,
26
+ 'motion_shift': False,
27
+ 'random_resize_aspect_ratio': [0.75, 1.35],
28
+ 'random_resize_scale': [0.3, 1.0],
29
+ 'reprob': 0.0},
30
+ 'folder': '/scratch-shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable',
31
+ 'loss': {'loss_exp': 1.0},
32
+ 'mask': [ { 'aspect_ratio': [0.75, 1.5],
33
+ 'full_complement': False,
34
+ 'max_keep': None,
35
+ 'max_temporal_keep': 1.0,
36
+ 'num_blocks': 8,
37
+ 'spatial_scale': [0.15, 0.15],
38
+ 'temporal_scale': [1.0, 1.0]},
39
+ { 'aspect_ratio': [0.75, 1.5],
40
+ 'full_complement': False,
41
+ 'max_keep': None,
42
+ 'max_temporal_keep': 1.0,
43
+ 'num_blocks': 2,
44
+ 'spatial_scale': [0.7, 0.7],
45
+ 'temporal_scale': [1.0, 1.0]}],
46
+ 'mem_per_gpu': '180G',
47
+ 'meta': { 'dtype': 'bfloat16',
48
+ 'knn_eval_epoch0': False,
49
+ 'knn_eval_freq': 5,
50
+ 'knn_eval_presets': [ { 'config': { 'batch_size': 64,
51
+ 'dataset_train': '/scratch-shared/mdorkenw/ucf101/ucfTrainTestlist/train.csv',
52
+ 'dataset_val': '/scratch-shared/mdorkenw/ucf101/ucfTrainTestlist/val.csv',
53
+ 'eval_videos_per_class': 25,
54
+ 'num_workers': 8,
55
+ 'pool_type': 'slot_temporal_concat',
56
+ 'train_videos_per_class': 100},
57
+ 'preset': 'ucf101'},
58
+ { 'config': { 'batch_size': 64,
59
+ 'dataset_train': '/scratch-shared/mdorkenw/20bn-something-something-v2/something-something-v2-annotations/train_coarse10.csv',
60
+ 'dataset_val': '/scratch-shared/mdorkenw/20bn-something-something-v2/something-something-v2-annotations/val_coarse10.csv',
61
+ 'eval_videos_per_class': 100,
62
+ 'linear_probe': True,
63
+ 'num_workers': 8,
64
+ 'pool_type': 'slot_temporal_concat',
65
+ 'train_videos_per_class': 500},
66
+ 'preset': 'ssv2_coarse10'}],
67
+ 'load_checkpoint': True,
68
+ 'read_checkpoint': None,
69
+ 'save_every_freq': 5,
70
+ 'seed': 239,
71
+ 'use_sdpa': True,
72
+ 'use_wandb': True,
73
+ 'wandb_project': 'vjepa_ujepaside'},
74
+ 'metrics': {'sigreg': {}, 'std': {}},
75
+ 'model': { 'model_name': 'ujepaside_large_patch14_capi_lvd1689m',
76
+ 'pred_depth': 6,
77
+ 'pred_embed_dim': 384,
78
+ 'pred_num_heads': 12,
79
+ 'predictor': 'v2_cross',
80
+ 'st_causal': False,
81
+ 'st_drop_path': 0.2,
82
+ 'st_flex_enable': False,
83
+ 'st_layer_scale_init': 1e-05,
84
+ 'st_num_slots': 16,
85
+ 'st_slots_causal_within_frame': False,
86
+ 'target_kind': 'frozen_2d',
87
+ 'target_type': 'vit_large_patch14_capi.lvd1689m',
88
+ 'temporal_spacing': 1.0,
89
+ 'uniform_power': True,
90
+ 'use_activation_checkpointing': True,
91
+ 'use_mask_tokens': True,
92
+ 'use_rope': True,
93
+ 'use_sdpa': True,
94
+ 'zero_init_mask_tokens': True},
95
+ 'nodes': 4,
96
+ 'optimization': { 'clip_grad': 3.0,
97
+ 'ema': [0.99925, 0.99925],
98
+ 'epochs': 100,
99
+ 'final_lr': 0.0001,
100
+ 'final_weight_decay': 0.04,
101
+ 'ipe': 300,
102
+ 'ipe_scale': 1.0,
103
+ 'lr': 0.0005,
104
+ 'start_lr': 0.0001,
105
+ 'warmup': 10,
106
+ 'weight_decay': 0.04},
107
+ 'tasks_per_node': 4}
108
+ INFO:root:Running pre-training of app: vjepa
109
+ [INFO ][2026-05-10 12:16:53][app.vjepa.train ][main ] which_dtype='bfloat16'
110
+ [INFO ][2026-05-10 12:16:53][app.vjepa.train ][main ] Disabling persistent_workers (incompatible with KNN eval)
111
+ [INFO ][2026-05-10 12:16:53][app.vjepa.train ][main ] NCCL_SOCKET_IFNAME=eno
112
+ [INFO ][2026-05-10 12:16:54][app.vjepa.train ][main ] Initialized (rank/world-size) 2/16, tasks_per_node=4
113
+ [INFO ][2026-05-10 12:16:55][root ][stage_datasets ] [local_rank 2/4] Staging kinetics_240 (targz_parts)
114
+ [INFO ][2026-05-10 12:16:55][root ][_stage_targz_parts ] [rank 2] Extracting 26/103 tar.gz parts to /scratch-node/dcanez.22614766/kinetics_240
115
+ [INFO ][2026-05-10 12:17:09][root ][_stage_targz_parts ] [local_rank 2] Extracted 2/26 parts
116
+ [INFO ][2026-05-10 12:17:23][root ][_stage_targz_parts ] [local_rank 2] Extracted 4/26 parts
117
+ [INFO ][2026-05-10 12:17:38][root ][_stage_targz_parts ] [local_rank 2] Extracted 6/26 parts
118
+ [INFO ][2026-05-10 12:17:53][root ][_stage_targz_parts ] [local_rank 2] Extracted 8/26 parts
119
+ [INFO ][2026-05-10 12:18:07][root ][_stage_targz_parts ] [local_rank 2] Extracted 10/26 parts
120
+ [INFO ][2026-05-10 12:18:21][root ][_stage_targz_parts ] [local_rank 2] Extracted 12/26 parts
121
+ [INFO ][2026-05-10 12:18:35][root ][_stage_targz_parts ] [local_rank 2] Extracted 14/26 parts
122
+ [INFO ][2026-05-10 12:18:49][root ][_stage_targz_parts ] [local_rank 2] Extracted 16/26 parts
123
+ [INFO ][2026-05-10 12:19:04][root ][_stage_targz_parts ] [local_rank 2] Extracted 18/26 parts
124
+ [INFO ][2026-05-10 12:19:18][root ][_stage_targz_parts ] [local_rank 2] Extracted 20/26 parts
125
+ [INFO ][2026-05-10 12:19:32][root ][_stage_targz_parts ] [local_rank 2] Extracted 22/26 parts
126
+ [INFO ][2026-05-10 12:19:46][root ][_stage_targz_parts ] [local_rank 2] Extracted 24/26 parts
127
+ [INFO ][2026-05-10 12:20:01][root ][_stage_targz_parts ] [local_rank 2] Extracted 26/26 parts
128
+ [INFO ][2026-05-10 12:20:01][root ][stage_datasets ] [local_rank 2/4] Staging ssv2 (multipart_tar)
129
+ [INFO ][2026-05-10 12:21:03][app.vjepa.train ][main ] Staged datasets: ['/scratch-node/dcanez.22614766/kinetics_240/train.csv', '/scratch-node/dcanez.22614766/ssv2/train.csv']
130
+ [INFO ][2026-05-10 12:23:31][experiments.stmodels.vision_transformers_v3][vit_large_patch14_capi_lvd1689m] Loading pretrained weights for vit_large_patch14_capi_lvd1689m from https://dl.fbaipublicfiles.com/capi/capi_vitl14_lvd.pth
131
+ [INFO ][2026-05-10 12:23:57][root ][_build_frozen_2d_target ] Loaded frozen_2d target from timm: vit_large_patch14_capi.lvd1689m
132
+ [INFO ][2026-05-10 12:23:57][root ][init_video_model ] ViTMultiSeqWrapper(
133
+ (backbone): UJEPAside(
134
+ (patch_embed): PatchEmbed3D(
135
+ (proj): Conv3d(3, 1024, kernel_size=(1, 14, 14), stride=(1, 14, 14))
136
+ )
137
+ (rope): CAPI2DRoPE()
138
+ (blocks): ModuleList(
139
+ (0-23): 24 x Block(
140
+ (norm1): RMSNorm((1024,), eps=1e-05, elementwise_affine=True)
141
+ (rope_impl): CAPI2DRoPE()
142
+ (norm2): RMSNorm((1024,), eps=1e-05, elementwise_affine=True)
143
+ (drop_path1): Identity()
144
+ (drop_path2): Identity()
145
+ (attn): Attention(
146
+ (qkv): Linear(in_features=1024, out_features=3072, bias=False)
147
+ (attn_drop): Dropout(p=0.0, inplace=False)
148
+ (proj): Linear(in_features=1024, out_features=1024, bias=False)
149
+ (proj_drop): Dropout(p=0.0, inplace=False)
150
+ (rope_impl): CAPI2DRoPE()
151
+ )
152
+ (mlp): MLP(
153
+ (fc1): Linear(in_features=1024, out_features=4096, bias=False)
154
+ (act): GELU(approximate='none')
155
+ (fc2): Linear(in_features=4096, out_features=1024, bias=False)
156
+ (drop): Dropout(p=0.0, inplace=False)
157
+ )
158
+ )
159
+ )
160
+ (norm): RMSNorm((1024,), eps=1e-05, elementwise_affine=False)
161
+ (st_blocks): ModuleList(
162
+ (0-23): 24 x Block(
163
+ (residual1): EfficientResidual(
164
+ (norm): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
165
+ (fn): Attention(
166
+ (q_proj): Linear(in_features=1024, out_features=1024, bias=False)
167
+ (k_proj): Linear(in_features=1024, out_features=1024, bias=False)
168
+ (v_proj): Linear(in_features=1024, out_features=1024, bias=False)
169
+ (proj): Linear(in_features=1024, out_features=1024, bias=False)
170
+ (rope_impl): CAPI3DRoPE()
171
+ )
172
+ )
173
+ (residual2): EfficientResidual(
174
+ (norm): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
175
+ (fn): MLP(
176
+ (fc1): Linear(in_features=1024, out_features=4096, bias=False)
177
+ (act): GELU(approximate='none')
178
+ (fc2): Linear(in_features=4096, out_features=1024, bias=False)
179
+ (drop): Dropout(p=0.0, inplace=False)
180
+ )
181
+ )
182
+ )
183
+ )
184
+ (st_rope): CAPI3DRoPE()
185
+ )
186
+ )
187
+ [INFO ][2026-05-10 12:23:57][root ][init_video_model ] UJEPAsidePredictorMultiSeqWrapper(
188
+ (backbone): PredictorV2(
189
+ (predictor_embed): Linear(in_features=1024, out_features=384, bias=True)
190
+ (mask_tokens): ParameterList(
191
+ (0): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
192
+ (1): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
193
+ (2): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
194
+ (3): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
195
+ )
196
+ (predictor_blocks): ModuleList(
197
+ (0-5): 6 x Block(
198
+ (residual1): EfficientResidual(
199
+ (norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
200
+ (fn): Attention(
201
+ (q_proj): Linear(in_features=384, out_features=384, bias=False)
202
+ (k_proj): Linear(in_features=384, out_features=384, bias=False)
203
+ (v_proj): Linear(in_features=384, out_features=384, bias=False)
204
+ (proj): Linear(in_features=384, out_features=384, bias=False)
205
+ (rope): Rope()
206
+ )
207
+ )
208
+ (residual2): EfficientResidual(
209
+ (norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
210
+ (fn): MLP(
211
+ (fc1): Linear(in_features=384, out_features=1536, bias=False)
212
+ (act): GELU(approximate='none')
213
+ (fc2): Linear(in_features=1536, out_features=384, bias=False)
214
+ (drop): Dropout(p=0.0, inplace=False)
215
+ )
216
+ )
217
+ )
218
+ )
219
+ (predictor_norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
220
+ (predictor_proj): Linear(in_features=384, out_features=1024, bias=True)
221
+ )
222
+ )
223
+ [INFO ][2026-05-10 12:23:57][root ][init_video_model ] MultiSeqWrapper(
224
+ (backbone): Frozen2DTargetWrapper(
225
+ (backbone): Eva(
226
+ (patch_embed): PatchEmbed(
227
+ (proj): Conv2d(3, 1024, kernel_size=(14, 14), stride=(14, 14))
228
+ (norm): Identity()
229
+ )
230
+ (pos_drop): Dropout(p=0.0, inplace=False)
231
+ (norm_pre): Identity()
232
+ (blocks): ModuleList(
233
+ (0-23): 24 x EvaBlock(
234
+ (norm1): RMSNorm((1024,), eps=1e-05, elementwise_affine=True)
235
+ (attn): EvaAttention(
236
+ (qkv): Linear(in_features=1024, out_features=3072, bias=False)
237
+ (q_norm): Identity()
238
+ (k_norm): Identity()
239
+ (attn_drop): Dropout(p=0.0, inplace=False)
240
+ (norm): Identity()
241
+ (proj): Linear(in_features=1024, out_features=1024, bias=True)
242
+ (proj_drop): Dropout(p=0.0, inplace=False)
243
+ )
244
+ (drop_path1): Identity()
245
+ (norm2): RMSNorm((1024,), eps=1e-05, elementwise_affine=True)
246
+ (mlp): Mlp(
247
+ (fc1): Linear(in_features=1024, out_features=4096, bias=True)
248
+ (act): GELU(approximate='none')
249
+ (drop1): Dropout(p=0.0, inplace=False)
250
+ (norm): Identity()
251
+ (fc2): Linear(in_features=4096, out_features=1024, bias=True)
252
+ (drop2): Dropout(p=0.0, inplace=False)
253
+ )
254
+ (drop_path2): Identity()
255
+ )
256
+ )
257
+ (norm): RMSNorm((1024,), eps=1e-05, elementwise_affine=True)
258
+ (fc_norm): Identity()
259
+ (head_drop): Dropout(p=0.0, inplace=False)
260
+ (head): Identity()
261
+ (rope): _CapiPatchRoPE()
262
+ )
263
+ )
264
+ )
265
+ [INFO ][2026-05-10 12:23:57][root ][init_video_model ] Encoder number of parameters: 302399488
266
+ [INFO ][2026-05-10 12:23:57][root ][init_video_model ] Predictor number of parameters: 11416192
267
+ [INFO ][2026-05-10 12:23:57][root ][init_video_model ] Target encoder number of parameters: 0
268
+ [INFO ][2026-05-10 12:24:07][root ][make_videodataset ] VideoDataset dataset created
269
+ [INFO ][2026-05-10 12:24:07][WeightedSampler ][__init__ ] Using DistributedWeightedSampler with rank 2 / 16
270
+ [INFO ][2026-05-10 12:24:07][root ][make_videodataset ] VideoDataset unsupervised data loader created
271
+ [INFO ][2026-05-10 12:24:07][app.vjepa.train ][main ] iterations per epoch/dataset length: 300/399
272
+ [INFO ][2026-05-10 12:24:07][app.vjepa.train ][main ] Wrapping models in DDP (rank 2)...
273
+ [INFO ][2026-05-10 12:24:08][app.vjepa.train ][main ] Initializing loader...
274
+ submitit ERROR (2026-05-10 16:19:44,648) - Submitted job triggered an exception
275
+ [ERROR ][2026-05-10 16:19:44][submitit ][process_job ] Submitted job triggered an exception
276
+ [ERROR ][2026-05-10 16:19:44][app.vjepa.train ][_save_crash ] Saved crash log to /scratch-shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/crash_rank2.log
vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/job_22614766/22614766_3_log.err ADDED
@@ -0,0 +1,41 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/timm/models/layers/__init__.py:49: FutureWarning: Importing from timm.models.layers is deprecated, please import via timm.layers
2
+ warnings.warn(f"Importing from {__name__} is deprecated, please import via timm.layers", FutureWarning)
3
+ [rank3]:[W510 12:19:56.018259668 ProcessGroupNCCL.cpp:5138] Guessing device ID based on global rank. This can cause a hang if rank to GPU mapping is heterogeneous. You can specify device_id in init_process_group()
4
+ /gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/vjepa/utils.py:796: FutureWarning: `torch.cuda.amp.GradScaler(args...)` is deprecated. Please use `torch.amp.GradScaler('cuda', args...)` instead.
5
+ scaler = torch.cuda.amp.GradScaler() if mixed_precision else None
6
+ /gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/vjepa/train.py:955: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
7
+ with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
8
+ /gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/vjepa/train.py:955: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
9
+ with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
10
+ /gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/vjepa/train.py:955: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
11
+ with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
12
+ /gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/vjepa/train.py:955: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
13
+ with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
14
+ submitit ERROR (2026-05-10 16:19:44,648) - Submitted job triggered an exception
15
+ Traceback (most recent call last):
16
+ File "<frozen runpy>", line 198, in _run_module_as_main
17
+ File "<frozen runpy>", line 88, in _run_code
18
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/_submit.py", line 11, in <module>
19
+ submitit_main()
20
+ ~~~~~~~~~~~~~^^
21
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 76, in submitit_main
22
+ process_job(args.folder)
23
+ ~~~~~~~~~~~^^^^^^^^^^^^^
24
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 69, in process_job
25
+ raise error
26
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 55, in process_job
27
+ result = delayed.result()
28
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/utils.py", line 137, in result
29
+ self._result = self.function(*self.args, **self.kwargs)
30
+ ~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^
31
+ File "/gpfs/home2/dcanez/vd/app/main_distributed.py", line 102, in __call__
32
+ app_main(app, args=params, resume_preempt=resume_preempt)
33
+ ~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
34
+ File "/gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/scaffold.py", line 17, in main
35
+ return importlib.import_module(f"app.{app}.train").main(args=args, resume_preempt=resume_preempt)
36
+ ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
37
+ File "/gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/vjepa/train.py", line 1189, in main
38
+ assert not np.isnan(loss), "loss is nan"
39
+ ^^^^^^^^^^^^^^^^^^
40
+ AssertionError: loss is nan
41
+ [rank3]:[W510 16:19:45.445996938 ProcessGroupNCCL.cpp:1553] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())
vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/job_22614766/22614766_3_log.out ADDED
@@ -0,0 +1,276 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ submitit INFO (2026-05-10 12:14:34,050) - Starting with JobEnvironment(job_id=22614766, hostname=gcn80.local.snellius.surf.nl, local_rank=3(4), node=0(4), global_rank=3(16))
2
+ submitit INFO (2026-05-10 12:14:34,051) - Loading pickle: /scratch-shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/job_22614766/22614766_submitted.pkl
3
+ INFO:root:loaded pretrain params...
4
+ { 'app': 'vjepa',
5
+ 'cpus_per_task': 16,
6
+ 'data': { 'batch_size': 64,
7
+ 'crop_size': 224,
8
+ 'dataset_fpcs': [16, 16],
9
+ 'dataset_type': 'VideoDataset',
10
+ 'datasets': [ '/scratch-shared/dcanez/data/kinetics/k400/train.csv',
11
+ '/scratch-shared/dcanez/data/ssv2/train.csv'],
12
+ 'datasets_weights': [0.65, 0.35],
13
+ 'fps': 4,
14
+ 'num_workers': 10,
15
+ 'patch_size': 14,
16
+ 'persistent_workers': True,
17
+ 'pin_mem': True,
18
+ 'stage': [ { 'dest': 'kinetics_240',
19
+ 'format': 'targz_parts',
20
+ 'src': '/scratch-shared/dcanez/data/kinetics/k400/tars_240/'},
21
+ { 'dest': 'ssv2',
22
+ 'format': 'multipart_tar',
23
+ 'src': '/scratch-nvme/ml-datasets/something-something-v2/'}],
24
+ 'tubelet_size': 1},
25
+ 'data_aug': { 'auto_augment': False,
26
+ 'motion_shift': False,
27
+ 'random_resize_aspect_ratio': [0.75, 1.35],
28
+ 'random_resize_scale': [0.3, 1.0],
29
+ 'reprob': 0.0},
30
+ 'folder': '/scratch-shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable',
31
+ 'loss': {'loss_exp': 1.0},
32
+ 'mask': [ { 'aspect_ratio': [0.75, 1.5],
33
+ 'full_complement': False,
34
+ 'max_keep': None,
35
+ 'max_temporal_keep': 1.0,
36
+ 'num_blocks': 8,
37
+ 'spatial_scale': [0.15, 0.15],
38
+ 'temporal_scale': [1.0, 1.0]},
39
+ { 'aspect_ratio': [0.75, 1.5],
40
+ 'full_complement': False,
41
+ 'max_keep': None,
42
+ 'max_temporal_keep': 1.0,
43
+ 'num_blocks': 2,
44
+ 'spatial_scale': [0.7, 0.7],
45
+ 'temporal_scale': [1.0, 1.0]}],
46
+ 'mem_per_gpu': '180G',
47
+ 'meta': { 'dtype': 'bfloat16',
48
+ 'knn_eval_epoch0': False,
49
+ 'knn_eval_freq': 5,
50
+ 'knn_eval_presets': [ { 'config': { 'batch_size': 64,
51
+ 'dataset_train': '/scratch-shared/mdorkenw/ucf101/ucfTrainTestlist/train.csv',
52
+ 'dataset_val': '/scratch-shared/mdorkenw/ucf101/ucfTrainTestlist/val.csv',
53
+ 'eval_videos_per_class': 25,
54
+ 'num_workers': 8,
55
+ 'pool_type': 'slot_temporal_concat',
56
+ 'train_videos_per_class': 100},
57
+ 'preset': 'ucf101'},
58
+ { 'config': { 'batch_size': 64,
59
+ 'dataset_train': '/scratch-shared/mdorkenw/20bn-something-something-v2/something-something-v2-annotations/train_coarse10.csv',
60
+ 'dataset_val': '/scratch-shared/mdorkenw/20bn-something-something-v2/something-something-v2-annotations/val_coarse10.csv',
61
+ 'eval_videos_per_class': 100,
62
+ 'linear_probe': True,
63
+ 'num_workers': 8,
64
+ 'pool_type': 'slot_temporal_concat',
65
+ 'train_videos_per_class': 500},
66
+ 'preset': 'ssv2_coarse10'}],
67
+ 'load_checkpoint': True,
68
+ 'read_checkpoint': None,
69
+ 'save_every_freq': 5,
70
+ 'seed': 239,
71
+ 'use_sdpa': True,
72
+ 'use_wandb': True,
73
+ 'wandb_project': 'vjepa_ujepaside'},
74
+ 'metrics': {'sigreg': {}, 'std': {}},
75
+ 'model': { 'model_name': 'ujepaside_large_patch14_capi_lvd1689m',
76
+ 'pred_depth': 6,
77
+ 'pred_embed_dim': 384,
78
+ 'pred_num_heads': 12,
79
+ 'predictor': 'v2_cross',
80
+ 'st_causal': False,
81
+ 'st_drop_path': 0.2,
82
+ 'st_flex_enable': False,
83
+ 'st_layer_scale_init': 1e-05,
84
+ 'st_num_slots': 16,
85
+ 'st_slots_causal_within_frame': False,
86
+ 'target_kind': 'frozen_2d',
87
+ 'target_type': 'vit_large_patch14_capi.lvd1689m',
88
+ 'temporal_spacing': 1.0,
89
+ 'uniform_power': True,
90
+ 'use_activation_checkpointing': True,
91
+ 'use_mask_tokens': True,
92
+ 'use_rope': True,
93
+ 'use_sdpa': True,
94
+ 'zero_init_mask_tokens': True},
95
+ 'nodes': 4,
96
+ 'optimization': { 'clip_grad': 3.0,
97
+ 'ema': [0.99925, 0.99925],
98
+ 'epochs': 100,
99
+ 'final_lr': 0.0001,
100
+ 'final_weight_decay': 0.04,
101
+ 'ipe': 300,
102
+ 'ipe_scale': 1.0,
103
+ 'lr': 0.0005,
104
+ 'start_lr': 0.0001,
105
+ 'warmup': 10,
106
+ 'weight_decay': 0.04},
107
+ 'tasks_per_node': 4}
108
+ INFO:root:Running pre-training of app: vjepa
109
+ [INFO ][2026-05-10 12:16:53][app.vjepa.train ][main ] which_dtype='bfloat16'
110
+ [INFO ][2026-05-10 12:16:53][app.vjepa.train ][main ] Disabling persistent_workers (incompatible with KNN eval)
111
+ [INFO ][2026-05-10 12:16:53][app.vjepa.train ][main ] NCCL_SOCKET_IFNAME=eno
112
+ [INFO ][2026-05-10 12:16:56][app.vjepa.train ][main ] Initialized (rank/world-size) 3/16, tasks_per_node=4
113
+ [INFO ][2026-05-10 12:16:56][root ][stage_datasets ] [local_rank 3/4] Staging kinetics_240 (targz_parts)
114
+ [INFO ][2026-05-10 12:16:56][root ][_stage_targz_parts ] [rank 3] Extracting 25/103 tar.gz parts to /scratch-node/dcanez.22614766/kinetics_240
115
+ [INFO ][2026-05-10 12:17:10][root ][_stage_targz_parts ] [local_rank 3] Extracted 2/25 parts
116
+ [INFO ][2026-05-10 12:17:25][root ][_stage_targz_parts ] [local_rank 3] Extracted 4/25 parts
117
+ [INFO ][2026-05-10 12:17:39][root ][_stage_targz_parts ] [local_rank 3] Extracted 6/25 parts
118
+ [INFO ][2026-05-10 12:17:53][root ][_stage_targz_parts ] [local_rank 3] Extracted 8/25 parts
119
+ [INFO ][2026-05-10 12:18:07][root ][_stage_targz_parts ] [local_rank 3] Extracted 10/25 parts
120
+ [INFO ][2026-05-10 12:18:22][root ][_stage_targz_parts ] [local_rank 3] Extracted 12/25 parts
121
+ [INFO ][2026-05-10 12:18:36][root ][_stage_targz_parts ] [local_rank 3] Extracted 14/25 parts
122
+ [INFO ][2026-05-10 12:18:50][root ][_stage_targz_parts ] [local_rank 3] Extracted 16/25 parts
123
+ [INFO ][2026-05-10 12:19:04][root ][_stage_targz_parts ] [local_rank 3] Extracted 18/25 parts
124
+ [INFO ][2026-05-10 12:19:19][root ][_stage_targz_parts ] [local_rank 3] Extracted 20/25 parts
125
+ [INFO ][2026-05-10 12:19:33][root ][_stage_targz_parts ] [local_rank 3] Extracted 22/25 parts
126
+ [INFO ][2026-05-10 12:19:47][root ][_stage_targz_parts ] [local_rank 3] Extracted 24/25 parts
127
+ [INFO ][2026-05-10 12:19:54][root ][_stage_targz_parts ] [local_rank 3] Extracted 25/25 parts
128
+ [INFO ][2026-05-10 12:19:54][root ][stage_datasets ] [local_rank 3/4] Staging ssv2 (multipart_tar)
129
+ [INFO ][2026-05-10 12:21:03][app.vjepa.train ][main ] Staged datasets: ['/scratch-node/dcanez.22614766/kinetics_240/train.csv', '/scratch-node/dcanez.22614766/ssv2/train.csv']
130
+ [INFO ][2026-05-10 12:23:31][experiments.stmodels.vision_transformers_v3][vit_large_patch14_capi_lvd1689m] Loading pretrained weights for vit_large_patch14_capi_lvd1689m from https://dl.fbaipublicfiles.com/capi/capi_vitl14_lvd.pth
131
+ [INFO ][2026-05-10 12:23:57][root ][_build_frozen_2d_target ] Loaded frozen_2d target from timm: vit_large_patch14_capi.lvd1689m
132
+ [INFO ][2026-05-10 12:23:57][root ][init_video_model ] ViTMultiSeqWrapper(
133
+ (backbone): UJEPAside(
134
+ (patch_embed): PatchEmbed3D(
135
+ (proj): Conv3d(3, 1024, kernel_size=(1, 14, 14), stride=(1, 14, 14))
136
+ )
137
+ (rope): CAPI2DRoPE()
138
+ (blocks): ModuleList(
139
+ (0-23): 24 x Block(
140
+ (norm1): RMSNorm((1024,), eps=1e-05, elementwise_affine=True)
141
+ (rope_impl): CAPI2DRoPE()
142
+ (norm2): RMSNorm((1024,), eps=1e-05, elementwise_affine=True)
143
+ (drop_path1): Identity()
144
+ (drop_path2): Identity()
145
+ (attn): Attention(
146
+ (qkv): Linear(in_features=1024, out_features=3072, bias=False)
147
+ (attn_drop): Dropout(p=0.0, inplace=False)
148
+ (proj): Linear(in_features=1024, out_features=1024, bias=False)
149
+ (proj_drop): Dropout(p=0.0, inplace=False)
150
+ (rope_impl): CAPI2DRoPE()
151
+ )
152
+ (mlp): MLP(
153
+ (fc1): Linear(in_features=1024, out_features=4096, bias=False)
154
+ (act): GELU(approximate='none')
155
+ (fc2): Linear(in_features=4096, out_features=1024, bias=False)
156
+ (drop): Dropout(p=0.0, inplace=False)
157
+ )
158
+ )
159
+ )
160
+ (norm): RMSNorm((1024,), eps=1e-05, elementwise_affine=False)
161
+ (st_blocks): ModuleList(
162
+ (0-23): 24 x Block(
163
+ (residual1): EfficientResidual(
164
+ (norm): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
165
+ (fn): Attention(
166
+ (q_proj): Linear(in_features=1024, out_features=1024, bias=False)
167
+ (k_proj): Linear(in_features=1024, out_features=1024, bias=False)
168
+ (v_proj): Linear(in_features=1024, out_features=1024, bias=False)
169
+ (proj): Linear(in_features=1024, out_features=1024, bias=False)
170
+ (rope_impl): CAPI3DRoPE()
171
+ )
172
+ )
173
+ (residual2): EfficientResidual(
174
+ (norm): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
175
+ (fn): MLP(
176
+ (fc1): Linear(in_features=1024, out_features=4096, bias=False)
177
+ (act): GELU(approximate='none')
178
+ (fc2): Linear(in_features=4096, out_features=1024, bias=False)
179
+ (drop): Dropout(p=0.0, inplace=False)
180
+ )
181
+ )
182
+ )
183
+ )
184
+ (st_rope): CAPI3DRoPE()
185
+ )
186
+ )
187
+ [INFO ][2026-05-10 12:23:57][root ][init_video_model ] UJEPAsidePredictorMultiSeqWrapper(
188
+ (backbone): PredictorV2(
189
+ (predictor_embed): Linear(in_features=1024, out_features=384, bias=True)
190
+ (mask_tokens): ParameterList(
191
+ (0): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
192
+ (1): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
193
+ (2): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
194
+ (3): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
195
+ )
196
+ (predictor_blocks): ModuleList(
197
+ (0-5): 6 x Block(
198
+ (residual1): EfficientResidual(
199
+ (norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
200
+ (fn): Attention(
201
+ (q_proj): Linear(in_features=384, out_features=384, bias=False)
202
+ (k_proj): Linear(in_features=384, out_features=384, bias=False)
203
+ (v_proj): Linear(in_features=384, out_features=384, bias=False)
204
+ (proj): Linear(in_features=384, out_features=384, bias=False)
205
+ (rope): Rope()
206
+ )
207
+ )
208
+ (residual2): EfficientResidual(
209
+ (norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
210
+ (fn): MLP(
211
+ (fc1): Linear(in_features=384, out_features=1536, bias=False)
212
+ (act): GELU(approximate='none')
213
+ (fc2): Linear(in_features=1536, out_features=384, bias=False)
214
+ (drop): Dropout(p=0.0, inplace=False)
215
+ )
216
+ )
217
+ )
218
+ )
219
+ (predictor_norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
220
+ (predictor_proj): Linear(in_features=384, out_features=1024, bias=True)
221
+ )
222
+ )
223
+ [INFO ][2026-05-10 12:23:57][root ][init_video_model ] MultiSeqWrapper(
224
+ (backbone): Frozen2DTargetWrapper(
225
+ (backbone): Eva(
226
+ (patch_embed): PatchEmbed(
227
+ (proj): Conv2d(3, 1024, kernel_size=(14, 14), stride=(14, 14))
228
+ (norm): Identity()
229
+ )
230
+ (pos_drop): Dropout(p=0.0, inplace=False)
231
+ (norm_pre): Identity()
232
+ (blocks): ModuleList(
233
+ (0-23): 24 x EvaBlock(
234
+ (norm1): RMSNorm((1024,), eps=1e-05, elementwise_affine=True)
235
+ (attn): EvaAttention(
236
+ (qkv): Linear(in_features=1024, out_features=3072, bias=False)
237
+ (q_norm): Identity()
238
+ (k_norm): Identity()
239
+ (attn_drop): Dropout(p=0.0, inplace=False)
240
+ (norm): Identity()
241
+ (proj): Linear(in_features=1024, out_features=1024, bias=True)
242
+ (proj_drop): Dropout(p=0.0, inplace=False)
243
+ )
244
+ (drop_path1): Identity()
245
+ (norm2): RMSNorm((1024,), eps=1e-05, elementwise_affine=True)
246
+ (mlp): Mlp(
247
+ (fc1): Linear(in_features=1024, out_features=4096, bias=True)
248
+ (act): GELU(approximate='none')
249
+ (drop1): Dropout(p=0.0, inplace=False)
250
+ (norm): Identity()
251
+ (fc2): Linear(in_features=4096, out_features=1024, bias=True)
252
+ (drop2): Dropout(p=0.0, inplace=False)
253
+ )
254
+ (drop_path2): Identity()
255
+ )
256
+ )
257
+ (norm): RMSNorm((1024,), eps=1e-05, elementwise_affine=True)
258
+ (fc_norm): Identity()
259
+ (head_drop): Dropout(p=0.0, inplace=False)
260
+ (head): Identity()
261
+ (rope): _CapiPatchRoPE()
262
+ )
263
+ )
264
+ )
265
+ [INFO ][2026-05-10 12:23:57][root ][init_video_model ] Encoder number of parameters: 302399488
266
+ [INFO ][2026-05-10 12:23:57][root ][init_video_model ] Predictor number of parameters: 11416192
267
+ [INFO ][2026-05-10 12:23:57][root ][init_video_model ] Target encoder number of parameters: 0
268
+ [INFO ][2026-05-10 12:24:07][root ][make_videodataset ] VideoDataset dataset created
269
+ [INFO ][2026-05-10 12:24:07][WeightedSampler ][__init__ ] Using DistributedWeightedSampler with rank 3 / 16
270
+ [INFO ][2026-05-10 12:24:07][root ][make_videodataset ] VideoDataset unsupervised data loader created
271
+ [INFO ][2026-05-10 12:24:07][app.vjepa.train ][main ] iterations per epoch/dataset length: 300/399
272
+ [INFO ][2026-05-10 12:24:07][app.vjepa.train ][main ] Wrapping models in DDP (rank 3)...
273
+ [INFO ][2026-05-10 12:24:08][app.vjepa.train ][main ] Initializing loader...
274
+ submitit ERROR (2026-05-10 16:19:44,648) - Submitted job triggered an exception
275
+ [ERROR ][2026-05-10 16:19:44][submitit ][process_job ] Submitted job triggered an exception
276
+ [ERROR ][2026-05-10 16:19:44][app.vjepa.train ][_save_crash ] Saved crash log to /scratch-shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/crash_rank3.log
vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/job_22614766/22614766_4_log.err ADDED
@@ -0,0 +1,180 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/timm/models/layers/__init__.py:49: FutureWarning: Importing from timm.models.layers is deprecated, please import via timm.layers
2
+ warnings.warn(f"Importing from {__name__} is deprecated, please import via timm.layers", FutureWarning)
3
+ [rank4]:[W510 12:20:52.615848849 ProcessGroupNCCL.cpp:5138] Guessing device ID based on global rank. This can cause a hang if rank to GPU mapping is heterogeneous. You can specify device_id in init_process_group()
4
+ /gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/vjepa/utils.py:796: FutureWarning: `torch.cuda.amp.GradScaler(args...)` is deprecated. Please use `torch.amp.GradScaler('cuda', args...)` instead.
5
+ scaler = torch.cuda.amp.GradScaler() if mixed_precision else None
6
+ /gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/vjepa/train.py:955: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
7
+ with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
8
+ [rank4]:[W510 16:19:49.193287972 TCPStore.cpp:125] [c10d] recvValue failed on SocketImpl(fd=3, addr=[gcn82.local.snellius.surf.nl]:43904, remote=[gcn80.local.snellius.surf.nl]:37129): Failed to recv, got 0 bytes. Connection was likely closed. Did the remote server shutdown or crash?
9
+ Exception raised from recvBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:682 (most recent call first):
10
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x14e4ff737fdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
11
+ frame #1: <unknown function> + 0x6a3325d (0x14e5438c625d in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
12
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x273 (0x14e5438c41f3 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
13
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x14e5009a6eec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
14
+ frame #4: <unknown function> + 0xed164 (0x14e63fe10164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
15
+ frame #5: <unknown function> + 0x8a19a (0x14e65088219a in /lib64/libc.so.6)
16
+ frame #6: <unknown function> + 0x10f100 (0x14e650907100 in /lib64/libc.so.6)
17
+
18
+ [rank4]:[W510 16:19:49.250674818 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 4] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Failed to recv, got 0 bytes. Connection was likely closed. Did the remote server shutdown or crash?
19
+ [rank4]:[W510 16:19:50.250831182 TCPStore.cpp:106] [c10d] sendBytes failed on SocketImpl(fd=3, addr=[gcn82.local.snellius.surf.nl]:43904, remote=[gcn80.local.snellius.surf.nl]:37129): Broken pipe
20
+ Exception raised from sendBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:653 (most recent call first):
21
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x14e4ff737fdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
22
+ frame #1: <unknown function> + 0x6a326d1 (0x14e5438c56d1 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
23
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x24d (0x14e5438c41cd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
24
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x14e5009a6eec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
25
+ frame #4: <unknown function> + 0xed164 (0x14e63fe10164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
26
+ frame #5: <unknown function> + 0x8a19a (0x14e65088219a in /lib64/libc.so.6)
27
+ frame #6: <unknown function> + 0x10f100 (0x14e650907100 in /lib64/libc.so.6)
28
+
29
+ [rank4]:[W510 16:19:50.254593415 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 4] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Broken pipe
30
+ [rank4]:[W510 16:19:51.254722169 TCPStore.cpp:106] [c10d] sendBytes failed on SocketImpl(fd=3, addr=[gcn82.local.snellius.surf.nl]:43904, remote=[gcn80.local.snellius.surf.nl]:37129): Broken pipe
31
+ Exception raised from sendBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:653 (most recent call first):
32
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x14e4ff737fdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
33
+ frame #1: <unknown function> + 0x6a326d1 (0x14e5438c56d1 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
34
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x24d (0x14e5438c41cd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
35
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x14e5009a6eec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
36
+ frame #4: <unknown function> + 0xed164 (0x14e63fe10164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
37
+ frame #5: <unknown function> + 0x8a19a (0x14e65088219a in /lib64/libc.so.6)
38
+ frame #6: <unknown function> + 0x10f100 (0x14e650907100 in /lib64/libc.so.6)
39
+
40
+ [rank4]:[W510 16:19:51.256917393 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 4] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Broken pipe
41
+ [rank4]:[W510 16:19:52.257044298 TCPStore.cpp:106] [c10d] sendBytes failed on SocketImpl(fd=3, addr=[gcn82.local.snellius.surf.nl]:43904, remote=[gcn80.local.snellius.surf.nl]:37129): Broken pipe
42
+ Exception raised from sendBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:653 (most recent call first):
43
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x14e4ff737fdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
44
+ frame #1: <unknown function> + 0x6a326d1 (0x14e5438c56d1 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
45
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x24d (0x14e5438c41cd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
46
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x14e5009a6eec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
47
+ frame #4: <unknown function> + 0xed164 (0x14e63fe10164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
48
+ frame #5: <unknown function> + 0x8a19a (0x14e65088219a in /lib64/libc.so.6)
49
+ frame #6: <unknown function> + 0x10f100 (0x14e650907100 in /lib64/libc.so.6)
50
+
51
+ [rank4]:[W510 16:19:52.259240942 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 4] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Broken pipe
52
+ [rank4]:[W510 16:19:53.259363187 TCPStore.cpp:106] [c10d] sendBytes failed on SocketImpl(fd=3, addr=[gcn82.local.snellius.surf.nl]:43904, remote=[gcn80.local.snellius.surf.nl]:37129): Broken pipe
53
+ Exception raised from sendBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:653 (most recent call first):
54
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x14e4ff737fdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
55
+ frame #1: <unknown function> + 0x6a326d1 (0x14e5438c56d1 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
56
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x24d (0x14e5438c41cd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
57
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x14e5009a6eec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
58
+ frame #4: <unknown function> + 0xed164 (0x14e63fe10164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
59
+ frame #5: <unknown function> + 0x8a19a (0x14e65088219a in /lib64/libc.so.6)
60
+ frame #6: <unknown function> + 0x10f100 (0x14e650907100 in /lib64/libc.so.6)
61
+
62
+ [rank4]:[W510 16:19:53.261586531 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 4] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Broken pipe
63
+ [rank4]:[W510 16:19:54.261712946 TCPStore.cpp:106] [c10d] sendBytes failed on SocketImpl(fd=3, addr=[gcn82.local.snellius.surf.nl]:43904, remote=[gcn80.local.snellius.surf.nl]:37129): Broken pipe
64
+ Exception raised from sendBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:653 (most recent call first):
65
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x14e4ff737fdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
66
+ frame #1: <unknown function> + 0x6a326d1 (0x14e5438c56d1 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
67
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x24d (0x14e5438c41cd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
68
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x14e5009a6eec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
69
+ frame #4: <unknown function> + 0xed164 (0x14e63fe10164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
70
+ frame #5: <unknown function> + 0x8a19a (0x14e65088219a in /lib64/libc.so.6)
71
+ frame #6: <unknown function> + 0x10f100 (0x14e650907100 in /lib64/libc.so.6)
72
+
73
+ [rank4]:[W510 16:19:54.263938389 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 4] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Broken pipe
74
+ [rank4]:[W510 16:19:55.264077644 TCPStore.cpp:106] [c10d] sendBytes failed on SocketImpl(fd=3, addr=[gcn82.local.snellius.surf.nl]:43904, remote=[gcn80.local.snellius.surf.nl]:37129): Broken pipe
75
+ Exception raised from sendBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:653 (most recent call first):
76
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x14e4ff737fdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
77
+ frame #1: <unknown function> + 0x6a326d1 (0x14e5438c56d1 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
78
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x24d (0x14e5438c41cd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
79
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x14e5009a6eec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
80
+ frame #4: <unknown function> + 0xed164 (0x14e63fe10164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
81
+ frame #5: <unknown function> + 0x8a19a (0x14e65088219a in /lib64/libc.so.6)
82
+ frame #6: <unknown function> + 0x10f100 (0x14e650907100 in /lib64/libc.so.6)
83
+
84
+ [rank4]:[W510 16:19:55.266358357 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 4] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Broken pipe
85
+ [rank4]:[W510 16:19:56.266502282 TCPStore.cpp:106] [c10d] sendBytes failed on SocketImpl(fd=3, addr=[gcn82.local.snellius.surf.nl]:43904, remote=[gcn80.local.snellius.surf.nl]:37129): Broken pipe
86
+ Exception raised from sendBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:653 (most recent call first):
87
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x14e4ff737fdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
88
+ frame #1: <unknown function> + 0x6a326d1 (0x14e5438c56d1 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
89
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x24d (0x14e5438c41cd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
90
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x14e5009a6eec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
91
+ frame #4: <unknown function> + 0xed164 (0x14e63fe10164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
92
+ frame #5: <unknown function> + 0x8a19a (0x14e65088219a in /lib64/libc.so.6)
93
+ frame #6: <unknown function> + 0x10f100 (0x14e650907100 in /lib64/libc.so.6)
94
+
95
+ [rank4]:[W510 16:19:56.268782175 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 4] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Broken pipe
96
+ [rank4]:[W510 16:19:57.268924191 TCPStore.cpp:106] [c10d] sendBytes failed on SocketImpl(fd=3, addr=[gcn82.local.snellius.surf.nl]:43904, remote=[gcn80.local.snellius.surf.nl]:37129): Broken pipe
97
+ Exception raised from sendBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:653 (most recent call first):
98
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x14e4ff737fdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
99
+ frame #1: <unknown function> + 0x6a326d1 (0x14e5438c56d1 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
100
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x24d (0x14e5438c41cd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
101
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x14e5009a6eec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
102
+ frame #4: <unknown function> + 0xed164 (0x14e63fe10164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
103
+ frame #5: <unknown function> + 0x8a19a (0x14e65088219a in /lib64/libc.so.6)
104
+ frame #6: <unknown function> + 0x10f100 (0x14e650907100 in /lib64/libc.so.6)
105
+
106
+ [rank4]:[W510 16:19:57.271270323 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 4] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Broken pipe
107
+ [rank4]:[W510 16:19:58.271408538 TCPStore.cpp:106] [c10d] sendBytes failed on SocketImpl(fd=3, addr=[gcn82.local.snellius.surf.nl]:43904, remote=[gcn80.local.snellius.surf.nl]:37129): Broken pipe
108
+ Exception raised from sendBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:653 (most recent call first):
109
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x14e4ff737fdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
110
+ frame #1: <unknown function> + 0x6a326d1 (0x14e5438c56d1 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
111
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x24d (0x14e5438c41cd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
112
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x14e5009a6eec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
113
+ frame #4: <unknown function> + 0xed164 (0x14e63fe10164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
114
+ frame #5: <unknown function> + 0x8a19a (0x14e65088219a in /lib64/libc.so.6)
115
+ frame #6: <unknown function> + 0x10f100 (0x14e650907100 in /lib64/libc.so.6)
116
+
117
+ [rank4]:[W510 16:19:58.273590672 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 4] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Broken pipe
118
+ [rank4]:[W510 16:19:59.273737087 TCPStore.cpp:106] [c10d] sendBytes failed on SocketImpl(fd=3, addr=[gcn82.local.snellius.surf.nl]:43904, remote=[gcn80.local.snellius.surf.nl]:37129): Broken pipe
119
+ Exception raised from sendBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:653 (most recent call first):
120
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x14e4ff737fdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
121
+ frame #1: <unknown function> + 0x6a326d1 (0x14e5438c56d1 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
122
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x24d (0x14e5438c41cd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
123
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x14e5009a6eec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
124
+ frame #4: <unknown function> + 0xed164 (0x14e63fe10164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
125
+ frame #5: <unknown function> + 0x8a19a (0x14e65088219a in /lib64/libc.so.6)
126
+ frame #6: <unknown function> + 0x10f100 (0x14e650907100 in /lib64/libc.so.6)
127
+
128
+ [rank4]:[W510 16:19:59.276033990 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 4] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Broken pipe
129
+ submitit WARNING (2026-05-10 16:20:00,106) - Bypassing signal SIGTERM
130
+ submitit WARNING (2026-05-10 16:20:00,118) - Bypassing signal SIGCONT
131
+ submitit ERROR (2026-05-10 16:20:00,118) - Submitted job triggered an exception
132
+ Traceback (most recent call last):
133
+ File "<frozen runpy>", line 198, in _run_module_as_main
134
+ File "<frozen runpy>", line 88, in _run_code
135
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/_submit.py", line 11, in <module>
136
+ submitit_main()
137
+ ~~~~~~~~~~~~~^^
138
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 76, in submitit_main
139
+ process_job(args.folder)
140
+ ~~~~~~~~~~~^^^^^^^^^^^^^
141
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 69, in process_job
142
+ raise error
143
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 55, in process_job
144
+ result = delayed.result()
145
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/utils.py", line 137, in result
146
+ self._result = self.function(*self.args, **self.kwargs)
147
+ ~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^
148
+ File "/gpfs/home2/dcanez/vd/app/main_distributed.py", line 102, in __call__
149
+ app_main(app, args=params, resume_preempt=resume_preempt)
150
+ ~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
151
+ File "/gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/scaffold.py", line 17, in main
152
+ return importlib.import_module(f"app.{app}.train").main(args=args, resume_preempt=resume_preempt)
153
+ ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
154
+ File "/gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/vjepa/train.py", line 1189, in main
155
+ assert not np.isnan(loss), "loss is nan"
156
+ ^^^^^^^^^^^^^^^^^^
157
+ AssertionError: loss is nan
158
+ [rank4]:[W510 16:20:00.276182975 TCPStore.cpp:106] [c10d] sendBytes failed on SocketImpl(fd=3, addr=[gcn82.local.snellius.surf.nl]:43904, remote=[gcn80.local.snellius.surf.nl]:37129): Broken pipe
159
+ Exception raised from sendBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:653 (most recent call first):
160
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x14e4ff737fdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
161
+ frame #1: <unknown function> + 0x6a326d1 (0x14e5438c56d1 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
162
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x24d (0x14e5438c41cd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
163
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x14e5009a6eec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
164
+ frame #4: <unknown function> + 0xed164 (0x14e63fe10164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
165
+ frame #5: <unknown function> + 0x8a19a (0x14e65088219a in /lib64/libc.so.6)
166
+ frame #6: <unknown function> + 0x10f100 (0x14e650907100 in /lib64/libc.so.6)
167
+
168
+ [rank4]:[W510 16:20:00.278511248 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 4] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Broken pipe
169
+ [rank4]:[W510 16:20:00.717128829 ProcessGroupNCCL.cpp:1553] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())
170
+ [rank4]:[W510 16:20:01.278653903 TCPStore.cpp:106] [c10d] sendBytes failed on SocketImpl(fd=3, addr=[gcn82.local.snellius.surf.nl]:43904, remote=[gcn80.local.snellius.surf.nl]:37129): Broken pipe
171
+ Exception raised from sendBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:653 (most recent call first):
172
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x14e4ff737fdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
173
+ frame #1: <unknown function> + 0x6a326d1 (0x14e5438c56d1 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
174
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x24d (0x14e5438c41cd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
175
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x14e5009a6eec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
176
+ frame #4: <unknown function> + 0xed164 (0x14e63fe10164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
177
+ frame #5: <unknown function> + 0x8a19a (0x14e65088219a in /lib64/libc.so.6)
178
+ frame #6: <unknown function> + 0x10f100 (0x14e650907100 in /lib64/libc.so.6)
179
+
180
+ [rank4]:[W510 16:20:01.280841577 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 4] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Broken pipe
vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/job_22614766/22614766_4_log.out ADDED
@@ -0,0 +1,284 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ submitit INFO (2026-05-10 12:14:34,049) - Starting with JobEnvironment(job_id=22614766, hostname=gcn82.local.snellius.surf.nl, local_rank=0(4), node=1(4), global_rank=4(16))
2
+ submitit INFO (2026-05-10 12:14:34,049) - Loading pickle: /scratch-shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/job_22614766/22614766_submitted.pkl
3
+ INFO:root:loaded pretrain params...
4
+ { 'app': 'vjepa',
5
+ 'cpus_per_task': 16,
6
+ 'data': { 'batch_size': 64,
7
+ 'crop_size': 224,
8
+ 'dataset_fpcs': [16, 16],
9
+ 'dataset_type': 'VideoDataset',
10
+ 'datasets': [ '/scratch-shared/dcanez/data/kinetics/k400/train.csv',
11
+ '/scratch-shared/dcanez/data/ssv2/train.csv'],
12
+ 'datasets_weights': [0.65, 0.35],
13
+ 'fps': 4,
14
+ 'num_workers': 10,
15
+ 'patch_size': 14,
16
+ 'persistent_workers': True,
17
+ 'pin_mem': True,
18
+ 'stage': [ { 'dest': 'kinetics_240',
19
+ 'format': 'targz_parts',
20
+ 'src': '/scratch-shared/dcanez/data/kinetics/k400/tars_240/'},
21
+ { 'dest': 'ssv2',
22
+ 'format': 'multipart_tar',
23
+ 'src': '/scratch-nvme/ml-datasets/something-something-v2/'}],
24
+ 'tubelet_size': 1},
25
+ 'data_aug': { 'auto_augment': False,
26
+ 'motion_shift': False,
27
+ 'random_resize_aspect_ratio': [0.75, 1.35],
28
+ 'random_resize_scale': [0.3, 1.0],
29
+ 'reprob': 0.0},
30
+ 'folder': '/scratch-shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable',
31
+ 'loss': {'loss_exp': 1.0},
32
+ 'mask': [ { 'aspect_ratio': [0.75, 1.5],
33
+ 'full_complement': False,
34
+ 'max_keep': None,
35
+ 'max_temporal_keep': 1.0,
36
+ 'num_blocks': 8,
37
+ 'spatial_scale': [0.15, 0.15],
38
+ 'temporal_scale': [1.0, 1.0]},
39
+ { 'aspect_ratio': [0.75, 1.5],
40
+ 'full_complement': False,
41
+ 'max_keep': None,
42
+ 'max_temporal_keep': 1.0,
43
+ 'num_blocks': 2,
44
+ 'spatial_scale': [0.7, 0.7],
45
+ 'temporal_scale': [1.0, 1.0]}],
46
+ 'mem_per_gpu': '180G',
47
+ 'meta': { 'dtype': 'bfloat16',
48
+ 'knn_eval_epoch0': False,
49
+ 'knn_eval_freq': 5,
50
+ 'knn_eval_presets': [ { 'config': { 'batch_size': 64,
51
+ 'dataset_train': '/scratch-shared/mdorkenw/ucf101/ucfTrainTestlist/train.csv',
52
+ 'dataset_val': '/scratch-shared/mdorkenw/ucf101/ucfTrainTestlist/val.csv',
53
+ 'eval_videos_per_class': 25,
54
+ 'num_workers': 8,
55
+ 'pool_type': 'slot_temporal_concat',
56
+ 'train_videos_per_class': 100},
57
+ 'preset': 'ucf101'},
58
+ { 'config': { 'batch_size': 64,
59
+ 'dataset_train': '/scratch-shared/mdorkenw/20bn-something-something-v2/something-something-v2-annotations/train_coarse10.csv',
60
+ 'dataset_val': '/scratch-shared/mdorkenw/20bn-something-something-v2/something-something-v2-annotations/val_coarse10.csv',
61
+ 'eval_videos_per_class': 100,
62
+ 'linear_probe': True,
63
+ 'num_workers': 8,
64
+ 'pool_type': 'slot_temporal_concat',
65
+ 'train_videos_per_class': 500},
66
+ 'preset': 'ssv2_coarse10'}],
67
+ 'load_checkpoint': True,
68
+ 'read_checkpoint': None,
69
+ 'save_every_freq': 5,
70
+ 'seed': 239,
71
+ 'use_sdpa': True,
72
+ 'use_wandb': True,
73
+ 'wandb_project': 'vjepa_ujepaside'},
74
+ 'metrics': {'sigreg': {}, 'std': {}},
75
+ 'model': { 'model_name': 'ujepaside_large_patch14_capi_lvd1689m',
76
+ 'pred_depth': 6,
77
+ 'pred_embed_dim': 384,
78
+ 'pred_num_heads': 12,
79
+ 'predictor': 'v2_cross',
80
+ 'st_causal': False,
81
+ 'st_drop_path': 0.2,
82
+ 'st_flex_enable': False,
83
+ 'st_layer_scale_init': 1e-05,
84
+ 'st_num_slots': 16,
85
+ 'st_slots_causal_within_frame': False,
86
+ 'target_kind': 'frozen_2d',
87
+ 'target_type': 'vit_large_patch14_capi.lvd1689m',
88
+ 'temporal_spacing': 1.0,
89
+ 'uniform_power': True,
90
+ 'use_activation_checkpointing': True,
91
+ 'use_mask_tokens': True,
92
+ 'use_rope': True,
93
+ 'use_sdpa': True,
94
+ 'zero_init_mask_tokens': True},
95
+ 'nodes': 4,
96
+ 'optimization': { 'clip_grad': 3.0,
97
+ 'ema': [0.99925, 0.99925],
98
+ 'epochs': 100,
99
+ 'final_lr': 0.0001,
100
+ 'final_weight_decay': 0.04,
101
+ 'ipe': 300,
102
+ 'ipe_scale': 1.0,
103
+ 'lr': 0.0005,
104
+ 'start_lr': 0.0001,
105
+ 'warmup': 10,
106
+ 'weight_decay': 0.04},
107
+ 'tasks_per_node': 4}
108
+ INFO:root:Running pre-training of app: vjepa
109
+ [INFO ][2026-05-10 12:16:53][app.vjepa.train ][main ] which_dtype='bfloat16'
110
+ [INFO ][2026-05-10 12:16:53][app.vjepa.train ][main ] Disabling persistent_workers (incompatible with KNN eval)
111
+ [INFO ][2026-05-10 12:16:53][app.vjepa.train ][main ] NCCL_SOCKET_IFNAME=eno
112
+ [INFO ][2026-05-10 12:16:55][app.vjepa.train ][main ] Initialized (rank/world-size) 4/16, tasks_per_node=4
113
+ [INFO ][2026-05-10 12:16:55][root ][stage_datasets ] [local_rank 0/4] Staging kinetics_240 (targz_parts)
114
+ [INFO ][2026-05-10 12:16:55][root ][_stage_targz_parts ] [rank 0] Extracting 26/103 tar.gz parts to /scratch-node/dcanez.22614766/kinetics_240
115
+ [INFO ][2026-05-10 12:17:09][root ][_stage_targz_parts ] [local_rank 0] Extracted 2/26 parts
116
+ [INFO ][2026-05-10 12:17:23][root ][_stage_targz_parts ] [local_rank 0] Extracted 4/26 parts
117
+ [INFO ][2026-05-10 12:17:37][root ][_stage_targz_parts ] [local_rank 0] Extracted 6/26 parts
118
+ [INFO ][2026-05-10 12:17:51][root ][_stage_targz_parts ] [local_rank 0] Extracted 8/26 parts
119
+ [INFO ][2026-05-10 12:18:05][root ][_stage_targz_parts ] [local_rank 0] Extracted 10/26 parts
120
+ [INFO ][2026-05-10 12:18:19][root ][_stage_targz_parts ] [local_rank 0] Extracted 12/26 parts
121
+ [INFO ][2026-05-10 12:18:33][root ][_stage_targz_parts ] [local_rank 0] Extracted 14/26 parts
122
+ [INFO ][2026-05-10 12:18:47][root ][_stage_targz_parts ] [local_rank 0] Extracted 16/26 parts
123
+ [INFO ][2026-05-10 12:19:02][root ][_stage_targz_parts ] [local_rank 0] Extracted 18/26 parts
124
+ [INFO ][2026-05-10 12:19:16][root ][_stage_targz_parts ] [local_rank 0] Extracted 20/26 parts
125
+ [INFO ][2026-05-10 12:19:30][root ][_stage_targz_parts ] [local_rank 0] Extracted 22/26 parts
126
+ [INFO ][2026-05-10 12:19:44][root ][_stage_targz_parts ] [local_rank 0] Extracted 24/26 parts
127
+ [INFO ][2026-05-10 12:19:58][root ][_stage_targz_parts ] [local_rank 0] Extracted 26/26 parts
128
+ [INFO ][2026-05-10 12:19:58][root ][stage_datasets ] [local_rank 0/4] Staging ssv2 (multipart_tar)
129
+ [INFO ][2026-05-10 12:19:58][root ][_stage_multipart_tar ] [rank 0] Extracting multipart tar (2 files) to /scratch-node/dcanez.22614766/ssv2
130
+ [INFO ][2026-05-10 12:21:02][root ][stage_datasets ] Data staging completed in 246.4s (4.1min)
131
+ [INFO ][2026-05-10 12:21:02][root ][_rewrite_csv ] Wrote local CSV: /scratch-node/dcanez.22614766/kinetics_240/train.csv (239789 entries)
132
+ [INFO ][2026-05-10 12:21:03][root ][_rewrite_csv ] Wrote local CSV: /scratch-node/dcanez.22614766/ssv2/train.csv (168913 entries)
133
+ [INFO ][2026-05-10 12:21:03][app.vjepa.train ][main ] Staged datasets: ['/scratch-node/dcanez.22614766/kinetics_240/train.csv', '/scratch-node/dcanez.22614766/ssv2/train.csv']
134
+ [INFO ][2026-05-10 12:23:31][experiments.stmodels.vision_transformers_v3][vit_large_patch14_capi_lvd1689m] Loading pretrained weights for vit_large_patch14_capi_lvd1689m from https://dl.fbaipublicfiles.com/capi/capi_vitl14_lvd.pth
135
+ [INFO ][2026-05-10 12:23:56][root ][_build_frozen_2d_target ] Loaded frozen_2d target from timm: vit_large_patch14_capi.lvd1689m
136
+ [INFO ][2026-05-10 12:23:56][root ][init_video_model ] ViTMultiSeqWrapper(
137
+ (backbone): UJEPAside(
138
+ (patch_embed): PatchEmbed3D(
139
+ (proj): Conv3d(3, 1024, kernel_size=(1, 14, 14), stride=(1, 14, 14))
140
+ )
141
+ (rope): CAPI2DRoPE()
142
+ (blocks): ModuleList(
143
+ (0-23): 24 x Block(
144
+ (norm1): RMSNorm((1024,), eps=1e-05, elementwise_affine=True)
145
+ (rope_impl): CAPI2DRoPE()
146
+ (norm2): RMSNorm((1024,), eps=1e-05, elementwise_affine=True)
147
+ (drop_path1): Identity()
148
+ (drop_path2): Identity()
149
+ (attn): Attention(
150
+ (qkv): Linear(in_features=1024, out_features=3072, bias=False)
151
+ (attn_drop): Dropout(p=0.0, inplace=False)
152
+ (proj): Linear(in_features=1024, out_features=1024, bias=False)
153
+ (proj_drop): Dropout(p=0.0, inplace=False)
154
+ (rope_impl): CAPI2DRoPE()
155
+ )
156
+ (mlp): MLP(
157
+ (fc1): Linear(in_features=1024, out_features=4096, bias=False)
158
+ (act): GELU(approximate='none')
159
+ (fc2): Linear(in_features=4096, out_features=1024, bias=False)
160
+ (drop): Dropout(p=0.0, inplace=False)
161
+ )
162
+ )
163
+ )
164
+ (norm): RMSNorm((1024,), eps=1e-05, elementwise_affine=False)
165
+ (st_blocks): ModuleList(
166
+ (0-23): 24 x Block(
167
+ (residual1): EfficientResidual(
168
+ (norm): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
169
+ (fn): Attention(
170
+ (q_proj): Linear(in_features=1024, out_features=1024, bias=False)
171
+ (k_proj): Linear(in_features=1024, out_features=1024, bias=False)
172
+ (v_proj): Linear(in_features=1024, out_features=1024, bias=False)
173
+ (proj): Linear(in_features=1024, out_features=1024, bias=False)
174
+ (rope_impl): CAPI3DRoPE()
175
+ )
176
+ )
177
+ (residual2): EfficientResidual(
178
+ (norm): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
179
+ (fn): MLP(
180
+ (fc1): Linear(in_features=1024, out_features=4096, bias=False)
181
+ (act): GELU(approximate='none')
182
+ (fc2): Linear(in_features=4096, out_features=1024, bias=False)
183
+ (drop): Dropout(p=0.0, inplace=False)
184
+ )
185
+ )
186
+ )
187
+ )
188
+ (st_rope): CAPI3DRoPE()
189
+ )
190
+ )
191
+ [INFO ][2026-05-10 12:23:56][root ][init_video_model ] UJEPAsidePredictorMultiSeqWrapper(
192
+ (backbone): PredictorV2(
193
+ (predictor_embed): Linear(in_features=1024, out_features=384, bias=True)
194
+ (mask_tokens): ParameterList(
195
+ (0): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
196
+ (1): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
197
+ (2): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
198
+ (3): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
199
+ )
200
+ (predictor_blocks): ModuleList(
201
+ (0-5): 6 x Block(
202
+ (residual1): EfficientResidual(
203
+ (norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
204
+ (fn): Attention(
205
+ (q_proj): Linear(in_features=384, out_features=384, bias=False)
206
+ (k_proj): Linear(in_features=384, out_features=384, bias=False)
207
+ (v_proj): Linear(in_features=384, out_features=384, bias=False)
208
+ (proj): Linear(in_features=384, out_features=384, bias=False)
209
+ (rope): Rope()
210
+ )
211
+ )
212
+ (residual2): EfficientResidual(
213
+ (norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
214
+ (fn): MLP(
215
+ (fc1): Linear(in_features=384, out_features=1536, bias=False)
216
+ (act): GELU(approximate='none')
217
+ (fc2): Linear(in_features=1536, out_features=384, bias=False)
218
+ (drop): Dropout(p=0.0, inplace=False)
219
+ )
220
+ )
221
+ )
222
+ )
223
+ (predictor_norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
224
+ (predictor_proj): Linear(in_features=384, out_features=1024, bias=True)
225
+ )
226
+ )
227
+ [INFO ][2026-05-10 12:23:56][root ][init_video_model ] MultiSeqWrapper(
228
+ (backbone): Frozen2DTargetWrapper(
229
+ (backbone): Eva(
230
+ (patch_embed): PatchEmbed(
231
+ (proj): Conv2d(3, 1024, kernel_size=(14, 14), stride=(14, 14))
232
+ (norm): Identity()
233
+ )
234
+ (pos_drop): Dropout(p=0.0, inplace=False)
235
+ (norm_pre): Identity()
236
+ (blocks): ModuleList(
237
+ (0-23): 24 x EvaBlock(
238
+ (norm1): RMSNorm((1024,), eps=1e-05, elementwise_affine=True)
239
+ (attn): EvaAttention(
240
+ (qkv): Linear(in_features=1024, out_features=3072, bias=False)
241
+ (q_norm): Identity()
242
+ (k_norm): Identity()
243
+ (attn_drop): Dropout(p=0.0, inplace=False)
244
+ (norm): Identity()
245
+ (proj): Linear(in_features=1024, out_features=1024, bias=True)
246
+ (proj_drop): Dropout(p=0.0, inplace=False)
247
+ )
248
+ (drop_path1): Identity()
249
+ (norm2): RMSNorm((1024,), eps=1e-05, elementwise_affine=True)
250
+ (mlp): Mlp(
251
+ (fc1): Linear(in_features=1024, out_features=4096, bias=True)
252
+ (act): GELU(approximate='none')
253
+ (drop1): Dropout(p=0.0, inplace=False)
254
+ (norm): Identity()
255
+ (fc2): Linear(in_features=4096, out_features=1024, bias=True)
256
+ (drop2): Dropout(p=0.0, inplace=False)
257
+ )
258
+ (drop_path2): Identity()
259
+ )
260
+ )
261
+ (norm): RMSNorm((1024,), eps=1e-05, elementwise_affine=True)
262
+ (fc_norm): Identity()
263
+ (head_drop): Dropout(p=0.0, inplace=False)
264
+ (head): Identity()
265
+ (rope): _CapiPatchRoPE()
266
+ )
267
+ )
268
+ )
269
+ [INFO ][2026-05-10 12:23:56][root ][init_video_model ] Encoder number of parameters: 302399488
270
+ [INFO ][2026-05-10 12:23:56][root ][init_video_model ] Predictor number of parameters: 11416192
271
+ [INFO ][2026-05-10 12:23:56][root ][init_video_model ] Target encoder number of parameters: 0
272
+ [INFO ][2026-05-10 12:24:07][root ][make_videodataset ] VideoDataset dataset created
273
+ [INFO ][2026-05-10 12:24:07][WeightedSampler ][__init__ ] Using DistributedWeightedSampler with rank 4 / 16
274
+ [INFO ][2026-05-10 12:24:07][root ][make_videodataset ] VideoDataset unsupervised data loader created
275
+ [INFO ][2026-05-10 12:24:07][app.vjepa.train ][main ] iterations per epoch/dataset length: 300/399
276
+ [INFO ][2026-05-10 12:24:07][app.vjepa.train ][main ] Wrapping models in DDP (rank 4)...
277
+ [INFO ][2026-05-10 12:24:09][app.vjepa.train ][main ] Initializing loader...
278
+ submitit WARNING (2026-05-10 16:20:00,106) - Bypassing signal SIGTERM
279
+ [WARNING ][2026-05-10 16:20:00][submitit ][bypass ] Bypassing signal SIGTERM
280
+ submitit WARNING (2026-05-10 16:20:00,118) - Bypassing signal SIGCONT
281
+ [WARNING ][2026-05-10 16:20:00][submitit ][bypass ] Bypassing signal SIGCONT
282
+ submitit ERROR (2026-05-10 16:20:00,118) - Submitted job triggered an exception
283
+ [ERROR ][2026-05-10 16:20:00][submitit ][process_job ] Submitted job triggered an exception
284
+ [ERROR ][2026-05-10 16:20:00][app.vjepa.train ][_save_crash ] Saved crash log to /scratch-shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/crash_rank4.log
vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/job_22614766/22614766_5_log.err ADDED
@@ -0,0 +1,158 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/timm/models/layers/__init__.py:49: FutureWarning: Importing from timm.models.layers is deprecated, please import via timm.layers
2
+ warnings.warn(f"Importing from {__name__} is deprecated, please import via timm.layers", FutureWarning)
3
+ [rank5]:[W510 12:19:58.114218222 ProcessGroupNCCL.cpp:5138] Guessing device ID based on global rank. This can cause a hang if rank to GPU mapping is heterogeneous. You can specify device_id in init_process_group()
4
+ /gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/vjepa/utils.py:796: FutureWarning: `torch.cuda.amp.GradScaler(args...)` is deprecated. Please use `torch.amp.GradScaler('cuda', args...)` instead.
5
+ scaler = torch.cuda.amp.GradScaler() if mixed_precision else None
6
+ /gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/vjepa/train.py:955: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
7
+ with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
8
+ [rank5]:[W510 16:19:49.194232775 TCPStore.cpp:125] [c10d] recvValue failed on SocketImpl(fd=3, addr=[gcn82.local.snellius.surf.nl]:43910, remote=[gcn80.local.snellius.surf.nl]:37129): Failed to recv, got 0 bytes. Connection was likely closed. Did the remote server shutdown or crash?
9
+ Exception raised from recvBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:682 (most recent call first):
10
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x14bc8bbfcfdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
11
+ frame #1: <unknown function> + 0x6a3325d (0x14bccfd8b25d in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
12
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x273 (0x14bccfd891f3 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
13
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x14bc8ce6beec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
14
+ frame #4: <unknown function> + 0xed164 (0x14bdcc2d5164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
15
+ frame #5: <unknown function> + 0x8a19a (0x14bddcd4719a in /lib64/libc.so.6)
16
+ frame #6: <unknown function> + 0x10f100 (0x14bddcdcc100 in /lib64/libc.so.6)
17
+
18
+ [rank5]:[W510 16:19:49.250678388 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 5] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Failed to recv, got 0 bytes. Connection was likely closed. Did the remote server shutdown or crash?
19
+ [rank5]:[W510 16:19:50.250824642 TCPStore.cpp:106] [c10d] sendBytes failed on SocketImpl(fd=3, addr=[gcn82.local.snellius.surf.nl]:43910, remote=[gcn80.local.snellius.surf.nl]:37129): Broken pipe
20
+ Exception raised from sendBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:653 (most recent call first):
21
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x14bc8bbfcfdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
22
+ frame #1: <unknown function> + 0x6a326d1 (0x14bccfd8a6d1 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
23
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x24d (0x14bccfd891cd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
24
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x14bc8ce6beec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
25
+ frame #4: <unknown function> + 0xed164 (0x14bdcc2d5164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
26
+ frame #5: <unknown function> + 0x8a19a (0x14bddcd4719a in /lib64/libc.so.6)
27
+ frame #6: <unknown function> + 0x10f100 (0x14bddcdcc100 in /lib64/libc.so.6)
28
+
29
+ [rank5]:[W510 16:19:50.254503495 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 5] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Broken pipe
30
+ [rank5]:[W510 16:19:51.254625920 TCPStore.cpp:106] [c10d] sendBytes failed on SocketImpl(fd=3, addr=[gcn82.local.snellius.surf.nl]:43910, remote=[gcn80.local.snellius.surf.nl]:37129): Broken pipe
31
+ Exception raised from sendBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:653 (most recent call first):
32
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x14bc8bbfcfdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
33
+ frame #1: <unknown function> + 0x6a326d1 (0x14bccfd8a6d1 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
34
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x24d (0x14bccfd891cd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
35
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x14bc8ce6beec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
36
+ frame #4: <unknown function> + 0xed164 (0x14bdcc2d5164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
37
+ frame #5: <unknown function> + 0x8a19a (0x14bddcd4719a in /lib64/libc.so.6)
38
+ frame #6: <unknown function> + 0x10f100 (0x14bddcdcc100 in /lib64/libc.so.6)
39
+
40
+ [rank5]:[W510 16:19:51.256900143 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 5] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Broken pipe
41
+ [rank5]:[W510 16:19:52.257043938 TCPStore.cpp:106] [c10d] sendBytes failed on SocketImpl(fd=3, addr=[gcn82.local.snellius.surf.nl]:43910, remote=[gcn80.local.snellius.surf.nl]:37129): Broken pipe
42
+ Exception raised from sendBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:653 (most recent call first):
43
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x14bc8bbfcfdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
44
+ frame #1: <unknown function> + 0x6a326d1 (0x14bccfd8a6d1 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
45
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x24d (0x14bccfd891cd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
46
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x14bc8ce6beec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
47
+ frame #4: <unknown function> + 0xed164 (0x14bdcc2d5164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
48
+ frame #5: <unknown function> + 0x8a19a (0x14bddcd4719a in /lib64/libc.so.6)
49
+ frame #6: <unknown function> + 0x10f100 (0x14bddcdcc100 in /lib64/libc.so.6)
50
+
51
+ [rank5]:[W510 16:19:52.259270682 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 5] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Broken pipe
52
+ [rank5]:[W510 16:19:53.259384977 TCPStore.cpp:106] [c10d] sendBytes failed on SocketImpl(fd=3, addr=[gcn82.local.snellius.surf.nl]:43910, remote=[gcn80.local.snellius.surf.nl]:37129): Broken pipe
53
+ Exception raised from sendBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:653 (most recent call first):
54
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x14bc8bbfcfdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
55
+ frame #1: <unknown function> + 0x6a326d1 (0x14bccfd8a6d1 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
56
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x24d (0x14bccfd891cd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
57
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x14bc8ce6beec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
58
+ frame #4: <unknown function> + 0xed164 (0x14bdcc2d5164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
59
+ frame #5: <unknown function> + 0x8a19a (0x14bddcd4719a in /lib64/libc.so.6)
60
+ frame #6: <unknown function> + 0x10f100 (0x14bddcdcc100 in /lib64/libc.so.6)
61
+
62
+ [rank5]:[W510 16:19:53.261599871 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 5] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Broken pipe
63
+ [rank5]:[W510 16:19:54.261725616 TCPStore.cpp:106] [c10d] sendBytes failed on SocketImpl(fd=3, addr=[gcn82.local.snellius.surf.nl]:43910, remote=[gcn80.local.snellius.surf.nl]:37129): Broken pipe
64
+ Exception raised from sendBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:653 (most recent call first):
65
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x14bc8bbfcfdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
66
+ frame #1: <unknown function> + 0x6a326d1 (0x14bccfd8a6d1 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
67
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x24d (0x14bccfd891cd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
68
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x14bc8ce6beec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
69
+ frame #4: <unknown function> + 0xed164 (0x14bdcc2d5164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
70
+ frame #5: <unknown function> + 0x8a19a (0x14bddcd4719a in /lib64/libc.so.6)
71
+ frame #6: <unknown function> + 0x10f100 (0x14bddcdcc100 in /lib64/libc.so.6)
72
+
73
+ [rank5]:[W510 16:19:54.264077558 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 5] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Broken pipe
74
+ [rank5]:[W510 16:19:55.264218893 TCPStore.cpp:106] [c10d] sendBytes failed on SocketImpl(fd=3, addr=[gcn82.local.snellius.surf.nl]:43910, remote=[gcn80.local.snellius.surf.nl]:37129): Broken pipe
75
+ Exception raised from sendBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:653 (most recent call first):
76
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x14bc8bbfcfdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
77
+ frame #1: <unknown function> + 0x6a326d1 (0x14bccfd8a6d1 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
78
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x24d (0x14bccfd891cd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
79
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x14bc8ce6beec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
80
+ frame #4: <unknown function> + 0xed164 (0x14bdcc2d5164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
81
+ frame #5: <unknown function> + 0x8a19a (0x14bddcd4719a in /lib64/libc.so.6)
82
+ frame #6: <unknown function> + 0x10f100 (0x14bddcdcc100 in /lib64/libc.so.6)
83
+
84
+ [rank5]:[W510 16:19:55.266477376 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 5] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Broken pipe
85
+ [rank5]:[W510 16:19:56.266606102 TCPStore.cpp:106] [c10d] sendBytes failed on SocketImpl(fd=3, addr=[gcn82.local.snellius.surf.nl]:43910, remote=[gcn80.local.snellius.surf.nl]:37129): Broken pipe
86
+ Exception raised from sendBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:653 (most recent call first):
87
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x14bc8bbfcfdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
88
+ frame #1: <unknown function> + 0x6a326d1 (0x14bccfd8a6d1 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
89
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x24d (0x14bccfd891cd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
90
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x14bc8ce6beec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
91
+ frame #4: <unknown function> + 0xed164 (0x14bdcc2d5164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
92
+ frame #5: <unknown function> + 0x8a19a (0x14bddcd4719a in /lib64/libc.so.6)
93
+ frame #6: <unknown function> + 0x10f100 (0x14bddcdcc100 in /lib64/libc.so.6)
94
+
95
+ [rank5]:[W510 16:19:56.268802325 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 5] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Broken pipe
96
+ [rank5]:[W510 16:19:57.268933681 TCPStore.cpp:106] [c10d] sendBytes failed on SocketImpl(fd=3, addr=[gcn82.local.snellius.surf.nl]:43910, remote=[gcn80.local.snellius.surf.nl]:37129): Broken pipe
97
+ Exception raised from sendBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:653 (most recent call first):
98
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x14bc8bbfcfdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
99
+ frame #1: <unknown function> + 0x6a326d1 (0x14bccfd8a6d1 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
100
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x24d (0x14bccfd891cd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
101
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x14bc8ce6beec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
102
+ frame #4: <unknown function> + 0xed164 (0x14bdcc2d5164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
103
+ frame #5: <unknown function> + 0x8a19a (0x14bddcd4719a in /lib64/libc.so.6)
104
+ frame #6: <unknown function> + 0x10f100 (0x14bddcdcc100 in /lib64/libc.so.6)
105
+
106
+ [rank5]:[W510 16:19:57.271200114 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 5] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Broken pipe
107
+ submitit WARNING (2026-05-10 16:19:58,222) - Bypassing signal SIGTERM
108
+ submitit WARNING (2026-05-10 16:19:58,230) - Bypassing signal SIGCONT
109
+ submitit ERROR (2026-05-10 16:19:58,230) - Submitted job triggered an exception
110
+ Traceback (most recent call last):
111
+ File "<frozen runpy>", line 198, in _run_module_as_main
112
+ File "<frozen runpy>", line 88, in _run_code
113
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/_submit.py", line 11, in <module>
114
+ submitit_main()
115
+ ~~~~~~~~~~~~~^^
116
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 76, in submitit_main
117
+ process_job(args.folder)
118
+ ~~~~~~~~~~~^^^^^^^^^^^^^
119
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 69, in process_job
120
+ raise error
121
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 55, in process_job
122
+ result = delayed.result()
123
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/utils.py", line 137, in result
124
+ self._result = self.function(*self.args, **self.kwargs)
125
+ ~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^
126
+ File "/gpfs/home2/dcanez/vd/app/main_distributed.py", line 102, in __call__
127
+ app_main(app, args=params, resume_preempt=resume_preempt)
128
+ ~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
129
+ File "/gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/scaffold.py", line 17, in main
130
+ return importlib.import_module(f"app.{app}.train").main(args=args, resume_preempt=resume_preempt)
131
+ ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
132
+ File "/gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/vjepa/train.py", line 1189, in main
133
+ assert not np.isnan(loss), "loss is nan"
134
+ ^^^^^^^^^^^^^^^^^^
135
+ AssertionError: loss is nan
136
+ [rank5]:[W510 16:19:58.271351009 TCPStore.cpp:106] [c10d] sendBytes failed on SocketImpl(fd=3, addr=[gcn82.local.snellius.surf.nl]:43910, remote=[gcn80.local.snellius.surf.nl]:37129): Broken pipe
137
+ Exception raised from sendBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:653 (most recent call first):
138
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x14bc8bbfcfdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
139
+ frame #1: <unknown function> + 0x6a326d1 (0x14bccfd8a6d1 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
140
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x24d (0x14bccfd891cd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
141
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x14bc8ce6beec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
142
+ frame #4: <unknown function> + 0xed164 (0x14bdcc2d5164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
143
+ frame #5: <unknown function> + 0x8a19a (0x14bddcd4719a in /lib64/libc.so.6)
144
+ frame #6: <unknown function> + 0x10f100 (0x14bddcdcc100 in /lib64/libc.so.6)
145
+
146
+ [rank5]:[W510 16:19:58.273646922 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 5] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Broken pipe
147
+ [rank5]:[W510 16:19:58.864755947 ProcessGroupNCCL.cpp:1553] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())
148
+ [rank5]:[W510 16:19:59.273790727 TCPStore.cpp:106] [c10d] sendBytes failed on SocketImpl(fd=3, addr=[gcn82.local.snellius.surf.nl]:43910, remote=[gcn80.local.snellius.surf.nl]:37129): Broken pipe
149
+ Exception raised from sendBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:653 (most recent call first):
150
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x14bc8bbfcfdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
151
+ frame #1: <unknown function> + 0x6a326d1 (0x14bccfd8a6d1 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
152
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x24d (0x14bccfd891cd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
153
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x14bc8ce6beec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
154
+ frame #4: <unknown function> + 0xed164 (0x14bdcc2d5164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
155
+ frame #5: <unknown function> + 0x8a19a (0x14bddcd4719a in /lib64/libc.so.6)
156
+ frame #6: <unknown function> + 0x10f100 (0x14bddcdcc100 in /lib64/libc.so.6)
157
+
158
+ [rank5]:[W510 16:19:59.276053390 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 5] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Broken pipe
vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/job_22614766/22614766_5_log.out ADDED
@@ -0,0 +1,280 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ submitit INFO (2026-05-10 12:14:34,049) - Starting with JobEnvironment(job_id=22614766, hostname=gcn82.local.snellius.surf.nl, local_rank=1(4), node=1(4), global_rank=5(16))
2
+ submitit INFO (2026-05-10 12:14:34,049) - Loading pickle: /scratch-shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/job_22614766/22614766_submitted.pkl
3
+ INFO:root:loaded pretrain params...
4
+ { 'app': 'vjepa',
5
+ 'cpus_per_task': 16,
6
+ 'data': { 'batch_size': 64,
7
+ 'crop_size': 224,
8
+ 'dataset_fpcs': [16, 16],
9
+ 'dataset_type': 'VideoDataset',
10
+ 'datasets': [ '/scratch-shared/dcanez/data/kinetics/k400/train.csv',
11
+ '/scratch-shared/dcanez/data/ssv2/train.csv'],
12
+ 'datasets_weights': [0.65, 0.35],
13
+ 'fps': 4,
14
+ 'num_workers': 10,
15
+ 'patch_size': 14,
16
+ 'persistent_workers': True,
17
+ 'pin_mem': True,
18
+ 'stage': [ { 'dest': 'kinetics_240',
19
+ 'format': 'targz_parts',
20
+ 'src': '/scratch-shared/dcanez/data/kinetics/k400/tars_240/'},
21
+ { 'dest': 'ssv2',
22
+ 'format': 'multipart_tar',
23
+ 'src': '/scratch-nvme/ml-datasets/something-something-v2/'}],
24
+ 'tubelet_size': 1},
25
+ 'data_aug': { 'auto_augment': False,
26
+ 'motion_shift': False,
27
+ 'random_resize_aspect_ratio': [0.75, 1.35],
28
+ 'random_resize_scale': [0.3, 1.0],
29
+ 'reprob': 0.0},
30
+ 'folder': '/scratch-shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable',
31
+ 'loss': {'loss_exp': 1.0},
32
+ 'mask': [ { 'aspect_ratio': [0.75, 1.5],
33
+ 'full_complement': False,
34
+ 'max_keep': None,
35
+ 'max_temporal_keep': 1.0,
36
+ 'num_blocks': 8,
37
+ 'spatial_scale': [0.15, 0.15],
38
+ 'temporal_scale': [1.0, 1.0]},
39
+ { 'aspect_ratio': [0.75, 1.5],
40
+ 'full_complement': False,
41
+ 'max_keep': None,
42
+ 'max_temporal_keep': 1.0,
43
+ 'num_blocks': 2,
44
+ 'spatial_scale': [0.7, 0.7],
45
+ 'temporal_scale': [1.0, 1.0]}],
46
+ 'mem_per_gpu': '180G',
47
+ 'meta': { 'dtype': 'bfloat16',
48
+ 'knn_eval_epoch0': False,
49
+ 'knn_eval_freq': 5,
50
+ 'knn_eval_presets': [ { 'config': { 'batch_size': 64,
51
+ 'dataset_train': '/scratch-shared/mdorkenw/ucf101/ucfTrainTestlist/train.csv',
52
+ 'dataset_val': '/scratch-shared/mdorkenw/ucf101/ucfTrainTestlist/val.csv',
53
+ 'eval_videos_per_class': 25,
54
+ 'num_workers': 8,
55
+ 'pool_type': 'slot_temporal_concat',
56
+ 'train_videos_per_class': 100},
57
+ 'preset': 'ucf101'},
58
+ { 'config': { 'batch_size': 64,
59
+ 'dataset_train': '/scratch-shared/mdorkenw/20bn-something-something-v2/something-something-v2-annotations/train_coarse10.csv',
60
+ 'dataset_val': '/scratch-shared/mdorkenw/20bn-something-something-v2/something-something-v2-annotations/val_coarse10.csv',
61
+ 'eval_videos_per_class': 100,
62
+ 'linear_probe': True,
63
+ 'num_workers': 8,
64
+ 'pool_type': 'slot_temporal_concat',
65
+ 'train_videos_per_class': 500},
66
+ 'preset': 'ssv2_coarse10'}],
67
+ 'load_checkpoint': True,
68
+ 'read_checkpoint': None,
69
+ 'save_every_freq': 5,
70
+ 'seed': 239,
71
+ 'use_sdpa': True,
72
+ 'use_wandb': True,
73
+ 'wandb_project': 'vjepa_ujepaside'},
74
+ 'metrics': {'sigreg': {}, 'std': {}},
75
+ 'model': { 'model_name': 'ujepaside_large_patch14_capi_lvd1689m',
76
+ 'pred_depth': 6,
77
+ 'pred_embed_dim': 384,
78
+ 'pred_num_heads': 12,
79
+ 'predictor': 'v2_cross',
80
+ 'st_causal': False,
81
+ 'st_drop_path': 0.2,
82
+ 'st_flex_enable': False,
83
+ 'st_layer_scale_init': 1e-05,
84
+ 'st_num_slots': 16,
85
+ 'st_slots_causal_within_frame': False,
86
+ 'target_kind': 'frozen_2d',
87
+ 'target_type': 'vit_large_patch14_capi.lvd1689m',
88
+ 'temporal_spacing': 1.0,
89
+ 'uniform_power': True,
90
+ 'use_activation_checkpointing': True,
91
+ 'use_mask_tokens': True,
92
+ 'use_rope': True,
93
+ 'use_sdpa': True,
94
+ 'zero_init_mask_tokens': True},
95
+ 'nodes': 4,
96
+ 'optimization': { 'clip_grad': 3.0,
97
+ 'ema': [0.99925, 0.99925],
98
+ 'epochs': 100,
99
+ 'final_lr': 0.0001,
100
+ 'final_weight_decay': 0.04,
101
+ 'ipe': 300,
102
+ 'ipe_scale': 1.0,
103
+ 'lr': 0.0005,
104
+ 'start_lr': 0.0001,
105
+ 'warmup': 10,
106
+ 'weight_decay': 0.04},
107
+ 'tasks_per_node': 4}
108
+ INFO:root:Running pre-training of app: vjepa
109
+ [INFO ][2026-05-10 12:16:53][app.vjepa.train ][main ] which_dtype='bfloat16'
110
+ [INFO ][2026-05-10 12:16:53][app.vjepa.train ][main ] Disabling persistent_workers (incompatible with KNN eval)
111
+ [INFO ][2026-05-10 12:16:53][app.vjepa.train ][main ] NCCL_SOCKET_IFNAME=eno
112
+ [INFO ][2026-05-10 12:16:56][app.vjepa.train ][main ] Initialized (rank/world-size) 5/16, tasks_per_node=4
113
+ [INFO ][2026-05-10 12:16:56][root ][stage_datasets ] [local_rank 1/4] Staging kinetics_240 (targz_parts)
114
+ [INFO ][2026-05-10 12:16:56][root ][_stage_targz_parts ] [rank 1] Extracting 26/103 tar.gz parts to /scratch-node/dcanez.22614766/kinetics_240
115
+ [INFO ][2026-05-10 12:17:08][root ][_stage_targz_parts ] [local_rank 1] Extracted 2/26 parts
116
+ [INFO ][2026-05-10 12:17:22][root ][_stage_targz_parts ] [local_rank 1] Extracted 4/26 parts
117
+ [INFO ][2026-05-10 12:17:36][root ][_stage_targz_parts ] [local_rank 1] Extracted 6/26 parts
118
+ [INFO ][2026-05-10 12:17:50][root ][_stage_targz_parts ] [local_rank 1] Extracted 8/26 parts
119
+ [INFO ][2026-05-10 12:18:04][root ][_stage_targz_parts ] [local_rank 1] Extracted 10/26 parts
120
+ [INFO ][2026-05-10 12:18:19][root ][_stage_targz_parts ] [local_rank 1] Extracted 12/26 parts
121
+ [INFO ][2026-05-10 12:18:32][root ][_stage_targz_parts ] [local_rank 1] Extracted 14/26 parts
122
+ [INFO ][2026-05-10 12:18:46][root ][_stage_targz_parts ] [local_rank 1] Extracted 16/26 parts
123
+ [INFO ][2026-05-10 12:19:00][root ][_stage_targz_parts ] [local_rank 1] Extracted 18/26 parts
124
+ [INFO ][2026-05-10 12:19:15][root ][_stage_targz_parts ] [local_rank 1] Extracted 20/26 parts
125
+ [INFO ][2026-05-10 12:19:29][root ][_stage_targz_parts ] [local_rank 1] Extracted 22/26 parts
126
+ [INFO ][2026-05-10 12:19:43][root ][_stage_targz_parts ] [local_rank 1] Extracted 24/26 parts
127
+ [INFO ][2026-05-10 12:19:57][root ][_stage_targz_parts ] [local_rank 1] Extracted 26/26 parts
128
+ [INFO ][2026-05-10 12:19:57][root ][stage_datasets ] [local_rank 1/4] Staging ssv2 (multipart_tar)
129
+ [INFO ][2026-05-10 12:21:03][app.vjepa.train ][main ] Staged datasets: ['/scratch-node/dcanez.22614766/kinetics_240/train.csv', '/scratch-node/dcanez.22614766/ssv2/train.csv']
130
+ [INFO ][2026-05-10 12:23:31][experiments.stmodels.vision_transformers_v3][vit_large_patch14_capi_lvd1689m] Loading pretrained weights for vit_large_patch14_capi_lvd1689m from https://dl.fbaipublicfiles.com/capi/capi_vitl14_lvd.pth
131
+ [INFO ][2026-05-10 12:23:56][root ][_build_frozen_2d_target ] Loaded frozen_2d target from timm: vit_large_patch14_capi.lvd1689m
132
+ [INFO ][2026-05-10 12:23:56][root ][init_video_model ] ViTMultiSeqWrapper(
133
+ (backbone): UJEPAside(
134
+ (patch_embed): PatchEmbed3D(
135
+ (proj): Conv3d(3, 1024, kernel_size=(1, 14, 14), stride=(1, 14, 14))
136
+ )
137
+ (rope): CAPI2DRoPE()
138
+ (blocks): ModuleList(
139
+ (0-23): 24 x Block(
140
+ (norm1): RMSNorm((1024,), eps=1e-05, elementwise_affine=True)
141
+ (rope_impl): CAPI2DRoPE()
142
+ (norm2): RMSNorm((1024,), eps=1e-05, elementwise_affine=True)
143
+ (drop_path1): Identity()
144
+ (drop_path2): Identity()
145
+ (attn): Attention(
146
+ (qkv): Linear(in_features=1024, out_features=3072, bias=False)
147
+ (attn_drop): Dropout(p=0.0, inplace=False)
148
+ (proj): Linear(in_features=1024, out_features=1024, bias=False)
149
+ (proj_drop): Dropout(p=0.0, inplace=False)
150
+ (rope_impl): CAPI2DRoPE()
151
+ )
152
+ (mlp): MLP(
153
+ (fc1): Linear(in_features=1024, out_features=4096, bias=False)
154
+ (act): GELU(approximate='none')
155
+ (fc2): Linear(in_features=4096, out_features=1024, bias=False)
156
+ (drop): Dropout(p=0.0, inplace=False)
157
+ )
158
+ )
159
+ )
160
+ (norm): RMSNorm((1024,), eps=1e-05, elementwise_affine=False)
161
+ (st_blocks): ModuleList(
162
+ (0-23): 24 x Block(
163
+ (residual1): EfficientResidual(
164
+ (norm): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
165
+ (fn): Attention(
166
+ (q_proj): Linear(in_features=1024, out_features=1024, bias=False)
167
+ (k_proj): Linear(in_features=1024, out_features=1024, bias=False)
168
+ (v_proj): Linear(in_features=1024, out_features=1024, bias=False)
169
+ (proj): Linear(in_features=1024, out_features=1024, bias=False)
170
+ (rope_impl): CAPI3DRoPE()
171
+ )
172
+ )
173
+ (residual2): EfficientResidual(
174
+ (norm): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
175
+ (fn): MLP(
176
+ (fc1): Linear(in_features=1024, out_features=4096, bias=False)
177
+ (act): GELU(approximate='none')
178
+ (fc2): Linear(in_features=4096, out_features=1024, bias=False)
179
+ (drop): Dropout(p=0.0, inplace=False)
180
+ )
181
+ )
182
+ )
183
+ )
184
+ (st_rope): CAPI3DRoPE()
185
+ )
186
+ )
187
+ [INFO ][2026-05-10 12:23:56][root ][init_video_model ] UJEPAsidePredictorMultiSeqWrapper(
188
+ (backbone): PredictorV2(
189
+ (predictor_embed): Linear(in_features=1024, out_features=384, bias=True)
190
+ (mask_tokens): ParameterList(
191
+ (0): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
192
+ (1): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
193
+ (2): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
194
+ (3): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
195
+ )
196
+ (predictor_blocks): ModuleList(
197
+ (0-5): 6 x Block(
198
+ (residual1): EfficientResidual(
199
+ (norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
200
+ (fn): Attention(
201
+ (q_proj): Linear(in_features=384, out_features=384, bias=False)
202
+ (k_proj): Linear(in_features=384, out_features=384, bias=False)
203
+ (v_proj): Linear(in_features=384, out_features=384, bias=False)
204
+ (proj): Linear(in_features=384, out_features=384, bias=False)
205
+ (rope): Rope()
206
+ )
207
+ )
208
+ (residual2): EfficientResidual(
209
+ (norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
210
+ (fn): MLP(
211
+ (fc1): Linear(in_features=384, out_features=1536, bias=False)
212
+ (act): GELU(approximate='none')
213
+ (fc2): Linear(in_features=1536, out_features=384, bias=False)
214
+ (drop): Dropout(p=0.0, inplace=False)
215
+ )
216
+ )
217
+ )
218
+ )
219
+ (predictor_norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
220
+ (predictor_proj): Linear(in_features=384, out_features=1024, bias=True)
221
+ )
222
+ )
223
+ [INFO ][2026-05-10 12:23:56][root ][init_video_model ] MultiSeqWrapper(
224
+ (backbone): Frozen2DTargetWrapper(
225
+ (backbone): Eva(
226
+ (patch_embed): PatchEmbed(
227
+ (proj): Conv2d(3, 1024, kernel_size=(14, 14), stride=(14, 14))
228
+ (norm): Identity()
229
+ )
230
+ (pos_drop): Dropout(p=0.0, inplace=False)
231
+ (norm_pre): Identity()
232
+ (blocks): ModuleList(
233
+ (0-23): 24 x EvaBlock(
234
+ (norm1): RMSNorm((1024,), eps=1e-05, elementwise_affine=True)
235
+ (attn): EvaAttention(
236
+ (qkv): Linear(in_features=1024, out_features=3072, bias=False)
237
+ (q_norm): Identity()
238
+ (k_norm): Identity()
239
+ (attn_drop): Dropout(p=0.0, inplace=False)
240
+ (norm): Identity()
241
+ (proj): Linear(in_features=1024, out_features=1024, bias=True)
242
+ (proj_drop): Dropout(p=0.0, inplace=False)
243
+ )
244
+ (drop_path1): Identity()
245
+ (norm2): RMSNorm((1024,), eps=1e-05, elementwise_affine=True)
246
+ (mlp): Mlp(
247
+ (fc1): Linear(in_features=1024, out_features=4096, bias=True)
248
+ (act): GELU(approximate='none')
249
+ (drop1): Dropout(p=0.0, inplace=False)
250
+ (norm): Identity()
251
+ (fc2): Linear(in_features=4096, out_features=1024, bias=True)
252
+ (drop2): Dropout(p=0.0, inplace=False)
253
+ )
254
+ (drop_path2): Identity()
255
+ )
256
+ )
257
+ (norm): RMSNorm((1024,), eps=1e-05, elementwise_affine=True)
258
+ (fc_norm): Identity()
259
+ (head_drop): Dropout(p=0.0, inplace=False)
260
+ (head): Identity()
261
+ (rope): _CapiPatchRoPE()
262
+ )
263
+ )
264
+ )
265
+ [INFO ][2026-05-10 12:23:56][root ][init_video_model ] Encoder number of parameters: 302399488
266
+ [INFO ][2026-05-10 12:23:56][root ][init_video_model ] Predictor number of parameters: 11416192
267
+ [INFO ][2026-05-10 12:23:56][root ][init_video_model ] Target encoder number of parameters: 0
268
+ [INFO ][2026-05-10 12:24:07][root ][make_videodataset ] VideoDataset dataset created
269
+ [INFO ][2026-05-10 12:24:07][WeightedSampler ][__init__ ] Using DistributedWeightedSampler with rank 5 / 16
270
+ [INFO ][2026-05-10 12:24:07][root ][make_videodataset ] VideoDataset unsupervised data loader created
271
+ [INFO ][2026-05-10 12:24:07][app.vjepa.train ][main ] iterations per epoch/dataset length: 300/399
272
+ [INFO ][2026-05-10 12:24:07][app.vjepa.train ][main ] Wrapping models in DDP (rank 5)...
273
+ [INFO ][2026-05-10 12:24:09][app.vjepa.train ][main ] Initializing loader...
274
+ submitit WARNING (2026-05-10 16:19:58,222) - Bypassing signal SIGTERM
275
+ [WARNING ][2026-05-10 16:19:58][submitit ][bypass ] Bypassing signal SIGTERM
276
+ submitit WARNING (2026-05-10 16:19:58,230) - Bypassing signal SIGCONT
277
+ [WARNING ][2026-05-10 16:19:58][submitit ][bypass ] Bypassing signal SIGCONT
278
+ submitit ERROR (2026-05-10 16:19:58,230) - Submitted job triggered an exception
279
+ [ERROR ][2026-05-10 16:19:58][submitit ][process_job ] Submitted job triggered an exception
280
+ [ERROR ][2026-05-10 16:19:58][app.vjepa.train ][_save_crash ] Saved crash log to /scratch-shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/crash_rank5.log
vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/job_22614766/22614766_6_log.err ADDED
@@ -0,0 +1,68 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/timm/models/layers/__init__.py:49: FutureWarning: Importing from timm.models.layers is deprecated, please import via timm.layers
2
+ warnings.warn(f"Importing from {__name__} is deprecated, please import via timm.layers", FutureWarning)
3
+ [rank6]:[W510 12:20:01.354617374 ProcessGroupNCCL.cpp:5138] Guessing device ID based on global rank. This can cause a hang if rank to GPU mapping is heterogeneous. You can specify device_id in init_process_group()
4
+ /gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/vjepa/utils.py:796: FutureWarning: `torch.cuda.amp.GradScaler(args...)` is deprecated. Please use `torch.amp.GradScaler('cuda', args...)` instead.
5
+ scaler = torch.cuda.amp.GradScaler() if mixed_precision else None
6
+ /gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/vjepa/train.py:955: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
7
+ with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
8
+ submitit ERROR (2026-05-10 16:19:48,628) - Submitted job triggered an exception
9
+ [rank6]:[W510 16:19:49.534536952 TCPStore.cpp:125] [c10d] recvValue failed on SocketImpl(fd=3, addr=[gcn82.local.snellius.surf.nl]:43922, remote=[gcn80.local.snellius.surf.nl]:37129): Failed to recv, got 0 bytes. Connection was likely closed. Did the remote server shutdown or crash?
10
+ Exception raised from recvBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:682 (most recent call first):
11
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x14620edb2fdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
12
+ frame #1: <unknown function> + 0x6a3325d (0x146252f4125d in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
13
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x273 (0x146252f3f1f3 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
14
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x146210021eec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
15
+ frame #4: <unknown function> + 0xed164 (0x14634f48b164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
16
+ frame #5: <unknown function> + 0x8a19a (0x14635fefd19a in /lib64/libc.so.6)
17
+ frame #6: <unknown function> + 0x10f100 (0x14635ff82100 in /lib64/libc.so.6)
18
+
19
+ [rank6]:[W510 16:19:49.536929964 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 6] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Failed to recv, got 0 bytes. Connection was likely closed. Did the remote server shutdown or crash?
20
+ [rank6]:[W510 16:19:50.537085079 TCPStore.cpp:106] [c10d] sendBytes failed on SocketImpl(fd=3, addr=[gcn82.local.snellius.surf.nl]:43922, remote=[gcn80.local.snellius.surf.nl]:37129): Broken pipe
21
+ Exception raised from sendBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:653 (most recent call first):
22
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x14620edb2fdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
23
+ frame #1: <unknown function> + 0x6a326d1 (0x146252f406d1 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
24
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x24d (0x146252f3f1cd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
25
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x146210021eec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
26
+ frame #4: <unknown function> + 0xed164 (0x14634f48b164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
27
+ frame #5: <unknown function> + 0x8a19a (0x14635fefd19a in /lib64/libc.so.6)
28
+ frame #6: <unknown function> + 0x10f100 (0x14635ff82100 in /lib64/libc.so.6)
29
+
30
+ [rank6]:[W510 16:19:50.540863871 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 6] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Broken pipe
31
+ Traceback (most recent call last):
32
+ File "<frozen runpy>", line 198, in _run_module_as_main
33
+ File "<frozen runpy>", line 88, in _run_code
34
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/_submit.py", line 11, in <module>
35
+ submitit_main()
36
+ ~~~~~~~~~~~~~^^
37
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 76, in submitit_main
38
+ process_job(args.folder)
39
+ ~~~~~~~~~~~^^^^^^^^^^^^^
40
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 69, in process_job
41
+ raise error
42
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 55, in process_job
43
+ result = delayed.result()
44
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/utils.py", line 137, in result
45
+ self._result = self.function(*self.args, **self.kwargs)
46
+ ~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^
47
+ File "/gpfs/home2/dcanez/vd/app/main_distributed.py", line 102, in __call__
48
+ app_main(app, args=params, resume_preempt=resume_preempt)
49
+ ~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
50
+ File "/gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/scaffold.py", line 17, in main
51
+ return importlib.import_module(f"app.{app}.train").main(args=args, resume_preempt=resume_preempt)
52
+ ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
53
+ File "/gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/vjepa/train.py", line 1189, in main
54
+ assert not np.isnan(loss), "loss is nan"
55
+ ^^^^^^^^^^^^^^^^^^
56
+ AssertionError: loss is nan
57
+ [rank6]:[W510 16:19:51.541031985 TCPStore.cpp:106] [c10d] sendBytes failed on SocketImpl(fd=3, addr=[gcn82.local.snellius.surf.nl]:43922, remote=[gcn80.local.snellius.surf.nl]:37129): Broken pipe
58
+ Exception raised from sendBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:653 (most recent call first):
59
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x14620edb2fdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
60
+ frame #1: <unknown function> + 0x6a326d1 (0x146252f406d1 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
61
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x24d (0x146252f3f1cd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
62
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x146210021eec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
63
+ frame #4: <unknown function> + 0xed164 (0x14634f48b164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
64
+ frame #5: <unknown function> + 0x8a19a (0x14635fefd19a in /lib64/libc.so.6)
65
+ frame #6: <unknown function> + 0x10f100 (0x14635ff82100 in /lib64/libc.so.6)
66
+
67
+ [rank6]:[W510 16:19:51.543425278 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 6] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Broken pipe
68
+ [rank6]:[W510 16:19:51.813700522 ProcessGroupNCCL.cpp:1553] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())
vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/job_22614766/22614766_6_log.out ADDED
@@ -0,0 +1,276 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ submitit INFO (2026-05-10 12:14:34,049) - Starting with JobEnvironment(job_id=22614766, hostname=gcn82.local.snellius.surf.nl, local_rank=2(4), node=1(4), global_rank=6(16))
2
+ submitit INFO (2026-05-10 12:14:34,049) - Loading pickle: /scratch-shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/job_22614766/22614766_submitted.pkl
3
+ INFO:root:loaded pretrain params...
4
+ { 'app': 'vjepa',
5
+ 'cpus_per_task': 16,
6
+ 'data': { 'batch_size': 64,
7
+ 'crop_size': 224,
8
+ 'dataset_fpcs': [16, 16],
9
+ 'dataset_type': 'VideoDataset',
10
+ 'datasets': [ '/scratch-shared/dcanez/data/kinetics/k400/train.csv',
11
+ '/scratch-shared/dcanez/data/ssv2/train.csv'],
12
+ 'datasets_weights': [0.65, 0.35],
13
+ 'fps': 4,
14
+ 'num_workers': 10,
15
+ 'patch_size': 14,
16
+ 'persistent_workers': True,
17
+ 'pin_mem': True,
18
+ 'stage': [ { 'dest': 'kinetics_240',
19
+ 'format': 'targz_parts',
20
+ 'src': '/scratch-shared/dcanez/data/kinetics/k400/tars_240/'},
21
+ { 'dest': 'ssv2',
22
+ 'format': 'multipart_tar',
23
+ 'src': '/scratch-nvme/ml-datasets/something-something-v2/'}],
24
+ 'tubelet_size': 1},
25
+ 'data_aug': { 'auto_augment': False,
26
+ 'motion_shift': False,
27
+ 'random_resize_aspect_ratio': [0.75, 1.35],
28
+ 'random_resize_scale': [0.3, 1.0],
29
+ 'reprob': 0.0},
30
+ 'folder': '/scratch-shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable',
31
+ 'loss': {'loss_exp': 1.0},
32
+ 'mask': [ { 'aspect_ratio': [0.75, 1.5],
33
+ 'full_complement': False,
34
+ 'max_keep': None,
35
+ 'max_temporal_keep': 1.0,
36
+ 'num_blocks': 8,
37
+ 'spatial_scale': [0.15, 0.15],
38
+ 'temporal_scale': [1.0, 1.0]},
39
+ { 'aspect_ratio': [0.75, 1.5],
40
+ 'full_complement': False,
41
+ 'max_keep': None,
42
+ 'max_temporal_keep': 1.0,
43
+ 'num_blocks': 2,
44
+ 'spatial_scale': [0.7, 0.7],
45
+ 'temporal_scale': [1.0, 1.0]}],
46
+ 'mem_per_gpu': '180G',
47
+ 'meta': { 'dtype': 'bfloat16',
48
+ 'knn_eval_epoch0': False,
49
+ 'knn_eval_freq': 5,
50
+ 'knn_eval_presets': [ { 'config': { 'batch_size': 64,
51
+ 'dataset_train': '/scratch-shared/mdorkenw/ucf101/ucfTrainTestlist/train.csv',
52
+ 'dataset_val': '/scratch-shared/mdorkenw/ucf101/ucfTrainTestlist/val.csv',
53
+ 'eval_videos_per_class': 25,
54
+ 'num_workers': 8,
55
+ 'pool_type': 'slot_temporal_concat',
56
+ 'train_videos_per_class': 100},
57
+ 'preset': 'ucf101'},
58
+ { 'config': { 'batch_size': 64,
59
+ 'dataset_train': '/scratch-shared/mdorkenw/20bn-something-something-v2/something-something-v2-annotations/train_coarse10.csv',
60
+ 'dataset_val': '/scratch-shared/mdorkenw/20bn-something-something-v2/something-something-v2-annotations/val_coarse10.csv',
61
+ 'eval_videos_per_class': 100,
62
+ 'linear_probe': True,
63
+ 'num_workers': 8,
64
+ 'pool_type': 'slot_temporal_concat',
65
+ 'train_videos_per_class': 500},
66
+ 'preset': 'ssv2_coarse10'}],
67
+ 'load_checkpoint': True,
68
+ 'read_checkpoint': None,
69
+ 'save_every_freq': 5,
70
+ 'seed': 239,
71
+ 'use_sdpa': True,
72
+ 'use_wandb': True,
73
+ 'wandb_project': 'vjepa_ujepaside'},
74
+ 'metrics': {'sigreg': {}, 'std': {}},
75
+ 'model': { 'model_name': 'ujepaside_large_patch14_capi_lvd1689m',
76
+ 'pred_depth': 6,
77
+ 'pred_embed_dim': 384,
78
+ 'pred_num_heads': 12,
79
+ 'predictor': 'v2_cross',
80
+ 'st_causal': False,
81
+ 'st_drop_path': 0.2,
82
+ 'st_flex_enable': False,
83
+ 'st_layer_scale_init': 1e-05,
84
+ 'st_num_slots': 16,
85
+ 'st_slots_causal_within_frame': False,
86
+ 'target_kind': 'frozen_2d',
87
+ 'target_type': 'vit_large_patch14_capi.lvd1689m',
88
+ 'temporal_spacing': 1.0,
89
+ 'uniform_power': True,
90
+ 'use_activation_checkpointing': True,
91
+ 'use_mask_tokens': True,
92
+ 'use_rope': True,
93
+ 'use_sdpa': True,
94
+ 'zero_init_mask_tokens': True},
95
+ 'nodes': 4,
96
+ 'optimization': { 'clip_grad': 3.0,
97
+ 'ema': [0.99925, 0.99925],
98
+ 'epochs': 100,
99
+ 'final_lr': 0.0001,
100
+ 'final_weight_decay': 0.04,
101
+ 'ipe': 300,
102
+ 'ipe_scale': 1.0,
103
+ 'lr': 0.0005,
104
+ 'start_lr': 0.0001,
105
+ 'warmup': 10,
106
+ 'weight_decay': 0.04},
107
+ 'tasks_per_node': 4}
108
+ INFO:root:Running pre-training of app: vjepa
109
+ [INFO ][2026-05-10 12:16:53][app.vjepa.train ][main ] which_dtype='bfloat16'
110
+ [INFO ][2026-05-10 12:16:53][app.vjepa.train ][main ] Disabling persistent_workers (incompatible with KNN eval)
111
+ [INFO ][2026-05-10 12:16:53][app.vjepa.train ][main ] NCCL_SOCKET_IFNAME=eno
112
+ [INFO ][2026-05-10 12:16:56][app.vjepa.train ][main ] Initialized (rank/world-size) 6/16, tasks_per_node=4
113
+ [INFO ][2026-05-10 12:16:56][root ][stage_datasets ] [local_rank 2/4] Staging kinetics_240 (targz_parts)
114
+ [INFO ][2026-05-10 12:16:56][root ][_stage_targz_parts ] [rank 2] Extracting 26/103 tar.gz parts to /scratch-node/dcanez.22614766/kinetics_240
115
+ [INFO ][2026-05-10 12:17:10][root ][_stage_targz_parts ] [local_rank 2] Extracted 2/26 parts
116
+ [INFO ][2026-05-10 12:17:23][root ][_stage_targz_parts ] [local_rank 2] Extracted 4/26 parts
117
+ [INFO ][2026-05-10 12:17:38][root ][_stage_targz_parts ] [local_rank 2] Extracted 6/26 parts
118
+ [INFO ][2026-05-10 12:17:53][root ][_stage_targz_parts ] [local_rank 2] Extracted 8/26 parts
119
+ [INFO ][2026-05-10 12:18:07][root ][_stage_targz_parts ] [local_rank 2] Extracted 10/26 parts
120
+ [INFO ][2026-05-10 12:18:21][root ][_stage_targz_parts ] [local_rank 2] Extracted 12/26 parts
121
+ [INFO ][2026-05-10 12:18:35][root ][_stage_targz_parts ] [local_rank 2] Extracted 14/26 parts
122
+ [INFO ][2026-05-10 12:18:49][root ][_stage_targz_parts ] [local_rank 2] Extracted 16/26 parts
123
+ [INFO ][2026-05-10 12:19:03][root ][_stage_targz_parts ] [local_rank 2] Extracted 18/26 parts
124
+ [INFO ][2026-05-10 12:19:18][root ][_stage_targz_parts ] [local_rank 2] Extracted 20/26 parts
125
+ [INFO ][2026-05-10 12:19:32][root ][_stage_targz_parts ] [local_rank 2] Extracted 22/26 parts
126
+ [INFO ][2026-05-10 12:19:46][root ][_stage_targz_parts ] [local_rank 2] Extracted 24/26 parts
127
+ [INFO ][2026-05-10 12:20:01][root ][_stage_targz_parts ] [local_rank 2] Extracted 26/26 parts
128
+ [INFO ][2026-05-10 12:20:01][root ][stage_datasets ] [local_rank 2/4] Staging ssv2 (multipart_tar)
129
+ [INFO ][2026-05-10 12:21:03][app.vjepa.train ][main ] Staged datasets: ['/scratch-node/dcanez.22614766/kinetics_240/train.csv', '/scratch-node/dcanez.22614766/ssv2/train.csv']
130
+ [INFO ][2026-05-10 12:23:31][experiments.stmodels.vision_transformers_v3][vit_large_patch14_capi_lvd1689m] Loading pretrained weights for vit_large_patch14_capi_lvd1689m from https://dl.fbaipublicfiles.com/capi/capi_vitl14_lvd.pth
131
+ [INFO ][2026-05-10 12:23:56][root ][_build_frozen_2d_target ] Loaded frozen_2d target from timm: vit_large_patch14_capi.lvd1689m
132
+ [INFO ][2026-05-10 12:23:56][root ][init_video_model ] ViTMultiSeqWrapper(
133
+ (backbone): UJEPAside(
134
+ (patch_embed): PatchEmbed3D(
135
+ (proj): Conv3d(3, 1024, kernel_size=(1, 14, 14), stride=(1, 14, 14))
136
+ )
137
+ (rope): CAPI2DRoPE()
138
+ (blocks): ModuleList(
139
+ (0-23): 24 x Block(
140
+ (norm1): RMSNorm((1024,), eps=1e-05, elementwise_affine=True)
141
+ (rope_impl): CAPI2DRoPE()
142
+ (norm2): RMSNorm((1024,), eps=1e-05, elementwise_affine=True)
143
+ (drop_path1): Identity()
144
+ (drop_path2): Identity()
145
+ (attn): Attention(
146
+ (qkv): Linear(in_features=1024, out_features=3072, bias=False)
147
+ (attn_drop): Dropout(p=0.0, inplace=False)
148
+ (proj): Linear(in_features=1024, out_features=1024, bias=False)
149
+ (proj_drop): Dropout(p=0.0, inplace=False)
150
+ (rope_impl): CAPI2DRoPE()
151
+ )
152
+ (mlp): MLP(
153
+ (fc1): Linear(in_features=1024, out_features=4096, bias=False)
154
+ (act): GELU(approximate='none')
155
+ (fc2): Linear(in_features=4096, out_features=1024, bias=False)
156
+ (drop): Dropout(p=0.0, inplace=False)
157
+ )
158
+ )
159
+ )
160
+ (norm): RMSNorm((1024,), eps=1e-05, elementwise_affine=False)
161
+ (st_blocks): ModuleList(
162
+ (0-23): 24 x Block(
163
+ (residual1): EfficientResidual(
164
+ (norm): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
165
+ (fn): Attention(
166
+ (q_proj): Linear(in_features=1024, out_features=1024, bias=False)
167
+ (k_proj): Linear(in_features=1024, out_features=1024, bias=False)
168
+ (v_proj): Linear(in_features=1024, out_features=1024, bias=False)
169
+ (proj): Linear(in_features=1024, out_features=1024, bias=False)
170
+ (rope_impl): CAPI3DRoPE()
171
+ )
172
+ )
173
+ (residual2): EfficientResidual(
174
+ (norm): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
175
+ (fn): MLP(
176
+ (fc1): Linear(in_features=1024, out_features=4096, bias=False)
177
+ (act): GELU(approximate='none')
178
+ (fc2): Linear(in_features=4096, out_features=1024, bias=False)
179
+ (drop): Dropout(p=0.0, inplace=False)
180
+ )
181
+ )
182
+ )
183
+ )
184
+ (st_rope): CAPI3DRoPE()
185
+ )
186
+ )
187
+ [INFO ][2026-05-10 12:23:56][root ][init_video_model ] UJEPAsidePredictorMultiSeqWrapper(
188
+ (backbone): PredictorV2(
189
+ (predictor_embed): Linear(in_features=1024, out_features=384, bias=True)
190
+ (mask_tokens): ParameterList(
191
+ (0): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
192
+ (1): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
193
+ (2): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
194
+ (3): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
195
+ )
196
+ (predictor_blocks): ModuleList(
197
+ (0-5): 6 x Block(
198
+ (residual1): EfficientResidual(
199
+ (norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
200
+ (fn): Attention(
201
+ (q_proj): Linear(in_features=384, out_features=384, bias=False)
202
+ (k_proj): Linear(in_features=384, out_features=384, bias=False)
203
+ (v_proj): Linear(in_features=384, out_features=384, bias=False)
204
+ (proj): Linear(in_features=384, out_features=384, bias=False)
205
+ (rope): Rope()
206
+ )
207
+ )
208
+ (residual2): EfficientResidual(
209
+ (norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
210
+ (fn): MLP(
211
+ (fc1): Linear(in_features=384, out_features=1536, bias=False)
212
+ (act): GELU(approximate='none')
213
+ (fc2): Linear(in_features=1536, out_features=384, bias=False)
214
+ (drop): Dropout(p=0.0, inplace=False)
215
+ )
216
+ )
217
+ )
218
+ )
219
+ (predictor_norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
220
+ (predictor_proj): Linear(in_features=384, out_features=1024, bias=True)
221
+ )
222
+ )
223
+ [INFO ][2026-05-10 12:23:56][root ][init_video_model ] MultiSeqWrapper(
224
+ (backbone): Frozen2DTargetWrapper(
225
+ (backbone): Eva(
226
+ (patch_embed): PatchEmbed(
227
+ (proj): Conv2d(3, 1024, kernel_size=(14, 14), stride=(14, 14))
228
+ (norm): Identity()
229
+ )
230
+ (pos_drop): Dropout(p=0.0, inplace=False)
231
+ (norm_pre): Identity()
232
+ (blocks): ModuleList(
233
+ (0-23): 24 x EvaBlock(
234
+ (norm1): RMSNorm((1024,), eps=1e-05, elementwise_affine=True)
235
+ (attn): EvaAttention(
236
+ (qkv): Linear(in_features=1024, out_features=3072, bias=False)
237
+ (q_norm): Identity()
238
+ (k_norm): Identity()
239
+ (attn_drop): Dropout(p=0.0, inplace=False)
240
+ (norm): Identity()
241
+ (proj): Linear(in_features=1024, out_features=1024, bias=True)
242
+ (proj_drop): Dropout(p=0.0, inplace=False)
243
+ )
244
+ (drop_path1): Identity()
245
+ (norm2): RMSNorm((1024,), eps=1e-05, elementwise_affine=True)
246
+ (mlp): Mlp(
247
+ (fc1): Linear(in_features=1024, out_features=4096, bias=True)
248
+ (act): GELU(approximate='none')
249
+ (drop1): Dropout(p=0.0, inplace=False)
250
+ (norm): Identity()
251
+ (fc2): Linear(in_features=4096, out_features=1024, bias=True)
252
+ (drop2): Dropout(p=0.0, inplace=False)
253
+ )
254
+ (drop_path2): Identity()
255
+ )
256
+ )
257
+ (norm): RMSNorm((1024,), eps=1e-05, elementwise_affine=True)
258
+ (fc_norm): Identity()
259
+ (head_drop): Dropout(p=0.0, inplace=False)
260
+ (head): Identity()
261
+ (rope): _CapiPatchRoPE()
262
+ )
263
+ )
264
+ )
265
+ [INFO ][2026-05-10 12:23:56][root ][init_video_model ] Encoder number of parameters: 302399488
266
+ [INFO ][2026-05-10 12:23:56][root ][init_video_model ] Predictor number of parameters: 11416192
267
+ [INFO ][2026-05-10 12:23:56][root ][init_video_model ] Target encoder number of parameters: 0
268
+ [INFO ][2026-05-10 12:24:07][root ][make_videodataset ] VideoDataset dataset created
269
+ [INFO ][2026-05-10 12:24:07][WeightedSampler ][__init__ ] Using DistributedWeightedSampler with rank 6 / 16
270
+ [INFO ][2026-05-10 12:24:07][root ][make_videodataset ] VideoDataset unsupervised data loader created
271
+ [INFO ][2026-05-10 12:24:07][app.vjepa.train ][main ] iterations per epoch/dataset length: 300/399
272
+ [INFO ][2026-05-10 12:24:07][app.vjepa.train ][main ] Wrapping models in DDP (rank 6)...
273
+ [INFO ][2026-05-10 12:24:09][app.vjepa.train ][main ] Initializing loader...
274
+ submitit ERROR (2026-05-10 16:19:48,628) - Submitted job triggered an exception
275
+ [ERROR ][2026-05-10 16:19:48][submitit ][process_job ] Submitted job triggered an exception
276
+ [ERROR ][2026-05-10 16:19:51][app.vjepa.train ][_save_crash ] Saved crash log to /scratch-shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/crash_rank6.log
vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/job_22614766/22614766_7_log.err ADDED
@@ -0,0 +1,68 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/timm/models/layers/__init__.py:49: FutureWarning: Importing from timm.models.layers is deprecated, please import via timm.layers
2
+ warnings.warn(f"Importing from {__name__} is deprecated, please import via timm.layers", FutureWarning)
3
+ [rank7]:[W510 12:19:56.166705489 ProcessGroupNCCL.cpp:5138] Guessing device ID based on global rank. This can cause a hang if rank to GPU mapping is heterogeneous. You can specify device_id in init_process_group()
4
+ /gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/vjepa/utils.py:796: FutureWarning: `torch.cuda.amp.GradScaler(args...)` is deprecated. Please use `torch.amp.GradScaler('cuda', args...)` instead.
5
+ scaler = torch.cuda.amp.GradScaler() if mixed_precision else None
6
+ /gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/vjepa/train.py:955: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
7
+ with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
8
+ [rank7]:[W510 16:19:49.193338512 TCPStore.cpp:125] [c10d] recvValue failed on SocketImpl(fd=3, addr=[gcn82.local.snellius.surf.nl]:43918, remote=[gcn80.local.snellius.surf.nl]:37129): Failed to recv, got 0 bytes. Connection was likely closed. Did the remote server shutdown or crash?
9
+ Exception raised from recvBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:682 (most recent call first):
10
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x14f93fe48fdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
11
+ frame #1: <unknown function> + 0x6a3325d (0x14f983fd725d in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
12
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x273 (0x14f983fd51f3 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
13
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x14f9410b7eec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
14
+ frame #4: <unknown function> + 0xed164 (0x14fa80521164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
15
+ frame #5: <unknown function> + 0x8a19a (0x14fa90f9319a in /lib64/libc.so.6)
16
+ frame #6: <unknown function> + 0x10f100 (0x14fa91018100 in /lib64/libc.so.6)
17
+
18
+ [rank7]:[W510 16:19:49.250685538 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 7] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Failed to recv, got 0 bytes. Connection was likely closed. Did the remote server shutdown or crash?
19
+ submitit ERROR (2026-05-10 16:19:50,186) - Submitted job triggered an exception
20
+ [rank7]:[W510 16:19:50.250857942 TCPStore.cpp:106] [c10d] sendBytes failed on SocketImpl(fd=3, addr=[gcn82.local.snellius.surf.nl]:43918, remote=[gcn80.local.snellius.surf.nl]:37129): Broken pipe
21
+ Exception raised from sendBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:653 (most recent call first):
22
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x14f93fe48fdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
23
+ frame #1: <unknown function> + 0x6a326d1 (0x14f983fd66d1 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
24
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x24d (0x14f983fd51cd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
25
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x14f9410b7eec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
26
+ frame #4: <unknown function> + 0xed164 (0x14fa80521164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
27
+ frame #5: <unknown function> + 0x8a19a (0x14fa90f9319a in /lib64/libc.so.6)
28
+ frame #6: <unknown function> + 0x10f100 (0x14fa91018100 in /lib64/libc.so.6)
29
+
30
+ [rank7]:[W510 16:19:50.254608224 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 7] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Broken pipe
31
+ Traceback (most recent call last):
32
+ File "<frozen runpy>", line 198, in _run_module_as_main
33
+ File "<frozen runpy>", line 88, in _run_code
34
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/_submit.py", line 11, in <module>
35
+ submitit_main()
36
+ ~~~~~~~~~~~~~^^
37
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 76, in submitit_main
38
+ process_job(args.folder)
39
+ ~~~~~~~~~~~^^^^^^^^^^^^^
40
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 69, in process_job
41
+ raise error
42
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 55, in process_job
43
+ result = delayed.result()
44
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/utils.py", line 137, in result
45
+ self._result = self.function(*self.args, **self.kwargs)
46
+ ~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^
47
+ File "/gpfs/home2/dcanez/vd/app/main_distributed.py", line 102, in __call__
48
+ app_main(app, args=params, resume_preempt=resume_preempt)
49
+ ~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
50
+ File "/gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/scaffold.py", line 17, in main
51
+ return importlib.import_module(f"app.{app}.train").main(args=args, resume_preempt=resume_preempt)
52
+ ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
53
+ File "/gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/vjepa/train.py", line 1189, in main
54
+ assert not np.isnan(loss), "loss is nan"
55
+ ^^^^^^^^^^^^^^^^^^
56
+ AssertionError: loss is nan
57
+ [rank7]:[W510 16:19:51.254784989 TCPStore.cpp:106] [c10d] sendBytes failed on SocketImpl(fd=3, addr=[gcn82.local.snellius.surf.nl]:43918, remote=[gcn80.local.snellius.surf.nl]:37129): Broken pipe
58
+ Exception raised from sendBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:653 (most recent call first):
59
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x14f93fe48fdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
60
+ frame #1: <unknown function> + 0x6a326d1 (0x14f983fd66d1 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
61
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x24d (0x14f983fd51cd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
62
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x14f9410b7eec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
63
+ frame #4: <unknown function> + 0xed164 (0x14fa80521164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
64
+ frame #5: <unknown function> + 0x8a19a (0x14fa90f9319a in /lib64/libc.so.6)
65
+ frame #6: <unknown function> + 0x10f100 (0x14fa91018100 in /lib64/libc.so.6)
66
+
67
+ [rank7]:[W510 16:19:51.257267151 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 7] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Broken pipe
68
+ [rank7]:[W510 16:19:51.813923330 ProcessGroupNCCL.cpp:1553] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())
vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/job_22614766/22614766_7_log.out ADDED
@@ -0,0 +1,276 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ submitit INFO (2026-05-10 12:14:34,049) - Starting with JobEnvironment(job_id=22614766, hostname=gcn82.local.snellius.surf.nl, local_rank=3(4), node=1(4), global_rank=7(16))
2
+ submitit INFO (2026-05-10 12:14:34,049) - Loading pickle: /scratch-shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/job_22614766/22614766_submitted.pkl
3
+ INFO:root:loaded pretrain params...
4
+ { 'app': 'vjepa',
5
+ 'cpus_per_task': 16,
6
+ 'data': { 'batch_size': 64,
7
+ 'crop_size': 224,
8
+ 'dataset_fpcs': [16, 16],
9
+ 'dataset_type': 'VideoDataset',
10
+ 'datasets': [ '/scratch-shared/dcanez/data/kinetics/k400/train.csv',
11
+ '/scratch-shared/dcanez/data/ssv2/train.csv'],
12
+ 'datasets_weights': [0.65, 0.35],
13
+ 'fps': 4,
14
+ 'num_workers': 10,
15
+ 'patch_size': 14,
16
+ 'persistent_workers': True,
17
+ 'pin_mem': True,
18
+ 'stage': [ { 'dest': 'kinetics_240',
19
+ 'format': 'targz_parts',
20
+ 'src': '/scratch-shared/dcanez/data/kinetics/k400/tars_240/'},
21
+ { 'dest': 'ssv2',
22
+ 'format': 'multipart_tar',
23
+ 'src': '/scratch-nvme/ml-datasets/something-something-v2/'}],
24
+ 'tubelet_size': 1},
25
+ 'data_aug': { 'auto_augment': False,
26
+ 'motion_shift': False,
27
+ 'random_resize_aspect_ratio': [0.75, 1.35],
28
+ 'random_resize_scale': [0.3, 1.0],
29
+ 'reprob': 0.0},
30
+ 'folder': '/scratch-shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable',
31
+ 'loss': {'loss_exp': 1.0},
32
+ 'mask': [ { 'aspect_ratio': [0.75, 1.5],
33
+ 'full_complement': False,
34
+ 'max_keep': None,
35
+ 'max_temporal_keep': 1.0,
36
+ 'num_blocks': 8,
37
+ 'spatial_scale': [0.15, 0.15],
38
+ 'temporal_scale': [1.0, 1.0]},
39
+ { 'aspect_ratio': [0.75, 1.5],
40
+ 'full_complement': False,
41
+ 'max_keep': None,
42
+ 'max_temporal_keep': 1.0,
43
+ 'num_blocks': 2,
44
+ 'spatial_scale': [0.7, 0.7],
45
+ 'temporal_scale': [1.0, 1.0]}],
46
+ 'mem_per_gpu': '180G',
47
+ 'meta': { 'dtype': 'bfloat16',
48
+ 'knn_eval_epoch0': False,
49
+ 'knn_eval_freq': 5,
50
+ 'knn_eval_presets': [ { 'config': { 'batch_size': 64,
51
+ 'dataset_train': '/scratch-shared/mdorkenw/ucf101/ucfTrainTestlist/train.csv',
52
+ 'dataset_val': '/scratch-shared/mdorkenw/ucf101/ucfTrainTestlist/val.csv',
53
+ 'eval_videos_per_class': 25,
54
+ 'num_workers': 8,
55
+ 'pool_type': 'slot_temporal_concat',
56
+ 'train_videos_per_class': 100},
57
+ 'preset': 'ucf101'},
58
+ { 'config': { 'batch_size': 64,
59
+ 'dataset_train': '/scratch-shared/mdorkenw/20bn-something-something-v2/something-something-v2-annotations/train_coarse10.csv',
60
+ 'dataset_val': '/scratch-shared/mdorkenw/20bn-something-something-v2/something-something-v2-annotations/val_coarse10.csv',
61
+ 'eval_videos_per_class': 100,
62
+ 'linear_probe': True,
63
+ 'num_workers': 8,
64
+ 'pool_type': 'slot_temporal_concat',
65
+ 'train_videos_per_class': 500},
66
+ 'preset': 'ssv2_coarse10'}],
67
+ 'load_checkpoint': True,
68
+ 'read_checkpoint': None,
69
+ 'save_every_freq': 5,
70
+ 'seed': 239,
71
+ 'use_sdpa': True,
72
+ 'use_wandb': True,
73
+ 'wandb_project': 'vjepa_ujepaside'},
74
+ 'metrics': {'sigreg': {}, 'std': {}},
75
+ 'model': { 'model_name': 'ujepaside_large_patch14_capi_lvd1689m',
76
+ 'pred_depth': 6,
77
+ 'pred_embed_dim': 384,
78
+ 'pred_num_heads': 12,
79
+ 'predictor': 'v2_cross',
80
+ 'st_causal': False,
81
+ 'st_drop_path': 0.2,
82
+ 'st_flex_enable': False,
83
+ 'st_layer_scale_init': 1e-05,
84
+ 'st_num_slots': 16,
85
+ 'st_slots_causal_within_frame': False,
86
+ 'target_kind': 'frozen_2d',
87
+ 'target_type': 'vit_large_patch14_capi.lvd1689m',
88
+ 'temporal_spacing': 1.0,
89
+ 'uniform_power': True,
90
+ 'use_activation_checkpointing': True,
91
+ 'use_mask_tokens': True,
92
+ 'use_rope': True,
93
+ 'use_sdpa': True,
94
+ 'zero_init_mask_tokens': True},
95
+ 'nodes': 4,
96
+ 'optimization': { 'clip_grad': 3.0,
97
+ 'ema': [0.99925, 0.99925],
98
+ 'epochs': 100,
99
+ 'final_lr': 0.0001,
100
+ 'final_weight_decay': 0.04,
101
+ 'ipe': 300,
102
+ 'ipe_scale': 1.0,
103
+ 'lr': 0.0005,
104
+ 'start_lr': 0.0001,
105
+ 'warmup': 10,
106
+ 'weight_decay': 0.04},
107
+ 'tasks_per_node': 4}
108
+ INFO:root:Running pre-training of app: vjepa
109
+ [INFO ][2026-05-10 12:16:53][app.vjepa.train ][main ] which_dtype='bfloat16'
110
+ [INFO ][2026-05-10 12:16:53][app.vjepa.train ][main ] Disabling persistent_workers (incompatible with KNN eval)
111
+ [INFO ][2026-05-10 12:16:53][app.vjepa.train ][main ] NCCL_SOCKET_IFNAME=eno
112
+ [INFO ][2026-05-10 12:16:57][app.vjepa.train ][main ] Initialized (rank/world-size) 7/16, tasks_per_node=4
113
+ [INFO ][2026-05-10 12:16:57][root ][stage_datasets ] [local_rank 3/4] Staging kinetics_240 (targz_parts)
114
+ [INFO ][2026-05-10 12:16:57][root ][_stage_targz_parts ] [rank 3] Extracting 25/103 tar.gz parts to /scratch-node/dcanez.22614766/kinetics_240
115
+ [INFO ][2026-05-10 12:17:11][root ][_stage_targz_parts ] [local_rank 3] Extracted 2/25 parts
116
+ [INFO ][2026-05-10 12:17:25][root ][_stage_targz_parts ] [local_rank 3] Extracted 4/25 parts
117
+ [INFO ][2026-05-10 12:17:39][root ][_stage_targz_parts ] [local_rank 3] Extracted 6/25 parts
118
+ [INFO ][2026-05-10 12:17:53][root ][_stage_targz_parts ] [local_rank 3] Extracted 8/25 parts
119
+ [INFO ][2026-05-10 12:18:07][root ][_stage_targz_parts ] [local_rank 3] Extracted 10/25 parts
120
+ [INFO ][2026-05-10 12:18:22][root ][_stage_targz_parts ] [local_rank 3] Extracted 12/25 parts
121
+ [INFO ][2026-05-10 12:18:36][root ][_stage_targz_parts ] [local_rank 3] Extracted 14/25 parts
122
+ [INFO ][2026-05-10 12:18:51][root ][_stage_targz_parts ] [local_rank 3] Extracted 16/25 parts
123
+ [INFO ][2026-05-10 12:19:04][root ][_stage_targz_parts ] [local_rank 3] Extracted 18/25 parts
124
+ [INFO ][2026-05-10 12:19:19][root ][_stage_targz_parts ] [local_rank 3] Extracted 20/25 parts
125
+ [INFO ][2026-05-10 12:19:33][root ][_stage_targz_parts ] [local_rank 3] Extracted 22/25 parts
126
+ [INFO ][2026-05-10 12:19:47][root ][_stage_targz_parts ] [local_rank 3] Extracted 24/25 parts
127
+ [INFO ][2026-05-10 12:19:55][root ][_stage_targz_parts ] [local_rank 3] Extracted 25/25 parts
128
+ [INFO ][2026-05-10 12:19:55][root ][stage_datasets ] [local_rank 3/4] Staging ssv2 (multipart_tar)
129
+ [INFO ][2026-05-10 12:21:03][app.vjepa.train ][main ] Staged datasets: ['/scratch-node/dcanez.22614766/kinetics_240/train.csv', '/scratch-node/dcanez.22614766/ssv2/train.csv']
130
+ [INFO ][2026-05-10 12:23:31][experiments.stmodels.vision_transformers_v3][vit_large_patch14_capi_lvd1689m] Loading pretrained weights for vit_large_patch14_capi_lvd1689m from https://dl.fbaipublicfiles.com/capi/capi_vitl14_lvd.pth
131
+ [INFO ][2026-05-10 12:23:56][root ][_build_frozen_2d_target ] Loaded frozen_2d target from timm: vit_large_patch14_capi.lvd1689m
132
+ [INFO ][2026-05-10 12:23:56][root ][init_video_model ] ViTMultiSeqWrapper(
133
+ (backbone): UJEPAside(
134
+ (patch_embed): PatchEmbed3D(
135
+ (proj): Conv3d(3, 1024, kernel_size=(1, 14, 14), stride=(1, 14, 14))
136
+ )
137
+ (rope): CAPI2DRoPE()
138
+ (blocks): ModuleList(
139
+ (0-23): 24 x Block(
140
+ (norm1): RMSNorm((1024,), eps=1e-05, elementwise_affine=True)
141
+ (rope_impl): CAPI2DRoPE()
142
+ (norm2): RMSNorm((1024,), eps=1e-05, elementwise_affine=True)
143
+ (drop_path1): Identity()
144
+ (drop_path2): Identity()
145
+ (attn): Attention(
146
+ (qkv): Linear(in_features=1024, out_features=3072, bias=False)
147
+ (attn_drop): Dropout(p=0.0, inplace=False)
148
+ (proj): Linear(in_features=1024, out_features=1024, bias=False)
149
+ (proj_drop): Dropout(p=0.0, inplace=False)
150
+ (rope_impl): CAPI2DRoPE()
151
+ )
152
+ (mlp): MLP(
153
+ (fc1): Linear(in_features=1024, out_features=4096, bias=False)
154
+ (act): GELU(approximate='none')
155
+ (fc2): Linear(in_features=4096, out_features=1024, bias=False)
156
+ (drop): Dropout(p=0.0, inplace=False)
157
+ )
158
+ )
159
+ )
160
+ (norm): RMSNorm((1024,), eps=1e-05, elementwise_affine=False)
161
+ (st_blocks): ModuleList(
162
+ (0-23): 24 x Block(
163
+ (residual1): EfficientResidual(
164
+ (norm): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
165
+ (fn): Attention(
166
+ (q_proj): Linear(in_features=1024, out_features=1024, bias=False)
167
+ (k_proj): Linear(in_features=1024, out_features=1024, bias=False)
168
+ (v_proj): Linear(in_features=1024, out_features=1024, bias=False)
169
+ (proj): Linear(in_features=1024, out_features=1024, bias=False)
170
+ (rope_impl): CAPI3DRoPE()
171
+ )
172
+ )
173
+ (residual2): EfficientResidual(
174
+ (norm): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
175
+ (fn): MLP(
176
+ (fc1): Linear(in_features=1024, out_features=4096, bias=False)
177
+ (act): GELU(approximate='none')
178
+ (fc2): Linear(in_features=4096, out_features=1024, bias=False)
179
+ (drop): Dropout(p=0.0, inplace=False)
180
+ )
181
+ )
182
+ )
183
+ )
184
+ (st_rope): CAPI3DRoPE()
185
+ )
186
+ )
187
+ [INFO ][2026-05-10 12:23:56][root ][init_video_model ] UJEPAsidePredictorMultiSeqWrapper(
188
+ (backbone): PredictorV2(
189
+ (predictor_embed): Linear(in_features=1024, out_features=384, bias=True)
190
+ (mask_tokens): ParameterList(
191
+ (0): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
192
+ (1): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
193
+ (2): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
194
+ (3): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
195
+ )
196
+ (predictor_blocks): ModuleList(
197
+ (0-5): 6 x Block(
198
+ (residual1): EfficientResidual(
199
+ (norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
200
+ (fn): Attention(
201
+ (q_proj): Linear(in_features=384, out_features=384, bias=False)
202
+ (k_proj): Linear(in_features=384, out_features=384, bias=False)
203
+ (v_proj): Linear(in_features=384, out_features=384, bias=False)
204
+ (proj): Linear(in_features=384, out_features=384, bias=False)
205
+ (rope): Rope()
206
+ )
207
+ )
208
+ (residual2): EfficientResidual(
209
+ (norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
210
+ (fn): MLP(
211
+ (fc1): Linear(in_features=384, out_features=1536, bias=False)
212
+ (act): GELU(approximate='none')
213
+ (fc2): Linear(in_features=1536, out_features=384, bias=False)
214
+ (drop): Dropout(p=0.0, inplace=False)
215
+ )
216
+ )
217
+ )
218
+ )
219
+ (predictor_norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
220
+ (predictor_proj): Linear(in_features=384, out_features=1024, bias=True)
221
+ )
222
+ )
223
+ [INFO ][2026-05-10 12:23:56][root ][init_video_model ] MultiSeqWrapper(
224
+ (backbone): Frozen2DTargetWrapper(
225
+ (backbone): Eva(
226
+ (patch_embed): PatchEmbed(
227
+ (proj): Conv2d(3, 1024, kernel_size=(14, 14), stride=(14, 14))
228
+ (norm): Identity()
229
+ )
230
+ (pos_drop): Dropout(p=0.0, inplace=False)
231
+ (norm_pre): Identity()
232
+ (blocks): ModuleList(
233
+ (0-23): 24 x EvaBlock(
234
+ (norm1): RMSNorm((1024,), eps=1e-05, elementwise_affine=True)
235
+ (attn): EvaAttention(
236
+ (qkv): Linear(in_features=1024, out_features=3072, bias=False)
237
+ (q_norm): Identity()
238
+ (k_norm): Identity()
239
+ (attn_drop): Dropout(p=0.0, inplace=False)
240
+ (norm): Identity()
241
+ (proj): Linear(in_features=1024, out_features=1024, bias=True)
242
+ (proj_drop): Dropout(p=0.0, inplace=False)
243
+ )
244
+ (drop_path1): Identity()
245
+ (norm2): RMSNorm((1024,), eps=1e-05, elementwise_affine=True)
246
+ (mlp): Mlp(
247
+ (fc1): Linear(in_features=1024, out_features=4096, bias=True)
248
+ (act): GELU(approximate='none')
249
+ (drop1): Dropout(p=0.0, inplace=False)
250
+ (norm): Identity()
251
+ (fc2): Linear(in_features=4096, out_features=1024, bias=True)
252
+ (drop2): Dropout(p=0.0, inplace=False)
253
+ )
254
+ (drop_path2): Identity()
255
+ )
256
+ )
257
+ (norm): RMSNorm((1024,), eps=1e-05, elementwise_affine=True)
258
+ (fc_norm): Identity()
259
+ (head_drop): Dropout(p=0.0, inplace=False)
260
+ (head): Identity()
261
+ (rope): _CapiPatchRoPE()
262
+ )
263
+ )
264
+ )
265
+ [INFO ][2026-05-10 12:23:56][root ][init_video_model ] Encoder number of parameters: 302399488
266
+ [INFO ][2026-05-10 12:23:56][root ][init_video_model ] Predictor number of parameters: 11416192
267
+ [INFO ][2026-05-10 12:23:56][root ][init_video_model ] Target encoder number of parameters: 0
268
+ [INFO ][2026-05-10 12:24:07][root ][make_videodataset ] VideoDataset dataset created
269
+ [INFO ][2026-05-10 12:24:07][WeightedSampler ][__init__ ] Using DistributedWeightedSampler with rank 7 / 16
270
+ [INFO ][2026-05-10 12:24:07][root ][make_videodataset ] VideoDataset unsupervised data loader created
271
+ [INFO ][2026-05-10 12:24:07][app.vjepa.train ][main ] iterations per epoch/dataset length: 300/399
272
+ [INFO ][2026-05-10 12:24:07][app.vjepa.train ][main ] Wrapping models in DDP (rank 7)...
273
+ [INFO ][2026-05-10 12:24:09][app.vjepa.train ][main ] Initializing loader...
274
+ submitit ERROR (2026-05-10 16:19:50,186) - Submitted job triggered an exception
275
+ [ERROR ][2026-05-10 16:19:50][submitit ][process_job ] Submitted job triggered an exception
276
+ [ERROR ][2026-05-10 16:19:51][app.vjepa.train ][_save_crash ] Saved crash log to /scratch-shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/crash_rank7.log
vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/job_22614766/22614766_8_log.err ADDED
@@ -0,0 +1,68 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/timm/models/layers/__init__.py:49: FutureWarning: Importing from timm.models.layers is deprecated, please import via timm.layers
2
+ warnings.warn(f"Importing from {__name__} is deprecated, please import via timm.layers", FutureWarning)
3
+ [rank8]:[W510 12:20:54.584744077 ProcessGroupNCCL.cpp:5138] Guessing device ID based on global rank. This can cause a hang if rank to GPU mapping is heterogeneous. You can specify device_id in init_process_group()
4
+ /gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/vjepa/utils.py:796: FutureWarning: `torch.cuda.amp.GradScaler(args...)` is deprecated. Please use `torch.amp.GradScaler('cuda', args...)` instead.
5
+ scaler = torch.cuda.amp.GradScaler() if mixed_precision else None
6
+ /gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/vjepa/train.py:955: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
7
+ with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
8
+ submitit ERROR (2026-05-10 16:19:48,627) - Submitted job triggered an exception
9
+ [rank8]:[W510 16:19:49.713101052 TCPStore.cpp:125] [c10d] recvValue failed on SocketImpl(fd=3, addr=[gcn85.local.snellius.surf.nl]:48716, remote=[gcn80.local.snellius.surf.nl]:37129): Failed to recv, got 0 bytes. Connection was likely closed. Did the remote server shutdown or crash?
10
+ Exception raised from recvBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:682 (most recent call first):
11
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x14ec81298fdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
12
+ frame #1: <unknown function> + 0x6a3325d (0x14ecc542725d in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
13
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x273 (0x14ecc54251f3 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
14
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x14ec82507eec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
15
+ frame #4: <unknown function> + 0xed164 (0x14edc1971164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
16
+ frame #5: <unknown function> + 0x8a19a (0x14edd23e319a in /lib64/libc.so.6)
17
+ frame #6: <unknown function> + 0x10f100 (0x14edd2468100 in /lib64/libc.so.6)
18
+
19
+ [rank8]:[W510 16:19:49.756860305 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 8] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Failed to recv, got 0 bytes. Connection was likely closed. Did the remote server shutdown or crash?
20
+ [rank8]:[W510 16:19:50.757024745 TCPStore.cpp:106] [c10d] sendBytes failed on SocketImpl(fd=3, addr=[gcn85.local.snellius.surf.nl]:48716, remote=[gcn80.local.snellius.surf.nl]:37129): Broken pipe
21
+ Exception raised from sendBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:653 (most recent call first):
22
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x14ec81298fdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
23
+ frame #1: <unknown function> + 0x6a326d1 (0x14ecc54266d1 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
24
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x24d (0x14ecc54251cd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
25
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x14ec82507eec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
26
+ frame #4: <unknown function> + 0xed164 (0x14edc1971164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
27
+ frame #5: <unknown function> + 0x8a19a (0x14edd23e319a in /lib64/libc.so.6)
28
+ frame #6: <unknown function> + 0x10f100 (0x14edd2468100 in /lib64/libc.so.6)
29
+
30
+ [rank8]:[W510 16:19:50.760610632 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 8] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Broken pipe
31
+ Traceback (most recent call last):
32
+ File "<frozen runpy>", line 198, in _run_module_as_main
33
+ File "<frozen runpy>", line 88, in _run_code
34
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/_submit.py", line 11, in <module>
35
+ submitit_main()
36
+ ~~~~~~~~~~~~~^^
37
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 76, in submitit_main
38
+ process_job(args.folder)
39
+ ~~~~~~~~~~~^^^^^^^^^^^^^
40
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 69, in process_job
41
+ raise error
42
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 55, in process_job
43
+ result = delayed.result()
44
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/utils.py", line 137, in result
45
+ self._result = self.function(*self.args, **self.kwargs)
46
+ ~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^
47
+ File "/gpfs/home2/dcanez/vd/app/main_distributed.py", line 102, in __call__
48
+ app_main(app, args=params, resume_preempt=resume_preempt)
49
+ ~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
50
+ File "/gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/scaffold.py", line 17, in main
51
+ return importlib.import_module(f"app.{app}.train").main(args=args, resume_preempt=resume_preempt)
52
+ ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
53
+ File "/gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/vjepa/train.py", line 1189, in main
54
+ assert not np.isnan(loss), "loss is nan"
55
+ ^^^^^^^^^^^^^^^^^^
56
+ AssertionError: loss is nan
57
+ [rank8]:[W510 16:19:51.760810102 TCPStore.cpp:106] [c10d] sendBytes failed on SocketImpl(fd=3, addr=[gcn85.local.snellius.surf.nl]:48716, remote=[gcn80.local.snellius.surf.nl]:37129): Broken pipe
58
+ Exception raised from sendBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:653 (most recent call first):
59
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x14ec81298fdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
60
+ frame #1: <unknown function> + 0x6a326d1 (0x14ecc54266d1 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
61
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x24d (0x14ecc54251cd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
62
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x14ec82507eec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
63
+ frame #4: <unknown function> + 0xed164 (0x14edc1971164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
64
+ frame #5: <unknown function> + 0x8a19a (0x14edd23e319a in /lib64/libc.so.6)
65
+ frame #6: <unknown function> + 0x10f100 (0x14edd2468100 in /lib64/libc.so.6)
66
+
67
+ [rank8]:[W510 16:19:51.763104707 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 8] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Broken pipe
68
+ [rank8]:[W510 16:19:51.202810034 ProcessGroupNCCL.cpp:1553] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())
vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/job_22614766/22614766_8_log.out ADDED
@@ -0,0 +1,280 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ submitit INFO (2026-05-10 12:14:34,049) - Starting with JobEnvironment(job_id=22614766, hostname=gcn85.local.snellius.surf.nl, local_rank=0(4), node=2(4), global_rank=8(16))
2
+ submitit INFO (2026-05-10 12:14:34,049) - Loading pickle: /scratch-shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/job_22614766/22614766_submitted.pkl
3
+ INFO:root:loaded pretrain params...
4
+ { 'app': 'vjepa',
5
+ 'cpus_per_task': 16,
6
+ 'data': { 'batch_size': 64,
7
+ 'crop_size': 224,
8
+ 'dataset_fpcs': [16, 16],
9
+ 'dataset_type': 'VideoDataset',
10
+ 'datasets': [ '/scratch-shared/dcanez/data/kinetics/k400/train.csv',
11
+ '/scratch-shared/dcanez/data/ssv2/train.csv'],
12
+ 'datasets_weights': [0.65, 0.35],
13
+ 'fps': 4,
14
+ 'num_workers': 10,
15
+ 'patch_size': 14,
16
+ 'persistent_workers': True,
17
+ 'pin_mem': True,
18
+ 'stage': [ { 'dest': 'kinetics_240',
19
+ 'format': 'targz_parts',
20
+ 'src': '/scratch-shared/dcanez/data/kinetics/k400/tars_240/'},
21
+ { 'dest': 'ssv2',
22
+ 'format': 'multipart_tar',
23
+ 'src': '/scratch-nvme/ml-datasets/something-something-v2/'}],
24
+ 'tubelet_size': 1},
25
+ 'data_aug': { 'auto_augment': False,
26
+ 'motion_shift': False,
27
+ 'random_resize_aspect_ratio': [0.75, 1.35],
28
+ 'random_resize_scale': [0.3, 1.0],
29
+ 'reprob': 0.0},
30
+ 'folder': '/scratch-shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable',
31
+ 'loss': {'loss_exp': 1.0},
32
+ 'mask': [ { 'aspect_ratio': [0.75, 1.5],
33
+ 'full_complement': False,
34
+ 'max_keep': None,
35
+ 'max_temporal_keep': 1.0,
36
+ 'num_blocks': 8,
37
+ 'spatial_scale': [0.15, 0.15],
38
+ 'temporal_scale': [1.0, 1.0]},
39
+ { 'aspect_ratio': [0.75, 1.5],
40
+ 'full_complement': False,
41
+ 'max_keep': None,
42
+ 'max_temporal_keep': 1.0,
43
+ 'num_blocks': 2,
44
+ 'spatial_scale': [0.7, 0.7],
45
+ 'temporal_scale': [1.0, 1.0]}],
46
+ 'mem_per_gpu': '180G',
47
+ 'meta': { 'dtype': 'bfloat16',
48
+ 'knn_eval_epoch0': False,
49
+ 'knn_eval_freq': 5,
50
+ 'knn_eval_presets': [ { 'config': { 'batch_size': 64,
51
+ 'dataset_train': '/scratch-shared/mdorkenw/ucf101/ucfTrainTestlist/train.csv',
52
+ 'dataset_val': '/scratch-shared/mdorkenw/ucf101/ucfTrainTestlist/val.csv',
53
+ 'eval_videos_per_class': 25,
54
+ 'num_workers': 8,
55
+ 'pool_type': 'slot_temporal_concat',
56
+ 'train_videos_per_class': 100},
57
+ 'preset': 'ucf101'},
58
+ { 'config': { 'batch_size': 64,
59
+ 'dataset_train': '/scratch-shared/mdorkenw/20bn-something-something-v2/something-something-v2-annotations/train_coarse10.csv',
60
+ 'dataset_val': '/scratch-shared/mdorkenw/20bn-something-something-v2/something-something-v2-annotations/val_coarse10.csv',
61
+ 'eval_videos_per_class': 100,
62
+ 'linear_probe': True,
63
+ 'num_workers': 8,
64
+ 'pool_type': 'slot_temporal_concat',
65
+ 'train_videos_per_class': 500},
66
+ 'preset': 'ssv2_coarse10'}],
67
+ 'load_checkpoint': True,
68
+ 'read_checkpoint': None,
69
+ 'save_every_freq': 5,
70
+ 'seed': 239,
71
+ 'use_sdpa': True,
72
+ 'use_wandb': True,
73
+ 'wandb_project': 'vjepa_ujepaside'},
74
+ 'metrics': {'sigreg': {}, 'std': {}},
75
+ 'model': { 'model_name': 'ujepaside_large_patch14_capi_lvd1689m',
76
+ 'pred_depth': 6,
77
+ 'pred_embed_dim': 384,
78
+ 'pred_num_heads': 12,
79
+ 'predictor': 'v2_cross',
80
+ 'st_causal': False,
81
+ 'st_drop_path': 0.2,
82
+ 'st_flex_enable': False,
83
+ 'st_layer_scale_init': 1e-05,
84
+ 'st_num_slots': 16,
85
+ 'st_slots_causal_within_frame': False,
86
+ 'target_kind': 'frozen_2d',
87
+ 'target_type': 'vit_large_patch14_capi.lvd1689m',
88
+ 'temporal_spacing': 1.0,
89
+ 'uniform_power': True,
90
+ 'use_activation_checkpointing': True,
91
+ 'use_mask_tokens': True,
92
+ 'use_rope': True,
93
+ 'use_sdpa': True,
94
+ 'zero_init_mask_tokens': True},
95
+ 'nodes': 4,
96
+ 'optimization': { 'clip_grad': 3.0,
97
+ 'ema': [0.99925, 0.99925],
98
+ 'epochs': 100,
99
+ 'final_lr': 0.0001,
100
+ 'final_weight_decay': 0.04,
101
+ 'ipe': 300,
102
+ 'ipe_scale': 1.0,
103
+ 'lr': 0.0005,
104
+ 'start_lr': 0.0001,
105
+ 'warmup': 10,
106
+ 'weight_decay': 0.04},
107
+ 'tasks_per_node': 4}
108
+ INFO:root:Running pre-training of app: vjepa
109
+ [INFO ][2026-05-10 12:16:54][app.vjepa.train ][main ] which_dtype='bfloat16'
110
+ [INFO ][2026-05-10 12:16:54][app.vjepa.train ][main ] Disabling persistent_workers (incompatible with KNN eval)
111
+ [INFO ][2026-05-10 12:16:54][app.vjepa.train ][main ] NCCL_SOCKET_IFNAME=eno
112
+ [INFO ][2026-05-10 12:16:56][app.vjepa.train ][main ] Initialized (rank/world-size) 8/16, tasks_per_node=4
113
+ [INFO ][2026-05-10 12:16:56][root ][stage_datasets ] [local_rank 0/4] Staging kinetics_240 (targz_parts)
114
+ [INFO ][2026-05-10 12:16:56][root ][_stage_targz_parts ] [rank 0] Extracting 26/103 tar.gz parts to /scratch-node/dcanez.22614766/kinetics_240
115
+ [INFO ][2026-05-10 12:17:10][root ][_stage_targz_parts ] [local_rank 0] Extracted 2/26 parts
116
+ [INFO ][2026-05-10 12:17:23][root ][_stage_targz_parts ] [local_rank 0] Extracted 4/26 parts
117
+ [INFO ][2026-05-10 12:17:37][root ][_stage_targz_parts ] [local_rank 0] Extracted 6/26 parts
118
+ [INFO ][2026-05-10 12:17:51][root ][_stage_targz_parts ] [local_rank 0] Extracted 8/26 parts
119
+ [INFO ][2026-05-10 12:18:05][root ][_stage_targz_parts ] [local_rank 0] Extracted 10/26 parts
120
+ [INFO ][2026-05-10 12:18:19][root ][_stage_targz_parts ] [local_rank 0] Extracted 12/26 parts
121
+ [INFO ][2026-05-10 12:18:33][root ][_stage_targz_parts ] [local_rank 0] Extracted 14/26 parts
122
+ [INFO ][2026-05-10 12:18:47][root ][_stage_targz_parts ] [local_rank 0] Extracted 16/26 parts
123
+ [INFO ][2026-05-10 12:19:01][root ][_stage_targz_parts ] [local_rank 0] Extracted 18/26 parts
124
+ [INFO ][2026-05-10 12:19:16][root ][_stage_targz_parts ] [local_rank 0] Extracted 20/26 parts
125
+ [INFO ][2026-05-10 12:19:30][root ][_stage_targz_parts ] [local_rank 0] Extracted 22/26 parts
126
+ [INFO ][2026-05-10 12:19:44][root ][_stage_targz_parts ] [local_rank 0] Extracted 24/26 parts
127
+ [INFO ][2026-05-10 12:19:58][root ][_stage_targz_parts ] [local_rank 0] Extracted 26/26 parts
128
+ [INFO ][2026-05-10 12:19:58][root ][stage_datasets ] [local_rank 0/4] Staging ssv2 (multipart_tar)
129
+ [INFO ][2026-05-10 12:19:58][root ][_stage_multipart_tar ] [rank 0] Extracting multipart tar (2 files) to /scratch-node/dcanez.22614766/ssv2
130
+ [INFO ][2026-05-10 12:21:02][root ][stage_datasets ] Data staging completed in 245.8s (4.1min)
131
+ [INFO ][2026-05-10 12:21:02][root ][_rewrite_csv ] Wrote local CSV: /scratch-node/dcanez.22614766/kinetics_240/train.csv (239789 entries)
132
+ [INFO ][2026-05-10 12:21:03][root ][_rewrite_csv ] Wrote local CSV: /scratch-node/dcanez.22614766/ssv2/train.csv (168913 entries)
133
+ [INFO ][2026-05-10 12:21:03][app.vjepa.train ][main ] Staged datasets: ['/scratch-node/dcanez.22614766/kinetics_240/train.csv', '/scratch-node/dcanez.22614766/ssv2/train.csv']
134
+ [INFO ][2026-05-10 12:23:31][experiments.stmodels.vision_transformers_v3][vit_large_patch14_capi_lvd1689m] Loading pretrained weights for vit_large_patch14_capi_lvd1689m from https://dl.fbaipublicfiles.com/capi/capi_vitl14_lvd.pth
135
+ [INFO ][2026-05-10 12:23:55][root ][_build_frozen_2d_target ] Loaded frozen_2d target from timm: vit_large_patch14_capi.lvd1689m
136
+ [INFO ][2026-05-10 12:23:55][root ][init_video_model ] ViTMultiSeqWrapper(
137
+ (backbone): UJEPAside(
138
+ (patch_embed): PatchEmbed3D(
139
+ (proj): Conv3d(3, 1024, kernel_size=(1, 14, 14), stride=(1, 14, 14))
140
+ )
141
+ (rope): CAPI2DRoPE()
142
+ (blocks): ModuleList(
143
+ (0-23): 24 x Block(
144
+ (norm1): RMSNorm((1024,), eps=1e-05, elementwise_affine=True)
145
+ (rope_impl): CAPI2DRoPE()
146
+ (norm2): RMSNorm((1024,), eps=1e-05, elementwise_affine=True)
147
+ (drop_path1): Identity()
148
+ (drop_path2): Identity()
149
+ (attn): Attention(
150
+ (qkv): Linear(in_features=1024, out_features=3072, bias=False)
151
+ (attn_drop): Dropout(p=0.0, inplace=False)
152
+ (proj): Linear(in_features=1024, out_features=1024, bias=False)
153
+ (proj_drop): Dropout(p=0.0, inplace=False)
154
+ (rope_impl): CAPI2DRoPE()
155
+ )
156
+ (mlp): MLP(
157
+ (fc1): Linear(in_features=1024, out_features=4096, bias=False)
158
+ (act): GELU(approximate='none')
159
+ (fc2): Linear(in_features=4096, out_features=1024, bias=False)
160
+ (drop): Dropout(p=0.0, inplace=False)
161
+ )
162
+ )
163
+ )
164
+ (norm): RMSNorm((1024,), eps=1e-05, elementwise_affine=False)
165
+ (st_blocks): ModuleList(
166
+ (0-23): 24 x Block(
167
+ (residual1): EfficientResidual(
168
+ (norm): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
169
+ (fn): Attention(
170
+ (q_proj): Linear(in_features=1024, out_features=1024, bias=False)
171
+ (k_proj): Linear(in_features=1024, out_features=1024, bias=False)
172
+ (v_proj): Linear(in_features=1024, out_features=1024, bias=False)
173
+ (proj): Linear(in_features=1024, out_features=1024, bias=False)
174
+ (rope_impl): CAPI3DRoPE()
175
+ )
176
+ )
177
+ (residual2): EfficientResidual(
178
+ (norm): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
179
+ (fn): MLP(
180
+ (fc1): Linear(in_features=1024, out_features=4096, bias=False)
181
+ (act): GELU(approximate='none')
182
+ (fc2): Linear(in_features=4096, out_features=1024, bias=False)
183
+ (drop): Dropout(p=0.0, inplace=False)
184
+ )
185
+ )
186
+ )
187
+ )
188
+ (st_rope): CAPI3DRoPE()
189
+ )
190
+ )
191
+ [INFO ][2026-05-10 12:23:55][root ][init_video_model ] UJEPAsidePredictorMultiSeqWrapper(
192
+ (backbone): PredictorV2(
193
+ (predictor_embed): Linear(in_features=1024, out_features=384, bias=True)
194
+ (mask_tokens): ParameterList(
195
+ (0): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
196
+ (1): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
197
+ (2): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
198
+ (3): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
199
+ )
200
+ (predictor_blocks): ModuleList(
201
+ (0-5): 6 x Block(
202
+ (residual1): EfficientResidual(
203
+ (norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
204
+ (fn): Attention(
205
+ (q_proj): Linear(in_features=384, out_features=384, bias=False)
206
+ (k_proj): Linear(in_features=384, out_features=384, bias=False)
207
+ (v_proj): Linear(in_features=384, out_features=384, bias=False)
208
+ (proj): Linear(in_features=384, out_features=384, bias=False)
209
+ (rope): Rope()
210
+ )
211
+ )
212
+ (residual2): EfficientResidual(
213
+ (norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
214
+ (fn): MLP(
215
+ (fc1): Linear(in_features=384, out_features=1536, bias=False)
216
+ (act): GELU(approximate='none')
217
+ (fc2): Linear(in_features=1536, out_features=384, bias=False)
218
+ (drop): Dropout(p=0.0, inplace=False)
219
+ )
220
+ )
221
+ )
222
+ )
223
+ (predictor_norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
224
+ (predictor_proj): Linear(in_features=384, out_features=1024, bias=True)
225
+ )
226
+ )
227
+ [INFO ][2026-05-10 12:23:55][root ][init_video_model ] MultiSeqWrapper(
228
+ (backbone): Frozen2DTargetWrapper(
229
+ (backbone): Eva(
230
+ (patch_embed): PatchEmbed(
231
+ (proj): Conv2d(3, 1024, kernel_size=(14, 14), stride=(14, 14))
232
+ (norm): Identity()
233
+ )
234
+ (pos_drop): Dropout(p=0.0, inplace=False)
235
+ (norm_pre): Identity()
236
+ (blocks): ModuleList(
237
+ (0-23): 24 x EvaBlock(
238
+ (norm1): RMSNorm((1024,), eps=1e-05, elementwise_affine=True)
239
+ (attn): EvaAttention(
240
+ (qkv): Linear(in_features=1024, out_features=3072, bias=False)
241
+ (q_norm): Identity()
242
+ (k_norm): Identity()
243
+ (attn_drop): Dropout(p=0.0, inplace=False)
244
+ (norm): Identity()
245
+ (proj): Linear(in_features=1024, out_features=1024, bias=True)
246
+ (proj_drop): Dropout(p=0.0, inplace=False)
247
+ )
248
+ (drop_path1): Identity()
249
+ (norm2): RMSNorm((1024,), eps=1e-05, elementwise_affine=True)
250
+ (mlp): Mlp(
251
+ (fc1): Linear(in_features=1024, out_features=4096, bias=True)
252
+ (act): GELU(approximate='none')
253
+ (drop1): Dropout(p=0.0, inplace=False)
254
+ (norm): Identity()
255
+ (fc2): Linear(in_features=4096, out_features=1024, bias=True)
256
+ (drop2): Dropout(p=0.0, inplace=False)
257
+ )
258
+ (drop_path2): Identity()
259
+ )
260
+ )
261
+ (norm): RMSNorm((1024,), eps=1e-05, elementwise_affine=True)
262
+ (fc_norm): Identity()
263
+ (head_drop): Dropout(p=0.0, inplace=False)
264
+ (head): Identity()
265
+ (rope): _CapiPatchRoPE()
266
+ )
267
+ )
268
+ )
269
+ [INFO ][2026-05-10 12:23:55][root ][init_video_model ] Encoder number of parameters: 302399488
270
+ [INFO ][2026-05-10 12:23:55][root ][init_video_model ] Predictor number of parameters: 11416192
271
+ [INFO ][2026-05-10 12:23:55][root ][init_video_model ] Target encoder number of parameters: 0
272
+ [INFO ][2026-05-10 12:23:56][root ][make_videodataset ] VideoDataset dataset created
273
+ [INFO ][2026-05-10 12:23:56][WeightedSampler ][__init__ ] Using DistributedWeightedSampler with rank 8 / 16
274
+ [INFO ][2026-05-10 12:23:56][root ][make_videodataset ] VideoDataset unsupervised data loader created
275
+ [INFO ][2026-05-10 12:23:56][app.vjepa.train ][main ] iterations per epoch/dataset length: 300/399
276
+ [INFO ][2026-05-10 12:23:56][app.vjepa.train ][main ] Wrapping models in DDP (rank 8)...
277
+ [INFO ][2026-05-10 12:24:08][app.vjepa.train ][main ] Initializing loader...
278
+ submitit ERROR (2026-05-10 16:19:48,627) - Submitted job triggered an exception
279
+ [ERROR ][2026-05-10 16:19:48][submitit ][process_job ] Submitted job triggered an exception
280
+ [ERROR ][2026-05-10 16:19:51][app.vjepa.train ][_save_crash ] Saved crash log to /scratch-shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/crash_rank8.log
vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/job_22614766/22614766_9_log.err ADDED
@@ -0,0 +1,191 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/timm/models/layers/__init__.py:49: FutureWarning: Importing from timm.models.layers is deprecated, please import via timm.layers
2
+ warnings.warn(f"Importing from {__name__} is deprecated, please import via timm.layers", FutureWarning)
3
+ [rank9]:[W510 12:19:58.520496896 ProcessGroupNCCL.cpp:5138] Guessing device ID based on global rank. This can cause a hang if rank to GPU mapping is heterogeneous. You can specify device_id in init_process_group()
4
+ /gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/vjepa/utils.py:796: FutureWarning: `torch.cuda.amp.GradScaler(args...)` is deprecated. Please use `torch.amp.GradScaler('cuda', args...)` instead.
5
+ scaler = torch.cuda.amp.GradScaler() if mixed_precision else None
6
+ /gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/vjepa/train.py:955: FutureWarning: `torch.cuda.amp.autocast(args...)` is deprecated. Please use `torch.amp.autocast('cuda', args...)` instead.
7
+ with torch.cuda.amp.autocast(dtype=dtype, enabled=mixed_precision):
8
+ [rank9]:[W510 16:19:49.713098522 TCPStore.cpp:125] [c10d] recvValue failed on SocketImpl(fd=3, addr=[gcn85.local.snellius.surf.nl]:48682, remote=[gcn80.local.snellius.surf.nl]:37129): Failed to recv, got 0 bytes. Connection was likely closed. Did the remote server shutdown or crash?
9
+ Exception raised from recvBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:682 (most recent call first):
10
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x148fe9d7dfdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
11
+ frame #1: <unknown function> + 0x6a3325d (0x14902df0c25d in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
12
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x273 (0x14902df0a1f3 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
13
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x148feafeceec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
14
+ frame #4: <unknown function> + 0xed164 (0x14912a456164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
15
+ frame #5: <unknown function> + 0x8a19a (0x14913aec819a in /lib64/libc.so.6)
16
+ frame #6: <unknown function> + 0x10f100 (0x14913af4d100 in /lib64/libc.so.6)
17
+
18
+ [rank9]:[W510 16:19:49.756857925 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 9] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Failed to recv, got 0 bytes. Connection was likely closed. Did the remote server shutdown or crash?
19
+ [rank9]:[W510 16:19:50.756994415 TCPStore.cpp:106] [c10d] sendBytes failed on SocketImpl(fd=3, addr=[gcn85.local.snellius.surf.nl]:48682, remote=[gcn80.local.snellius.surf.nl]:37129): Broken pipe
20
+ Exception raised from sendBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:653 (most recent call first):
21
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x148fe9d7dfdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
22
+ frame #1: <unknown function> + 0x6a326d1 (0x14902df0b6d1 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
23
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x24d (0x14902df0a1cd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
24
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x148feafeceec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
25
+ frame #4: <unknown function> + 0xed164 (0x14912a456164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
26
+ frame #5: <unknown function> + 0x8a19a (0x14913aec819a in /lib64/libc.so.6)
27
+ frame #6: <unknown function> + 0x10f100 (0x14913af4d100 in /lib64/libc.so.6)
28
+
29
+ [rank9]:[W510 16:19:50.760565092 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 9] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Broken pipe
30
+ [rank9]:[W510 16:19:51.760703343 TCPStore.cpp:106] [c10d] sendBytes failed on SocketImpl(fd=3, addr=[gcn85.local.snellius.surf.nl]:48682, remote=[gcn80.local.snellius.surf.nl]:37129): Broken pipe
31
+ Exception raised from sendBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:653 (most recent call first):
32
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x148fe9d7dfdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
33
+ frame #1: <unknown function> + 0x6a326d1 (0x14902df0b6d1 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
34
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x24d (0x14902df0a1cd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
35
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x148feafeceec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
36
+ frame #4: <unknown function> + 0xed164 (0x14912a456164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
37
+ frame #5: <unknown function> + 0x8a19a (0x14913aec819a in /lib64/libc.so.6)
38
+ frame #6: <unknown function> + 0x10f100 (0x14913af4d100 in /lib64/libc.so.6)
39
+
40
+ [rank9]:[W510 16:19:51.762962548 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 9] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Broken pipe
41
+ [rank9]:[W510 16:19:52.763081018 TCPStore.cpp:106] [c10d] sendBytes failed on SocketImpl(fd=3, addr=[gcn85.local.snellius.surf.nl]:48682, remote=[gcn80.local.snellius.surf.nl]:37129): Broken pipe
42
+ Exception raised from sendBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:653 (most recent call first):
43
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x148fe9d7dfdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
44
+ frame #1: <unknown function> + 0x6a326d1 (0x14902df0b6d1 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
45
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x24d (0x14902df0a1cd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
46
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x148feafeceec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
47
+ frame #4: <unknown function> + 0xed164 (0x14912a456164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
48
+ frame #5: <unknown function> + 0x8a19a (0x14913aec819a in /lib64/libc.so.6)
49
+ frame #6: <unknown function> + 0x10f100 (0x14913af4d100 in /lib64/libc.so.6)
50
+
51
+ [rank9]:[W510 16:19:52.765328454 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 9] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Broken pipe
52
+ [rank9]:[W510 16:19:53.765437783 TCPStore.cpp:106] [c10d] sendBytes failed on SocketImpl(fd=3, addr=[gcn85.local.snellius.surf.nl]:48682, remote=[gcn80.local.snellius.surf.nl]:37129): Broken pipe
53
+ Exception raised from sendBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:653 (most recent call first):
54
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x148fe9d7dfdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
55
+ frame #1: <unknown function> + 0x6a326d1 (0x14902df0b6d1 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
56
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x24d (0x14902df0a1cd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
57
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x148feafeceec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
58
+ frame #4: <unknown function> + 0xed164 (0x14912a456164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
59
+ frame #5: <unknown function> + 0x8a19a (0x14913aec819a in /lib64/libc.so.6)
60
+ frame #6: <unknown function> + 0x10f100 (0x14913af4d100 in /lib64/libc.so.6)
61
+
62
+ [rank9]:[W510 16:19:53.767857408 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 9] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Broken pipe
63
+ [rank9]:[W510 16:19:54.767966367 TCPStore.cpp:106] [c10d] sendBytes failed on SocketImpl(fd=3, addr=[gcn85.local.snellius.surf.nl]:48682, remote=[gcn80.local.snellius.surf.nl]:37129): Broken pipe
64
+ Exception raised from sendBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:653 (most recent call first):
65
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x148fe9d7dfdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
66
+ frame #1: <unknown function> + 0x6a326d1 (0x14902df0b6d1 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
67
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x24d (0x14902df0a1cd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
68
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x148feafeceec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
69
+ frame #4: <unknown function> + 0xed164 (0x14912a456164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
70
+ frame #5: <unknown function> + 0x8a19a (0x14913aec819a in /lib64/libc.so.6)
71
+ frame #6: <unknown function> + 0x10f100 (0x14913af4d100 in /lib64/libc.so.6)
72
+
73
+ [rank9]:[W510 16:19:54.770239623 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 9] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Broken pipe
74
+ [rank9]:[W510 16:19:55.770352823 TCPStore.cpp:106] [c10d] sendBytes failed on SocketImpl(fd=3, addr=[gcn85.local.snellius.surf.nl]:48682, remote=[gcn80.local.snellius.surf.nl]:37129): Broken pipe
75
+ Exception raised from sendBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:653 (most recent call first):
76
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x148fe9d7dfdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
77
+ frame #1: <unknown function> + 0x6a326d1 (0x14902df0b6d1 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
78
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x24d (0x14902df0a1cd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
79
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x148feafeceec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
80
+ frame #4: <unknown function> + 0xed164 (0x14912a456164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
81
+ frame #5: <unknown function> + 0x8a19a (0x14913aec819a in /lib64/libc.so.6)
82
+ frame #6: <unknown function> + 0x10f100 (0x14913af4d100 in /lib64/libc.so.6)
83
+
84
+ [rank9]:[W510 16:19:55.772627358 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 9] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Broken pipe
85
+ [rank9]:[W510 16:19:56.772765518 TCPStore.cpp:106] [c10d] sendBytes failed on SocketImpl(fd=3, addr=[gcn85.local.snellius.surf.nl]:48682, remote=[gcn80.local.snellius.surf.nl]:37129): Broken pipe
86
+ Exception raised from sendBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:653 (most recent call first):
87
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x148fe9d7dfdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
88
+ frame #1: <unknown function> + 0x6a326d1 (0x14902df0b6d1 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
89
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x24d (0x14902df0a1cd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
90
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x148feafeceec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
91
+ frame #4: <unknown function> + 0xed164 (0x14912a456164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
92
+ frame #5: <unknown function> + 0x8a19a (0x14913aec819a in /lib64/libc.so.6)
93
+ frame #6: <unknown function> + 0x10f100 (0x14913af4d100 in /lib64/libc.so.6)
94
+
95
+ [rank9]:[W510 16:19:56.775215702 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 9] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Broken pipe
96
+ [rank9]:[W510 16:19:57.775350242 TCPStore.cpp:106] [c10d] sendBytes failed on SocketImpl(fd=3, addr=[gcn85.local.snellius.surf.nl]:48682, remote=[gcn80.local.snellius.surf.nl]:37129): Broken pipe
97
+ Exception raised from sendBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:653 (most recent call first):
98
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x148fe9d7dfdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
99
+ frame #1: <unknown function> + 0x6a326d1 (0x14902df0b6d1 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
100
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x24d (0x14902df0a1cd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
101
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x148feafeceec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
102
+ frame #4: <unknown function> + 0xed164 (0x14912a456164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
103
+ frame #5: <unknown function> + 0x8a19a (0x14913aec819a in /lib64/libc.so.6)
104
+ frame #6: <unknown function> + 0x10f100 (0x14913af4d100 in /lib64/libc.so.6)
105
+
106
+ [rank9]:[W510 16:19:57.777628537 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 9] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Broken pipe
107
+ [rank9]:[W510 16:19:58.777757757 TCPStore.cpp:106] [c10d] sendBytes failed on SocketImpl(fd=3, addr=[gcn85.local.snellius.surf.nl]:48682, remote=[gcn80.local.snellius.surf.nl]:37129): Broken pipe
108
+ Exception raised from sendBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:653 (most recent call first):
109
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x148fe9d7dfdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
110
+ frame #1: <unknown function> + 0x6a326d1 (0x14902df0b6d1 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
111
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x24d (0x14902df0a1cd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
112
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x148feafeceec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
113
+ frame #4: <unknown function> + 0xed164 (0x14912a456164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
114
+ frame #5: <unknown function> + 0x8a19a (0x14913aec819a in /lib64/libc.so.6)
115
+ frame #6: <unknown function> + 0x10f100 (0x14913af4d100 in /lib64/libc.so.6)
116
+
117
+ [rank9]:[W510 16:19:58.779960573 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 9] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Broken pipe
118
+ [rank9]:[W510 16:19:59.780087782 TCPStore.cpp:106] [c10d] sendBytes failed on SocketImpl(fd=3, addr=[gcn85.local.snellius.surf.nl]:48682, remote=[gcn80.local.snellius.surf.nl]:37129): Broken pipe
119
+ Exception raised from sendBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:653 (most recent call first):
120
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x148fe9d7dfdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
121
+ frame #1: <unknown function> + 0x6a326d1 (0x14902df0b6d1 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
122
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x24d (0x14902df0a1cd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
123
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x148feafeceec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
124
+ frame #4: <unknown function> + 0xed164 (0x14912a456164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
125
+ frame #5: <unknown function> + 0x8a19a (0x14913aec819a in /lib64/libc.so.6)
126
+ frame #6: <unknown function> + 0x10f100 (0x14913af4d100 in /lib64/libc.so.6)
127
+
128
+ [rank9]:[W510 16:19:59.782502357 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 9] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Broken pipe
129
+ submitit WARNING (2026-05-10 16:20:00,106) - Bypassing signal SIGTERM
130
+ submitit WARNING (2026-05-10 16:20:00,176) - Bypassing signal SIGCONT
131
+ submitit ERROR (2026-05-10 16:20:00,176) - Submitted job triggered an exception
132
+ Traceback (most recent call last):
133
+ File "<frozen runpy>", line 198, in _run_module_as_main
134
+ File "<frozen runpy>", line 88, in _run_code
135
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/_submit.py", line 11, in <module>
136
+ submitit_main()
137
+ ~~~~~~~~~~~~~^^
138
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 76, in submitit_main
139
+ process_job(args.folder)
140
+ ~~~~~~~~~~~^^^^^^^^^^^^^
141
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 69, in process_job
142
+ raise error
143
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/submission.py", line 55, in process_job
144
+ result = delayed.result()
145
+ File "/scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/submitit/core/utils.py", line 137, in result
146
+ self._result = self.function(*self.args, **self.kwargs)
147
+ ~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^
148
+ File "/gpfs/home2/dcanez/vd/app/main_distributed.py", line 102, in __call__
149
+ app_main(app, args=params, resume_preempt=resume_preempt)
150
+ ~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
151
+ File "/gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/scaffold.py", line 17, in main
152
+ return importlib.import_module(f"app.{app}.train").main(args=args, resume_preempt=resume_preempt)
153
+ ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
154
+ File "/gpfs/scratch1/shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/code/app/vjepa/train.py", line 1189, in main
155
+ assert not np.isnan(loss), "loss is nan"
156
+ ^^^^^^^^^^^^^^^^^^
157
+ AssertionError: loss is nan
158
+ [rank9]:[W510 16:20:00.782633087 TCPStore.cpp:106] [c10d] sendBytes failed on SocketImpl(fd=3, addr=[gcn85.local.snellius.surf.nl]:48682, remote=[gcn80.local.snellius.surf.nl]:37129): Broken pipe
159
+ Exception raised from sendBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:653 (most recent call first):
160
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x148fe9d7dfdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
161
+ frame #1: <unknown function> + 0x6a326d1 (0x14902df0b6d1 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
162
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x24d (0x14902df0a1cd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
163
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x148feafeceec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
164
+ frame #4: <unknown function> + 0xed164 (0x14912a456164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
165
+ frame #5: <unknown function> + 0x8a19a (0x14913aec819a in /lib64/libc.so.6)
166
+ frame #6: <unknown function> + 0x10f100 (0x14913af4d100 in /lib64/libc.so.6)
167
+
168
+ [rank9]:[W510 16:20:00.784970142 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 9] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Broken pipe
169
+ [rank9]:[W510 16:20:00.236003417 ProcessGroupNCCL.cpp:1553] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())
170
+ [rank9]:[W510 16:20:01.785104182 TCPStore.cpp:106] [c10d] sendBytes failed on SocketImpl(fd=3, addr=[gcn85.local.snellius.surf.nl]:48682, remote=[gcn80.local.snellius.surf.nl]:37129): Broken pipe
171
+ Exception raised from sendBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:653 (most recent call first):
172
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x148fe9d7dfdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
173
+ frame #1: <unknown function> + 0x6a326d1 (0x14902df0b6d1 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
174
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x24d (0x14902df0a1cd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
175
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x148feafeceec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
176
+ frame #4: <unknown function> + 0xed164 (0x14912a456164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
177
+ frame #5: <unknown function> + 0x8a19a (0x14913aec819a in /lib64/libc.so.6)
178
+ frame #6: <unknown function> + 0x10f100 (0x14913af4d100 in /lib64/libc.so.6)
179
+
180
+ [rank9]:[W510 16:20:01.787343787 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 9] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Broken pipe
181
+ [rank9]:[W510 16:20:02.787484377 TCPStore.cpp:106] [c10d] sendBytes failed on SocketImpl(fd=3, addr=[gcn85.local.snellius.surf.nl]:48682, remote=[gcn80.local.snellius.surf.nl]:37129): Broken pipe
182
+ Exception raised from sendBytes at /pytorch/torch/csrc/distributed/c10d/Utils.hpp:653 (most recent call first):
183
+ frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x148fe9d7dfdd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libc10.so)
184
+ frame #1: <unknown function> + 0x6a326d1 (0x14902df0b6d1 in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
185
+ frame #2: c10d::TCPStore::check(std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > > const&) + 0x24d (0x14902df0a1cd in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cpu.so)
186
+ frame #3: c10d::ProcessGroupNCCL::HeartbeatMonitor::runLoop() + 0x44c (0x148feafeceec in /scratch-shared/dcanez/.cache/uv5/virtualenvs/vd/lib/python3.13/site-packages/torch/lib/libtorch_cuda.so)
187
+ frame #4: <unknown function> + 0xed164 (0x14912a456164 in /sw/arch/RHEL9/EB_production/2025/software/GCCcore/14.2.0/lib64/libstdc++.so.6)
188
+ frame #5: <unknown function> + 0x8a19a (0x14913aec819a in /lib64/libc.so.6)
189
+ frame #6: <unknown function> + 0x10f100 (0x14913af4d100 in /lib64/libc.so.6)
190
+
191
+ [rank9]:[W510 16:20:02.789916262 ProcessGroupNCCL.cpp:1802] [PG ID 0 PG GUID 0(default_pg) Rank 9] Failed to check the "should dump" flag on TCPStore, (maybe TCPStore server has shut down too early), with error: Broken pipe
vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/job_22614766/22614766_9_log.out ADDED
@@ -0,0 +1,280 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ submitit INFO (2026-05-10 12:14:34,049) - Starting with JobEnvironment(job_id=22614766, hostname=gcn85.local.snellius.surf.nl, local_rank=1(4), node=2(4), global_rank=9(16))
2
+ submitit INFO (2026-05-10 12:14:34,049) - Loading pickle: /scratch-shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/job_22614766/22614766_submitted.pkl
3
+ INFO:root:loaded pretrain params...
4
+ { 'app': 'vjepa',
5
+ 'cpus_per_task': 16,
6
+ 'data': { 'batch_size': 64,
7
+ 'crop_size': 224,
8
+ 'dataset_fpcs': [16, 16],
9
+ 'dataset_type': 'VideoDataset',
10
+ 'datasets': [ '/scratch-shared/dcanez/data/kinetics/k400/train.csv',
11
+ '/scratch-shared/dcanez/data/ssv2/train.csv'],
12
+ 'datasets_weights': [0.65, 0.35],
13
+ 'fps': 4,
14
+ 'num_workers': 10,
15
+ 'patch_size': 14,
16
+ 'persistent_workers': True,
17
+ 'pin_mem': True,
18
+ 'stage': [ { 'dest': 'kinetics_240',
19
+ 'format': 'targz_parts',
20
+ 'src': '/scratch-shared/dcanez/data/kinetics/k400/tars_240/'},
21
+ { 'dest': 'ssv2',
22
+ 'format': 'multipart_tar',
23
+ 'src': '/scratch-nvme/ml-datasets/something-something-v2/'}],
24
+ 'tubelet_size': 1},
25
+ 'data_aug': { 'auto_augment': False,
26
+ 'motion_shift': False,
27
+ 'random_resize_aspect_ratio': [0.75, 1.35],
28
+ 'random_resize_scale': [0.3, 1.0],
29
+ 'reprob': 0.0},
30
+ 'folder': '/scratch-shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable',
31
+ 'loss': {'loss_exp': 1.0},
32
+ 'mask': [ { 'aspect_ratio': [0.75, 1.5],
33
+ 'full_complement': False,
34
+ 'max_keep': None,
35
+ 'max_temporal_keep': 1.0,
36
+ 'num_blocks': 8,
37
+ 'spatial_scale': [0.15, 0.15],
38
+ 'temporal_scale': [1.0, 1.0]},
39
+ { 'aspect_ratio': [0.75, 1.5],
40
+ 'full_complement': False,
41
+ 'max_keep': None,
42
+ 'max_temporal_keep': 1.0,
43
+ 'num_blocks': 2,
44
+ 'spatial_scale': [0.7, 0.7],
45
+ 'temporal_scale': [1.0, 1.0]}],
46
+ 'mem_per_gpu': '180G',
47
+ 'meta': { 'dtype': 'bfloat16',
48
+ 'knn_eval_epoch0': False,
49
+ 'knn_eval_freq': 5,
50
+ 'knn_eval_presets': [ { 'config': { 'batch_size': 64,
51
+ 'dataset_train': '/scratch-shared/mdorkenw/ucf101/ucfTrainTestlist/train.csv',
52
+ 'dataset_val': '/scratch-shared/mdorkenw/ucf101/ucfTrainTestlist/val.csv',
53
+ 'eval_videos_per_class': 25,
54
+ 'num_workers': 8,
55
+ 'pool_type': 'slot_temporal_concat',
56
+ 'train_videos_per_class': 100},
57
+ 'preset': 'ucf101'},
58
+ { 'config': { 'batch_size': 64,
59
+ 'dataset_train': '/scratch-shared/mdorkenw/20bn-something-something-v2/something-something-v2-annotations/train_coarse10.csv',
60
+ 'dataset_val': '/scratch-shared/mdorkenw/20bn-something-something-v2/something-something-v2-annotations/val_coarse10.csv',
61
+ 'eval_videos_per_class': 100,
62
+ 'linear_probe': True,
63
+ 'num_workers': 8,
64
+ 'pool_type': 'slot_temporal_concat',
65
+ 'train_videos_per_class': 500},
66
+ 'preset': 'ssv2_coarse10'}],
67
+ 'load_checkpoint': True,
68
+ 'read_checkpoint': None,
69
+ 'save_every_freq': 5,
70
+ 'seed': 239,
71
+ 'use_sdpa': True,
72
+ 'use_wandb': True,
73
+ 'wandb_project': 'vjepa_ujepaside'},
74
+ 'metrics': {'sigreg': {}, 'std': {}},
75
+ 'model': { 'model_name': 'ujepaside_large_patch14_capi_lvd1689m',
76
+ 'pred_depth': 6,
77
+ 'pred_embed_dim': 384,
78
+ 'pred_num_heads': 12,
79
+ 'predictor': 'v2_cross',
80
+ 'st_causal': False,
81
+ 'st_drop_path': 0.2,
82
+ 'st_flex_enable': False,
83
+ 'st_layer_scale_init': 1e-05,
84
+ 'st_num_slots': 16,
85
+ 'st_slots_causal_within_frame': False,
86
+ 'target_kind': 'frozen_2d',
87
+ 'target_type': 'vit_large_patch14_capi.lvd1689m',
88
+ 'temporal_spacing': 1.0,
89
+ 'uniform_power': True,
90
+ 'use_activation_checkpointing': True,
91
+ 'use_mask_tokens': True,
92
+ 'use_rope': True,
93
+ 'use_sdpa': True,
94
+ 'zero_init_mask_tokens': True},
95
+ 'nodes': 4,
96
+ 'optimization': { 'clip_grad': 3.0,
97
+ 'ema': [0.99925, 0.99925],
98
+ 'epochs': 100,
99
+ 'final_lr': 0.0001,
100
+ 'final_weight_decay': 0.04,
101
+ 'ipe': 300,
102
+ 'ipe_scale': 1.0,
103
+ 'lr': 0.0005,
104
+ 'start_lr': 0.0001,
105
+ 'warmup': 10,
106
+ 'weight_decay': 0.04},
107
+ 'tasks_per_node': 4}
108
+ INFO:root:Running pre-training of app: vjepa
109
+ [INFO ][2026-05-10 12:16:54][app.vjepa.train ][main ] which_dtype='bfloat16'
110
+ [INFO ][2026-05-10 12:16:54][app.vjepa.train ][main ] Disabling persistent_workers (incompatible with KNN eval)
111
+ [INFO ][2026-05-10 12:16:54][app.vjepa.train ][main ] NCCL_SOCKET_IFNAME=eno
112
+ [INFO ][2026-05-10 12:16:56][app.vjepa.train ][main ] Initialized (rank/world-size) 9/16, tasks_per_node=4
113
+ [INFO ][2026-05-10 12:16:56][root ][stage_datasets ] [local_rank 1/4] Staging kinetics_240 (targz_parts)
114
+ [INFO ][2026-05-10 12:16:56][root ][_stage_targz_parts ] [rank 1] Extracting 26/103 tar.gz parts to /scratch-node/dcanez.22614766/kinetics_240
115
+ [INFO ][2026-05-10 12:17:08][root ][_stage_targz_parts ] [local_rank 1] Extracted 2/26 parts
116
+ [INFO ][2026-05-10 12:17:22][root ][_stage_targz_parts ] [local_rank 1] Extracted 4/26 parts
117
+ [INFO ][2026-05-10 12:17:36][root ][_stage_targz_parts ] [local_rank 1] Extracted 6/26 parts
118
+ [INFO ][2026-05-10 12:17:50][root ][_stage_targz_parts ] [local_rank 1] Extracted 8/26 parts
119
+ [INFO ][2026-05-10 12:18:04][root ][_stage_targz_parts ] [local_rank 1] Extracted 10/26 parts
120
+ [INFO ][2026-05-10 12:18:19][root ][_stage_targz_parts ] [local_rank 1] Extracted 12/26 parts
121
+ [INFO ][2026-05-10 12:18:32][root ][_stage_targz_parts ] [local_rank 1] Extracted 14/26 parts
122
+ [INFO ][2026-05-10 12:18:46][root ][_stage_targz_parts ] [local_rank 1] Extracted 16/26 parts
123
+ [INFO ][2026-05-10 12:19:00][root ][_stage_targz_parts ] [local_rank 1] Extracted 18/26 parts
124
+ [INFO ][2026-05-10 12:19:14][root ][_stage_targz_parts ] [local_rank 1] Extracted 20/26 parts
125
+ [INFO ][2026-05-10 12:19:29][root ][_stage_targz_parts ] [local_rank 1] Extracted 22/26 parts
126
+ [INFO ][2026-05-10 12:19:43][root ][_stage_targz_parts ] [local_rank 1] Extracted 24/26 parts
127
+ [INFO ][2026-05-10 12:19:57][root ][_stage_targz_parts ] [local_rank 1] Extracted 26/26 parts
128
+ [INFO ][2026-05-10 12:19:57][root ][stage_datasets ] [local_rank 1/4] Staging ssv2 (multipart_tar)
129
+ [INFO ][2026-05-10 12:21:03][app.vjepa.train ][main ] Staged datasets: ['/scratch-node/dcanez.22614766/kinetics_240/train.csv', '/scratch-node/dcanez.22614766/ssv2/train.csv']
130
+ [INFO ][2026-05-10 12:23:31][experiments.stmodels.vision_transformers_v3][vit_large_patch14_capi_lvd1689m] Loading pretrained weights for vit_large_patch14_capi_lvd1689m from https://dl.fbaipublicfiles.com/capi/capi_vitl14_lvd.pth
131
+ [INFO ][2026-05-10 12:23:45][root ][_build_frozen_2d_target ] Loaded frozen_2d target from timm: vit_large_patch14_capi.lvd1689m
132
+ [INFO ][2026-05-10 12:23:46][root ][init_video_model ] ViTMultiSeqWrapper(
133
+ (backbone): UJEPAside(
134
+ (patch_embed): PatchEmbed3D(
135
+ (proj): Conv3d(3, 1024, kernel_size=(1, 14, 14), stride=(1, 14, 14))
136
+ )
137
+ (rope): CAPI2DRoPE()
138
+ (blocks): ModuleList(
139
+ (0-23): 24 x Block(
140
+ (norm1): RMSNorm((1024,), eps=1e-05, elementwise_affine=True)
141
+ (rope_impl): CAPI2DRoPE()
142
+ (norm2): RMSNorm((1024,), eps=1e-05, elementwise_affine=True)
143
+ (drop_path1): Identity()
144
+ (drop_path2): Identity()
145
+ (attn): Attention(
146
+ (qkv): Linear(in_features=1024, out_features=3072, bias=False)
147
+ (attn_drop): Dropout(p=0.0, inplace=False)
148
+ (proj): Linear(in_features=1024, out_features=1024, bias=False)
149
+ (proj_drop): Dropout(p=0.0, inplace=False)
150
+ (rope_impl): CAPI2DRoPE()
151
+ )
152
+ (mlp): MLP(
153
+ (fc1): Linear(in_features=1024, out_features=4096, bias=False)
154
+ (act): GELU(approximate='none')
155
+ (fc2): Linear(in_features=4096, out_features=1024, bias=False)
156
+ (drop): Dropout(p=0.0, inplace=False)
157
+ )
158
+ )
159
+ )
160
+ (norm): RMSNorm((1024,), eps=1e-05, elementwise_affine=False)
161
+ (st_blocks): ModuleList(
162
+ (0-23): 24 x Block(
163
+ (residual1): EfficientResidual(
164
+ (norm): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
165
+ (fn): Attention(
166
+ (q_proj): Linear(in_features=1024, out_features=1024, bias=False)
167
+ (k_proj): Linear(in_features=1024, out_features=1024, bias=False)
168
+ (v_proj): Linear(in_features=1024, out_features=1024, bias=False)
169
+ (proj): Linear(in_features=1024, out_features=1024, bias=False)
170
+ (rope_impl): CAPI3DRoPE()
171
+ )
172
+ )
173
+ (residual2): EfficientResidual(
174
+ (norm): LayerNorm((1024,), eps=1e-06, elementwise_affine=True)
175
+ (fn): MLP(
176
+ (fc1): Linear(in_features=1024, out_features=4096, bias=False)
177
+ (act): GELU(approximate='none')
178
+ (fc2): Linear(in_features=4096, out_features=1024, bias=False)
179
+ (drop): Dropout(p=0.0, inplace=False)
180
+ )
181
+ )
182
+ )
183
+ )
184
+ (st_rope): CAPI3DRoPE()
185
+ )
186
+ )
187
+ [INFO ][2026-05-10 12:23:46][root ][init_video_model ] UJEPAsidePredictorMultiSeqWrapper(
188
+ (backbone): PredictorV2(
189
+ (predictor_embed): Linear(in_features=1024, out_features=384, bias=True)
190
+ (mask_tokens): ParameterList(
191
+ (0): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
192
+ (1): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
193
+ (2): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
194
+ (3): Parameter containing: [torch.float32 of size 1x384 (cuda:0)]
195
+ )
196
+ (predictor_blocks): ModuleList(
197
+ (0-5): 6 x Block(
198
+ (residual1): EfficientResidual(
199
+ (norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
200
+ (fn): Attention(
201
+ (q_proj): Linear(in_features=384, out_features=384, bias=False)
202
+ (k_proj): Linear(in_features=384, out_features=384, bias=False)
203
+ (v_proj): Linear(in_features=384, out_features=384, bias=False)
204
+ (proj): Linear(in_features=384, out_features=384, bias=False)
205
+ (rope): Rope()
206
+ )
207
+ )
208
+ (residual2): EfficientResidual(
209
+ (norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
210
+ (fn): MLP(
211
+ (fc1): Linear(in_features=384, out_features=1536, bias=False)
212
+ (act): GELU(approximate='none')
213
+ (fc2): Linear(in_features=1536, out_features=384, bias=False)
214
+ (drop): Dropout(p=0.0, inplace=False)
215
+ )
216
+ )
217
+ )
218
+ )
219
+ (predictor_norm): LayerNorm((384,), eps=1e-05, elementwise_affine=True)
220
+ (predictor_proj): Linear(in_features=384, out_features=1024, bias=True)
221
+ )
222
+ )
223
+ [INFO ][2026-05-10 12:23:46][root ][init_video_model ] MultiSeqWrapper(
224
+ (backbone): Frozen2DTargetWrapper(
225
+ (backbone): Eva(
226
+ (patch_embed): PatchEmbed(
227
+ (proj): Conv2d(3, 1024, kernel_size=(14, 14), stride=(14, 14))
228
+ (norm): Identity()
229
+ )
230
+ (pos_drop): Dropout(p=0.0, inplace=False)
231
+ (norm_pre): Identity()
232
+ (blocks): ModuleList(
233
+ (0-23): 24 x EvaBlock(
234
+ (norm1): RMSNorm((1024,), eps=1e-05, elementwise_affine=True)
235
+ (attn): EvaAttention(
236
+ (qkv): Linear(in_features=1024, out_features=3072, bias=False)
237
+ (q_norm): Identity()
238
+ (k_norm): Identity()
239
+ (attn_drop): Dropout(p=0.0, inplace=False)
240
+ (norm): Identity()
241
+ (proj): Linear(in_features=1024, out_features=1024, bias=True)
242
+ (proj_drop): Dropout(p=0.0, inplace=False)
243
+ )
244
+ (drop_path1): Identity()
245
+ (norm2): RMSNorm((1024,), eps=1e-05, elementwise_affine=True)
246
+ (mlp): Mlp(
247
+ (fc1): Linear(in_features=1024, out_features=4096, bias=True)
248
+ (act): GELU(approximate='none')
249
+ (drop1): Dropout(p=0.0, inplace=False)
250
+ (norm): Identity()
251
+ (fc2): Linear(in_features=4096, out_features=1024, bias=True)
252
+ (drop2): Dropout(p=0.0, inplace=False)
253
+ )
254
+ (drop_path2): Identity()
255
+ )
256
+ )
257
+ (norm): RMSNorm((1024,), eps=1e-05, elementwise_affine=True)
258
+ (fc_norm): Identity()
259
+ (head_drop): Dropout(p=0.0, inplace=False)
260
+ (head): Identity()
261
+ (rope): _CapiPatchRoPE()
262
+ )
263
+ )
264
+ )
265
+ [INFO ][2026-05-10 12:23:46][root ][init_video_model ] Encoder number of parameters: 302399488
266
+ [INFO ][2026-05-10 12:23:46][root ][init_video_model ] Predictor number of parameters: 11416192
267
+ [INFO ][2026-05-10 12:23:46][root ][init_video_model ] Target encoder number of parameters: 0
268
+ [INFO ][2026-05-10 12:23:49][root ][make_videodataset ] VideoDataset dataset created
269
+ [INFO ][2026-05-10 12:23:49][WeightedSampler ][__init__ ] Using DistributedWeightedSampler with rank 9 / 16
270
+ [INFO ][2026-05-10 12:23:49][root ][make_videodataset ] VideoDataset unsupervised data loader created
271
+ [INFO ][2026-05-10 12:23:49][app.vjepa.train ][main ] iterations per epoch/dataset length: 300/399
272
+ [INFO ][2026-05-10 12:23:49][app.vjepa.train ][main ] Wrapping models in DDP (rank 9)...
273
+ [INFO ][2026-05-10 12:24:08][app.vjepa.train ][main ] Initializing loader...
274
+ submitit WARNING (2026-05-10 16:20:00,106) - Bypassing signal SIGTERM
275
+ [WARNING ][2026-05-10 16:20:00][submitit ][bypass ] Bypassing signal SIGTERM
276
+ submitit WARNING (2026-05-10 16:20:00,176) - Bypassing signal SIGCONT
277
+ [WARNING ][2026-05-10 16:20:00][submitit ][bypass ] Bypassing signal SIGCONT
278
+ submitit ERROR (2026-05-10 16:20:00,176) - Submitted job triggered an exception
279
+ [ERROR ][2026-05-10 16:20:00][submitit ][process_job ] Submitted job triggered an exception
280
+ [ERROR ][2026-05-10 16:20:00][app.vjepa.train ][_save_crash ] Saved crash log to /scratch-shared/dcanez/runs/vjepa/ujepaside_capi_lvd/c002_vitl_k16_simple_cross_stable/crash_rank9.log