twanghcmut/backup-foundation-physics / logs /batch_attempt1_oom.log
twanghcmut's picture
download
raw
34.2 kB
18:57:05 INFO fpgm.datagen.batch.gpu_pool: gpu pool: discovered 4 device(s): [0, 1, 2, 3] (measured free MB: {0: 20538, 1: 24659, 2: 3451, 3: 11677})
18:57:05 INFO fpgm.datagen.batch.summary: batch summary: resumed 1 episode record(s) from /srv/data/foundation-physics-graph-model/outputs/datagen/batch_summary.json
18:57:05 INFO fpgm.datagen.batch.runner: batch: skipping 1 episode(s) already 'ok' in /srv/data/foundation-physics-graph-model/outputs/datagen/batch_summary.json
18:57:05 INFO fpgm.datagen.batch.runner: batch: pilot -- running 7/61 episode(s) first
18:57:05 INFO fpgm.datagen.batch.runner: batch: 7 episode(s) -> 4 shard(s) (concurrency=4, 4 GPU device(s) visible: (0, 1, 2, 3))
18:57:06 INFO fpgm.datagen.batch.gpu_pool: gpu claim: device 0 claimed (measured free 20538 MB >= min 16000 MB); 1/1 workers now claimed on this device
18:57:06 INFO fpgm.datagen.batch.gpu_pool: gpu claim: device 1 claimed (measured free 24659 MB >= min 16000 MB); 1/1 workers now claimed on this device
18:57:06 INFO fpgm.datagen.batch.gpu_pool: gpu claim: no device available yet (attempt 1; measured free MB: {0: 20538, 1: 24659, 2: 3451, 3: 11677}; claims: {0: 1, 1: 1, 2: 0, 3: 0}); sleeping 20s
18:57:08 INFO fpgm.datagen.batch.worker.0: shard 0: starting on cuda:0 (visible as cuda:0), 2 episode(s), 8 thread(s)
18:57:08 INFO fpgm.datagen.batch.worker.1: shard 1: starting on cuda:1 (visible as cuda:0), 2 episode(s), 8 thread(s)
18:57:09 INFO fpgm.datagen.pipeline: extrinsics candidate 'optimized_cameras_json': Bundle-adjusted world->camera extrinsic from cameras.json's optimized_extrinsics (Phase 0 measured this byte-identical to SceneFlowClip.extrinsic). Expected to win the gate since it is jointly optimized against the whole scene rather than a single onboard frame estimate -- but that expectation is exactly what the gate verifies, not assumes.
18:57:09 INFO fpgm.datagen.pipeline: extrinsics candidate 'trajectory_h5_row0': DROID's per-frame 6-vector cam2world at trajectory row 0, converted via cam2world_vector_to_world2cam. Row 0 stands in for the whole episode under the fixed-camera assumption; max translation spread across all 128 rows was 0.00 mm (a large spread here would mean this single-row candidate is not representative even on its own terms).
18:57:10 INFO fpgm.datagen.pipeline: extrinsics candidate 'optimized_cameras_json': Bundle-adjusted world->camera extrinsic from cameras.json's optimized_extrinsics (Phase 0 measured this byte-identical to SceneFlowClip.extrinsic). Expected to win the gate since it is jointly optimized against the whole scene rather than a single onboard frame estimate -- but that expectation is exactly what the gate verifies, not assumes.
18:57:10 INFO fpgm.datagen.pipeline: extrinsics candidate 'trajectory_h5_row0': DROID's per-frame 6-vector cam2world at trajectory row 0, converted via cam2world_vector_to_world2cam. Row 0 stands in for the whole episode under the fixed-camera assumption; max translation spread across all 147 rows was 0.00 mm (a large spread here would mean this single-row candidate is not representative even on its own terms).
18:57:10 INFO fpgm.datagen.pipeline: picked GPU 0 (21141 MB free >= 16000 MB required)
18:57:10 INFO fpgm.datagen.robot_buffers: robot_buffers: cache hit for AUTOLab+0d4edc83+2023-10-21-19h-07m-04s/22008760, skipping render
18:57:10 INFO fpgm.datagen.pipeline: robot alignment gate (prompt='robotic arm', sam_tracked=127, dropped=0, recall>=0.85, systematic_shift<=2.0px):
-> optimized_cameras_json n=127 recall(median/p10)=0.950/0.876 shift(gated, systematic px)=1.24 shift(scatter, diagnostic px, std)=9.7 shift(scatter, diagnostic px, median|.|)=4.2 iou(mean, diagnostic)=0.768
trajectory_h5_row0 n=127 recall(median/p10)=0.940/0.858 shift(gated, systematic px)=6.20 shift(scatter, diagnostic px, std)=9.8 shift(scatter, diagnostic px, median|.|)=7.0 iou(mean, diagnostic)=0.755
PASSED
18:57:10 INFO fpgm.datagen.pipeline: picked GPU 1 (25262 MB free >= 16000 MB required)
18:57:11 INFO fpgm.pipeline.frames: extracted 146 frames [0, 145] stride=1 from 22008760.mp4 -> /tmp/fpgm_frames_pvqhg5zr
/home/quang/miniconda3/envs/fpgm/lib/python3.12/site-packages/timm/models/layers/__init__.py:49: FutureWarning: Importing from timm.models.layers is deprecated, please import via timm.layers
warnings.warn(f"Importing from {__name__} is deprecated, please import via timm.layers", FutureWarning)
18:57:13 INFO fpgm.segmentation.sam3: building SAM 3.1 predictor (use_fa3=False)
18:57:26 INFO fpgm.datagen.batch.gpu_pool: gpu claim: no device available yet (attempt 2; measured free MB: {0: 20538, 1: 21692, 2: 3451, 3: 11677}; claims: {0: 1, 1: 1, 2: 0, 3: 0}); sleeping 20s
INFO 2026-08-01 18:57:29,088 4055173 sam3_multiplex_base.py: 336: `setting max_num_objects` to 16 -- creating num_obj_for_compile=1 objects for torch.compile cache
18:57:32 INFO fpgm.segmentation.sam3: applied sam3 compatibility shim: init_state does not accept ['offload_state_to_cpu']
18:57:32 INFO fpgm.segmentation.sam3: applied sam3 compatibility shim: _build_sam2_output was discarding refined masks on frames with no prior cache entry, which breaks point-prompt video propagation
dynamic_multimask_via_stability is reset to False in the multiplex model
frame loading (image folder) [rank=0]: 0%| | 0/146 [00:00<?, ?it/s]INFO 2026-08-01 18:57:32,357 4055173 sam3_base_predictor.py: 146: started new session 65ce0823-e5d9-423b-a158-522dc6492fa1
INFO 2026-08-01 18:57:32,358 4055173 sam3_multiplex_tracking.py:1677: Running add_prompt on frame 73
frame loading (image folder) [rank=0]: 2%|▏ | 3/146 [00:00<00:04, 29.43it/s] frame loading (image folder) [rank=0]: 5%|▌ | 8/146 [00:00<00:03, 35.99it/s] frame loading (image folder) [rank=0]: 8%|▊ | 12/146 [00:00<00:03, 34.72it/s] frame loading (image folder) [rank=0]: 11%|█ | 16/146 [00:00<00:03, 35.96it/s] frame loading (image folder) [rank=0]: 14%|█▎ | 20/146 [00:00<00:04, 28.96it/s]INFO 2026-08-01 18:57:33,089 4055173 sam3_multiplex_tracking.py:2316: Running full VG propagation (reverse=False).
propagate_in_video: 0%| | 0/73 [00:00<?, ?it/s] frame loading (image folder) [rank=0]: 16%|█▋ | 24/146 [00:00<00:04, 25.51it/s] frame loading (image folder) [rank=0]: 18%|█▊ | 27/146 [00:00<00:04, 24.23it/s] frame loading (image folder) [rank=0]: 21%|██ | 30/146 [00:01<00:04, 23.35it/s] frame loading (image folder) [rank=0]: 23%|██▎ | 34/146 [00:01<00:04, 25.20it/s] frame loading (image folder) [rank=0]: 25%|██▌ | 37/146 [00:01<00:04, 25.44it/s] frame loading (image folder) [rank=0]: 28%|██▊ | 41/146 [00:01<00:04, 25.76it/s] frame loading (image folder) [rank=0]: 31%|███ | 45/146 [00:01<00:04, 24.85it/s] frame loading (image folder) [rank=0]: 33%|███▎ | 48/146 [00:01<00:03, 24.71it/s] frame loading (image folder) [rank=0]: 35%|███▍ | 51/146 [00:02<00:04, 20.70it/s]
propagate_in_video: 1%|▏ | 1/73 [00:01<01:40, 1.39s/it] frame loading (image folder) [rank=0]: 38%|███▊ | 55/146 [00:02<00:04, 22.55it/s] frame loading (image folder) [rank=0]: 40%|███▉ | 58/146 [00:02<00:03, 22.76it/s] frame loading (image folder) [rank=0]: 42%|████▏ | 61/146 [00:02<00:04, 21.23it/s] frame loading (image folder) [rank=0]: 44%|████▍ | 64/146 [00:02<00:03, 22.90it/s] frame loading (image folder) [rank=0]: 64%|██████▎ | 93/146 [00:02<00:00, 83.10it/s] frame loading (image folder) [rank=0]: 71%|███████ | 103/146 [00:02<00:00, 66.88it/s] frame loading (image folder) [rank=0]: 76%|███████▌ | 111/146 [00:03<00:00, 55.85it/s]
propagate_in_video: 11%|█ | 8/73 [00:02<00:17, 3.64it/s] frame loading (image folder) [rank=0]: 81%|████████ | 118/146 [00:03<00:00, 47.40it/s] frame loading (image folder) [rank=0]: 85%|████████▍ | 124/146 [00:03<00:00, 40.59it/s] frame loading (image folder) [rank=0]: 88%|████████▊ | 129/146 [00:03<00:00, 36.16it/s] propagate_in_video: 32%|███▏ | 23/73 [00:03<00:06, 7.26it/s]
INFO 2026-08-01 18:57:36,260 4055173 sam3_base_predictor.py: 305: propagation ended in session 65ce0823-e5d9-423b-a158-522dc6492fa1
INFO 2026-08-01 18:57:36,546 4055173 sam3_base_predictor.py: 398: empty_cache freed 2717908992 bytes (free_pct 0.8% -> 2.6%, reserved 23938990080 -> 21221081088 bytes)
INFO 2026-08-01 18:57:36,546 4055173 sam3_base_predictor.py: 410: removed session 65ce0823-e5d9-423b-a158-522dc6492fa1
frame loading (image folder) [rank=0]: 92%|█████████▏| 134/146 [00:04<00:00, 23.26it/s] frame loading (image folder) [rank=0]: 92%|█████████▏| 134/146 [00:04<00:00, 31.67it/s]
18:57:36 WARNING fpgm.datagen.pipeline: stage s2 failed for AUTOLab+0d4edc83+2023-10-21-19h-07m-18s/ext1: CUDA out of memory. Tried to allocate 1.27 GiB. GPU 0 has a total capacity of 139.81 GiB of which 1.12 GiB is free. Including non-PyTorch memory, this process has 22.97 GiB memory in use. Of the allocated memory 20.16 GiB is allocated by PyTorch, and 2.13 GiB is reserved by PyTorch but unallocated. If reserved but unallocated memory is large try setting PYTORCH_ALLOC_CONF=expandable_segments:True to avoid fragmentation. See documentation for Memory Management (https://pytorch.org/docs/stable/notes/cuda.html#environment-variables)
18:57:36 INFO fpgm.datagen.batch.worker.1: shard 1: AUTOLab+0d4edc83+2023-10-21-19h-07m-18s -> failed (27.8s)
18:57:37 INFO fpgm.datagen.pipeline: extrinsics candidate 'optimized_cameras_json': Bundle-adjusted world->camera extrinsic from cameras.json's optimized_extrinsics (Phase 0 measured this byte-identical to SceneFlowClip.extrinsic). Expected to win the gate since it is jointly optimized against the whole scene rather than a single onboard frame estimate -- but that expectation is exactly what the gate verifies, not assumes.
18:57:37 INFO fpgm.datagen.pipeline: extrinsics candidate 'trajectory_h5_row0': DROID's per-frame 6-vector cam2world at trajectory row 0, converted via cam2world_vector_to_world2cam. Row 0 stands in for the whole episode under the fixed-camera assumption; max translation spread across all 199 rows was 0.00 mm (a large spread here would mean this single-row candidate is not representative even on its own terms).
18:57:37 WARNING fpgm.datagen.pipeline: no GPU with >= 16000 MB free (attempt 1); retrying in 20s
18:57:47 INFO fpgm.datagen.batch.gpu_pool: gpu claim: no device available yet (attempt 3; measured free MB: {0: 20538, 1: 3742, 2: 3451, 3: 11677}; claims: {0: 1, 1: 1, 2: 0, 3: 0}); sleeping 20s
18:57:58 WARNING fpgm.datagen.pipeline: no GPU with >= 16000 MB free (attempt 2); retrying in 20s
18:57:58 INFO fpgm.datagen.pipeline: wrote /srv/data/foundation-physics-graph-model/outputs/datagen/AUTOLab+0d4edc83+2023-10-21-19h-07m-04s/22008760/master/robot_overlay.mp4
18:57:59 INFO fpgm.datagen.background_depth: background depth: reusing cached /srv/data/foundation-physics-graph-model/outputs/datagen/AUTOLab+0d4edc83+2023-10-21-19h-07m-04s/22008760/master/background_depth.h5
18:58:06 INFO fpgm.datagen.pipeline: picked GPU 0 (21141 MB free >= 16000 MB required)
18:58:06 INFO fpgm.datagen.model_registry: model registry: building sam3 on cuda
18:58:06 INFO fpgm.datagen.model_registry: model registry: building tapnext on cuda
18:58:07 INFO fpgm.datagen.object_masks: prompt_masks: reusing cached results for 2 roles
18:58:07 INFO fpgm.datagen.batch.gpu_pool: gpu claim: no device available yet (attempt 4; measured free MB: {0: 20538, 1: 3742, 2: 3451, 3: 11677}; claims: {0: 1, 1: 1, 2: 0, 3: 0}); sleeping 20s
18:58:07 INFO fpgm.datagen.pipeline: picked GPU 0 (21141 MB free >= 16000 MB required)
18:58:07 INFO fpgm.datagen.model_registry: model registry: building sam3 on cuda
18:58:07 INFO fpgm.datagen.model_registry: model registry: building tapnext on cuda
18:58:08 INFO fpgm.datagen.object_masks: prompt_masks: reusing cached results for 2 roles
18:58:09 INFO fpgm.datagen.static_span: static span brick: frames 0-30 (30 frames) via tracks, max step 0.77 px, drift 1.50 px
18:58:09 INFO fpgm.datagen.static_span: static span drawer: frames 0-30 (30 frames) via tracks, max step 0.65 px, drift 0.81 px
18:58:09 INFO fpgm.robot.kinematics: flange table built: g=0.0 offset=[0.0, 0.00024, 0.11498] sep=0.0941; g=1.0 offset=[0.0, 0.0003, 0.12627] sep=0.0162
18:58:09 INFO fpgm.geometry.rigid_lifting: rigid seed: frame 0 with 300 anchor points (mean confidence 1.00)
18:58:10 INFO fpgm.geometry.rigid_lifting: rigid PnP solved 127/127 frames (median inlier fraction 100%, seed planarity 0.4001)
18:58:10 INFO fpgm.datagen.geometry_ops: build_observed_surface_cloud('drawer'): prismatic precondition passed (residual_m=0.00045 rotation_range_deg=1.903 n_frames=101) -- geometry is OBSERVED POINTS, no orientation fit anywhere in this path
18:58:10 INFO fpgm.geometry.rigid_lifting: rigid seed: frame 0 with 300 anchor points (mean confidence 1.00)
18:58:10 INFO fpgm.datagen.geometry_ops: build_observed_surface_cloud('drawer'): robot carve -- seed frame 0 mask/robot overlap BEFORE carving = 0/49393 px (0.00% of the seed mask); carved from every frame in static_span=(0, 30) before unioning (robot_seg supplied)
18:58:10 INFO fpgm.datagen.geometry_ops: build_observed_surface_cloud('drawer'): mask union px, robot-carved=3433 vs. pre-carve=3433 (+0 px removed by the carve)
18:58:10 INFO fpgm.datagen.geometry_ops: build_observed_surface_cloud('drawer'): static_span=(0, 30) -- mask union 3433 px, 3433 valid (coverage 100.0%, of which 3433 genuinely measured/ANCHOR = 100.0% of the mask)
18:58:10 INFO fpgm.datagen.pipeline: S6: drawer geometry built as an OBSERVED point cloud (no fitted mesh, no fitted orientation): 3433 points, coverage=100.0% of the mask (100.0% genuinely measured)
18:58:10 INFO fpgm.datagen.geometry_ops: fit_support_plane('drawer'): frame 0 -- 3077 cloud points, top 15% band -> 2353 points
18:58:10 INFO fpgm.datagen.geometry_ops: fit_support_plane('drawer'): fit ok -- normal=[0.3019, 0.0422, 0.9524] offset=0.2202 rms=0.0054m n_inliers=1139/2353 footprint_ratio=0.88
18:58:10 INFO fpgm.datagen.pipeline: S6: brick -- pose_noise_mm=0.518 silhouette_iou(mean over valid)=0.748 rigid_lift={'planarity': 0.17042300800645957, 'n_rejected_implausible': 0, 'n_solved': 89, 'n_frames': 127} pose_source_counts={0: 87, 3: 40}
18:58:10 INFO fpgm.datagen.pipeline: S6: drawer -- pose_noise_mm=0.137 silhouette_iou(mean over valid)=0.353 rigid_lift={'planarity': 0.40009308398436977, 'n_rejected_implausible': 0, 'n_solved': 127, 'n_frames': 127} pose_source_counts={0: 127}
18:58:11 INFO fpgm.datagen.pipeline: S6: brick -- parent bridging via drawer: ['frames 87-126: parent=drawer attach rejected (pre-gap relative-pose score=0.064 < threshold=0.5)'] pose_source_counts={0: 87, 3: 40} silhouette_iou(mean over valid)=0.748
18:58:11 INFO fpgm.datagen.events: EventLabeler.label: 40/127 frames missing pose data, gap-filling for windowed computations (see EventTimeline.valid for which)
18:58:12 INFO fpgm.datagen.object_poses: wrote /srv/data/foundation-physics-graph-model/outputs/datagen/AUTOLab+0d4edc83+2023-10-21-19h-07m-04s/22008760/master/object_poses_debug.mp4
18:58:12 INFO fpgm.datagen.geometry_ops: wrote /srv/data/foundation-physics-graph-model/outputs/datagen/AUTOLab+0d4edc83+2023-10-21-19h-07m-04s/22008760/master/poses.npz
18:58:12 INFO fpgm.datagen.geometry_ops: wrote /srv/data/foundation-physics-graph-model/outputs/datagen/AUTOLab+0d4edc83+2023-10-21-19h-07m-04s/22008760/master/events.json
18:58:13 INFO fpgm.datagen.export_vace: S8: AUTOLab+0d4edc83+2023-10-21-19h-07m-04s/22008760 -- 127 video frames -> 3 window(s) (window=81 stride=40): [(0, 81), (40, 121), (46, 127)] -- mesh objects=('brick',), point-cloud objects=('drawer',)
18:58:13 INFO fpgm.datagen.export_vace: S8: carved 'drawer''s reference-frame (0) footprint out of the background plate -- 25609/399360 export pixels (6.4%)
18:58:13 INFO fpgm.datagen.export_vace: S8: window [0, 81) up to date, skipping
18:58:13 INFO fpgm.datagen.export_vace: S8: window [40, 121) up to date, skipping
18:58:13 INFO fpgm.datagen.export_vace: S8: window [46, 127) up to date, skipping
18:58:13 INFO fpgm.datagen.batch.worker.0: shard 0: AUTOLab+0d4edc83+2023-10-21-19h-07m-04s -> ok (64.9s)
18:58:13 INFO fpgm.datagen.pipeline: extrinsics candidate 'optimized_cameras_json': Bundle-adjusted world->camera extrinsic from cameras.json's optimized_extrinsics (Phase 0 measured this byte-identical to SceneFlowClip.extrinsic). Expected to win the gate since it is jointly optimized against the whole scene rather than a single onboard frame estimate -- but that expectation is exactly what the gate verifies, not assumes.
18:58:13 INFO fpgm.datagen.pipeline: extrinsics candidate 'trajectory_h5_row0': DROID's per-frame 6-vector cam2world at trajectory row 0, converted via cam2world_vector_to_world2cam. Row 0 stands in for the whole episode under the fixed-camera assumption; max translation spread across all 166 rows was 0.00 mm (a large spread here would mean this single-row candidate is not representative even on its own terms).
18:58:14 INFO fpgm.datagen.pipeline: picked GPU 0 (21141 MB free >= 16000 MB required)
18:58:15 INFO fpgm.pipeline.frames: extracted 165 frames [0, 164] stride=1 from 22008760.mp4 -> /tmp/fpgm_frames_t45x5ypd
/home/quang/miniconda3/envs/fpgm/lib/python3.12/site-packages/timm/models/layers/__init__.py:49: FutureWarning: Importing from timm.models.layers is deprecated, please import via timm.layers
warnings.warn(f"Importing from {__name__} is deprecated, please import via timm.layers", FutureWarning)
18:58:16 INFO fpgm.segmentation.sam3: building SAM 3.1 predictor (use_fa3=False)
18:58:18 WARNING fpgm.datagen.pipeline: no GPU with >= 16000 MB free (attempt 3); retrying in 20s
INFO 2026-08-01 18:58:33,405 4055167 sam3_multiplex_base.py: 336: `setting max_num_objects` to 16 -- creating num_obj_for_compile=1 objects for torch.compile cache
18:58:36 INFO fpgm.segmentation.sam3: applied sam3 compatibility shim: init_state does not accept ['offload_state_to_cpu']
18:58:36 INFO fpgm.segmentation.sam3: applied sam3 compatibility shim: _build_sam2_output was discarding refined masks on frames with no prior cache entry, which breaks point-prompt video propagation
dynamic_multimask_via_stability is reset to False in the multiplex model
frame loading (image folder) [rank=0]: 0%| | 0/165 [00:00<?, ?it/s]INFO 2026-08-01 18:58:36,840 4055167 sam3_base_predictor.py: 146: started new session 49367040-02fe-4735-a729-84328a4a7a12
INFO 2026-08-01 18:58:36,841 4055167 sam3_multiplex_tracking.py:1677: Running add_prompt on frame 82
frame loading (image folder) [rank=0]: 1%| | 2/165 [00:00<00:09, 16.37it/s] frame loading (image folder) [rank=0]: 2%|▏ | 4/165 [00:00<00:09, 17.34it/s] frame loading (image folder) [rank=0]: 5%|▌ | 9/165 [00:00<00:05, 30.66it/s] frame loading (image folder) [rank=0]: 8%|▊ | 13/165 [00:00<00:05, 29.41it/s] frame loading (image folder) [rank=0]: 10%|█ | 17/165 [00:00<00:04, 30.14it/s]INFO 2026-08-01 18:58:37,520 4055167 sam3_multiplex_tracking.py:2316: Running full VG propagation (reverse=False).
propagate_in_video: 0%| | 0/83 [00:00<?, ?it/s] frame loading (image folder) [rank=0]: 13%|█▎ | 21/165 [00:00<00:05, 25.65it/s] frame loading (image folder) [rank=0]: 15%|█▍ | 24/165 [00:00<00:06, 22.36it/s] frame loading (image folder) [rank=0]: 17%|█▋ | 28/165 [00:01<00:05, 24.50it/s] frame loading (image folder) [rank=0]: 20%|██ | 33/165 [00:01<00:04, 28.84it/s] frame loading (image folder) [rank=0]: 22%|██▏ | 37/165 [00:01<00:04, 30.15it/s] frame loading (image folder) [rank=0]: 25%|██▍ | 41/165 [00:01<00:03, 31.17it/s] frame loading (image folder) [rank=0]: 27%|██▋ | 45/165 [00:01<00:03, 32.66it/s] frame loading (image folder) [rank=0]: 30%|██▉ | 49/165 [00:01<00:03, 32.85it/s] frame loading (image folder) [rank=0]: 33%|███▎ | 54/165 [00:01<00:03, 36.81it/s]18:58:38 INFO fpgm.datagen.pipeline: picked GPU 1 (21076 MB free >= 16000 MB required)
frame loading (image folder) [rank=0]: 35%|███▌ | 58/165 [00:02<00:03, 27.34it/s] frame loading (image folder) [rank=0]: 38%|███▊ | 62/165 [00:02<00:03, 26.04it/s]
propagate_in_video: 1%| | 1/83 [00:01<02:08, 1.57s/it] frame loading (image folder) [rank=0]: 39%|███▉ | 65/165 [00:02<00:04, 23.74it/s] frame loading (image folder) [rank=0]: 41%|████ | 68/165 [00:02<00:05, 19.07it/s] frame loading (image folder) [rank=0]: 43%|████▎ | 71/165 [00:02<00:05, 17.77it/s] frame loading (image folder) [rank=0]: 45%|████▍ | 74/165 [00:02<00:04, 19.70it/s]18:58:39 INFO fpgm.pipeline.frames: extracted 198 frames [0, 197] stride=1 from 22008760.mp4 -> /tmp/fpgm_frames_1q3hoin2
18:58:39 INFO fpgm.segmentation.sam3: building SAM 3.1 predictor (use_fa3=False)
frame loading (image folder) [rank=0]: 47%|████▋ | 77/165 [00:03<00:04, 18.38it/s] frame loading (image folder) [rank=0]: 48%|████▊ | 80/165 [00:03<00:04, 17.72it/s] frame loading (image folder) [rank=0]: 69%|██████▉ | 114/165 [00:03<00:00, 76.88it/s] propagate_in_video: 17%|█▋ | 14/83 [00:03<00:15, 4.59it/s]
INFO 2026-08-01 18:58:40,569 4055167 sam3_base_predictor.py: 305: propagation ended in session 49367040-02fe-4735-a729-84328a4a7a12
INFO 2026-08-01 18:58:41,039 4055167 sam3_base_predictor.py: 398: empty_cache freed 0 bytes (free_pct 0.4% -> 0.4%, reserved 20201865216 -> 20201865216 bytes)
INFO 2026-08-01 18:58:41,039 4055167 sam3_base_predictor.py: 410: removed session 49367040-02fe-4735-a729-84328a4a7a12
frame loading (image folder) [rank=0]: 75%|███████▍ | 123/165 [00:04<00:01, 29.10it/s]
18:58:41 WARNING fpgm.datagen.pipeline: stage s2 failed for AUTOLab+0d4edc83+2023-10-21-19h-08m-41s/ext1: CUDA out of memory. Tried to allocate 1.27 GiB. GPU 0 has a total capacity of 139.81 GiB of which 577.38 MiB is free. Process 3647385 has 6.89 GiB memory in use. Process 3866682 has 1.50 GiB memory in use. Process 3866681 has 1.50 GiB memory in use. Process 3866679 has 1.50 GiB memory in use. Process 3866680 has 1.50 GiB memory in use. Including non-PyTorch memory, this process has 19.49 GiB memory in use. Of the allocated memory 17.31 GiB is allocated by PyTorch, and 1.50 GiB is reserved by PyTorch but unallocated. If reserved but unallocated memory is large try setting PYTORCH_ALLOC_CONF=expandable_segments:True to avoid fragmentation. See documentation for Memory Management (https://pytorch.org/docs/stable/notes/cuda.html#environment-variables)
18:58:41 INFO fpgm.datagen.batch.worker.0: shard 0: AUTOLab+0d4edc83+2023-10-21-19h-08m-41s -> failed (27.6s)
18:58:41 INFO fpgm.datagen.batch.worker.0: shard 0: done, 4 model construction(s) this process
INFO 2026-08-01 18:58:55,377 4055173 sam3_multiplex_base.py: 336: `setting max_num_objects` to 16 -- creating num_obj_for_compile=1 objects for torch.compile cache
18:58:58 INFO fpgm.segmentation.sam3: applied sam3 compatibility shim: init_state does not accept ['offload_state_to_cpu']
18:58:58 INFO fpgm.segmentation.sam3: applied sam3 compatibility shim: _build_sam2_output was discarding refined masks on frames with no prior cache entry, which breaks point-prompt video propagation
dynamic_multimask_via_stability is reset to False in the multiplex model
frame loading (image folder) [rank=0]: 0%| | 0/198 [00:00<?, ?it/s]INFO 2026-08-01 18:58:59,052 4055173 sam3_base_predictor.py: 146: started new session eb0fda83-9061-4e48-8fb3-517d5453df56
INFO 2026-08-01 18:58:59,053 4055173 sam3_multiplex_tracking.py:1677: Running add_prompt on frame 99
frame loading (image folder) [rank=0]: 3%|▎ | 5/198 [00:00<00:04, 43.31it/s] frame loading (image folder) [rank=0]: 5%|▌ | 10/198 [00:00<00:04, 39.33it/s]INFO 2026-08-01 18:58:59,386 4055173 sam3_multiplex_tracking.py:2316: Running full VG propagation (reverse=False).
propagate_in_video: 0%| | 0/99 [00:00<?, ?it/s] frame loading (image folder) [rank=0]: 7%|▋ | 14/198 [00:00<00:05, 33.06it/s] frame loading (image folder) [rank=0]: 9%|▉ | 18/198 [00:00<00:05, 31.90it/s] frame loading (image folder) [rank=0]: 11%|█ | 22/198 [00:00<00:05, 32.98it/s] frame loading (image folder) [rank=0]: 14%|█▎ | 27/198 [00:00<00:04, 34.75it/s] frame loading (image folder) [rank=0]: 16%|█▌ | 31/198 [00:00<00:05, 30.45it/s] frame loading (image folder) [rank=0]: 18%|█▊ | 35/198 [00:01<00:05, 31.50it/s] frame loading (image folder) [rank=0]: 20%|██ | 40/198 [00:01<00:04, 34.53it/s]
propagate_in_video: 1%| | 1/99 [00:00<01:28, 1.11it/s] frame loading (image folder) [rank=0]: 22%|██▏ | 44/198 [00:01<00:04, 34.73it/s]
propagate_in_video: 12%|█▏ | 12/99 [00:01<00:05, 15.89it/s] frame loading (image folder) [rank=0]: 24%|██▍ | 48/198 [00:01<00:04, 35.56it/s] frame loading (image folder) [rank=0]: 26%|██▋ | 52/198 [00:01<00:04, 31.41it/s] frame loading (image folder) [rank=0]: 28%|██▊ | 56/198 [00:01<00:05, 27.07it/s] frame loading (image folder) [rank=0]: 30%|██▉ | 59/198 [00:01<00:05, 25.42it/s] frame loading (image folder) [rank=0]: 32%|███▏ | 63/198 [00:02<00:04, 27.23it/s] frame loading (image folder) [rank=0]: 34%|███▍ | 67/198 [00:02<00:04, 30.08it/s] frame loading (image folder) [rank=0]: 36%|███▌ | 71/198 [00:02<00:03, 31.84it/s] frame loading (image folder) [rank=0]: 38%|███▊ | 75/198 [00:02<00:03, 33.94it/s] frame loading (image folder) [rank=0]: 41%|████ | 81/198 [00:02<00:02, 39.72it/s] frame loading (image folder) [rank=0]: 43%|████▎ | 86/198 [00:02<00:03, 33.23it/s]
propagate_in_video: 19%|█▉ | 19/99 [00:02<00:09, 8.41it/s] frame loading (image folder) [rank=0]: 45%|████▌ | 90/198 [00:02<00:03, 30.56it/s]
propagate_in_video: 23%|██▎ | 23/99 [00:02<00:07, 10.22it/s] frame loading (image folder) [rank=0]: 47%|████▋ | 94/198 [00:02<00:03, 30.26it/s] frame loading (image folder) [rank=0]: 67%|██████▋ | 132/198 [00:03<00:00, 106.13it/s] frame loading (image folder) [rank=0]: 73%|███████▎ | 144/198 [00:03<00:00, 65.81it/s] frame loading (image folder) [rank=0]: 78%|███████▊ | 154/198 [00:03<00:00, 55.59it/s]
propagate_in_video: 30%|███ | 30/99 [00:03<00:08, 7.76it/s] frame loading (image folder) [rank=0]: 82%|████████▏ | 162/198 [00:04<00:00, 42.72it/s] frame loading (image folder) [rank=0]: 85%|████████▍ | 168/198 [00:04<00:00, 39.22it/s] frame loading (image folder) [rank=0]: 88%|████████▊ | 174/198 [00:04<00:00, 37.21it/s] frame loading (image folder) [rank=0]: 90%|█████████ | 179/198 [00:04<00:00, 36.91it/s]
propagate_in_video: 46%|████▋ | 46/99 [00:04<00:04, 12.79it/s] frame loading (image folder) [rank=0]: 93%|█████████▎| 184/198 [00:04<00:00, 35.65it/s] frame loading (image folder) [rank=0]: 95%|█████████▍| 188/198 [00:04<00:00, 36.10it/s] frame loading (image folder) [rank=0]: 97%|█████████▋| 192/198 [00:04<00:00, 34.69it/s] frame loading (image folder) [rank=0]: 99%|█████████▉| 196/198 [00:05<00:00, 31.12it/s] frame loading (image folder) [rank=0]: 100%|██████████| 198/198 [00:05<00:00, 37.97it/s]
propagate_in_video: 63%|██████▎ | 62/99 [00:04<00:02, 16.56it/s]
propagate_in_video: 70%|██████▉ | 69/99 [00:05<00:01, 19.85it/s]
propagate_in_video: 76%|███████▌ | 75/99 [00:05<00:01, 23.03it/s]
propagate_in_video: 81%|████████ | 80/99 [00:05<00:01, 16.24it/s]
propagate_in_video: 88%|████████▊ | 87/99 [00:05<00:00, 20.59it/s]
propagate_in_video: 93%|█████████▎| 92/99 [00:06<00:00, 23.11it/s]
propagate_in_video: 98%|█████████▊| 97/99 [00:06<00:00, 18.60it/s] propagate_in_video: 100%|██████████| 99/99 [00:06<00:00, 15.17it/s]
INFO 2026-08-01 18:59:05,912 4055173 sam3_multiplex_tracking.py: 580: Bucket utilization rate: 6.25%, subscription rate: 100.00%
INFO 2026-08-01 18:59:05,912 4055173 sam3_multiplex_tracking.py:2316: Running full VG propagation (reverse=True).
propagate_in_video: 0%| | 0/99 [00:00<?, ?it/s] propagate_in_video: 1%| | 1/99 [00:00<00:58, 1.66it/s] propagate_in_video: 4%|▍ | 4/99 [00:01<00:29, 3.27it/s] propagate_in_video: 11%|█ | 11/99 [00:01<00:08, 10.59it/s] propagate_in_video: 17%|█▋ | 17/99 [00:01<00:04, 17.22it/s] propagate_in_video: 22%|██▏ | 22/99 [00:02<00:06, 12.10it/s] propagate_in_video: 27%|██▋ | 27/99 [00:02<00:04, 16.13it/s] propagate_in_video: 31%|███▏ | 31/99 [00:02<00:03, 19.27it/s] propagate_in_video: 36%|███▋ | 36/99 [00:03<00:05, 12.52it/s] propagate_in_video: 41%|████▏ | 41/99 [00:03<00:03, 16.34it/s] propagate_in_video: 46%|████▋ | 46/99 [00:03<00:02, 20.31it/s] propagate_in_video: 53%|█████▎ | 52/99 [00:03<00:03, 14.00it/s] propagate_in_video: 59%|█████▊ | 58/99 [00:04<00:02, 18.81it/s] propagate_in_video: 65%|██████▍ | 64/99 [00:04<00:01, 23.84it/s] propagate_in_video: 70%|██████▉ | 69/99 [00:04<00:02, 14.91it/s] propagate_in_video: 76%|███████▌ | 75/99 [00:04<00:01, 19.58it/s] propagate_in_video: 82%|████████▏ | 81/99 [00:05<00:00, 24.59it/s] propagate_in_video: 87%|████████▋ | 86/99 [00:05<00:00, 14.11it/s] propagate_in_video: 93%|█████████▎| 92/99 [00:05<00:00, 18.34it/s] propagate_in_video: 98%|█████████▊| 97/99 [00:06<00:00, 21.73it/s] propagate_in_video: 100%|██████████| 99/99 [00:06<00:00, 16.25it/s]
INFO 2026-08-01 18:59:12,008 4055173 sam3_multiplex_tracking.py: 580: Bucket utilization rate: 6.25%, subscription rate: 100.00%
INFO 2026-08-01 18:59:12,008 4055173 sam3_base_predictor.py: 305: propagation ended in session eb0fda83-9061-4e48-8fb3-517d5453df56
INFO 2026-08-01 18:59:12,535 4055173 sam3_base_predictor.py: 398: empty_cache freed 25312624640 bytes (free_pct 6.1% -> 22.9%, reserved 33592180736 -> 8279556096 bytes)
INFO 2026-08-01 18:59:12,535 4055173 sam3_base_predictor.py: 410: removed session eb0fda83-9061-4e48-8fb3-517d5453df56
18:59:12 INFO fpgm.datagen.robot_buffers: robot_buffers: SAM3 prompt 'robot arm and gripper' tracked 13/198 frames with a non-empty mask
18:59:12 INFO fpgm.datagen.robot_buffers: robot_buffers: cached 13 SAM masks to /srv/data/foundation-physics-graph-model/outputs/datagen/AUTOLab+0d4edc83+2023-10-21-19h-08m-58s/22008760/master/robot_buffers/sam_masks_full.npz
18:59:27 INFO fpgm.datagen.robot_buffers: robot_buffers: robot alignment gate (prompt='robot arm and gripper', sam_tracked=13, dropped=0, recall>=0.85, systematic_shift<=2.0px):
optimized_cameras_json n= 13 recall(median/p10)=0.944/0.921 shift(gated, systematic px)=16.42 shift(scatter, diagnostic px, std)=4.6 shift(scatter, diagnostic px, median|.|)=17.5 iou(mean, diagnostic)=0.719
-> trajectory_h5_row0 n= 13 recall(median/p10)=0.956/0.940 shift(gated, systematic px)=14.21 shift(scatter, diagnostic px, std)=5.2 shift(scatter, diagnostic px, median|.|)=15.0 iou(mean, diagnostic)=0.719
FAILED
18:59:27 ERROR fpgm.datagen.robot_buffers: ROBOT ALIGNMENT GATE FAILED: best candidate 'trajectory_h5_row0' scored median recall 0.956 (gate 0.85) and systematic shift 14.21px (gate 2.0) against SAM3 'robot arm and gripper' masks. Treat every buffer in this stage as UNVERIFIED. (uuid=AUTOLab+0d4edc83+2023-10-21-19h-08m-58s camera=22008760 recall=0.956 shift=14.21px)
19:00:39 INFO fpgm.datagen.pipeline: robot alignment gate (prompt='robot arm and gripper', sam_tracked=13, dropped=0, recall>=0.85, systematic_shift<=2.0px):
optimized_cameras_json n= 13 recall(median/p10)=0.944/0.921 shift(gated, systematic px)=16.42 shift(scatter, diagnostic px, std)=4.6 shift(scatter, diagnostic px, median|.|)=17.5 iou(mean, diagnostic)=0.719
-> trajectory_h5_row0 n= 13 recall(median/p10)=0.956/0.940 shift(gated, systematic px)=14.21 shift(scatter, diagnostic px, std)=5.2 shift(scatter, diagnostic px, median|.|)=15.0 iou(mean, diagnostic)=0.719
FAILED
19:00:39 WARNING fpgm.datagen.pipeline: S2 robot alignment gate FAILED for AUTOLab+0d4edc83+2023-10-21-19h-08m-58s/22008760 -- extrinsics/FK/URDF alignment is NOT verified; every buffer from this stage is UNVERIFIED.
/home/quang/miniconda3/envs/fpgm/lib/python3.12/multiprocessing/resource_tracker.py:279: UserWarning: resource_tracker: There appear to be 7 leaked semaphore objects to clean up at shutdown
warnings.warn('resource_tracker: There appear to be %d '

Xet Storage Details

Size:
34.2 kB
·
Xet hash:
3f619d080ae90ccb516c9178d6730f6b636a27ce649aa01315820914aa7e2e83

Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.