[run.sh] distributed_sharded auto-set NUM_PROCESSES=1 (all visible GPUs) [run.sh] Launch mode=distributed_sharded (DeepSpeed ZeRO-3) 2026-04-07 13:37:34.662042: I tensorflow/core/util/port.cc:153] oneDNN custom operations are on. You may see slightly different numerical results due to floating-point round-off errors from different computation orders. To turn them off, set the environment variable `TF_ENABLE_ONEDNN_OPTS=0`. 2026-04-07 13:37:37.839913: I tensorflow/core/platform/cpu_feature_guard.cc:210] This TensorFlow binary is optimized to use available CPU instructions in performance-critical operations. To enable the following instructions: AVX2 AVX512F AVX512_VNNI AVX512_BF16 AVX512_FP16 AVX_VNNI AMX_TILE AMX_INT8 AMX_BF16 FMA, in other operations, rebuild TensorFlow with the appropriate compiler flags. 2026-04-07 13:37:42.467616: I tensorflow/core/util/port.cc:153] oneDNN custom operations are on. You may see slightly different numerical results due to floating-point round-off errors from different computation orders. To turn them off, set the environment variable `TF_ENABLE_ONEDNN_OPTS=0`. 2026-04-07 13:37:42.473892: I external/local_xla/xla/tsl/cuda/cudart_stub.cc:31] Could not find cuda drivers on your machine, GPU will not be used. [train.py] Disabled flash/mem-efficient SDP kernels; using math SDP backend. [2026-04-07 13:38:10,838][root][INFO] - /g/data/rr81/aev/bin/x86_64-conda-linux-gnu-cc -march=nocona -mtune=haswell -ftree-vectorize -fPIC -fstack-protector-strong -fno-plt -O2 -ffunction-sections -pipe -isystem /g/data/rr81/aev/include -I/g/data/rr81/aev/targets/x86_64-linux/include -I/g/data/rr81/aev/targets/x86_64-linux/include/cccl -DNDEBUG -D_FORTIFY_SOURCE=2 -O2 -isystem /g/data/rr81/aev/include -I/g/data/rr81/aev/targets/x86_64-linux/include -I/g/data/rr81/aev/targets/x86_64-linux/include/cccl -fPIC -march=nocona -mtune=haswell -ftree-vectorize -fPIC -fstack-protector-strong -fno-plt -O2 -ffunction-sections -pipe -isystem /g/data/rr81/aev/include -I/g/data/rr81/aev/targets/x86_64-linux/include -I/g/data/rr81/aev/targets/x86_64-linux/include/cccl -c /scratch/rr81/ma5430/tmp/tmpspb_o7mh/test.c -o /scratch/rr81/ma5430/tmp/tmpspb_o7mh/test.o [2026-04-07 13:38:24,727][root][INFO] - /g/data/rr81/aev/bin/x86_64-conda-linux-gnu-cc -Wl,-O2 -Wl,--sort-common -Wl,--as-needed -Wl,-z,relro -Wl,-z,now -Wl,--disable-new-dtags -Wl,--gc-sections -Wl,-rpath,/g/data/rr81/aev/lib -Wl,-rpath-link,/g/data/rr81/aev/lib -L/g/data/rr81/aev/lib -L/g/data/rr81/aev/targets/x86_64-linux/lib -L/g/data/rr81/aev/targets/x86_64-linux/lib/stubs /scratch/rr81/ma5430/tmp/tmpspb_o7mh/test.o -laio -o /scratch/rr81/ma5430/tmp/tmpspb_o7mh/a.out [2026-04-07 13:38:25,328][root][INFO] - /g/data/rr81/aev/bin/x86_64-conda-linux-gnu-cc -march=nocona -mtune=haswell -ftree-vectorize -fPIC -fstack-protector-strong -fno-plt -O2 -ffunction-sections -pipe -isystem /g/data/rr81/aev/include -I/g/data/rr81/aev/targets/x86_64-linux/include -I/g/data/rr81/aev/targets/x86_64-linux/include/cccl -DNDEBUG -D_FORTIFY_SOURCE=2 -O2 -isystem /g/data/rr81/aev/include -I/g/data/rr81/aev/targets/x86_64-linux/include -I/g/data/rr81/aev/targets/x86_64-linux/include/cccl -fPIC -march=nocona -mtune=haswell -ftree-vectorize -fPIC -fstack-protector-strong -fno-plt -O2 -ffunction-sections -pipe -isystem /g/data/rr81/aev/include -I/g/data/rr81/aev/targets/x86_64-linux/include -I/g/data/rr81/aev/targets/x86_64-linux/include/cccl -c /scratch/rr81/ma5430/tmp/tmp6efs5r08/test.c -o /scratch/rr81/ma5430/tmp/tmp6efs5r08/test.o [2026-04-07 13:38:25,385][root][INFO] - /g/data/rr81/aev/bin/x86_64-conda-linux-gnu-cc -Wl,-O2 -Wl,--sort-common -Wl,--as-needed -Wl,-z,relro -Wl,-z,now -Wl,--disable-new-dtags -Wl,--gc-sections -Wl,-rpath,/g/data/rr81/aev/lib -Wl,-rpath-link,/g/data/rr81/aev/lib -L/g/data/rr81/aev/lib -L/g/data/rr81/aev/targets/x86_64-linux/lib -L/g/data/rr81/aev/targets/x86_64-linux/lib/stubs /scratch/rr81/ma5430/tmp/tmp6efs5r08/test.o -L/g/data/rr81/aev -L/g/data/rr81/aev/lib64 -lcufile -o /scratch/rr81/ma5430/tmp/tmp6efs5r08/a.out [2026-04-07 13:38:25,551][root][INFO] - /g/data/rr81/aev/bin/x86_64-conda-linux-gnu-cc -march=nocona -mtune=haswell -ftree-vectorize -fPIC -fstack-protector-strong -fno-plt -O2 -ffunction-sections -pipe -isystem /g/data/rr81/aev/include -I/g/data/rr81/aev/targets/x86_64-linux/include -I/g/data/rr81/aev/targets/x86_64-linux/include/cccl -DNDEBUG -D_FORTIFY_SOURCE=2 -O2 -isystem /g/data/rr81/aev/include -I/g/data/rr81/aev/targets/x86_64-linux/include -I/g/data/rr81/aev/targets/x86_64-linux/include/cccl -fPIC -march=nocona -mtune=haswell -ftree-vectorize -fPIC -fstack-protector-strong -fno-plt -O2 -ffunction-sections -pipe -isystem /g/data/rr81/aev/include -I/g/data/rr81/aev/targets/x86_64-linux/include -I/g/data/rr81/aev/targets/x86_64-linux/include/cccl -c /scratch/rr81/ma5430/tmp/tmp0ddzq53f/test.c -o /scratch/rr81/ma5430/tmp/tmp0ddzq53f/test.o [2026-04-07 13:38:25,609][root][INFO] - /g/data/rr81/aev/bin/x86_64-conda-linux-gnu-cc -Wl,-O2 -Wl,--sort-common -Wl,--as-needed -Wl,-z,relro -Wl,-z,now -Wl,--disable-new-dtags -Wl,--gc-sections -Wl,-rpath,/g/data/rr81/aev/lib -Wl,-rpath-link,/g/data/rr81/aev/lib -L/g/data/rr81/aev/lib -L/g/data/rr81/aev/targets/x86_64-linux/lib -L/g/data/rr81/aev/targets/x86_64-linux/lib/stubs /scratch/rr81/ma5430/tmp/tmp0ddzq53f/test.o -laio -o /scratch/rr81/ma5430/tmp/tmp0ddzq53f/a.out [2026-04-07 13:38:28,753][accelerate.utils.other][WARNING] - Detected kernel version 4.18.0, which is below the recommended minimum of 5.5.0; this can cause the process to hang. It is recommended to upgrade the kernel to the minimum version or higher. [2026-04-07 13:38:28,754][trainer.accelerators.base_accelerator][INFO] - Setting seed 42 [2026-04-07 13:38:28,771][trainer.accelerators.base_accelerator][INFO] - Initialized accelerator: rank=0 [2026-04-07 13:38:28,777][__main__][INFO] - Config can be found in logs/v5/reward_model/step_sana_sana_sprint_0_6b_1024_variable-t_lr1e-5_step-8000_filter2_time951/config.yaml [2026-04-07 13:38:28,778][__main__][INFO] - Loading task [2026-04-07 13:38:29,730][__main__][INFO] - Loading model `torch_dtype` is deprecated! Use `dtype` instead! Loading checkpoint shards: 0%| | 0/2 [00:00 sys.exit(main()) ^^^^^^ File "/g/data/rr81/aev/lib/python3.11/site-packages/accelerate/commands/accelerate_cli.py", line 50, in main args.func(args) File "/g/data/rr81/aev/lib/python3.11/site-packages/accelerate/commands/launch.py", line 1281, in launch_command simple_launcher(args) File "/g/data/rr81/aev/lib/python3.11/site-packages/accelerate/commands/launch.py", line 869, in simple_launcher raise subprocess.CalledProcessError(returncode=process.returncode, cmd=cmd) subprocess.CalledProcessError: Command '['/g/data/rr81/aev/bin/python3.11', 'trainer/scripts/train.py', '--config-path', '/g/data/rr81/LPO/lrm/lrm_sana/trainer/conf', '--config-name', 'step_sana_base', 'accelerator.mixed_precision=BF16', 'model.model_profile=sana_sprint_0_6b_1024', 'model.pretrained_model_name_or_path=Efficient-Large-Model/Sana_Sprint_0.6B_1024px_diffusers', 'model.image_size=1024', 'accelerator.run_name=step_sana_sana_sprint_0_6b_1024_variable-t_lr1e-5_step-8000_filter2_time951', 'accelerator.log_with=null', 'accelerator=deepspeed', 'optimizer=dummy', 'lr_scheduler=dummy', 'criterion.is_distributed=true', 'accelerator.deepspeed.zero_optimization.stage=3', 'accelerator.deepspeed.gradient_accumulation_steps=1', 'dataset.pseudo_preference_path=/g/data/rr81/LPO/lrm/lrm_sana/vqa_aes_clip_score_mp.csv', 'dataset.valid_split_name=validation_unique', 'dataset.test_split_name=test_unique']' returned non-zero exit status 1.