W0611 06:08:18.948000 1077320 site-packages/torch/distributed/run.py:766] W0611 06:08:18.948000 1077320 site-packages/torch/distributed/run.py:766] ***************************************** W0611 06:08:18.948000 1077320 site-packages/torch/distributed/run.py:766] Setting OMP_NUM_THREADS environment variable for each process to be 1 in default, to avoid your system being overloaded, please further tune the variable for optimal performance in your application as needed. W0611 06:08:18.948000 1077320 site-packages/torch/distributed/run.py:766] ***************************************** Trainer._get_train_sampler replaced with custom implementation. Trainer._get_train_sampler replaced with custom implementation. Trainer._get_train_sampler replaced with custom implementation. Trainer._get_train_sampler replaced with custom implementation. Trainer._get_train_sampler replaced with custom implementation. Trainer._get_train_sampler replaced with custom implementation. Trainer._get_train_sampler replaced with custom implementation. Trainer._get_train_sampler replaced with custom implementation. You are attempting to use Flash Attention 2.0 with a model not initialized on GPU. Make sure to move the model to GPU after initializing it on CPU with `model.to('cuda')`. You are attempting to use Flash Attention 2.0 with a model not initialized on GPU. Make sure to move the model to GPU after initializing it on CPU with `model.to('cuda')`. You are attempting to use Flash Attention 2.0 with a model not initialized on GPU. Make sure to move the model to GPU after initializing it on CPU with `model.to('cuda')`. You are attempting to use Flash Attention 2.0 with a model not initialized on GPU. Make sure to move the model to GPU after initializing it on CPU with `model.to('cuda')`. You are attempting to use Flash Attention 2.0 with a model not initialized on GPU. Make sure to move the model to GPU after initializing it on CPU with `model.to('cuda')`. You are attempting to use Flash Attention 2.0 with a model not initialized on GPU. Make sure to move the model to GPU after initializing it on CPU with `model.to('cuda')`. You are attempting to use Flash Attention 2.0 with a model not initialized on GPU. Make sure to move the model to GPU after initializing it on CPU with `model.to('cuda')`. You are attempting to use Flash Attention 2.0 with a model not initialized on GPU. Make sure to move the model to GPU after initializing it on CPU with `model.to('cuda')`. StartRecentKVCache: 8, 48 StartRecentKVCache: 8, 48 StartRecentKVCache: 8, 48 StartRecentKVCache: 8, 48 StartRecentKVCache: 8, 48 StartRecentKVCache: 8, 48 StartRecentKVCache: 8, 48 StartRecentKVCache: 8, 48 Loading checkpoint shards: 0%| | 0/2 [00:00