=== full run start Mon Sep 14 16:44:32 UTC 2026 === [2026-09-14 16:44:37,556] [WARNING] [runner.py:232:fetch_hostfile] Unable to find hostfile, will proceed with training with local resources only. [2026-09-14 16:44:37,556] [INFO] [runner.py:630:main] cmd = /opt/conda/bin/python -u -m deepspeed.launcher.launch --world_info=eyJsb2NhbGhvc3QiOiBbMCwgMV19 --master_addr=127.0.0.1 --master_port=29500 --enable_each_rank_log=None --log_level=info scripts/training/train.py --model_config configs/model/model_300m.yaml --tokenizer_path tokenizer --train_data_dir /workspace/data/pretrain/train --val_data_dir /workspace/data/pretrain/val_small --output_dir outputs/devops-300m-4096-bf16-vast --max_steps 4000 --warmup_steps 40 --learning_rate 6e-4 --batch_size 1 --grad_accum 16 --logging_steps 10 --save_steps 2000 --eval_steps 2000 --deepspeed_config configs/training/deepspeed_zero2_bf16.json --bf16 --seq_length 4096 --run_name devops-300m-4096-bf16-vast [2026-09-14 16:44:42,440] [INFO] [launch.py:155:main] 0 NCCL_DEBUG=INFO [2026-09-14 16:44:42,440] [INFO] [launch.py:162:main] WORLD INFO DICT: {'localhost': [0, 1]} [2026-09-14 16:44:42,440] [INFO] [launch.py:168:main] nnodes=1, num_local_procs=2, node_rank=0 [2026-09-14 16:44:42,440] [INFO] [launch.py:179:main] global_rank_mapping=defaultdict(, {'localhost': [0, 1]}) [2026-09-14 16:44:42,440] [INFO] [launch.py:180:main] dist_world_size=2 [2026-09-14 16:44:42,440] [INFO] [launch.py:184:main] Setting CUDA_VISIBLE_DEVICES=0,1 [2026-09-14 16:44:42,441] [INFO] [launch.py:272:main] process 10692 spawned with command: ['/opt/conda/bin/python', '-u', 'scripts/training/train.py', '--local_rank=0', '--model_config', 'configs/model/model_300m.yaml', '--tokenizer_path', 'tokenizer', '--train_data_dir', '/workspace/data/pretrain/train', '--val_data_dir', '/workspace/data/pretrain/val_small', '--output_dir', 'outputs/devops-300m-4096-bf16-vast', '--max_steps', '4000', '--warmup_steps', '40', '--learning_rate', '6e-4', '--batch_size', '1', '--grad_accum', '16', '--logging_steps', '10', '--save_steps', '2000', '--eval_steps', '2000', '--deepspeed_config', 'configs/training/deepspeed_zero2_bf16.json', '--bf16', '--seq_length', '4096', '--run_name', 'devops-300m-4096-bf16-vast'] [2026-09-14 16:44:42,441] [INFO] [launch.py:272:main] process 10693 spawned with command: ['/opt/conda/bin/python', '-u', 'scripts/training/train.py', '--local_rank=1', '--model_config', 'configs/model/model_300m.yaml', '--tokenizer_path', 'tokenizer', '--train_data_dir', '/workspace/data/pretrain/train', '--val_data_dir', '/workspace/data/pretrain/val_small', '--output_dir', 'outputs/devops-300m-4096-bf16-vast', '--max_steps', '4000', '--warmup_steps', '40', '--learning_rate', '6e-4', '--batch_size', '1', '--grad_accum', '16', '--logging_steps', '10', '--save_steps', '2000', '--eval_steps', '2000', '--deepspeed_config', 'configs/training/deepspeed_zero2_bf16.json', '--bf16', '--seq_length', '4096', '--run_name', 'devops-300m-4096-bf16-vast'] [Tokenizer] vocab_size=65536 [Tokenizer] vocab_size=65536 [Dataset] 1509 files, 132,068 non-overlapping samples, seq_len=4096 [Dataset] 10 files, 22 non-overlapping samples, seq_len=4096 [Dataset] 1509 files, 132,068 non-overlapping samples, seq_len=4096 [Dataset] 10 files, 22 non-overlapping samples, seq_len=4096 [transformers] LlamaForCausalLM has generative capabilities, as `prepare_inputs_for_generation` is explicitly defined. However, it doesn't directly inherit from `GenerationMixin`. From šŸ‘‰v4.50šŸ‘ˆ onwards, `PreTrainedModel` will NOT inherit from `GenerationMixin`, and this model will lose the ability to call `generate` and other related functions. - If you're using `trust_remote_code=True`, you can get rid of this warning by loading the model with an auto class. See https://huggingface.co/docs/transformers/en/model_doc/auto#auto-classes - If you are the owner of the model architecture code, please modify your model class such that it inherits from `GenerationMixin` (after `PreTrainedModel`, otherwise you'll get an exception). - If you are not the owner of the model architecture class, please contact the model code owner to update it. [transformers] LlamaForCausalLM has generative capabilities, as `prepare_inputs_for_generation` is explicitly defined. However, it doesn't directly inherit from `GenerationMixin`. From šŸ‘‰v4.50šŸ‘ˆ onwards, `PreTrainedModel` will NOT inherit from `GenerationMixin`, and this model will lose the ability to call `generate` and other related functions. - If you're using `trust_remote_code=True`, you can get rid of this warning by loading the model with an auto class. See https://huggingface.co/docs/transformers/en/model_doc/auto#auto-classes - If you are the owner of the model architecture code, please modify your model class such that it inherits from `GenerationMixin` (after `PreTrainedModel`, otherwise you'll get an exception). - If you are not the owner of the model architecture class, please contact the model code owner to update it. [Model] 287.3M parameters [Env] 2 GPUs detected, using DeepSpeed (launch with: deepspeed --master_port 29500 train.py ...) [Model] 287.3M parameters [Env] 2 GPUs detected, using DeepSpeed (launch with: deepspeed --master_port 29500 train.py ...) 81d666a3c511:10692:10692 [0] NCCL INFO Bootstrap : Using eth0:172.17.0.4<0> 81d666a3c511:10692:10692 [0] NCCL INFO NET/Plugin: No plugin found (libnccl-net.so) 81d666a3c511:10692:10692 [0] NCCL INFO NET/Plugin: Plugin load returned 2 : libnccl-net.so: cannot open shared object file: No such file or directory : when loading libnccl-net.so 81d666a3c511:10692:10692 [0] NCCL INFO NET/Plugin: Using internal network plugin. 81d666a3c511:10692:10692 [0] NCCL INFO cudaDriverVersion 13020 NCCL version 2.21.5+cuda12.4 81d666a3c511:10693:10693 [1] NCCL INFO cudaDriverVersion 13020 81d666a3c511:10693:10693 [1] NCCL INFO Bootstrap : Using eth0:172.17.0.4<0> 81d666a3c511:10693:10693 [1] NCCL INFO NET/Plugin: No plugin found (libnccl-net.so) 81d666a3c511:10693:10693 [1] NCCL INFO NET/Plugin: Plugin load returned 2 : libnccl-net.so: cannot open shared object file: No such file or directory : when loading libnccl-net.so 81d666a3c511:10693:10693 [1] NCCL INFO NET/Plugin: Using internal network plugin. 81d666a3c511:10693:10962 [1] NCCL INFO Failed to open libibverbs.so[.1] 81d666a3c511:10693:10962 [1] NCCL INFO NET/Socket : Using [0]eth0:172.17.0.4<0> 81d666a3c511:10693:10962 [1] NCCL INFO Using non-device net plugin version 0 81d666a3c511:10693:10962 [1] NCCL INFO Using network Socket 81d666a3c511:10692:10964 [0] NCCL INFO Failed to open libibverbs.so[.1] 81d666a3c511:10692:10964 [0] NCCL INFO NET/Socket : Using [0]eth0:172.17.0.4<0> 81d666a3c511:10692:10964 [0] NCCL INFO Using non-device net plugin version 0 81d666a3c511:10692:10964 [0] NCCL INFO Using network Socket 81d666a3c511:10692:10964 [0] NCCL INFO ncclCommInitRank comm 0x1adc3b90 rank 0 nranks 2 cudaDev 0 nvmlDev 0 busId d2000 commId 0x86d35acf72d855e7 - Init START 81d666a3c511:10693:10962 [1] NCCL INFO ncclCommInitRank comm 0x1aa33b30 rank 1 nranks 2 cudaDev 1 nvmlDev 1 busId d6000 commId 0x86d35acf72d855e7 - Init START 81d666a3c511:10692:10964 [0] NCCL INFO P2P is disabled between connected GPUs 1 and 0. You can repress this message with NCCL_IGNORE_DISABLED_P2P=1. 81d666a3c511:10692:10964 [0] NCCL INFO P2P is disabled between connected GPUs 0 and 1. You can repress this message with NCCL_IGNORE_DISABLED_P2P=1. 81d666a3c511:10692:10964 [0] NCCL INFO P2P is disabled between connected GPUs 1 and 0. You can repress this message with NCCL_IGNORE_DISABLED_P2P=1. 81d666a3c511:10692:10964 [0] NCCL INFO P2P is disabled between connected GPUs 0 and 1. You can repress this message with NCCL_IGNORE_DISABLED_P2P=1. 81d666a3c511:10692:10964 [0] NCCL INFO Setting affinity for GPU 0 to ffffffff,00000000,ffffffff,00000000 81d666a3c511:10693:10962 [1] NCCL INFO P2P is disabled between connected GPUs 1 and 0. You can repress this message with NCCL_IGNORE_DISABLED_P2P=1. 81d666a3c511:10693:10962 [1] NCCL INFO P2P is disabled between connected GPUs 0 and 1. You can repress this message with NCCL_IGNORE_DISABLED_P2P=1. 81d666a3c511:10693:10962 [1] NCCL INFO P2P is disabled between connected GPUs 1 and 0. You can repress this message with NCCL_IGNORE_DISABLED_P2P=1. 81d666a3c511:10693:10962 [1] NCCL INFO P2P is disabled between connected GPUs 0 and 1. You can repress this message with NCCL_IGNORE_DISABLED_P2P=1. 81d666a3c511:10693:10962 [1] NCCL INFO Setting affinity for GPU 1 to ffffffff,00000000,ffffffff,00000000 81d666a3c511:10692:10964 [0] NCCL INFO comm 0x1adc3b90 rank 0 nRanks 2 nNodes 1 localRanks 2 localRank 0 MNNVL 0 81d666a3c511:10693:10962 [1] NCCL INFO comm 0x1aa33b30 rank 1 nRanks 2 nNodes 1 localRanks 2 localRank 1 MNNVL 0 81d666a3c511:10692:10964 [0] NCCL INFO Channel 00/02 : 0 1 81d666a3c511:10692:10964 [0] NCCL INFO Channel 01/02 : 0 1 81d666a3c511:10692:10964 [0] NCCL INFO Trees [0] 1/-1/-1->0->-1 [1] 1/-1/-1->0->-1 81d666a3c511:10692:10964 [0] NCCL INFO P2P Chunksize set to 131072 81d666a3c511:10693:10962 [1] NCCL INFO Trees [0] -1/-1/-1->1->0 [1] -1/-1/-1->1->0 81d666a3c511:10693:10962 [1] NCCL INFO P2P Chunksize set to 131072 81d666a3c511:10692:10964 [0] NCCL INFO Channel 00 : 0[0] -> 1[1] via SHM/direct/direct 81d666a3c511:10692:10964 [0] NCCL INFO Channel 01 : 0[0] -> 1[1] via SHM/direct/direct 81d666a3c511:10693:10962 [1] NCCL INFO Channel 00 : 1[1] -> 0[0] via SHM/direct/direct 81d666a3c511:10693:10962 [1] NCCL INFO Channel 01 : 1[1] -> 0[0] via SHM/direct/direct 81d666a3c511:10692:10964 [0] NCCL INFO Connected all rings 81d666a3c511:10692:10964 [0] NCCL INFO Connected all trees 81d666a3c511:10693:10962 [1] NCCL INFO Connected all rings 81d666a3c511:10693:10962 [1] NCCL INFO Connected all trees 81d666a3c511:10693:10962 [1] NCCL INFO threadThresholds 8/8/64 | 16/8/64 | 512 | 512 81d666a3c511:10693:10962 [1] NCCL INFO 2 coll channels, 2 collnet channels, 0 nvls channels, 2 p2p channels, 2 p2p channels per peer 81d666a3c511:10692:10964 [0] NCCL INFO threadThresholds 8/8/64 | 16/8/64 | 512 | 512 81d666a3c511:10692:10964 [0] NCCL INFO 2 coll channels, 2 collnet channels, 0 nvls channels, 2 p2p channels, 2 p2p channels per peer 81d666a3c511:10692:10964 [0] NCCL INFO TUNER/Plugin: Plugin load returned 2 : libnccl-net.so: cannot open shared object file: No such file or directory : when loading libnccl-tuner.so 81d666a3c511:10693:10962 [1] NCCL INFO TUNER/Plugin: Plugin load returned 2 : libnccl-net.so: cannot open shared object file: No such file or directory : when loading libnccl-tuner.so 81d666a3c511:10692:10964 [0] NCCL INFO TUNER/Plugin: Using internal tuner plugin. 81d666a3c511:10693:10962 [1] NCCL INFO TUNER/Plugin: Using internal tuner plugin. 81d666a3c511:10692:10964 [0] NCCL INFO ncclCommInitRank comm 0x1adc3b90 rank 0 nranks 2 cudaDev 0 nvmlDev 0 busId d2000 commId 0x86d35acf72d855e7 - Init COMPLETE 81d666a3c511:10693:10962 [1] NCCL INFO ncclCommInitRank comm 0x1aa33b30 rank 1 nranks 2 cudaDev 1 nvmlDev 1 busId d6000 commId 0x86d35acf72d855e7 - Init COMPLETE [Train] Starting: 4000 steps, batch=1Ɨ16Ɨworld_size, lr=0.0006, seq_len=4096 [Train] Starting: 4000 steps, batch=1Ɨ16Ɨworld_size, lr=0.0006, seq_len=4096 [RANK 0] Gradient accumulation steps mismatch: GradientAccumulationPlugin has 1, DeepSpeed config has 16. Using DeepSpeed's value. 81d666a3c511:10693:11106 [1] NCCL INFO Using non-device net plugin version 0 81d666a3c511:10693:11106 [1] NCCL INFO Using network Socket 81d666a3c511:10692:11110 [0] NCCL INFO Using non-device net plugin version 0 81d666a3c511:10692:11110 [0] NCCL INFO Using network Socket 81d666a3c511:10693:11106 [1] NCCL INFO bootstrapSplit: comm 0xc798460 parent 0x1aa33b30 rank 1 nranks 2 color 1530306504 key 1 prev 0 next 0 - DONE 81d666a3c511:10693:11106 [1] NCCL INFO ncclCommSplit comm 0xc798460 rank 1 nranks 2 cudaDev 1 nvmlDev 1 busId d6000 parent 0x1aa33b30 color 1530306504 key 1 commId 0xbc296a4dcdede661 - Init START 81d666a3c511:10692:11110 [0] NCCL INFO bootstrapSplit: comm 0xcf26230 parent 0x1adc3b90 rank 0 nranks 2 color 1530306504 key 0 prev 1 next 1 - DONE 81d666a3c511:10692:11110 [0] NCCL INFO ncclCommSplit comm 0xcf26230 rank 0 nranks 2 cudaDev 0 nvmlDev 0 busId d2000 parent 0x1adc3b90 color 1530306504 key 0 commId 0xbc296a4dcdede661 - Init START 81d666a3c511:10693:11106 [1] NCCL INFO P2P is disabled between connected GPUs 1 and 0. You can repress this message with NCCL_IGNORE_DISABLED_P2P=1. 81d666a3c511:10693:11106 [1] NCCL INFO P2P is disabled between connected GPUs 0 and 1. You can repress this message with NCCL_IGNORE_DISABLED_P2P=1. 81d666a3c511:10693:11106 [1] NCCL INFO P2P is disabled between connected GPUs 1 and 0. You can repress this message with NCCL_IGNORE_DISABLED_P2P=1. 81d666a3c511:10693:11106 [1] NCCL INFO P2P is disabled between connected GPUs 0 and 1. You can repress this message with NCCL_IGNORE_DISABLED_P2P=1. 81d666a3c511:10693:11106 [1] NCCL INFO Setting affinity for GPU 1 to ffffffff,00000000,ffffffff,00000000 81d666a3c511:10692:11110 [0] NCCL INFO P2P is disabled between connected GPUs 1 and 0. You can repress this message with NCCL_IGNORE_DISABLED_P2P=1. 81d666a3c511:10692:11110 [0] NCCL INFO P2P is disabled between connected GPUs 0 and 1. You can repress this message with NCCL_IGNORE_DISABLED_P2P=1. 81d666a3c511:10692:11110 [0] NCCL INFO P2P is disabled between connected GPUs 1 and 0. You can repress this message with NCCL_IGNORE_DISABLED_P2P=1. 81d666a3c511:10692:11110 [0] NCCL INFO P2P is disabled between connected GPUs 0 and 1. You can repress this message with NCCL_IGNORE_DISABLED_P2P=1. 81d666a3c511:10692:11110 [0] NCCL INFO Setting affinity for GPU 0 to ffffffff,00000000,ffffffff,00000000 81d666a3c511:10693:11106 [1] NCCL INFO comm 0xc798460 rank 1 nRanks 2 nNodes 1 localRanks 2 localRank 1 MNNVL 0 81d666a3c511:10692:11110 [0] NCCL INFO comm 0xcf26230 rank 0 nRanks 2 nNodes 1 localRanks 2 localRank 0 MNNVL 0 81d666a3c511:10693:11106 [1] NCCL INFO Trees [0] -1/-1/-1->1->0 [1] -1/-1/-1->1->0 81d666a3c511:10693:11106 [1] NCCL INFO P2P Chunksize set to 131072 81d666a3c511:10692:11110 [0] NCCL INFO Channel 00/02 : 0 1 81d666a3c511:10692:11110 [0] NCCL INFO Channel 01/02 : 0 1 81d666a3c511:10692:11110 [0] NCCL INFO Trees [0] 1/-1/-1->0->-1 [1] 1/-1/-1->0->-1 81d666a3c511:10692:11110 [0] NCCL INFO P2P Chunksize set to 131072 81d666a3c511:10693:11106 [1] NCCL INFO Channel 00 : 1[1] -> 0[0] via SHM/direct/direct 81d666a3c511:10693:11106 [1] NCCL INFO Channel 01 : 1[1] -> 0[0] via SHM/direct/direct 81d666a3c511:10692:11110 [0] NCCL INFO Channel 00 : 0[0] -> 1[1] via SHM/direct/direct 81d666a3c511:10692:11110 [0] NCCL INFO Channel 01 : 0[0] -> 1[1] via SHM/direct/direct 81d666a3c511:10693:11106 [1] NCCL INFO Connected all rings 81d666a3c511:10693:11106 [1] NCCL INFO Connected all trees 81d666a3c511:10693:11106 [1] NCCL INFO threadThresholds 8/8/64 | 16/8/64 | 512 | 512 81d666a3c511:10693:11106 [1] NCCL INFO 2 coll channels, 2 collnet channels, 0 nvls channels, 2 p2p channels, 2 p2p channels per peer 81d666a3c511:10692:11110 [0] NCCL INFO Connected all rings 81d666a3c511:10692:11110 [0] NCCL INFO Connected all trees 81d666a3c511:10692:11110 [0] NCCL INFO threadThresholds 8/8/64 | 16/8/64 | 512 | 512 81d666a3c511:10692:11110 [0] NCCL INFO 2 coll channels, 2 collnet channels, 0 nvls channels, 2 p2p channels, 2 p2p channels per peer 81d666a3c511:10692:11110 [0] NCCL INFO ncclCommSplit comm 0xcf26230 rank 0 nranks 2 cudaDev 0 nvmlDev 0 busId d2000 parent 0x1adc3b90 color 1530306504 key 0 commId 0xbc296a4dcdede661 - Init COMPLETE 81d666a3c511:10693:11106 [1] NCCL INFO ncclCommSplit comm 0xc798460 rank 1 nranks 2 cudaDev 1 nvmlDev 1 busId d6000 parent 0x1aa33b30 color 1530306504 key 1 commId 0xbc296a4dcdede661 - Init COMPLETE [2026-09-14 16:44:56,871] [WARNING] [lr_schedules.py:693:get_lr] Attempting to get learning rate from scheduler before it has started [2026-09-14 16:44:57,091] [WARNING] [lr_schedules.py:693:get_lr] Attempting to get learning rate from scheduler before it has started 0%| | 0/4000 [00:00