==================================================================== [run] flame38m_dense_local mode=dense TEMPORAL=0 H=256 heads=16 ffn=1422 experts=64 topk=6 moe_ffn=176 shared=352 [run] tasks=arc_challenge,arc_easy,hellaswag,openbookqa,piqa,winogrande,boolq,copa,lambada_openai,sciq out=/workspace/FLAME-MoE/results/phase0/runs/flame38m_dense_local/lmeval_1e18_0shot ==================================================================== /usr/local/lib/python3.11/dist-packages/torch/cuda/__init__.py:58: FutureWarning: The pynvml package is deprecated. Please install nvidia-ml-py instead. If you did not install pynvml directly, please report this to the maintainers of the package that installed pynvml for you. import pynvml # type: ignore[import] [run_lmeval] stubbed VLM backends (hf_vlms, vllm_vlms) /usr/local/lib/python3.11/dist-packages/torch/cuda/__init__.py:58: FutureWarning: The pynvml package is deprecated. Please install nvidia-ml-py instead. If you did not install pynvml directly, please report this to the maintainers of the package that installed pynvml for you. import pynvml # type: ignore[import] /usr/local/lib/python3.11/dist-packages/requests/__init__.py:113: RequestsDependencyWarning: urllib3 (2.2.3) or chardet (6.0.0.post1)/charset_normalizer (3.3.2) doesn't match a supported version! warnings.warn( /usr/local/lib/python3.11/dist-packages/torch/cuda/__init__.py:58: FutureWarning: The pynvml package is deprecated. Please install nvidia-ml-py instead. If you did not install pynvml directly, please report this to the maintainers of the package that installed pynvml for you. import pynvml # type: ignore[import] /workspace/FLAME-MoE/.venv/lib/python3.11/site-packages/transformer_engine/pytorch/attention.py:108: UserWarning: To use flash-attn v3, please use the following commands to install: (1) pip install "git+https://github.com/Dao-AILab/flash-attention.git#egg=flashattn-hopper&subdirectory=hopper" (2) python_path=`python -c "import site; print(site.getsitepackages()[0])"` (3) mkdir -p $python_path/flashattn_hopper (4) wget -P $python_path/flashattn_hopper https://raw.githubusercontent.com/Dao-AILab/flash-attention/main/hopper/flash_attn_interface.py warnings.warn( /workspace/FLAME-MoE/Megatron-LM/megatron/core/models/gpt/gpt_layer_specs.py:54: UserWarning: Apex is not installed. Falling back to Torch Norm warnings.warn('Apex is not installed. Falling back to Torch Norm') /workspace/FLAME-MoE/Megatron-LM/megatron/core/models/retro/encoder_spec.py:47: UserWarning: Apex is not installed. Falling back to Torch Norm warnings.warn(f'Apex is not installed. Falling back to Torch Norm') /workspace/FLAME-MoE/Megatron-LM/megatron/core/models/retro/decoder_spec.py:39: UserWarning: Apex is not installed. Falling back to Torch Norm warnings.warn(f'Apex is not installed. Falling back to Torch Norm') 2026-07-23:02:30:31,897 INFO [__main__.py:279] Verbosity set to INFO 2026-07-23:02:31:02,419 INFO [__init__.py:459] The tag 'arc_ca' is already registered as a group, this tag will not be registered. This may affect tasks you want to call. 2026-07-23:02:31:02,433 INFO [__init__.py:459] The tag 'arc_ca' is already registered as a group, this tag will not be registered. This may affect tasks you want to call. 2026-07-23:02:31:18,032 INFO [__main__.py:376] Selected Tasks: ['arc_challenge', 'arc_easy', 'boolq', 'copa', 'hellaswag', 'lambada_openai', 'openbookqa', 'piqa', 'sciq', 'winogrande'] 2026-07-23:02:31:18,053 INFO [evaluator.py:164] Setting random seed to 42 | Setting numpy seed to 42 | Setting torch manual seed to 42 | Setting fewshot manual seed to 42 2026-07-23:02:31:18,053 INFO [evaluator.py:201] Initializing megatron_lm model, with arguments: {} using world size: 1, data-parallel size: 1, context-parallel size: 1, hierarchical context-parallel sizes: Nonetensor-model-parallel size: 1, encoder-tensor-model-parallel size: 0, pipeline-model-parallel size: 1, encoder-pipeline-model-parallel size: 0 WARNING: overriding default arguments for exit_on_missing_checkpoint:True with exit_on_missing_checkpoint:False setting global batch size to 32 Number of virtual stages per pipeline stage: None accumulate and all-reduce gradients in fp32 for bfloat16 data type. using torch.bfloat16 for parameters ... ------------------------ arguments ------------------------ account_for_embedding_in_pipeline_split ......... False account_for_loss_in_pipeline_split .............. False accumulate_allreduce_grads_in_fp32 .............. True adam_beta1 ...................................... 0.9 adam_beta2 ...................................... 0.999 adam_eps ........................................ 1e-08 add_bias_linear ................................. False add_position_embedding .......................... True add_qkv_bias .................................... False adlr_autoresume ................................. False adlr_autoresume_interval ........................ 1000 align_grad_reduce ............................... True align_param_gather .............................. False app_tag_run_name ................................ None app_tag_run_version ............................. 0.0.0 apply_layernorm_1p .............................. False apply_query_key_layer_scaling ................... False apply_residual_connection_post_layernorm ........ False apply_rope_fusion ............................... True async_save ...................................... None async_tensor_model_parallel_allreduce ........... True attention_backend ............................... AttnBackend.auto attention_dropout ............................... 0.0 attention_softmax_in_fp32 ....................... False auto_detect_ckpt_format ......................... False barrier_with_L1_time ............................ True bert_binary_head ................................ True bert_embedder_type .............................. megatron bert_load ....................................... None bf16 ............................................ True bias_dropout_fusion ............................. True bias_gelu_fusion ................................ False bias_swiglu_fusion .............................. True biencoder_projection_dim ........................ 0 biencoder_shared_query_context_model ............ False block_data_path ................................. None calc_ft_timeouts ................................ False calculate_per_token_loss ........................ False check_for_large_grads ........................... False check_for_nan_in_loss_and_grad .................. True check_for_spiky_loss ............................ False check_weight_hash_across_dp_replicas_interval ... None ckpt_assume_constant_structure .................. False ckpt_convert_format ............................. None ckpt_convert_save ............................... None ckpt_convert_update_legacy_dist_opt_format ...... False ckpt_format ..................................... torch_dist ckpt_fully_parallel_load ........................ False ckpt_fully_parallel_save ........................ True ckpt_fully_parallel_save_deprecated ............. False ckpt_step ....................................... None classes_fraction ................................ 1.0 clip_grad ....................................... 1.0 clone_scatter_output_in_embedding ............... True config_logger_dir ............................... consumed_train_samples .......................... 0 consumed_valid_samples .......................... 0 context_parallel_size ........................... 1 cp_comm_type .................................... ['p2p'] create_attention_mask_in_dataloader ............. True cross_entropy_loss_fusion ....................... False cuda_graph_warmup_steps ......................... 3 data_args_path .................................. None data_cache_path ................................. None data_parallel_random_init ....................... False data_parallel_sharding_strategy ................. no_shard data_parallel_size .............................. 1 data_path ....................................... None data_per_class_fraction ......................... 1.0 data_sharding ................................... True dataloader_type ................................. single ddp_average_in_collective ....................... False ddp_bucket_size ................................. None ddp_num_buckets ................................. None ddp_pad_buckets_for_high_nccl_busbw ............. False decoder_first_pipeline_num_layers ............... None decoder_last_pipeline_num_layers ................ None decoder_num_layers .............................. None decoder_seq_length .............................. None decoupled_lr .................................... None decoupled_min_lr ................................ None decrease_batch_size_if_needed ................... False defer_embedding_wgrad_compute ................... False deprecated_use_mcore_models ..................... False deterministic_mode .............................. False dino_bottleneck_size ............................ 256 dino_freeze_last_layer .......................... 1 dino_head_hidden_size ........................... 2048 dino_local_crops_number ......................... 10 dino_local_img_size ............................. 96 dino_norm_last_layer ............................ False dino_teacher_temp ............................... 0.07 dino_warmup_teacher_temp ........................ 0.04 dino_warmup_teacher_temp_epochs ................. 30 disable_straggler_on_startup .................... False dist_ckpt_format_deprecated ..................... None dist_ckpt_strictness ............................ assume_ok_unexpected distribute_saved_activations .................... False distributed_backend ............................. nccl distributed_timeout_minutes ..................... 10 embedding_path .................................. None empty_unused_memory_level ....................... 0 enable_cuda_graph ............................... False enable_ft_package ............................... False enable_gloo_process_groups ...................... True enable_one_logger ............................... True encoder_num_layers .............................. 9 encoder_pipeline_model_parallel_size ............ 0 encoder_seq_length .............................. 2048 encoder_tensor_model_parallel_size .............. 0 end_weight_decay ................................ 0.01 eod_mask_loss ................................... False error_injection_rate ............................ 0 error_injection_type ............................ transient_error eval_interval ................................... 1000 eval_iters ...................................... 100 evidence_data_path .............................. None exit_duration_in_mins ........................... None exit_interval ................................... None exit_on_missing_checkpoint ...................... False exit_signal_handler ............................. False exp_avg_dtype ................................... torch.float32 exp_avg_sq_dtype ................................ torch.float32 expert_model_parallel_size ...................... 1 expert_tensor_parallel_size ..................... 1 ffn_hidden_size ................................. 1422 finetune ........................................ False flash_decode .................................... False fp16 ............................................ False fp16_lm_cross_entropy ........................... False fp32_residual_connection ........................ False fp8 ............................................. None fp8_amax_compute_algo ........................... most_recent fp8_amax_history_len ............................ 1 fp8_interval .................................... 1 fp8_margin ...................................... 0 fp8_param_gather ................................ False fp8_wgrad ....................................... True global_batch_size ............................... 32 grad_reduce_in_bf16 ............................. False gradient_accumulation_fusion .................... False gradient_reduce_div_fusion ...................... True group_query_attention ........................... False head_lr_mult .................................... 1.0 hidden_dropout .................................. 0.0 hidden_size ..................................... 256 hierarchical_context_parallel_sizes ............. None hybrid_attention_ratio .......................... 0.0 hybrid_mlp_ratio ................................ 0.0 hybrid_override_pattern ......................... None hysteresis ...................................... 2 ict_head_size ................................... None ict_load ........................................ None img_h ........................................... 224 img_w ........................................... 224 indexer_batch_size .............................. 128 indexer_log_interval ............................ 1000 inference_batch_times_seqlen_threshold .......... -1 inference_max_batch_size ........................ 8 inference_max_seq_length ........................ 2560 inference_rng_tracker ........................... False init_method_std ................................. 0.02 init_method_xavier_uniform ...................... False init_model_with_meta_device ..................... False initial_loss_scale .............................. 4294967296 iter_per_epoch .................................. 1250 iterations_to_skip .............................. [] keep_fp8_transpose_cache_when_using_custom_fsdp . False kv_channels ..................................... 16 kv_lora_rank .................................... 32 lazy_mpu_init ................................... None load ............................................ /workspace/FLAME-MoE/results/phase0/runs/flame38m_dense_local/ckpt local_rank ...................................... 0 log_interval .................................... 100 log_loss_scale_to_tensorboard ................... True log_memory_to_tensorboard ....................... False log_num_zeros_in_grad ........................... False log_params_norm ................................. False log_progress .................................... False log_straggler ................................... False log_throughput .................................. False log_timers_to_tensorboard ....................... False log_validation_ppl_to_tensorboard ............... False log_world_size_to_tensorboard ................... False logging_level ................................... None loss_scale ...................................... None loss_scale_window ............................... 1000 lr .............................................. None lr_decay_iters .................................. None lr_decay_samples ................................ None lr_decay_style .................................. linear lr_warmup_fraction .............................. None lr_warmup_init .................................. 0.0 lr_warmup_iters ................................. 0 lr_warmup_samples ............................... 0 lr_wsd_decay_iters .............................. None lr_wsd_decay_samples ............................ None lr_wsd_decay_style .............................. exponential main_grads_dtype ................................ torch.float32 main_params_dtype ............................... torch.float32 make_vocab_size_divisible_by .................... 128 manual_gc ....................................... False manual_gc_eval .................................. True manual_gc_interval .............................. 0 mask_factor ..................................... 1.0 mask_prob ....................................... 0.15 mask_type ....................................... random masked_softmax_fusion ........................... True max_position_embeddings ......................... 2048 max_tokens_to_oom ............................... 10000000 memory_snapshot_path ............................ snapshot.pickle merge_file ...................................... None micro_batch_size ................................ 32 microbatch_group_size_per_vp_stage .............. None min_loss_scale .................................. 1.0 min_lr .......................................... 0.0 mmap_bin_files .................................. True mock_data ....................................... False moe_aux_loss_coeff .............................. 0.0 moe_enable_deepep ............................... False moe_expert_capacity_factor ...................... None moe_extended_tp ................................. False moe_ffn_hidden_size ............................. 1422 moe_grouped_gemm ................................ False moe_input_jitter_eps ............................ None moe_layer_freq .................................. 1 moe_layer_recompute ............................. False moe_pad_expert_input_to_capacity ................ False moe_per_layer_logging ........................... False moe_permute_fusion .............................. False moe_router_bias_update_rate ..................... 0.001 moe_router_dtype ................................ None moe_router_enable_expert_bias ................... False moe_router_group_topk ........................... None moe_router_load_balancing_type .................. aux_loss moe_router_num_groups ........................... None moe_router_pre_softmax .......................... False moe_router_score_function ....................... softmax moe_router_topk ................................. 2 moe_router_topk_scaling_factor .................. None moe_shared_expert_intermediate_size ............. None moe_shared_expert_overlap ....................... False moe_token_dispatcher_type ....................... allgather moe_token_drop_policy ........................... probs moe_use_legacy_grouped_gemm ..................... False moe_use_upcycling ............................... False moe_z_loss_coeff ................................ None mscale .......................................... 1.0 mscale_all_dim .................................. 1.0 multi_latent_attention .......................... False nccl_communicator_config_path ................... None no_load_optim ................................... True no_load_rng ..................................... True no_persist_layer_norm ........................... False no_save_optim ................................... None no_save_rng ..................................... None non_persistent_ckpt_type ........................ None non_persistent_global_ckpt_dir .................. None non_persistent_local_ckpt_algo .................. fully_parallel non_persistent_local_ckpt_dir ................... None non_persistent_save_interval .................... None norm_epsilon .................................... 1e-06 normalization ................................... RMSNorm num_attention_heads ............................. 16 num_channels .................................... 3 num_classes ..................................... 1000 num_dataset_builder_threads ..................... 1 num_distributed_optimizer_instances ............. 1 num_experts ..................................... None num_layers ...................................... 9 num_layers_per_virtual_pipeline_stage ........... None num_query_groups ................................ 1 num_virtual_stages_per_pipeline_rank ............ None num_workers ..................................... 2 one_logger_async ................................ False one_logger_project .............................. megatron-lm one_logger_run_name ............................. None onnx_safe ....................................... None openai_gelu ..................................... False optimizer ....................................... adam optimizer_cpu_offload ........................... False optimizer_offload_fraction ...................... 1.0 output_bert_embeddings .......................... False overlap_cpu_optimizer_d2h_h2d ................... False overlap_grad_reduce ............................. False overlap_p2p_comm ................................ False overlap_p2p_comm_warmup_flush ................... False overlap_param_gather ............................ False overlap_param_gather_with_optimizer_step ........ False override_opt_param_scheduler .................... False params_dtype .................................... torch.bfloat16 patch_dim ....................................... 16 per_split_data_args_path ........................ None perform_initialization .......................... True pin_cpu_grads ................................... True pin_cpu_params .................................. True pipeline_model_parallel_comm_backend ............ None pipeline_model_parallel_size .................... 1 pipeline_model_parallel_split_rank .............. None position_embedding_type ......................... rope pretrained_checkpoint ........................... None profile ......................................... False profile_ranks ................................... [0] profile_step_end ................................ 12 profile_step_start .............................. 10 q_lora_rank ..................................... None qk_head_dim ..................................... 128 qk_layernorm .................................... False qk_pos_emb_head_dim ............................. 64 query_in_block_prob ............................. 0.1 rampup_batch_size ............................... None rank ............................................ 0 recompute_granularity ........................... None recompute_method ................................ None recompute_num_layers ............................ None record_memory_history ........................... False relative_attention_max_distance ................. 128 relative_attention_num_buckets .................. 32 replication ..................................... False replication_factor .............................. 2 replication_jump ................................ None rerun_mode ...................................... disabled reset_attention_mask ............................ False reset_position_ids .............................. False result_rejected_tracker_filename ................ None retriever_report_topk_accuracies ................ [] retriever_score_scaling ......................... False retriever_seq_length ............................ 256 retro_add_retriever ............................. False retro_attention_gate ............................ 1 retro_cyclic_train_iters ........................ None retro_encoder_attention_dropout ................. 0.1 retro_encoder_hidden_dropout .................... 0.1 retro_encoder_layers ............................ 2 retro_num_neighbors ............................. 2 retro_num_retrieved_chunks ...................... 2 retro_project_dir ............................... None retro_verify_neighbor_count ..................... True rope_scaling_factor ............................. 8.0 rotary_base ..................................... 10000 rotary_interleaved .............................. False rotary_percent .................................. 1.0 rotary_scaling_factor ........................... 1.0 rotary_seq_len_interpolation_factor ............. None s3_cache_path ................................... None sample_rate ..................................... 1.0 save ............................................ None save_interval ................................... None scatter_gather_tensors_in_pipeline .............. True seed ............................................ 42 seq_length ...................................... 2048 sequence_parallel ............................... False sgd_momentum .................................... 0.9 short_seq_prob .................................. 0.1 skip_train ...................................... False skipped_train_samples ........................... 0 spec ............................................ None split ........................................... None squared_relu .................................... False start_weight_decay .............................. 0.01 straggler_ctrlr_port ............................ 65535 straggler_minmax_count .......................... 1 suggested_communication_unit_size ............... 400000000 swiglu .......................................... True swin_backbone_type .............................. tiny te_rng_tracker .................................. False tensor_model_parallel_size ...................... 1 tensorboard_dir ................................. None tensorboard_log_interval ........................ 1 tensorboard_queue_size .......................... 1000 test_data_path .................................. None test_mode ....................................... False tiktoken_num_special_tokens ..................... 1000 tiktoken_pattern ................................ None tiktoken_special_tokens ......................... None timing_log_level ................................ 0 timing_log_option ............................... minmax titles_data_path ................................ None tokenizer_model ................................. EleutherAI/pythia-12b tokenizer_type .................................. HuggingFaceTokenizer tp_comm_bootstrap_backend ....................... nccl tp_comm_bulk_dgrad .............................. True tp_comm_bulk_wgrad .............................. True tp_comm_overlap ................................. False tp_comm_overlap_ag .............................. True tp_comm_overlap_cfg ............................. None tp_comm_overlap_rs .............................. True tp_comm_overlap_rs_dgrad ........................ False tp_comm_split_ag ................................ True tp_comm_split_rs ................................ True train_data_path ................................. None train_iters ..................................... None train_samples ................................... None train_sync_interval ............................. None transformer_impl ................................ transformer_engine transformer_pipeline_model_parallel_size ........ 1 untie_embeddings_and_output_weights ............. True use_checkpoint_args ............................. False use_checkpoint_opt_param_scheduler .............. False use_cpu_initialization .......................... None use_custom_fsdp ................................. False use_dist_ckpt ................................... True use_dist_ckpt_deprecated ........................ False use_distributed_optimizer ....................... True use_flash_attn .................................. False use_legacy_models ............................... False use_mp_args_from_checkpoint_args ................ False use_one_sent_docs ............................... False use_persistent_ckpt_worker ...................... False use_precision_aware_optimizer ................... False use_pytorch_profiler ............................ False use_ring_exchange_p2p ........................... False use_rope_scaling ................................ False use_rotary_position_embeddings .................. False use_tokenizer_model_from_checkpoint_args ........ True use_torch_fsdp2 ................................. False use_torch_optimizer_for_cpu_offload ............. False use_tp_pp_dp_mapping ............................ False v_head_dim ...................................... 128 valid_data_path ................................. None variable_seq_lengths ............................ False virtual_pipeline_model_parallel_size ............ None vision_backbone_type ............................ vit vision_pretraining .............................. False vision_pretraining_type ......................... classify vocab_extra_ids ................................. 0 vocab_file ...................................... None vocab_size ...................................... None wandb_exp_name .................................. wandb_project ................................... wandb_save_dir .................................. weight_decay .................................... 0.01 weight_decay_incr_style ......................... constant wgrad_deferral_limit ............................ 0 world_size ...................................... 1 yaml_cfg ........................................ None -------------------- end of arguments --------------------- 2026-07-23:02:31:18,068 INFO [num_microbatches_calculator.py:228] setting number of microbatches to constant 1 > building HuggingFaceTokenizer tokenizer ... 2026-07-23:02:31:18,260 INFO [_client.py:1038] HTTP Request: HEAD https://huggingface.co/EleutherAI/pythia-12b/resolve/main/config.json "HTTP/1.1 307 Temporary Redirect" 2026-07-23:02:31:18,272 INFO [_client.py:1038] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/EleutherAI/pythia-12b/bb1e3e710cdf6b524461d543cfb5ba773f0a81b6/config.json "HTTP/1.1 200 OK" 2026-07-23:02:31:18,410 INFO [_client.py:1038] HTTP Request: HEAD https://huggingface.co/EleutherAI/pythia-12b/resolve/main/tokenizer_config.json "HTTP/1.1 307 Temporary Redirect" 2026-07-23:02:31:18,423 INFO [_client.py:1038] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/EleutherAI/pythia-12b/bb1e3e710cdf6b524461d543cfb5ba773f0a81b6/tokenizer_config.json "HTTP/1.1 200 OK" 2026-07-23:02:31:18,524 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/models/EleutherAI/pythia-12b/tree/main/additional_chat_templates?recursive=false&expand=false "HTTP/1.1 404 Not Found" 2026-07-23:02:31:18,594 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/models/EleutherAI/pythia-12b/tree/main?recursive=true&expand=false "HTTP/1.1 200 OK" > padded vocab (size: 50277) with 27 dummy tokens (new size: 50304) WARNING: one_logger package is required to enable e2e metrics tracking. please go to https://confluence.nvidia.com/display/MLWFO/Package+Repositories for details to install it 2026-07-23:02:31:18,803 WARNING [rerun_state_machine.py:239] RerunStateMachine initialized in mode RerunMode.DISABLED > initializing torch distributed ... [W723 02:31:18.321220886 CUDAAllocatorConfig.h:28] Warning: expandable_segments not supported on this platform (function operator()) > initialized tensor model parallel with size 1 > initialized pipeline model parallel with size 1 > setting random seeds to 42 ... > compiling dataset index builder ... make: Entering directory '/workspace/FLAME-MoE/Megatron-LM/megatron/core/datasets' make: Nothing to be done for 'default'. make: Leaving directory '/workspace/FLAME-MoE/Megatron-LM/megatron/core/datasets' >>> done with dataset index builder. Compilation time: 0.731 seconds > compiling and loading fused kernels ... >>> done with compiling and loading fused kernels. Compilation time: 0.348 seconds building GPT model ... /workspace/FLAME-MoE/Megatron-LM/megatron/core/models/gpt/gpt_layer_specs.py:81: UserWarning: The fp8 argument in "get_gpt_layer_with_transformer_engine_spec" has been deprecated and will be removed soon. Please update your code accordingly. warnings.warn( > number of parameters on (tensor, pipeline) model parallel rank (0, 0): 37948672 /workspace/FLAME-MoE/Megatron-LM/megatron/core/transformer/transformer_layer.py:339: UserWarning: TransformerLayer._get_layer_offset is deprecated.Please use get_transformer_layer_offset instead. warnings.warn( loading distributed checkpoint from /workspace/FLAME-MoE/results/phase0/runs/flame38m_dense_local/ckpt at iteration 2121 /workspace/FLAME-MoE/Megatron-LM/megatron/core/dist_checkpointing/strategies/torch.py:847: FutureWarning: `load_state_dict` is deprecated and will be removed in future versions. Please use `load` instead. checkpoint.load_state_dict( /usr/local/lib/python3.11/dist-packages/torch/distributed/checkpoint/planner_helpers.py:306: FutureWarning: Please use DTensor instead and we are deprecating ShardedTensor. device = getattr(value, "device", None) /workspace/FLAME-MoE/.venv/lib/python3.11/site-packages/transformer_engine/pytorch/module/base.py:566: FutureWarning: You are using `torch.load` with `weights_only=False` (the current default value), which uses the default pickle module implicitly. It is possible to construct malicious pickle data which will execute arbitrary code during unpickling (See https://github.com/pytorch/pytorch/blob/main/SECURITY.md#untrusted-models for more details). In a future release, the default value for `weights_only` will be flipped to `True`. This limits the functions that could be executed during unpickling. Arbitrary objects will no longer be allowed to be loaded via this mode unless they are explicitly allowlisted by the user via `torch.serialization.add_safe_globals`. We recommend you start setting `weights_only=True` for any use case where you don't have full control of the loaded file. Please open an issue on GitHub for any issues related to this experimental feature. state = torch.load(state, map_location="cuda") checkpoint version 3.0 2026-07-23:02:31:20,977 WARNING [rerun_state_machine.py:811] RerunStateMachine disabled via CLI, ignoring machine state saved in checkpoint successfully loaded checkpoint from /workspace/FLAME-MoE/results/phase0/runs/flame38m_dense_local/ckpt [ t 1/1, p 1/1 ] at iteration 2121 2026-07-23:02:31:21,059 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc "HTTP/1.1 200 OK" 2026-07-23:02:31:21,297 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc "HTTP/1.1 200 OK" 2026-07-23:02:31:21,358 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc/revision/210d026faf9955653af8916fad021475a3f00453 "HTTP/1.1 200 OK" 2026-07-23:02:31:21,425 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc/tree/210d026faf9955653af8916fad021475a3f00453?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:31:21,714 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc/tree/210d026faf9955653af8916fad021475a3f00453/ARC-Challenge?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:31:22,191 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc "HTTP/1.1 200 OK" 2026-07-23:02:31:22,420 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc "HTTP/1.1 200 OK" 2026-07-23:02:31:22,719 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc/tree/210d026faf9955653af8916fad021475a3f00453/ARC-Easy?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:31:22,862 WARNING [task.py:799] [Task: boolq] metric acc is defined, but aggregation is not. using default aggregation=mean 2026-07-23:02:31:22,862 WARNING [task.py:811] [Task: boolq] metric acc is defined, but higher_is_better is not. using default higher_is_better=True 2026-07-23:02:31:22,919 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue "HTTP/1.1 200 OK" 2026-07-23:02:31:23,140 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue "HTTP/1.1 200 OK" 2026-07-23:02:31:23,200 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/revision/3de24cf8022e94f4ee4b9d55a6f539891524d646 "HTTP/1.1 200 OK" 2026-07-23:02:31:23,262 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/tree/3de24cf8022e94f4ee4b9d55a6f539891524d646?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:31:23,520 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/tree/3de24cf8022e94f4ee4b9d55a6f539891524d646/axb?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:31:23,588 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/revision/3de24cf8022e94f4ee4b9d55a6f539891524d646 "HTTP/1.1 200 OK" 2026-07-23:02:31:23,669 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/tree/3de24cf8022e94f4ee4b9d55a6f539891524d646/boolq?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:31:23,933 WARNING [task.py:799] [Task: copa] metric acc is defined, but aggregation is not. using default aggregation=mean 2026-07-23:02:31:23,933 WARNING [task.py:811] [Task: copa] metric acc is defined, but higher_is_better is not. using default higher_is_better=True 2026-07-23:02:31:23,996 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue "HTTP/1.1 200 OK" 2026-07-23:02:31:24,247 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue "HTTP/1.1 200 OK" 2026-07-23:02:31:24,507 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/tree/3de24cf8022e94f4ee4b9d55a6f539891524d646/copa?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:31:24,655 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag "HTTP/1.1 200 OK" 2026-07-23:02:31:24,892 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag "HTTP/1.1 200 OK" 2026-07-23:02:31:24,953 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag/revision/218ec52e09a7e7462a5400043bb9a69a41d06b76 "HTTP/1.1 200 OK" 2026-07-23:02:31:25,013 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag/tree/218ec52e09a7e7462a5400043bb9a69a41d06b76?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:31:25,251 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag/tree/218ec52e09a7e7462a5400043bb9a69a41d06b76/data?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:31:25,316 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag/revision/218ec52e09a7e7462a5400043bb9a69a41d06b76 "HTTP/1.1 200 OK" 2026-07-23:02:31:27,443 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai "HTTP/1.1 200 OK" 2026-07-23:02:31:27,677 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai "HTTP/1.1 200 OK" 2026-07-23:02:31:27,780 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/revision/900124bf3b8235c6daf21033af9948b3f07346c4 "HTTP/1.1 200 OK" 2026-07-23:02:31:27,845 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/tree/900124bf3b8235c6daf21033af9948b3f07346c4?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:31:28,103 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/tree/900124bf3b8235c6daf21033af9948b3f07346c4/default%2Ftest?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:31:28,166 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/tree/900124bf3b8235c6daf21033af9948b3f07346c4/default?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:31:28,228 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/revision/900124bf3b8235c6daf21033af9948b3f07346c4 "HTTP/1.1 200 OK" 2026-07-23:02:31:28,248 WARNING [task.py:325] [Task: lambada_openai] has_training_docs and has_validation_docs are False, using test_docs as fewshot_docs but this is not recommended. 2026-07-23:02:31:28,248 WARNING [task.py:325] [Task: lambada_openai] has_training_docs and has_validation_docs are False, using test_docs as fewshot_docs but this is not recommended. 2026-07-23:02:31:28,508 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa "HTTP/1.1 200 OK" 2026-07-23:02:31:28,728 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa "HTTP/1.1 200 OK" 2026-07-23:02:31:28,793 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/revision/388097ea7776314e93a529163e0fea805b8a6454 "HTTP/1.1 200 OK" 2026-07-23:02:31:28,859 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/tree/388097ea7776314e93a529163e0fea805b8a6454?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:31:29,111 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/tree/388097ea7776314e93a529163e0fea805b8a6454/additional?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:31:29,175 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/revision/388097ea7776314e93a529163e0fea805b8a6454 "HTTP/1.1 200 OK" 2026-07-23:02:31:29,270 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/tree/388097ea7776314e93a529163e0fea805b8a6454/main?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:31:29,816 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa "HTTP/1.1 200 OK" 2026-07-23:02:31:29,884 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa "HTTP/1.1 200 OK" 2026-07-23:02:31:29,942 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa/tree/main?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:31:29,994 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/datasets/ybisk/piqa/resolve/main/piqa.py "HTTP/1.1 307 Temporary Redirect" 2026-07-23:02:31:30,005 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/resolve-cache/datasets/ybisk/piqa/2e8ac2dffd59bac8c3c6714948f4c551a0848bb0/piqa.py?%2Fdatasets%2Fybisk%2Fpiqa%2Fresolve%2Fmain%2Fpiqa.py=&etag=%220f0d96172f9bf7aa10069374b3236ef083cfeb8e%22 "HTTP/1.1 200 OK" 2026-07-23:02:31:30,679 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa/revision/main "HTTP/1.1 200 OK" 2026-07-23:02:31:30,741 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa/tree/main?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:31:31,178 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq "HTTP/1.1 200 OK" 2026-07-23:02:31:31,417 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq "HTTP/1.1 200 OK" 2026-07-23:02:31:31,478 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq/revision/2c94ad3e1aafab77146f384e23536f97a4849815 "HTTP/1.1 200 OK" 2026-07-23:02:31:31,541 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq/tree/2c94ad3e1aafab77146f384e23536f97a4849815?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:31:31,776 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq/tree/2c94ad3e1aafab77146f384e23536f97a4849815/data?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:31:31,844 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq/revision/2c94ad3e1aafab77146f384e23536f97a4849815 "HTTP/1.1 200 OK" 2026-07-23:02:31:32,274 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande "HTTP/1.1 200 OK" 2026-07-23:02:31:32,511 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande "HTTP/1.1 200 OK" 2026-07-23:02:31:32,568 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/revision/01e74176c63542e6b0bcb004dcdea22d94fb67b5 "HTTP/1.1 200 OK" 2026-07-23:02:31:32,628 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/tree/01e74176c63542e6b0bcb004dcdea22d94fb67b5?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:31:32,854 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/tree/01e74176c63542e6b0bcb004dcdea22d94fb67b5/winogrande_debiased?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:31:32,916 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/revision/01e74176c63542e6b0bcb004dcdea22d94fb67b5 "HTTP/1.1 200 OK" 2026-07-23:02:31:33,001 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/tree/01e74176c63542e6b0bcb004dcdea22d94fb67b5/winogrande_xl?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:31:33,644 WARNING [evaluator.py:270] Overwriting default num_fewshot of winogrande from None to 0 2026-07-23:02:31:33,644 WARNING [evaluator.py:270] Overwriting default num_fewshot of sciq from None to 0 2026-07-23:02:31:33,645 WARNING [evaluator.py:270] Overwriting default num_fewshot of piqa from None to 0 2026-07-23:02:31:33,645 WARNING [evaluator.py:270] Overwriting default num_fewshot of openbookqa from None to 0 2026-07-23:02:31:33,645 WARNING [evaluator.py:270] Overwriting default num_fewshot of lambada_openai from None to 0 2026-07-23:02:31:33,645 WARNING [evaluator.py:270] Overwriting default num_fewshot of hellaswag from None to 0 2026-07-23:02:31:33,645 WARNING [evaluator.py:270] Overwriting default num_fewshot of copa from None to 0 2026-07-23:02:31:33,645 WARNING [evaluator.py:270] Overwriting default num_fewshot of boolq from None to 0 2026-07-23:02:31:33,645 WARNING [evaluator.py:270] Overwriting default num_fewshot of arc_easy from None to 0 2026-07-23:02:31:33,645 WARNING [evaluator.py:270] Overwriting default num_fewshot of arc_challenge from None to 0 2026-07-23:02:31:33,674 INFO [task.py:415] Building contexts for winogrande on rank 0... 0%| | 0/1267 [00:00 /workspace/FLAME-MoE/results/ablations/flame1e18_downstream.csv ==================================================================== [run] flame38m_g1_moe mode=moe TEMPORAL=0 H=256 heads=16 ffn=1368 experts=64 topk=6 moe_ffn=176 shared=352 [run] tasks=arc_challenge,arc_easy,hellaswag,openbookqa,piqa,winogrande,boolq,copa,lambada_openai,sciq out=/workspace/FLAME-MoE/results/phase0/runs/flame38m_g1_moe/lmeval_1e18_0shot ==================================================================== /usr/local/lib/python3.11/dist-packages/torch/cuda/__init__.py:58: FutureWarning: The pynvml package is deprecated. Please install nvidia-ml-py instead. If you did not install pynvml directly, please report this to the maintainers of the package that installed pynvml for you. import pynvml # type: ignore[import] [run_lmeval] stubbed VLM backends (hf_vlms, vllm_vlms) /usr/local/lib/python3.11/dist-packages/torch/cuda/__init__.py:58: FutureWarning: The pynvml package is deprecated. Please install nvidia-ml-py instead. If you did not install pynvml directly, please report this to the maintainers of the package that installed pynvml for you. import pynvml # type: ignore[import] /usr/local/lib/python3.11/dist-packages/requests/__init__.py:113: RequestsDependencyWarning: urllib3 (2.2.3) or chardet (6.0.0.post1)/charset_normalizer (3.3.2) doesn't match a supported version! warnings.warn( /usr/local/lib/python3.11/dist-packages/torch/cuda/__init__.py:58: FutureWarning: The pynvml package is deprecated. Please install nvidia-ml-py instead. If you did not install pynvml directly, please report this to the maintainers of the package that installed pynvml for you. import pynvml # type: ignore[import] /workspace/FLAME-MoE/.venv/lib/python3.11/site-packages/transformer_engine/pytorch/attention.py:108: UserWarning: To use flash-attn v3, please use the following commands to install: (1) pip install "git+https://github.com/Dao-AILab/flash-attention.git#egg=flashattn-hopper&subdirectory=hopper" (2) python_path=`python -c "import site; print(site.getsitepackages()[0])"` (3) mkdir -p $python_path/flashattn_hopper (4) wget -P $python_path/flashattn_hopper https://raw.githubusercontent.com/Dao-AILab/flash-attention/main/hopper/flash_attn_interface.py warnings.warn( /workspace/FLAME-MoE/Megatron-LM/megatron/core/models/gpt/gpt_layer_specs.py:54: UserWarning: Apex is not installed. Falling back to Torch Norm warnings.warn('Apex is not installed. Falling back to Torch Norm') /workspace/FLAME-MoE/Megatron-LM/megatron/core/models/retro/encoder_spec.py:47: UserWarning: Apex is not installed. Falling back to Torch Norm warnings.warn(f'Apex is not installed. Falling back to Torch Norm') /workspace/FLAME-MoE/Megatron-LM/megatron/core/models/retro/decoder_spec.py:39: UserWarning: Apex is not installed. Falling back to Torch Norm warnings.warn(f'Apex is not installed. Falling back to Torch Norm') 2026-07-23:02:37:02,085 INFO [__main__.py:279] Verbosity set to INFO 2026-07-23:02:37:31,624 INFO [__init__.py:459] The tag 'arc_ca' is already registered as a group, this tag will not be registered. This may affect tasks you want to call. 2026-07-23:02:37:31,638 INFO [__init__.py:459] The tag 'arc_ca' is already registered as a group, this tag will not be registered. This may affect tasks you want to call. 2026-07-23:02:37:44,933 INFO [__main__.py:376] Selected Tasks: ['arc_challenge', 'arc_easy', 'boolq', 'copa', 'hellaswag', 'lambada_openai', 'openbookqa', 'piqa', 'sciq', 'winogrande'] 2026-07-23:02:37:44,947 INFO [evaluator.py:164] Setting random seed to 42 | Setting numpy seed to 42 | Setting torch manual seed to 42 | Setting fewshot manual seed to 42 2026-07-23:02:37:44,947 INFO [evaluator.py:201] Initializing megatron_lm model, with arguments: {} using world size: 1, data-parallel size: 1, context-parallel size: 1, hierarchical context-parallel sizes: Nonetensor-model-parallel size: 1, encoder-tensor-model-parallel size: 0, pipeline-model-parallel size: 1, encoder-pipeline-model-parallel size: 0 WARNING: overriding default arguments for exit_on_missing_checkpoint:True with exit_on_missing_checkpoint:False setting global batch size to 32 Number of virtual stages per pipeline stage: None accumulate and all-reduce gradients in fp32 for bfloat16 data type. using torch.bfloat16 for parameters ... ------------------------ arguments ------------------------ account_for_embedding_in_pipeline_split ......... False account_for_loss_in_pipeline_split .............. False accumulate_allreduce_grads_in_fp32 .............. True adam_beta1 ...................................... 0.9 adam_beta2 ...................................... 0.999 adam_eps ........................................ 1e-08 add_bias_linear ................................. False add_position_embedding .......................... True add_qkv_bias .................................... False adlr_autoresume ................................. False adlr_autoresume_interval ........................ 1000 align_grad_reduce ............................... True align_param_gather .............................. False app_tag_run_name ................................ None app_tag_run_version ............................. 0.0.0 apply_layernorm_1p .............................. False apply_query_key_layer_scaling ................... False apply_residual_connection_post_layernorm ........ False apply_rope_fusion ............................... True async_save ...................................... None async_tensor_model_parallel_allreduce ........... True attention_backend ............................... AttnBackend.auto attention_dropout ............................... 0.0 attention_softmax_in_fp32 ....................... False auto_detect_ckpt_format ......................... False barrier_with_L1_time ............................ True bert_binary_head ................................ True bert_embedder_type .............................. megatron bert_load ....................................... None bf16 ............................................ True bias_dropout_fusion ............................. True bias_gelu_fusion ................................ False bias_swiglu_fusion .............................. True biencoder_projection_dim ........................ 0 biencoder_shared_query_context_model ............ False block_data_path ................................. None calc_ft_timeouts ................................ False calculate_per_token_loss ........................ False check_for_large_grads ........................... False check_for_nan_in_loss_and_grad .................. True check_for_spiky_loss ............................ False check_weight_hash_across_dp_replicas_interval ... None ckpt_assume_constant_structure .................. False ckpt_convert_format ............................. None ckpt_convert_save ............................... None ckpt_convert_update_legacy_dist_opt_format ...... False ckpt_format ..................................... torch_dist ckpt_fully_parallel_load ........................ False ckpt_fully_parallel_save ........................ True ckpt_fully_parallel_save_deprecated ............. False ckpt_step ....................................... None classes_fraction ................................ 1.0 clip_grad ....................................... 1.0 clone_scatter_output_in_embedding ............... True config_logger_dir ............................... consumed_train_samples .......................... 0 consumed_valid_samples .......................... 0 context_parallel_size ........................... 1 cp_comm_type .................................... ['p2p'] create_attention_mask_in_dataloader ............. True cross_entropy_loss_fusion ....................... False cuda_graph_warmup_steps ......................... 3 data_args_path .................................. None data_cache_path ................................. None data_parallel_random_init ....................... False data_parallel_sharding_strategy ................. no_shard data_parallel_size .............................. 1 data_path ....................................... None data_per_class_fraction ......................... 1.0 data_sharding ................................... True dataloader_type ................................. single ddp_average_in_collective ....................... False ddp_bucket_size ................................. None ddp_num_buckets ................................. None ddp_pad_buckets_for_high_nccl_busbw ............. False decoder_first_pipeline_num_layers ............... None decoder_last_pipeline_num_layers ................ None decoder_num_layers .............................. None decoder_seq_length .............................. None decoupled_lr .................................... None decoupled_min_lr ................................ None decrease_batch_size_if_needed ................... False defer_embedding_wgrad_compute ................... False deprecated_use_mcore_models ..................... False deterministic_mode .............................. False dino_bottleneck_size ............................ 256 dino_freeze_last_layer .......................... 1 dino_head_hidden_size ........................... 2048 dino_local_crops_number ......................... 10 dino_local_img_size ............................. 96 dino_norm_last_layer ............................ False dino_teacher_temp ............................... 0.07 dino_warmup_teacher_temp ........................ 0.04 dino_warmup_teacher_temp_epochs ................. 30 disable_straggler_on_startup .................... False dist_ckpt_format_deprecated ..................... None dist_ckpt_strictness ............................ assume_ok_unexpected distribute_saved_activations .................... False distributed_backend ............................. nccl distributed_timeout_minutes ..................... 10 embedding_path .................................. None empty_unused_memory_level ....................... 0 enable_cuda_graph ............................... False enable_ft_package ............................... False enable_gloo_process_groups ...................... True enable_one_logger ............................... True encoder_num_layers .............................. 9 encoder_pipeline_model_parallel_size ............ 0 encoder_seq_length .............................. 2048 encoder_tensor_model_parallel_size .............. 0 end_weight_decay ................................ 0.01 eod_mask_loss ................................... False error_injection_rate ............................ 0 error_injection_type ............................ transient_error eval_interval ................................... 1000 eval_iters ...................................... 100 evidence_data_path .............................. None exit_duration_in_mins ........................... None exit_interval ................................... None exit_on_missing_checkpoint ...................... False exit_signal_handler ............................. False exp_avg_dtype ................................... torch.float32 exp_avg_sq_dtype ................................ torch.float32 expert_model_parallel_size ...................... 1 expert_tensor_parallel_size ..................... 1 ffn_hidden_size ................................. 1368 finetune ........................................ False flash_decode .................................... False fp16 ............................................ False fp16_lm_cross_entropy ........................... False fp32_residual_connection ........................ False fp8 ............................................. None fp8_amax_compute_algo ........................... most_recent fp8_amax_history_len ............................ 1 fp8_interval .................................... 1 fp8_margin ...................................... 0 fp8_param_gather ................................ False fp8_wgrad ....................................... True global_batch_size ............................... 32 grad_reduce_in_bf16 ............................. False gradient_accumulation_fusion .................... False gradient_reduce_div_fusion ...................... True group_query_attention ........................... False head_lr_mult .................................... 1.0 hidden_dropout .................................. 0.0 hidden_size ..................................... 256 hierarchical_context_parallel_sizes ............. None hybrid_attention_ratio .......................... 0.0 hybrid_mlp_ratio ................................ 0.0 hybrid_override_pattern ......................... None hysteresis ...................................... 2 ict_head_size ................................... None ict_load ........................................ None img_h ........................................... 224 img_w ........................................... 224 indexer_batch_size .............................. 128 indexer_log_interval ............................ 1000 inference_batch_times_seqlen_threshold .......... -1 inference_max_batch_size ........................ 8 inference_max_seq_length ........................ 2560 inference_rng_tracker ........................... False init_method_std ................................. 0.02 init_method_xavier_uniform ...................... False init_model_with_meta_device ..................... False initial_loss_scale .............................. 4294967296 iter_per_epoch .................................. 1250 iterations_to_skip .............................. [] keep_fp8_transpose_cache_when_using_custom_fsdp . False kv_channels ..................................... 16 kv_lora_rank .................................... 32 lazy_mpu_init ................................... None load ............................................ /workspace/FLAME-MoE/results/phase0/runs/flame38m_g1_moe/ckpt local_rank ...................................... 0 log_interval .................................... 100 log_loss_scale_to_tensorboard ................... True log_memory_to_tensorboard ....................... False log_num_zeros_in_grad ........................... False log_params_norm ................................. False log_progress .................................... False log_straggler ................................... False log_throughput .................................. False log_timers_to_tensorboard ....................... False log_validation_ppl_to_tensorboard ............... False log_world_size_to_tensorboard ................... False logging_level ................................... None loss_scale ...................................... None loss_scale_window ............................... 1000 lr .............................................. None lr_decay_iters .................................. None lr_decay_samples ................................ None lr_decay_style .................................. linear lr_warmup_fraction .............................. None lr_warmup_init .................................. 0.0 lr_warmup_iters ................................. 0 lr_warmup_samples ............................... 0 lr_wsd_decay_iters .............................. None lr_wsd_decay_samples ............................ None lr_wsd_decay_style .............................. exponential main_grads_dtype ................................ torch.float32 main_params_dtype ............................... torch.float32 make_vocab_size_divisible_by .................... 128 manual_gc ....................................... False manual_gc_eval .................................. True manual_gc_interval .............................. 0 mask_factor ..................................... 1.0 mask_prob ....................................... 0.15 mask_type ....................................... random masked_softmax_fusion ........................... True max_position_embeddings ......................... 2048 max_tokens_to_oom ............................... 10000000 memory_snapshot_path ............................ snapshot.pickle merge_file ...................................... None micro_batch_size ................................ 32 microbatch_group_size_per_vp_stage .............. None min_loss_scale .................................. 1.0 min_lr .......................................... 0.0 mmap_bin_files .................................. True mock_data ....................................... False moe_aux_loss_coeff .............................. 0.01 moe_enable_deepep ............................... False moe_expert_capacity_factor ...................... None moe_extended_tp ................................. False moe_ffn_hidden_size ............................. 176 moe_grouped_gemm ................................ True moe_input_jitter_eps ............................ None moe_layer_freq .................................. [0, 1, 1, 1, 1, 1, 1, 1, 1] moe_layer_recompute ............................. False moe_pad_expert_input_to_capacity ................ False moe_per_layer_logging ........................... False moe_permute_fusion .............................. False moe_router_bias_update_rate ..................... 0.001 moe_router_dtype ................................ fp32 moe_router_enable_expert_bias ................... False moe_router_group_topk ........................... None moe_router_load_balancing_type .................. aux_loss moe_router_num_groups ........................... None moe_router_pre_softmax .......................... True moe_router_score_function ....................... softmax moe_router_topk ................................. 6 moe_router_topk_scaling_factor .................. None moe_shared_expert_intermediate_size ............. 352 moe_shared_expert_overlap ....................... False moe_token_dispatcher_type ....................... alltoall moe_token_drop_policy ........................... probs moe_use_legacy_grouped_gemm ..................... False moe_use_upcycling ............................... False moe_z_loss_coeff ................................ 0.001 mscale .......................................... 1.0 mscale_all_dim .................................. 1.0 multi_latent_attention .......................... False nccl_communicator_config_path ................... None no_load_optim ................................... True no_load_rng ..................................... True no_persist_layer_norm ........................... False no_save_optim ................................... None no_save_rng ..................................... None non_persistent_ckpt_type ........................ None non_persistent_global_ckpt_dir .................. None non_persistent_local_ckpt_algo .................. fully_parallel non_persistent_local_ckpt_dir ................... None non_persistent_save_interval .................... None norm_epsilon .................................... 1e-06 normalization ................................... RMSNorm num_attention_heads ............................. 16 num_channels .................................... 3 num_classes ..................................... 1000 num_dataset_builder_threads ..................... 1 num_distributed_optimizer_instances ............. 1 num_experts ..................................... 64 num_layers ...................................... 9 num_layers_per_virtual_pipeline_stage ........... None num_query_groups ................................ 1 num_virtual_stages_per_pipeline_rank ............ None num_workers ..................................... 2 one_logger_async ................................ False one_logger_project .............................. megatron-lm one_logger_run_name ............................. None onnx_safe ....................................... None openai_gelu ..................................... False optimizer ....................................... adam optimizer_cpu_offload ........................... False optimizer_offload_fraction ...................... 1.0 output_bert_embeddings .......................... False overlap_cpu_optimizer_d2h_h2d ................... False overlap_grad_reduce ............................. False overlap_p2p_comm ................................ False overlap_p2p_comm_warmup_flush ................... False overlap_param_gather ............................ False overlap_param_gather_with_optimizer_step ........ False override_opt_param_scheduler .................... False params_dtype .................................... torch.bfloat16 patch_dim ....................................... 16 per_split_data_args_path ........................ None perform_initialization .......................... True pin_cpu_grads ................................... True pin_cpu_params .................................. True pipeline_model_parallel_comm_backend ............ None pipeline_model_parallel_size .................... 1 pipeline_model_parallel_split_rank .............. None position_embedding_type ......................... rope pretrained_checkpoint ........................... None profile ......................................... False profile_ranks ................................... [0] profile_step_end ................................ 12 profile_step_start .............................. 10 q_lora_rank ..................................... None qk_head_dim ..................................... 128 qk_layernorm .................................... False qk_pos_emb_head_dim ............................. 64 query_in_block_prob ............................. 0.1 rampup_batch_size ............................... None rank ............................................ 0 recompute_granularity ........................... None recompute_method ................................ None recompute_num_layers ............................ None record_memory_history ........................... False relative_attention_max_distance ................. 128 relative_attention_num_buckets .................. 32 replication ..................................... False replication_factor .............................. 2 replication_jump ................................ None rerun_mode ...................................... disabled reset_attention_mask ............................ False reset_position_ids .............................. False result_rejected_tracker_filename ................ None retriever_report_topk_accuracies ................ [] retriever_score_scaling ......................... False retriever_seq_length ............................ 256 retro_add_retriever ............................. False retro_attention_gate ............................ 1 retro_cyclic_train_iters ........................ None retro_encoder_attention_dropout ................. 0.1 retro_encoder_hidden_dropout .................... 0.1 retro_encoder_layers ............................ 2 retro_num_neighbors ............................. 2 retro_num_retrieved_chunks ...................... 2 retro_project_dir ............................... None retro_verify_neighbor_count ..................... True rope_scaling_factor ............................. 8.0 rotary_base ..................................... 10000 rotary_interleaved .............................. False rotary_percent .................................. 1.0 rotary_scaling_factor ........................... 1.0 rotary_seq_len_interpolation_factor ............. None s3_cache_path ................................... None sample_rate ..................................... 1.0 save ............................................ None save_interval ................................... None scatter_gather_tensors_in_pipeline .............. True seed ............................................ 42 seq_length ...................................... 2048 sequence_parallel ............................... False sgd_momentum .................................... 0.9 short_seq_prob .................................. 0.1 skip_train ...................................... False skipped_train_samples ........................... 0 spec ............................................ None split ........................................... None squared_relu .................................... False start_weight_decay .............................. 0.01 straggler_ctrlr_port ............................ 65535 straggler_minmax_count .......................... 1 suggested_communication_unit_size ............... 400000000 swiglu .......................................... True swin_backbone_type .............................. tiny te_rng_tracker .................................. False tensor_model_parallel_size ...................... 1 tensorboard_dir ................................. None tensorboard_log_interval ........................ 1 tensorboard_queue_size .......................... 1000 test_data_path .................................. None test_mode ....................................... False tiktoken_num_special_tokens ..................... 1000 tiktoken_pattern ................................ None tiktoken_special_tokens ......................... None timing_log_level ................................ 0 timing_log_option ............................... minmax titles_data_path ................................ None tokenizer_model ................................. EleutherAI/pythia-12b tokenizer_type .................................. HuggingFaceTokenizer tp_comm_bootstrap_backend ....................... nccl tp_comm_bulk_dgrad .............................. True tp_comm_bulk_wgrad .............................. True tp_comm_overlap ................................. False tp_comm_overlap_ag .............................. True tp_comm_overlap_cfg ............................. None tp_comm_overlap_rs .............................. True tp_comm_overlap_rs_dgrad ........................ False tp_comm_split_ag ................................ True tp_comm_split_rs ................................ True train_data_path ................................. None train_iters ..................................... None train_samples ................................... None train_sync_interval ............................. None transformer_impl ................................ transformer_engine transformer_pipeline_model_parallel_size ........ 1 untie_embeddings_and_output_weights ............. True use_checkpoint_args ............................. False use_checkpoint_opt_param_scheduler .............. False use_cpu_initialization .......................... None use_custom_fsdp ................................. False use_dist_ckpt ................................... True use_dist_ckpt_deprecated ........................ False use_distributed_optimizer ....................... True use_flash_attn .................................. False use_legacy_models ............................... False use_mp_args_from_checkpoint_args ................ False use_one_sent_docs ............................... False use_persistent_ckpt_worker ...................... False use_precision_aware_optimizer ................... False use_pytorch_profiler ............................ False use_ring_exchange_p2p ........................... False use_rope_scaling ................................ False use_rotary_position_embeddings .................. False use_tokenizer_model_from_checkpoint_args ........ True use_torch_fsdp2 ................................. False use_torch_optimizer_for_cpu_offload ............. False use_tp_pp_dp_mapping ............................ False v_head_dim ...................................... 128 valid_data_path ................................. None variable_seq_lengths ............................ False virtual_pipeline_model_parallel_size ............ None vision_backbone_type ............................ vit vision_pretraining .............................. False vision_pretraining_type ......................... classify vocab_extra_ids ................................. 0 vocab_file ...................................... None vocab_size ...................................... None wandb_exp_name .................................. wandb_project ................................... wandb_save_dir .................................. weight_decay .................................... 0.01 weight_decay_incr_style ......................... constant wgrad_deferral_limit ............................ 0 world_size ...................................... 1 yaml_cfg ........................................ None -------------------- end of arguments --------------------- 2026-07-23:02:37:44,953 INFO [num_microbatches_calculator.py:228] setting number of microbatches to constant 1 > building HuggingFaceTokenizer tokenizer ... 2026-07-23:02:37:45,122 INFO [_client.py:1038] HTTP Request: HEAD https://huggingface.co/EleutherAI/pythia-12b/resolve/main/config.json "HTTP/1.1 307 Temporary Redirect" 2026-07-23:02:37:45,133 INFO [_client.py:1038] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/EleutherAI/pythia-12b/bb1e3e710cdf6b524461d543cfb5ba773f0a81b6/config.json "HTTP/1.1 200 OK" 2026-07-23:02:37:45,256 INFO [_client.py:1038] HTTP Request: HEAD https://huggingface.co/EleutherAI/pythia-12b/resolve/main/tokenizer_config.json "HTTP/1.1 307 Temporary Redirect" 2026-07-23:02:37:45,269 INFO [_client.py:1038] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/EleutherAI/pythia-12b/bb1e3e710cdf6b524461d543cfb5ba773f0a81b6/tokenizer_config.json "HTTP/1.1 200 OK" 2026-07-23:02:37:45,371 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/models/EleutherAI/pythia-12b/tree/main/additional_chat_templates?recursive=false&expand=false "HTTP/1.1 404 Not Found" 2026-07-23:02:37:45,438 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/models/EleutherAI/pythia-12b/tree/main?recursive=true&expand=false "HTTP/1.1 200 OK" > padded vocab (size: 50277) with 27 dummy tokens (new size: 50304) WARNING: one_logger package is required to enable e2e metrics tracking. please go to https://confluence.nvidia.com/display/MLWFO/Package+Repositories for details to install it 2026-07-23:02:37:45,639 WARNING [rerun_state_machine.py:239] RerunStateMachine initialized in mode RerunMode.DISABLED > initializing torch distributed ... [W723 02:37:45.156429844 CUDAAllocatorConfig.h:28] Warning: expandable_segments not supported on this platform (function operator()) > initialized tensor model parallel with size 1 > initialized pipeline model parallel with size 1 > setting random seeds to 42 ... > compiling dataset index builder ... make: Entering directory '/workspace/FLAME-MoE/Megatron-LM/megatron/core/datasets' make: Nothing to be done for 'default'. make: Leaving directory '/workspace/FLAME-MoE/Megatron-LM/megatron/core/datasets' >>> done with dataset index builder. Compilation time: 0.698 seconds > compiling and loading fused kernels ... >>> done with compiling and loading fused kernels. Compilation time: 0.355 seconds building GPT model ... > number of parameters on (tensor, pipeline) model parallel rank (0, 0): 100670208 /workspace/FLAME-MoE/Megatron-LM/megatron/core/transformer/transformer_layer.py:339: UserWarning: TransformerLayer._get_layer_offset is deprecated.Please use get_transformer_layer_offset instead. warnings.warn( loading distributed checkpoint from /workspace/FLAME-MoE/results/phase0/runs/flame38m_g1_moe/ckpt at iteration 2121 /workspace/FLAME-MoE/Megatron-LM/megatron/core/dist_checkpointing/strategies/torch.py:847: FutureWarning: `load_state_dict` is deprecated and will be removed in future versions. Please use `load` instead. checkpoint.load_state_dict( /usr/local/lib/python3.11/dist-packages/torch/distributed/checkpoint/planner_helpers.py:306: FutureWarning: Please use DTensor instead and we are deprecating ShardedTensor. device = getattr(value, "device", None) /workspace/FLAME-MoE/.venv/lib/python3.11/site-packages/transformer_engine/pytorch/module/base.py:566: FutureWarning: You are using `torch.load` with `weights_only=False` (the current default value), which uses the default pickle module implicitly. It is possible to construct malicious pickle data which will execute arbitrary code during unpickling (See https://github.com/pytorch/pytorch/blob/main/SECURITY.md#untrusted-models for more details). In a future release, the default value for `weights_only` will be flipped to `True`. This limits the functions that could be executed during unpickling. Arbitrary objects will no longer be allowed to be loaded via this mode unless they are explicitly allowlisted by the user via `torch.serialization.add_safe_globals`. We recommend you start setting `weights_only=True` for any use case where you don't have full control of the loaded file. Please open an issue on GitHub for any issues related to this experimental feature. state = torch.load(state, map_location="cuda") checkpoint version 3.0 2026-07-23:02:37:48,270 WARNING [rerun_state_machine.py:811] RerunStateMachine disabled via CLI, ignoring machine state saved in checkpoint successfully loaded checkpoint from /workspace/FLAME-MoE/results/phase0/runs/flame38m_g1_moe/ckpt [ t 1/1, p 1/1 ] at iteration 2121 2026-07-23:02:37:48,348 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc "HTTP/1.1 200 OK" 2026-07-23:02:37:48,582 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc "HTTP/1.1 200 OK" 2026-07-23:02:37:48,656 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc/revision/210d026faf9955653af8916fad021475a3f00453 "HTTP/1.1 200 OK" 2026-07-23:02:37:48,710 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc/tree/210d026faf9955653af8916fad021475a3f00453?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:37:48,992 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc/tree/210d026faf9955653af8916fad021475a3f00453/ARC-Challenge?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:37:49,506 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc "HTTP/1.1 200 OK" 2026-07-23:02:37:49,725 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc "HTTP/1.1 200 OK" 2026-07-23:02:37:49,949 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc/tree/210d026faf9955653af8916fad021475a3f00453/ARC-Easy?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:37:50,088 WARNING [task.py:799] [Task: boolq] metric acc is defined, but aggregation is not. using default aggregation=mean 2026-07-23:02:37:50,088 WARNING [task.py:811] [Task: boolq] metric acc is defined, but higher_is_better is not. using default higher_is_better=True 2026-07-23:02:37:50,148 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue "HTTP/1.1 200 OK" 2026-07-23:02:37:50,375 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue "HTTP/1.1 200 OK" 2026-07-23:02:37:50,439 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/revision/3de24cf8022e94f4ee4b9d55a6f539891524d646 "HTTP/1.1 200 OK" 2026-07-23:02:37:50,500 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/tree/3de24cf8022e94f4ee4b9d55a6f539891524d646?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:37:50,803 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/tree/3de24cf8022e94f4ee4b9d55a6f539891524d646/axb?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:37:50,867 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/revision/3de24cf8022e94f4ee4b9d55a6f539891524d646 "HTTP/1.1 200 OK" 2026-07-23:02:37:50,949 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/tree/3de24cf8022e94f4ee4b9d55a6f539891524d646/boolq?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:37:51,197 WARNING [task.py:799] [Task: copa] metric acc is defined, but aggregation is not. using default aggregation=mean 2026-07-23:02:37:51,197 WARNING [task.py:811] [Task: copa] metric acc is defined, but higher_is_better is not. using default higher_is_better=True 2026-07-23:02:37:51,258 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue "HTTP/1.1 200 OK" 2026-07-23:02:37:51,492 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue "HTTP/1.1 200 OK" 2026-07-23:02:37:51,828 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/tree/3de24cf8022e94f4ee4b9d55a6f539891524d646/copa?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:37:51,961 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag "HTTP/1.1 200 OK" 2026-07-23:02:37:52,210 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag "HTTP/1.1 200 OK" 2026-07-23:02:37:52,336 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag/revision/218ec52e09a7e7462a5400043bb9a69a41d06b76 "HTTP/1.1 200 OK" 2026-07-23:02:37:52,403 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag/tree/218ec52e09a7e7462a5400043bb9a69a41d06b76?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:37:52,621 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag/tree/218ec52e09a7e7462a5400043bb9a69a41d06b76/data?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:37:52,686 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag/revision/218ec52e09a7e7462a5400043bb9a69a41d06b76 "HTTP/1.1 200 OK" 2026-07-23:02:37:54,981 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai "HTTP/1.1 200 OK" 2026-07-23:02:37:55,276 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai "HTTP/1.1 200 OK" 2026-07-23:02:37:55,341 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/revision/900124bf3b8235c6daf21033af9948b3f07346c4 "HTTP/1.1 200 OK" 2026-07-23:02:37:55,404 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/tree/900124bf3b8235c6daf21033af9948b3f07346c4?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:37:55,637 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/tree/900124bf3b8235c6daf21033af9948b3f07346c4/default%2Ftest?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:37:55,701 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/tree/900124bf3b8235c6daf21033af9948b3f07346c4/default?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:37:55,765 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/revision/900124bf3b8235c6daf21033af9948b3f07346c4 "HTTP/1.1 200 OK" 2026-07-23:02:37:55,783 WARNING [task.py:325] [Task: lambada_openai] has_training_docs and has_validation_docs are False, using test_docs as fewshot_docs but this is not recommended. 2026-07-23:02:37:55,784 WARNING [task.py:325] [Task: lambada_openai] has_training_docs and has_validation_docs are False, using test_docs as fewshot_docs but this is not recommended. 2026-07-23:02:37:55,948 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa "HTTP/1.1 200 OK" 2026-07-23:02:37:56,171 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa "HTTP/1.1 200 OK" 2026-07-23:02:37:56,237 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/revision/388097ea7776314e93a529163e0fea805b8a6454 "HTTP/1.1 200 OK" 2026-07-23:02:37:56,300 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/tree/388097ea7776314e93a529163e0fea805b8a6454?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:37:56,547 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/tree/388097ea7776314e93a529163e0fea805b8a6454/additional?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:37:56,607 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/revision/388097ea7776314e93a529163e0fea805b8a6454 "HTTP/1.1 200 OK" 2026-07-23:02:37:56,687 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/tree/388097ea7776314e93a529163e0fea805b8a6454/main?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:37:56,966 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa "HTTP/1.1 200 OK" 2026-07-23:02:37:57,028 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa "HTTP/1.1 200 OK" 2026-07-23:02:37:57,092 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa/tree/main?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:37:57,152 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/datasets/ybisk/piqa/resolve/main/piqa.py "HTTP/1.1 307 Temporary Redirect" 2026-07-23:02:37:57,165 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/resolve-cache/datasets/ybisk/piqa/2e8ac2dffd59bac8c3c6714948f4c551a0848bb0/piqa.py?%2Fdatasets%2Fybisk%2Fpiqa%2Fresolve%2Fmain%2Fpiqa.py=&etag=%220f0d96172f9bf7aa10069374b3236ef083cfeb8e%22 "HTTP/1.1 200 OK" 2026-07-23:02:37:57,807 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa/revision/main "HTTP/1.1 200 OK" 2026-07-23:02:37:57,864 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa/tree/main?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:37:58,305 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq "HTTP/1.1 200 OK" 2026-07-23:02:37:58,528 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq "HTTP/1.1 200 OK" 2026-07-23:02:37:58,584 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq/revision/2c94ad3e1aafab77146f384e23536f97a4849815 "HTTP/1.1 200 OK" 2026-07-23:02:37:58,638 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq/tree/2c94ad3e1aafab77146f384e23536f97a4849815?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:37:58,849 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq/tree/2c94ad3e1aafab77146f384e23536f97a4849815/data?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:37:58,911 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq/revision/2c94ad3e1aafab77146f384e23536f97a4849815 "HTTP/1.1 200 OK" 2026-07-23:02:37:59,343 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande "HTTP/1.1 200 OK" 2026-07-23:02:37:59,571 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande "HTTP/1.1 200 OK" 2026-07-23:02:37:59,631 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/revision/01e74176c63542e6b0bcb004dcdea22d94fb67b5 "HTTP/1.1 200 OK" 2026-07-23:02:37:59,688 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/tree/01e74176c63542e6b0bcb004dcdea22d94fb67b5?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:37:59,928 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/tree/01e74176c63542e6b0bcb004dcdea22d94fb67b5/winogrande_debiased?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:37:59,987 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/revision/01e74176c63542e6b0bcb004dcdea22d94fb67b5 "HTTP/1.1 200 OK" 2026-07-23:02:38:00,076 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/tree/01e74176c63542e6b0bcb004dcdea22d94fb67b5/winogrande_xl?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:38:00,718 WARNING [evaluator.py:270] Overwriting default num_fewshot of winogrande from None to 0 2026-07-23:02:38:00,718 WARNING [evaluator.py:270] Overwriting default num_fewshot of sciq from None to 0 2026-07-23:02:38:00,718 WARNING [evaluator.py:270] Overwriting default num_fewshot of piqa from None to 0 2026-07-23:02:38:00,718 WARNING [evaluator.py:270] Overwriting default num_fewshot of openbookqa from None to 0 2026-07-23:02:38:00,718 WARNING [evaluator.py:270] Overwriting default num_fewshot of lambada_openai from None to 0 2026-07-23:02:38:00,718 WARNING [evaluator.py:270] Overwriting default num_fewshot of hellaswag from None to 0 2026-07-23:02:38:00,718 WARNING [evaluator.py:270] Overwriting default num_fewshot of copa from None to 0 2026-07-23:02:38:00,718 WARNING [evaluator.py:270] Overwriting default num_fewshot of boolq from None to 0 2026-07-23:02:38:00,718 WARNING [evaluator.py:270] Overwriting default num_fewshot of arc_easy from None to 0 2026-07-23:02:38:00,718 WARNING [evaluator.py:270] Overwriting default num_fewshot of arc_challenge from None to 0 2026-07-23:02:38:00,748 INFO [task.py:415] Building contexts for winogrande on rank 0... 0%| | 0/1267 [00:00 /workspace/FLAME-MoE/results/ablations/flame1e18_downstream.csv ==================================================================== [run] flame38m_g1_moe_s2 mode=moe TEMPORAL=0 H=256 heads=16 ffn=1368 experts=64 topk=6 moe_ffn=176 shared=352 [run] tasks=arc_challenge,arc_easy,hellaswag,openbookqa,piqa,winogrande,boolq,copa,lambada_openai,sciq out=/workspace/FLAME-MoE/results/phase0/runs/flame38m_g1_moe_s2/lmeval_1e18_0shot ==================================================================== /usr/local/lib/python3.11/dist-packages/torch/cuda/__init__.py:58: FutureWarning: The pynvml package is deprecated. Please install nvidia-ml-py instead. If you did not install pynvml directly, please report this to the maintainers of the package that installed pynvml for you. import pynvml # type: ignore[import] [run_lmeval] stubbed VLM backends (hf_vlms, vllm_vlms) /usr/local/lib/python3.11/dist-packages/torch/cuda/__init__.py:58: FutureWarning: The pynvml package is deprecated. Please install nvidia-ml-py instead. If you did not install pynvml directly, please report this to the maintainers of the package that installed pynvml for you. import pynvml # type: ignore[import] /usr/local/lib/python3.11/dist-packages/requests/__init__.py:113: RequestsDependencyWarning: urllib3 (2.2.3) or chardet (6.0.0.post1)/charset_normalizer (3.3.2) doesn't match a supported version! warnings.warn( /usr/local/lib/python3.11/dist-packages/torch/cuda/__init__.py:58: FutureWarning: The pynvml package is deprecated. Please install nvidia-ml-py instead. If you did not install pynvml directly, please report this to the maintainers of the package that installed pynvml for you. import pynvml # type: ignore[import] /workspace/FLAME-MoE/.venv/lib/python3.11/site-packages/transformer_engine/pytorch/attention.py:108: UserWarning: To use flash-attn v3, please use the following commands to install: (1) pip install "git+https://github.com/Dao-AILab/flash-attention.git#egg=flashattn-hopper&subdirectory=hopper" (2) python_path=`python -c "import site; print(site.getsitepackages()[0])"` (3) mkdir -p $python_path/flashattn_hopper (4) wget -P $python_path/flashattn_hopper https://raw.githubusercontent.com/Dao-AILab/flash-attention/main/hopper/flash_attn_interface.py warnings.warn( /workspace/FLAME-MoE/Megatron-LM/megatron/core/models/gpt/gpt_layer_specs.py:54: UserWarning: Apex is not installed. Falling back to Torch Norm warnings.warn('Apex is not installed. Falling back to Torch Norm') /workspace/FLAME-MoE/Megatron-LM/megatron/core/models/retro/encoder_spec.py:47: UserWarning: Apex is not installed. Falling back to Torch Norm warnings.warn(f'Apex is not installed. Falling back to Torch Norm') /workspace/FLAME-MoE/Megatron-LM/megatron/core/models/retro/decoder_spec.py:39: UserWarning: Apex is not installed. Falling back to Torch Norm warnings.warn(f'Apex is not installed. Falling back to Torch Norm') 2026-07-23:02:45:29,643 INFO [__main__.py:279] Verbosity set to INFO 2026-07-23:02:45:59,499 INFO [__init__.py:459] The tag 'arc_ca' is already registered as a group, this tag will not be registered. This may affect tasks you want to call. 2026-07-23:02:45:59,513 INFO [__init__.py:459] The tag 'arc_ca' is already registered as a group, this tag will not be registered. This may affect tasks you want to call. 2026-07-23:02:46:16,153 INFO [__main__.py:376] Selected Tasks: ['arc_challenge', 'arc_easy', 'boolq', 'copa', 'hellaswag', 'lambada_openai', 'openbookqa', 'piqa', 'sciq', 'winogrande'] 2026-07-23:02:46:16,176 INFO [evaluator.py:164] Setting random seed to 42 | Setting numpy seed to 42 | Setting torch manual seed to 42 | Setting fewshot manual seed to 42 2026-07-23:02:46:16,176 INFO [evaluator.py:201] Initializing megatron_lm model, with arguments: {} using world size: 1, data-parallel size: 1, context-parallel size: 1, hierarchical context-parallel sizes: Nonetensor-model-parallel size: 1, encoder-tensor-model-parallel size: 0, pipeline-model-parallel size: 1, encoder-pipeline-model-parallel size: 0 WARNING: overriding default arguments for exit_on_missing_checkpoint:True with exit_on_missing_checkpoint:False setting global batch size to 32 Number of virtual stages per pipeline stage: None accumulate and all-reduce gradients in fp32 for bfloat16 data type. using torch.bfloat16 for parameters ... ------------------------ arguments ------------------------ account_for_embedding_in_pipeline_split ......... False account_for_loss_in_pipeline_split .............. False accumulate_allreduce_grads_in_fp32 .............. True adam_beta1 ...................................... 0.9 adam_beta2 ...................................... 0.999 adam_eps ........................................ 1e-08 add_bias_linear ................................. False add_position_embedding .......................... True add_qkv_bias .................................... False adlr_autoresume ................................. False adlr_autoresume_interval ........................ 1000 align_grad_reduce ............................... True align_param_gather .............................. False app_tag_run_name ................................ None app_tag_run_version ............................. 0.0.0 apply_layernorm_1p .............................. False apply_query_key_layer_scaling ................... False apply_residual_connection_post_layernorm ........ False apply_rope_fusion ............................... True async_save ...................................... None async_tensor_model_parallel_allreduce ........... True attention_backend ............................... AttnBackend.auto attention_dropout ............................... 0.0 attention_softmax_in_fp32 ....................... False auto_detect_ckpt_format ......................... False barrier_with_L1_time ............................ True bert_binary_head ................................ True bert_embedder_type .............................. megatron bert_load ....................................... None bf16 ............................................ True bias_dropout_fusion ............................. True bias_gelu_fusion ................................ False bias_swiglu_fusion .............................. True biencoder_projection_dim ........................ 0 biencoder_shared_query_context_model ............ False block_data_path ................................. None calc_ft_timeouts ................................ False calculate_per_token_loss ........................ False check_for_large_grads ........................... False check_for_nan_in_loss_and_grad .................. True check_for_spiky_loss ............................ False check_weight_hash_across_dp_replicas_interval ... None ckpt_assume_constant_structure .................. False ckpt_convert_format ............................. None ckpt_convert_save ............................... None ckpt_convert_update_legacy_dist_opt_format ...... False ckpt_format ..................................... torch_dist ckpt_fully_parallel_load ........................ False ckpt_fully_parallel_save ........................ True ckpt_fully_parallel_save_deprecated ............. False ckpt_step ....................................... None classes_fraction ................................ 1.0 clip_grad ....................................... 1.0 clone_scatter_output_in_embedding ............... True config_logger_dir ............................... consumed_train_samples .......................... 0 consumed_valid_samples .......................... 0 context_parallel_size ........................... 1 cp_comm_type .................................... ['p2p'] create_attention_mask_in_dataloader ............. True cross_entropy_loss_fusion ....................... False cuda_graph_warmup_steps ......................... 3 data_args_path .................................. None data_cache_path ................................. None data_parallel_random_init ....................... False data_parallel_sharding_strategy ................. no_shard data_parallel_size .............................. 1 data_path ....................................... None data_per_class_fraction ......................... 1.0 data_sharding ................................... True dataloader_type ................................. single ddp_average_in_collective ....................... False ddp_bucket_size ................................. None ddp_num_buckets ................................. None ddp_pad_buckets_for_high_nccl_busbw ............. False decoder_first_pipeline_num_layers ............... None decoder_last_pipeline_num_layers ................ None decoder_num_layers .............................. None decoder_seq_length .............................. None decoupled_lr .................................... None decoupled_min_lr ................................ None decrease_batch_size_if_needed ................... False defer_embedding_wgrad_compute ................... False deprecated_use_mcore_models ..................... False deterministic_mode .............................. False dino_bottleneck_size ............................ 256 dino_freeze_last_layer .......................... 1 dino_head_hidden_size ........................... 2048 dino_local_crops_number ......................... 10 dino_local_img_size ............................. 96 dino_norm_last_layer ............................ False dino_teacher_temp ............................... 0.07 dino_warmup_teacher_temp ........................ 0.04 dino_warmup_teacher_temp_epochs ................. 30 disable_straggler_on_startup .................... False dist_ckpt_format_deprecated ..................... None dist_ckpt_strictness ............................ assume_ok_unexpected distribute_saved_activations .................... False distributed_backend ............................. nccl distributed_timeout_minutes ..................... 10 embedding_path .................................. None empty_unused_memory_level ....................... 0 enable_cuda_graph ............................... False enable_ft_package ............................... False enable_gloo_process_groups ...................... True enable_one_logger ............................... True encoder_num_layers .............................. 9 encoder_pipeline_model_parallel_size ............ 0 encoder_seq_length .............................. 2048 encoder_tensor_model_parallel_size .............. 0 end_weight_decay ................................ 0.01 eod_mask_loss ................................... False error_injection_rate ............................ 0 error_injection_type ............................ transient_error eval_interval ................................... 1000 eval_iters ...................................... 100 evidence_data_path .............................. None exit_duration_in_mins ........................... None exit_interval ................................... None exit_on_missing_checkpoint ...................... False exit_signal_handler ............................. False exp_avg_dtype ................................... torch.float32 exp_avg_sq_dtype ................................ torch.float32 expert_model_parallel_size ...................... 1 expert_tensor_parallel_size ..................... 1 ffn_hidden_size ................................. 1368 finetune ........................................ False flash_decode .................................... False fp16 ............................................ False fp16_lm_cross_entropy ........................... False fp32_residual_connection ........................ False fp8 ............................................. None fp8_amax_compute_algo ........................... most_recent fp8_amax_history_len ............................ 1 fp8_interval .................................... 1 fp8_margin ...................................... 0 fp8_param_gather ................................ False fp8_wgrad ....................................... True global_batch_size ............................... 32 grad_reduce_in_bf16 ............................. False gradient_accumulation_fusion .................... False gradient_reduce_div_fusion ...................... True group_query_attention ........................... False head_lr_mult .................................... 1.0 hidden_dropout .................................. 0.0 hidden_size ..................................... 256 hierarchical_context_parallel_sizes ............. None hybrid_attention_ratio .......................... 0.0 hybrid_mlp_ratio ................................ 0.0 hybrid_override_pattern ......................... None hysteresis ...................................... 2 ict_head_size ................................... None ict_load ........................................ None img_h ........................................... 224 img_w ........................................... 224 indexer_batch_size .............................. 128 indexer_log_interval ............................ 1000 inference_batch_times_seqlen_threshold .......... -1 inference_max_batch_size ........................ 8 inference_max_seq_length ........................ 2560 inference_rng_tracker ........................... False init_method_std ................................. 0.02 init_method_xavier_uniform ...................... False init_model_with_meta_device ..................... False initial_loss_scale .............................. 4294967296 iter_per_epoch .................................. 1250 iterations_to_skip .............................. [] keep_fp8_transpose_cache_when_using_custom_fsdp . False kv_channels ..................................... 16 kv_lora_rank .................................... 32 lazy_mpu_init ................................... None load ............................................ /workspace/FLAME-MoE/results/phase0/runs/flame38m_g1_moe_s2/ckpt local_rank ...................................... 0 log_interval .................................... 100 log_loss_scale_to_tensorboard ................... True log_memory_to_tensorboard ....................... False log_num_zeros_in_grad ........................... False log_params_norm ................................. False log_progress .................................... False log_straggler ................................... False log_throughput .................................. False log_timers_to_tensorboard ....................... False log_validation_ppl_to_tensorboard ............... False log_world_size_to_tensorboard ................... False logging_level ................................... None loss_scale ...................................... None loss_scale_window ............................... 1000 lr .............................................. None lr_decay_iters .................................. None lr_decay_samples ................................ None lr_decay_style .................................. linear lr_warmup_fraction .............................. None lr_warmup_init .................................. 0.0 lr_warmup_iters ................................. 0 lr_warmup_samples ............................... 0 lr_wsd_decay_iters .............................. None lr_wsd_decay_samples ............................ None lr_wsd_decay_style .............................. exponential main_grads_dtype ................................ torch.float32 main_params_dtype ............................... torch.float32 make_vocab_size_divisible_by .................... 128 manual_gc ....................................... False manual_gc_eval .................................. True manual_gc_interval .............................. 0 mask_factor ..................................... 1.0 mask_prob ....................................... 0.15 mask_type ....................................... random masked_softmax_fusion ........................... True max_position_embeddings ......................... 2048 max_tokens_to_oom ............................... 10000000 memory_snapshot_path ............................ snapshot.pickle merge_file ...................................... None micro_batch_size ................................ 32 microbatch_group_size_per_vp_stage .............. None min_loss_scale .................................. 1.0 min_lr .......................................... 0.0 mmap_bin_files .................................. True mock_data ....................................... False moe_aux_loss_coeff .............................. 0.01 moe_enable_deepep ............................... False moe_expert_capacity_factor ...................... None moe_extended_tp ................................. False moe_ffn_hidden_size ............................. 176 moe_grouped_gemm ................................ True moe_input_jitter_eps ............................ None moe_layer_freq .................................. [0, 1, 1, 1, 1, 1, 1, 1, 1] moe_layer_recompute ............................. False moe_pad_expert_input_to_capacity ................ False moe_per_layer_logging ........................... False moe_permute_fusion .............................. False moe_router_bias_update_rate ..................... 0.001 moe_router_dtype ................................ fp32 moe_router_enable_expert_bias ................... False moe_router_group_topk ........................... None moe_router_load_balancing_type .................. aux_loss moe_router_num_groups ........................... None moe_router_pre_softmax .......................... True moe_router_score_function ....................... softmax moe_router_topk ................................. 6 moe_router_topk_scaling_factor .................. None moe_shared_expert_intermediate_size ............. 352 moe_shared_expert_overlap ....................... False moe_token_dispatcher_type ....................... alltoall moe_token_drop_policy ........................... probs moe_use_legacy_grouped_gemm ..................... False moe_use_upcycling ............................... False moe_z_loss_coeff ................................ 0.001 mscale .......................................... 1.0 mscale_all_dim .................................. 1.0 multi_latent_attention .......................... False nccl_communicator_config_path ................... None no_load_optim ................................... True no_load_rng ..................................... True no_persist_layer_norm ........................... False no_save_optim ................................... None no_save_rng ..................................... None non_persistent_ckpt_type ........................ None non_persistent_global_ckpt_dir .................. None non_persistent_local_ckpt_algo .................. fully_parallel non_persistent_local_ckpt_dir ................... None non_persistent_save_interval .................... None norm_epsilon .................................... 1e-06 normalization ................................... RMSNorm num_attention_heads ............................. 16 num_channels .................................... 3 num_classes ..................................... 1000 num_dataset_builder_threads ..................... 1 num_distributed_optimizer_instances ............. 1 num_experts ..................................... 64 num_layers ...................................... 9 num_layers_per_virtual_pipeline_stage ........... None num_query_groups ................................ 1 num_virtual_stages_per_pipeline_rank ............ None num_workers ..................................... 2 one_logger_async ................................ False one_logger_project .............................. megatron-lm one_logger_run_name ............................. None onnx_safe ....................................... None openai_gelu ..................................... False optimizer ....................................... adam optimizer_cpu_offload ........................... False optimizer_offload_fraction ...................... 1.0 output_bert_embeddings .......................... False overlap_cpu_optimizer_d2h_h2d ................... False overlap_grad_reduce ............................. False overlap_p2p_comm ................................ False overlap_p2p_comm_warmup_flush ................... False overlap_param_gather ............................ False overlap_param_gather_with_optimizer_step ........ False override_opt_param_scheduler .................... False params_dtype .................................... torch.bfloat16 patch_dim ....................................... 16 per_split_data_args_path ........................ None perform_initialization .......................... True pin_cpu_grads ................................... True pin_cpu_params .................................. True pipeline_model_parallel_comm_backend ............ None pipeline_model_parallel_size .................... 1 pipeline_model_parallel_split_rank .............. None position_embedding_type ......................... rope pretrained_checkpoint ........................... None profile ......................................... False profile_ranks ................................... [0] profile_step_end ................................ 12 profile_step_start .............................. 10 q_lora_rank ..................................... None qk_head_dim ..................................... 128 qk_layernorm .................................... False qk_pos_emb_head_dim ............................. 64 query_in_block_prob ............................. 0.1 rampup_batch_size ............................... None rank ............................................ 0 recompute_granularity ........................... None recompute_method ................................ None recompute_num_layers ............................ None record_memory_history ........................... False relative_attention_max_distance ................. 128 relative_attention_num_buckets .................. 32 replication ..................................... False replication_factor .............................. 2 replication_jump ................................ None rerun_mode ...................................... disabled reset_attention_mask ............................ False reset_position_ids .............................. False result_rejected_tracker_filename ................ None retriever_report_topk_accuracies ................ [] retriever_score_scaling ......................... False retriever_seq_length ............................ 256 retro_add_retriever ............................. False retro_attention_gate ............................ 1 retro_cyclic_train_iters ........................ None retro_encoder_attention_dropout ................. 0.1 retro_encoder_hidden_dropout .................... 0.1 retro_encoder_layers ............................ 2 retro_num_neighbors ............................. 2 retro_num_retrieved_chunks ...................... 2 retro_project_dir ............................... None retro_verify_neighbor_count ..................... True rope_scaling_factor ............................. 8.0 rotary_base ..................................... 10000 rotary_interleaved .............................. False rotary_percent .................................. 1.0 rotary_scaling_factor ........................... 1.0 rotary_seq_len_interpolation_factor ............. None s3_cache_path ................................... None sample_rate ..................................... 1.0 save ............................................ None save_interval ................................... None scatter_gather_tensors_in_pipeline .............. True seed ............................................ 42 seq_length ...................................... 2048 sequence_parallel ............................... False sgd_momentum .................................... 0.9 short_seq_prob .................................. 0.1 skip_train ...................................... False skipped_train_samples ........................... 0 spec ............................................ None split ........................................... None squared_relu .................................... False start_weight_decay .............................. 0.01 straggler_ctrlr_port ............................ 65535 straggler_minmax_count .......................... 1 suggested_communication_unit_size ............... 400000000 swiglu .......................................... True swin_backbone_type .............................. tiny te_rng_tracker .................................. False tensor_model_parallel_size ...................... 1 tensorboard_dir ................................. None tensorboard_log_interval ........................ 1 tensorboard_queue_size .......................... 1000 test_data_path .................................. None test_mode ....................................... False tiktoken_num_special_tokens ..................... 1000 tiktoken_pattern ................................ None tiktoken_special_tokens ......................... None timing_log_level ................................ 0 timing_log_option ............................... minmax titles_data_path ................................ None tokenizer_model ................................. EleutherAI/pythia-12b tokenizer_type .................................. HuggingFaceTokenizer tp_comm_bootstrap_backend ....................... nccl tp_comm_bulk_dgrad .............................. True tp_comm_bulk_wgrad .............................. True tp_comm_overlap ................................. False tp_comm_overlap_ag .............................. True tp_comm_overlap_cfg ............................. None tp_comm_overlap_rs .............................. True tp_comm_overlap_rs_dgrad ........................ False tp_comm_split_ag ................................ True tp_comm_split_rs ................................ True train_data_path ................................. None train_iters ..................................... None train_samples ................................... None train_sync_interval ............................. None transformer_impl ................................ transformer_engine transformer_pipeline_model_parallel_size ........ 1 untie_embeddings_and_output_weights ............. True use_checkpoint_args ............................. False use_checkpoint_opt_param_scheduler .............. False use_cpu_initialization .......................... None use_custom_fsdp ................................. False use_dist_ckpt ................................... True use_dist_ckpt_deprecated ........................ False use_distributed_optimizer ....................... True use_flash_attn .................................. False use_legacy_models ............................... False use_mp_args_from_checkpoint_args ................ False use_one_sent_docs ............................... False use_persistent_ckpt_worker ...................... False use_precision_aware_optimizer ................... False use_pytorch_profiler ............................ False use_ring_exchange_p2p ........................... False use_rope_scaling ................................ False use_rotary_position_embeddings .................. False use_tokenizer_model_from_checkpoint_args ........ True use_torch_fsdp2 ................................. False use_torch_optimizer_for_cpu_offload ............. False use_tp_pp_dp_mapping ............................ False v_head_dim ...................................... 128 valid_data_path ................................. None variable_seq_lengths ............................ False virtual_pipeline_model_parallel_size ............ None vision_backbone_type ............................ vit vision_pretraining .............................. False vision_pretraining_type ......................... classify vocab_extra_ids ................................. 0 vocab_file ...................................... None vocab_size ...................................... None wandb_exp_name .................................. wandb_project ................................... wandb_save_dir .................................. weight_decay .................................... 0.01 weight_decay_incr_style ......................... constant wgrad_deferral_limit ............................ 0 world_size ...................................... 1 yaml_cfg ........................................ None -------------------- end of arguments --------------------- 2026-07-23:02:46:16,191 INFO [num_microbatches_calculator.py:228] setting number of microbatches to constant 1 > building HuggingFaceTokenizer tokenizer ... 2026-07-23:02:46:16,400 INFO [_client.py:1038] HTTP Request: HEAD https://huggingface.co/EleutherAI/pythia-12b/resolve/main/config.json "HTTP/1.1 307 Temporary Redirect" 2026-07-23:02:46:16,412 INFO [_client.py:1038] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/EleutherAI/pythia-12b/bb1e3e710cdf6b524461d543cfb5ba773f0a81b6/config.json "HTTP/1.1 200 OK" 2026-07-23:02:46:16,598 INFO [_client.py:1038] HTTP Request: HEAD https://huggingface.co/EleutherAI/pythia-12b/resolve/main/tokenizer_config.json "HTTP/1.1 307 Temporary Redirect" 2026-07-23:02:46:16,611 INFO [_client.py:1038] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/EleutherAI/pythia-12b/bb1e3e710cdf6b524461d543cfb5ba773f0a81b6/tokenizer_config.json "HTTP/1.1 200 OK" 2026-07-23:02:46:16,711 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/models/EleutherAI/pythia-12b/tree/main/additional_chat_templates?recursive=false&expand=false "HTTP/1.1 404 Not Found" 2026-07-23:02:46:16,782 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/models/EleutherAI/pythia-12b/tree/main?recursive=true&expand=false "HTTP/1.1 200 OK" > padded vocab (size: 50277) with 27 dummy tokens (new size: 50304) WARNING: one_logger package is required to enable e2e metrics tracking. please go to https://confluence.nvidia.com/display/MLWFO/Package+Repositories for details to install it 2026-07-23:02:46:16,978 WARNING [rerun_state_machine.py:239] RerunStateMachine initialized in mode RerunMode.DISABLED > initializing torch distributed ... [W723 02:46:16.496283563 CUDAAllocatorConfig.h:28] Warning: expandable_segments not supported on this platform (function operator()) > initialized tensor model parallel with size 1 > initialized pipeline model parallel with size 1 > setting random seeds to 42 ... > compiling dataset index builder ... make: Entering directory '/workspace/FLAME-MoE/Megatron-LM/megatron/core/datasets' make: Nothing to be done for 'default'. make: Leaving directory '/workspace/FLAME-MoE/Megatron-LM/megatron/core/datasets' >>> done with dataset index builder. Compilation time: 0.697 seconds > compiling and loading fused kernels ... >>> done with compiling and loading fused kernels. Compilation time: 0.378 seconds building GPT model ... > number of parameters on (tensor, pipeline) model parallel rank (0, 0): 100670208 /workspace/FLAME-MoE/Megatron-LM/megatron/core/transformer/transformer_layer.py:339: UserWarning: TransformerLayer._get_layer_offset is deprecated.Please use get_transformer_layer_offset instead. warnings.warn( loading distributed checkpoint from /workspace/FLAME-MoE/results/phase0/runs/flame38m_g1_moe_s2/ckpt at iteration 2121 /workspace/FLAME-MoE/Megatron-LM/megatron/core/dist_checkpointing/strategies/torch.py:847: FutureWarning: `load_state_dict` is deprecated and will be removed in future versions. Please use `load` instead. checkpoint.load_state_dict( /usr/local/lib/python3.11/dist-packages/torch/distributed/checkpoint/planner_helpers.py:306: FutureWarning: Please use DTensor instead and we are deprecating ShardedTensor. device = getattr(value, "device", None) /workspace/FLAME-MoE/.venv/lib/python3.11/site-packages/transformer_engine/pytorch/module/base.py:566: FutureWarning: You are using `torch.load` with `weights_only=False` (the current default value), which uses the default pickle module implicitly. It is possible to construct malicious pickle data which will execute arbitrary code during unpickling (See https://github.com/pytorch/pytorch/blob/main/SECURITY.md#untrusted-models for more details). In a future release, the default value for `weights_only` will be flipped to `True`. This limits the functions that could be executed during unpickling. Arbitrary objects will no longer be allowed to be loaded via this mode unless they are explicitly allowlisted by the user via `torch.serialization.add_safe_globals`. We recommend you start setting `weights_only=True` for any use case where you don't have full control of the loaded file. Please open an issue on GitHub for any issues related to this experimental feature. state = torch.load(state, map_location="cuda") checkpoint version 3.0 2026-07-23:02:46:22,407 WARNING [rerun_state_machine.py:811] RerunStateMachine disabled via CLI, ignoring machine state saved in checkpoint successfully loaded checkpoint from /workspace/FLAME-MoE/results/phase0/runs/flame38m_g1_moe_s2/ckpt [ t 1/1, p 1/1 ] at iteration 2121 2026-07-23:02:46:22,577 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc "HTTP/1.1 200 OK" 2026-07-23:02:46:22,835 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc "HTTP/1.1 200 OK" 2026-07-23:02:46:22,900 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc/revision/210d026faf9955653af8916fad021475a3f00453 "HTTP/1.1 200 OK" 2026-07-23:02:46:22,959 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc/tree/210d026faf9955653af8916fad021475a3f00453?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:46:23,248 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc/tree/210d026faf9955653af8916fad021475a3f00453/ARC-Challenge?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:46:23,785 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc "HTTP/1.1 200 OK" 2026-07-23:02:46:24,018 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc "HTTP/1.1 200 OK" 2026-07-23:02:46:24,310 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc/tree/210d026faf9955653af8916fad021475a3f00453/ARC-Easy?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:46:24,454 WARNING [task.py:799] [Task: boolq] metric acc is defined, but aggregation is not. using default aggregation=mean 2026-07-23:02:46:24,455 WARNING [task.py:811] [Task: boolq] metric acc is defined, but higher_is_better is not. using default higher_is_better=True 2026-07-23:02:46:24,514 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue "HTTP/1.1 200 OK" 2026-07-23:02:46:24,755 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue "HTTP/1.1 200 OK" 2026-07-23:02:46:24,817 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/revision/3de24cf8022e94f4ee4b9d55a6f539891524d646 "HTTP/1.1 200 OK" 2026-07-23:02:46:24,880 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/tree/3de24cf8022e94f4ee4b9d55a6f539891524d646?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:46:25,227 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/tree/3de24cf8022e94f4ee4b9d55a6f539891524d646/axb?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:46:25,298 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/revision/3de24cf8022e94f4ee4b9d55a6f539891524d646 "HTTP/1.1 200 OK" 2026-07-23:02:46:25,372 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/tree/3de24cf8022e94f4ee4b9d55a6f539891524d646/boolq?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:46:25,625 WARNING [task.py:799] [Task: copa] metric acc is defined, but aggregation is not. using default aggregation=mean 2026-07-23:02:46:25,625 WARNING [task.py:811] [Task: copa] metric acc is defined, but higher_is_better is not. using default higher_is_better=True 2026-07-23:02:46:25,686 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue "HTTP/1.1 200 OK" 2026-07-23:02:46:25,921 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue "HTTP/1.1 200 OK" 2026-07-23:02:46:26,235 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/tree/3de24cf8022e94f4ee4b9d55a6f539891524d646/copa?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:46:26,374 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag "HTTP/1.1 200 OK" 2026-07-23:02:46:26,637 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag "HTTP/1.1 200 OK" 2026-07-23:02:46:26,702 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag/revision/218ec52e09a7e7462a5400043bb9a69a41d06b76 "HTTP/1.1 200 OK" 2026-07-23:02:46:26,761 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag/tree/218ec52e09a7e7462a5400043bb9a69a41d06b76?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:46:27,065 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag/tree/218ec52e09a7e7462a5400043bb9a69a41d06b76/data?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:46:27,126 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag/revision/218ec52e09a7e7462a5400043bb9a69a41d06b76 "HTTP/1.1 200 OK" 2026-07-23:02:46:29,441 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai "HTTP/1.1 200 OK" 2026-07-23:02:46:29,685 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai "HTTP/1.1 200 OK" 2026-07-23:02:46:29,741 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/revision/900124bf3b8235c6daf21033af9948b3f07346c4 "HTTP/1.1 200 OK" 2026-07-23:02:46:29,801 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/tree/900124bf3b8235c6daf21033af9948b3f07346c4?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:46:30,027 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/tree/900124bf3b8235c6daf21033af9948b3f07346c4/default%2Ftest?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:46:30,094 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/tree/900124bf3b8235c6daf21033af9948b3f07346c4/default?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:46:30,161 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/revision/900124bf3b8235c6daf21033af9948b3f07346c4 "HTTP/1.1 200 OK" 2026-07-23:02:46:30,180 WARNING [task.py:325] [Task: lambada_openai] has_training_docs and has_validation_docs are False, using test_docs as fewshot_docs but this is not recommended. 2026-07-23:02:46:30,180 WARNING [task.py:325] [Task: lambada_openai] has_training_docs and has_validation_docs are False, using test_docs as fewshot_docs but this is not recommended. 2026-07-23:02:46:30,335 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa "HTTP/1.1 200 OK" 2026-07-23:02:46:30,584 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa "HTTP/1.1 200 OK" 2026-07-23:02:46:30,646 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/revision/388097ea7776314e93a529163e0fea805b8a6454 "HTTP/1.1 200 OK" 2026-07-23:02:46:30,703 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/tree/388097ea7776314e93a529163e0fea805b8a6454?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:46:30,927 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/tree/388097ea7776314e93a529163e0fea805b8a6454/additional?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:46:30,984 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/revision/388097ea7776314e93a529163e0fea805b8a6454 "HTTP/1.1 200 OK" 2026-07-23:02:46:31,070 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/tree/388097ea7776314e93a529163e0fea805b8a6454/main?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:46:31,340 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa "HTTP/1.1 200 OK" 2026-07-23:02:46:31,402 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa "HTTP/1.1 200 OK" 2026-07-23:02:46:31,460 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa/tree/main?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:46:31,528 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/datasets/ybisk/piqa/resolve/main/piqa.py "HTTP/1.1 307 Temporary Redirect" 2026-07-23:02:46:31,542 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/resolve-cache/datasets/ybisk/piqa/2e8ac2dffd59bac8c3c6714948f4c551a0848bb0/piqa.py?%2Fdatasets%2Fybisk%2Fpiqa%2Fresolve%2Fmain%2Fpiqa.py=&etag=%220f0d96172f9bf7aa10069374b3236ef083cfeb8e%22 "HTTP/1.1 200 OK" 2026-07-23:02:46:32,080 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa/revision/main "HTTP/1.1 200 OK" 2026-07-23:02:46:32,144 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa/tree/main?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:46:32,577 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq "HTTP/1.1 200 OK" 2026-07-23:02:46:32,815 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq "HTTP/1.1 200 OK" 2026-07-23:02:46:32,875 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq/revision/2c94ad3e1aafab77146f384e23536f97a4849815 "HTTP/1.1 200 OK" 2026-07-23:02:46:32,933 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq/tree/2c94ad3e1aafab77146f384e23536f97a4849815?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:46:33,177 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq/tree/2c94ad3e1aafab77146f384e23536f97a4849815/data?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:46:33,239 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq/revision/2c94ad3e1aafab77146f384e23536f97a4849815 "HTTP/1.1 200 OK" 2026-07-23:02:46:33,650 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande "HTTP/1.1 200 OK" 2026-07-23:02:46:33,864 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande "HTTP/1.1 200 OK" 2026-07-23:02:46:33,928 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/revision/01e74176c63542e6b0bcb004dcdea22d94fb67b5 "HTTP/1.1 200 OK" 2026-07-23:02:46:33,991 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/tree/01e74176c63542e6b0bcb004dcdea22d94fb67b5?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:46:34,262 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/tree/01e74176c63542e6b0bcb004dcdea22d94fb67b5/winogrande_debiased?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:46:34,322 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/revision/01e74176c63542e6b0bcb004dcdea22d94fb67b5 "HTTP/1.1 200 OK" 2026-07-23:02:46:34,414 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/tree/01e74176c63542e6b0bcb004dcdea22d94fb67b5/winogrande_xl?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:46:35,051 WARNING [evaluator.py:270] Overwriting default num_fewshot of winogrande from None to 0 2026-07-23:02:46:35,051 WARNING [evaluator.py:270] Overwriting default num_fewshot of sciq from None to 0 2026-07-23:02:46:35,051 WARNING [evaluator.py:270] Overwriting default num_fewshot of piqa from None to 0 2026-07-23:02:46:35,051 WARNING [evaluator.py:270] Overwriting default num_fewshot of openbookqa from None to 0 2026-07-23:02:46:35,051 WARNING [evaluator.py:270] Overwriting default num_fewshot of lambada_openai from None to 0 2026-07-23:02:46:35,051 WARNING [evaluator.py:270] Overwriting default num_fewshot of hellaswag from None to 0 2026-07-23:02:46:35,051 WARNING [evaluator.py:270] Overwriting default num_fewshot of copa from None to 0 2026-07-23:02:46:35,052 WARNING [evaluator.py:270] Overwriting default num_fewshot of boolq from None to 0 2026-07-23:02:46:35,052 WARNING [evaluator.py:270] Overwriting default num_fewshot of arc_easy from None to 0 2026-07-23:02:46:35,052 WARNING [evaluator.py:270] Overwriting default num_fewshot of arc_challenge from None to 0 2026-07-23:02:46:35,081 INFO [task.py:415] Building contexts for winogrande on rank 0... 0%| | 0/1267 [00:00 /workspace/FLAME-MoE/results/ablations/flame1e18_downstream.csv ==================================================================== [run] flame38m_g1_moe_s3 mode=moe TEMPORAL=0 H=256 heads=16 ffn=1368 experts=64 topk=6 moe_ffn=176 shared=352 [run] tasks=arc_challenge,arc_easy,hellaswag,openbookqa,piqa,winogrande,boolq,copa,lambada_openai,sciq out=/workspace/FLAME-MoE/results/phase0/runs/flame38m_g1_moe_s3/lmeval_1e18_0shot ==================================================================== /usr/local/lib/python3.11/dist-packages/torch/cuda/__init__.py:58: FutureWarning: The pynvml package is deprecated. Please install nvidia-ml-py instead. If you did not install pynvml directly, please report this to the maintainers of the package that installed pynvml for you. import pynvml # type: ignore[import] [run_lmeval] stubbed VLM backends (hf_vlms, vllm_vlms) /usr/local/lib/python3.11/dist-packages/torch/cuda/__init__.py:58: FutureWarning: The pynvml package is deprecated. Please install nvidia-ml-py instead. If you did not install pynvml directly, please report this to the maintainers of the package that installed pynvml for you. import pynvml # type: ignore[import] /usr/local/lib/python3.11/dist-packages/requests/__init__.py:113: RequestsDependencyWarning: urllib3 (2.2.3) or chardet (6.0.0.post1)/charset_normalizer (3.3.2) doesn't match a supported version! warnings.warn( /usr/local/lib/python3.11/dist-packages/torch/cuda/__init__.py:58: FutureWarning: The pynvml package is deprecated. Please install nvidia-ml-py instead. If you did not install pynvml directly, please report this to the maintainers of the package that installed pynvml for you. import pynvml # type: ignore[import] /workspace/FLAME-MoE/.venv/lib/python3.11/site-packages/transformer_engine/pytorch/attention.py:108: UserWarning: To use flash-attn v3, please use the following commands to install: (1) pip install "git+https://github.com/Dao-AILab/flash-attention.git#egg=flashattn-hopper&subdirectory=hopper" (2) python_path=`python -c "import site; print(site.getsitepackages()[0])"` (3) mkdir -p $python_path/flashattn_hopper (4) wget -P $python_path/flashattn_hopper https://raw.githubusercontent.com/Dao-AILab/flash-attention/main/hopper/flash_attn_interface.py warnings.warn( /workspace/FLAME-MoE/Megatron-LM/megatron/core/models/gpt/gpt_layer_specs.py:54: UserWarning: Apex is not installed. Falling back to Torch Norm warnings.warn('Apex is not installed. Falling back to Torch Norm') /workspace/FLAME-MoE/Megatron-LM/megatron/core/models/retro/encoder_spec.py:47: UserWarning: Apex is not installed. Falling back to Torch Norm warnings.warn(f'Apex is not installed. Falling back to Torch Norm') /workspace/FLAME-MoE/Megatron-LM/megatron/core/models/retro/decoder_spec.py:39: UserWarning: Apex is not installed. Falling back to Torch Norm warnings.warn(f'Apex is not installed. Falling back to Torch Norm') 2026-07-23:02:54:07,760 INFO [__main__.py:279] Verbosity set to INFO 2026-07-23:02:54:37,203 INFO [__init__.py:459] The tag 'arc_ca' is already registered as a group, this tag will not be registered. This may affect tasks you want to call. 2026-07-23:02:54:37,215 INFO [__init__.py:459] The tag 'arc_ca' is already registered as a group, this tag will not be registered. This may affect tasks you want to call. 2026-07-23:02:54:53,260 INFO [__main__.py:376] Selected Tasks: ['arc_challenge', 'arc_easy', 'boolq', 'copa', 'hellaswag', 'lambada_openai', 'openbookqa', 'piqa', 'sciq', 'winogrande'] 2026-07-23:02:54:53,276 INFO [evaluator.py:164] Setting random seed to 42 | Setting numpy seed to 42 | Setting torch manual seed to 42 | Setting fewshot manual seed to 42 2026-07-23:02:54:53,276 INFO [evaluator.py:201] Initializing megatron_lm model, with arguments: {} using world size: 1, data-parallel size: 1, context-parallel size: 1, hierarchical context-parallel sizes: Nonetensor-model-parallel size: 1, encoder-tensor-model-parallel size: 0, pipeline-model-parallel size: 1, encoder-pipeline-model-parallel size: 0 WARNING: overriding default arguments for exit_on_missing_checkpoint:True with exit_on_missing_checkpoint:False setting global batch size to 32 Number of virtual stages per pipeline stage: None accumulate and all-reduce gradients in fp32 for bfloat16 data type. using torch.bfloat16 for parameters ... ------------------------ arguments ------------------------ account_for_embedding_in_pipeline_split ......... False account_for_loss_in_pipeline_split .............. False accumulate_allreduce_grads_in_fp32 .............. True adam_beta1 ...................................... 0.9 adam_beta2 ...................................... 0.999 adam_eps ........................................ 1e-08 add_bias_linear ................................. False add_position_embedding .......................... True add_qkv_bias .................................... False adlr_autoresume ................................. False adlr_autoresume_interval ........................ 1000 align_grad_reduce ............................... True align_param_gather .............................. False app_tag_run_name ................................ None app_tag_run_version ............................. 0.0.0 apply_layernorm_1p .............................. False apply_query_key_layer_scaling ................... False apply_residual_connection_post_layernorm ........ False apply_rope_fusion ............................... True async_save ...................................... None async_tensor_model_parallel_allreduce ........... True attention_backend ............................... AttnBackend.auto attention_dropout ............................... 0.0 attention_softmax_in_fp32 ....................... False auto_detect_ckpt_format ......................... False barrier_with_L1_time ............................ True bert_binary_head ................................ True bert_embedder_type .............................. megatron bert_load ....................................... None bf16 ............................................ True bias_dropout_fusion ............................. True bias_gelu_fusion ................................ False bias_swiglu_fusion .............................. True biencoder_projection_dim ........................ 0 biencoder_shared_query_context_model ............ False block_data_path ................................. None calc_ft_timeouts ................................ False calculate_per_token_loss ........................ False check_for_large_grads ........................... False check_for_nan_in_loss_and_grad .................. True check_for_spiky_loss ............................ False check_weight_hash_across_dp_replicas_interval ... None ckpt_assume_constant_structure .................. False ckpt_convert_format ............................. None ckpt_convert_save ............................... None ckpt_convert_update_legacy_dist_opt_format ...... False ckpt_format ..................................... torch_dist ckpt_fully_parallel_load ........................ False ckpt_fully_parallel_save ........................ True ckpt_fully_parallel_save_deprecated ............. False ckpt_step ....................................... None classes_fraction ................................ 1.0 clip_grad ....................................... 1.0 clone_scatter_output_in_embedding ............... True config_logger_dir ............................... consumed_train_samples .......................... 0 consumed_valid_samples .......................... 0 context_parallel_size ........................... 1 cp_comm_type .................................... ['p2p'] create_attention_mask_in_dataloader ............. True cross_entropy_loss_fusion ....................... False cuda_graph_warmup_steps ......................... 3 data_args_path .................................. None data_cache_path ................................. None data_parallel_random_init ....................... False data_parallel_sharding_strategy ................. no_shard data_parallel_size .............................. 1 data_path ....................................... None data_per_class_fraction ......................... 1.0 data_sharding ................................... True dataloader_type ................................. single ddp_average_in_collective ....................... False ddp_bucket_size ................................. None ddp_num_buckets ................................. None ddp_pad_buckets_for_high_nccl_busbw ............. False decoder_first_pipeline_num_layers ............... None decoder_last_pipeline_num_layers ................ None decoder_num_layers .............................. None decoder_seq_length .............................. None decoupled_lr .................................... None decoupled_min_lr ................................ None decrease_batch_size_if_needed ................... False defer_embedding_wgrad_compute ................... False deprecated_use_mcore_models ..................... False deterministic_mode .............................. False dino_bottleneck_size ............................ 256 dino_freeze_last_layer .......................... 1 dino_head_hidden_size ........................... 2048 dino_local_crops_number ......................... 10 dino_local_img_size ............................. 96 dino_norm_last_layer ............................ False dino_teacher_temp ............................... 0.07 dino_warmup_teacher_temp ........................ 0.04 dino_warmup_teacher_temp_epochs ................. 30 disable_straggler_on_startup .................... False dist_ckpt_format_deprecated ..................... None dist_ckpt_strictness ............................ assume_ok_unexpected distribute_saved_activations .................... False distributed_backend ............................. nccl distributed_timeout_minutes ..................... 10 embedding_path .................................. None empty_unused_memory_level ....................... 0 enable_cuda_graph ............................... False enable_ft_package ............................... False enable_gloo_process_groups ...................... True enable_one_logger ............................... True encoder_num_layers .............................. 9 encoder_pipeline_model_parallel_size ............ 0 encoder_seq_length .............................. 2048 encoder_tensor_model_parallel_size .............. 0 end_weight_decay ................................ 0.01 eod_mask_loss ................................... False error_injection_rate ............................ 0 error_injection_type ............................ transient_error eval_interval ................................... 1000 eval_iters ...................................... 100 evidence_data_path .............................. None exit_duration_in_mins ........................... None exit_interval ................................... None exit_on_missing_checkpoint ...................... False exit_signal_handler ............................. False exp_avg_dtype ................................... torch.float32 exp_avg_sq_dtype ................................ torch.float32 expert_model_parallel_size ...................... 1 expert_tensor_parallel_size ..................... 1 ffn_hidden_size ................................. 1368 finetune ........................................ False flash_decode .................................... False fp16 ............................................ False fp16_lm_cross_entropy ........................... False fp32_residual_connection ........................ False fp8 ............................................. None fp8_amax_compute_algo ........................... most_recent fp8_amax_history_len ............................ 1 fp8_interval .................................... 1 fp8_margin ...................................... 0 fp8_param_gather ................................ False fp8_wgrad ....................................... True global_batch_size ............................... 32 grad_reduce_in_bf16 ............................. False gradient_accumulation_fusion .................... False gradient_reduce_div_fusion ...................... True group_query_attention ........................... False head_lr_mult .................................... 1.0 hidden_dropout .................................. 0.0 hidden_size ..................................... 256 hierarchical_context_parallel_sizes ............. None hybrid_attention_ratio .......................... 0.0 hybrid_mlp_ratio ................................ 0.0 hybrid_override_pattern ......................... None hysteresis ...................................... 2 ict_head_size ................................... None ict_load ........................................ None img_h ........................................... 224 img_w ........................................... 224 indexer_batch_size .............................. 128 indexer_log_interval ............................ 1000 inference_batch_times_seqlen_threshold .......... -1 inference_max_batch_size ........................ 8 inference_max_seq_length ........................ 2560 inference_rng_tracker ........................... False init_method_std ................................. 0.02 init_method_xavier_uniform ...................... False init_model_with_meta_device ..................... False initial_loss_scale .............................. 4294967296 iter_per_epoch .................................. 1250 iterations_to_skip .............................. [] keep_fp8_transpose_cache_when_using_custom_fsdp . False kv_channels ..................................... 16 kv_lora_rank .................................... 32 lazy_mpu_init ................................... None load ............................................ /workspace/FLAME-MoE/results/phase0/runs/flame38m_g1_moe_s3/ckpt local_rank ...................................... 0 log_interval .................................... 100 log_loss_scale_to_tensorboard ................... True log_memory_to_tensorboard ....................... False log_num_zeros_in_grad ........................... False log_params_norm ................................. False log_progress .................................... False log_straggler ................................... False log_throughput .................................. False log_timers_to_tensorboard ....................... False log_validation_ppl_to_tensorboard ............... False log_world_size_to_tensorboard ................... False logging_level ................................... None loss_scale ...................................... None loss_scale_window ............................... 1000 lr .............................................. None lr_decay_iters .................................. None lr_decay_samples ................................ None lr_decay_style .................................. linear lr_warmup_fraction .............................. None lr_warmup_init .................................. 0.0 lr_warmup_iters ................................. 0 lr_warmup_samples ............................... 0 lr_wsd_decay_iters .............................. None lr_wsd_decay_samples ............................ None lr_wsd_decay_style .............................. exponential main_grads_dtype ................................ torch.float32 main_params_dtype ............................... torch.float32 make_vocab_size_divisible_by .................... 128 manual_gc ....................................... False manual_gc_eval .................................. True manual_gc_interval .............................. 0 mask_factor ..................................... 1.0 mask_prob ....................................... 0.15 mask_type ....................................... random masked_softmax_fusion ........................... True max_position_embeddings ......................... 2048 max_tokens_to_oom ............................... 10000000 memory_snapshot_path ............................ snapshot.pickle merge_file ...................................... None micro_batch_size ................................ 32 microbatch_group_size_per_vp_stage .............. None min_loss_scale .................................. 1.0 min_lr .......................................... 0.0 mmap_bin_files .................................. True mock_data ....................................... False moe_aux_loss_coeff .............................. 0.01 moe_enable_deepep ............................... False moe_expert_capacity_factor ...................... None moe_extended_tp ................................. False moe_ffn_hidden_size ............................. 176 moe_grouped_gemm ................................ True moe_input_jitter_eps ............................ None moe_layer_freq .................................. [0, 1, 1, 1, 1, 1, 1, 1, 1] moe_layer_recompute ............................. False moe_pad_expert_input_to_capacity ................ False moe_per_layer_logging ........................... False moe_permute_fusion .............................. False moe_router_bias_update_rate ..................... 0.001 moe_router_dtype ................................ fp32 moe_router_enable_expert_bias ................... False moe_router_group_topk ........................... None moe_router_load_balancing_type .................. aux_loss moe_router_num_groups ........................... None moe_router_pre_softmax .......................... True moe_router_score_function ....................... softmax moe_router_topk ................................. 6 moe_router_topk_scaling_factor .................. None moe_shared_expert_intermediate_size ............. 352 moe_shared_expert_overlap ....................... False moe_token_dispatcher_type ....................... alltoall moe_token_drop_policy ........................... probs moe_use_legacy_grouped_gemm ..................... False moe_use_upcycling ............................... False moe_z_loss_coeff ................................ 0.001 mscale .......................................... 1.0 mscale_all_dim .................................. 1.0 multi_latent_attention .......................... False nccl_communicator_config_path ................... None no_load_optim ................................... True no_load_rng ..................................... True no_persist_layer_norm ........................... False no_save_optim ................................... None no_save_rng ..................................... None non_persistent_ckpt_type ........................ None non_persistent_global_ckpt_dir .................. None non_persistent_local_ckpt_algo .................. fully_parallel non_persistent_local_ckpt_dir ................... None non_persistent_save_interval .................... None norm_epsilon .................................... 1e-06 normalization ................................... RMSNorm num_attention_heads ............................. 16 num_channels .................................... 3 num_classes ..................................... 1000 num_dataset_builder_threads ..................... 1 num_distributed_optimizer_instances ............. 1 num_experts ..................................... 64 num_layers ...................................... 9 num_layers_per_virtual_pipeline_stage ........... None num_query_groups ................................ 1 num_virtual_stages_per_pipeline_rank ............ None num_workers ..................................... 2 one_logger_async ................................ False one_logger_project .............................. megatron-lm one_logger_run_name ............................. None onnx_safe ....................................... None openai_gelu ..................................... False optimizer ....................................... adam optimizer_cpu_offload ........................... False optimizer_offload_fraction ...................... 1.0 output_bert_embeddings .......................... False overlap_cpu_optimizer_d2h_h2d ................... False overlap_grad_reduce ............................. False overlap_p2p_comm ................................ False overlap_p2p_comm_warmup_flush ................... False overlap_param_gather ............................ False overlap_param_gather_with_optimizer_step ........ False override_opt_param_scheduler .................... False params_dtype .................................... torch.bfloat16 patch_dim ....................................... 16 per_split_data_args_path ........................ None perform_initialization .......................... True pin_cpu_grads ................................... True pin_cpu_params .................................. True pipeline_model_parallel_comm_backend ............ None pipeline_model_parallel_size .................... 1 pipeline_model_parallel_split_rank .............. None position_embedding_type ......................... rope pretrained_checkpoint ........................... None profile ......................................... False profile_ranks ................................... [0] profile_step_end ................................ 12 profile_step_start .............................. 10 q_lora_rank ..................................... None qk_head_dim ..................................... 128 qk_layernorm .................................... False qk_pos_emb_head_dim ............................. 64 query_in_block_prob ............................. 0.1 rampup_batch_size ............................... None rank ............................................ 0 recompute_granularity ........................... None recompute_method ................................ None recompute_num_layers ............................ None record_memory_history ........................... False relative_attention_max_distance ................. 128 relative_attention_num_buckets .................. 32 replication ..................................... False replication_factor .............................. 2 replication_jump ................................ None rerun_mode ...................................... disabled reset_attention_mask ............................ False reset_position_ids .............................. False result_rejected_tracker_filename ................ None retriever_report_topk_accuracies ................ [] retriever_score_scaling ......................... False retriever_seq_length ............................ 256 retro_add_retriever ............................. False retro_attention_gate ............................ 1 retro_cyclic_train_iters ........................ None retro_encoder_attention_dropout ................. 0.1 retro_encoder_hidden_dropout .................... 0.1 retro_encoder_layers ............................ 2 retro_num_neighbors ............................. 2 retro_num_retrieved_chunks ...................... 2 retro_project_dir ............................... None retro_verify_neighbor_count ..................... True rope_scaling_factor ............................. 8.0 rotary_base ..................................... 10000 rotary_interleaved .............................. False rotary_percent .................................. 1.0 rotary_scaling_factor ........................... 1.0 rotary_seq_len_interpolation_factor ............. None s3_cache_path ................................... None sample_rate ..................................... 1.0 save ............................................ None save_interval ................................... None scatter_gather_tensors_in_pipeline .............. True seed ............................................ 42 seq_length ...................................... 2048 sequence_parallel ............................... False sgd_momentum .................................... 0.9 short_seq_prob .................................. 0.1 skip_train ...................................... False skipped_train_samples ........................... 0 spec ............................................ None split ........................................... None squared_relu .................................... False start_weight_decay .............................. 0.01 straggler_ctrlr_port ............................ 65535 straggler_minmax_count .......................... 1 suggested_communication_unit_size ............... 400000000 swiglu .......................................... True swin_backbone_type .............................. tiny te_rng_tracker .................................. False tensor_model_parallel_size ...................... 1 tensorboard_dir ................................. None tensorboard_log_interval ........................ 1 tensorboard_queue_size .......................... 1000 test_data_path .................................. None test_mode ....................................... False tiktoken_num_special_tokens ..................... 1000 tiktoken_pattern ................................ None tiktoken_special_tokens ......................... None timing_log_level ................................ 0 timing_log_option ............................... minmax titles_data_path ................................ None tokenizer_model ................................. EleutherAI/pythia-12b tokenizer_type .................................. HuggingFaceTokenizer tp_comm_bootstrap_backend ....................... nccl tp_comm_bulk_dgrad .............................. True tp_comm_bulk_wgrad .............................. True tp_comm_overlap ................................. False tp_comm_overlap_ag .............................. True tp_comm_overlap_cfg ............................. None tp_comm_overlap_rs .............................. True tp_comm_overlap_rs_dgrad ........................ False tp_comm_split_ag ................................ True tp_comm_split_rs ................................ True train_data_path ................................. None train_iters ..................................... None train_samples ................................... None train_sync_interval ............................. None transformer_impl ................................ transformer_engine transformer_pipeline_model_parallel_size ........ 1 untie_embeddings_and_output_weights ............. True use_checkpoint_args ............................. False use_checkpoint_opt_param_scheduler .............. False use_cpu_initialization .......................... None use_custom_fsdp ................................. False use_dist_ckpt ................................... True use_dist_ckpt_deprecated ........................ False use_distributed_optimizer ....................... True use_flash_attn .................................. False use_legacy_models ............................... False use_mp_args_from_checkpoint_args ................ False use_one_sent_docs ............................... False use_persistent_ckpt_worker ...................... False use_precision_aware_optimizer ................... False use_pytorch_profiler ............................ False use_ring_exchange_p2p ........................... False use_rope_scaling ................................ False use_rotary_position_embeddings .................. False use_tokenizer_model_from_checkpoint_args ........ True use_torch_fsdp2 ................................. False use_torch_optimizer_for_cpu_offload ............. False use_tp_pp_dp_mapping ............................ False v_head_dim ...................................... 128 valid_data_path ................................. None variable_seq_lengths ............................ False virtual_pipeline_model_parallel_size ............ None vision_backbone_type ............................ vit vision_pretraining .............................. False vision_pretraining_type ......................... classify vocab_extra_ids ................................. 0 vocab_file ...................................... None vocab_size ...................................... None wandb_exp_name .................................. wandb_project ................................... wandb_save_dir .................................. weight_decay .................................... 0.01 weight_decay_incr_style ......................... constant wgrad_deferral_limit ............................ 0 world_size ...................................... 1 yaml_cfg ........................................ None -------------------- end of arguments --------------------- 2026-07-23:02:54:53,290 INFO [num_microbatches_calculator.py:228] setting number of microbatches to constant 1 > building HuggingFaceTokenizer tokenizer ... 2026-07-23:02:54:53,480 INFO [_client.py:1038] HTTP Request: HEAD https://huggingface.co/EleutherAI/pythia-12b/resolve/main/config.json "HTTP/1.1 307 Temporary Redirect" 2026-07-23:02:54:53,493 INFO [_client.py:1038] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/EleutherAI/pythia-12b/bb1e3e710cdf6b524461d543cfb5ba773f0a81b6/config.json "HTTP/1.1 200 OK" 2026-07-23:02:54:53,618 INFO [_client.py:1038] HTTP Request: HEAD https://huggingface.co/EleutherAI/pythia-12b/resolve/main/tokenizer_config.json "HTTP/1.1 307 Temporary Redirect" 2026-07-23:02:54:53,631 INFO [_client.py:1038] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/EleutherAI/pythia-12b/bb1e3e710cdf6b524461d543cfb5ba773f0a81b6/tokenizer_config.json "HTTP/1.1 200 OK" 2026-07-23:02:54:53,737 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/models/EleutherAI/pythia-12b/tree/main/additional_chat_templates?recursive=false&expand=false "HTTP/1.1 404 Not Found" 2026-07-23:02:54:53,804 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/models/EleutherAI/pythia-12b/tree/main?recursive=true&expand=false "HTTP/1.1 200 OK" > padded vocab (size: 50277) with 27 dummy tokens (new size: 50304) WARNING: one_logger package is required to enable e2e metrics tracking. please go to https://confluence.nvidia.com/display/MLWFO/Package+Repositories for details to install it 2026-07-23:02:54:54,002 WARNING [rerun_state_machine.py:239] RerunStateMachine initialized in mode RerunMode.DISABLED > initializing torch distributed ... [W723 02:54:54.520510420 CUDAAllocatorConfig.h:28] Warning: expandable_segments not supported on this platform (function operator()) > initialized tensor model parallel with size 1 > initialized pipeline model parallel with size 1 > setting random seeds to 42 ... > compiling dataset index builder ... make: Entering directory '/workspace/FLAME-MoE/Megatron-LM/megatron/core/datasets' make: Nothing to be done for 'default'. make: Leaving directory '/workspace/FLAME-MoE/Megatron-LM/megatron/core/datasets' >>> done with dataset index builder. Compilation time: 0.557 seconds > compiling and loading fused kernels ... >>> done with compiling and loading fused kernels. Compilation time: 0.356 seconds building GPT model ... > number of parameters on (tensor, pipeline) model parallel rank (0, 0): 100670208 /workspace/FLAME-MoE/Megatron-LM/megatron/core/transformer/transformer_layer.py:339: UserWarning: TransformerLayer._get_layer_offset is deprecated.Please use get_transformer_layer_offset instead. warnings.warn( loading distributed checkpoint from /workspace/FLAME-MoE/results/phase0/runs/flame38m_g1_moe_s3/ckpt at iteration 2121 /workspace/FLAME-MoE/Megatron-LM/megatron/core/dist_checkpointing/strategies/torch.py:847: FutureWarning: `load_state_dict` is deprecated and will be removed in future versions. Please use `load` instead. checkpoint.load_state_dict( /usr/local/lib/python3.11/dist-packages/torch/distributed/checkpoint/planner_helpers.py:306: FutureWarning: Please use DTensor instead and we are deprecating ShardedTensor. device = getattr(value, "device", None) /workspace/FLAME-MoE/.venv/lib/python3.11/site-packages/transformer_engine/pytorch/module/base.py:566: FutureWarning: You are using `torch.load` with `weights_only=False` (the current default value), which uses the default pickle module implicitly. It is possible to construct malicious pickle data which will execute arbitrary code during unpickling (See https://github.com/pytorch/pytorch/blob/main/SECURITY.md#untrusted-models for more details). In a future release, the default value for `weights_only` will be flipped to `True`. This limits the functions that could be executed during unpickling. Arbitrary objects will no longer be allowed to be loaded via this mode unless they are explicitly allowlisted by the user via `torch.serialization.add_safe_globals`. We recommend you start setting `weights_only=True` for any use case where you don't have full control of the loaded file. Please open an issue on GitHub for any issues related to this experimental feature. state = torch.load(state, map_location="cuda") checkpoint version 3.0 2026-07-23:02:54:59,094 WARNING [rerun_state_machine.py:811] RerunStateMachine disabled via CLI, ignoring machine state saved in checkpoint successfully loaded checkpoint from /workspace/FLAME-MoE/results/phase0/runs/flame38m_g1_moe_s3/ckpt [ t 1/1, p 1/1 ] at iteration 2121 2026-07-23:02:54:59,231 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc "HTTP/1.1 200 OK" 2026-07-23:02:54:59,466 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc "HTTP/1.1 200 OK" 2026-07-23:02:54:59,568 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc/revision/210d026faf9955653af8916fad021475a3f00453 "HTTP/1.1 200 OK" 2026-07-23:02:54:59,632 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc/tree/210d026faf9955653af8916fad021475a3f00453?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:54:59,866 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc/tree/210d026faf9955653af8916fad021475a3f00453/ARC-Challenge?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:55:00,410 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc "HTTP/1.1 200 OK" 2026-07-23:02:55:00,652 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc "HTTP/1.1 200 OK" 2026-07-23:02:55:00,907 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc/tree/210d026faf9955653af8916fad021475a3f00453/ARC-Easy?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:55:01,048 WARNING [task.py:799] [Task: boolq] metric acc is defined, but aggregation is not. using default aggregation=mean 2026-07-23:02:55:01,048 WARNING [task.py:811] [Task: boolq] metric acc is defined, but higher_is_better is not. using default higher_is_better=True 2026-07-23:02:55:01,131 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue "HTTP/1.1 200 OK" 2026-07-23:02:55:01,372 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue "HTTP/1.1 200 OK" 2026-07-23:02:55:01,433 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/revision/3de24cf8022e94f4ee4b9d55a6f539891524d646 "HTTP/1.1 200 OK" 2026-07-23:02:55:01,496 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/tree/3de24cf8022e94f4ee4b9d55a6f539891524d646?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:55:01,753 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/tree/3de24cf8022e94f4ee4b9d55a6f539891524d646/axb?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:55:01,880 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/revision/3de24cf8022e94f4ee4b9d55a6f539891524d646 "HTTP/1.1 200 OK" 2026-07-23:02:55:01,954 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/tree/3de24cf8022e94f4ee4b9d55a6f539891524d646/boolq?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:55:02,194 WARNING [task.py:799] [Task: copa] metric acc is defined, but aggregation is not. using default aggregation=mean 2026-07-23:02:55:02,194 WARNING [task.py:811] [Task: copa] metric acc is defined, but higher_is_better is not. using default higher_is_better=True 2026-07-23:02:55:02,276 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue "HTTP/1.1 200 OK" 2026-07-23:02:55:02,506 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue "HTTP/1.1 200 OK" 2026-07-23:02:55:02,784 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/tree/3de24cf8022e94f4ee4b9d55a6f539891524d646/copa?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:55:02,930 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag "HTTP/1.1 200 OK" 2026-07-23:02:55:03,183 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag "HTTP/1.1 200 OK" 2026-07-23:02:55:03,246 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag/revision/218ec52e09a7e7462a5400043bb9a69a41d06b76 "HTTP/1.1 200 OK" 2026-07-23:02:55:03,305 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag/tree/218ec52e09a7e7462a5400043bb9a69a41d06b76?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:55:03,529 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag/tree/218ec52e09a7e7462a5400043bb9a69a41d06b76/data?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:55:03,608 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag/revision/218ec52e09a7e7462a5400043bb9a69a41d06b76 "HTTP/1.1 200 OK" 2026-07-23:02:55:05,910 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai "HTTP/1.1 200 OK" 2026-07-23:02:55:06,138 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai "HTTP/1.1 200 OK" 2026-07-23:02:55:06,201 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/revision/900124bf3b8235c6daf21033af9948b3f07346c4 "HTTP/1.1 200 OK" 2026-07-23:02:55:06,260 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/tree/900124bf3b8235c6daf21033af9948b3f07346c4?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:55:06,476 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/tree/900124bf3b8235c6daf21033af9948b3f07346c4/default%2Ftest?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:55:06,541 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/tree/900124bf3b8235c6daf21033af9948b3f07346c4/default?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:55:06,601 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/revision/900124bf3b8235c6daf21033af9948b3f07346c4 "HTTP/1.1 200 OK" 2026-07-23:02:55:06,620 WARNING [task.py:325] [Task: lambada_openai] has_training_docs and has_validation_docs are False, using test_docs as fewshot_docs but this is not recommended. 2026-07-23:02:55:06,620 WARNING [task.py:325] [Task: lambada_openai] has_training_docs and has_validation_docs are False, using test_docs as fewshot_docs but this is not recommended. 2026-07-23:02:55:06,782 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa "HTTP/1.1 200 OK" 2026-07-23:02:55:07,020 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa "HTTP/1.1 200 OK" 2026-07-23:02:55:07,093 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/revision/388097ea7776314e93a529163e0fea805b8a6454 "HTTP/1.1 200 OK" 2026-07-23:02:55:07,153 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/tree/388097ea7776314e93a529163e0fea805b8a6454?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:55:07,366 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/tree/388097ea7776314e93a529163e0fea805b8a6454/additional?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:55:07,444 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/revision/388097ea7776314e93a529163e0fea805b8a6454 "HTTP/1.1 200 OK" 2026-07-23:02:55:07,525 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/tree/388097ea7776314e93a529163e0fea805b8a6454/main?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:55:07,781 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa "HTTP/1.1 200 OK" 2026-07-23:02:55:07,839 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa "HTTP/1.1 200 OK" 2026-07-23:02:55:07,942 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa/tree/main?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:55:08,050 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/datasets/ybisk/piqa/resolve/main/piqa.py "HTTP/1.1 307 Temporary Redirect" 2026-07-23:02:55:08,062 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/resolve-cache/datasets/ybisk/piqa/2e8ac2dffd59bac8c3c6714948f4c551a0848bb0/piqa.py?%2Fdatasets%2Fybisk%2Fpiqa%2Fresolve%2Fmain%2Fpiqa.py=&etag=%220f0d96172f9bf7aa10069374b3236ef083cfeb8e%22 "HTTP/1.1 200 OK" 2026-07-23:02:55:08,633 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa/revision/main "HTTP/1.1 200 OK" 2026-07-23:02:55:08,697 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa/tree/main?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:55:09,132 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq "HTTP/1.1 200 OK" 2026-07-23:02:55:09,363 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq "HTTP/1.1 200 OK" 2026-07-23:02:55:09,424 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq/revision/2c94ad3e1aafab77146f384e23536f97a4849815 "HTTP/1.1 200 OK" 2026-07-23:02:55:09,565 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq/tree/2c94ad3e1aafab77146f384e23536f97a4849815?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:55:09,769 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq/tree/2c94ad3e1aafab77146f384e23536f97a4849815/data?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:55:09,830 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq/revision/2c94ad3e1aafab77146f384e23536f97a4849815 "HTTP/1.1 200 OK" 2026-07-23:02:55:10,237 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande "HTTP/1.1 200 OK" 2026-07-23:02:55:10,460 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande "HTTP/1.1 200 OK" 2026-07-23:02:55:10,521 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/revision/01e74176c63542e6b0bcb004dcdea22d94fb67b5 "HTTP/1.1 200 OK" 2026-07-23:02:55:10,581 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/tree/01e74176c63542e6b0bcb004dcdea22d94fb67b5?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:55:10,888 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/tree/01e74176c63542e6b0bcb004dcdea22d94fb67b5/winogrande_debiased?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:55:10,949 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/revision/01e74176c63542e6b0bcb004dcdea22d94fb67b5 "HTTP/1.1 200 OK" 2026-07-23:02:55:11,031 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/tree/01e74176c63542e6b0bcb004dcdea22d94fb67b5/winogrande_xl?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:02:55:11,661 WARNING [evaluator.py:270] Overwriting default num_fewshot of winogrande from None to 0 2026-07-23:02:55:11,661 WARNING [evaluator.py:270] Overwriting default num_fewshot of sciq from None to 0 2026-07-23:02:55:11,661 WARNING [evaluator.py:270] Overwriting default num_fewshot of piqa from None to 0 2026-07-23:02:55:11,661 WARNING [evaluator.py:270] Overwriting default num_fewshot of openbookqa from None to 0 2026-07-23:02:55:11,661 WARNING [evaluator.py:270] Overwriting default num_fewshot of lambada_openai from None to 0 2026-07-23:02:55:11,661 WARNING [evaluator.py:270] Overwriting default num_fewshot of hellaswag from None to 0 2026-07-23:02:55:11,662 WARNING [evaluator.py:270] Overwriting default num_fewshot of copa from None to 0 2026-07-23:02:55:11,662 WARNING [evaluator.py:270] Overwriting default num_fewshot of boolq from None to 0 2026-07-23:02:55:11,662 WARNING [evaluator.py:270] Overwriting default num_fewshot of arc_easy from None to 0 2026-07-23:02:55:11,662 WARNING [evaluator.py:270] Overwriting default num_fewshot of arc_challenge from None to 0 2026-07-23:02:55:11,696 INFO [task.py:415] Building contexts for winogrande on rank 0... 0%| | 0/1267 [00:00 /workspace/FLAME-MoE/results/ablations/flame1e18_downstream.csv ==================================================================== [run] flame38m_g1_temporal mode=temporal TEMPORAL=1 H=256 heads=16 ffn=1368 experts=64 topk=6 moe_ffn=176 shared=352 [run] tasks=arc_challenge,arc_easy,hellaswag,openbookqa,piqa,winogrande,boolq,copa,lambada_openai,sciq out=/workspace/FLAME-MoE/results/phase0/runs/flame38m_g1_temporal/lmeval_1e18_0shot ==================================================================== /usr/local/lib/python3.11/dist-packages/torch/cuda/__init__.py:58: FutureWarning: The pynvml package is deprecated. Please install nvidia-ml-py instead. If you did not install pynvml directly, please report this to the maintainers of the package that installed pynvml for you. import pynvml # type: ignore[import] [run_lmeval] stubbed VLM backends (hf_vlms, vllm_vlms) /usr/local/lib/python3.11/dist-packages/torch/cuda/__init__.py:58: FutureWarning: The pynvml package is deprecated. Please install nvidia-ml-py instead. If you did not install pynvml directly, please report this to the maintainers of the package that installed pynvml for you. import pynvml # type: ignore[import] /usr/local/lib/python3.11/dist-packages/torch/cuda/__init__.py:58: FutureWarning: The pynvml package is deprecated. Please install nvidia-ml-py instead. If you did not install pynvml directly, please report this to the maintainers of the package that installed pynvml for you. import pynvml # type: ignore[import] /workspace/FLAME-MoE/.venv/lib/python3.11/site-packages/transformer_engine/pytorch/attention.py:108: UserWarning: To use flash-attn v3, please use the following commands to install: (1) pip install "git+https://github.com/Dao-AILab/flash-attention.git#egg=flashattn-hopper&subdirectory=hopper" (2) python_path=`python -c "import site; print(site.getsitepackages()[0])"` (3) mkdir -p $python_path/flashattn_hopper (4) wget -P $python_path/flashattn_hopper https://raw.githubusercontent.com/Dao-AILab/flash-attention/main/hopper/flash_attn_interface.py warnings.warn( [temporal] rolling-residency router installed (evict=min_logit) [run_lmeval] installed temporal residency router (native regime) /usr/local/lib/python3.11/dist-packages/requests/__init__.py:113: RequestsDependencyWarning: urllib3 (2.2.3) or chardet (6.0.0.post1)/charset_normalizer (3.3.2) doesn't match a supported version! warnings.warn( 2026-07-23:03:02:41,889 INFO [__main__.py:279] Verbosity set to INFO 2026-07-23:03:03:12,718 INFO [__init__.py:459] The tag 'arc_ca' is already registered as a group, this tag will not be registered. This may affect tasks you want to call. 2026-07-23:03:03:12,732 INFO [__init__.py:459] The tag 'arc_ca' is already registered as a group, this tag will not be registered. This may affect tasks you want to call. 2026-07-23:03:03:29,874 INFO [__main__.py:376] Selected Tasks: ['arc_challenge', 'arc_easy', 'boolq', 'copa', 'hellaswag', 'lambada_openai', 'openbookqa', 'piqa', 'sciq', 'winogrande'] 2026-07-23:03:03:29,903 INFO [evaluator.py:164] Setting random seed to 42 | Setting numpy seed to 42 | Setting torch manual seed to 42 | Setting fewshot manual seed to 42 2026-07-23:03:03:29,904 INFO [evaluator.py:201] Initializing megatron_lm model, with arguments: {} using world size: 1, data-parallel size: 1, context-parallel size: 1, hierarchical context-parallel sizes: Nonetensor-model-parallel size: 1, encoder-tensor-model-parallel size: 0, pipeline-model-parallel size: 1, encoder-pipeline-model-parallel size: 0 WARNING: overriding default arguments for exit_on_missing_checkpoint:True with exit_on_missing_checkpoint:False setting global batch size to 32 Number of virtual stages per pipeline stage: None accumulate and all-reduce gradients in fp32 for bfloat16 data type. using torch.bfloat16 for parameters ... ------------------------ arguments ------------------------ account_for_embedding_in_pipeline_split ......... False account_for_loss_in_pipeline_split .............. False accumulate_allreduce_grads_in_fp32 .............. True adam_beta1 ...................................... 0.9 adam_beta2 ...................................... 0.999 adam_eps ........................................ 1e-08 add_bias_linear ................................. False add_position_embedding .......................... True add_qkv_bias .................................... False adlr_autoresume ................................. False adlr_autoresume_interval ........................ 1000 align_grad_reduce ............................... True align_param_gather .............................. False app_tag_run_name ................................ None app_tag_run_version ............................. 0.0.0 apply_layernorm_1p .............................. False apply_query_key_layer_scaling ................... False apply_residual_connection_post_layernorm ........ False apply_rope_fusion ............................... True async_save ...................................... None async_tensor_model_parallel_allreduce ........... True attention_backend ............................... AttnBackend.auto attention_dropout ............................... 0.0 attention_softmax_in_fp32 ....................... False auto_detect_ckpt_format ......................... False barrier_with_L1_time ............................ True bert_binary_head ................................ True bert_embedder_type .............................. megatron bert_load ....................................... None bf16 ............................................ True bias_dropout_fusion ............................. True bias_gelu_fusion ................................ False bias_swiglu_fusion .............................. True biencoder_projection_dim ........................ 0 biencoder_shared_query_context_model ............ False block_data_path ................................. None calc_ft_timeouts ................................ False calculate_per_token_loss ........................ False check_for_large_grads ........................... False check_for_nan_in_loss_and_grad .................. True check_for_spiky_loss ............................ False check_weight_hash_across_dp_replicas_interval ... None ckpt_assume_constant_structure .................. False ckpt_convert_format ............................. None ckpt_convert_save ............................... None ckpt_convert_update_legacy_dist_opt_format ...... False ckpt_format ..................................... torch_dist ckpt_fully_parallel_load ........................ False ckpt_fully_parallel_save ........................ True ckpt_fully_parallel_save_deprecated ............. False ckpt_step ....................................... None classes_fraction ................................ 1.0 clip_grad ....................................... 1.0 clone_scatter_output_in_embedding ............... True config_logger_dir ............................... consumed_train_samples .......................... 0 consumed_valid_samples .......................... 0 context_parallel_size ........................... 1 cp_comm_type .................................... ['p2p'] create_attention_mask_in_dataloader ............. True cross_entropy_loss_fusion ....................... False cuda_graph_warmup_steps ......................... 3 data_args_path .................................. None data_cache_path ................................. None data_parallel_random_init ....................... False data_parallel_sharding_strategy ................. no_shard data_parallel_size .............................. 1 data_path ....................................... None data_per_class_fraction ......................... 1.0 data_sharding ................................... True dataloader_type ................................. single ddp_average_in_collective ....................... False ddp_bucket_size ................................. None ddp_num_buckets ................................. None ddp_pad_buckets_for_high_nccl_busbw ............. False decoder_first_pipeline_num_layers ............... None decoder_last_pipeline_num_layers ................ None decoder_num_layers .............................. None decoder_seq_length .............................. None decoupled_lr .................................... None decoupled_min_lr ................................ None decrease_batch_size_if_needed ................... False defer_embedding_wgrad_compute ................... False deprecated_use_mcore_models ..................... False deterministic_mode .............................. False dino_bottleneck_size ............................ 256 dino_freeze_last_layer .......................... 1 dino_head_hidden_size ........................... 2048 dino_local_crops_number ......................... 10 dino_local_img_size ............................. 96 dino_norm_last_layer ............................ False dino_teacher_temp ............................... 0.07 dino_warmup_teacher_temp ........................ 0.04 dino_warmup_teacher_temp_epochs ................. 30 disable_straggler_on_startup .................... False dist_ckpt_format_deprecated ..................... None dist_ckpt_strictness ............................ assume_ok_unexpected distribute_saved_activations .................... False distributed_backend ............................. nccl distributed_timeout_minutes ..................... 10 embedding_path .................................. None empty_unused_memory_level ....................... 0 enable_cuda_graph ............................... False enable_ft_package ............................... False enable_gloo_process_groups ...................... True enable_one_logger ............................... True encoder_num_layers .............................. 9 encoder_pipeline_model_parallel_size ............ 0 encoder_seq_length .............................. 2048 encoder_tensor_model_parallel_size .............. 0 end_weight_decay ................................ 0.01 eod_mask_loss ................................... False error_injection_rate ............................ 0 error_injection_type ............................ transient_error eval_interval ................................... 1000 eval_iters ...................................... 100 evidence_data_path .............................. None exit_duration_in_mins ........................... None exit_interval ................................... None exit_on_missing_checkpoint ...................... False exit_signal_handler ............................. False exp_avg_dtype ................................... torch.float32 exp_avg_sq_dtype ................................ torch.float32 expert_model_parallel_size ...................... 1 expert_tensor_parallel_size ..................... 1 ffn_hidden_size ................................. 1368 finetune ........................................ False flash_decode .................................... False fp16 ............................................ False fp16_lm_cross_entropy ........................... False fp32_residual_connection ........................ False fp8 ............................................. None fp8_amax_compute_algo ........................... most_recent fp8_amax_history_len ............................ 1 fp8_interval .................................... 1 fp8_margin ...................................... 0 fp8_param_gather ................................ False fp8_wgrad ....................................... True global_batch_size ............................... 32 grad_reduce_in_bf16 ............................. False gradient_accumulation_fusion .................... False gradient_reduce_div_fusion ...................... True group_query_attention ........................... False head_lr_mult .................................... 1.0 hidden_dropout .................................. 0.0 hidden_size ..................................... 256 hierarchical_context_parallel_sizes ............. None hybrid_attention_ratio .......................... 0.0 hybrid_mlp_ratio ................................ 0.0 hybrid_override_pattern ......................... None hysteresis ...................................... 2 ict_head_size ................................... None ict_load ........................................ None img_h ........................................... 224 img_w ........................................... 224 indexer_batch_size .............................. 128 indexer_log_interval ............................ 1000 inference_batch_times_seqlen_threshold .......... -1 inference_max_batch_size ........................ 8 inference_max_seq_length ........................ 2560 inference_rng_tracker ........................... False init_method_std ................................. 0.02 init_method_xavier_uniform ...................... False init_model_with_meta_device ..................... False initial_loss_scale .............................. 4294967296 iter_per_epoch .................................. 1250 iterations_to_skip .............................. [] keep_fp8_transpose_cache_when_using_custom_fsdp . False kv_channels ..................................... 16 kv_lora_rank .................................... 32 lazy_mpu_init ................................... None load ............................................ /workspace/FLAME-MoE/results/phase0/runs/flame38m_g1_temporal/ckpt local_rank ...................................... 0 log_interval .................................... 100 log_loss_scale_to_tensorboard ................... True log_memory_to_tensorboard ....................... False log_num_zeros_in_grad ........................... False log_params_norm ................................. False log_progress .................................... False log_straggler ................................... False log_throughput .................................. False log_timers_to_tensorboard ....................... False log_validation_ppl_to_tensorboard ............... False log_world_size_to_tensorboard ................... False logging_level ................................... None loss_scale ...................................... None loss_scale_window ............................... 1000 lr .............................................. None lr_decay_iters .................................. None lr_decay_samples ................................ None lr_decay_style .................................. linear lr_warmup_fraction .............................. None lr_warmup_init .................................. 0.0 lr_warmup_iters ................................. 0 lr_warmup_samples ............................... 0 lr_wsd_decay_iters .............................. None lr_wsd_decay_samples ............................ None lr_wsd_decay_style .............................. exponential main_grads_dtype ................................ torch.float32 main_params_dtype ............................... torch.float32 make_vocab_size_divisible_by .................... 128 manual_gc ....................................... False manual_gc_eval .................................. True manual_gc_interval .............................. 0 mask_factor ..................................... 1.0 mask_prob ....................................... 0.15 mask_type ....................................... random masked_softmax_fusion ........................... True max_position_embeddings ......................... 2048 max_tokens_to_oom ............................... 10000000 memory_snapshot_path ............................ snapshot.pickle merge_file ...................................... None micro_batch_size ................................ 32 microbatch_group_size_per_vp_stage .............. None min_loss_scale .................................. 1.0 min_lr .......................................... 0.0 mmap_bin_files .................................. True mock_data ....................................... False moe_aux_loss_coeff .............................. 0.01 moe_enable_deepep ............................... False moe_expert_capacity_factor ...................... None moe_extended_tp ................................. False moe_ffn_hidden_size ............................. 176 moe_grouped_gemm ................................ True moe_input_jitter_eps ............................ None moe_layer_freq .................................. [0, 1, 1, 1, 1, 1, 1, 1, 1] moe_layer_recompute ............................. False moe_pad_expert_input_to_capacity ................ False moe_per_layer_logging ........................... False moe_permute_fusion .............................. False moe_router_bias_update_rate ..................... 0.001 moe_router_dtype ................................ fp32 moe_router_enable_expert_bias ................... False moe_router_group_topk ........................... None moe_router_load_balancing_type .................. aux_loss moe_router_num_groups ........................... None moe_router_pre_softmax .......................... True moe_router_score_function ....................... softmax moe_router_topk ................................. 6 moe_router_topk_scaling_factor .................. None moe_shared_expert_intermediate_size ............. 352 moe_shared_expert_overlap ....................... False moe_token_dispatcher_type ....................... alltoall moe_token_drop_policy ........................... probs moe_use_legacy_grouped_gemm ..................... False moe_use_upcycling ............................... False moe_z_loss_coeff ................................ 0.001 mscale .......................................... 1.0 mscale_all_dim .................................. 1.0 multi_latent_attention .......................... False nccl_communicator_config_path ................... None no_load_optim ................................... True no_load_rng ..................................... True no_persist_layer_norm ........................... False no_save_optim ................................... None no_save_rng ..................................... None non_persistent_ckpt_type ........................ None non_persistent_global_ckpt_dir .................. None non_persistent_local_ckpt_algo .................. fully_parallel non_persistent_local_ckpt_dir ................... None non_persistent_save_interval .................... None norm_epsilon .................................... 1e-06 normalization ................................... RMSNorm num_attention_heads ............................. 16 num_channels .................................... 3 num_classes ..................................... 1000 num_dataset_builder_threads ..................... 1 num_distributed_optimizer_instances ............. 1 num_experts ..................................... 64 num_layers ...................................... 9 num_layers_per_virtual_pipeline_stage ........... None num_query_groups ................................ 1 num_virtual_stages_per_pipeline_rank ............ None num_workers ..................................... 2 one_logger_async ................................ False one_logger_project .............................. megatron-lm one_logger_run_name ............................. None onnx_safe ....................................... None openai_gelu ..................................... False optimizer ....................................... adam optimizer_cpu_offload ........................... False optimizer_offload_fraction ...................... 1.0 output_bert_embeddings .......................... False overlap_cpu_optimizer_d2h_h2d ................... False overlap_grad_reduce ............................. False overlap_p2p_comm ................................ False overlap_p2p_comm_warmup_flush ................... False overlap_param_gather ............................ False overlap_param_gather_with_optimizer_step ........ False override_opt_param_scheduler .................... False params_dtype .................................... torch.bfloat16 patch_dim ....................................... 16 per_split_data_args_path ........................ None perform_initialization .......................... True pin_cpu_grads ................................... True pin_cpu_params .................................. True pipeline_model_parallel_comm_backend ............ None pipeline_model_parallel_size .................... 1 pipeline_model_parallel_split_rank .............. None position_embedding_type ......................... rope pretrained_checkpoint ........................... None profile ......................................... False profile_ranks ................................... [0] profile_step_end ................................ 12 profile_step_start .............................. 10 q_lora_rank ..................................... None qk_head_dim ..................................... 128 qk_layernorm .................................... False qk_pos_emb_head_dim ............................. 64 query_in_block_prob ............................. 0.1 rampup_batch_size ............................... None rank ............................................ 0 recompute_granularity ........................... None recompute_method ................................ None recompute_num_layers ............................ None record_memory_history ........................... False relative_attention_max_distance ................. 128 relative_attention_num_buckets .................. 32 replication ..................................... False replication_factor .............................. 2 replication_jump ................................ None rerun_mode ...................................... disabled reset_attention_mask ............................ False reset_position_ids .............................. False result_rejected_tracker_filename ................ None retriever_report_topk_accuracies ................ [] retriever_score_scaling ......................... False retriever_seq_length ............................ 256 retro_add_retriever ............................. False retro_attention_gate ............................ 1 retro_cyclic_train_iters ........................ None retro_encoder_attention_dropout ................. 0.1 retro_encoder_hidden_dropout .................... 0.1 retro_encoder_layers ............................ 2 retro_num_neighbors ............................. 2 retro_num_retrieved_chunks ...................... 2 retro_project_dir ............................... None retro_verify_neighbor_count ..................... True rope_scaling_factor ............................. 8.0 rotary_base ..................................... 10000 rotary_interleaved .............................. False rotary_percent .................................. 1.0 rotary_scaling_factor ........................... 1.0 rotary_seq_len_interpolation_factor ............. None s3_cache_path ................................... None sample_rate ..................................... 1.0 save ............................................ None save_interval ................................... None scatter_gather_tensors_in_pipeline .............. True seed ............................................ 42 seq_length ...................................... 2048 sequence_parallel ............................... False sgd_momentum .................................... 0.9 short_seq_prob .................................. 0.1 skip_train ...................................... False skipped_train_samples ........................... 0 spec ............................................ None split ........................................... None squared_relu .................................... False start_weight_decay .............................. 0.01 straggler_ctrlr_port ............................ 65535 straggler_minmax_count .......................... 1 suggested_communication_unit_size ............... 400000000 swiglu .......................................... True swin_backbone_type .............................. tiny te_rng_tracker .................................. False tensor_model_parallel_size ...................... 1 tensorboard_dir ................................. None tensorboard_log_interval ........................ 1 tensorboard_queue_size .......................... 1000 test_data_path .................................. None test_mode ....................................... False tiktoken_num_special_tokens ..................... 1000 tiktoken_pattern ................................ None tiktoken_special_tokens ......................... None timing_log_level ................................ 0 timing_log_option ............................... minmax titles_data_path ................................ None tokenizer_model ................................. EleutherAI/pythia-12b tokenizer_type .................................. HuggingFaceTokenizer tp_comm_bootstrap_backend ....................... nccl tp_comm_bulk_dgrad .............................. True tp_comm_bulk_wgrad .............................. True tp_comm_overlap ................................. False tp_comm_overlap_ag .............................. True tp_comm_overlap_cfg ............................. None tp_comm_overlap_rs .............................. True tp_comm_overlap_rs_dgrad ........................ False tp_comm_split_ag ................................ True tp_comm_split_rs ................................ True train_data_path ................................. None train_iters ..................................... None train_samples ................................... None train_sync_interval ............................. None transformer_impl ................................ transformer_engine transformer_pipeline_model_parallel_size ........ 1 untie_embeddings_and_output_weights ............. True use_checkpoint_args ............................. False use_checkpoint_opt_param_scheduler .............. False use_cpu_initialization .......................... None use_custom_fsdp ................................. False use_dist_ckpt ................................... True use_dist_ckpt_deprecated ........................ False use_distributed_optimizer ....................... True use_flash_attn .................................. False use_legacy_models ............................... False use_mp_args_from_checkpoint_args ................ False use_one_sent_docs ............................... False use_persistent_ckpt_worker ...................... False use_precision_aware_optimizer ................... False use_pytorch_profiler ............................ False use_ring_exchange_p2p ........................... False use_rope_scaling ................................ False use_rotary_position_embeddings .................. False use_tokenizer_model_from_checkpoint_args ........ True use_torch_fsdp2 ................................. False use_torch_optimizer_for_cpu_offload ............. False use_tp_pp_dp_mapping ............................ False v_head_dim ...................................... 128 valid_data_path ................................. None variable_seq_lengths ............................ False virtual_pipeline_model_parallel_size ............ None vision_backbone_type ............................ vit vision_pretraining .............................. False vision_pretraining_type ......................... classify vocab_extra_ids ................................. 0 vocab_file ...................................... None vocab_size ...................................... None wandb_exp_name .................................. wandb_project ................................... wandb_save_dir .................................. weight_decay .................................... 0.01 weight_decay_incr_style ......................... constant wgrad_deferral_limit ............................ 0 world_size ...................................... 1 yaml_cfg ........................................ None -------------------- end of arguments --------------------- 2026-07-23:03:03:29,921 INFO [num_microbatches_calculator.py:228] setting number of microbatches to constant 1 > building HuggingFaceTokenizer tokenizer ... 2026-07-23:03:03:30,125 INFO [_client.py:1038] HTTP Request: HEAD https://huggingface.co/EleutherAI/pythia-12b/resolve/main/config.json "HTTP/1.1 307 Temporary Redirect" 2026-07-23:03:03:30,137 INFO [_client.py:1038] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/EleutherAI/pythia-12b/bb1e3e710cdf6b524461d543cfb5ba773f0a81b6/config.json "HTTP/1.1 200 OK" 2026-07-23:03:03:30,265 INFO [_client.py:1038] HTTP Request: HEAD https://huggingface.co/EleutherAI/pythia-12b/resolve/main/tokenizer_config.json "HTTP/1.1 307 Temporary Redirect" 2026-07-23:03:03:30,277 INFO [_client.py:1038] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/EleutherAI/pythia-12b/bb1e3e710cdf6b524461d543cfb5ba773f0a81b6/tokenizer_config.json "HTTP/1.1 200 OK" 2026-07-23:03:03:30,387 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/models/EleutherAI/pythia-12b/tree/main/additional_chat_templates?recursive=false&expand=false "HTTP/1.1 404 Not Found" 2026-07-23:03:03:30,461 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/models/EleutherAI/pythia-12b/tree/main?recursive=true&expand=false "HTTP/1.1 200 OK" > padded vocab (size: 50277) with 27 dummy tokens (new size: 50304) WARNING: one_logger package is required to enable e2e metrics tracking. please go to https://confluence.nvidia.com/display/MLWFO/Package+Repositories for details to install it 2026-07-23:03:03:30,656 WARNING [rerun_state_machine.py:239] RerunStateMachine initialized in mode RerunMode.DISABLED > initializing torch distributed ... [W723 03:03:30.173472448 CUDAAllocatorConfig.h:28] Warning: expandable_segments not supported on this platform (function operator()) > initialized tensor model parallel with size 1 > initialized pipeline model parallel with size 1 > setting random seeds to 42 ... > compiling dataset index builder ... make: Entering directory '/workspace/FLAME-MoE/Megatron-LM/megatron/core/datasets' make: Nothing to be done for 'default'. make: Leaving directory '/workspace/FLAME-MoE/Megatron-LM/megatron/core/datasets' >>> done with dataset index builder. Compilation time: 0.695 seconds > compiling and loading fused kernels ... >>> done with compiling and loading fused kernels. Compilation time: 0.342 seconds building GPT model ... > number of parameters on (tensor, pipeline) model parallel rank (0, 0): 100670208 /workspace/FLAME-MoE/Megatron-LM/megatron/core/transformer/transformer_layer.py:339: UserWarning: TransformerLayer._get_layer_offset is deprecated.Please use get_transformer_layer_offset instead. warnings.warn( loading distributed checkpoint from /workspace/FLAME-MoE/results/phase0/runs/flame38m_g1_temporal/ckpt at iteration 2121 /workspace/FLAME-MoE/Megatron-LM/megatron/core/dist_checkpointing/strategies/torch.py:847: FutureWarning: `load_state_dict` is deprecated and will be removed in future versions. Please use `load` instead. checkpoint.load_state_dict( /usr/local/lib/python3.11/dist-packages/torch/distributed/checkpoint/planner_helpers.py:306: FutureWarning: Please use DTensor instead and we are deprecating ShardedTensor. device = getattr(value, "device", None) /workspace/FLAME-MoE/.venv/lib/python3.11/site-packages/transformer_engine/pytorch/module/base.py:566: FutureWarning: You are using `torch.load` with `weights_only=False` (the current default value), which uses the default pickle module implicitly. It is possible to construct malicious pickle data which will execute arbitrary code during unpickling (See https://github.com/pytorch/pytorch/blob/main/SECURITY.md#untrusted-models for more details). In a future release, the default value for `weights_only` will be flipped to `True`. This limits the functions that could be executed during unpickling. Arbitrary objects will no longer be allowed to be loaded via this mode unless they are explicitly allowlisted by the user via `torch.serialization.add_safe_globals`. We recommend you start setting `weights_only=True` for any use case where you don't have full control of the loaded file. Please open an issue on GitHub for any issues related to this experimental feature. state = torch.load(state, map_location="cuda") checkpoint version 3.0 2026-07-23:03:03:35,539 WARNING [rerun_state_machine.py:811] RerunStateMachine disabled via CLI, ignoring machine state saved in checkpoint successfully loaded checkpoint from /workspace/FLAME-MoE/results/phase0/runs/flame38m_g1_temporal/ckpt [ t 1/1, p 1/1 ] at iteration 2121 2026-07-23:03:03:35,678 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc "HTTP/1.1 200 OK" 2026-07-23:03:03:35,917 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc "HTTP/1.1 200 OK" 2026-07-23:03:03:35,983 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc/revision/210d026faf9955653af8916fad021475a3f00453 "HTTP/1.1 200 OK" 2026-07-23:03:03:36,054 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc/tree/210d026faf9955653af8916fad021475a3f00453?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:03:36,354 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc/tree/210d026faf9955653af8916fad021475a3f00453/ARC-Challenge?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:03:36,903 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc "HTTP/1.1 200 OK" 2026-07-23:03:03:37,186 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc "HTTP/1.1 200 OK" 2026-07-23:03:03:37,490 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc/tree/210d026faf9955653af8916fad021475a3f00453/ARC-Easy?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:03:37,632 WARNING [task.py:799] [Task: boolq] metric acc is defined, but aggregation is not. using default aggregation=mean 2026-07-23:03:03:37,632 WARNING [task.py:811] [Task: boolq] metric acc is defined, but higher_is_better is not. using default higher_is_better=True 2026-07-23:03:03:37,690 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue "HTTP/1.1 200 OK" 2026-07-23:03:03:37,913 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue "HTTP/1.1 200 OK" 2026-07-23:03:03:37,988 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/revision/3de24cf8022e94f4ee4b9d55a6f539891524d646 "HTTP/1.1 200 OK" 2026-07-23:03:03:38,048 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/tree/3de24cf8022e94f4ee4b9d55a6f539891524d646?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:03:38,407 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/tree/3de24cf8022e94f4ee4b9d55a6f539891524d646/axb?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:03:38,471 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/revision/3de24cf8022e94f4ee4b9d55a6f539891524d646 "HTTP/1.1 200 OK" 2026-07-23:03:03:38,541 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/tree/3de24cf8022e94f4ee4b9d55a6f539891524d646/boolq?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:03:38,792 WARNING [task.py:799] [Task: copa] metric acc is defined, but aggregation is not. using default aggregation=mean 2026-07-23:03:03:38,792 WARNING [task.py:811] [Task: copa] metric acc is defined, but higher_is_better is not. using default higher_is_better=True 2026-07-23:03:03:38,852 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue "HTTP/1.1 200 OK" 2026-07-23:03:03:39,102 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue "HTTP/1.1 200 OK" 2026-07-23:03:03:39,472 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/tree/3de24cf8022e94f4ee4b9d55a6f539891524d646/copa?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:03:39,615 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag "HTTP/1.1 200 OK" 2026-07-23:03:03:39,844 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag "HTTP/1.1 200 OK" 2026-07-23:03:03:39,905 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag/revision/218ec52e09a7e7462a5400043bb9a69a41d06b76 "HTTP/1.1 200 OK" 2026-07-23:03:03:39,979 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag/tree/218ec52e09a7e7462a5400043bb9a69a41d06b76?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:03:40,189 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag/tree/218ec52e09a7e7462a5400043bb9a69a41d06b76/data?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:03:40,256 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag/revision/218ec52e09a7e7462a5400043bb9a69a41d06b76 "HTTP/1.1 200 OK" 2026-07-23:03:03:42,544 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai "HTTP/1.1 200 OK" 2026-07-23:03:03:42,777 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai "HTTP/1.1 200 OK" 2026-07-23:03:03:42,840 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/revision/900124bf3b8235c6daf21033af9948b3f07346c4 "HTTP/1.1 200 OK" 2026-07-23:03:03:42,901 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/tree/900124bf3b8235c6daf21033af9948b3f07346c4?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:03:43,144 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/tree/900124bf3b8235c6daf21033af9948b3f07346c4/default%2Ftest?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:03:43,204 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/tree/900124bf3b8235c6daf21033af9948b3f07346c4/default?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:03:43,271 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/revision/900124bf3b8235c6daf21033af9948b3f07346c4 "HTTP/1.1 200 OK" 2026-07-23:03:03:43,289 WARNING [task.py:325] [Task: lambada_openai] has_training_docs and has_validation_docs are False, using test_docs as fewshot_docs but this is not recommended. 2026-07-23:03:03:43,289 WARNING [task.py:325] [Task: lambada_openai] has_training_docs and has_validation_docs are False, using test_docs as fewshot_docs but this is not recommended. 2026-07-23:03:03:43,451 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa "HTTP/1.1 200 OK" 2026-07-23:03:03:43,764 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa "HTTP/1.1 200 OK" 2026-07-23:03:03:43,825 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/revision/388097ea7776314e93a529163e0fea805b8a6454 "HTTP/1.1 200 OK" 2026-07-23:03:03:43,901 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/tree/388097ea7776314e93a529163e0fea805b8a6454?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:03:44,128 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/tree/388097ea7776314e93a529163e0fea805b8a6454/additional?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:03:44,194 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/revision/388097ea7776314e93a529163e0fea805b8a6454 "HTTP/1.1 200 OK" 2026-07-23:03:03:44,287 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/tree/388097ea7776314e93a529163e0fea805b8a6454/main?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:03:44,573 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa "HTTP/1.1 200 OK" 2026-07-23:03:03:44,634 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa "HTTP/1.1 200 OK" 2026-07-23:03:03:44,741 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa/tree/main?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:03:44,812 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/datasets/ybisk/piqa/resolve/main/piqa.py "HTTP/1.1 307 Temporary Redirect" 2026-07-23:03:03:44,824 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/resolve-cache/datasets/ybisk/piqa/2e8ac2dffd59bac8c3c6714948f4c551a0848bb0/piqa.py?%2Fdatasets%2Fybisk%2Fpiqa%2Fresolve%2Fmain%2Fpiqa.py=&etag=%220f0d96172f9bf7aa10069374b3236ef083cfeb8e%22 "HTTP/1.1 200 OK" 2026-07-23:03:03:45,379 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa/revision/main "HTTP/1.1 200 OK" 2026-07-23:03:03:45,439 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa/tree/main?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:03:45,862 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq "HTTP/1.1 200 OK" 2026-07-23:03:03:46,083 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq "HTTP/1.1 200 OK" 2026-07-23:03:03:46,142 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq/revision/2c94ad3e1aafab77146f384e23536f97a4849815 "HTTP/1.1 200 OK" 2026-07-23:03:03:46,203 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq/tree/2c94ad3e1aafab77146f384e23536f97a4849815?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:03:46,412 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq/tree/2c94ad3e1aafab77146f384e23536f97a4849815/data?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:03:46,531 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq/revision/2c94ad3e1aafab77146f384e23536f97a4849815 "HTTP/1.1 200 OK" 2026-07-23:03:03:47,026 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande "HTTP/1.1 200 OK" 2026-07-23:03:03:47,265 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande "HTTP/1.1 200 OK" 2026-07-23:03:03:47,321 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/revision/01e74176c63542e6b0bcb004dcdea22d94fb67b5 "HTTP/1.1 200 OK" 2026-07-23:03:03:47,384 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/tree/01e74176c63542e6b0bcb004dcdea22d94fb67b5?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:03:47,635 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/tree/01e74176c63542e6b0bcb004dcdea22d94fb67b5/winogrande_debiased?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:03:47,695 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/revision/01e74176c63542e6b0bcb004dcdea22d94fb67b5 "HTTP/1.1 200 OK" 2026-07-23:03:03:47,786 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/tree/01e74176c63542e6b0bcb004dcdea22d94fb67b5/winogrande_xl?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:03:48,416 WARNING [evaluator.py:270] Overwriting default num_fewshot of winogrande from None to 0 2026-07-23:03:03:48,416 WARNING [evaluator.py:270] Overwriting default num_fewshot of sciq from None to 0 2026-07-23:03:03:48,416 WARNING [evaluator.py:270] Overwriting default num_fewshot of piqa from None to 0 2026-07-23:03:03:48,416 WARNING [evaluator.py:270] Overwriting default num_fewshot of openbookqa from None to 0 2026-07-23:03:03:48,416 WARNING [evaluator.py:270] Overwriting default num_fewshot of lambada_openai from None to 0 2026-07-23:03:03:48,416 WARNING [evaluator.py:270] Overwriting default num_fewshot of hellaswag from None to 0 2026-07-23:03:03:48,416 WARNING [evaluator.py:270] Overwriting default num_fewshot of copa from None to 0 2026-07-23:03:03:48,416 WARNING [evaluator.py:270] Overwriting default num_fewshot of boolq from None to 0 2026-07-23:03:03:48,416 WARNING [evaluator.py:270] Overwriting default num_fewshot of arc_easy from None to 0 2026-07-23:03:03:48,416 WARNING [evaluator.py:270] Overwriting default num_fewshot of arc_challenge from None to 0 2026-07-23:03:03:48,453 INFO [task.py:415] Building contexts for winogrande on rank 0... 0%| | 0/1267 [00:00 /workspace/FLAME-MoE/results/ablations/flame1e18_downstream.csv ==================================================================== [run] flame38m_g1_temporal_s2 mode=temporal TEMPORAL=1 H=256 heads=16 ffn=1368 experts=64 topk=6 moe_ffn=176 shared=352 [run] tasks=arc_challenge,arc_easy,hellaswag,openbookqa,piqa,winogrande,boolq,copa,lambada_openai,sciq out=/workspace/FLAME-MoE/results/phase0/runs/flame38m_g1_temporal_s2/lmeval_1e18_0shot ==================================================================== /usr/local/lib/python3.11/dist-packages/torch/cuda/__init__.py:58: FutureWarning: The pynvml package is deprecated. Please install nvidia-ml-py instead. If you did not install pynvml directly, please report this to the maintainers of the package that installed pynvml for you. import pynvml # type: ignore[import] [run_lmeval] stubbed VLM backends (hf_vlms, vllm_vlms) /usr/local/lib/python3.11/dist-packages/torch/cuda/__init__.py:58: FutureWarning: The pynvml package is deprecated. Please install nvidia-ml-py instead. If you did not install pynvml directly, please report this to the maintainers of the package that installed pynvml for you. import pynvml # type: ignore[import] /usr/local/lib/python3.11/dist-packages/torch/cuda/__init__.py:58: FutureWarning: The pynvml package is deprecated. Please install nvidia-ml-py instead. If you did not install pynvml directly, please report this to the maintainers of the package that installed pynvml for you. import pynvml # type: ignore[import] /workspace/FLAME-MoE/.venv/lib/python3.11/site-packages/transformer_engine/pytorch/attention.py:108: UserWarning: To use flash-attn v3, please use the following commands to install: (1) pip install "git+https://github.com/Dao-AILab/flash-attention.git#egg=flashattn-hopper&subdirectory=hopper" (2) python_path=`python -c "import site; print(site.getsitepackages()[0])"` (3) mkdir -p $python_path/flashattn_hopper (4) wget -P $python_path/flashattn_hopper https://raw.githubusercontent.com/Dao-AILab/flash-attention/main/hopper/flash_attn_interface.py warnings.warn( [temporal] rolling-residency router installed (evict=min_logit) [run_lmeval] installed temporal residency router (native regime) /usr/local/lib/python3.11/dist-packages/requests/__init__.py:113: RequestsDependencyWarning: urllib3 (2.2.3) or chardet (6.0.0.post1)/charset_normalizer (3.3.2) doesn't match a supported version! warnings.warn( 2026-07-23:03:14:22,863 INFO [__main__.py:279] Verbosity set to INFO 2026-07-23:03:15:04,024 INFO [__init__.py:459] The tag 'arc_ca' is already registered as a group, this tag will not be registered. This may affect tasks you want to call. 2026-07-23:03:15:04,037 INFO [__init__.py:459] The tag 'arc_ca' is already registered as a group, this tag will not be registered. This may affect tasks you want to call. 2026-07-23:03:15:21,387 INFO [__main__.py:376] Selected Tasks: ['arc_challenge', 'arc_easy', 'boolq', 'copa', 'hellaswag', 'lambada_openai', 'openbookqa', 'piqa', 'sciq', 'winogrande'] 2026-07-23:03:15:21,407 INFO [evaluator.py:164] Setting random seed to 42 | Setting numpy seed to 42 | Setting torch manual seed to 42 | Setting fewshot manual seed to 42 2026-07-23:03:15:21,407 INFO [evaluator.py:201] Initializing megatron_lm model, with arguments: {} using world size: 1, data-parallel size: 1, context-parallel size: 1, hierarchical context-parallel sizes: Nonetensor-model-parallel size: 1, encoder-tensor-model-parallel size: 0, pipeline-model-parallel size: 1, encoder-pipeline-model-parallel size: 0 WARNING: overriding default arguments for exit_on_missing_checkpoint:True with exit_on_missing_checkpoint:False setting global batch size to 32 Number of virtual stages per pipeline stage: None accumulate and all-reduce gradients in fp32 for bfloat16 data type. using torch.bfloat16 for parameters ... ------------------------ arguments ------------------------ account_for_embedding_in_pipeline_split ......... False account_for_loss_in_pipeline_split .............. False accumulate_allreduce_grads_in_fp32 .............. True adam_beta1 ...................................... 0.9 adam_beta2 ...................................... 0.999 adam_eps ........................................ 1e-08 add_bias_linear ................................. False add_position_embedding .......................... True add_qkv_bias .................................... False adlr_autoresume ................................. False adlr_autoresume_interval ........................ 1000 align_grad_reduce ............................... True align_param_gather .............................. False app_tag_run_name ................................ None app_tag_run_version ............................. 0.0.0 apply_layernorm_1p .............................. False apply_query_key_layer_scaling ................... False apply_residual_connection_post_layernorm ........ False apply_rope_fusion ............................... True async_save ...................................... None async_tensor_model_parallel_allreduce ........... True attention_backend ............................... AttnBackend.auto attention_dropout ............................... 0.0 attention_softmax_in_fp32 ....................... False auto_detect_ckpt_format ......................... False barrier_with_L1_time ............................ True bert_binary_head ................................ True bert_embedder_type .............................. megatron bert_load ....................................... None bf16 ............................................ True bias_dropout_fusion ............................. True bias_gelu_fusion ................................ False bias_swiglu_fusion .............................. True biencoder_projection_dim ........................ 0 biencoder_shared_query_context_model ............ False block_data_path ................................. None calc_ft_timeouts ................................ False calculate_per_token_loss ........................ False check_for_large_grads ........................... False check_for_nan_in_loss_and_grad .................. True check_for_spiky_loss ............................ False check_weight_hash_across_dp_replicas_interval ... None ckpt_assume_constant_structure .................. False ckpt_convert_format ............................. None ckpt_convert_save ............................... None ckpt_convert_update_legacy_dist_opt_format ...... False ckpt_format ..................................... torch_dist ckpt_fully_parallel_load ........................ False ckpt_fully_parallel_save ........................ True ckpt_fully_parallel_save_deprecated ............. False ckpt_step ....................................... None classes_fraction ................................ 1.0 clip_grad ....................................... 1.0 clone_scatter_output_in_embedding ............... True config_logger_dir ............................... consumed_train_samples .......................... 0 consumed_valid_samples .......................... 0 context_parallel_size ........................... 1 cp_comm_type .................................... ['p2p'] create_attention_mask_in_dataloader ............. True cross_entropy_loss_fusion ....................... False cuda_graph_warmup_steps ......................... 3 data_args_path .................................. None data_cache_path ................................. None data_parallel_random_init ....................... False data_parallel_sharding_strategy ................. no_shard data_parallel_size .............................. 1 data_path ....................................... None data_per_class_fraction ......................... 1.0 data_sharding ................................... True dataloader_type ................................. single ddp_average_in_collective ....................... False ddp_bucket_size ................................. None ddp_num_buckets ................................. None ddp_pad_buckets_for_high_nccl_busbw ............. False decoder_first_pipeline_num_layers ............... None decoder_last_pipeline_num_layers ................ None decoder_num_layers .............................. None decoder_seq_length .............................. None decoupled_lr .................................... None decoupled_min_lr ................................ None decrease_batch_size_if_needed ................... False defer_embedding_wgrad_compute ................... False deprecated_use_mcore_models ..................... False deterministic_mode .............................. False dino_bottleneck_size ............................ 256 dino_freeze_last_layer .......................... 1 dino_head_hidden_size ........................... 2048 dino_local_crops_number ......................... 10 dino_local_img_size ............................. 96 dino_norm_last_layer ............................ False dino_teacher_temp ............................... 0.07 dino_warmup_teacher_temp ........................ 0.04 dino_warmup_teacher_temp_epochs ................. 30 disable_straggler_on_startup .................... False dist_ckpt_format_deprecated ..................... None dist_ckpt_strictness ............................ assume_ok_unexpected distribute_saved_activations .................... False distributed_backend ............................. nccl distributed_timeout_minutes ..................... 10 embedding_path .................................. None empty_unused_memory_level ....................... 0 enable_cuda_graph ............................... False enable_ft_package ............................... False enable_gloo_process_groups ...................... True enable_one_logger ............................... True encoder_num_layers .............................. 9 encoder_pipeline_model_parallel_size ............ 0 encoder_seq_length .............................. 2048 encoder_tensor_model_parallel_size .............. 0 end_weight_decay ................................ 0.01 eod_mask_loss ................................... False error_injection_rate ............................ 0 error_injection_type ............................ transient_error eval_interval ................................... 1000 eval_iters ...................................... 100 evidence_data_path .............................. None exit_duration_in_mins ........................... None exit_interval ................................... None exit_on_missing_checkpoint ...................... False exit_signal_handler ............................. False exp_avg_dtype ................................... torch.float32 exp_avg_sq_dtype ................................ torch.float32 expert_model_parallel_size ...................... 1 expert_tensor_parallel_size ..................... 1 ffn_hidden_size ................................. 1368 finetune ........................................ False flash_decode .................................... False fp16 ............................................ False fp16_lm_cross_entropy ........................... False fp32_residual_connection ........................ False fp8 ............................................. None fp8_amax_compute_algo ........................... most_recent fp8_amax_history_len ............................ 1 fp8_interval .................................... 1 fp8_margin ...................................... 0 fp8_param_gather ................................ False fp8_wgrad ....................................... True global_batch_size ............................... 32 grad_reduce_in_bf16 ............................. False gradient_accumulation_fusion .................... False gradient_reduce_div_fusion ...................... True group_query_attention ........................... False head_lr_mult .................................... 1.0 hidden_dropout .................................. 0.0 hidden_size ..................................... 256 hierarchical_context_parallel_sizes ............. None hybrid_attention_ratio .......................... 0.0 hybrid_mlp_ratio ................................ 0.0 hybrid_override_pattern ......................... None hysteresis ...................................... 2 ict_head_size ................................... None ict_load ........................................ None img_h ........................................... 224 img_w ........................................... 224 indexer_batch_size .............................. 128 indexer_log_interval ............................ 1000 inference_batch_times_seqlen_threshold .......... -1 inference_max_batch_size ........................ 8 inference_max_seq_length ........................ 2560 inference_rng_tracker ........................... False init_method_std ................................. 0.02 init_method_xavier_uniform ...................... False init_model_with_meta_device ..................... False initial_loss_scale .............................. 4294967296 iter_per_epoch .................................. 1250 iterations_to_skip .............................. [] keep_fp8_transpose_cache_when_using_custom_fsdp . False kv_channels ..................................... 16 kv_lora_rank .................................... 32 lazy_mpu_init ................................... None load ............................................ /workspace/FLAME-MoE/results/phase0/runs/flame38m_g1_temporal_s2/ckpt local_rank ...................................... 0 log_interval .................................... 100 log_loss_scale_to_tensorboard ................... True log_memory_to_tensorboard ....................... False log_num_zeros_in_grad ........................... False log_params_norm ................................. False log_progress .................................... False log_straggler ................................... False log_throughput .................................. False log_timers_to_tensorboard ....................... False log_validation_ppl_to_tensorboard ............... False log_world_size_to_tensorboard ................... False logging_level ................................... None loss_scale ...................................... None loss_scale_window ............................... 1000 lr .............................................. None lr_decay_iters .................................. None lr_decay_samples ................................ None lr_decay_style .................................. linear lr_warmup_fraction .............................. None lr_warmup_init .................................. 0.0 lr_warmup_iters ................................. 0 lr_warmup_samples ............................... 0 lr_wsd_decay_iters .............................. None lr_wsd_decay_samples ............................ None lr_wsd_decay_style .............................. exponential main_grads_dtype ................................ torch.float32 main_params_dtype ............................... torch.float32 make_vocab_size_divisible_by .................... 128 manual_gc ....................................... False manual_gc_eval .................................. True manual_gc_interval .............................. 0 mask_factor ..................................... 1.0 mask_prob ....................................... 0.15 mask_type ....................................... random masked_softmax_fusion ........................... True max_position_embeddings ......................... 2048 max_tokens_to_oom ............................... 10000000 memory_snapshot_path ............................ snapshot.pickle merge_file ...................................... None micro_batch_size ................................ 32 microbatch_group_size_per_vp_stage .............. None min_loss_scale .................................. 1.0 min_lr .......................................... 0.0 mmap_bin_files .................................. True mock_data ....................................... False moe_aux_loss_coeff .............................. 0.01 moe_enable_deepep ............................... False moe_expert_capacity_factor ...................... None moe_extended_tp ................................. False moe_ffn_hidden_size ............................. 176 moe_grouped_gemm ................................ True moe_input_jitter_eps ............................ None moe_layer_freq .................................. [0, 1, 1, 1, 1, 1, 1, 1, 1] moe_layer_recompute ............................. False moe_pad_expert_input_to_capacity ................ False moe_per_layer_logging ........................... False moe_permute_fusion .............................. False moe_router_bias_update_rate ..................... 0.001 moe_router_dtype ................................ fp32 moe_router_enable_expert_bias ................... False moe_router_group_topk ........................... None moe_router_load_balancing_type .................. aux_loss moe_router_num_groups ........................... None moe_router_pre_softmax .......................... True moe_router_score_function ....................... softmax moe_router_topk ................................. 6 moe_router_topk_scaling_factor .................. None moe_shared_expert_intermediate_size ............. 352 moe_shared_expert_overlap ....................... False moe_token_dispatcher_type ....................... alltoall moe_token_drop_policy ........................... probs moe_use_legacy_grouped_gemm ..................... False moe_use_upcycling ............................... False moe_z_loss_coeff ................................ 0.001 mscale .......................................... 1.0 mscale_all_dim .................................. 1.0 multi_latent_attention .......................... False nccl_communicator_config_path ................... None no_load_optim ................................... True no_load_rng ..................................... True no_persist_layer_norm ........................... False no_save_optim ................................... None no_save_rng ..................................... None non_persistent_ckpt_type ........................ None non_persistent_global_ckpt_dir .................. None non_persistent_local_ckpt_algo .................. fully_parallel non_persistent_local_ckpt_dir ................... None non_persistent_save_interval .................... None norm_epsilon .................................... 1e-06 normalization ................................... RMSNorm num_attention_heads ............................. 16 num_channels .................................... 3 num_classes ..................................... 1000 num_dataset_builder_threads ..................... 1 num_distributed_optimizer_instances ............. 1 num_experts ..................................... 64 num_layers ...................................... 9 num_layers_per_virtual_pipeline_stage ........... None num_query_groups ................................ 1 num_virtual_stages_per_pipeline_rank ............ None num_workers ..................................... 2 one_logger_async ................................ False one_logger_project .............................. megatron-lm one_logger_run_name ............................. None onnx_safe ....................................... None openai_gelu ..................................... False optimizer ....................................... adam optimizer_cpu_offload ........................... False optimizer_offload_fraction ...................... 1.0 output_bert_embeddings .......................... False overlap_cpu_optimizer_d2h_h2d ................... False overlap_grad_reduce ............................. False overlap_p2p_comm ................................ False overlap_p2p_comm_warmup_flush ................... False overlap_param_gather ............................ False overlap_param_gather_with_optimizer_step ........ False override_opt_param_scheduler .................... False params_dtype .................................... torch.bfloat16 patch_dim ....................................... 16 per_split_data_args_path ........................ None perform_initialization .......................... True pin_cpu_grads ................................... True pin_cpu_params .................................. True pipeline_model_parallel_comm_backend ............ None pipeline_model_parallel_size .................... 1 pipeline_model_parallel_split_rank .............. None position_embedding_type ......................... rope pretrained_checkpoint ........................... None profile ......................................... False profile_ranks ................................... [0] profile_step_end ................................ 12 profile_step_start .............................. 10 q_lora_rank ..................................... None qk_head_dim ..................................... 128 qk_layernorm .................................... False qk_pos_emb_head_dim ............................. 64 query_in_block_prob ............................. 0.1 rampup_batch_size ............................... None rank ............................................ 0 recompute_granularity ........................... None recompute_method ................................ None recompute_num_layers ............................ None record_memory_history ........................... False relative_attention_max_distance ................. 128 relative_attention_num_buckets .................. 32 replication ..................................... False replication_factor .............................. 2 replication_jump ................................ None rerun_mode ...................................... disabled reset_attention_mask ............................ False reset_position_ids .............................. False result_rejected_tracker_filename ................ None retriever_report_topk_accuracies ................ [] retriever_score_scaling ......................... False retriever_seq_length ............................ 256 retro_add_retriever ............................. False retro_attention_gate ............................ 1 retro_cyclic_train_iters ........................ None retro_encoder_attention_dropout ................. 0.1 retro_encoder_hidden_dropout .................... 0.1 retro_encoder_layers ............................ 2 retro_num_neighbors ............................. 2 retro_num_retrieved_chunks ...................... 2 retro_project_dir ............................... None retro_verify_neighbor_count ..................... True rope_scaling_factor ............................. 8.0 rotary_base ..................................... 10000 rotary_interleaved .............................. False rotary_percent .................................. 1.0 rotary_scaling_factor ........................... 1.0 rotary_seq_len_interpolation_factor ............. None s3_cache_path ................................... None sample_rate ..................................... 1.0 save ............................................ None save_interval ................................... None scatter_gather_tensors_in_pipeline .............. True seed ............................................ 42 seq_length ...................................... 2048 sequence_parallel ............................... False sgd_momentum .................................... 0.9 short_seq_prob .................................. 0.1 skip_train ...................................... False skipped_train_samples ........................... 0 spec ............................................ None split ........................................... None squared_relu .................................... False start_weight_decay .............................. 0.01 straggler_ctrlr_port ............................ 65535 straggler_minmax_count .......................... 1 suggested_communication_unit_size ............... 400000000 swiglu .......................................... True swin_backbone_type .............................. tiny te_rng_tracker .................................. False tensor_model_parallel_size ...................... 1 tensorboard_dir ................................. None tensorboard_log_interval ........................ 1 tensorboard_queue_size .......................... 1000 test_data_path .................................. None test_mode ....................................... False tiktoken_num_special_tokens ..................... 1000 tiktoken_pattern ................................ None tiktoken_special_tokens ......................... None timing_log_level ................................ 0 timing_log_option ............................... minmax titles_data_path ................................ None tokenizer_model ................................. EleutherAI/pythia-12b tokenizer_type .................................. HuggingFaceTokenizer tp_comm_bootstrap_backend ....................... nccl tp_comm_bulk_dgrad .............................. True tp_comm_bulk_wgrad .............................. True tp_comm_overlap ................................. False tp_comm_overlap_ag .............................. True tp_comm_overlap_cfg ............................. None tp_comm_overlap_rs .............................. True tp_comm_overlap_rs_dgrad ........................ False tp_comm_split_ag ................................ True tp_comm_split_rs ................................ True train_data_path ................................. None train_iters ..................................... None train_samples ................................... None train_sync_interval ............................. None transformer_impl ................................ transformer_engine transformer_pipeline_model_parallel_size ........ 1 untie_embeddings_and_output_weights ............. True use_checkpoint_args ............................. False use_checkpoint_opt_param_scheduler .............. False use_cpu_initialization .......................... None use_custom_fsdp ................................. False use_dist_ckpt ................................... True use_dist_ckpt_deprecated ........................ False use_distributed_optimizer ....................... True use_flash_attn .................................. False use_legacy_models ............................... False use_mp_args_from_checkpoint_args ................ False use_one_sent_docs ............................... False use_persistent_ckpt_worker ...................... False use_precision_aware_optimizer ................... False use_pytorch_profiler ............................ False use_ring_exchange_p2p ........................... False use_rope_scaling ................................ False use_rotary_position_embeddings .................. False use_tokenizer_model_from_checkpoint_args ........ True use_torch_fsdp2 ................................. False use_torch_optimizer_for_cpu_offload ............. False use_tp_pp_dp_mapping ............................ False v_head_dim ...................................... 128 valid_data_path ................................. None variable_seq_lengths ............................ False virtual_pipeline_model_parallel_size ............ None vision_backbone_type ............................ vit vision_pretraining .............................. False vision_pretraining_type ......................... classify vocab_extra_ids ................................. 0 vocab_file ...................................... None vocab_size ...................................... None wandb_exp_name .................................. wandb_project ................................... wandb_save_dir .................................. weight_decay .................................... 0.01 weight_decay_incr_style ......................... constant wgrad_deferral_limit ............................ 0 world_size ...................................... 1 yaml_cfg ........................................ None -------------------- end of arguments --------------------- 2026-07-23:03:15:21,424 INFO [num_microbatches_calculator.py:228] setting number of microbatches to constant 1 > building HuggingFaceTokenizer tokenizer ... 2026-07-23:03:15:21,653 INFO [_client.py:1038] HTTP Request: HEAD https://huggingface.co/EleutherAI/pythia-12b/resolve/main/config.json "HTTP/1.1 307 Temporary Redirect" 2026-07-23:03:15:21,666 INFO [_client.py:1038] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/EleutherAI/pythia-12b/bb1e3e710cdf6b524461d543cfb5ba773f0a81b6/config.json "HTTP/1.1 200 OK" 2026-07-23:03:15:21,803 INFO [_client.py:1038] HTTP Request: HEAD https://huggingface.co/EleutherAI/pythia-12b/resolve/main/tokenizer_config.json "HTTP/1.1 307 Temporary Redirect" 2026-07-23:03:15:21,816 INFO [_client.py:1038] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/EleutherAI/pythia-12b/bb1e3e710cdf6b524461d543cfb5ba773f0a81b6/tokenizer_config.json "HTTP/1.1 200 OK" 2026-07-23:03:15:21,914 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/models/EleutherAI/pythia-12b/tree/main/additional_chat_templates?recursive=false&expand=false "HTTP/1.1 404 Not Found" 2026-07-23:03:15:21,981 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/models/EleutherAI/pythia-12b/tree/main?recursive=true&expand=false "HTTP/1.1 200 OK" > padded vocab (size: 50277) with 27 dummy tokens (new size: 50304) WARNING: one_logger package is required to enable e2e metrics tracking. please go to https://confluence.nvidia.com/display/MLWFO/Package+Repositories for details to install it 2026-07-23:03:15:22,188 WARNING [rerun_state_machine.py:239] RerunStateMachine initialized in mode RerunMode.DISABLED > initializing torch distributed ... [W723 03:15:22.706138567 CUDAAllocatorConfig.h:28] Warning: expandable_segments not supported on this platform (function operator()) > initialized tensor model parallel with size 1 > initialized pipeline model parallel with size 1 > setting random seeds to 42 ... > compiling dataset index builder ... make: Entering directory '/workspace/FLAME-MoE/Megatron-LM/megatron/core/datasets' make: Nothing to be done for 'default'. make: Leaving directory '/workspace/FLAME-MoE/Megatron-LM/megatron/core/datasets' >>> done with dataset index builder. Compilation time: 0.686 seconds > compiling and loading fused kernels ... >>> done with compiling and loading fused kernels. Compilation time: 0.359 seconds building GPT model ... > number of parameters on (tensor, pipeline) model parallel rank (0, 0): 100670208 /workspace/FLAME-MoE/Megatron-LM/megatron/core/transformer/transformer_layer.py:339: UserWarning: TransformerLayer._get_layer_offset is deprecated.Please use get_transformer_layer_offset instead. warnings.warn( loading distributed checkpoint from /workspace/FLAME-MoE/results/phase0/runs/flame38m_g1_temporal_s2/ckpt at iteration 2121 /workspace/FLAME-MoE/Megatron-LM/megatron/core/dist_checkpointing/strategies/torch.py:847: FutureWarning: `load_state_dict` is deprecated and will be removed in future versions. Please use `load` instead. checkpoint.load_state_dict( /usr/local/lib/python3.11/dist-packages/torch/distributed/checkpoint/planner_helpers.py:306: FutureWarning: Please use DTensor instead and we are deprecating ShardedTensor. device = getattr(value, "device", None) /workspace/FLAME-MoE/.venv/lib/python3.11/site-packages/transformer_engine/pytorch/module/base.py:566: FutureWarning: You are using `torch.load` with `weights_only=False` (the current default value), which uses the default pickle module implicitly. It is possible to construct malicious pickle data which will execute arbitrary code during unpickling (See https://github.com/pytorch/pytorch/blob/main/SECURITY.md#untrusted-models for more details). In a future release, the default value for `weights_only` will be flipped to `True`. This limits the functions that could be executed during unpickling. Arbitrary objects will no longer be allowed to be loaded via this mode unless they are explicitly allowlisted by the user via `torch.serialization.add_safe_globals`. We recommend you start setting `weights_only=True` for any use case where you don't have full control of the loaded file. Please open an issue on GitHub for any issues related to this experimental feature. state = torch.load(state, map_location="cuda") checkpoint version 3.0 2026-07-23:03:15:27,643 WARNING [rerun_state_machine.py:811] RerunStateMachine disabled via CLI, ignoring machine state saved in checkpoint successfully loaded checkpoint from /workspace/FLAME-MoE/results/phase0/runs/flame38m_g1_temporal_s2/ckpt [ t 1/1, p 1/1 ] at iteration 2121 2026-07-23:03:15:27,774 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc "HTTP/1.1 200 OK" 2026-07-23:03:15:27,999 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc "HTTP/1.1 200 OK" 2026-07-23:03:15:28,065 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc/revision/210d026faf9955653af8916fad021475a3f00453 "HTTP/1.1 200 OK" 2026-07-23:03:15:28,124 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc/tree/210d026faf9955653af8916fad021475a3f00453?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:15:28,380 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc/tree/210d026faf9955653af8916fad021475a3f00453/ARC-Challenge?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:15:28,927 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc "HTTP/1.1 200 OK" 2026-07-23:03:15:29,166 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc "HTTP/1.1 200 OK" 2026-07-23:03:15:29,461 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc/tree/210d026faf9955653af8916fad021475a3f00453/ARC-Easy?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:15:29,606 WARNING [task.py:799] [Task: boolq] metric acc is defined, but aggregation is not. using default aggregation=mean 2026-07-23:03:15:29,606 WARNING [task.py:811] [Task: boolq] metric acc is defined, but higher_is_better is not. using default higher_is_better=True 2026-07-23:03:15:29,667 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue "HTTP/1.1 200 OK" 2026-07-23:03:15:29,903 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue "HTTP/1.1 200 OK" 2026-07-23:03:15:29,965 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/revision/3de24cf8022e94f4ee4b9d55a6f539891524d646 "HTTP/1.1 200 OK" 2026-07-23:03:15:30,037 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/tree/3de24cf8022e94f4ee4b9d55a6f539891524d646?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:15:30,301 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/tree/3de24cf8022e94f4ee4b9d55a6f539891524d646/axb?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:15:30,423 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/revision/3de24cf8022e94f4ee4b9d55a6f539891524d646 "HTTP/1.1 200 OK" 2026-07-23:03:15:30,501 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/tree/3de24cf8022e94f4ee4b9d55a6f539891524d646/boolq?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:15:30,753 WARNING [task.py:799] [Task: copa] metric acc is defined, but aggregation is not. using default aggregation=mean 2026-07-23:03:15:30,753 WARNING [task.py:811] [Task: copa] metric acc is defined, but higher_is_better is not. using default higher_is_better=True 2026-07-23:03:15:30,813 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue "HTTP/1.1 200 OK" 2026-07-23:03:15:31,056 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue "HTTP/1.1 200 OK" 2026-07-23:03:15:31,372 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/tree/3de24cf8022e94f4ee4b9d55a6f539891524d646/copa?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:15:31,525 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag "HTTP/1.1 200 OK" 2026-07-23:03:15:31,754 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag "HTTP/1.1 200 OK" 2026-07-23:03:15:31,812 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag/revision/218ec52e09a7e7462a5400043bb9a69a41d06b76 "HTTP/1.1 200 OK" 2026-07-23:03:15:31,963 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag/tree/218ec52e09a7e7462a5400043bb9a69a41d06b76?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:15:32,174 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag/tree/218ec52e09a7e7462a5400043bb9a69a41d06b76/data?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:15:32,234 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag/revision/218ec52e09a7e7462a5400043bb9a69a41d06b76 "HTTP/1.1 200 OK" 2026-07-23:03:15:34,713 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai "HTTP/1.1 200 OK" 2026-07-23:03:15:35,023 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai "HTTP/1.1 200 OK" 2026-07-23:03:15:35,083 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/revision/900124bf3b8235c6daf21033af9948b3f07346c4 "HTTP/1.1 200 OK" 2026-07-23:03:15:35,149 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/tree/900124bf3b8235c6daf21033af9948b3f07346c4?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:15:35,355 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/tree/900124bf3b8235c6daf21033af9948b3f07346c4/default%2Ftest?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:15:35,418 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/tree/900124bf3b8235c6daf21033af9948b3f07346c4/default?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:15:35,482 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/revision/900124bf3b8235c6daf21033af9948b3f07346c4 "HTTP/1.1 200 OK" 2026-07-23:03:15:35,510 WARNING [task.py:325] [Task: lambada_openai] has_training_docs and has_validation_docs are False, using test_docs as fewshot_docs but this is not recommended. 2026-07-23:03:15:35,510 WARNING [task.py:325] [Task: lambada_openai] has_training_docs and has_validation_docs are False, using test_docs as fewshot_docs but this is not recommended. 2026-07-23:03:15:35,676 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa "HTTP/1.1 200 OK" 2026-07-23:03:15:35,916 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa "HTTP/1.1 200 OK" 2026-07-23:03:15:36,021 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/revision/388097ea7776314e93a529163e0fea805b8a6454 "HTTP/1.1 200 OK" 2026-07-23:03:15:36,091 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/tree/388097ea7776314e93a529163e0fea805b8a6454?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:15:36,335 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/tree/388097ea7776314e93a529163e0fea805b8a6454/additional?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:15:36,398 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/revision/388097ea7776314e93a529163e0fea805b8a6454 "HTTP/1.1 200 OK" 2026-07-23:03:15:36,486 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/tree/388097ea7776314e93a529163e0fea805b8a6454/main?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:15:36,761 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa "HTTP/1.1 200 OK" 2026-07-23:03:15:36,824 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa "HTTP/1.1 200 OK" 2026-07-23:03:15:36,882 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa/tree/main?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:15:36,949 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/datasets/ybisk/piqa/resolve/main/piqa.py "HTTP/1.1 307 Temporary Redirect" 2026-07-23:03:15:36,963 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/resolve-cache/datasets/ybisk/piqa/2e8ac2dffd59bac8c3c6714948f4c551a0848bb0/piqa.py?%2Fdatasets%2Fybisk%2Fpiqa%2Fresolve%2Fmain%2Fpiqa.py=&etag=%220f0d96172f9bf7aa10069374b3236ef083cfeb8e%22 "HTTP/1.1 200 OK" 2026-07-23:03:15:37,593 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa/revision/main "HTTP/1.1 200 OK" 2026-07-23:03:15:37,655 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa/tree/main?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:15:38,159 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq "HTTP/1.1 200 OK" 2026-07-23:03:15:38,411 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq "HTTP/1.1 200 OK" 2026-07-23:03:15:38,536 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq/revision/2c94ad3e1aafab77146f384e23536f97a4849815 "HTTP/1.1 200 OK" 2026-07-23:03:15:38,599 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq/tree/2c94ad3e1aafab77146f384e23536f97a4849815?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:15:38,802 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq/tree/2c94ad3e1aafab77146f384e23536f97a4849815/data?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:15:38,868 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq/revision/2c94ad3e1aafab77146f384e23536f97a4849815 "HTTP/1.1 200 OK" 2026-07-23:03:15:39,375 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande "HTTP/1.1 200 OK" 2026-07-23:03:15:39,617 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande "HTTP/1.1 200 OK" 2026-07-23:03:15:39,680 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/revision/01e74176c63542e6b0bcb004dcdea22d94fb67b5 "HTTP/1.1 200 OK" 2026-07-23:03:15:39,746 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/tree/01e74176c63542e6b0bcb004dcdea22d94fb67b5?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:15:40,039 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/tree/01e74176c63542e6b0bcb004dcdea22d94fb67b5/winogrande_debiased?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:15:40,099 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/revision/01e74176c63542e6b0bcb004dcdea22d94fb67b5 "HTTP/1.1 200 OK" 2026-07-23:03:15:40,263 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/tree/01e74176c63542e6b0bcb004dcdea22d94fb67b5/winogrande_xl?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:15:40,950 WARNING [evaluator.py:270] Overwriting default num_fewshot of winogrande from None to 0 2026-07-23:03:15:40,950 WARNING [evaluator.py:270] Overwriting default num_fewshot of sciq from None to 0 2026-07-23:03:15:40,950 WARNING [evaluator.py:270] Overwriting default num_fewshot of piqa from None to 0 2026-07-23:03:15:40,950 WARNING [evaluator.py:270] Overwriting default num_fewshot of openbookqa from None to 0 2026-07-23:03:15:40,950 WARNING [evaluator.py:270] Overwriting default num_fewshot of lambada_openai from None to 0 2026-07-23:03:15:40,950 WARNING [evaluator.py:270] Overwriting default num_fewshot of hellaswag from None to 0 2026-07-23:03:15:40,950 WARNING [evaluator.py:270] Overwriting default num_fewshot of copa from None to 0 2026-07-23:03:15:40,950 WARNING [evaluator.py:270] Overwriting default num_fewshot of boolq from None to 0 2026-07-23:03:15:40,950 WARNING [evaluator.py:270] Overwriting default num_fewshot of arc_easy from None to 0 2026-07-23:03:15:40,950 WARNING [evaluator.py:270] Overwriting default num_fewshot of arc_challenge from None to 0 2026-07-23:03:15:40,972 INFO [task.py:415] Building contexts for winogrande on rank 0... 0%| | 0/1267 [00:00 /workspace/FLAME-MoE/results/ablations/flame1e18_downstream.csv ==================================================================== [run] flame38m_g1_temporal_s3 mode=temporal TEMPORAL=1 H=256 heads=16 ffn=1368 experts=64 topk=6 moe_ffn=176 shared=352 [run] tasks=arc_challenge,arc_easy,hellaswag,openbookqa,piqa,winogrande,boolq,copa,lambada_openai,sciq out=/workspace/FLAME-MoE/results/phase0/runs/flame38m_g1_temporal_s3/lmeval_1e18_0shot ==================================================================== /usr/local/lib/python3.11/dist-packages/torch/cuda/__init__.py:58: FutureWarning: The pynvml package is deprecated. Please install nvidia-ml-py instead. If you did not install pynvml directly, please report this to the maintainers of the package that installed pynvml for you. import pynvml # type: ignore[import] [run_lmeval] stubbed VLM backends (hf_vlms, vllm_vlms) /usr/local/lib/python3.11/dist-packages/torch/cuda/__init__.py:58: FutureWarning: The pynvml package is deprecated. Please install nvidia-ml-py instead. If you did not install pynvml directly, please report this to the maintainers of the package that installed pynvml for you. import pynvml # type: ignore[import] /usr/local/lib/python3.11/dist-packages/torch/cuda/__init__.py:58: FutureWarning: The pynvml package is deprecated. Please install nvidia-ml-py instead. If you did not install pynvml directly, please report this to the maintainers of the package that installed pynvml for you. import pynvml # type: ignore[import] /workspace/FLAME-MoE/.venv/lib/python3.11/site-packages/transformer_engine/pytorch/attention.py:108: UserWarning: To use flash-attn v3, please use the following commands to install: (1) pip install "git+https://github.com/Dao-AILab/flash-attention.git#egg=flashattn-hopper&subdirectory=hopper" (2) python_path=`python -c "import site; print(site.getsitepackages()[0])"` (3) mkdir -p $python_path/flashattn_hopper (4) wget -P $python_path/flashattn_hopper https://raw.githubusercontent.com/Dao-AILab/flash-attention/main/hopper/flash_attn_interface.py warnings.warn( [temporal] rolling-residency router installed (evict=min_logit) [run_lmeval] installed temporal residency router (native regime) /usr/local/lib/python3.11/dist-packages/requests/__init__.py:113: RequestsDependencyWarning: urllib3 (2.2.3) or chardet (6.0.0.post1)/charset_normalizer (3.3.2) doesn't match a supported version! warnings.warn( 2026-07-23:03:23:22,781 INFO [__main__.py:279] Verbosity set to INFO 2026-07-23:03:23:59,018 INFO [__init__.py:459] The tag 'arc_ca' is already registered as a group, this tag will not be registered. This may affect tasks you want to call. 2026-07-23:03:23:59,032 INFO [__init__.py:459] The tag 'arc_ca' is already registered as a group, this tag will not be registered. This may affect tasks you want to call. 2026-07-23:03:24:19,486 INFO [__main__.py:376] Selected Tasks: ['arc_challenge', 'arc_easy', 'boolq', 'copa', 'hellaswag', 'lambada_openai', 'openbookqa', 'piqa', 'sciq', 'winogrande'] 2026-07-23:03:24:19,527 INFO [evaluator.py:164] Setting random seed to 42 | Setting numpy seed to 42 | Setting torch manual seed to 42 | Setting fewshot manual seed to 42 2026-07-23:03:24:19,527 INFO [evaluator.py:201] Initializing megatron_lm model, with arguments: {} using world size: 1, data-parallel size: 1, context-parallel size: 1, hierarchical context-parallel sizes: Nonetensor-model-parallel size: 1, encoder-tensor-model-parallel size: 0, pipeline-model-parallel size: 1, encoder-pipeline-model-parallel size: 0 WARNING: overriding default arguments for exit_on_missing_checkpoint:True with exit_on_missing_checkpoint:False setting global batch size to 32 Number of virtual stages per pipeline stage: None accumulate and all-reduce gradients in fp32 for bfloat16 data type. using torch.bfloat16 for parameters ... ------------------------ arguments ------------------------ account_for_embedding_in_pipeline_split ......... False account_for_loss_in_pipeline_split .............. False accumulate_allreduce_grads_in_fp32 .............. True adam_beta1 ...................................... 0.9 adam_beta2 ...................................... 0.999 adam_eps ........................................ 1e-08 add_bias_linear ................................. False add_position_embedding .......................... True add_qkv_bias .................................... False adlr_autoresume ................................. False adlr_autoresume_interval ........................ 1000 align_grad_reduce ............................... True align_param_gather .............................. False app_tag_run_name ................................ None app_tag_run_version ............................. 0.0.0 apply_layernorm_1p .............................. False apply_query_key_layer_scaling ................... False apply_residual_connection_post_layernorm ........ False apply_rope_fusion ............................... True async_save ...................................... None async_tensor_model_parallel_allreduce ........... True attention_backend ............................... AttnBackend.auto attention_dropout ............................... 0.0 attention_softmax_in_fp32 ....................... False auto_detect_ckpt_format ......................... False barrier_with_L1_time ............................ True bert_binary_head ................................ True bert_embedder_type .............................. megatron bert_load ....................................... None bf16 ............................................ True bias_dropout_fusion ............................. True bias_gelu_fusion ................................ False bias_swiglu_fusion .............................. True biencoder_projection_dim ........................ 0 biencoder_shared_query_context_model ............ False block_data_path ................................. None calc_ft_timeouts ................................ False calculate_per_token_loss ........................ False check_for_large_grads ........................... False check_for_nan_in_loss_and_grad .................. True check_for_spiky_loss ............................ False check_weight_hash_across_dp_replicas_interval ... None ckpt_assume_constant_structure .................. False ckpt_convert_format ............................. None ckpt_convert_save ............................... None ckpt_convert_update_legacy_dist_opt_format ...... False ckpt_format ..................................... torch_dist ckpt_fully_parallel_load ........................ False ckpt_fully_parallel_save ........................ True ckpt_fully_parallel_save_deprecated ............. False ckpt_step ....................................... None classes_fraction ................................ 1.0 clip_grad ....................................... 1.0 clone_scatter_output_in_embedding ............... True config_logger_dir ............................... consumed_train_samples .......................... 0 consumed_valid_samples .......................... 0 context_parallel_size ........................... 1 cp_comm_type .................................... ['p2p'] create_attention_mask_in_dataloader ............. True cross_entropy_loss_fusion ....................... False cuda_graph_warmup_steps ......................... 3 data_args_path .................................. None data_cache_path ................................. None data_parallel_random_init ....................... False data_parallel_sharding_strategy ................. no_shard data_parallel_size .............................. 1 data_path ....................................... None data_per_class_fraction ......................... 1.0 data_sharding ................................... True dataloader_type ................................. single ddp_average_in_collective ....................... False ddp_bucket_size ................................. None ddp_num_buckets ................................. None ddp_pad_buckets_for_high_nccl_busbw ............. False decoder_first_pipeline_num_layers ............... None decoder_last_pipeline_num_layers ................ None decoder_num_layers .............................. None decoder_seq_length .............................. None decoupled_lr .................................... None decoupled_min_lr ................................ None decrease_batch_size_if_needed ................... False defer_embedding_wgrad_compute ................... False deprecated_use_mcore_models ..................... False deterministic_mode .............................. False dino_bottleneck_size ............................ 256 dino_freeze_last_layer .......................... 1 dino_head_hidden_size ........................... 2048 dino_local_crops_number ......................... 10 dino_local_img_size ............................. 96 dino_norm_last_layer ............................ False dino_teacher_temp ............................... 0.07 dino_warmup_teacher_temp ........................ 0.04 dino_warmup_teacher_temp_epochs ................. 30 disable_straggler_on_startup .................... False dist_ckpt_format_deprecated ..................... None dist_ckpt_strictness ............................ assume_ok_unexpected distribute_saved_activations .................... False distributed_backend ............................. nccl distributed_timeout_minutes ..................... 10 embedding_path .................................. None empty_unused_memory_level ....................... 0 enable_cuda_graph ............................... False enable_ft_package ............................... False enable_gloo_process_groups ...................... True enable_one_logger ............................... True encoder_num_layers .............................. 9 encoder_pipeline_model_parallel_size ............ 0 encoder_seq_length .............................. 2048 encoder_tensor_model_parallel_size .............. 0 end_weight_decay ................................ 0.01 eod_mask_loss ................................... False error_injection_rate ............................ 0 error_injection_type ............................ transient_error eval_interval ................................... 1000 eval_iters ...................................... 100 evidence_data_path .............................. None exit_duration_in_mins ........................... None exit_interval ................................... None exit_on_missing_checkpoint ...................... False exit_signal_handler ............................. False exp_avg_dtype ................................... torch.float32 exp_avg_sq_dtype ................................ torch.float32 expert_model_parallel_size ...................... 1 expert_tensor_parallel_size ..................... 1 ffn_hidden_size ................................. 1368 finetune ........................................ False flash_decode .................................... False fp16 ............................................ False fp16_lm_cross_entropy ........................... False fp32_residual_connection ........................ False fp8 ............................................. None fp8_amax_compute_algo ........................... most_recent fp8_amax_history_len ............................ 1 fp8_interval .................................... 1 fp8_margin ...................................... 0 fp8_param_gather ................................ False fp8_wgrad ....................................... True global_batch_size ............................... 32 grad_reduce_in_bf16 ............................. False gradient_accumulation_fusion .................... False gradient_reduce_div_fusion ...................... True group_query_attention ........................... False head_lr_mult .................................... 1.0 hidden_dropout .................................. 0.0 hidden_size ..................................... 256 hierarchical_context_parallel_sizes ............. None hybrid_attention_ratio .......................... 0.0 hybrid_mlp_ratio ................................ 0.0 hybrid_override_pattern ......................... None hysteresis ...................................... 2 ict_head_size ................................... None ict_load ........................................ None img_h ........................................... 224 img_w ........................................... 224 indexer_batch_size .............................. 128 indexer_log_interval ............................ 1000 inference_batch_times_seqlen_threshold .......... -1 inference_max_batch_size ........................ 8 inference_max_seq_length ........................ 2560 inference_rng_tracker ........................... False init_method_std ................................. 0.02 init_method_xavier_uniform ...................... False init_model_with_meta_device ..................... False initial_loss_scale .............................. 4294967296 iter_per_epoch .................................. 1250 iterations_to_skip .............................. [] keep_fp8_transpose_cache_when_using_custom_fsdp . False kv_channels ..................................... 16 kv_lora_rank .................................... 32 lazy_mpu_init ................................... None load ............................................ /workspace/FLAME-MoE/results/phase0/runs/flame38m_g1_temporal_s3/ckpt local_rank ...................................... 0 log_interval .................................... 100 log_loss_scale_to_tensorboard ................... True log_memory_to_tensorboard ....................... False log_num_zeros_in_grad ........................... False log_params_norm ................................. False log_progress .................................... False log_straggler ................................... False log_throughput .................................. False log_timers_to_tensorboard ....................... False log_validation_ppl_to_tensorboard ............... False log_world_size_to_tensorboard ................... False logging_level ................................... None loss_scale ...................................... None loss_scale_window ............................... 1000 lr .............................................. None lr_decay_iters .................................. None lr_decay_samples ................................ None lr_decay_style .................................. linear lr_warmup_fraction .............................. None lr_warmup_init .................................. 0.0 lr_warmup_iters ................................. 0 lr_warmup_samples ............................... 0 lr_wsd_decay_iters .............................. None lr_wsd_decay_samples ............................ None lr_wsd_decay_style .............................. exponential main_grads_dtype ................................ torch.float32 main_params_dtype ............................... torch.float32 make_vocab_size_divisible_by .................... 128 manual_gc ....................................... False manual_gc_eval .................................. True manual_gc_interval .............................. 0 mask_factor ..................................... 1.0 mask_prob ....................................... 0.15 mask_type ....................................... random masked_softmax_fusion ........................... True max_position_embeddings ......................... 2048 max_tokens_to_oom ............................... 10000000 memory_snapshot_path ............................ snapshot.pickle merge_file ...................................... None micro_batch_size ................................ 32 microbatch_group_size_per_vp_stage .............. None min_loss_scale .................................. 1.0 min_lr .......................................... 0.0 mmap_bin_files .................................. True mock_data ....................................... False moe_aux_loss_coeff .............................. 0.01 moe_enable_deepep ............................... False moe_expert_capacity_factor ...................... None moe_extended_tp ................................. False moe_ffn_hidden_size ............................. 176 moe_grouped_gemm ................................ True moe_input_jitter_eps ............................ None moe_layer_freq .................................. [0, 1, 1, 1, 1, 1, 1, 1, 1] moe_layer_recompute ............................. False moe_pad_expert_input_to_capacity ................ False moe_per_layer_logging ........................... False moe_permute_fusion .............................. False moe_router_bias_update_rate ..................... 0.001 moe_router_dtype ................................ fp32 moe_router_enable_expert_bias ................... False moe_router_group_topk ........................... None moe_router_load_balancing_type .................. aux_loss moe_router_num_groups ........................... None moe_router_pre_softmax .......................... True moe_router_score_function ....................... softmax moe_router_topk ................................. 6 moe_router_topk_scaling_factor .................. None moe_shared_expert_intermediate_size ............. 352 moe_shared_expert_overlap ....................... False moe_token_dispatcher_type ....................... alltoall moe_token_drop_policy ........................... probs moe_use_legacy_grouped_gemm ..................... False moe_use_upcycling ............................... False moe_z_loss_coeff ................................ 0.001 mscale .......................................... 1.0 mscale_all_dim .................................. 1.0 multi_latent_attention .......................... False nccl_communicator_config_path ................... None no_load_optim ................................... True no_load_rng ..................................... True no_persist_layer_norm ........................... False no_save_optim ................................... None no_save_rng ..................................... None non_persistent_ckpt_type ........................ None non_persistent_global_ckpt_dir .................. None non_persistent_local_ckpt_algo .................. fully_parallel non_persistent_local_ckpt_dir ................... None non_persistent_save_interval .................... None norm_epsilon .................................... 1e-06 normalization ................................... RMSNorm num_attention_heads ............................. 16 num_channels .................................... 3 num_classes ..................................... 1000 num_dataset_builder_threads ..................... 1 num_distributed_optimizer_instances ............. 1 num_experts ..................................... 64 num_layers ...................................... 9 num_layers_per_virtual_pipeline_stage ........... None num_query_groups ................................ 1 num_virtual_stages_per_pipeline_rank ............ None num_workers ..................................... 2 one_logger_async ................................ False one_logger_project .............................. megatron-lm one_logger_run_name ............................. None onnx_safe ....................................... None openai_gelu ..................................... False optimizer ....................................... adam optimizer_cpu_offload ........................... False optimizer_offload_fraction ...................... 1.0 output_bert_embeddings .......................... False overlap_cpu_optimizer_d2h_h2d ................... False overlap_grad_reduce ............................. False overlap_p2p_comm ................................ False overlap_p2p_comm_warmup_flush ................... False overlap_param_gather ............................ False overlap_param_gather_with_optimizer_step ........ False override_opt_param_scheduler .................... False params_dtype .................................... torch.bfloat16 patch_dim ....................................... 16 per_split_data_args_path ........................ None perform_initialization .......................... True pin_cpu_grads ................................... True pin_cpu_params .................................. True pipeline_model_parallel_comm_backend ............ None pipeline_model_parallel_size .................... 1 pipeline_model_parallel_split_rank .............. None position_embedding_type ......................... rope pretrained_checkpoint ........................... None profile ......................................... False profile_ranks ................................... [0] profile_step_end ................................ 12 profile_step_start .............................. 10 q_lora_rank ..................................... None qk_head_dim ..................................... 128 qk_layernorm .................................... False qk_pos_emb_head_dim ............................. 64 query_in_block_prob ............................. 0.1 rampup_batch_size ............................... None rank ............................................ 0 recompute_granularity ........................... None recompute_method ................................ None recompute_num_layers ............................ None record_memory_history ........................... False relative_attention_max_distance ................. 128 relative_attention_num_buckets .................. 32 replication ..................................... False replication_factor .............................. 2 replication_jump ................................ None rerun_mode ...................................... disabled reset_attention_mask ............................ False reset_position_ids .............................. False result_rejected_tracker_filename ................ None retriever_report_topk_accuracies ................ [] retriever_score_scaling ......................... False retriever_seq_length ............................ 256 retro_add_retriever ............................. False retro_attention_gate ............................ 1 retro_cyclic_train_iters ........................ None retro_encoder_attention_dropout ................. 0.1 retro_encoder_hidden_dropout .................... 0.1 retro_encoder_layers ............................ 2 retro_num_neighbors ............................. 2 retro_num_retrieved_chunks ...................... 2 retro_project_dir ............................... None retro_verify_neighbor_count ..................... True rope_scaling_factor ............................. 8.0 rotary_base ..................................... 10000 rotary_interleaved .............................. False rotary_percent .................................. 1.0 rotary_scaling_factor ........................... 1.0 rotary_seq_len_interpolation_factor ............. None s3_cache_path ................................... None sample_rate ..................................... 1.0 save ............................................ None save_interval ................................... None scatter_gather_tensors_in_pipeline .............. True seed ............................................ 42 seq_length ...................................... 2048 sequence_parallel ............................... False sgd_momentum .................................... 0.9 short_seq_prob .................................. 0.1 skip_train ...................................... False skipped_train_samples ........................... 0 spec ............................................ None split ........................................... None squared_relu .................................... False start_weight_decay .............................. 0.01 straggler_ctrlr_port ............................ 65535 straggler_minmax_count .......................... 1 suggested_communication_unit_size ............... 400000000 swiglu .......................................... True swin_backbone_type .............................. tiny te_rng_tracker .................................. False tensor_model_parallel_size ...................... 1 tensorboard_dir ................................. None tensorboard_log_interval ........................ 1 tensorboard_queue_size .......................... 1000 test_data_path .................................. None test_mode ....................................... False tiktoken_num_special_tokens ..................... 1000 tiktoken_pattern ................................ None tiktoken_special_tokens ......................... None timing_log_level ................................ 0 timing_log_option ............................... minmax titles_data_path ................................ None tokenizer_model ................................. EleutherAI/pythia-12b tokenizer_type .................................. HuggingFaceTokenizer tp_comm_bootstrap_backend ....................... nccl tp_comm_bulk_dgrad .............................. True tp_comm_bulk_wgrad .............................. True tp_comm_overlap ................................. False tp_comm_overlap_ag .............................. True tp_comm_overlap_cfg ............................. None tp_comm_overlap_rs .............................. True tp_comm_overlap_rs_dgrad ........................ False tp_comm_split_ag ................................ True tp_comm_split_rs ................................ True train_data_path ................................. None train_iters ..................................... None train_samples ................................... None train_sync_interval ............................. None transformer_impl ................................ transformer_engine transformer_pipeline_model_parallel_size ........ 1 untie_embeddings_and_output_weights ............. True use_checkpoint_args ............................. False use_checkpoint_opt_param_scheduler .............. False use_cpu_initialization .......................... None use_custom_fsdp ................................. False use_dist_ckpt ................................... True use_dist_ckpt_deprecated ........................ False use_distributed_optimizer ....................... True use_flash_attn .................................. False use_legacy_models ............................... False use_mp_args_from_checkpoint_args ................ False use_one_sent_docs ............................... False use_persistent_ckpt_worker ...................... False use_precision_aware_optimizer ................... False use_pytorch_profiler ............................ False use_ring_exchange_p2p ........................... False use_rope_scaling ................................ False use_rotary_position_embeddings .................. False use_tokenizer_model_from_checkpoint_args ........ True use_torch_fsdp2 ................................. False use_torch_optimizer_for_cpu_offload ............. False use_tp_pp_dp_mapping ............................ False v_head_dim ...................................... 128 valid_data_path ................................. None variable_seq_lengths ............................ False virtual_pipeline_model_parallel_size ............ None vision_backbone_type ............................ vit vision_pretraining .............................. False vision_pretraining_type ......................... classify vocab_extra_ids ................................. 0 vocab_file ...................................... None vocab_size ...................................... None wandb_exp_name .................................. wandb_project ................................... wandb_save_dir .................................. weight_decay .................................... 0.01 weight_decay_incr_style ......................... constant wgrad_deferral_limit ............................ 0 world_size ...................................... 1 yaml_cfg ........................................ None -------------------- end of arguments --------------------- 2026-07-23:03:24:19,545 INFO [num_microbatches_calculator.py:228] setting number of microbatches to constant 1 > building HuggingFaceTokenizer tokenizer ... 2026-07-23:03:24:19,747 INFO [_client.py:1038] HTTP Request: HEAD https://huggingface.co/EleutherAI/pythia-12b/resolve/main/config.json "HTTP/1.1 307 Temporary Redirect" 2026-07-23:03:24:19,758 INFO [_client.py:1038] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/EleutherAI/pythia-12b/bb1e3e710cdf6b524461d543cfb5ba773f0a81b6/config.json "HTTP/1.1 200 OK" 2026-07-23:03:24:19,957 INFO [_client.py:1038] HTTP Request: HEAD https://huggingface.co/EleutherAI/pythia-12b/resolve/main/tokenizer_config.json "HTTP/1.1 307 Temporary Redirect" 2026-07-23:03:24:19,968 INFO [_client.py:1038] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/EleutherAI/pythia-12b/bb1e3e710cdf6b524461d543cfb5ba773f0a81b6/tokenizer_config.json "HTTP/1.1 200 OK" 2026-07-23:03:24:20,082 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/models/EleutherAI/pythia-12b/tree/main/additional_chat_templates?recursive=false&expand=false "HTTP/1.1 404 Not Found" 2026-07-23:03:24:20,149 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/models/EleutherAI/pythia-12b/tree/main?recursive=true&expand=false "HTTP/1.1 200 OK" > padded vocab (size: 50277) with 27 dummy tokens (new size: 50304) WARNING: one_logger package is required to enable e2e metrics tracking. please go to https://confluence.nvidia.com/display/MLWFO/Package+Repositories for details to install it 2026-07-23:03:24:20,371 WARNING [rerun_state_machine.py:239] RerunStateMachine initialized in mode RerunMode.DISABLED > initializing torch distributed ... [W723 03:24:20.888974960 CUDAAllocatorConfig.h:28] Warning: expandable_segments not supported on this platform (function operator()) > initialized tensor model parallel with size 1 > initialized pipeline model parallel with size 1 > setting random seeds to 42 ... > compiling dataset index builder ... make: Entering directory '/workspace/FLAME-MoE/Megatron-LM/megatron/core/datasets' make: Nothing to be done for 'default'. make: Leaving directory '/workspace/FLAME-MoE/Megatron-LM/megatron/core/datasets' >>> done with dataset index builder. Compilation time: 1.119 seconds > compiling and loading fused kernels ... >>> done with compiling and loading fused kernels. Compilation time: 0.343 seconds building GPT model ... > number of parameters on (tensor, pipeline) model parallel rank (0, 0): 100670208 /workspace/FLAME-MoE/Megatron-LM/megatron/core/transformer/transformer_layer.py:339: UserWarning: TransformerLayer._get_layer_offset is deprecated.Please use get_transformer_layer_offset instead. warnings.warn( loading distributed checkpoint from /workspace/FLAME-MoE/results/phase0/runs/flame38m_g1_temporal_s3/ckpt at iteration 2121 /workspace/FLAME-MoE/Megatron-LM/megatron/core/dist_checkpointing/strategies/torch.py:847: FutureWarning: `load_state_dict` is deprecated and will be removed in future versions. Please use `load` instead. checkpoint.load_state_dict( /usr/local/lib/python3.11/dist-packages/torch/distributed/checkpoint/planner_helpers.py:306: FutureWarning: Please use DTensor instead and we are deprecating ShardedTensor. device = getattr(value, "device", None) /workspace/FLAME-MoE/.venv/lib/python3.11/site-packages/transformer_engine/pytorch/module/base.py:566: FutureWarning: You are using `torch.load` with `weights_only=False` (the current default value), which uses the default pickle module implicitly. It is possible to construct malicious pickle data which will execute arbitrary code during unpickling (See https://github.com/pytorch/pytorch/blob/main/SECURITY.md#untrusted-models for more details). In a future release, the default value for `weights_only` will be flipped to `True`. This limits the functions that could be executed during unpickling. Arbitrary objects will no longer be allowed to be loaded via this mode unless they are explicitly allowlisted by the user via `torch.serialization.add_safe_globals`. We recommend you start setting `weights_only=True` for any use case where you don't have full control of the loaded file. Please open an issue on GitHub for any issues related to this experimental feature. state = torch.load(state, map_location="cuda") checkpoint version 3.0 2026-07-23:03:24:26,681 WARNING [rerun_state_machine.py:811] RerunStateMachine disabled via CLI, ignoring machine state saved in checkpoint successfully loaded checkpoint from /workspace/FLAME-MoE/results/phase0/runs/flame38m_g1_temporal_s3/ckpt [ t 1/1, p 1/1 ] at iteration 2121 2026-07-23:03:24:26,811 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc "HTTP/1.1 200 OK" 2026-07-23:03:24:27,052 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc "HTTP/1.1 200 OK" 2026-07-23:03:24:27,117 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc/revision/210d026faf9955653af8916fad021475a3f00453 "HTTP/1.1 200 OK" 2026-07-23:03:24:27,173 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc/tree/210d026faf9955653af8916fad021475a3f00453?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:24:27,466 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc/tree/210d026faf9955653af8916fad021475a3f00453/ARC-Challenge?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:24:28,030 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc "HTTP/1.1 200 OK" 2026-07-23:03:24:28,291 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc "HTTP/1.1 200 OK" 2026-07-23:03:24:28,594 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc/tree/210d026faf9955653af8916fad021475a3f00453/ARC-Easy?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:24:28,739 WARNING [task.py:799] [Task: boolq] metric acc is defined, but aggregation is not. using default aggregation=mean 2026-07-23:03:24:28,739 WARNING [task.py:811] [Task: boolq] metric acc is defined, but higher_is_better is not. using default higher_is_better=True 2026-07-23:03:24:28,809 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue "HTTP/1.1 200 OK" 2026-07-23:03:24:29,046 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue "HTTP/1.1 200 OK" 2026-07-23:03:24:29,107 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/revision/3de24cf8022e94f4ee4b9d55a6f539891524d646 "HTTP/1.1 200 OK" 2026-07-23:03:24:29,173 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/tree/3de24cf8022e94f4ee4b9d55a6f539891524d646?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:24:29,460 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/tree/3de24cf8022e94f4ee4b9d55a6f539891524d646/axb?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:24:29,519 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/revision/3de24cf8022e94f4ee4b9d55a6f539891524d646 "HTTP/1.1 200 OK" 2026-07-23:03:24:29,596 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/tree/3de24cf8022e94f4ee4b9d55a6f539891524d646/boolq?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:24:29,858 WARNING [task.py:799] [Task: copa] metric acc is defined, but aggregation is not. using default aggregation=mean 2026-07-23:03:24:29,858 WARNING [task.py:811] [Task: copa] metric acc is defined, but higher_is_better is not. using default higher_is_better=True 2026-07-23:03:24:29,917 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue "HTTP/1.1 200 OK" 2026-07-23:03:24:30,158 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue "HTTP/1.1 200 OK" 2026-07-23:03:24:30,462 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/tree/3de24cf8022e94f4ee4b9d55a6f539891524d646/copa?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:24:30,642 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag "HTTP/1.1 200 OK" 2026-07-23:03:24:30,878 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag "HTTP/1.1 200 OK" 2026-07-23:03:24:30,948 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag/revision/218ec52e09a7e7462a5400043bb9a69a41d06b76 "HTTP/1.1 200 OK" 2026-07-23:03:24:31,016 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag/tree/218ec52e09a7e7462a5400043bb9a69a41d06b76?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:24:31,220 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag/tree/218ec52e09a7e7462a5400043bb9a69a41d06b76/data?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:24:31,287 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag/revision/218ec52e09a7e7462a5400043bb9a69a41d06b76 "HTTP/1.1 200 OK" 2026-07-23:03:24:33,779 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai "HTTP/1.1 200 OK" 2026-07-23:03:24:34,023 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai "HTTP/1.1 200 OK" 2026-07-23:03:24:34,087 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/revision/900124bf3b8235c6daf21033af9948b3f07346c4 "HTTP/1.1 200 OK" 2026-07-23:03:24:34,154 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/tree/900124bf3b8235c6daf21033af9948b3f07346c4?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:24:34,384 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/tree/900124bf3b8235c6daf21033af9948b3f07346c4/default%2Ftest?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:24:34,450 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/tree/900124bf3b8235c6daf21033af9948b3f07346c4/default?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:24:34,533 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/revision/900124bf3b8235c6daf21033af9948b3f07346c4 "HTTP/1.1 200 OK" 2026-07-23:03:24:34,566 WARNING [task.py:325] [Task: lambada_openai] has_training_docs and has_validation_docs are False, using test_docs as fewshot_docs but this is not recommended. 2026-07-23:03:24:34,567 WARNING [task.py:325] [Task: lambada_openai] has_training_docs and has_validation_docs are False, using test_docs as fewshot_docs but this is not recommended. 2026-07-23:03:24:34,722 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa "HTTP/1.1 200 OK" 2026-07-23:03:24:34,957 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa "HTTP/1.1 200 OK" 2026-07-23:03:24:35,022 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/revision/388097ea7776314e93a529163e0fea805b8a6454 "HTTP/1.1 200 OK" 2026-07-23:03:24:35,080 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/tree/388097ea7776314e93a529163e0fea805b8a6454?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:24:35,347 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/tree/388097ea7776314e93a529163e0fea805b8a6454/additional?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:24:35,404 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/revision/388097ea7776314e93a529163e0fea805b8a6454 "HTTP/1.1 200 OK" 2026-07-23:03:24:35,488 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/tree/388097ea7776314e93a529163e0fea805b8a6454/main?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:24:35,778 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa "HTTP/1.1 200 OK" 2026-07-23:03:24:35,841 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa "HTTP/1.1 200 OK" 2026-07-23:03:24:35,901 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa/tree/main?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:24:35,963 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/datasets/ybisk/piqa/resolve/main/piqa.py "HTTP/1.1 307 Temporary Redirect" 2026-07-23:03:24:35,976 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/resolve-cache/datasets/ybisk/piqa/2e8ac2dffd59bac8c3c6714948f4c551a0848bb0/piqa.py?%2Fdatasets%2Fybisk%2Fpiqa%2Fresolve%2Fmain%2Fpiqa.py=&etag=%220f0d96172f9bf7aa10069374b3236ef083cfeb8e%22 "HTTP/1.1 200 OK" 2026-07-23:03:24:36,551 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa/revision/main "HTTP/1.1 200 OK" 2026-07-23:03:24:36,613 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa/tree/main?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:24:37,132 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq "HTTP/1.1 200 OK" 2026-07-23:03:24:37,360 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq "HTTP/1.1 200 OK" 2026-07-23:03:24:37,423 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq/revision/2c94ad3e1aafab77146f384e23536f97a4849815 "HTTP/1.1 200 OK" 2026-07-23:03:24:37,488 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq/tree/2c94ad3e1aafab77146f384e23536f97a4849815?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:24:37,736 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq/tree/2c94ad3e1aafab77146f384e23536f97a4849815/data?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:24:37,814 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq/revision/2c94ad3e1aafab77146f384e23536f97a4849815 "HTTP/1.1 200 OK" 2026-07-23:03:24:38,232 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande "HTTP/1.1 200 OK" 2026-07-23:03:24:38,473 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande "HTTP/1.1 200 OK" 2026-07-23:03:24:38,538 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/revision/01e74176c63542e6b0bcb004dcdea22d94fb67b5 "HTTP/1.1 200 OK" 2026-07-23:03:24:38,624 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/tree/01e74176c63542e6b0bcb004dcdea22d94fb67b5?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:24:38,868 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/tree/01e74176c63542e6b0bcb004dcdea22d94fb67b5/winogrande_debiased?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:24:38,931 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/revision/01e74176c63542e6b0bcb004dcdea22d94fb67b5 "HTTP/1.1 200 OK" 2026-07-23:03:24:39,025 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/tree/01e74176c63542e6b0bcb004dcdea22d94fb67b5/winogrande_xl?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:24:39,681 WARNING [evaluator.py:270] Overwriting default num_fewshot of winogrande from None to 0 2026-07-23:03:24:39,681 WARNING [evaluator.py:270] Overwriting default num_fewshot of sciq from None to 0 2026-07-23:03:24:39,681 WARNING [evaluator.py:270] Overwriting default num_fewshot of piqa from None to 0 2026-07-23:03:24:39,681 WARNING [evaluator.py:270] Overwriting default num_fewshot of openbookqa from None to 0 2026-07-23:03:24:39,681 WARNING [evaluator.py:270] Overwriting default num_fewshot of lambada_openai from None to 0 2026-07-23:03:24:39,681 WARNING [evaluator.py:270] Overwriting default num_fewshot of hellaswag from None to 0 2026-07-23:03:24:39,681 WARNING [evaluator.py:270] Overwriting default num_fewshot of copa from None to 0 2026-07-23:03:24:39,681 WARNING [evaluator.py:270] Overwriting default num_fewshot of boolq from None to 0 2026-07-23:03:24:39,681 WARNING [evaluator.py:270] Overwriting default num_fewshot of arc_easy from None to 0 2026-07-23:03:24:39,682 WARNING [evaluator.py:270] Overwriting default num_fewshot of arc_challenge from None to 0 2026-07-23:03:24:39,726 INFO [task.py:415] Building contexts for winogrande on rank 0... 0%| | 0/1267 [00:00 /workspace/FLAME-MoE/results/ablations/flame1e18_downstream.csv ==================================================================== [run] flame38m_g3_moe mode=moe TEMPORAL=0 H=256 heads=16 ffn=1368 experts=192 topk=18 moe_ffn=58 shared=352 [run] tasks=arc_challenge,arc_easy,hellaswag,openbookqa,piqa,winogrande,boolq,copa,lambada_openai,sciq out=/workspace/FLAME-MoE/results/phase0/runs/flame38m_g3_moe/lmeval_1e18_0shot ==================================================================== /usr/local/lib/python3.11/dist-packages/torch/cuda/__init__.py:58: FutureWarning: The pynvml package is deprecated. Please install nvidia-ml-py instead. If you did not install pynvml directly, please report this to the maintainers of the package that installed pynvml for you. import pynvml # type: ignore[import] [run_lmeval] stubbed VLM backends (hf_vlms, vllm_vlms) /usr/local/lib/python3.11/dist-packages/torch/cuda/__init__.py:58: FutureWarning: The pynvml package is deprecated. Please install nvidia-ml-py instead. If you did not install pynvml directly, please report this to the maintainers of the package that installed pynvml for you. import pynvml # type: ignore[import] /usr/local/lib/python3.11/dist-packages/requests/__init__.py:113: RequestsDependencyWarning: urllib3 (2.2.3) or chardet (6.0.0.post1)/charset_normalizer (3.3.2) doesn't match a supported version! warnings.warn( /usr/local/lib/python3.11/dist-packages/torch/cuda/__init__.py:58: FutureWarning: The pynvml package is deprecated. Please install nvidia-ml-py instead. If you did not install pynvml directly, please report this to the maintainers of the package that installed pynvml for you. import pynvml # type: ignore[import] /workspace/FLAME-MoE/.venv/lib/python3.11/site-packages/transformer_engine/pytorch/attention.py:108: UserWarning: To use flash-attn v3, please use the following commands to install: (1) pip install "git+https://github.com/Dao-AILab/flash-attention.git#egg=flashattn-hopper&subdirectory=hopper" (2) python_path=`python -c "import site; print(site.getsitepackages()[0])"` (3) mkdir -p $python_path/flashattn_hopper (4) wget -P $python_path/flashattn_hopper https://raw.githubusercontent.com/Dao-AILab/flash-attention/main/hopper/flash_attn_interface.py warnings.warn( /workspace/FLAME-MoE/Megatron-LM/megatron/core/models/gpt/gpt_layer_specs.py:54: UserWarning: Apex is not installed. Falling back to Torch Norm warnings.warn('Apex is not installed. Falling back to Torch Norm') /workspace/FLAME-MoE/Megatron-LM/megatron/core/models/retro/encoder_spec.py:47: UserWarning: Apex is not installed. Falling back to Torch Norm warnings.warn(f'Apex is not installed. Falling back to Torch Norm') /workspace/FLAME-MoE/Megatron-LM/megatron/core/models/retro/decoder_spec.py:39: UserWarning: Apex is not installed. Falling back to Torch Norm warnings.warn(f'Apex is not installed. Falling back to Torch Norm') 2026-07-23:03:32:20,889 INFO [__main__.py:279] Verbosity set to INFO 2026-07-23:03:32:51,387 INFO [__init__.py:459] The tag 'arc_ca' is already registered as a group, this tag will not be registered. This may affect tasks you want to call. 2026-07-23:03:32:51,403 INFO [__init__.py:459] The tag 'arc_ca' is already registered as a group, this tag will not be registered. This may affect tasks you want to call. 2026-07-23:03:33:07,162 INFO [__main__.py:376] Selected Tasks: ['arc_challenge', 'arc_easy', 'boolq', 'copa', 'hellaswag', 'lambada_openai', 'openbookqa', 'piqa', 'sciq', 'winogrande'] 2026-07-23:03:33:07,180 INFO [evaluator.py:164] Setting random seed to 42 | Setting numpy seed to 42 | Setting torch manual seed to 42 | Setting fewshot manual seed to 42 2026-07-23:03:33:07,180 INFO [evaluator.py:201] Initializing megatron_lm model, with arguments: {} using world size: 1, data-parallel size: 1, context-parallel size: 1, hierarchical context-parallel sizes: Nonetensor-model-parallel size: 1, encoder-tensor-model-parallel size: 0, pipeline-model-parallel size: 1, encoder-pipeline-model-parallel size: 0 WARNING: overriding default arguments for exit_on_missing_checkpoint:True with exit_on_missing_checkpoint:False setting global batch size to 32 Number of virtual stages per pipeline stage: None accumulate and all-reduce gradients in fp32 for bfloat16 data type. using torch.bfloat16 for parameters ... ------------------------ arguments ------------------------ account_for_embedding_in_pipeline_split ......... False account_for_loss_in_pipeline_split .............. False accumulate_allreduce_grads_in_fp32 .............. True adam_beta1 ...................................... 0.9 adam_beta2 ...................................... 0.999 adam_eps ........................................ 1e-08 add_bias_linear ................................. False add_position_embedding .......................... True add_qkv_bias .................................... False adlr_autoresume ................................. False adlr_autoresume_interval ........................ 1000 align_grad_reduce ............................... True align_param_gather .............................. False app_tag_run_name ................................ None app_tag_run_version ............................. 0.0.0 apply_layernorm_1p .............................. False apply_query_key_layer_scaling ................... False apply_residual_connection_post_layernorm ........ False apply_rope_fusion ............................... True async_save ...................................... None async_tensor_model_parallel_allreduce ........... True attention_backend ............................... AttnBackend.auto attention_dropout ............................... 0.0 attention_softmax_in_fp32 ....................... False auto_detect_ckpt_format ......................... False barrier_with_L1_time ............................ True bert_binary_head ................................ True bert_embedder_type .............................. megatron bert_load ....................................... None bf16 ............................................ True bias_dropout_fusion ............................. True bias_gelu_fusion ................................ False bias_swiglu_fusion .............................. True biencoder_projection_dim ........................ 0 biencoder_shared_query_context_model ............ False block_data_path ................................. None calc_ft_timeouts ................................ False calculate_per_token_loss ........................ False check_for_large_grads ........................... False check_for_nan_in_loss_and_grad .................. True check_for_spiky_loss ............................ False check_weight_hash_across_dp_replicas_interval ... None ckpt_assume_constant_structure .................. False ckpt_convert_format ............................. None ckpt_convert_save ............................... None ckpt_convert_update_legacy_dist_opt_format ...... False ckpt_format ..................................... torch_dist ckpt_fully_parallel_load ........................ False ckpt_fully_parallel_save ........................ True ckpt_fully_parallel_save_deprecated ............. False ckpt_step ....................................... None classes_fraction ................................ 1.0 clip_grad ....................................... 1.0 clone_scatter_output_in_embedding ............... True config_logger_dir ............................... consumed_train_samples .......................... 0 consumed_valid_samples .......................... 0 context_parallel_size ........................... 1 cp_comm_type .................................... ['p2p'] create_attention_mask_in_dataloader ............. True cross_entropy_loss_fusion ....................... False cuda_graph_warmup_steps ......................... 3 data_args_path .................................. None data_cache_path ................................. None data_parallel_random_init ....................... False data_parallel_sharding_strategy ................. no_shard data_parallel_size .............................. 1 data_path ....................................... None data_per_class_fraction ......................... 1.0 data_sharding ................................... True dataloader_type ................................. single ddp_average_in_collective ....................... False ddp_bucket_size ................................. None ddp_num_buckets ................................. None ddp_pad_buckets_for_high_nccl_busbw ............. False decoder_first_pipeline_num_layers ............... None decoder_last_pipeline_num_layers ................ None decoder_num_layers .............................. None decoder_seq_length .............................. None decoupled_lr .................................... None decoupled_min_lr ................................ None decrease_batch_size_if_needed ................... False defer_embedding_wgrad_compute ................... False deprecated_use_mcore_models ..................... False deterministic_mode .............................. False dino_bottleneck_size ............................ 256 dino_freeze_last_layer .......................... 1 dino_head_hidden_size ........................... 2048 dino_local_crops_number ......................... 10 dino_local_img_size ............................. 96 dino_norm_last_layer ............................ False dino_teacher_temp ............................... 0.07 dino_warmup_teacher_temp ........................ 0.04 dino_warmup_teacher_temp_epochs ................. 30 disable_straggler_on_startup .................... False dist_ckpt_format_deprecated ..................... None dist_ckpt_strictness ............................ assume_ok_unexpected distribute_saved_activations .................... False distributed_backend ............................. nccl distributed_timeout_minutes ..................... 10 embedding_path .................................. None empty_unused_memory_level ....................... 0 enable_cuda_graph ............................... False enable_ft_package ............................... False enable_gloo_process_groups ...................... True enable_one_logger ............................... True encoder_num_layers .............................. 9 encoder_pipeline_model_parallel_size ............ 0 encoder_seq_length .............................. 2048 encoder_tensor_model_parallel_size .............. 0 end_weight_decay ................................ 0.01 eod_mask_loss ................................... False error_injection_rate ............................ 0 error_injection_type ............................ transient_error eval_interval ................................... 1000 eval_iters ...................................... 100 evidence_data_path .............................. None exit_duration_in_mins ........................... None exit_interval ................................... None exit_on_missing_checkpoint ...................... False exit_signal_handler ............................. False exp_avg_dtype ................................... torch.float32 exp_avg_sq_dtype ................................ torch.float32 expert_model_parallel_size ...................... 1 expert_tensor_parallel_size ..................... 1 ffn_hidden_size ................................. 1368 finetune ........................................ False flash_decode .................................... False fp16 ............................................ False fp16_lm_cross_entropy ........................... False fp32_residual_connection ........................ False fp8 ............................................. None fp8_amax_compute_algo ........................... most_recent fp8_amax_history_len ............................ 1 fp8_interval .................................... 1 fp8_margin ...................................... 0 fp8_param_gather ................................ False fp8_wgrad ....................................... True global_batch_size ............................... 32 grad_reduce_in_bf16 ............................. False gradient_accumulation_fusion .................... False gradient_reduce_div_fusion ...................... True group_query_attention ........................... False head_lr_mult .................................... 1.0 hidden_dropout .................................. 0.0 hidden_size ..................................... 256 hierarchical_context_parallel_sizes ............. None hybrid_attention_ratio .......................... 0.0 hybrid_mlp_ratio ................................ 0.0 hybrid_override_pattern ......................... None hysteresis ...................................... 2 ict_head_size ................................... None ict_load ........................................ None img_h ........................................... 224 img_w ........................................... 224 indexer_batch_size .............................. 128 indexer_log_interval ............................ 1000 inference_batch_times_seqlen_threshold .......... -1 inference_max_batch_size ........................ 8 inference_max_seq_length ........................ 2560 inference_rng_tracker ........................... False init_method_std ................................. 0.02 init_method_xavier_uniform ...................... False init_model_with_meta_device ..................... False initial_loss_scale .............................. 4294967296 iter_per_epoch .................................. 1250 iterations_to_skip .............................. [] keep_fp8_transpose_cache_when_using_custom_fsdp . False kv_channels ..................................... 16 kv_lora_rank .................................... 32 lazy_mpu_init ................................... None load ............................................ /workspace/FLAME-MoE/results/phase0/runs/flame38m_g3_moe/ckpt local_rank ...................................... 0 log_interval .................................... 100 log_loss_scale_to_tensorboard ................... True log_memory_to_tensorboard ....................... False log_num_zeros_in_grad ........................... False log_params_norm ................................. False log_progress .................................... False log_straggler ................................... False log_throughput .................................. False log_timers_to_tensorboard ....................... False log_validation_ppl_to_tensorboard ............... False log_world_size_to_tensorboard ................... False logging_level ................................... None loss_scale ...................................... None loss_scale_window ............................... 1000 lr .............................................. None lr_decay_iters .................................. None lr_decay_samples ................................ None lr_decay_style .................................. linear lr_warmup_fraction .............................. None lr_warmup_init .................................. 0.0 lr_warmup_iters ................................. 0 lr_warmup_samples ............................... 0 lr_wsd_decay_iters .............................. None lr_wsd_decay_samples ............................ None lr_wsd_decay_style .............................. exponential main_grads_dtype ................................ torch.float32 main_params_dtype ............................... torch.float32 make_vocab_size_divisible_by .................... 128 manual_gc ....................................... False manual_gc_eval .................................. True manual_gc_interval .............................. 0 mask_factor ..................................... 1.0 mask_prob ....................................... 0.15 mask_type ....................................... random masked_softmax_fusion ........................... True max_position_embeddings ......................... 2048 max_tokens_to_oom ............................... 10000000 memory_snapshot_path ............................ snapshot.pickle merge_file ...................................... None micro_batch_size ................................ 32 microbatch_group_size_per_vp_stage .............. None min_loss_scale .................................. 1.0 min_lr .......................................... 0.0 mmap_bin_files .................................. True mock_data ....................................... False moe_aux_loss_coeff .............................. 0.01 moe_enable_deepep ............................... False moe_expert_capacity_factor ...................... None moe_extended_tp ................................. False moe_ffn_hidden_size ............................. 58 moe_grouped_gemm ................................ True moe_input_jitter_eps ............................ None moe_layer_freq .................................. [0, 1, 1, 1, 1, 1, 1, 1, 1] moe_layer_recompute ............................. False moe_pad_expert_input_to_capacity ................ False moe_per_layer_logging ........................... False moe_permute_fusion .............................. False moe_router_bias_update_rate ..................... 0.001 moe_router_dtype ................................ fp32 moe_router_enable_expert_bias ................... False moe_router_group_topk ........................... None moe_router_load_balancing_type .................. aux_loss moe_router_num_groups ........................... None moe_router_pre_softmax .......................... True moe_router_score_function ....................... softmax moe_router_topk ................................. 18 moe_router_topk_scaling_factor .................. None moe_shared_expert_intermediate_size ............. 352 moe_shared_expert_overlap ....................... False moe_token_dispatcher_type ....................... alltoall moe_token_drop_policy ........................... probs moe_use_legacy_grouped_gemm ..................... False moe_use_upcycling ............................... False moe_z_loss_coeff ................................ 0.001 mscale .......................................... 1.0 mscale_all_dim .................................. 1.0 multi_latent_attention .......................... False nccl_communicator_config_path ................... None no_load_optim ................................... True no_load_rng ..................................... True no_persist_layer_norm ........................... False no_save_optim ................................... None no_save_rng ..................................... None non_persistent_ckpt_type ........................ None non_persistent_global_ckpt_dir .................. None non_persistent_local_ckpt_algo .................. fully_parallel non_persistent_local_ckpt_dir ................... None non_persistent_save_interval .................... None norm_epsilon .................................... 1e-06 normalization ................................... RMSNorm num_attention_heads ............................. 16 num_channels .................................... 3 num_classes ..................................... 1000 num_dataset_builder_threads ..................... 1 num_distributed_optimizer_instances ............. 1 num_experts ..................................... 192 num_layers ...................................... 9 num_layers_per_virtual_pipeline_stage ........... None num_query_groups ................................ 1 num_virtual_stages_per_pipeline_rank ............ None num_workers ..................................... 2 one_logger_async ................................ False one_logger_project .............................. megatron-lm one_logger_run_name ............................. None onnx_safe ....................................... None openai_gelu ..................................... False optimizer ....................................... adam optimizer_cpu_offload ........................... False optimizer_offload_fraction ...................... 1.0 output_bert_embeddings .......................... False overlap_cpu_optimizer_d2h_h2d ................... False overlap_grad_reduce ............................. False overlap_p2p_comm ................................ False overlap_p2p_comm_warmup_flush ................... False overlap_param_gather ............................ False overlap_param_gather_with_optimizer_step ........ False override_opt_param_scheduler .................... False params_dtype .................................... torch.bfloat16 patch_dim ....................................... 16 per_split_data_args_path ........................ None perform_initialization .......................... True pin_cpu_grads ................................... True pin_cpu_params .................................. True pipeline_model_parallel_comm_backend ............ None pipeline_model_parallel_size .................... 1 pipeline_model_parallel_split_rank .............. None position_embedding_type ......................... rope pretrained_checkpoint ........................... None profile ......................................... False profile_ranks ................................... [0] profile_step_end ................................ 12 profile_step_start .............................. 10 q_lora_rank ..................................... None qk_head_dim ..................................... 128 qk_layernorm .................................... False qk_pos_emb_head_dim ............................. 64 query_in_block_prob ............................. 0.1 rampup_batch_size ............................... None rank ............................................ 0 recompute_granularity ........................... None recompute_method ................................ None recompute_num_layers ............................ None record_memory_history ........................... False relative_attention_max_distance ................. 128 relative_attention_num_buckets .................. 32 replication ..................................... False replication_factor .............................. 2 replication_jump ................................ None rerun_mode ...................................... disabled reset_attention_mask ............................ False reset_position_ids .............................. False result_rejected_tracker_filename ................ None retriever_report_topk_accuracies ................ [] retriever_score_scaling ......................... False retriever_seq_length ............................ 256 retro_add_retriever ............................. False retro_attention_gate ............................ 1 retro_cyclic_train_iters ........................ None retro_encoder_attention_dropout ................. 0.1 retro_encoder_hidden_dropout .................... 0.1 retro_encoder_layers ............................ 2 retro_num_neighbors ............................. 2 retro_num_retrieved_chunks ...................... 2 retro_project_dir ............................... None retro_verify_neighbor_count ..................... True rope_scaling_factor ............................. 8.0 rotary_base ..................................... 10000 rotary_interleaved .............................. False rotary_percent .................................. 1.0 rotary_scaling_factor ........................... 1.0 rotary_seq_len_interpolation_factor ............. None s3_cache_path ................................... None sample_rate ..................................... 1.0 save ............................................ None save_interval ................................... None scatter_gather_tensors_in_pipeline .............. True seed ............................................ 42 seq_length ...................................... 2048 sequence_parallel ............................... False sgd_momentum .................................... 0.9 short_seq_prob .................................. 0.1 skip_train ...................................... False skipped_train_samples ........................... 0 spec ............................................ None split ........................................... None squared_relu .................................... False start_weight_decay .............................. 0.01 straggler_ctrlr_port ............................ 65535 straggler_minmax_count .......................... 1 suggested_communication_unit_size ............... 400000000 swiglu .......................................... True swin_backbone_type .............................. tiny te_rng_tracker .................................. False tensor_model_parallel_size ...................... 1 tensorboard_dir ................................. None tensorboard_log_interval ........................ 1 tensorboard_queue_size .......................... 1000 test_data_path .................................. None test_mode ....................................... False tiktoken_num_special_tokens ..................... 1000 tiktoken_pattern ................................ None tiktoken_special_tokens ......................... None timing_log_level ................................ 0 timing_log_option ............................... minmax titles_data_path ................................ None tokenizer_model ................................. EleutherAI/pythia-12b tokenizer_type .................................. HuggingFaceTokenizer tp_comm_bootstrap_backend ....................... nccl tp_comm_bulk_dgrad .............................. True tp_comm_bulk_wgrad .............................. True tp_comm_overlap ................................. False tp_comm_overlap_ag .............................. True tp_comm_overlap_cfg ............................. None tp_comm_overlap_rs .............................. True tp_comm_overlap_rs_dgrad ........................ False tp_comm_split_ag ................................ True tp_comm_split_rs ................................ True train_data_path ................................. None train_iters ..................................... None train_samples ................................... None train_sync_interval ............................. None transformer_impl ................................ transformer_engine transformer_pipeline_model_parallel_size ........ 1 untie_embeddings_and_output_weights ............. True use_checkpoint_args ............................. False use_checkpoint_opt_param_scheduler .............. False use_cpu_initialization .......................... None use_custom_fsdp ................................. False use_dist_ckpt ................................... True use_dist_ckpt_deprecated ........................ False use_distributed_optimizer ....................... True use_flash_attn .................................. False use_legacy_models ............................... False use_mp_args_from_checkpoint_args ................ False use_one_sent_docs ............................... False use_persistent_ckpt_worker ...................... False use_precision_aware_optimizer ................... False use_pytorch_profiler ............................ False use_ring_exchange_p2p ........................... False use_rope_scaling ................................ False use_rotary_position_embeddings .................. False use_tokenizer_model_from_checkpoint_args ........ True use_torch_fsdp2 ................................. False use_torch_optimizer_for_cpu_offload ............. False use_tp_pp_dp_mapping ............................ False v_head_dim ...................................... 128 valid_data_path ................................. None variable_seq_lengths ............................ False virtual_pipeline_model_parallel_size ............ None vision_backbone_type ............................ vit vision_pretraining .............................. False vision_pretraining_type ......................... classify vocab_extra_ids ................................. 0 vocab_file ...................................... None vocab_size ...................................... None wandb_exp_name .................................. wandb_project ................................... wandb_save_dir .................................. weight_decay .................................... 0.01 weight_decay_incr_style ......................... constant wgrad_deferral_limit ............................ 0 world_size ...................................... 1 yaml_cfg ........................................ None -------------------- end of arguments --------------------- 2026-07-23:03:33:07,195 INFO [num_microbatches_calculator.py:228] setting number of microbatches to constant 1 > building HuggingFaceTokenizer tokenizer ... 2026-07-23:03:33:07,396 INFO [_client.py:1038] HTTP Request: HEAD https://huggingface.co/EleutherAI/pythia-12b/resolve/main/config.json "HTTP/1.1 307 Temporary Redirect" 2026-07-23:03:33:07,408 INFO [_client.py:1038] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/EleutherAI/pythia-12b/bb1e3e710cdf6b524461d543cfb5ba773f0a81b6/config.json "HTTP/1.1 200 OK" 2026-07-23:03:33:07,520 INFO [_client.py:1038] HTTP Request: HEAD https://huggingface.co/EleutherAI/pythia-12b/resolve/main/tokenizer_config.json "HTTP/1.1 307 Temporary Redirect" 2026-07-23:03:33:07,532 INFO [_client.py:1038] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/EleutherAI/pythia-12b/bb1e3e710cdf6b524461d543cfb5ba773f0a81b6/tokenizer_config.json "HTTP/1.1 200 OK" 2026-07-23:03:33:07,632 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/models/EleutherAI/pythia-12b/tree/main/additional_chat_templates?recursive=false&expand=false "HTTP/1.1 404 Not Found" 2026-07-23:03:33:07,701 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/models/EleutherAI/pythia-12b/tree/main?recursive=true&expand=false "HTTP/1.1 200 OK" > padded vocab (size: 50277) with 27 dummy tokens (new size: 50304) WARNING: one_logger package is required to enable e2e metrics tracking. please go to https://confluence.nvidia.com/display/MLWFO/Package+Repositories for details to install it 2026-07-23:03:33:07,894 WARNING [rerun_state_machine.py:239] RerunStateMachine initialized in mode RerunMode.DISABLED > initializing torch distributed ... [W723 03:33:07.411571411 CUDAAllocatorConfig.h:28] Warning: expandable_segments not supported on this platform (function operator()) > initialized tensor model parallel with size 1 > initialized pipeline model parallel with size 1 > setting random seeds to 42 ... > compiling dataset index builder ... make: Entering directory '/workspace/FLAME-MoE/Megatron-LM/megatron/core/datasets' make: Nothing to be done for 'default'. make: Leaving directory '/workspace/FLAME-MoE/Megatron-LM/megatron/core/datasets' >>> done with dataset index builder. Compilation time: 0.661 seconds > compiling and loading fused kernels ... >>> done with compiling and loading fused kernels. Compilation time: 0.334 seconds building GPT model ... > number of parameters on (tensor, pipeline) model parallel rank (0, 0): 100145920 /workspace/FLAME-MoE/Megatron-LM/megatron/core/transformer/transformer_layer.py:339: UserWarning: TransformerLayer._get_layer_offset is deprecated.Please use get_transformer_layer_offset instead. warnings.warn( loading distributed checkpoint from /workspace/FLAME-MoE/results/phase0/runs/flame38m_g3_moe/ckpt at iteration 2121 /workspace/FLAME-MoE/Megatron-LM/megatron/core/dist_checkpointing/strategies/torch.py:847: FutureWarning: `load_state_dict` is deprecated and will be removed in future versions. Please use `load` instead. checkpoint.load_state_dict( /usr/local/lib/python3.11/dist-packages/torch/distributed/checkpoint/planner_helpers.py:306: FutureWarning: Please use DTensor instead and we are deprecating ShardedTensor. device = getattr(value, "device", None) /workspace/FLAME-MoE/.venv/lib/python3.11/site-packages/transformer_engine/pytorch/module/base.py:566: FutureWarning: You are using `torch.load` with `weights_only=False` (the current default value), which uses the default pickle module implicitly. It is possible to construct malicious pickle data which will execute arbitrary code during unpickling (See https://github.com/pytorch/pytorch/blob/main/SECURITY.md#untrusted-models for more details). In a future release, the default value for `weights_only` will be flipped to `True`. This limits the functions that could be executed during unpickling. Arbitrary objects will no longer be allowed to be loaded via this mode unless they are explicitly allowlisted by the user via `torch.serialization.add_safe_globals`. We recommend you start setting `weights_only=True` for any use case where you don't have full control of the loaded file. Please open an issue on GitHub for any issues related to this experimental feature. state = torch.load(state, map_location="cuda") checkpoint version 3.0 2026-07-23:03:33:14,071 WARNING [rerun_state_machine.py:811] RerunStateMachine disabled via CLI, ignoring machine state saved in checkpoint successfully loaded checkpoint from /workspace/FLAME-MoE/results/phase0/runs/flame38m_g3_moe/ckpt [ t 1/1, p 1/1 ] at iteration 2121 2026-07-23:03:33:14,210 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc "HTTP/1.1 200 OK" 2026-07-23:03:33:14,460 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc "HTTP/1.1 200 OK" 2026-07-23:03:33:14,528 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc/revision/210d026faf9955653af8916fad021475a3f00453 "HTTP/1.1 200 OK" 2026-07-23:03:33:14,594 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc/tree/210d026faf9955653af8916fad021475a3f00453?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:33:14,845 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc/tree/210d026faf9955653af8916fad021475a3f00453/ARC-Challenge?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:33:15,367 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc "HTTP/1.1 200 OK" 2026-07-23:03:33:15,584 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc "HTTP/1.1 200 OK" 2026-07-23:03:33:15,870 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc/tree/210d026faf9955653af8916fad021475a3f00453/ARC-Easy?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:33:16,010 WARNING [task.py:799] [Task: boolq] metric acc is defined, but aggregation is not. using default aggregation=mean 2026-07-23:03:33:16,011 WARNING [task.py:811] [Task: boolq] metric acc is defined, but higher_is_better is not. using default higher_is_better=True 2026-07-23:03:33:16,108 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue "HTTP/1.1 200 OK" 2026-07-23:03:33:16,339 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue "HTTP/1.1 200 OK" 2026-07-23:03:33:16,406 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/revision/3de24cf8022e94f4ee4b9d55a6f539891524d646 "HTTP/1.1 200 OK" 2026-07-23:03:33:16,467 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/tree/3de24cf8022e94f4ee4b9d55a6f539891524d646?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:33:16,784 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/tree/3de24cf8022e94f4ee4b9d55a6f539891524d646/axb?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:33:16,846 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/revision/3de24cf8022e94f4ee4b9d55a6f539891524d646 "HTTP/1.1 200 OK" 2026-07-23:03:33:16,920 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/tree/3de24cf8022e94f4ee4b9d55a6f539891524d646/boolq?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:33:17,166 WARNING [task.py:799] [Task: copa] metric acc is defined, but aggregation is not. using default aggregation=mean 2026-07-23:03:33:17,166 WARNING [task.py:811] [Task: copa] metric acc is defined, but higher_is_better is not. using default higher_is_better=True 2026-07-23:03:33:17,233 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue "HTTP/1.1 200 OK" 2026-07-23:03:33:17,497 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue "HTTP/1.1 200 OK" 2026-07-23:03:33:18,147 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/tree/3de24cf8022e94f4ee4b9d55a6f539891524d646/copa?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:33:18,294 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag "HTTP/1.1 200 OK" 2026-07-23:03:33:18,532 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag "HTTP/1.1 200 OK" 2026-07-23:03:33:18,590 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag/revision/218ec52e09a7e7462a5400043bb9a69a41d06b76 "HTTP/1.1 200 OK" 2026-07-23:03:33:18,652 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag/tree/218ec52e09a7e7462a5400043bb9a69a41d06b76?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:33:18,901 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag/tree/218ec52e09a7e7462a5400043bb9a69a41d06b76/data?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:33:18,962 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag/revision/218ec52e09a7e7462a5400043bb9a69a41d06b76 "HTTP/1.1 200 OK" 2026-07-23:03:33:20,979 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai "HTTP/1.1 200 OK" 2026-07-23:03:33:21,203 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai "HTTP/1.1 200 OK" 2026-07-23:03:33:21,264 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/revision/900124bf3b8235c6daf21033af9948b3f07346c4 "HTTP/1.1 200 OK" 2026-07-23:03:33:21,363 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/tree/900124bf3b8235c6daf21033af9948b3f07346c4?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:33:21,588 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/tree/900124bf3b8235c6daf21033af9948b3f07346c4/default%2Ftest?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:33:21,656 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/tree/900124bf3b8235c6daf21033af9948b3f07346c4/default?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:33:21,714 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/revision/900124bf3b8235c6daf21033af9948b3f07346c4 "HTTP/1.1 200 OK" 2026-07-23:03:33:21,734 WARNING [task.py:325] [Task: lambada_openai] has_training_docs and has_validation_docs are False, using test_docs as fewshot_docs but this is not recommended. 2026-07-23:03:33:21,734 WARNING [task.py:325] [Task: lambada_openai] has_training_docs and has_validation_docs are False, using test_docs as fewshot_docs but this is not recommended. 2026-07-23:03:33:21,895 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa "HTTP/1.1 200 OK" 2026-07-23:03:33:22,136 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa "HTTP/1.1 200 OK" 2026-07-23:03:33:22,202 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/revision/388097ea7776314e93a529163e0fea805b8a6454 "HTTP/1.1 200 OK" 2026-07-23:03:33:22,265 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/tree/388097ea7776314e93a529163e0fea805b8a6454?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:33:22,539 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/tree/388097ea7776314e93a529163e0fea805b8a6454/additional?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:33:22,604 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/revision/388097ea7776314e93a529163e0fea805b8a6454 "HTTP/1.1 200 OK" 2026-07-23:03:33:22,691 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/tree/388097ea7776314e93a529163e0fea805b8a6454/main?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:33:22,959 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa "HTTP/1.1 200 OK" 2026-07-23:03:33:23,018 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa "HTTP/1.1 200 OK" 2026-07-23:03:33:23,081 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa/tree/main?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:33:23,146 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/datasets/ybisk/piqa/resolve/main/piqa.py "HTTP/1.1 307 Temporary Redirect" 2026-07-23:03:33:23,159 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/resolve-cache/datasets/ybisk/piqa/2e8ac2dffd59bac8c3c6714948f4c551a0848bb0/piqa.py?%2Fdatasets%2Fybisk%2Fpiqa%2Fresolve%2Fmain%2Fpiqa.py=&etag=%220f0d96172f9bf7aa10069374b3236ef083cfeb8e%22 "HTTP/1.1 200 OK" 2026-07-23:03:33:23,684 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa/revision/main "HTTP/1.1 200 OK" 2026-07-23:03:33:23,747 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa/tree/main?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:33:24,185 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq "HTTP/1.1 200 OK" 2026-07-23:03:33:24,422 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq "HTTP/1.1 200 OK" 2026-07-23:03:33:24,542 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq/revision/2c94ad3e1aafab77146f384e23536f97a4849815 "HTTP/1.1 200 OK" 2026-07-23:03:33:24,609 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq/tree/2c94ad3e1aafab77146f384e23536f97a4849815?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:33:24,824 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq/tree/2c94ad3e1aafab77146f384e23536f97a4849815/data?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:33:24,885 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq/revision/2c94ad3e1aafab77146f384e23536f97a4849815 "HTTP/1.1 200 OK" 2026-07-23:03:33:25,305 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande "HTTP/1.1 200 OK" 2026-07-23:03:33:25,562 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande "HTTP/1.1 200 OK" 2026-07-23:03:33:25,620 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/revision/01e74176c63542e6b0bcb004dcdea22d94fb67b5 "HTTP/1.1 200 OK" 2026-07-23:03:33:25,688 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/tree/01e74176c63542e6b0bcb004dcdea22d94fb67b5?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:33:25,917 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/tree/01e74176c63542e6b0bcb004dcdea22d94fb67b5/winogrande_debiased?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:33:25,983 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/revision/01e74176c63542e6b0bcb004dcdea22d94fb67b5 "HTTP/1.1 200 OK" 2026-07-23:03:33:26,080 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/tree/01e74176c63542e6b0bcb004dcdea22d94fb67b5/winogrande_xl?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:33:26,714 WARNING [evaluator.py:270] Overwriting default num_fewshot of winogrande from None to 0 2026-07-23:03:33:26,714 WARNING [evaluator.py:270] Overwriting default num_fewshot of sciq from None to 0 2026-07-23:03:33:26,714 WARNING [evaluator.py:270] Overwriting default num_fewshot of piqa from None to 0 2026-07-23:03:33:26,714 WARNING [evaluator.py:270] Overwriting default num_fewshot of openbookqa from None to 0 2026-07-23:03:33:26,714 WARNING [evaluator.py:270] Overwriting default num_fewshot of lambada_openai from None to 0 2026-07-23:03:33:26,714 WARNING [evaluator.py:270] Overwriting default num_fewshot of hellaswag from None to 0 2026-07-23:03:33:26,714 WARNING [evaluator.py:270] Overwriting default num_fewshot of copa from None to 0 2026-07-23:03:33:26,714 WARNING [evaluator.py:270] Overwriting default num_fewshot of boolq from None to 0 2026-07-23:03:33:26,714 WARNING [evaluator.py:270] Overwriting default num_fewshot of arc_easy from None to 0 2026-07-23:03:33:26,714 WARNING [evaluator.py:270] Overwriting default num_fewshot of arc_challenge from None to 0 2026-07-23:03:33:26,745 INFO [task.py:415] Building contexts for winogrande on rank 0... 0%| | 0/1267 [00:00 /workspace/FLAME-MoE/results/ablations/flame1e18_downstream.csv ==================================================================== [run] flame38m_g3_moe_s2 mode=moe TEMPORAL=0 H=256 heads=16 ffn=1368 experts=192 topk=18 moe_ffn=58 shared=352 [run] tasks=arc_challenge,arc_easy,hellaswag,openbookqa,piqa,winogrande,boolq,copa,lambada_openai,sciq out=/workspace/FLAME-MoE/results/phase0/runs/flame38m_g3_moe_s2/lmeval_1e18_0shot ==================================================================== /usr/local/lib/python3.11/dist-packages/torch/cuda/__init__.py:58: FutureWarning: The pynvml package is deprecated. Please install nvidia-ml-py instead. If you did not install pynvml directly, please report this to the maintainers of the package that installed pynvml for you. import pynvml # type: ignore[import] [run_lmeval] stubbed VLM backends (hf_vlms, vllm_vlms) /usr/local/lib/python3.11/dist-packages/torch/cuda/__init__.py:58: FutureWarning: The pynvml package is deprecated. Please install nvidia-ml-py instead. If you did not install pynvml directly, please report this to the maintainers of the package that installed pynvml for you. import pynvml # type: ignore[import] /usr/local/lib/python3.11/dist-packages/requests/__init__.py:113: RequestsDependencyWarning: urllib3 (2.2.3) or chardet (6.0.0.post1)/charset_normalizer (3.3.2) doesn't match a supported version! warnings.warn( /usr/local/lib/python3.11/dist-packages/torch/cuda/__init__.py:58: FutureWarning: The pynvml package is deprecated. Please install nvidia-ml-py instead. If you did not install pynvml directly, please report this to the maintainers of the package that installed pynvml for you. import pynvml # type: ignore[import] /workspace/FLAME-MoE/.venv/lib/python3.11/site-packages/transformer_engine/pytorch/attention.py:108: UserWarning: To use flash-attn v3, please use the following commands to install: (1) pip install "git+https://github.com/Dao-AILab/flash-attention.git#egg=flashattn-hopper&subdirectory=hopper" (2) python_path=`python -c "import site; print(site.getsitepackages()[0])"` (3) mkdir -p $python_path/flashattn_hopper (4) wget -P $python_path/flashattn_hopper https://raw.githubusercontent.com/Dao-AILab/flash-attention/main/hopper/flash_attn_interface.py warnings.warn( /workspace/FLAME-MoE/Megatron-LM/megatron/core/models/gpt/gpt_layer_specs.py:54: UserWarning: Apex is not installed. Falling back to Torch Norm warnings.warn('Apex is not installed. Falling back to Torch Norm') /workspace/FLAME-MoE/Megatron-LM/megatron/core/models/retro/encoder_spec.py:47: UserWarning: Apex is not installed. Falling back to Torch Norm warnings.warn(f'Apex is not installed. Falling back to Torch Norm') /workspace/FLAME-MoE/Megatron-LM/megatron/core/models/retro/decoder_spec.py:39: UserWarning: Apex is not installed. Falling back to Torch Norm warnings.warn(f'Apex is not installed. Falling back to Torch Norm') 2026-07-23:03:42:53,691 INFO [__main__.py:279] Verbosity set to INFO 2026-07-23:03:43:23,492 INFO [__init__.py:459] The tag 'arc_ca' is already registered as a group, this tag will not be registered. This may affect tasks you want to call. 2026-07-23:03:43:23,511 INFO [__init__.py:459] The tag 'arc_ca' is already registered as a group, this tag will not be registered. This may affect tasks you want to call. 2026-07-23:03:43:41,361 INFO [__main__.py:376] Selected Tasks: ['arc_challenge', 'arc_easy', 'boolq', 'copa', 'hellaswag', 'lambada_openai', 'openbookqa', 'piqa', 'sciq', 'winogrande'] 2026-07-23:03:43:41,381 INFO [evaluator.py:164] Setting random seed to 42 | Setting numpy seed to 42 | Setting torch manual seed to 42 | Setting fewshot manual seed to 42 2026-07-23:03:43:41,381 INFO [evaluator.py:201] Initializing megatron_lm model, with arguments: {} using world size: 1, data-parallel size: 1, context-parallel size: 1, hierarchical context-parallel sizes: Nonetensor-model-parallel size: 1, encoder-tensor-model-parallel size: 0, pipeline-model-parallel size: 1, encoder-pipeline-model-parallel size: 0 WARNING: overriding default arguments for exit_on_missing_checkpoint:True with exit_on_missing_checkpoint:False setting global batch size to 32 Number of virtual stages per pipeline stage: None accumulate and all-reduce gradients in fp32 for bfloat16 data type. using torch.bfloat16 for parameters ... ------------------------ arguments ------------------------ account_for_embedding_in_pipeline_split ......... False account_for_loss_in_pipeline_split .............. False accumulate_allreduce_grads_in_fp32 .............. True adam_beta1 ...................................... 0.9 adam_beta2 ...................................... 0.999 adam_eps ........................................ 1e-08 add_bias_linear ................................. False add_position_embedding .......................... True add_qkv_bias .................................... False adlr_autoresume ................................. False adlr_autoresume_interval ........................ 1000 align_grad_reduce ............................... True align_param_gather .............................. False app_tag_run_name ................................ None app_tag_run_version ............................. 0.0.0 apply_layernorm_1p .............................. False apply_query_key_layer_scaling ................... False apply_residual_connection_post_layernorm ........ False apply_rope_fusion ............................... True async_save ...................................... None async_tensor_model_parallel_allreduce ........... True attention_backend ............................... AttnBackend.auto attention_dropout ............................... 0.0 attention_softmax_in_fp32 ....................... False auto_detect_ckpt_format ......................... False barrier_with_L1_time ............................ True bert_binary_head ................................ True bert_embedder_type .............................. megatron bert_load ....................................... None bf16 ............................................ True bias_dropout_fusion ............................. True bias_gelu_fusion ................................ False bias_swiglu_fusion .............................. True biencoder_projection_dim ........................ 0 biencoder_shared_query_context_model ............ False block_data_path ................................. None calc_ft_timeouts ................................ False calculate_per_token_loss ........................ False check_for_large_grads ........................... False check_for_nan_in_loss_and_grad .................. True check_for_spiky_loss ............................ False check_weight_hash_across_dp_replicas_interval ... None ckpt_assume_constant_structure .................. False ckpt_convert_format ............................. None ckpt_convert_save ............................... None ckpt_convert_update_legacy_dist_opt_format ...... False ckpt_format ..................................... torch_dist ckpt_fully_parallel_load ........................ False ckpt_fully_parallel_save ........................ True ckpt_fully_parallel_save_deprecated ............. False ckpt_step ....................................... None classes_fraction ................................ 1.0 clip_grad ....................................... 1.0 clone_scatter_output_in_embedding ............... True config_logger_dir ............................... consumed_train_samples .......................... 0 consumed_valid_samples .......................... 0 context_parallel_size ........................... 1 cp_comm_type .................................... ['p2p'] create_attention_mask_in_dataloader ............. True cross_entropy_loss_fusion ....................... False cuda_graph_warmup_steps ......................... 3 data_args_path .................................. None data_cache_path ................................. None data_parallel_random_init ....................... False data_parallel_sharding_strategy ................. no_shard data_parallel_size .............................. 1 data_path ....................................... None data_per_class_fraction ......................... 1.0 data_sharding ................................... True dataloader_type ................................. single ddp_average_in_collective ....................... False ddp_bucket_size ................................. None ddp_num_buckets ................................. None ddp_pad_buckets_for_high_nccl_busbw ............. False decoder_first_pipeline_num_layers ............... None decoder_last_pipeline_num_layers ................ None decoder_num_layers .............................. None decoder_seq_length .............................. None decoupled_lr .................................... None decoupled_min_lr ................................ None decrease_batch_size_if_needed ................... False defer_embedding_wgrad_compute ................... False deprecated_use_mcore_models ..................... False deterministic_mode .............................. False dino_bottleneck_size ............................ 256 dino_freeze_last_layer .......................... 1 dino_head_hidden_size ........................... 2048 dino_local_crops_number ......................... 10 dino_local_img_size ............................. 96 dino_norm_last_layer ............................ False dino_teacher_temp ............................... 0.07 dino_warmup_teacher_temp ........................ 0.04 dino_warmup_teacher_temp_epochs ................. 30 disable_straggler_on_startup .................... False dist_ckpt_format_deprecated ..................... None dist_ckpt_strictness ............................ assume_ok_unexpected distribute_saved_activations .................... False distributed_backend ............................. nccl distributed_timeout_minutes ..................... 10 embedding_path .................................. None empty_unused_memory_level ....................... 0 enable_cuda_graph ............................... False enable_ft_package ............................... False enable_gloo_process_groups ...................... True enable_one_logger ............................... True encoder_num_layers .............................. 9 encoder_pipeline_model_parallel_size ............ 0 encoder_seq_length .............................. 2048 encoder_tensor_model_parallel_size .............. 0 end_weight_decay ................................ 0.01 eod_mask_loss ................................... False error_injection_rate ............................ 0 error_injection_type ............................ transient_error eval_interval ................................... 1000 eval_iters ...................................... 100 evidence_data_path .............................. None exit_duration_in_mins ........................... None exit_interval ................................... None exit_on_missing_checkpoint ...................... False exit_signal_handler ............................. False exp_avg_dtype ................................... torch.float32 exp_avg_sq_dtype ................................ torch.float32 expert_model_parallel_size ...................... 1 expert_tensor_parallel_size ..................... 1 ffn_hidden_size ................................. 1368 finetune ........................................ False flash_decode .................................... False fp16 ............................................ False fp16_lm_cross_entropy ........................... False fp32_residual_connection ........................ False fp8 ............................................. None fp8_amax_compute_algo ........................... most_recent fp8_amax_history_len ............................ 1 fp8_interval .................................... 1 fp8_margin ...................................... 0 fp8_param_gather ................................ False fp8_wgrad ....................................... True global_batch_size ............................... 32 grad_reduce_in_bf16 ............................. False gradient_accumulation_fusion .................... False gradient_reduce_div_fusion ...................... True group_query_attention ........................... False head_lr_mult .................................... 1.0 hidden_dropout .................................. 0.0 hidden_size ..................................... 256 hierarchical_context_parallel_sizes ............. None hybrid_attention_ratio .......................... 0.0 hybrid_mlp_ratio ................................ 0.0 hybrid_override_pattern ......................... None hysteresis ...................................... 2 ict_head_size ................................... None ict_load ........................................ None img_h ........................................... 224 img_w ........................................... 224 indexer_batch_size .............................. 128 indexer_log_interval ............................ 1000 inference_batch_times_seqlen_threshold .......... -1 inference_max_batch_size ........................ 8 inference_max_seq_length ........................ 2560 inference_rng_tracker ........................... False init_method_std ................................. 0.02 init_method_xavier_uniform ...................... False init_model_with_meta_device ..................... False initial_loss_scale .............................. 4294967296 iter_per_epoch .................................. 1250 iterations_to_skip .............................. [] keep_fp8_transpose_cache_when_using_custom_fsdp . False kv_channels ..................................... 16 kv_lora_rank .................................... 32 lazy_mpu_init ................................... None load ............................................ /workspace/FLAME-MoE/results/phase0/runs/flame38m_g3_moe_s2/ckpt local_rank ...................................... 0 log_interval .................................... 100 log_loss_scale_to_tensorboard ................... True log_memory_to_tensorboard ....................... False log_num_zeros_in_grad ........................... False log_params_norm ................................. False log_progress .................................... False log_straggler ................................... False log_throughput .................................. False log_timers_to_tensorboard ....................... False log_validation_ppl_to_tensorboard ............... False log_world_size_to_tensorboard ................... False logging_level ................................... None loss_scale ...................................... None loss_scale_window ............................... 1000 lr .............................................. None lr_decay_iters .................................. None lr_decay_samples ................................ None lr_decay_style .................................. linear lr_warmup_fraction .............................. None lr_warmup_init .................................. 0.0 lr_warmup_iters ................................. 0 lr_warmup_samples ............................... 0 lr_wsd_decay_iters .............................. None lr_wsd_decay_samples ............................ None lr_wsd_decay_style .............................. exponential main_grads_dtype ................................ torch.float32 main_params_dtype ............................... torch.float32 make_vocab_size_divisible_by .................... 128 manual_gc ....................................... False manual_gc_eval .................................. True manual_gc_interval .............................. 0 mask_factor ..................................... 1.0 mask_prob ....................................... 0.15 mask_type ....................................... random masked_softmax_fusion ........................... True max_position_embeddings ......................... 2048 max_tokens_to_oom ............................... 10000000 memory_snapshot_path ............................ snapshot.pickle merge_file ...................................... None micro_batch_size ................................ 32 microbatch_group_size_per_vp_stage .............. None min_loss_scale .................................. 1.0 min_lr .......................................... 0.0 mmap_bin_files .................................. True mock_data ....................................... False moe_aux_loss_coeff .............................. 0.01 moe_enable_deepep ............................... False moe_expert_capacity_factor ...................... None moe_extended_tp ................................. False moe_ffn_hidden_size ............................. 58 moe_grouped_gemm ................................ True moe_input_jitter_eps ............................ None moe_layer_freq .................................. [0, 1, 1, 1, 1, 1, 1, 1, 1] moe_layer_recompute ............................. False moe_pad_expert_input_to_capacity ................ False moe_per_layer_logging ........................... False moe_permute_fusion .............................. False moe_router_bias_update_rate ..................... 0.001 moe_router_dtype ................................ fp32 moe_router_enable_expert_bias ................... False moe_router_group_topk ........................... None moe_router_load_balancing_type .................. aux_loss moe_router_num_groups ........................... None moe_router_pre_softmax .......................... True moe_router_score_function ....................... softmax moe_router_topk ................................. 18 moe_router_topk_scaling_factor .................. None moe_shared_expert_intermediate_size ............. 352 moe_shared_expert_overlap ....................... False moe_token_dispatcher_type ....................... alltoall moe_token_drop_policy ........................... probs moe_use_legacy_grouped_gemm ..................... False moe_use_upcycling ............................... False moe_z_loss_coeff ................................ 0.001 mscale .......................................... 1.0 mscale_all_dim .................................. 1.0 multi_latent_attention .......................... False nccl_communicator_config_path ................... None no_load_optim ................................... True no_load_rng ..................................... True no_persist_layer_norm ........................... False no_save_optim ................................... None no_save_rng ..................................... None non_persistent_ckpt_type ........................ None non_persistent_global_ckpt_dir .................. None non_persistent_local_ckpt_algo .................. fully_parallel non_persistent_local_ckpt_dir ................... None non_persistent_save_interval .................... None norm_epsilon .................................... 1e-06 normalization ................................... RMSNorm num_attention_heads ............................. 16 num_channels .................................... 3 num_classes ..................................... 1000 num_dataset_builder_threads ..................... 1 num_distributed_optimizer_instances ............. 1 num_experts ..................................... 192 num_layers ...................................... 9 num_layers_per_virtual_pipeline_stage ........... None num_query_groups ................................ 1 num_virtual_stages_per_pipeline_rank ............ None num_workers ..................................... 2 one_logger_async ................................ False one_logger_project .............................. megatron-lm one_logger_run_name ............................. None onnx_safe ....................................... None openai_gelu ..................................... False optimizer ....................................... adam optimizer_cpu_offload ........................... False optimizer_offload_fraction ...................... 1.0 output_bert_embeddings .......................... False overlap_cpu_optimizer_d2h_h2d ................... False overlap_grad_reduce ............................. False overlap_p2p_comm ................................ False overlap_p2p_comm_warmup_flush ................... False overlap_param_gather ............................ False overlap_param_gather_with_optimizer_step ........ False override_opt_param_scheduler .................... False params_dtype .................................... torch.bfloat16 patch_dim ....................................... 16 per_split_data_args_path ........................ None perform_initialization .......................... True pin_cpu_grads ................................... True pin_cpu_params .................................. True pipeline_model_parallel_comm_backend ............ None pipeline_model_parallel_size .................... 1 pipeline_model_parallel_split_rank .............. None position_embedding_type ......................... rope pretrained_checkpoint ........................... None profile ......................................... False profile_ranks ................................... [0] profile_step_end ................................ 12 profile_step_start .............................. 10 q_lora_rank ..................................... None qk_head_dim ..................................... 128 qk_layernorm .................................... False qk_pos_emb_head_dim ............................. 64 query_in_block_prob ............................. 0.1 rampup_batch_size ............................... None rank ............................................ 0 recompute_granularity ........................... None recompute_method ................................ None recompute_num_layers ............................ None record_memory_history ........................... False relative_attention_max_distance ................. 128 relative_attention_num_buckets .................. 32 replication ..................................... False replication_factor .............................. 2 replication_jump ................................ None rerun_mode ...................................... disabled reset_attention_mask ............................ False reset_position_ids .............................. False result_rejected_tracker_filename ................ None retriever_report_topk_accuracies ................ [] retriever_score_scaling ......................... False retriever_seq_length ............................ 256 retro_add_retriever ............................. False retro_attention_gate ............................ 1 retro_cyclic_train_iters ........................ None retro_encoder_attention_dropout ................. 0.1 retro_encoder_hidden_dropout .................... 0.1 retro_encoder_layers ............................ 2 retro_num_neighbors ............................. 2 retro_num_retrieved_chunks ...................... 2 retro_project_dir ............................... None retro_verify_neighbor_count ..................... True rope_scaling_factor ............................. 8.0 rotary_base ..................................... 10000 rotary_interleaved .............................. False rotary_percent .................................. 1.0 rotary_scaling_factor ........................... 1.0 rotary_seq_len_interpolation_factor ............. None s3_cache_path ................................... None sample_rate ..................................... 1.0 save ............................................ None save_interval ................................... None scatter_gather_tensors_in_pipeline .............. True seed ............................................ 42 seq_length ...................................... 2048 sequence_parallel ............................... False sgd_momentum .................................... 0.9 short_seq_prob .................................. 0.1 skip_train ...................................... False skipped_train_samples ........................... 0 spec ............................................ None split ........................................... None squared_relu .................................... False start_weight_decay .............................. 0.01 straggler_ctrlr_port ............................ 65535 straggler_minmax_count .......................... 1 suggested_communication_unit_size ............... 400000000 swiglu .......................................... True swin_backbone_type .............................. tiny te_rng_tracker .................................. False tensor_model_parallel_size ...................... 1 tensorboard_dir ................................. None tensorboard_log_interval ........................ 1 tensorboard_queue_size .......................... 1000 test_data_path .................................. None test_mode ....................................... False tiktoken_num_special_tokens ..................... 1000 tiktoken_pattern ................................ None tiktoken_special_tokens ......................... None timing_log_level ................................ 0 timing_log_option ............................... minmax titles_data_path ................................ None tokenizer_model ................................. EleutherAI/pythia-12b tokenizer_type .................................. HuggingFaceTokenizer tp_comm_bootstrap_backend ....................... nccl tp_comm_bulk_dgrad .............................. True tp_comm_bulk_wgrad .............................. True tp_comm_overlap ................................. False tp_comm_overlap_ag .............................. True tp_comm_overlap_cfg ............................. None tp_comm_overlap_rs .............................. True tp_comm_overlap_rs_dgrad ........................ False tp_comm_split_ag ................................ True tp_comm_split_rs ................................ True train_data_path ................................. None train_iters ..................................... None train_samples ................................... None train_sync_interval ............................. None transformer_impl ................................ transformer_engine transformer_pipeline_model_parallel_size ........ 1 untie_embeddings_and_output_weights ............. True use_checkpoint_args ............................. False use_checkpoint_opt_param_scheduler .............. False use_cpu_initialization .......................... None use_custom_fsdp ................................. False use_dist_ckpt ................................... True use_dist_ckpt_deprecated ........................ False use_distributed_optimizer ....................... True use_flash_attn .................................. False use_legacy_models ............................... False use_mp_args_from_checkpoint_args ................ False use_one_sent_docs ............................... False use_persistent_ckpt_worker ...................... False use_precision_aware_optimizer ................... False use_pytorch_profiler ............................ False use_ring_exchange_p2p ........................... False use_rope_scaling ................................ False use_rotary_position_embeddings .................. False use_tokenizer_model_from_checkpoint_args ........ True use_torch_fsdp2 ................................. False use_torch_optimizer_for_cpu_offload ............. False use_tp_pp_dp_mapping ............................ False v_head_dim ...................................... 128 valid_data_path ................................. None variable_seq_lengths ............................ False virtual_pipeline_model_parallel_size ............ None vision_backbone_type ............................ vit vision_pretraining .............................. False vision_pretraining_type ......................... classify vocab_extra_ids ................................. 0 vocab_file ...................................... None vocab_size ...................................... None wandb_exp_name .................................. wandb_project ................................... wandb_save_dir .................................. weight_decay .................................... 0.01 weight_decay_incr_style ......................... constant wgrad_deferral_limit ............................ 0 world_size ...................................... 1 yaml_cfg ........................................ None -------------------- end of arguments --------------------- 2026-07-23:03:43:41,396 INFO [num_microbatches_calculator.py:228] setting number of microbatches to constant 1 > building HuggingFaceTokenizer tokenizer ... 2026-07-23:03:43:41,632 INFO [_client.py:1038] HTTP Request: HEAD https://huggingface.co/EleutherAI/pythia-12b/resolve/main/config.json "HTTP/1.1 307 Temporary Redirect" 2026-07-23:03:43:41,644 INFO [_client.py:1038] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/EleutherAI/pythia-12b/bb1e3e710cdf6b524461d543cfb5ba773f0a81b6/config.json "HTTP/1.1 200 OK" 2026-07-23:03:43:41,779 INFO [_client.py:1038] HTTP Request: HEAD https://huggingface.co/EleutherAI/pythia-12b/resolve/main/tokenizer_config.json "HTTP/1.1 307 Temporary Redirect" 2026-07-23:03:43:41,792 INFO [_client.py:1038] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/EleutherAI/pythia-12b/bb1e3e710cdf6b524461d543cfb5ba773f0a81b6/tokenizer_config.json "HTTP/1.1 200 OK" 2026-07-23:03:43:41,897 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/models/EleutherAI/pythia-12b/tree/main/additional_chat_templates?recursive=false&expand=false "HTTP/1.1 404 Not Found" 2026-07-23:03:43:41,969 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/models/EleutherAI/pythia-12b/tree/main?recursive=true&expand=false "HTTP/1.1 200 OK" > padded vocab (size: 50277) with 27 dummy tokens (new size: 50304) WARNING: one_logger package is required to enable e2e metrics tracking. please go to https://confluence.nvidia.com/display/MLWFO/Package+Repositories for details to install it 2026-07-23:03:43:42,165 WARNING [rerun_state_machine.py:239] RerunStateMachine initialized in mode RerunMode.DISABLED > initializing torch distributed ... [W723 03:43:42.682600685 CUDAAllocatorConfig.h:28] Warning: expandable_segments not supported on this platform (function operator()) > initialized tensor model parallel with size 1 > initialized pipeline model parallel with size 1 > setting random seeds to 42 ... > compiling dataset index builder ... make: Entering directory '/workspace/FLAME-MoE/Megatron-LM/megatron/core/datasets' make: Nothing to be done for 'default'. make: Leaving directory '/workspace/FLAME-MoE/Megatron-LM/megatron/core/datasets' >>> done with dataset index builder. Compilation time: 0.784 seconds > compiling and loading fused kernels ... >>> done with compiling and loading fused kernels. Compilation time: 0.316 seconds building GPT model ... > number of parameters on (tensor, pipeline) model parallel rank (0, 0): 100145920 /workspace/FLAME-MoE/Megatron-LM/megatron/core/transformer/transformer_layer.py:339: UserWarning: TransformerLayer._get_layer_offset is deprecated.Please use get_transformer_layer_offset instead. warnings.warn( loading distributed checkpoint from /workspace/FLAME-MoE/results/phase0/runs/flame38m_g3_moe_s2/ckpt at iteration 2121 /workspace/FLAME-MoE/Megatron-LM/megatron/core/dist_checkpointing/strategies/torch.py:847: FutureWarning: `load_state_dict` is deprecated and will be removed in future versions. Please use `load` instead. checkpoint.load_state_dict( /usr/local/lib/python3.11/dist-packages/torch/distributed/checkpoint/planner_helpers.py:306: FutureWarning: Please use DTensor instead and we are deprecating ShardedTensor. device = getattr(value, "device", None) /workspace/FLAME-MoE/.venv/lib/python3.11/site-packages/transformer_engine/pytorch/module/base.py:566: FutureWarning: You are using `torch.load` with `weights_only=False` (the current default value), which uses the default pickle module implicitly. It is possible to construct malicious pickle data which will execute arbitrary code during unpickling (See https://github.com/pytorch/pytorch/blob/main/SECURITY.md#untrusted-models for more details). In a future release, the default value for `weights_only` will be flipped to `True`. This limits the functions that could be executed during unpickling. Arbitrary objects will no longer be allowed to be loaded via this mode unless they are explicitly allowlisted by the user via `torch.serialization.add_safe_globals`. We recommend you start setting `weights_only=True` for any use case where you don't have full control of the loaded file. Please open an issue on GitHub for any issues related to this experimental feature. state = torch.load(state, map_location="cuda") checkpoint version 3.0 2026-07-23:03:43:48,551 WARNING [rerun_state_machine.py:811] RerunStateMachine disabled via CLI, ignoring machine state saved in checkpoint successfully loaded checkpoint from /workspace/FLAME-MoE/results/phase0/runs/flame38m_g3_moe_s2/ckpt [ t 1/1, p 1/1 ] at iteration 2121 2026-07-23:03:43:48,679 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc "HTTP/1.1 200 OK" 2026-07-23:03:43:48,902 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc "HTTP/1.1 200 OK" 2026-07-23:03:43:48,967 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc/revision/210d026faf9955653af8916fad021475a3f00453 "HTTP/1.1 200 OK" 2026-07-23:03:43:49,031 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc/tree/210d026faf9955653af8916fad021475a3f00453?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:43:49,278 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc/tree/210d026faf9955653af8916fad021475a3f00453/ARC-Challenge?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:43:49,955 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc "HTTP/1.1 200 OK" 2026-07-23:03:43:50,182 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc "HTTP/1.1 200 OK" 2026-07-23:03:43:50,445 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc/tree/210d026faf9955653af8916fad021475a3f00453/ARC-Easy?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:43:50,587 WARNING [task.py:799] [Task: boolq] metric acc is defined, but aggregation is not. using default aggregation=mean 2026-07-23:03:43:50,587 WARNING [task.py:811] [Task: boolq] metric acc is defined, but higher_is_better is not. using default higher_is_better=True 2026-07-23:03:43:50,650 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue "HTTP/1.1 200 OK" 2026-07-23:03:43:50,884 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue "HTTP/1.1 200 OK" 2026-07-23:03:43:50,947 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/revision/3de24cf8022e94f4ee4b9d55a6f539891524d646 "HTTP/1.1 200 OK" 2026-07-23:03:43:51,076 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/tree/3de24cf8022e94f4ee4b9d55a6f539891524d646?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:43:51,346 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/tree/3de24cf8022e94f4ee4b9d55a6f539891524d646/axb?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:43:51,418 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/revision/3de24cf8022e94f4ee4b9d55a6f539891524d646 "HTTP/1.1 200 OK" 2026-07-23:03:43:51,485 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/tree/3de24cf8022e94f4ee4b9d55a6f539891524d646/boolq?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:43:51,729 WARNING [task.py:799] [Task: copa] metric acc is defined, but aggregation is not. using default aggregation=mean 2026-07-23:03:43:51,729 WARNING [task.py:811] [Task: copa] metric acc is defined, but higher_is_better is not. using default higher_is_better=True 2026-07-23:03:43:51,790 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue "HTTP/1.1 200 OK" 2026-07-23:03:43:52,011 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue "HTTP/1.1 200 OK" 2026-07-23:03:43:52,563 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/tree/3de24cf8022e94f4ee4b9d55a6f539891524d646/copa?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:43:52,705 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag "HTTP/1.1 200 OK" 2026-07-23:03:43:52,936 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag "HTTP/1.1 200 OK" 2026-07-23:03:43:52,998 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag/revision/218ec52e09a7e7462a5400043bb9a69a41d06b76 "HTTP/1.1 200 OK" 2026-07-23:03:43:53,065 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag/tree/218ec52e09a7e7462a5400043bb9a69a41d06b76?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:43:53,344 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag/tree/218ec52e09a7e7462a5400043bb9a69a41d06b76/data?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:43:53,414 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag/revision/218ec52e09a7e7462a5400043bb9a69a41d06b76 "HTTP/1.1 200 OK" 2026-07-23:03:43:55,497 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai "HTTP/1.1 200 OK" 2026-07-23:03:43:55,746 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai "HTTP/1.1 200 OK" 2026-07-23:03:43:55,801 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/revision/900124bf3b8235c6daf21033af9948b3f07346c4 "HTTP/1.1 200 OK" 2026-07-23:03:43:55,865 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/tree/900124bf3b8235c6daf21033af9948b3f07346c4?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:43:56,117 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/tree/900124bf3b8235c6daf21033af9948b3f07346c4/default%2Ftest?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:43:56,187 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/tree/900124bf3b8235c6daf21033af9948b3f07346c4/default?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:43:56,246 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/revision/900124bf3b8235c6daf21033af9948b3f07346c4 "HTTP/1.1 200 OK" 2026-07-23:03:43:56,264 WARNING [task.py:325] [Task: lambada_openai] has_training_docs and has_validation_docs are False, using test_docs as fewshot_docs but this is not recommended. 2026-07-23:03:43:56,264 WARNING [task.py:325] [Task: lambada_openai] has_training_docs and has_validation_docs are False, using test_docs as fewshot_docs but this is not recommended. 2026-07-23:03:43:56,434 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa "HTTP/1.1 200 OK" 2026-07-23:03:43:56,661 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa "HTTP/1.1 200 OK" 2026-07-23:03:43:56,720 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/revision/388097ea7776314e93a529163e0fea805b8a6454 "HTTP/1.1 200 OK" 2026-07-23:03:43:56,784 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/tree/388097ea7776314e93a529163e0fea805b8a6454?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:43:57,025 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/tree/388097ea7776314e93a529163e0fea805b8a6454/additional?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:43:57,083 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/revision/388097ea7776314e93a529163e0fea805b8a6454 "HTTP/1.1 200 OK" 2026-07-23:03:43:57,175 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/tree/388097ea7776314e93a529163e0fea805b8a6454/main?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:43:57,432 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa "HTTP/1.1 200 OK" 2026-07-23:03:43:57,487 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa "HTTP/1.1 200 OK" 2026-07-23:03:43:57,551 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa/tree/main?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:43:57,631 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/datasets/ybisk/piqa/resolve/main/piqa.py "HTTP/1.1 307 Temporary Redirect" 2026-07-23:03:43:57,643 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/resolve-cache/datasets/ybisk/piqa/2e8ac2dffd59bac8c3c6714948f4c551a0848bb0/piqa.py?%2Fdatasets%2Fybisk%2Fpiqa%2Fresolve%2Fmain%2Fpiqa.py=&etag=%220f0d96172f9bf7aa10069374b3236ef083cfeb8e%22 "HTTP/1.1 200 OK" 2026-07-23:03:43:58,275 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa/revision/main "HTTP/1.1 200 OK" 2026-07-23:03:43:58,339 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa/tree/main?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:43:58,762 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq "HTTP/1.1 200 OK" 2026-07-23:03:43:59,002 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq "HTTP/1.1 200 OK" 2026-07-23:03:43:59,061 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq/revision/2c94ad3e1aafab77146f384e23536f97a4849815 "HTTP/1.1 200 OK" 2026-07-23:03:43:59,125 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq/tree/2c94ad3e1aafab77146f384e23536f97a4849815?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:43:59,439 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq/tree/2c94ad3e1aafab77146f384e23536f97a4849815/data?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:43:59,500 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq/revision/2c94ad3e1aafab77146f384e23536f97a4849815 "HTTP/1.1 200 OK" 2026-07-23:03:43:59,920 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande "HTTP/1.1 200 OK" 2026-07-23:03:44:00,140 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande "HTTP/1.1 200 OK" 2026-07-23:03:44:00,238 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/revision/01e74176c63542e6b0bcb004dcdea22d94fb67b5 "HTTP/1.1 200 OK" 2026-07-23:03:44:00,298 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/tree/01e74176c63542e6b0bcb004dcdea22d94fb67b5?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:44:00,697 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/tree/01e74176c63542e6b0bcb004dcdea22d94fb67b5/winogrande_debiased?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:44:00,763 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/revision/01e74176c63542e6b0bcb004dcdea22d94fb67b5 "HTTP/1.1 200 OK" 2026-07-23:03:44:00,851 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/tree/01e74176c63542e6b0bcb004dcdea22d94fb67b5/winogrande_xl?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:44:01,492 WARNING [evaluator.py:270] Overwriting default num_fewshot of winogrande from None to 0 2026-07-23:03:44:01,492 WARNING [evaluator.py:270] Overwriting default num_fewshot of sciq from None to 0 2026-07-23:03:44:01,492 WARNING [evaluator.py:270] Overwriting default num_fewshot of piqa from None to 0 2026-07-23:03:44:01,492 WARNING [evaluator.py:270] Overwriting default num_fewshot of openbookqa from None to 0 2026-07-23:03:44:01,492 WARNING [evaluator.py:270] Overwriting default num_fewshot of lambada_openai from None to 0 2026-07-23:03:44:01,492 WARNING [evaluator.py:270] Overwriting default num_fewshot of hellaswag from None to 0 2026-07-23:03:44:01,492 WARNING [evaluator.py:270] Overwriting default num_fewshot of copa from None to 0 2026-07-23:03:44:01,492 WARNING [evaluator.py:270] Overwriting default num_fewshot of boolq from None to 0 2026-07-23:03:44:01,492 WARNING [evaluator.py:270] Overwriting default num_fewshot of arc_easy from None to 0 2026-07-23:03:44:01,492 WARNING [evaluator.py:270] Overwriting default num_fewshot of arc_challenge from None to 0 2026-07-23:03:44:01,527 INFO [task.py:415] Building contexts for winogrande on rank 0... 0%| | 0/1267 [00:00 /workspace/FLAME-MoE/results/ablations/flame1e18_downstream.csv ==================================================================== [run] flame38m_g3_moe_s3 mode=moe TEMPORAL=0 H=256 heads=16 ffn=1368 experts=192 topk=18 moe_ffn=58 shared=352 [run] tasks=arc_challenge,arc_easy,hellaswag,openbookqa,piqa,winogrande,boolq,copa,lambada_openai,sciq out=/workspace/FLAME-MoE/results/phase0/runs/flame38m_g3_moe_s3/lmeval_1e18_0shot ==================================================================== /usr/local/lib/python3.11/dist-packages/torch/cuda/__init__.py:58: FutureWarning: The pynvml package is deprecated. Please install nvidia-ml-py instead. If you did not install pynvml directly, please report this to the maintainers of the package that installed pynvml for you. import pynvml # type: ignore[import] [run_lmeval] stubbed VLM backends (hf_vlms, vllm_vlms) /usr/local/lib/python3.11/dist-packages/torch/cuda/__init__.py:58: FutureWarning: The pynvml package is deprecated. Please install nvidia-ml-py instead. If you did not install pynvml directly, please report this to the maintainers of the package that installed pynvml for you. import pynvml # type: ignore[import] /usr/local/lib/python3.11/dist-packages/requests/__init__.py:113: RequestsDependencyWarning: urllib3 (2.2.3) or chardet (6.0.0.post1)/charset_normalizer (3.3.2) doesn't match a supported version! warnings.warn( /usr/local/lib/python3.11/dist-packages/torch/cuda/__init__.py:58: FutureWarning: The pynvml package is deprecated. Please install nvidia-ml-py instead. If you did not install pynvml directly, please report this to the maintainers of the package that installed pynvml for you. import pynvml # type: ignore[import] /workspace/FLAME-MoE/.venv/lib/python3.11/site-packages/transformer_engine/pytorch/attention.py:108: UserWarning: To use flash-attn v3, please use the following commands to install: (1) pip install "git+https://github.com/Dao-AILab/flash-attention.git#egg=flashattn-hopper&subdirectory=hopper" (2) python_path=`python -c "import site; print(site.getsitepackages()[0])"` (3) mkdir -p $python_path/flashattn_hopper (4) wget -P $python_path/flashattn_hopper https://raw.githubusercontent.com/Dao-AILab/flash-attention/main/hopper/flash_attn_interface.py warnings.warn( /workspace/FLAME-MoE/Megatron-LM/megatron/core/models/gpt/gpt_layer_specs.py:54: UserWarning: Apex is not installed. Falling back to Torch Norm warnings.warn('Apex is not installed. Falling back to Torch Norm') /workspace/FLAME-MoE/Megatron-LM/megatron/core/models/retro/encoder_spec.py:47: UserWarning: Apex is not installed. Falling back to Torch Norm warnings.warn(f'Apex is not installed. Falling back to Torch Norm') /workspace/FLAME-MoE/Megatron-LM/megatron/core/models/retro/decoder_spec.py:39: UserWarning: Apex is not installed. Falling back to Torch Norm warnings.warn(f'Apex is not installed. Falling back to Torch Norm') 2026-07-23:03:53:36,201 INFO [__main__.py:279] Verbosity set to INFO 2026-07-23:03:54:07,014 INFO [__init__.py:459] The tag 'arc_ca' is already registered as a group, this tag will not be registered. This may affect tasks you want to call. 2026-07-23:03:54:07,029 INFO [__init__.py:459] The tag 'arc_ca' is already registered as a group, this tag will not be registered. This may affect tasks you want to call. 2026-07-23:03:54:24,172 INFO [__main__.py:376] Selected Tasks: ['arc_challenge', 'arc_easy', 'boolq', 'copa', 'hellaswag', 'lambada_openai', 'openbookqa', 'piqa', 'sciq', 'winogrande'] 2026-07-23:03:54:24,195 INFO [evaluator.py:164] Setting random seed to 42 | Setting numpy seed to 42 | Setting torch manual seed to 42 | Setting fewshot manual seed to 42 2026-07-23:03:54:24,195 INFO [evaluator.py:201] Initializing megatron_lm model, with arguments: {} using world size: 1, data-parallel size: 1, context-parallel size: 1, hierarchical context-parallel sizes: Nonetensor-model-parallel size: 1, encoder-tensor-model-parallel size: 0, pipeline-model-parallel size: 1, encoder-pipeline-model-parallel size: 0 WARNING: overriding default arguments for exit_on_missing_checkpoint:True with exit_on_missing_checkpoint:False setting global batch size to 32 Number of virtual stages per pipeline stage: None accumulate and all-reduce gradients in fp32 for bfloat16 data type. using torch.bfloat16 for parameters ... ------------------------ arguments ------------------------ account_for_embedding_in_pipeline_split ......... False account_for_loss_in_pipeline_split .............. False accumulate_allreduce_grads_in_fp32 .............. True adam_beta1 ...................................... 0.9 adam_beta2 ...................................... 0.999 adam_eps ........................................ 1e-08 add_bias_linear ................................. False add_position_embedding .......................... True add_qkv_bias .................................... False adlr_autoresume ................................. False adlr_autoresume_interval ........................ 1000 align_grad_reduce ............................... True align_param_gather .............................. False app_tag_run_name ................................ None app_tag_run_version ............................. 0.0.0 apply_layernorm_1p .............................. False apply_query_key_layer_scaling ................... False apply_residual_connection_post_layernorm ........ False apply_rope_fusion ............................... True async_save ...................................... None async_tensor_model_parallel_allreduce ........... True attention_backend ............................... AttnBackend.auto attention_dropout ............................... 0.0 attention_softmax_in_fp32 ....................... False auto_detect_ckpt_format ......................... False barrier_with_L1_time ............................ True bert_binary_head ................................ True bert_embedder_type .............................. megatron bert_load ....................................... None bf16 ............................................ True bias_dropout_fusion ............................. True bias_gelu_fusion ................................ False bias_swiglu_fusion .............................. True biencoder_projection_dim ........................ 0 biencoder_shared_query_context_model ............ False block_data_path ................................. None calc_ft_timeouts ................................ False calculate_per_token_loss ........................ False check_for_large_grads ........................... False check_for_nan_in_loss_and_grad .................. True check_for_spiky_loss ............................ False check_weight_hash_across_dp_replicas_interval ... None ckpt_assume_constant_structure .................. False ckpt_convert_format ............................. None ckpt_convert_save ............................... None ckpt_convert_update_legacy_dist_opt_format ...... False ckpt_format ..................................... torch_dist ckpt_fully_parallel_load ........................ False ckpt_fully_parallel_save ........................ True ckpt_fully_parallel_save_deprecated ............. False ckpt_step ....................................... None classes_fraction ................................ 1.0 clip_grad ....................................... 1.0 clone_scatter_output_in_embedding ............... True config_logger_dir ............................... consumed_train_samples .......................... 0 consumed_valid_samples .......................... 0 context_parallel_size ........................... 1 cp_comm_type .................................... ['p2p'] create_attention_mask_in_dataloader ............. True cross_entropy_loss_fusion ....................... False cuda_graph_warmup_steps ......................... 3 data_args_path .................................. None data_cache_path ................................. None data_parallel_random_init ....................... False data_parallel_sharding_strategy ................. no_shard data_parallel_size .............................. 1 data_path ....................................... None data_per_class_fraction ......................... 1.0 data_sharding ................................... True dataloader_type ................................. single ddp_average_in_collective ....................... False ddp_bucket_size ................................. None ddp_num_buckets ................................. None ddp_pad_buckets_for_high_nccl_busbw ............. False decoder_first_pipeline_num_layers ............... None decoder_last_pipeline_num_layers ................ None decoder_num_layers .............................. None decoder_seq_length .............................. None decoupled_lr .................................... None decoupled_min_lr ................................ None decrease_batch_size_if_needed ................... False defer_embedding_wgrad_compute ................... False deprecated_use_mcore_models ..................... False deterministic_mode .............................. False dino_bottleneck_size ............................ 256 dino_freeze_last_layer .......................... 1 dino_head_hidden_size ........................... 2048 dino_local_crops_number ......................... 10 dino_local_img_size ............................. 96 dino_norm_last_layer ............................ False dino_teacher_temp ............................... 0.07 dino_warmup_teacher_temp ........................ 0.04 dino_warmup_teacher_temp_epochs ................. 30 disable_straggler_on_startup .................... False dist_ckpt_format_deprecated ..................... None dist_ckpt_strictness ............................ assume_ok_unexpected distribute_saved_activations .................... False distributed_backend ............................. nccl distributed_timeout_minutes ..................... 10 embedding_path .................................. None empty_unused_memory_level ....................... 0 enable_cuda_graph ............................... False enable_ft_package ............................... False enable_gloo_process_groups ...................... True enable_one_logger ............................... True encoder_num_layers .............................. 9 encoder_pipeline_model_parallel_size ............ 0 encoder_seq_length .............................. 2048 encoder_tensor_model_parallel_size .............. 0 end_weight_decay ................................ 0.01 eod_mask_loss ................................... False error_injection_rate ............................ 0 error_injection_type ............................ transient_error eval_interval ................................... 1000 eval_iters ...................................... 100 evidence_data_path .............................. None exit_duration_in_mins ........................... None exit_interval ................................... None exit_on_missing_checkpoint ...................... False exit_signal_handler ............................. False exp_avg_dtype ................................... torch.float32 exp_avg_sq_dtype ................................ torch.float32 expert_model_parallel_size ...................... 1 expert_tensor_parallel_size ..................... 1 ffn_hidden_size ................................. 1368 finetune ........................................ False flash_decode .................................... False fp16 ............................................ False fp16_lm_cross_entropy ........................... False fp32_residual_connection ........................ False fp8 ............................................. None fp8_amax_compute_algo ........................... most_recent fp8_amax_history_len ............................ 1 fp8_interval .................................... 1 fp8_margin ...................................... 0 fp8_param_gather ................................ False fp8_wgrad ....................................... True global_batch_size ............................... 32 grad_reduce_in_bf16 ............................. False gradient_accumulation_fusion .................... False gradient_reduce_div_fusion ...................... True group_query_attention ........................... False head_lr_mult .................................... 1.0 hidden_dropout .................................. 0.0 hidden_size ..................................... 256 hierarchical_context_parallel_sizes ............. None hybrid_attention_ratio .......................... 0.0 hybrid_mlp_ratio ................................ 0.0 hybrid_override_pattern ......................... None hysteresis ...................................... 2 ict_head_size ................................... None ict_load ........................................ None img_h ........................................... 224 img_w ........................................... 224 indexer_batch_size .............................. 128 indexer_log_interval ............................ 1000 inference_batch_times_seqlen_threshold .......... -1 inference_max_batch_size ........................ 8 inference_max_seq_length ........................ 2560 inference_rng_tracker ........................... False init_method_std ................................. 0.02 init_method_xavier_uniform ...................... False init_model_with_meta_device ..................... False initial_loss_scale .............................. 4294967296 iter_per_epoch .................................. 1250 iterations_to_skip .............................. [] keep_fp8_transpose_cache_when_using_custom_fsdp . False kv_channels ..................................... 16 kv_lora_rank .................................... 32 lazy_mpu_init ................................... None load ............................................ /workspace/FLAME-MoE/results/phase0/runs/flame38m_g3_moe_s3/ckpt local_rank ...................................... 0 log_interval .................................... 100 log_loss_scale_to_tensorboard ................... True log_memory_to_tensorboard ....................... False log_num_zeros_in_grad ........................... False log_params_norm ................................. False log_progress .................................... False log_straggler ................................... False log_throughput .................................. False log_timers_to_tensorboard ....................... False log_validation_ppl_to_tensorboard ............... False log_world_size_to_tensorboard ................... False logging_level ................................... None loss_scale ...................................... None loss_scale_window ............................... 1000 lr .............................................. None lr_decay_iters .................................. None lr_decay_samples ................................ None lr_decay_style .................................. linear lr_warmup_fraction .............................. None lr_warmup_init .................................. 0.0 lr_warmup_iters ................................. 0 lr_warmup_samples ............................... 0 lr_wsd_decay_iters .............................. None lr_wsd_decay_samples ............................ None lr_wsd_decay_style .............................. exponential main_grads_dtype ................................ torch.float32 main_params_dtype ............................... torch.float32 make_vocab_size_divisible_by .................... 128 manual_gc ....................................... False manual_gc_eval .................................. True manual_gc_interval .............................. 0 mask_factor ..................................... 1.0 mask_prob ....................................... 0.15 mask_type ....................................... random masked_softmax_fusion ........................... True max_position_embeddings ......................... 2048 max_tokens_to_oom ............................... 10000000 memory_snapshot_path ............................ snapshot.pickle merge_file ...................................... None micro_batch_size ................................ 32 microbatch_group_size_per_vp_stage .............. None min_loss_scale .................................. 1.0 min_lr .......................................... 0.0 mmap_bin_files .................................. True mock_data ....................................... False moe_aux_loss_coeff .............................. 0.01 moe_enable_deepep ............................... False moe_expert_capacity_factor ...................... None moe_extended_tp ................................. False moe_ffn_hidden_size ............................. 58 moe_grouped_gemm ................................ True moe_input_jitter_eps ............................ None moe_layer_freq .................................. [0, 1, 1, 1, 1, 1, 1, 1, 1] moe_layer_recompute ............................. False moe_pad_expert_input_to_capacity ................ False moe_per_layer_logging ........................... False moe_permute_fusion .............................. False moe_router_bias_update_rate ..................... 0.001 moe_router_dtype ................................ fp32 moe_router_enable_expert_bias ................... False moe_router_group_topk ........................... None moe_router_load_balancing_type .................. aux_loss moe_router_num_groups ........................... None moe_router_pre_softmax .......................... True moe_router_score_function ....................... softmax moe_router_topk ................................. 18 moe_router_topk_scaling_factor .................. None moe_shared_expert_intermediate_size ............. 352 moe_shared_expert_overlap ....................... False moe_token_dispatcher_type ....................... alltoall moe_token_drop_policy ........................... probs moe_use_legacy_grouped_gemm ..................... False moe_use_upcycling ............................... False moe_z_loss_coeff ................................ 0.001 mscale .......................................... 1.0 mscale_all_dim .................................. 1.0 multi_latent_attention .......................... False nccl_communicator_config_path ................... None no_load_optim ................................... True no_load_rng ..................................... True no_persist_layer_norm ........................... False no_save_optim ................................... None no_save_rng ..................................... None non_persistent_ckpt_type ........................ None non_persistent_global_ckpt_dir .................. None non_persistent_local_ckpt_algo .................. fully_parallel non_persistent_local_ckpt_dir ................... None non_persistent_save_interval .................... None norm_epsilon .................................... 1e-06 normalization ................................... RMSNorm num_attention_heads ............................. 16 num_channels .................................... 3 num_classes ..................................... 1000 num_dataset_builder_threads ..................... 1 num_distributed_optimizer_instances ............. 1 num_experts ..................................... 192 num_layers ...................................... 9 num_layers_per_virtual_pipeline_stage ........... None num_query_groups ................................ 1 num_virtual_stages_per_pipeline_rank ............ None num_workers ..................................... 2 one_logger_async ................................ False one_logger_project .............................. megatron-lm one_logger_run_name ............................. None onnx_safe ....................................... None openai_gelu ..................................... False optimizer ....................................... adam optimizer_cpu_offload ........................... False optimizer_offload_fraction ...................... 1.0 output_bert_embeddings .......................... False overlap_cpu_optimizer_d2h_h2d ................... False overlap_grad_reduce ............................. False overlap_p2p_comm ................................ False overlap_p2p_comm_warmup_flush ................... False overlap_param_gather ............................ False overlap_param_gather_with_optimizer_step ........ False override_opt_param_scheduler .................... False params_dtype .................................... torch.bfloat16 patch_dim ....................................... 16 per_split_data_args_path ........................ None perform_initialization .......................... True pin_cpu_grads ................................... True pin_cpu_params .................................. True pipeline_model_parallel_comm_backend ............ None pipeline_model_parallel_size .................... 1 pipeline_model_parallel_split_rank .............. None position_embedding_type ......................... rope pretrained_checkpoint ........................... None profile ......................................... False profile_ranks ................................... [0] profile_step_end ................................ 12 profile_step_start .............................. 10 q_lora_rank ..................................... None qk_head_dim ..................................... 128 qk_layernorm .................................... False qk_pos_emb_head_dim ............................. 64 query_in_block_prob ............................. 0.1 rampup_batch_size ............................... None rank ............................................ 0 recompute_granularity ........................... None recompute_method ................................ None recompute_num_layers ............................ None record_memory_history ........................... False relative_attention_max_distance ................. 128 relative_attention_num_buckets .................. 32 replication ..................................... False replication_factor .............................. 2 replication_jump ................................ None rerun_mode ...................................... disabled reset_attention_mask ............................ False reset_position_ids .............................. False result_rejected_tracker_filename ................ None retriever_report_topk_accuracies ................ [] retriever_score_scaling ......................... False retriever_seq_length ............................ 256 retro_add_retriever ............................. False retro_attention_gate ............................ 1 retro_cyclic_train_iters ........................ None retro_encoder_attention_dropout ................. 0.1 retro_encoder_hidden_dropout .................... 0.1 retro_encoder_layers ............................ 2 retro_num_neighbors ............................. 2 retro_num_retrieved_chunks ...................... 2 retro_project_dir ............................... None retro_verify_neighbor_count ..................... True rope_scaling_factor ............................. 8.0 rotary_base ..................................... 10000 rotary_interleaved .............................. False rotary_percent .................................. 1.0 rotary_scaling_factor ........................... 1.0 rotary_seq_len_interpolation_factor ............. None s3_cache_path ................................... None sample_rate ..................................... 1.0 save ............................................ None save_interval ................................... None scatter_gather_tensors_in_pipeline .............. True seed ............................................ 42 seq_length ...................................... 2048 sequence_parallel ............................... False sgd_momentum .................................... 0.9 short_seq_prob .................................. 0.1 skip_train ...................................... False skipped_train_samples ........................... 0 spec ............................................ None split ........................................... None squared_relu .................................... False start_weight_decay .............................. 0.01 straggler_ctrlr_port ............................ 65535 straggler_minmax_count .......................... 1 suggested_communication_unit_size ............... 400000000 swiglu .......................................... True swin_backbone_type .............................. tiny te_rng_tracker .................................. False tensor_model_parallel_size ...................... 1 tensorboard_dir ................................. None tensorboard_log_interval ........................ 1 tensorboard_queue_size .......................... 1000 test_data_path .................................. None test_mode ....................................... False tiktoken_num_special_tokens ..................... 1000 tiktoken_pattern ................................ None tiktoken_special_tokens ......................... None timing_log_level ................................ 0 timing_log_option ............................... minmax titles_data_path ................................ None tokenizer_model ................................. EleutherAI/pythia-12b tokenizer_type .................................. HuggingFaceTokenizer tp_comm_bootstrap_backend ....................... nccl tp_comm_bulk_dgrad .............................. True tp_comm_bulk_wgrad .............................. True tp_comm_overlap ................................. False tp_comm_overlap_ag .............................. True tp_comm_overlap_cfg ............................. None tp_comm_overlap_rs .............................. True tp_comm_overlap_rs_dgrad ........................ False tp_comm_split_ag ................................ True tp_comm_split_rs ................................ True train_data_path ................................. None train_iters ..................................... None train_samples ................................... None train_sync_interval ............................. None transformer_impl ................................ transformer_engine transformer_pipeline_model_parallel_size ........ 1 untie_embeddings_and_output_weights ............. True use_checkpoint_args ............................. False use_checkpoint_opt_param_scheduler .............. False use_cpu_initialization .......................... None use_custom_fsdp ................................. False use_dist_ckpt ................................... True use_dist_ckpt_deprecated ........................ False use_distributed_optimizer ....................... True use_flash_attn .................................. False use_legacy_models ............................... False use_mp_args_from_checkpoint_args ................ False use_one_sent_docs ............................... False use_persistent_ckpt_worker ...................... False use_precision_aware_optimizer ................... False use_pytorch_profiler ............................ False use_ring_exchange_p2p ........................... False use_rope_scaling ................................ False use_rotary_position_embeddings .................. False use_tokenizer_model_from_checkpoint_args ........ True use_torch_fsdp2 ................................. False use_torch_optimizer_for_cpu_offload ............. False use_tp_pp_dp_mapping ............................ False v_head_dim ...................................... 128 valid_data_path ................................. None variable_seq_lengths ............................ False virtual_pipeline_model_parallel_size ............ None vision_backbone_type ............................ vit vision_pretraining .............................. False vision_pretraining_type ......................... classify vocab_extra_ids ................................. 0 vocab_file ...................................... None vocab_size ...................................... None wandb_exp_name .................................. wandb_project ................................... wandb_save_dir .................................. weight_decay .................................... 0.01 weight_decay_incr_style ......................... constant wgrad_deferral_limit ............................ 0 world_size ...................................... 1 yaml_cfg ........................................ None -------------------- end of arguments --------------------- 2026-07-23:03:54:24,213 INFO [num_microbatches_calculator.py:228] setting number of microbatches to constant 1 > building HuggingFaceTokenizer tokenizer ... 2026-07-23:03:54:24,404 INFO [_client.py:1038] HTTP Request: HEAD https://huggingface.co/EleutherAI/pythia-12b/resolve/main/config.json "HTTP/1.1 307 Temporary Redirect" 2026-07-23:03:54:24,416 INFO [_client.py:1038] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/EleutherAI/pythia-12b/bb1e3e710cdf6b524461d543cfb5ba773f0a81b6/config.json "HTTP/1.1 200 OK" 2026-07-23:03:54:24,551 INFO [_client.py:1038] HTTP Request: HEAD https://huggingface.co/EleutherAI/pythia-12b/resolve/main/tokenizer_config.json "HTTP/1.1 307 Temporary Redirect" 2026-07-23:03:54:24,564 INFO [_client.py:1038] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/EleutherAI/pythia-12b/bb1e3e710cdf6b524461d543cfb5ba773f0a81b6/tokenizer_config.json "HTTP/1.1 200 OK" 2026-07-23:03:54:24,667 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/models/EleutherAI/pythia-12b/tree/main/additional_chat_templates?recursive=false&expand=false "HTTP/1.1 404 Not Found" 2026-07-23:03:54:24,733 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/models/EleutherAI/pythia-12b/tree/main?recursive=true&expand=false "HTTP/1.1 200 OK" > padded vocab (size: 50277) with 27 dummy tokens (new size: 50304) WARNING: one_logger package is required to enable e2e metrics tracking. please go to https://confluence.nvidia.com/display/MLWFO/Package+Repositories for details to install it 2026-07-23:03:54:24,944 WARNING [rerun_state_machine.py:239] RerunStateMachine initialized in mode RerunMode.DISABLED > initializing torch distributed ... [W723 03:54:24.462009380 CUDAAllocatorConfig.h:28] Warning: expandable_segments not supported on this platform (function operator()) > initialized tensor model parallel with size 1 > initialized pipeline model parallel with size 1 > setting random seeds to 42 ... > compiling dataset index builder ... make: Entering directory '/workspace/FLAME-MoE/Megatron-LM/megatron/core/datasets' make: Nothing to be done for 'default'. make: Leaving directory '/workspace/FLAME-MoE/Megatron-LM/megatron/core/datasets' >>> done with dataset index builder. Compilation time: 0.728 seconds > compiling and loading fused kernels ... >>> done with compiling and loading fused kernels. Compilation time: 0.387 seconds building GPT model ... > number of parameters on (tensor, pipeline) model parallel rank (0, 0): 100145920 /workspace/FLAME-MoE/Megatron-LM/megatron/core/transformer/transformer_layer.py:339: UserWarning: TransformerLayer._get_layer_offset is deprecated.Please use get_transformer_layer_offset instead. warnings.warn( loading distributed checkpoint from /workspace/FLAME-MoE/results/phase0/runs/flame38m_g3_moe_s3/ckpt at iteration 2121 /workspace/FLAME-MoE/Megatron-LM/megatron/core/dist_checkpointing/strategies/torch.py:847: FutureWarning: `load_state_dict` is deprecated and will be removed in future versions. Please use `load` instead. checkpoint.load_state_dict( /usr/local/lib/python3.11/dist-packages/torch/distributed/checkpoint/planner_helpers.py:306: FutureWarning: Please use DTensor instead and we are deprecating ShardedTensor. device = getattr(value, "device", None) /workspace/FLAME-MoE/.venv/lib/python3.11/site-packages/transformer_engine/pytorch/module/base.py:566: FutureWarning: You are using `torch.load` with `weights_only=False` (the current default value), which uses the default pickle module implicitly. It is possible to construct malicious pickle data which will execute arbitrary code during unpickling (See https://github.com/pytorch/pytorch/blob/main/SECURITY.md#untrusted-models for more details). In a future release, the default value for `weights_only` will be flipped to `True`. This limits the functions that could be executed during unpickling. Arbitrary objects will no longer be allowed to be loaded via this mode unless they are explicitly allowlisted by the user via `torch.serialization.add_safe_globals`. We recommend you start setting `weights_only=True` for any use case where you don't have full control of the loaded file. Please open an issue on GitHub for any issues related to this experimental feature. state = torch.load(state, map_location="cuda") checkpoint version 3.0 2026-07-23:03:54:31,692 WARNING [rerun_state_machine.py:811] RerunStateMachine disabled via CLI, ignoring machine state saved in checkpoint successfully loaded checkpoint from /workspace/FLAME-MoE/results/phase0/runs/flame38m_g3_moe_s3/ckpt [ t 1/1, p 1/1 ] at iteration 2121 2026-07-23:03:54:31,808 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc "HTTP/1.1 200 OK" 2026-07-23:03:54:32,047 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc "HTTP/1.1 200 OK" 2026-07-23:03:54:32,112 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc/revision/210d026faf9955653af8916fad021475a3f00453 "HTTP/1.1 200 OK" 2026-07-23:03:54:32,178 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc/tree/210d026faf9955653af8916fad021475a3f00453?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:54:32,493 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc/tree/210d026faf9955653af8916fad021475a3f00453/ARC-Challenge?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:54:32,980 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc "HTTP/1.1 200 OK" 2026-07-23:03:54:33,219 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc "HTTP/1.1 200 OK" 2026-07-23:03:54:33,453 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc/tree/210d026faf9955653af8916fad021475a3f00453/ARC-Easy?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:54:33,591 WARNING [task.py:799] [Task: boolq] metric acc is defined, but aggregation is not. using default aggregation=mean 2026-07-23:03:54:33,592 WARNING [task.py:811] [Task: boolq] metric acc is defined, but higher_is_better is not. using default higher_is_better=True 2026-07-23:03:54:33,731 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue "HTTP/1.1 200 OK" 2026-07-23:03:54:33,972 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue "HTTP/1.1 200 OK" 2026-07-23:03:54:34,033 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/revision/3de24cf8022e94f4ee4b9d55a6f539891524d646 "HTTP/1.1 200 OK" 2026-07-23:03:54:34,093 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/tree/3de24cf8022e94f4ee4b9d55a6f539891524d646?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:54:34,439 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/tree/3de24cf8022e94f4ee4b9d55a6f539891524d646/axb?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:54:34,501 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/revision/3de24cf8022e94f4ee4b9d55a6f539891524d646 "HTTP/1.1 200 OK" 2026-07-23:03:54:34,577 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/tree/3de24cf8022e94f4ee4b9d55a6f539891524d646/boolq?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:54:34,823 WARNING [task.py:799] [Task: copa] metric acc is defined, but aggregation is not. using default aggregation=mean 2026-07-23:03:54:34,823 WARNING [task.py:811] [Task: copa] metric acc is defined, but higher_is_better is not. using default higher_is_better=True 2026-07-23:03:54:34,887 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue "HTTP/1.1 200 OK" 2026-07-23:03:54:35,133 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue "HTTP/1.1 200 OK" 2026-07-23:03:54:35,720 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/tree/3de24cf8022e94f4ee4b9d55a6f539891524d646/copa?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:54:35,861 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag "HTTP/1.1 200 OK" 2026-07-23:03:54:36,091 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag "HTTP/1.1 200 OK" 2026-07-23:03:54:36,151 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag/revision/218ec52e09a7e7462a5400043bb9a69a41d06b76 "HTTP/1.1 200 OK" 2026-07-23:03:54:36,214 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag/tree/218ec52e09a7e7462a5400043bb9a69a41d06b76?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:54:36,491 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag/tree/218ec52e09a7e7462a5400043bb9a69a41d06b76/data?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:54:36,565 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag/revision/218ec52e09a7e7462a5400043bb9a69a41d06b76 "HTTP/1.1 200 OK" 2026-07-23:03:54:38,580 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai "HTTP/1.1 200 OK" 2026-07-23:03:54:38,836 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai "HTTP/1.1 200 OK" 2026-07-23:03:54:38,896 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/revision/900124bf3b8235c6daf21033af9948b3f07346c4 "HTTP/1.1 200 OK" 2026-07-23:03:54:38,953 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/tree/900124bf3b8235c6daf21033af9948b3f07346c4?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:54:39,166 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/tree/900124bf3b8235c6daf21033af9948b3f07346c4/default%2Ftest?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:54:39,231 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/tree/900124bf3b8235c6daf21033af9948b3f07346c4/default?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:54:39,315 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/revision/900124bf3b8235c6daf21033af9948b3f07346c4 "HTTP/1.1 200 OK" 2026-07-23:03:54:39,334 WARNING [task.py:325] [Task: lambada_openai] has_training_docs and has_validation_docs are False, using test_docs as fewshot_docs but this is not recommended. 2026-07-23:03:54:39,334 WARNING [task.py:325] [Task: lambada_openai] has_training_docs and has_validation_docs are False, using test_docs as fewshot_docs but this is not recommended. 2026-07-23:03:54:39,499 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa "HTTP/1.1 200 OK" 2026-07-23:03:54:39,743 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa "HTTP/1.1 200 OK" 2026-07-23:03:54:39,853 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/revision/388097ea7776314e93a529163e0fea805b8a6454 "HTTP/1.1 200 OK" 2026-07-23:03:54:39,914 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/tree/388097ea7776314e93a529163e0fea805b8a6454?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:54:40,159 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/tree/388097ea7776314e93a529163e0fea805b8a6454/additional?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:54:40,231 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/revision/388097ea7776314e93a529163e0fea805b8a6454 "HTTP/1.1 200 OK" 2026-07-23:03:54:40,313 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/tree/388097ea7776314e93a529163e0fea805b8a6454/main?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:54:40,523 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa "HTTP/1.1 200 OK" 2026-07-23:03:54:40,581 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa "HTTP/1.1 200 OK" 2026-07-23:03:54:40,641 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa/tree/main?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:54:40,705 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/datasets/ybisk/piqa/resolve/main/piqa.py "HTTP/1.1 307 Temporary Redirect" 2026-07-23:03:54:40,715 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/resolve-cache/datasets/ybisk/piqa/2e8ac2dffd59bac8c3c6714948f4c551a0848bb0/piqa.py?%2Fdatasets%2Fybisk%2Fpiqa%2Fresolve%2Fmain%2Fpiqa.py=&etag=%220f0d96172f9bf7aa10069374b3236ef083cfeb8e%22 "HTTP/1.1 200 OK" 2026-07-23:03:54:41,249 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa/revision/main "HTTP/1.1 200 OK" 2026-07-23:03:54:41,309 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa/tree/main?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:54:41,739 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq "HTTP/1.1 200 OK" 2026-07-23:03:54:41,971 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq "HTTP/1.1 200 OK" 2026-07-23:03:54:42,032 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq/revision/2c94ad3e1aafab77146f384e23536f97a4849815 "HTTP/1.1 200 OK" 2026-07-23:03:54:42,136 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq/tree/2c94ad3e1aafab77146f384e23536f97a4849815?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:54:42,372 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq/tree/2c94ad3e1aafab77146f384e23536f97a4849815/data?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:54:42,434 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq/revision/2c94ad3e1aafab77146f384e23536f97a4849815 "HTTP/1.1 200 OK" 2026-07-23:03:54:42,886 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande "HTTP/1.1 200 OK" 2026-07-23:03:54:43,131 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande "HTTP/1.1 200 OK" 2026-07-23:03:54:43,190 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/revision/01e74176c63542e6b0bcb004dcdea22d94fb67b5 "HTTP/1.1 200 OK" 2026-07-23:03:54:43,249 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/tree/01e74176c63542e6b0bcb004dcdea22d94fb67b5?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:54:43,492 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/tree/01e74176c63542e6b0bcb004dcdea22d94fb67b5/winogrande_debiased?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:54:43,556 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/revision/01e74176c63542e6b0bcb004dcdea22d94fb67b5 "HTTP/1.1 200 OK" 2026-07-23:03:54:43,641 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/tree/01e74176c63542e6b0bcb004dcdea22d94fb67b5/winogrande_xl?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:03:54:44,286 WARNING [evaluator.py:270] Overwriting default num_fewshot of winogrande from None to 0 2026-07-23:03:54:44,286 WARNING [evaluator.py:270] Overwriting default num_fewshot of sciq from None to 0 2026-07-23:03:54:44,286 WARNING [evaluator.py:270] Overwriting default num_fewshot of piqa from None to 0 2026-07-23:03:54:44,286 WARNING [evaluator.py:270] Overwriting default num_fewshot of openbookqa from None to 0 2026-07-23:03:54:44,286 WARNING [evaluator.py:270] Overwriting default num_fewshot of lambada_openai from None to 0 2026-07-23:03:54:44,286 WARNING [evaluator.py:270] Overwriting default num_fewshot of hellaswag from None to 0 2026-07-23:03:54:44,286 WARNING [evaluator.py:270] Overwriting default num_fewshot of copa from None to 0 2026-07-23:03:54:44,286 WARNING [evaluator.py:270] Overwriting default num_fewshot of boolq from None to 0 2026-07-23:03:54:44,286 WARNING [evaluator.py:270] Overwriting default num_fewshot of arc_easy from None to 0 2026-07-23:03:54:44,286 WARNING [evaluator.py:270] Overwriting default num_fewshot of arc_challenge from None to 0 2026-07-23:03:54:44,323 INFO [task.py:415] Building contexts for winogrande on rank 0... 0%| | 0/1267 [00:00 /workspace/FLAME-MoE/results/ablations/flame1e18_downstream.csv ==================================================================== [run] flame38m_g3_temporal mode=temporal TEMPORAL=1 H=256 heads=16 ffn=1368 experts=192 topk=18 moe_ffn=58 shared=352 [run] tasks=arc_challenge,arc_easy,hellaswag,openbookqa,piqa,winogrande,boolq,copa,lambada_openai,sciq out=/workspace/FLAME-MoE/results/phase0/runs/flame38m_g3_temporal/lmeval_1e18_0shot ==================================================================== /usr/local/lib/python3.11/dist-packages/torch/cuda/__init__.py:58: FutureWarning: The pynvml package is deprecated. Please install nvidia-ml-py instead. If you did not install pynvml directly, please report this to the maintainers of the package that installed pynvml for you. import pynvml # type: ignore[import] [run_lmeval] stubbed VLM backends (hf_vlms, vllm_vlms) /usr/local/lib/python3.11/dist-packages/torch/cuda/__init__.py:58: FutureWarning: The pynvml package is deprecated. Please install nvidia-ml-py instead. If you did not install pynvml directly, please report this to the maintainers of the package that installed pynvml for you. import pynvml # type: ignore[import] /usr/local/lib/python3.11/dist-packages/torch/cuda/__init__.py:58: FutureWarning: The pynvml package is deprecated. Please install nvidia-ml-py instead. If you did not install pynvml directly, please report this to the maintainers of the package that installed pynvml for you. import pynvml # type: ignore[import] /workspace/FLAME-MoE/.venv/lib/python3.11/site-packages/transformer_engine/pytorch/attention.py:108: UserWarning: To use flash-attn v3, please use the following commands to install: (1) pip install "git+https://github.com/Dao-AILab/flash-attention.git#egg=flashattn-hopper&subdirectory=hopper" (2) python_path=`python -c "import site; print(site.getsitepackages()[0])"` (3) mkdir -p $python_path/flashattn_hopper (4) wget -P $python_path/flashattn_hopper https://raw.githubusercontent.com/Dao-AILab/flash-attention/main/hopper/flash_attn_interface.py warnings.warn( [temporal] rolling-residency router installed (evict=min_logit) [run_lmeval] installed temporal residency router (native regime) /usr/local/lib/python3.11/dist-packages/requests/__init__.py:113: RequestsDependencyWarning: urllib3 (2.2.3) or chardet (6.0.0.post1)/charset_normalizer (3.3.2) doesn't match a supported version! warnings.warn( 2026-07-23:04:04:05,111 INFO [__main__.py:279] Verbosity set to INFO 2026-07-23:04:04:33,904 INFO [__init__.py:459] The tag 'arc_ca' is already registered as a group, this tag will not be registered. This may affect tasks you want to call. 2026-07-23:04:04:33,919 INFO [__init__.py:459] The tag 'arc_ca' is already registered as a group, this tag will not be registered. This may affect tasks you want to call. 2026-07-23:04:04:49,337 INFO [__main__.py:376] Selected Tasks: ['arc_challenge', 'arc_easy', 'boolq', 'copa', 'hellaswag', 'lambada_openai', 'openbookqa', 'piqa', 'sciq', 'winogrande'] 2026-07-23:04:04:49,355 INFO [evaluator.py:164] Setting random seed to 42 | Setting numpy seed to 42 | Setting torch manual seed to 42 | Setting fewshot manual seed to 42 2026-07-23:04:04:49,355 INFO [evaluator.py:201] Initializing megatron_lm model, with arguments: {} using world size: 1, data-parallel size: 1, context-parallel size: 1, hierarchical context-parallel sizes: Nonetensor-model-parallel size: 1, encoder-tensor-model-parallel size: 0, pipeline-model-parallel size: 1, encoder-pipeline-model-parallel size: 0 WARNING: overriding default arguments for exit_on_missing_checkpoint:True with exit_on_missing_checkpoint:False setting global batch size to 32 Number of virtual stages per pipeline stage: None accumulate and all-reduce gradients in fp32 for bfloat16 data type. using torch.bfloat16 for parameters ... ------------------------ arguments ------------------------ account_for_embedding_in_pipeline_split ......... False account_for_loss_in_pipeline_split .............. False accumulate_allreduce_grads_in_fp32 .............. True adam_beta1 ...................................... 0.9 adam_beta2 ...................................... 0.999 adam_eps ........................................ 1e-08 add_bias_linear ................................. False add_position_embedding .......................... True add_qkv_bias .................................... False adlr_autoresume ................................. False adlr_autoresume_interval ........................ 1000 align_grad_reduce ............................... True align_param_gather .............................. False app_tag_run_name ................................ None app_tag_run_version ............................. 0.0.0 apply_layernorm_1p .............................. False apply_query_key_layer_scaling ................... False apply_residual_connection_post_layernorm ........ False apply_rope_fusion ............................... True async_save ...................................... None async_tensor_model_parallel_allreduce ........... True attention_backend ............................... AttnBackend.auto attention_dropout ............................... 0.0 attention_softmax_in_fp32 ....................... False auto_detect_ckpt_format ......................... False barrier_with_L1_time ............................ True bert_binary_head ................................ True bert_embedder_type .............................. megatron bert_load ....................................... None bf16 ............................................ True bias_dropout_fusion ............................. True bias_gelu_fusion ................................ False bias_swiglu_fusion .............................. True biencoder_projection_dim ........................ 0 biencoder_shared_query_context_model ............ False block_data_path ................................. None calc_ft_timeouts ................................ False calculate_per_token_loss ........................ False check_for_large_grads ........................... False check_for_nan_in_loss_and_grad .................. True check_for_spiky_loss ............................ False check_weight_hash_across_dp_replicas_interval ... None ckpt_assume_constant_structure .................. False ckpt_convert_format ............................. None ckpt_convert_save ............................... None ckpt_convert_update_legacy_dist_opt_format ...... False ckpt_format ..................................... torch_dist ckpt_fully_parallel_load ........................ False ckpt_fully_parallel_save ........................ True ckpt_fully_parallel_save_deprecated ............. False ckpt_step ....................................... None classes_fraction ................................ 1.0 clip_grad ....................................... 1.0 clone_scatter_output_in_embedding ............... True config_logger_dir ............................... consumed_train_samples .......................... 0 consumed_valid_samples .......................... 0 context_parallel_size ........................... 1 cp_comm_type .................................... ['p2p'] create_attention_mask_in_dataloader ............. True cross_entropy_loss_fusion ....................... False cuda_graph_warmup_steps ......................... 3 data_args_path .................................. None data_cache_path ................................. None data_parallel_random_init ....................... False data_parallel_sharding_strategy ................. no_shard data_parallel_size .............................. 1 data_path ....................................... None data_per_class_fraction ......................... 1.0 data_sharding ................................... True dataloader_type ................................. single ddp_average_in_collective ....................... False ddp_bucket_size ................................. None ddp_num_buckets ................................. None ddp_pad_buckets_for_high_nccl_busbw ............. False decoder_first_pipeline_num_layers ............... None decoder_last_pipeline_num_layers ................ None decoder_num_layers .............................. None decoder_seq_length .............................. None decoupled_lr .................................... None decoupled_min_lr ................................ None decrease_batch_size_if_needed ................... False defer_embedding_wgrad_compute ................... False deprecated_use_mcore_models ..................... False deterministic_mode .............................. False dino_bottleneck_size ............................ 256 dino_freeze_last_layer .......................... 1 dino_head_hidden_size ........................... 2048 dino_local_crops_number ......................... 10 dino_local_img_size ............................. 96 dino_norm_last_layer ............................ False dino_teacher_temp ............................... 0.07 dino_warmup_teacher_temp ........................ 0.04 dino_warmup_teacher_temp_epochs ................. 30 disable_straggler_on_startup .................... False dist_ckpt_format_deprecated ..................... None dist_ckpt_strictness ............................ assume_ok_unexpected distribute_saved_activations .................... False distributed_backend ............................. nccl distributed_timeout_minutes ..................... 10 embedding_path .................................. None empty_unused_memory_level ....................... 0 enable_cuda_graph ............................... False enable_ft_package ............................... False enable_gloo_process_groups ...................... True enable_one_logger ............................... True encoder_num_layers .............................. 9 encoder_pipeline_model_parallel_size ............ 0 encoder_seq_length .............................. 2048 encoder_tensor_model_parallel_size .............. 0 end_weight_decay ................................ 0.01 eod_mask_loss ................................... False error_injection_rate ............................ 0 error_injection_type ............................ transient_error eval_interval ................................... 1000 eval_iters ...................................... 100 evidence_data_path .............................. None exit_duration_in_mins ........................... None exit_interval ................................... None exit_on_missing_checkpoint ...................... False exit_signal_handler ............................. False exp_avg_dtype ................................... torch.float32 exp_avg_sq_dtype ................................ torch.float32 expert_model_parallel_size ...................... 1 expert_tensor_parallel_size ..................... 1 ffn_hidden_size ................................. 1368 finetune ........................................ False flash_decode .................................... False fp16 ............................................ False fp16_lm_cross_entropy ........................... False fp32_residual_connection ........................ False fp8 ............................................. None fp8_amax_compute_algo ........................... most_recent fp8_amax_history_len ............................ 1 fp8_interval .................................... 1 fp8_margin ...................................... 0 fp8_param_gather ................................ False fp8_wgrad ....................................... True global_batch_size ............................... 32 grad_reduce_in_bf16 ............................. False gradient_accumulation_fusion .................... False gradient_reduce_div_fusion ...................... True group_query_attention ........................... False head_lr_mult .................................... 1.0 hidden_dropout .................................. 0.0 hidden_size ..................................... 256 hierarchical_context_parallel_sizes ............. None hybrid_attention_ratio .......................... 0.0 hybrid_mlp_ratio ................................ 0.0 hybrid_override_pattern ......................... None hysteresis ...................................... 2 ict_head_size ................................... None ict_load ........................................ None img_h ........................................... 224 img_w ........................................... 224 indexer_batch_size .............................. 128 indexer_log_interval ............................ 1000 inference_batch_times_seqlen_threshold .......... -1 inference_max_batch_size ........................ 8 inference_max_seq_length ........................ 2560 inference_rng_tracker ........................... False init_method_std ................................. 0.02 init_method_xavier_uniform ...................... False init_model_with_meta_device ..................... False initial_loss_scale .............................. 4294967296 iter_per_epoch .................................. 1250 iterations_to_skip .............................. [] keep_fp8_transpose_cache_when_using_custom_fsdp . False kv_channels ..................................... 16 kv_lora_rank .................................... 32 lazy_mpu_init ................................... None load ............................................ /workspace/FLAME-MoE/results/phase0/runs/flame38m_g3_temporal/ckpt local_rank ...................................... 0 log_interval .................................... 100 log_loss_scale_to_tensorboard ................... True log_memory_to_tensorboard ....................... False log_num_zeros_in_grad ........................... False log_params_norm ................................. False log_progress .................................... False log_straggler ................................... False log_throughput .................................. False log_timers_to_tensorboard ....................... False log_validation_ppl_to_tensorboard ............... False log_world_size_to_tensorboard ................... False logging_level ................................... None loss_scale ...................................... None loss_scale_window ............................... 1000 lr .............................................. None lr_decay_iters .................................. None lr_decay_samples ................................ None lr_decay_style .................................. linear lr_warmup_fraction .............................. None lr_warmup_init .................................. 0.0 lr_warmup_iters ................................. 0 lr_warmup_samples ............................... 0 lr_wsd_decay_iters .............................. None lr_wsd_decay_samples ............................ None lr_wsd_decay_style .............................. exponential main_grads_dtype ................................ torch.float32 main_params_dtype ............................... torch.float32 make_vocab_size_divisible_by .................... 128 manual_gc ....................................... False manual_gc_eval .................................. True manual_gc_interval .............................. 0 mask_factor ..................................... 1.0 mask_prob ....................................... 0.15 mask_type ....................................... random masked_softmax_fusion ........................... True max_position_embeddings ......................... 2048 max_tokens_to_oom ............................... 10000000 memory_snapshot_path ............................ snapshot.pickle merge_file ...................................... None micro_batch_size ................................ 32 microbatch_group_size_per_vp_stage .............. None min_loss_scale .................................. 1.0 min_lr .......................................... 0.0 mmap_bin_files .................................. True mock_data ....................................... False moe_aux_loss_coeff .............................. 0.01 moe_enable_deepep ............................... False moe_expert_capacity_factor ...................... None moe_extended_tp ................................. False moe_ffn_hidden_size ............................. 58 moe_grouped_gemm ................................ True moe_input_jitter_eps ............................ None moe_layer_freq .................................. [0, 1, 1, 1, 1, 1, 1, 1, 1] moe_layer_recompute ............................. False moe_pad_expert_input_to_capacity ................ False moe_per_layer_logging ........................... False moe_permute_fusion .............................. False moe_router_bias_update_rate ..................... 0.001 moe_router_dtype ................................ fp32 moe_router_enable_expert_bias ................... False moe_router_group_topk ........................... None moe_router_load_balancing_type .................. aux_loss moe_router_num_groups ........................... None moe_router_pre_softmax .......................... True moe_router_score_function ....................... softmax moe_router_topk ................................. 18 moe_router_topk_scaling_factor .................. None moe_shared_expert_intermediate_size ............. 352 moe_shared_expert_overlap ....................... False moe_token_dispatcher_type ....................... alltoall moe_token_drop_policy ........................... probs moe_use_legacy_grouped_gemm ..................... False moe_use_upcycling ............................... False moe_z_loss_coeff ................................ 0.001 mscale .......................................... 1.0 mscale_all_dim .................................. 1.0 multi_latent_attention .......................... False nccl_communicator_config_path ................... None no_load_optim ................................... True no_load_rng ..................................... True no_persist_layer_norm ........................... False no_save_optim ................................... None no_save_rng ..................................... None non_persistent_ckpt_type ........................ None non_persistent_global_ckpt_dir .................. None non_persistent_local_ckpt_algo .................. fully_parallel non_persistent_local_ckpt_dir ................... None non_persistent_save_interval .................... None norm_epsilon .................................... 1e-06 normalization ................................... RMSNorm num_attention_heads ............................. 16 num_channels .................................... 3 num_classes ..................................... 1000 num_dataset_builder_threads ..................... 1 num_distributed_optimizer_instances ............. 1 num_experts ..................................... 192 num_layers ...................................... 9 num_layers_per_virtual_pipeline_stage ........... None num_query_groups ................................ 1 num_virtual_stages_per_pipeline_rank ............ None num_workers ..................................... 2 one_logger_async ................................ False one_logger_project .............................. megatron-lm one_logger_run_name ............................. None onnx_safe ....................................... None openai_gelu ..................................... False optimizer ....................................... adam optimizer_cpu_offload ........................... False optimizer_offload_fraction ...................... 1.0 output_bert_embeddings .......................... False overlap_cpu_optimizer_d2h_h2d ................... False overlap_grad_reduce ............................. False overlap_p2p_comm ................................ False overlap_p2p_comm_warmup_flush ................... False overlap_param_gather ............................ False overlap_param_gather_with_optimizer_step ........ False override_opt_param_scheduler .................... False params_dtype .................................... torch.bfloat16 patch_dim ....................................... 16 per_split_data_args_path ........................ None perform_initialization .......................... True pin_cpu_grads ................................... True pin_cpu_params .................................. True pipeline_model_parallel_comm_backend ............ None pipeline_model_parallel_size .................... 1 pipeline_model_parallel_split_rank .............. None position_embedding_type ......................... rope pretrained_checkpoint ........................... None profile ......................................... False profile_ranks ................................... [0] profile_step_end ................................ 12 profile_step_start .............................. 10 q_lora_rank ..................................... None qk_head_dim ..................................... 128 qk_layernorm .................................... False qk_pos_emb_head_dim ............................. 64 query_in_block_prob ............................. 0.1 rampup_batch_size ............................... None rank ............................................ 0 recompute_granularity ........................... None recompute_method ................................ None recompute_num_layers ............................ None record_memory_history ........................... False relative_attention_max_distance ................. 128 relative_attention_num_buckets .................. 32 replication ..................................... False replication_factor .............................. 2 replication_jump ................................ None rerun_mode ...................................... disabled reset_attention_mask ............................ False reset_position_ids .............................. False result_rejected_tracker_filename ................ None retriever_report_topk_accuracies ................ [] retriever_score_scaling ......................... False retriever_seq_length ............................ 256 retro_add_retriever ............................. False retro_attention_gate ............................ 1 retro_cyclic_train_iters ........................ None retro_encoder_attention_dropout ................. 0.1 retro_encoder_hidden_dropout .................... 0.1 retro_encoder_layers ............................ 2 retro_num_neighbors ............................. 2 retro_num_retrieved_chunks ...................... 2 retro_project_dir ............................... None retro_verify_neighbor_count ..................... True rope_scaling_factor ............................. 8.0 rotary_base ..................................... 10000 rotary_interleaved .............................. False rotary_percent .................................. 1.0 rotary_scaling_factor ........................... 1.0 rotary_seq_len_interpolation_factor ............. None s3_cache_path ................................... None sample_rate ..................................... 1.0 save ............................................ None save_interval ................................... None scatter_gather_tensors_in_pipeline .............. True seed ............................................ 42 seq_length ...................................... 2048 sequence_parallel ............................... False sgd_momentum .................................... 0.9 short_seq_prob .................................. 0.1 skip_train ...................................... False skipped_train_samples ........................... 0 spec ............................................ None split ........................................... None squared_relu .................................... False start_weight_decay .............................. 0.01 straggler_ctrlr_port ............................ 65535 straggler_minmax_count .......................... 1 suggested_communication_unit_size ............... 400000000 swiglu .......................................... True swin_backbone_type .............................. tiny te_rng_tracker .................................. False tensor_model_parallel_size ...................... 1 tensorboard_dir ................................. None tensorboard_log_interval ........................ 1 tensorboard_queue_size .......................... 1000 test_data_path .................................. None test_mode ....................................... False tiktoken_num_special_tokens ..................... 1000 tiktoken_pattern ................................ None tiktoken_special_tokens ......................... None timing_log_level ................................ 0 timing_log_option ............................... minmax titles_data_path ................................ None tokenizer_model ................................. EleutherAI/pythia-12b tokenizer_type .................................. HuggingFaceTokenizer tp_comm_bootstrap_backend ....................... nccl tp_comm_bulk_dgrad .............................. True tp_comm_bulk_wgrad .............................. True tp_comm_overlap ................................. False tp_comm_overlap_ag .............................. True tp_comm_overlap_cfg ............................. None tp_comm_overlap_rs .............................. True tp_comm_overlap_rs_dgrad ........................ False tp_comm_split_ag ................................ True tp_comm_split_rs ................................ True train_data_path ................................. None train_iters ..................................... None train_samples ................................... None train_sync_interval ............................. None transformer_impl ................................ transformer_engine transformer_pipeline_model_parallel_size ........ 1 untie_embeddings_and_output_weights ............. True use_checkpoint_args ............................. False use_checkpoint_opt_param_scheduler .............. False use_cpu_initialization .......................... None use_custom_fsdp ................................. False use_dist_ckpt ................................... True use_dist_ckpt_deprecated ........................ False use_distributed_optimizer ....................... True use_flash_attn .................................. False use_legacy_models ............................... False use_mp_args_from_checkpoint_args ................ False use_one_sent_docs ............................... False use_persistent_ckpt_worker ...................... False use_precision_aware_optimizer ................... False use_pytorch_profiler ............................ False use_ring_exchange_p2p ........................... False use_rope_scaling ................................ False use_rotary_position_embeddings .................. False use_tokenizer_model_from_checkpoint_args ........ True use_torch_fsdp2 ................................. False use_torch_optimizer_for_cpu_offload ............. False use_tp_pp_dp_mapping ............................ False v_head_dim ...................................... 128 valid_data_path ................................. None variable_seq_lengths ............................ False virtual_pipeline_model_parallel_size ............ None vision_backbone_type ............................ vit vision_pretraining .............................. False vision_pretraining_type ......................... classify vocab_extra_ids ................................. 0 vocab_file ...................................... None vocab_size ...................................... None wandb_exp_name .................................. wandb_project ................................... wandb_save_dir .................................. weight_decay .................................... 0.01 weight_decay_incr_style ......................... constant wgrad_deferral_limit ............................ 0 world_size ...................................... 1 yaml_cfg ........................................ None -------------------- end of arguments --------------------- 2026-07-23:04:04:49,370 INFO [num_microbatches_calculator.py:228] setting number of microbatches to constant 1 > building HuggingFaceTokenizer tokenizer ... 2026-07-23:04:04:49,565 INFO [_client.py:1038] HTTP Request: HEAD https://huggingface.co/EleutherAI/pythia-12b/resolve/main/config.json "HTTP/1.1 307 Temporary Redirect" 2026-07-23:04:04:49,578 INFO [_client.py:1038] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/EleutherAI/pythia-12b/bb1e3e710cdf6b524461d543cfb5ba773f0a81b6/config.json "HTTP/1.1 200 OK" 2026-07-23:04:04:49,703 INFO [_client.py:1038] HTTP Request: HEAD https://huggingface.co/EleutherAI/pythia-12b/resolve/main/tokenizer_config.json "HTTP/1.1 307 Temporary Redirect" 2026-07-23:04:04:49,716 INFO [_client.py:1038] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/EleutherAI/pythia-12b/bb1e3e710cdf6b524461d543cfb5ba773f0a81b6/tokenizer_config.json "HTTP/1.1 200 OK" 2026-07-23:04:04:49,812 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/models/EleutherAI/pythia-12b/tree/main/additional_chat_templates?recursive=false&expand=false "HTTP/1.1 404 Not Found" 2026-07-23:04:04:49,876 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/models/EleutherAI/pythia-12b/tree/main?recursive=true&expand=false "HTTP/1.1 200 OK" > padded vocab (size: 50277) with 27 dummy tokens (new size: 50304) WARNING: one_logger package is required to enable e2e metrics tracking. please go to https://confluence.nvidia.com/display/MLWFO/Package+Repositories for details to install it 2026-07-23:04:04:50,076 WARNING [rerun_state_machine.py:239] RerunStateMachine initialized in mode RerunMode.DISABLED > initializing torch distributed ... [W723 04:04:50.593571249 CUDAAllocatorConfig.h:28] Warning: expandable_segments not supported on this platform (function operator()) > initialized tensor model parallel with size 1 > initialized pipeline model parallel with size 1 > setting random seeds to 42 ... > compiling dataset index builder ... make: Entering directory '/workspace/FLAME-MoE/Megatron-LM/megatron/core/datasets' make: Nothing to be done for 'default'. make: Leaving directory '/workspace/FLAME-MoE/Megatron-LM/megatron/core/datasets' >>> done with dataset index builder. Compilation time: 0.655 seconds > compiling and loading fused kernels ... >>> done with compiling and loading fused kernels. Compilation time: 0.350 seconds building GPT model ... > number of parameters on (tensor, pipeline) model parallel rank (0, 0): 100145920 /workspace/FLAME-MoE/Megatron-LM/megatron/core/transformer/transformer_layer.py:339: UserWarning: TransformerLayer._get_layer_offset is deprecated.Please use get_transformer_layer_offset instead. warnings.warn( loading distributed checkpoint from /workspace/FLAME-MoE/results/phase0/runs/flame38m_g3_temporal/ckpt at iteration 2121 /workspace/FLAME-MoE/Megatron-LM/megatron/core/dist_checkpointing/strategies/torch.py:847: FutureWarning: `load_state_dict` is deprecated and will be removed in future versions. Please use `load` instead. checkpoint.load_state_dict( /usr/local/lib/python3.11/dist-packages/torch/distributed/checkpoint/planner_helpers.py:306: FutureWarning: Please use DTensor instead and we are deprecating ShardedTensor. device = getattr(value, "device", None) /workspace/FLAME-MoE/.venv/lib/python3.11/site-packages/transformer_engine/pytorch/module/base.py:566: FutureWarning: You are using `torch.load` with `weights_only=False` (the current default value), which uses the default pickle module implicitly. It is possible to construct malicious pickle data which will execute arbitrary code during unpickling (See https://github.com/pytorch/pytorch/blob/main/SECURITY.md#untrusted-models for more details). In a future release, the default value for `weights_only` will be flipped to `True`. This limits the functions that could be executed during unpickling. Arbitrary objects will no longer be allowed to be loaded via this mode unless they are explicitly allowlisted by the user via `torch.serialization.add_safe_globals`. We recommend you start setting `weights_only=True` for any use case where you don't have full control of the loaded file. Please open an issue on GitHub for any issues related to this experimental feature. state = torch.load(state, map_location="cuda") checkpoint version 3.0 2026-07-23:04:04:56,483 WARNING [rerun_state_machine.py:811] RerunStateMachine disabled via CLI, ignoring machine state saved in checkpoint successfully loaded checkpoint from /workspace/FLAME-MoE/results/phase0/runs/flame38m_g3_temporal/ckpt [ t 1/1, p 1/1 ] at iteration 2121 2026-07-23:04:04:56,630 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc "HTTP/1.1 200 OK" 2026-07-23:04:04:56,854 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc "HTTP/1.1 200 OK" 2026-07-23:04:04:56,912 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc/revision/210d026faf9955653af8916fad021475a3f00453 "HTTP/1.1 200 OK" 2026-07-23:04:04:56,970 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc/tree/210d026faf9955653af8916fad021475a3f00453?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:04:57,214 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc/tree/210d026faf9955653af8916fad021475a3f00453/ARC-Challenge?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:04:57,711 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc "HTTP/1.1 200 OK" 2026-07-23:04:04:57,955 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc "HTTP/1.1 200 OK" 2026-07-23:04:04:58,197 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc/tree/210d026faf9955653af8916fad021475a3f00453/ARC-Easy?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:04:58,339 WARNING [task.py:799] [Task: boolq] metric acc is defined, but aggregation is not. using default aggregation=mean 2026-07-23:04:04:58,339 WARNING [task.py:811] [Task: boolq] metric acc is defined, but higher_is_better is not. using default higher_is_better=True 2026-07-23:04:04:58,402 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue "HTTP/1.1 200 OK" 2026-07-23:04:04:58,630 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue "HTTP/1.1 200 OK" 2026-07-23:04:04:58,699 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/revision/3de24cf8022e94f4ee4b9d55a6f539891524d646 "HTTP/1.1 200 OK" 2026-07-23:04:04:58,764 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/tree/3de24cf8022e94f4ee4b9d55a6f539891524d646?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:04:59,053 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/tree/3de24cf8022e94f4ee4b9d55a6f539891524d646/axb?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:04:59,120 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/revision/3de24cf8022e94f4ee4b9d55a6f539891524d646 "HTTP/1.1 200 OK" 2026-07-23:04:04:59,190 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/tree/3de24cf8022e94f4ee4b9d55a6f539891524d646/boolq?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:04:59,433 WARNING [task.py:799] [Task: copa] metric acc is defined, but aggregation is not. using default aggregation=mean 2026-07-23:04:04:59,433 WARNING [task.py:811] [Task: copa] metric acc is defined, but higher_is_better is not. using default higher_is_better=True 2026-07-23:04:04:59,497 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue "HTTP/1.1 200 OK" 2026-07-23:04:04:59,734 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue "HTTP/1.1 200 OK" 2026-07-23:04:05:00,031 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/tree/3de24cf8022e94f4ee4b9d55a6f539891524d646/copa?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:05:00,183 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag "HTTP/1.1 200 OK" 2026-07-23:04:05:00,432 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag "HTTP/1.1 200 OK" 2026-07-23:04:05:00,492 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag/revision/218ec52e09a7e7462a5400043bb9a69a41d06b76 "HTTP/1.1 200 OK" 2026-07-23:04:05:00,549 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag/tree/218ec52e09a7e7462a5400043bb9a69a41d06b76?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:05:00,756 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag/tree/218ec52e09a7e7462a5400043bb9a69a41d06b76/data?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:05:00,817 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag/revision/218ec52e09a7e7462a5400043bb9a69a41d06b76 "HTTP/1.1 200 OK" 2026-07-23:04:05:03,212 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai "HTTP/1.1 200 OK" 2026-07-23:04:05:03,424 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai "HTTP/1.1 200 OK" 2026-07-23:04:05:03,491 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/revision/900124bf3b8235c6daf21033af9948b3f07346c4 "HTTP/1.1 200 OK" 2026-07-23:04:05:03,555 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/tree/900124bf3b8235c6daf21033af9948b3f07346c4?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:05:03,781 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/tree/900124bf3b8235c6daf21033af9948b3f07346c4/default%2Ftest?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:05:03,845 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/tree/900124bf3b8235c6daf21033af9948b3f07346c4/default?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:05:03,907 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/revision/900124bf3b8235c6daf21033af9948b3f07346c4 "HTTP/1.1 200 OK" 2026-07-23:04:05:03,924 WARNING [task.py:325] [Task: lambada_openai] has_training_docs and has_validation_docs are False, using test_docs as fewshot_docs but this is not recommended. 2026-07-23:04:05:03,924 WARNING [task.py:325] [Task: lambada_openai] has_training_docs and has_validation_docs are False, using test_docs as fewshot_docs but this is not recommended. 2026-07-23:04:05:04,075 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa "HTTP/1.1 200 OK" 2026-07-23:04:05:04,305 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa "HTTP/1.1 200 OK" 2026-07-23:04:05:04,437 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/revision/388097ea7776314e93a529163e0fea805b8a6454 "HTTP/1.1 200 OK" 2026-07-23:04:05:04,499 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/tree/388097ea7776314e93a529163e0fea805b8a6454?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:05:04,717 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/tree/388097ea7776314e93a529163e0fea805b8a6454/additional?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:05:04,776 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/revision/388097ea7776314e93a529163e0fea805b8a6454 "HTTP/1.1 200 OK" 2026-07-23:04:05:04,866 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/tree/388097ea7776314e93a529163e0fea805b8a6454/main?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:05:05,153 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa "HTTP/1.1 200 OK" 2026-07-23:04:05:05,217 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa "HTTP/1.1 200 OK" 2026-07-23:04:05:05,283 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa/tree/main?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:05:05,342 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/datasets/ybisk/piqa/resolve/main/piqa.py "HTTP/1.1 307 Temporary Redirect" 2026-07-23:04:05:05,354 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/resolve-cache/datasets/ybisk/piqa/2e8ac2dffd59bac8c3c6714948f4c551a0848bb0/piqa.py?%2Fdatasets%2Fybisk%2Fpiqa%2Fresolve%2Fmain%2Fpiqa.py=&etag=%220f0d96172f9bf7aa10069374b3236ef083cfeb8e%22 "HTTP/1.1 200 OK" 2026-07-23:04:05:05,872 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa/revision/main "HTTP/1.1 200 OK" 2026-07-23:04:05:05,937 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa/tree/main?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:05:06,385 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq "HTTP/1.1 200 OK" 2026-07-23:04:05:06,598 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq "HTTP/1.1 200 OK" 2026-07-23:04:05:06,662 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq/revision/2c94ad3e1aafab77146f384e23536f97a4849815 "HTTP/1.1 200 OK" 2026-07-23:04:05:06,727 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq/tree/2c94ad3e1aafab77146f384e23536f97a4849815?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:05:06,931 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq/tree/2c94ad3e1aafab77146f384e23536f97a4849815/data?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:05:06,989 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq/revision/2c94ad3e1aafab77146f384e23536f97a4849815 "HTTP/1.1 200 OK" 2026-07-23:04:05:07,419 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande "HTTP/1.1 200 OK" 2026-07-23:04:05:07,644 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande "HTTP/1.1 200 OK" 2026-07-23:04:05:07,705 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/revision/01e74176c63542e6b0bcb004dcdea22d94fb67b5 "HTTP/1.1 200 OK" 2026-07-23:04:05:07,764 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/tree/01e74176c63542e6b0bcb004dcdea22d94fb67b5?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:05:07,999 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/tree/01e74176c63542e6b0bcb004dcdea22d94fb67b5/winogrande_debiased?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:05:08,061 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/revision/01e74176c63542e6b0bcb004dcdea22d94fb67b5 "HTTP/1.1 200 OK" 2026-07-23:04:05:08,159 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/tree/01e74176c63542e6b0bcb004dcdea22d94fb67b5/winogrande_xl?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:05:08,818 WARNING [evaluator.py:270] Overwriting default num_fewshot of winogrande from None to 0 2026-07-23:04:05:08,818 WARNING [evaluator.py:270] Overwriting default num_fewshot of sciq from None to 0 2026-07-23:04:05:08,818 WARNING [evaluator.py:270] Overwriting default num_fewshot of piqa from None to 0 2026-07-23:04:05:08,818 WARNING [evaluator.py:270] Overwriting default num_fewshot of openbookqa from None to 0 2026-07-23:04:05:08,818 WARNING [evaluator.py:270] Overwriting default num_fewshot of lambada_openai from None to 0 2026-07-23:04:05:08,818 WARNING [evaluator.py:270] Overwriting default num_fewshot of hellaswag from None to 0 2026-07-23:04:05:08,818 WARNING [evaluator.py:270] Overwriting default num_fewshot of copa from None to 0 2026-07-23:04:05:08,818 WARNING [evaluator.py:270] Overwriting default num_fewshot of boolq from None to 0 2026-07-23:04:05:08,818 WARNING [evaluator.py:270] Overwriting default num_fewshot of arc_easy from None to 0 2026-07-23:04:05:08,818 WARNING [evaluator.py:270] Overwriting default num_fewshot of arc_challenge from None to 0 2026-07-23:04:05:08,848 INFO [task.py:415] Building contexts for winogrande on rank 0... 0%| | 0/1267 [00:00 /workspace/FLAME-MoE/results/ablations/flame1e18_downstream.csv ==================================================================== [run] flame38m_g3_temporal_s2 mode=temporal TEMPORAL=1 H=256 heads=16 ffn=1368 experts=192 topk=18 moe_ffn=58 shared=352 [run] tasks=arc_challenge,arc_easy,hellaswag,openbookqa,piqa,winogrande,boolq,copa,lambada_openai,sciq out=/workspace/FLAME-MoE/results/phase0/runs/flame38m_g3_temporal_s2/lmeval_1e18_0shot ==================================================================== /usr/local/lib/python3.11/dist-packages/torch/cuda/__init__.py:58: FutureWarning: The pynvml package is deprecated. Please install nvidia-ml-py instead. If you did not install pynvml directly, please report this to the maintainers of the package that installed pynvml for you. import pynvml # type: ignore[import] [run_lmeval] stubbed VLM backends (hf_vlms, vllm_vlms) /usr/local/lib/python3.11/dist-packages/torch/cuda/__init__.py:58: FutureWarning: The pynvml package is deprecated. Please install nvidia-ml-py instead. If you did not install pynvml directly, please report this to the maintainers of the package that installed pynvml for you. import pynvml # type: ignore[import] /usr/local/lib/python3.11/dist-packages/torch/cuda/__init__.py:58: FutureWarning: The pynvml package is deprecated. Please install nvidia-ml-py instead. If you did not install pynvml directly, please report this to the maintainers of the package that installed pynvml for you. import pynvml # type: ignore[import] /workspace/FLAME-MoE/.venv/lib/python3.11/site-packages/transformer_engine/pytorch/attention.py:108: UserWarning: To use flash-attn v3, please use the following commands to install: (1) pip install "git+https://github.com/Dao-AILab/flash-attention.git#egg=flashattn-hopper&subdirectory=hopper" (2) python_path=`python -c "import site; print(site.getsitepackages()[0])"` (3) mkdir -p $python_path/flashattn_hopper (4) wget -P $python_path/flashattn_hopper https://raw.githubusercontent.com/Dao-AILab/flash-attention/main/hopper/flash_attn_interface.py warnings.warn( [temporal] rolling-residency router installed (evict=min_logit) [run_lmeval] installed temporal residency router (native regime) /usr/local/lib/python3.11/dist-packages/requests/__init__.py:113: RequestsDependencyWarning: urllib3 (2.2.3) or chardet (6.0.0.post1)/charset_normalizer (3.3.2) doesn't match a supported version! warnings.warn( 2026-07-23:04:14:21,944 INFO [__main__.py:279] Verbosity set to INFO 2026-07-23:04:14:53,378 INFO [__init__.py:459] The tag 'arc_ca' is already registered as a group, this tag will not be registered. This may affect tasks you want to call. 2026-07-23:04:14:53,390 INFO [__init__.py:459] The tag 'arc_ca' is already registered as a group, this tag will not be registered. This may affect tasks you want to call. 2026-07-23:04:15:09,050 INFO [__main__.py:376] Selected Tasks: ['arc_challenge', 'arc_easy', 'boolq', 'copa', 'hellaswag', 'lambada_openai', 'openbookqa', 'piqa', 'sciq', 'winogrande'] 2026-07-23:04:15:09,076 INFO [evaluator.py:164] Setting random seed to 42 | Setting numpy seed to 42 | Setting torch manual seed to 42 | Setting fewshot manual seed to 42 2026-07-23:04:15:09,076 INFO [evaluator.py:201] Initializing megatron_lm model, with arguments: {} using world size: 1, data-parallel size: 1, context-parallel size: 1, hierarchical context-parallel sizes: Nonetensor-model-parallel size: 1, encoder-tensor-model-parallel size: 0, pipeline-model-parallel size: 1, encoder-pipeline-model-parallel size: 0 WARNING: overriding default arguments for exit_on_missing_checkpoint:True with exit_on_missing_checkpoint:False setting global batch size to 32 Number of virtual stages per pipeline stage: None accumulate and all-reduce gradients in fp32 for bfloat16 data type. using torch.bfloat16 for parameters ... ------------------------ arguments ------------------------ account_for_embedding_in_pipeline_split ......... False account_for_loss_in_pipeline_split .............. False accumulate_allreduce_grads_in_fp32 .............. True adam_beta1 ...................................... 0.9 adam_beta2 ...................................... 0.999 adam_eps ........................................ 1e-08 add_bias_linear ................................. False add_position_embedding .......................... True add_qkv_bias .................................... False adlr_autoresume ................................. False adlr_autoresume_interval ........................ 1000 align_grad_reduce ............................... True align_param_gather .............................. False app_tag_run_name ................................ None app_tag_run_version ............................. 0.0.0 apply_layernorm_1p .............................. False apply_query_key_layer_scaling ................... False apply_residual_connection_post_layernorm ........ False apply_rope_fusion ............................... True async_save ...................................... None async_tensor_model_parallel_allreduce ........... True attention_backend ............................... AttnBackend.auto attention_dropout ............................... 0.0 attention_softmax_in_fp32 ....................... False auto_detect_ckpt_format ......................... False barrier_with_L1_time ............................ True bert_binary_head ................................ True bert_embedder_type .............................. megatron bert_load ....................................... None bf16 ............................................ True bias_dropout_fusion ............................. True bias_gelu_fusion ................................ False bias_swiglu_fusion .............................. True biencoder_projection_dim ........................ 0 biencoder_shared_query_context_model ............ False block_data_path ................................. None calc_ft_timeouts ................................ False calculate_per_token_loss ........................ False check_for_large_grads ........................... False check_for_nan_in_loss_and_grad .................. True check_for_spiky_loss ............................ False check_weight_hash_across_dp_replicas_interval ... None ckpt_assume_constant_structure .................. False ckpt_convert_format ............................. None ckpt_convert_save ............................... None ckpt_convert_update_legacy_dist_opt_format ...... False ckpt_format ..................................... torch_dist ckpt_fully_parallel_load ........................ False ckpt_fully_parallel_save ........................ True ckpt_fully_parallel_save_deprecated ............. False ckpt_step ....................................... None classes_fraction ................................ 1.0 clip_grad ....................................... 1.0 clone_scatter_output_in_embedding ............... True config_logger_dir ............................... consumed_train_samples .......................... 0 consumed_valid_samples .......................... 0 context_parallel_size ........................... 1 cp_comm_type .................................... ['p2p'] create_attention_mask_in_dataloader ............. True cross_entropy_loss_fusion ....................... False cuda_graph_warmup_steps ......................... 3 data_args_path .................................. None data_cache_path ................................. None data_parallel_random_init ....................... False data_parallel_sharding_strategy ................. no_shard data_parallel_size .............................. 1 data_path ....................................... None data_per_class_fraction ......................... 1.0 data_sharding ................................... True dataloader_type ................................. single ddp_average_in_collective ....................... False ddp_bucket_size ................................. None ddp_num_buckets ................................. None ddp_pad_buckets_for_high_nccl_busbw ............. False decoder_first_pipeline_num_layers ............... None decoder_last_pipeline_num_layers ................ None decoder_num_layers .............................. None decoder_seq_length .............................. None decoupled_lr .................................... None decoupled_min_lr ................................ None decrease_batch_size_if_needed ................... False defer_embedding_wgrad_compute ................... False deprecated_use_mcore_models ..................... False deterministic_mode .............................. False dino_bottleneck_size ............................ 256 dino_freeze_last_layer .......................... 1 dino_head_hidden_size ........................... 2048 dino_local_crops_number ......................... 10 dino_local_img_size ............................. 96 dino_norm_last_layer ............................ False dino_teacher_temp ............................... 0.07 dino_warmup_teacher_temp ........................ 0.04 dino_warmup_teacher_temp_epochs ................. 30 disable_straggler_on_startup .................... False dist_ckpt_format_deprecated ..................... None dist_ckpt_strictness ............................ assume_ok_unexpected distribute_saved_activations .................... False distributed_backend ............................. nccl distributed_timeout_minutes ..................... 10 embedding_path .................................. None empty_unused_memory_level ....................... 0 enable_cuda_graph ............................... False enable_ft_package ............................... False enable_gloo_process_groups ...................... True enable_one_logger ............................... True encoder_num_layers .............................. 9 encoder_pipeline_model_parallel_size ............ 0 encoder_seq_length .............................. 2048 encoder_tensor_model_parallel_size .............. 0 end_weight_decay ................................ 0.01 eod_mask_loss ................................... False error_injection_rate ............................ 0 error_injection_type ............................ transient_error eval_interval ................................... 1000 eval_iters ...................................... 100 evidence_data_path .............................. None exit_duration_in_mins ........................... None exit_interval ................................... None exit_on_missing_checkpoint ...................... False exit_signal_handler ............................. False exp_avg_dtype ................................... torch.float32 exp_avg_sq_dtype ................................ torch.float32 expert_model_parallel_size ...................... 1 expert_tensor_parallel_size ..................... 1 ffn_hidden_size ................................. 1368 finetune ........................................ False flash_decode .................................... False fp16 ............................................ False fp16_lm_cross_entropy ........................... False fp32_residual_connection ........................ False fp8 ............................................. None fp8_amax_compute_algo ........................... most_recent fp8_amax_history_len ............................ 1 fp8_interval .................................... 1 fp8_margin ...................................... 0 fp8_param_gather ................................ False fp8_wgrad ....................................... True global_batch_size ............................... 32 grad_reduce_in_bf16 ............................. False gradient_accumulation_fusion .................... False gradient_reduce_div_fusion ...................... True group_query_attention ........................... False head_lr_mult .................................... 1.0 hidden_dropout .................................. 0.0 hidden_size ..................................... 256 hierarchical_context_parallel_sizes ............. None hybrid_attention_ratio .......................... 0.0 hybrid_mlp_ratio ................................ 0.0 hybrid_override_pattern ......................... None hysteresis ...................................... 2 ict_head_size ................................... None ict_load ........................................ None img_h ........................................... 224 img_w ........................................... 224 indexer_batch_size .............................. 128 indexer_log_interval ............................ 1000 inference_batch_times_seqlen_threshold .......... -1 inference_max_batch_size ........................ 8 inference_max_seq_length ........................ 2560 inference_rng_tracker ........................... False init_method_std ................................. 0.02 init_method_xavier_uniform ...................... False init_model_with_meta_device ..................... False initial_loss_scale .............................. 4294967296 iter_per_epoch .................................. 1250 iterations_to_skip .............................. [] keep_fp8_transpose_cache_when_using_custom_fsdp . False kv_channels ..................................... 16 kv_lora_rank .................................... 32 lazy_mpu_init ................................... None load ............................................ /workspace/FLAME-MoE/results/phase0/runs/flame38m_g3_temporal_s2/ckpt local_rank ...................................... 0 log_interval .................................... 100 log_loss_scale_to_tensorboard ................... True log_memory_to_tensorboard ....................... False log_num_zeros_in_grad ........................... False log_params_norm ................................. False log_progress .................................... False log_straggler ................................... False log_throughput .................................. False log_timers_to_tensorboard ....................... False log_validation_ppl_to_tensorboard ............... False log_world_size_to_tensorboard ................... False logging_level ................................... None loss_scale ...................................... None loss_scale_window ............................... 1000 lr .............................................. None lr_decay_iters .................................. None lr_decay_samples ................................ None lr_decay_style .................................. linear lr_warmup_fraction .............................. None lr_warmup_init .................................. 0.0 lr_warmup_iters ................................. 0 lr_warmup_samples ............................... 0 lr_wsd_decay_iters .............................. None lr_wsd_decay_samples ............................ None lr_wsd_decay_style .............................. exponential main_grads_dtype ................................ torch.float32 main_params_dtype ............................... torch.float32 make_vocab_size_divisible_by .................... 128 manual_gc ....................................... False manual_gc_eval .................................. True manual_gc_interval .............................. 0 mask_factor ..................................... 1.0 mask_prob ....................................... 0.15 mask_type ....................................... random masked_softmax_fusion ........................... True max_position_embeddings ......................... 2048 max_tokens_to_oom ............................... 10000000 memory_snapshot_path ............................ snapshot.pickle merge_file ...................................... None micro_batch_size ................................ 32 microbatch_group_size_per_vp_stage .............. None min_loss_scale .................................. 1.0 min_lr .......................................... 0.0 mmap_bin_files .................................. True mock_data ....................................... False moe_aux_loss_coeff .............................. 0.01 moe_enable_deepep ............................... False moe_expert_capacity_factor ...................... None moe_extended_tp ................................. False moe_ffn_hidden_size ............................. 58 moe_grouped_gemm ................................ True moe_input_jitter_eps ............................ None moe_layer_freq .................................. [0, 1, 1, 1, 1, 1, 1, 1, 1] moe_layer_recompute ............................. False moe_pad_expert_input_to_capacity ................ False moe_per_layer_logging ........................... False moe_permute_fusion .............................. False moe_router_bias_update_rate ..................... 0.001 moe_router_dtype ................................ fp32 moe_router_enable_expert_bias ................... False moe_router_group_topk ........................... None moe_router_load_balancing_type .................. aux_loss moe_router_num_groups ........................... None moe_router_pre_softmax .......................... True moe_router_score_function ....................... softmax moe_router_topk ................................. 18 moe_router_topk_scaling_factor .................. None moe_shared_expert_intermediate_size ............. 352 moe_shared_expert_overlap ....................... False moe_token_dispatcher_type ....................... alltoall moe_token_drop_policy ........................... probs moe_use_legacy_grouped_gemm ..................... False moe_use_upcycling ............................... False moe_z_loss_coeff ................................ 0.001 mscale .......................................... 1.0 mscale_all_dim .................................. 1.0 multi_latent_attention .......................... False nccl_communicator_config_path ................... None no_load_optim ................................... True no_load_rng ..................................... True no_persist_layer_norm ........................... False no_save_optim ................................... None no_save_rng ..................................... None non_persistent_ckpt_type ........................ None non_persistent_global_ckpt_dir .................. None non_persistent_local_ckpt_algo .................. fully_parallel non_persistent_local_ckpt_dir ................... None non_persistent_save_interval .................... None norm_epsilon .................................... 1e-06 normalization ................................... RMSNorm num_attention_heads ............................. 16 num_channels .................................... 3 num_classes ..................................... 1000 num_dataset_builder_threads ..................... 1 num_distributed_optimizer_instances ............. 1 num_experts ..................................... 192 num_layers ...................................... 9 num_layers_per_virtual_pipeline_stage ........... None num_query_groups ................................ 1 num_virtual_stages_per_pipeline_rank ............ None num_workers ..................................... 2 one_logger_async ................................ False one_logger_project .............................. megatron-lm one_logger_run_name ............................. None onnx_safe ....................................... None openai_gelu ..................................... False optimizer ....................................... adam optimizer_cpu_offload ........................... False optimizer_offload_fraction ...................... 1.0 output_bert_embeddings .......................... False overlap_cpu_optimizer_d2h_h2d ................... False overlap_grad_reduce ............................. False overlap_p2p_comm ................................ False overlap_p2p_comm_warmup_flush ................... False overlap_param_gather ............................ False overlap_param_gather_with_optimizer_step ........ False override_opt_param_scheduler .................... False params_dtype .................................... torch.bfloat16 patch_dim ....................................... 16 per_split_data_args_path ........................ None perform_initialization .......................... True pin_cpu_grads ................................... True pin_cpu_params .................................. True pipeline_model_parallel_comm_backend ............ None pipeline_model_parallel_size .................... 1 pipeline_model_parallel_split_rank .............. None position_embedding_type ......................... rope pretrained_checkpoint ........................... None profile ......................................... False profile_ranks ................................... [0] profile_step_end ................................ 12 profile_step_start .............................. 10 q_lora_rank ..................................... None qk_head_dim ..................................... 128 qk_layernorm .................................... False qk_pos_emb_head_dim ............................. 64 query_in_block_prob ............................. 0.1 rampup_batch_size ............................... None rank ............................................ 0 recompute_granularity ........................... None recompute_method ................................ None recompute_num_layers ............................ None record_memory_history ........................... False relative_attention_max_distance ................. 128 relative_attention_num_buckets .................. 32 replication ..................................... False replication_factor .............................. 2 replication_jump ................................ None rerun_mode ...................................... disabled reset_attention_mask ............................ False reset_position_ids .............................. False result_rejected_tracker_filename ................ None retriever_report_topk_accuracies ................ [] retriever_score_scaling ......................... False retriever_seq_length ............................ 256 retro_add_retriever ............................. False retro_attention_gate ............................ 1 retro_cyclic_train_iters ........................ None retro_encoder_attention_dropout ................. 0.1 retro_encoder_hidden_dropout .................... 0.1 retro_encoder_layers ............................ 2 retro_num_neighbors ............................. 2 retro_num_retrieved_chunks ...................... 2 retro_project_dir ............................... None retro_verify_neighbor_count ..................... True rope_scaling_factor ............................. 8.0 rotary_base ..................................... 10000 rotary_interleaved .............................. False rotary_percent .................................. 1.0 rotary_scaling_factor ........................... 1.0 rotary_seq_len_interpolation_factor ............. None s3_cache_path ................................... None sample_rate ..................................... 1.0 save ............................................ None save_interval ................................... None scatter_gather_tensors_in_pipeline .............. True seed ............................................ 42 seq_length ...................................... 2048 sequence_parallel ............................... False sgd_momentum .................................... 0.9 short_seq_prob .................................. 0.1 skip_train ...................................... False skipped_train_samples ........................... 0 spec ............................................ None split ........................................... None squared_relu .................................... False start_weight_decay .............................. 0.01 straggler_ctrlr_port ............................ 65535 straggler_minmax_count .......................... 1 suggested_communication_unit_size ............... 400000000 swiglu .......................................... True swin_backbone_type .............................. tiny te_rng_tracker .................................. False tensor_model_parallel_size ...................... 1 tensorboard_dir ................................. None tensorboard_log_interval ........................ 1 tensorboard_queue_size .......................... 1000 test_data_path .................................. None test_mode ....................................... False tiktoken_num_special_tokens ..................... 1000 tiktoken_pattern ................................ None tiktoken_special_tokens ......................... None timing_log_level ................................ 0 timing_log_option ............................... minmax titles_data_path ................................ None tokenizer_model ................................. EleutherAI/pythia-12b tokenizer_type .................................. HuggingFaceTokenizer tp_comm_bootstrap_backend ....................... nccl tp_comm_bulk_dgrad .............................. True tp_comm_bulk_wgrad .............................. True tp_comm_overlap ................................. False tp_comm_overlap_ag .............................. True tp_comm_overlap_cfg ............................. None tp_comm_overlap_rs .............................. True tp_comm_overlap_rs_dgrad ........................ False tp_comm_split_ag ................................ True tp_comm_split_rs ................................ True train_data_path ................................. None train_iters ..................................... None train_samples ................................... None train_sync_interval ............................. None transformer_impl ................................ transformer_engine transformer_pipeline_model_parallel_size ........ 1 untie_embeddings_and_output_weights ............. True use_checkpoint_args ............................. False use_checkpoint_opt_param_scheduler .............. False use_cpu_initialization .......................... None use_custom_fsdp ................................. False use_dist_ckpt ................................... True use_dist_ckpt_deprecated ........................ False use_distributed_optimizer ....................... True use_flash_attn .................................. False use_legacy_models ............................... False use_mp_args_from_checkpoint_args ................ False use_one_sent_docs ............................... False use_persistent_ckpt_worker ...................... False use_precision_aware_optimizer ................... False use_pytorch_profiler ............................ False use_ring_exchange_p2p ........................... False use_rope_scaling ................................ False use_rotary_position_embeddings .................. False use_tokenizer_model_from_checkpoint_args ........ True use_torch_fsdp2 ................................. False use_torch_optimizer_for_cpu_offload ............. False use_tp_pp_dp_mapping ............................ False v_head_dim ...................................... 128 valid_data_path ................................. None variable_seq_lengths ............................ False virtual_pipeline_model_parallel_size ............ None vision_backbone_type ............................ vit vision_pretraining .............................. False vision_pretraining_type ......................... classify vocab_extra_ids ................................. 0 vocab_file ...................................... None vocab_size ...................................... None wandb_exp_name .................................. wandb_project ................................... wandb_save_dir .................................. weight_decay .................................... 0.01 weight_decay_incr_style ......................... constant wgrad_deferral_limit ............................ 0 world_size ...................................... 1 yaml_cfg ........................................ None -------------------- end of arguments --------------------- 2026-07-23:04:15:09,093 INFO [num_microbatches_calculator.py:228] setting number of microbatches to constant 1 > building HuggingFaceTokenizer tokenizer ... 2026-07-23:04:15:09,308 INFO [_client.py:1038] HTTP Request: HEAD https://huggingface.co/EleutherAI/pythia-12b/resolve/main/config.json "HTTP/1.1 307 Temporary Redirect" 2026-07-23:04:15:09,320 INFO [_client.py:1038] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/EleutherAI/pythia-12b/bb1e3e710cdf6b524461d543cfb5ba773f0a81b6/config.json "HTTP/1.1 200 OK" 2026-07-23:04:15:09,443 INFO [_client.py:1038] HTTP Request: HEAD https://huggingface.co/EleutherAI/pythia-12b/resolve/main/tokenizer_config.json "HTTP/1.1 307 Temporary Redirect" 2026-07-23:04:15:09,456 INFO [_client.py:1038] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/EleutherAI/pythia-12b/bb1e3e710cdf6b524461d543cfb5ba773f0a81b6/tokenizer_config.json "HTTP/1.1 200 OK" 2026-07-23:04:15:09,551 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/models/EleutherAI/pythia-12b/tree/main/additional_chat_templates?recursive=false&expand=false "HTTP/1.1 404 Not Found" 2026-07-23:04:15:09,613 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/models/EleutherAI/pythia-12b/tree/main?recursive=true&expand=false "HTTP/1.1 200 OK" > padded vocab (size: 50277) with 27 dummy tokens (new size: 50304) WARNING: one_logger package is required to enable e2e metrics tracking. please go to https://confluence.nvidia.com/display/MLWFO/Package+Repositories for details to install it 2026-07-23:04:15:09,802 WARNING [rerun_state_machine.py:239] RerunStateMachine initialized in mode RerunMode.DISABLED > initializing torch distributed ... [W723 04:15:09.319744698 CUDAAllocatorConfig.h:28] Warning: expandable_segments not supported on this platform (function operator()) > initialized tensor model parallel with size 1 > initialized pipeline model parallel with size 1 > setting random seeds to 42 ... > compiling dataset index builder ... make: Entering directory '/workspace/FLAME-MoE/Megatron-LM/megatron/core/datasets' make: Nothing to be done for 'default'. make: Leaving directory '/workspace/FLAME-MoE/Megatron-LM/megatron/core/datasets' >>> done with dataset index builder. Compilation time: 0.649 seconds > compiling and loading fused kernels ... >>> done with compiling and loading fused kernels. Compilation time: 0.330 seconds building GPT model ... > number of parameters on (tensor, pipeline) model parallel rank (0, 0): 100145920 /workspace/FLAME-MoE/Megatron-LM/megatron/core/transformer/transformer_layer.py:339: UserWarning: TransformerLayer._get_layer_offset is deprecated.Please use get_transformer_layer_offset instead. warnings.warn( loading distributed checkpoint from /workspace/FLAME-MoE/results/phase0/runs/flame38m_g3_temporal_s2/ckpt at iteration 2121 /workspace/FLAME-MoE/Megatron-LM/megatron/core/dist_checkpointing/strategies/torch.py:847: FutureWarning: `load_state_dict` is deprecated and will be removed in future versions. Please use `load` instead. checkpoint.load_state_dict( /usr/local/lib/python3.11/dist-packages/torch/distributed/checkpoint/planner_helpers.py:306: FutureWarning: Please use DTensor instead and we are deprecating ShardedTensor. device = getattr(value, "device", None) /workspace/FLAME-MoE/.venv/lib/python3.11/site-packages/transformer_engine/pytorch/module/base.py:566: FutureWarning: You are using `torch.load` with `weights_only=False` (the current default value), which uses the default pickle module implicitly. It is possible to construct malicious pickle data which will execute arbitrary code during unpickling (See https://github.com/pytorch/pytorch/blob/main/SECURITY.md#untrusted-models for more details). In a future release, the default value for `weights_only` will be flipped to `True`. This limits the functions that could be executed during unpickling. Arbitrary objects will no longer be allowed to be loaded via this mode unless they are explicitly allowlisted by the user via `torch.serialization.add_safe_globals`. We recommend you start setting `weights_only=True` for any use case where you don't have full control of the loaded file. Please open an issue on GitHub for any issues related to this experimental feature. state = torch.load(state, map_location="cuda") checkpoint version 3.0 2026-07-23:04:15:16,476 WARNING [rerun_state_machine.py:811] RerunStateMachine disabled via CLI, ignoring machine state saved in checkpoint successfully loaded checkpoint from /workspace/FLAME-MoE/results/phase0/runs/flame38m_g3_temporal_s2/ckpt [ t 1/1, p 1/1 ] at iteration 2121 2026-07-23:04:15:16,604 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc "HTTP/1.1 200 OK" 2026-07-23:04:15:16,835 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc "HTTP/1.1 200 OK" 2026-07-23:04:15:16,943 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc/revision/210d026faf9955653af8916fad021475a3f00453 "HTTP/1.1 200 OK" 2026-07-23:04:15:17,015 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc/tree/210d026faf9955653af8916fad021475a3f00453?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:15:17,255 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc/tree/210d026faf9955653af8916fad021475a3f00453/ARC-Challenge?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:15:17,787 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc "HTTP/1.1 200 OK" 2026-07-23:04:15:18,031 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc "HTTP/1.1 200 OK" 2026-07-23:04:15:18,292 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc/tree/210d026faf9955653af8916fad021475a3f00453/ARC-Easy?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:15:18,435 WARNING [task.py:799] [Task: boolq] metric acc is defined, but aggregation is not. using default aggregation=mean 2026-07-23:04:15:18,435 WARNING [task.py:811] [Task: boolq] metric acc is defined, but higher_is_better is not. using default higher_is_better=True 2026-07-23:04:15:18,498 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue "HTTP/1.1 200 OK" 2026-07-23:04:15:18,722 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue "HTTP/1.1 200 OK" 2026-07-23:04:15:18,781 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/revision/3de24cf8022e94f4ee4b9d55a6f539891524d646 "HTTP/1.1 200 OK" 2026-07-23:04:15:18,844 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/tree/3de24cf8022e94f4ee4b9d55a6f539891524d646?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:15:19,163 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/tree/3de24cf8022e94f4ee4b9d55a6f539891524d646/axb?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:15:19,223 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/revision/3de24cf8022e94f4ee4b9d55a6f539891524d646 "HTTP/1.1 200 OK" 2026-07-23:04:15:19,295 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/tree/3de24cf8022e94f4ee4b9d55a6f539891524d646/boolq?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:15:19,544 WARNING [task.py:799] [Task: copa] metric acc is defined, but aggregation is not. using default aggregation=mean 2026-07-23:04:15:19,544 WARNING [task.py:811] [Task: copa] metric acc is defined, but higher_is_better is not. using default higher_is_better=True 2026-07-23:04:15:19,606 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue "HTTP/1.1 200 OK" 2026-07-23:04:15:19,838 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue "HTTP/1.1 200 OK" 2026-07-23:04:15:20,119 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/tree/3de24cf8022e94f4ee4b9d55a6f539891524d646/copa?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:15:20,346 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag "HTTP/1.1 200 OK" 2026-07-23:04:15:20,633 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag "HTTP/1.1 200 OK" 2026-07-23:04:15:20,699 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag/revision/218ec52e09a7e7462a5400043bb9a69a41d06b76 "HTTP/1.1 200 OK" 2026-07-23:04:15:20,764 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag/tree/218ec52e09a7e7462a5400043bb9a69a41d06b76?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:15:20,969 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag/tree/218ec52e09a7e7462a5400043bb9a69a41d06b76/data?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:15:21,034 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag/revision/218ec52e09a7e7462a5400043bb9a69a41d06b76 "HTTP/1.1 200 OK" 2026-07-23:04:15:23,469 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai "HTTP/1.1 200 OK" 2026-07-23:04:15:23,691 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai "HTTP/1.1 200 OK" 2026-07-23:04:15:23,755 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/revision/900124bf3b8235c6daf21033af9948b3f07346c4 "HTTP/1.1 200 OK" 2026-07-23:04:15:23,810 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/tree/900124bf3b8235c6daf21033af9948b3f07346c4?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:15:24,106 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/tree/900124bf3b8235c6daf21033af9948b3f07346c4/default%2Ftest?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:15:24,168 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/tree/900124bf3b8235c6daf21033af9948b3f07346c4/default?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:15:24,231 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/revision/900124bf3b8235c6daf21033af9948b3f07346c4 "HTTP/1.1 200 OK" 2026-07-23:04:15:24,251 WARNING [task.py:325] [Task: lambada_openai] has_training_docs and has_validation_docs are False, using test_docs as fewshot_docs but this is not recommended. 2026-07-23:04:15:24,251 WARNING [task.py:325] [Task: lambada_openai] has_training_docs and has_validation_docs are False, using test_docs as fewshot_docs but this is not recommended. 2026-07-23:04:15:24,402 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa "HTTP/1.1 200 OK" 2026-07-23:04:15:24,632 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa "HTTP/1.1 200 OK" 2026-07-23:04:15:24,694 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/revision/388097ea7776314e93a529163e0fea805b8a6454 "HTTP/1.1 200 OK" 2026-07-23:04:15:24,752 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/tree/388097ea7776314e93a529163e0fea805b8a6454?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:15:25,013 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/tree/388097ea7776314e93a529163e0fea805b8a6454/additional?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:15:25,077 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/revision/388097ea7776314e93a529163e0fea805b8a6454 "HTTP/1.1 200 OK" 2026-07-23:04:15:25,166 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/tree/388097ea7776314e93a529163e0fea805b8a6454/main?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:15:25,506 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa "HTTP/1.1 200 OK" 2026-07-23:04:15:25,565 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa "HTTP/1.1 200 OK" 2026-07-23:04:15:25,626 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa/tree/main?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:15:25,690 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/datasets/ybisk/piqa/resolve/main/piqa.py "HTTP/1.1 307 Temporary Redirect" 2026-07-23:04:15:25,702 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/resolve-cache/datasets/ybisk/piqa/2e8ac2dffd59bac8c3c6714948f4c551a0848bb0/piqa.py?%2Fdatasets%2Fybisk%2Fpiqa%2Fresolve%2Fmain%2Fpiqa.py=&etag=%220f0d96172f9bf7aa10069374b3236ef083cfeb8e%22 "HTTP/1.1 200 OK" 2026-07-23:04:15:26,264 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa/revision/main "HTTP/1.1 200 OK" 2026-07-23:04:15:26,324 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa/tree/main?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:15:26,898 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq "HTTP/1.1 200 OK" 2026-07-23:04:15:27,134 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq "HTTP/1.1 200 OK" 2026-07-23:04:15:27,189 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq/revision/2c94ad3e1aafab77146f384e23536f97a4849815 "HTTP/1.1 200 OK" 2026-07-23:04:15:27,251 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq/tree/2c94ad3e1aafab77146f384e23536f97a4849815?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:15:27,483 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq/tree/2c94ad3e1aafab77146f384e23536f97a4849815/data?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:15:27,591 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq/revision/2c94ad3e1aafab77146f384e23536f97a4849815 "HTTP/1.1 200 OK" 2026-07-23:04:15:28,036 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande "HTTP/1.1 200 OK" 2026-07-23:04:15:28,263 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande "HTTP/1.1 200 OK" 2026-07-23:04:15:28,326 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/revision/01e74176c63542e6b0bcb004dcdea22d94fb67b5 "HTTP/1.1 200 OK" 2026-07-23:04:15:28,389 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/tree/01e74176c63542e6b0bcb004dcdea22d94fb67b5?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:15:28,621 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/tree/01e74176c63542e6b0bcb004dcdea22d94fb67b5/winogrande_debiased?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:15:28,677 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/revision/01e74176c63542e6b0bcb004dcdea22d94fb67b5 "HTTP/1.1 200 OK" 2026-07-23:04:15:28,773 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/tree/01e74176c63542e6b0bcb004dcdea22d94fb67b5/winogrande_xl?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:15:29,415 WARNING [evaluator.py:270] Overwriting default num_fewshot of winogrande from None to 0 2026-07-23:04:15:29,415 WARNING [evaluator.py:270] Overwriting default num_fewshot of sciq from None to 0 2026-07-23:04:15:29,415 WARNING [evaluator.py:270] Overwriting default num_fewshot of piqa from None to 0 2026-07-23:04:15:29,416 WARNING [evaluator.py:270] Overwriting default num_fewshot of openbookqa from None to 0 2026-07-23:04:15:29,416 WARNING [evaluator.py:270] Overwriting default num_fewshot of lambada_openai from None to 0 2026-07-23:04:15:29,416 WARNING [evaluator.py:270] Overwriting default num_fewshot of hellaswag from None to 0 2026-07-23:04:15:29,416 WARNING [evaluator.py:270] Overwriting default num_fewshot of copa from None to 0 2026-07-23:04:15:29,416 WARNING [evaluator.py:270] Overwriting default num_fewshot of boolq from None to 0 2026-07-23:04:15:29,416 WARNING [evaluator.py:270] Overwriting default num_fewshot of arc_easy from None to 0 2026-07-23:04:15:29,416 WARNING [evaluator.py:270] Overwriting default num_fewshot of arc_challenge from None to 0 2026-07-23:04:15:29,463 INFO [task.py:415] Building contexts for winogrande on rank 0... 0%| | 0/1267 [00:00 /workspace/FLAME-MoE/results/ablations/flame1e18_downstream.csv ==================================================================== [run] flame38m_g3_temporal_s3 mode=temporal TEMPORAL=1 H=256 heads=16 ffn=1368 experts=192 topk=18 moe_ffn=58 shared=352 [run] tasks=arc_challenge,arc_easy,hellaswag,openbookqa,piqa,winogrande,boolq,copa,lambada_openai,sciq out=/workspace/FLAME-MoE/results/phase0/runs/flame38m_g3_temporal_s3/lmeval_1e18_0shot ==================================================================== /usr/local/lib/python3.11/dist-packages/torch/cuda/__init__.py:58: FutureWarning: The pynvml package is deprecated. Please install nvidia-ml-py instead. If you did not install pynvml directly, please report this to the maintainers of the package that installed pynvml for you. import pynvml # type: ignore[import] [run_lmeval] stubbed VLM backends (hf_vlms, vllm_vlms) /usr/local/lib/python3.11/dist-packages/torch/cuda/__init__.py:58: FutureWarning: The pynvml package is deprecated. Please install nvidia-ml-py instead. If you did not install pynvml directly, please report this to the maintainers of the package that installed pynvml for you. import pynvml # type: ignore[import] /usr/local/lib/python3.11/dist-packages/torch/cuda/__init__.py:58: FutureWarning: The pynvml package is deprecated. Please install nvidia-ml-py instead. If you did not install pynvml directly, please report this to the maintainers of the package that installed pynvml for you. import pynvml # type: ignore[import] /workspace/FLAME-MoE/.venv/lib/python3.11/site-packages/transformer_engine/pytorch/attention.py:108: UserWarning: To use flash-attn v3, please use the following commands to install: (1) pip install "git+https://github.com/Dao-AILab/flash-attention.git#egg=flashattn-hopper&subdirectory=hopper" (2) python_path=`python -c "import site; print(site.getsitepackages()[0])"` (3) mkdir -p $python_path/flashattn_hopper (4) wget -P $python_path/flashattn_hopper https://raw.githubusercontent.com/Dao-AILab/flash-attention/main/hopper/flash_attn_interface.py warnings.warn( [temporal] rolling-residency router installed (evict=min_logit) [run_lmeval] installed temporal residency router (native regime) /usr/local/lib/python3.11/dist-packages/requests/__init__.py:113: RequestsDependencyWarning: urllib3 (2.2.3) or chardet (6.0.0.post1)/charset_normalizer (3.3.2) doesn't match a supported version! warnings.warn( 2026-07-23:04:24:35,695 INFO [__main__.py:279] Verbosity set to INFO 2026-07-23:04:25:04,578 INFO [__init__.py:459] The tag 'arc_ca' is already registered as a group, this tag will not be registered. This may affect tasks you want to call. 2026-07-23:04:25:04,589 INFO [__init__.py:459] The tag 'arc_ca' is already registered as a group, this tag will not be registered. This may affect tasks you want to call. 2026-07-23:04:25:21,410 INFO [__main__.py:376] Selected Tasks: ['arc_challenge', 'arc_easy', 'boolq', 'copa', 'hellaswag', 'lambada_openai', 'openbookqa', 'piqa', 'sciq', 'winogrande'] 2026-07-23:04:25:21,438 INFO [evaluator.py:164] Setting random seed to 42 | Setting numpy seed to 42 | Setting torch manual seed to 42 | Setting fewshot manual seed to 42 2026-07-23:04:25:21,438 INFO [evaluator.py:201] Initializing megatron_lm model, with arguments: {} using world size: 1, data-parallel size: 1, context-parallel size: 1, hierarchical context-parallel sizes: Nonetensor-model-parallel size: 1, encoder-tensor-model-parallel size: 0, pipeline-model-parallel size: 1, encoder-pipeline-model-parallel size: 0 WARNING: overriding default arguments for exit_on_missing_checkpoint:True with exit_on_missing_checkpoint:False setting global batch size to 32 Number of virtual stages per pipeline stage: None accumulate and all-reduce gradients in fp32 for bfloat16 data type. using torch.bfloat16 for parameters ... ------------------------ arguments ------------------------ account_for_embedding_in_pipeline_split ......... False account_for_loss_in_pipeline_split .............. False accumulate_allreduce_grads_in_fp32 .............. True adam_beta1 ...................................... 0.9 adam_beta2 ...................................... 0.999 adam_eps ........................................ 1e-08 add_bias_linear ................................. False add_position_embedding .......................... True add_qkv_bias .................................... False adlr_autoresume ................................. False adlr_autoresume_interval ........................ 1000 align_grad_reduce ............................... True align_param_gather .............................. False app_tag_run_name ................................ None app_tag_run_version ............................. 0.0.0 apply_layernorm_1p .............................. False apply_query_key_layer_scaling ................... False apply_residual_connection_post_layernorm ........ False apply_rope_fusion ............................... True async_save ...................................... None async_tensor_model_parallel_allreduce ........... True attention_backend ............................... AttnBackend.auto attention_dropout ............................... 0.0 attention_softmax_in_fp32 ....................... False auto_detect_ckpt_format ......................... False barrier_with_L1_time ............................ True bert_binary_head ................................ True bert_embedder_type .............................. megatron bert_load ....................................... None bf16 ............................................ True bias_dropout_fusion ............................. True bias_gelu_fusion ................................ False bias_swiglu_fusion .............................. True biencoder_projection_dim ........................ 0 biencoder_shared_query_context_model ............ False block_data_path ................................. None calc_ft_timeouts ................................ False calculate_per_token_loss ........................ False check_for_large_grads ........................... False check_for_nan_in_loss_and_grad .................. True check_for_spiky_loss ............................ False check_weight_hash_across_dp_replicas_interval ... None ckpt_assume_constant_structure .................. False ckpt_convert_format ............................. None ckpt_convert_save ............................... None ckpt_convert_update_legacy_dist_opt_format ...... False ckpt_format ..................................... torch_dist ckpt_fully_parallel_load ........................ False ckpt_fully_parallel_save ........................ True ckpt_fully_parallel_save_deprecated ............. False ckpt_step ....................................... None classes_fraction ................................ 1.0 clip_grad ....................................... 1.0 clone_scatter_output_in_embedding ............... True config_logger_dir ............................... consumed_train_samples .......................... 0 consumed_valid_samples .......................... 0 context_parallel_size ........................... 1 cp_comm_type .................................... ['p2p'] create_attention_mask_in_dataloader ............. True cross_entropy_loss_fusion ....................... False cuda_graph_warmup_steps ......................... 3 data_args_path .................................. None data_cache_path ................................. None data_parallel_random_init ....................... False data_parallel_sharding_strategy ................. no_shard data_parallel_size .............................. 1 data_path ....................................... None data_per_class_fraction ......................... 1.0 data_sharding ................................... True dataloader_type ................................. single ddp_average_in_collective ....................... False ddp_bucket_size ................................. None ddp_num_buckets ................................. None ddp_pad_buckets_for_high_nccl_busbw ............. False decoder_first_pipeline_num_layers ............... None decoder_last_pipeline_num_layers ................ None decoder_num_layers .............................. None decoder_seq_length .............................. None decoupled_lr .................................... None decoupled_min_lr ................................ None decrease_batch_size_if_needed ................... False defer_embedding_wgrad_compute ................... False deprecated_use_mcore_models ..................... False deterministic_mode .............................. False dino_bottleneck_size ............................ 256 dino_freeze_last_layer .......................... 1 dino_head_hidden_size ........................... 2048 dino_local_crops_number ......................... 10 dino_local_img_size ............................. 96 dino_norm_last_layer ............................ False dino_teacher_temp ............................... 0.07 dino_warmup_teacher_temp ........................ 0.04 dino_warmup_teacher_temp_epochs ................. 30 disable_straggler_on_startup .................... False dist_ckpt_format_deprecated ..................... None dist_ckpt_strictness ............................ assume_ok_unexpected distribute_saved_activations .................... False distributed_backend ............................. nccl distributed_timeout_minutes ..................... 10 embedding_path .................................. None empty_unused_memory_level ....................... 0 enable_cuda_graph ............................... False enable_ft_package ............................... False enable_gloo_process_groups ...................... True enable_one_logger ............................... True encoder_num_layers .............................. 9 encoder_pipeline_model_parallel_size ............ 0 encoder_seq_length .............................. 2048 encoder_tensor_model_parallel_size .............. 0 end_weight_decay ................................ 0.01 eod_mask_loss ................................... False error_injection_rate ............................ 0 error_injection_type ............................ transient_error eval_interval ................................... 1000 eval_iters ...................................... 100 evidence_data_path .............................. None exit_duration_in_mins ........................... None exit_interval ................................... None exit_on_missing_checkpoint ...................... False exit_signal_handler ............................. False exp_avg_dtype ................................... torch.float32 exp_avg_sq_dtype ................................ torch.float32 expert_model_parallel_size ...................... 1 expert_tensor_parallel_size ..................... 1 ffn_hidden_size ................................. 1368 finetune ........................................ False flash_decode .................................... False fp16 ............................................ False fp16_lm_cross_entropy ........................... False fp32_residual_connection ........................ False fp8 ............................................. None fp8_amax_compute_algo ........................... most_recent fp8_amax_history_len ............................ 1 fp8_interval .................................... 1 fp8_margin ...................................... 0 fp8_param_gather ................................ False fp8_wgrad ....................................... True global_batch_size ............................... 32 grad_reduce_in_bf16 ............................. False gradient_accumulation_fusion .................... False gradient_reduce_div_fusion ...................... True group_query_attention ........................... False head_lr_mult .................................... 1.0 hidden_dropout .................................. 0.0 hidden_size ..................................... 256 hierarchical_context_parallel_sizes ............. None hybrid_attention_ratio .......................... 0.0 hybrid_mlp_ratio ................................ 0.0 hybrid_override_pattern ......................... None hysteresis ...................................... 2 ict_head_size ................................... None ict_load ........................................ None img_h ........................................... 224 img_w ........................................... 224 indexer_batch_size .............................. 128 indexer_log_interval ............................ 1000 inference_batch_times_seqlen_threshold .......... -1 inference_max_batch_size ........................ 8 inference_max_seq_length ........................ 2560 inference_rng_tracker ........................... False init_method_std ................................. 0.02 init_method_xavier_uniform ...................... False init_model_with_meta_device ..................... False initial_loss_scale .............................. 4294967296 iter_per_epoch .................................. 1250 iterations_to_skip .............................. [] keep_fp8_transpose_cache_when_using_custom_fsdp . False kv_channels ..................................... 16 kv_lora_rank .................................... 32 lazy_mpu_init ................................... None load ............................................ /workspace/FLAME-MoE/results/phase0/runs/flame38m_g3_temporal_s3/ckpt local_rank ...................................... 0 log_interval .................................... 100 log_loss_scale_to_tensorboard ................... True log_memory_to_tensorboard ....................... False log_num_zeros_in_grad ........................... False log_params_norm ................................. False log_progress .................................... False log_straggler ................................... False log_throughput .................................. False log_timers_to_tensorboard ....................... False log_validation_ppl_to_tensorboard ............... False log_world_size_to_tensorboard ................... False logging_level ................................... None loss_scale ...................................... None loss_scale_window ............................... 1000 lr .............................................. None lr_decay_iters .................................. None lr_decay_samples ................................ None lr_decay_style .................................. linear lr_warmup_fraction .............................. None lr_warmup_init .................................. 0.0 lr_warmup_iters ................................. 0 lr_warmup_samples ............................... 0 lr_wsd_decay_iters .............................. None lr_wsd_decay_samples ............................ None lr_wsd_decay_style .............................. exponential main_grads_dtype ................................ torch.float32 main_params_dtype ............................... torch.float32 make_vocab_size_divisible_by .................... 128 manual_gc ....................................... False manual_gc_eval .................................. True manual_gc_interval .............................. 0 mask_factor ..................................... 1.0 mask_prob ....................................... 0.15 mask_type ....................................... random masked_softmax_fusion ........................... True max_position_embeddings ......................... 2048 max_tokens_to_oom ............................... 10000000 memory_snapshot_path ............................ snapshot.pickle merge_file ...................................... None micro_batch_size ................................ 32 microbatch_group_size_per_vp_stage .............. None min_loss_scale .................................. 1.0 min_lr .......................................... 0.0 mmap_bin_files .................................. True mock_data ....................................... False moe_aux_loss_coeff .............................. 0.01 moe_enable_deepep ............................... False moe_expert_capacity_factor ...................... None moe_extended_tp ................................. False moe_ffn_hidden_size ............................. 58 moe_grouped_gemm ................................ True moe_input_jitter_eps ............................ None moe_layer_freq .................................. [0, 1, 1, 1, 1, 1, 1, 1, 1] moe_layer_recompute ............................. False moe_pad_expert_input_to_capacity ................ False moe_per_layer_logging ........................... False moe_permute_fusion .............................. False moe_router_bias_update_rate ..................... 0.001 moe_router_dtype ................................ fp32 moe_router_enable_expert_bias ................... False moe_router_group_topk ........................... None moe_router_load_balancing_type .................. aux_loss moe_router_num_groups ........................... None moe_router_pre_softmax .......................... True moe_router_score_function ....................... softmax moe_router_topk ................................. 18 moe_router_topk_scaling_factor .................. None moe_shared_expert_intermediate_size ............. 352 moe_shared_expert_overlap ....................... False moe_token_dispatcher_type ....................... alltoall moe_token_drop_policy ........................... probs moe_use_legacy_grouped_gemm ..................... False moe_use_upcycling ............................... False moe_z_loss_coeff ................................ 0.001 mscale .......................................... 1.0 mscale_all_dim .................................. 1.0 multi_latent_attention .......................... False nccl_communicator_config_path ................... None no_load_optim ................................... True no_load_rng ..................................... True no_persist_layer_norm ........................... False no_save_optim ................................... None no_save_rng ..................................... None non_persistent_ckpt_type ........................ None non_persistent_global_ckpt_dir .................. None non_persistent_local_ckpt_algo .................. fully_parallel non_persistent_local_ckpt_dir ................... None non_persistent_save_interval .................... None norm_epsilon .................................... 1e-06 normalization ................................... RMSNorm num_attention_heads ............................. 16 num_channels .................................... 3 num_classes ..................................... 1000 num_dataset_builder_threads ..................... 1 num_distributed_optimizer_instances ............. 1 num_experts ..................................... 192 num_layers ...................................... 9 num_layers_per_virtual_pipeline_stage ........... None num_query_groups ................................ 1 num_virtual_stages_per_pipeline_rank ............ None num_workers ..................................... 2 one_logger_async ................................ False one_logger_project .............................. megatron-lm one_logger_run_name ............................. None onnx_safe ....................................... None openai_gelu ..................................... False optimizer ....................................... adam optimizer_cpu_offload ........................... False optimizer_offload_fraction ...................... 1.0 output_bert_embeddings .......................... False overlap_cpu_optimizer_d2h_h2d ................... False overlap_grad_reduce ............................. False overlap_p2p_comm ................................ False overlap_p2p_comm_warmup_flush ................... False overlap_param_gather ............................ False overlap_param_gather_with_optimizer_step ........ False override_opt_param_scheduler .................... False params_dtype .................................... torch.bfloat16 patch_dim ....................................... 16 per_split_data_args_path ........................ None perform_initialization .......................... True pin_cpu_grads ................................... True pin_cpu_params .................................. True pipeline_model_parallel_comm_backend ............ None pipeline_model_parallel_size .................... 1 pipeline_model_parallel_split_rank .............. None position_embedding_type ......................... rope pretrained_checkpoint ........................... None profile ......................................... False profile_ranks ................................... [0] profile_step_end ................................ 12 profile_step_start .............................. 10 q_lora_rank ..................................... None qk_head_dim ..................................... 128 qk_layernorm .................................... False qk_pos_emb_head_dim ............................. 64 query_in_block_prob ............................. 0.1 rampup_batch_size ............................... None rank ............................................ 0 recompute_granularity ........................... None recompute_method ................................ None recompute_num_layers ............................ None record_memory_history ........................... False relative_attention_max_distance ................. 128 relative_attention_num_buckets .................. 32 replication ..................................... False replication_factor .............................. 2 replication_jump ................................ None rerun_mode ...................................... disabled reset_attention_mask ............................ False reset_position_ids .............................. False result_rejected_tracker_filename ................ None retriever_report_topk_accuracies ................ [] retriever_score_scaling ......................... False retriever_seq_length ............................ 256 retro_add_retriever ............................. False retro_attention_gate ............................ 1 retro_cyclic_train_iters ........................ None retro_encoder_attention_dropout ................. 0.1 retro_encoder_hidden_dropout .................... 0.1 retro_encoder_layers ............................ 2 retro_num_neighbors ............................. 2 retro_num_retrieved_chunks ...................... 2 retro_project_dir ............................... None retro_verify_neighbor_count ..................... True rope_scaling_factor ............................. 8.0 rotary_base ..................................... 10000 rotary_interleaved .............................. False rotary_percent .................................. 1.0 rotary_scaling_factor ........................... 1.0 rotary_seq_len_interpolation_factor ............. None s3_cache_path ................................... None sample_rate ..................................... 1.0 save ............................................ None save_interval ................................... None scatter_gather_tensors_in_pipeline .............. True seed ............................................ 42 seq_length ...................................... 2048 sequence_parallel ............................... False sgd_momentum .................................... 0.9 short_seq_prob .................................. 0.1 skip_train ...................................... False skipped_train_samples ........................... 0 spec ............................................ None split ........................................... None squared_relu .................................... False start_weight_decay .............................. 0.01 straggler_ctrlr_port ............................ 65535 straggler_minmax_count .......................... 1 suggested_communication_unit_size ............... 400000000 swiglu .......................................... True swin_backbone_type .............................. tiny te_rng_tracker .................................. False tensor_model_parallel_size ...................... 1 tensorboard_dir ................................. None tensorboard_log_interval ........................ 1 tensorboard_queue_size .......................... 1000 test_data_path .................................. None test_mode ....................................... False tiktoken_num_special_tokens ..................... 1000 tiktoken_pattern ................................ None tiktoken_special_tokens ......................... None timing_log_level ................................ 0 timing_log_option ............................... minmax titles_data_path ................................ None tokenizer_model ................................. EleutherAI/pythia-12b tokenizer_type .................................. HuggingFaceTokenizer tp_comm_bootstrap_backend ....................... nccl tp_comm_bulk_dgrad .............................. True tp_comm_bulk_wgrad .............................. True tp_comm_overlap ................................. False tp_comm_overlap_ag .............................. True tp_comm_overlap_cfg ............................. None tp_comm_overlap_rs .............................. True tp_comm_overlap_rs_dgrad ........................ False tp_comm_split_ag ................................ True tp_comm_split_rs ................................ True train_data_path ................................. None train_iters ..................................... None train_samples ................................... None train_sync_interval ............................. None transformer_impl ................................ transformer_engine transformer_pipeline_model_parallel_size ........ 1 untie_embeddings_and_output_weights ............. True use_checkpoint_args ............................. False use_checkpoint_opt_param_scheduler .............. False use_cpu_initialization .......................... None use_custom_fsdp ................................. False use_dist_ckpt ................................... True use_dist_ckpt_deprecated ........................ False use_distributed_optimizer ....................... True use_flash_attn .................................. False use_legacy_models ............................... False use_mp_args_from_checkpoint_args ................ False use_one_sent_docs ............................... False use_persistent_ckpt_worker ...................... False use_precision_aware_optimizer ................... False use_pytorch_profiler ............................ False use_ring_exchange_p2p ........................... False use_rope_scaling ................................ False use_rotary_position_embeddings .................. False use_tokenizer_model_from_checkpoint_args ........ True use_torch_fsdp2 ................................. False use_torch_optimizer_for_cpu_offload ............. False use_tp_pp_dp_mapping ............................ False v_head_dim ...................................... 128 valid_data_path ................................. None variable_seq_lengths ............................ False virtual_pipeline_model_parallel_size ............ None vision_backbone_type ............................ vit vision_pretraining .............................. False vision_pretraining_type ......................... classify vocab_extra_ids ................................. 0 vocab_file ...................................... None vocab_size ...................................... None wandb_exp_name .................................. wandb_project ................................... wandb_save_dir .................................. weight_decay .................................... 0.01 weight_decay_incr_style ......................... constant wgrad_deferral_limit ............................ 0 world_size ...................................... 1 yaml_cfg ........................................ None -------------------- end of arguments --------------------- 2026-07-23:04:25:21,454 INFO [num_microbatches_calculator.py:228] setting number of microbatches to constant 1 > building HuggingFaceTokenizer tokenizer ... 2026-07-23:04:25:21,648 INFO [_client.py:1038] HTTP Request: HEAD https://huggingface.co/EleutherAI/pythia-12b/resolve/main/config.json "HTTP/1.1 307 Temporary Redirect" 2026-07-23:04:25:21,660 INFO [_client.py:1038] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/EleutherAI/pythia-12b/bb1e3e710cdf6b524461d543cfb5ba773f0a81b6/config.json "HTTP/1.1 200 OK" 2026-07-23:04:25:21,787 INFO [_client.py:1038] HTTP Request: HEAD https://huggingface.co/EleutherAI/pythia-12b/resolve/main/tokenizer_config.json "HTTP/1.1 307 Temporary Redirect" 2026-07-23:04:25:21,800 INFO [_client.py:1038] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/EleutherAI/pythia-12b/bb1e3e710cdf6b524461d543cfb5ba773f0a81b6/tokenizer_config.json "HTTP/1.1 200 OK" 2026-07-23:04:25:21,904 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/models/EleutherAI/pythia-12b/tree/main/additional_chat_templates?recursive=false&expand=false "HTTP/1.1 404 Not Found" 2026-07-23:04:25:21,974 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/models/EleutherAI/pythia-12b/tree/main?recursive=true&expand=false "HTTP/1.1 200 OK" > padded vocab (size: 50277) with 27 dummy tokens (new size: 50304) WARNING: one_logger package is required to enable e2e metrics tracking. please go to https://confluence.nvidia.com/display/MLWFO/Package+Repositories for details to install it 2026-07-23:04:25:22,185 WARNING [rerun_state_machine.py:239] RerunStateMachine initialized in mode RerunMode.DISABLED > initializing torch distributed ... [W723 04:25:22.703270232 CUDAAllocatorConfig.h:28] Warning: expandable_segments not supported on this platform (function operator()) > initialized tensor model parallel with size 1 > initialized pipeline model parallel with size 1 > setting random seeds to 42 ... > compiling dataset index builder ... make: Entering directory '/workspace/FLAME-MoE/Megatron-LM/megatron/core/datasets' make: Nothing to be done for 'default'. make: Leaving directory '/workspace/FLAME-MoE/Megatron-LM/megatron/core/datasets' >>> done with dataset index builder. Compilation time: 0.695 seconds > compiling and loading fused kernels ... >>> done with compiling and loading fused kernels. Compilation time: 0.354 seconds building GPT model ... > number of parameters on (tensor, pipeline) model parallel rank (0, 0): 100145920 /workspace/FLAME-MoE/Megatron-LM/megatron/core/transformer/transformer_layer.py:339: UserWarning: TransformerLayer._get_layer_offset is deprecated.Please use get_transformer_layer_offset instead. warnings.warn( loading distributed checkpoint from /workspace/FLAME-MoE/results/phase0/runs/flame38m_g3_temporal_s3/ckpt at iteration 2121 /workspace/FLAME-MoE/Megatron-LM/megatron/core/dist_checkpointing/strategies/torch.py:847: FutureWarning: `load_state_dict` is deprecated and will be removed in future versions. Please use `load` instead. checkpoint.load_state_dict( /usr/local/lib/python3.11/dist-packages/torch/distributed/checkpoint/planner_helpers.py:306: FutureWarning: Please use DTensor instead and we are deprecating ShardedTensor. device = getattr(value, "device", None) /workspace/FLAME-MoE/.venv/lib/python3.11/site-packages/transformer_engine/pytorch/module/base.py:566: FutureWarning: You are using `torch.load` with `weights_only=False` (the current default value), which uses the default pickle module implicitly. It is possible to construct malicious pickle data which will execute arbitrary code during unpickling (See https://github.com/pytorch/pytorch/blob/main/SECURITY.md#untrusted-models for more details). In a future release, the default value for `weights_only` will be flipped to `True`. This limits the functions that could be executed during unpickling. Arbitrary objects will no longer be allowed to be loaded via this mode unless they are explicitly allowlisted by the user via `torch.serialization.add_safe_globals`. We recommend you start setting `weights_only=True` for any use case where you don't have full control of the loaded file. Please open an issue on GitHub for any issues related to this experimental feature. state = torch.load(state, map_location="cuda") checkpoint version 3.0 2026-07-23:04:25:28,934 WARNING [rerun_state_machine.py:811] RerunStateMachine disabled via CLI, ignoring machine state saved in checkpoint successfully loaded checkpoint from /workspace/FLAME-MoE/results/phase0/runs/flame38m_g3_temporal_s3/ckpt [ t 1/1, p 1/1 ] at iteration 2121 2026-07-23:04:25:29,061 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc "HTTP/1.1 200 OK" 2026-07-23:04:25:29,299 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc "HTTP/1.1 200 OK" 2026-07-23:04:25:29,360 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc/revision/210d026faf9955653af8916fad021475a3f00453 "HTTP/1.1 200 OK" 2026-07-23:04:25:29,418 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc/tree/210d026faf9955653af8916fad021475a3f00453?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:25:29,706 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc/tree/210d026faf9955653af8916fad021475a3f00453/ARC-Challenge?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:25:30,236 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc "HTTP/1.1 200 OK" 2026-07-23:04:25:30,487 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc "HTTP/1.1 200 OK" 2026-07-23:04:25:30,754 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc/tree/210d026faf9955653af8916fad021475a3f00453/ARC-Easy?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:25:30,897 WARNING [task.py:799] [Task: boolq] metric acc is defined, but aggregation is not. using default aggregation=mean 2026-07-23:04:25:30,897 WARNING [task.py:811] [Task: boolq] metric acc is defined, but higher_is_better is not. using default higher_is_better=True 2026-07-23:04:25:30,963 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue "HTTP/1.1 200 OK" 2026-07-23:04:25:31,200 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue "HTTP/1.1 200 OK" 2026-07-23:04:25:31,264 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/revision/3de24cf8022e94f4ee4b9d55a6f539891524d646 "HTTP/1.1 200 OK" 2026-07-23:04:25:31,330 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/tree/3de24cf8022e94f4ee4b9d55a6f539891524d646?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:25:31,655 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/tree/3de24cf8022e94f4ee4b9d55a6f539891524d646/axb?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:25:31,737 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/revision/3de24cf8022e94f4ee4b9d55a6f539891524d646 "HTTP/1.1 200 OK" 2026-07-23:04:25:31,818 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/tree/3de24cf8022e94f4ee4b9d55a6f539891524d646/boolq?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:25:32,059 WARNING [task.py:799] [Task: copa] metric acc is defined, but aggregation is not. using default aggregation=mean 2026-07-23:04:25:32,059 WARNING [task.py:811] [Task: copa] metric acc is defined, but higher_is_better is not. using default higher_is_better=True 2026-07-23:04:25:32,120 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue "HTTP/1.1 200 OK" 2026-07-23:04:25:32,360 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue "HTTP/1.1 200 OK" 2026-07-23:04:25:32,635 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/tree/3de24cf8022e94f4ee4b9d55a6f539891524d646/copa?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:25:32,784 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag "HTTP/1.1 200 OK" 2026-07-23:04:25:33,025 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag "HTTP/1.1 200 OK" 2026-07-23:04:25:33,089 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag/revision/218ec52e09a7e7462a5400043bb9a69a41d06b76 "HTTP/1.1 200 OK" 2026-07-23:04:25:33,150 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag/tree/218ec52e09a7e7462a5400043bb9a69a41d06b76?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:25:33,359 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag/tree/218ec52e09a7e7462a5400043bb9a69a41d06b76/data?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:25:33,419 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag/revision/218ec52e09a7e7462a5400043bb9a69a41d06b76 "HTTP/1.1 200 OK" 2026-07-23:04:25:35,785 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai "HTTP/1.1 200 OK" 2026-07-23:04:25:36,062 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai "HTTP/1.1 200 OK" 2026-07-23:04:25:36,120 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/revision/900124bf3b8235c6daf21033af9948b3f07346c4 "HTTP/1.1 200 OK" 2026-07-23:04:25:36,179 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/tree/900124bf3b8235c6daf21033af9948b3f07346c4?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:25:36,432 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/tree/900124bf3b8235c6daf21033af9948b3f07346c4/default%2Ftest?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:25:36,492 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/tree/900124bf3b8235c6daf21033af9948b3f07346c4/default?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:25:36,555 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/revision/900124bf3b8235c6daf21033af9948b3f07346c4 "HTTP/1.1 200 OK" 2026-07-23:04:25:36,574 WARNING [task.py:325] [Task: lambada_openai] has_training_docs and has_validation_docs are False, using test_docs as fewshot_docs but this is not recommended. 2026-07-23:04:25:36,574 WARNING [task.py:325] [Task: lambada_openai] has_training_docs and has_validation_docs are False, using test_docs as fewshot_docs but this is not recommended. 2026-07-23:04:25:36,748 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa "HTTP/1.1 200 OK" 2026-07-23:04:25:36,990 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa "HTTP/1.1 200 OK" 2026-07-23:04:25:37,053 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/revision/388097ea7776314e93a529163e0fea805b8a6454 "HTTP/1.1 200 OK" 2026-07-23:04:25:37,114 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/tree/388097ea7776314e93a529163e0fea805b8a6454?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:25:37,324 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/tree/388097ea7776314e93a529163e0fea805b8a6454/additional?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:25:37,389 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/revision/388097ea7776314e93a529163e0fea805b8a6454 "HTTP/1.1 200 OK" 2026-07-23:04:25:37,480 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/tree/388097ea7776314e93a529163e0fea805b8a6454/main?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:25:37,747 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa "HTTP/1.1 200 OK" 2026-07-23:04:25:37,807 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa "HTTP/1.1 200 OK" 2026-07-23:04:25:37,870 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa/tree/main?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:25:37,926 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/datasets/ybisk/piqa/resolve/main/piqa.py "HTTP/1.1 307 Temporary Redirect" 2026-07-23:04:25:37,938 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/resolve-cache/datasets/ybisk/piqa/2e8ac2dffd59bac8c3c6714948f4c551a0848bb0/piqa.py?%2Fdatasets%2Fybisk%2Fpiqa%2Fresolve%2Fmain%2Fpiqa.py=&etag=%220f0d96172f9bf7aa10069374b3236ef083cfeb8e%22 "HTTP/1.1 200 OK" 2026-07-23:04:25:38,456 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa/revision/main "HTTP/1.1 200 OK" 2026-07-23:04:25:38,521 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa/tree/main?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:25:38,973 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq "HTTP/1.1 200 OK" 2026-07-23:04:25:39,206 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq "HTTP/1.1 200 OK" 2026-07-23:04:25:39,265 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq/revision/2c94ad3e1aafab77146f384e23536f97a4849815 "HTTP/1.1 200 OK" 2026-07-23:04:25:39,325 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq/tree/2c94ad3e1aafab77146f384e23536f97a4849815?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:25:39,541 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq/tree/2c94ad3e1aafab77146f384e23536f97a4849815/data?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:25:39,685 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq/revision/2c94ad3e1aafab77146f384e23536f97a4849815 "HTTP/1.1 200 OK" 2026-07-23:04:25:40,059 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande "HTTP/1.1 200 OK" 2026-07-23:04:25:40,330 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande "HTTP/1.1 200 OK" 2026-07-23:04:25:40,396 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/revision/01e74176c63542e6b0bcb004dcdea22d94fb67b5 "HTTP/1.1 200 OK" 2026-07-23:04:25:40,457 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/tree/01e74176c63542e6b0bcb004dcdea22d94fb67b5?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:25:40,712 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/tree/01e74176c63542e6b0bcb004dcdea22d94fb67b5/winogrande_debiased?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:25:40,779 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/revision/01e74176c63542e6b0bcb004dcdea22d94fb67b5 "HTTP/1.1 200 OK" 2026-07-23:04:25:40,864 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/tree/01e74176c63542e6b0bcb004dcdea22d94fb67b5/winogrande_xl?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:25:41,513 WARNING [evaluator.py:270] Overwriting default num_fewshot of winogrande from None to 0 2026-07-23:04:25:41,513 WARNING [evaluator.py:270] Overwriting default num_fewshot of sciq from None to 0 2026-07-23:04:25:41,513 WARNING [evaluator.py:270] Overwriting default num_fewshot of piqa from None to 0 2026-07-23:04:25:41,513 WARNING [evaluator.py:270] Overwriting default num_fewshot of openbookqa from None to 0 2026-07-23:04:25:41,513 WARNING [evaluator.py:270] Overwriting default num_fewshot of lambada_openai from None to 0 2026-07-23:04:25:41,513 WARNING [evaluator.py:270] Overwriting default num_fewshot of hellaswag from None to 0 2026-07-23:04:25:41,513 WARNING [evaluator.py:270] Overwriting default num_fewshot of copa from None to 0 2026-07-23:04:25:41,513 WARNING [evaluator.py:270] Overwriting default num_fewshot of boolq from None to 0 2026-07-23:04:25:41,513 WARNING [evaluator.py:270] Overwriting default num_fewshot of arc_easy from None to 0 2026-07-23:04:25:41,513 WARNING [evaluator.py:270] Overwriting default num_fewshot of arc_challenge from None to 0 2026-07-23:04:25:41,560 INFO [task.py:415] Building contexts for winogrande on rank 0... 0%| | 0/1267 [00:00 /workspace/FLAME-MoE/results/ablations/flame1e18_downstream.csv ==================================================================== [run] flame192_g3_moe mode=moe TEMPORAL=0 H=192 heads=12 ffn=1026 experts=192 topk=18 moe_ffn=44 shared=264 [run] tasks=arc_challenge,arc_easy,hellaswag,openbookqa,piqa,winogrande,boolq,copa,lambada_openai,sciq out=/workspace/FLAME-MoE/results/phase0/runs/flame192_g3_moe/lmeval_1e18_0shot ==================================================================== /usr/local/lib/python3.11/dist-packages/torch/cuda/__init__.py:58: FutureWarning: The pynvml package is deprecated. Please install nvidia-ml-py instead. If you did not install pynvml directly, please report this to the maintainers of the package that installed pynvml for you. import pynvml # type: ignore[import] [run_lmeval] stubbed VLM backends (hf_vlms, vllm_vlms) /usr/local/lib/python3.11/dist-packages/torch/cuda/__init__.py:58: FutureWarning: The pynvml package is deprecated. Please install nvidia-ml-py instead. If you did not install pynvml directly, please report this to the maintainers of the package that installed pynvml for you. import pynvml # type: ignore[import] /usr/local/lib/python3.11/dist-packages/requests/__init__.py:113: RequestsDependencyWarning: urllib3 (2.2.3) or chardet (6.0.0.post1)/charset_normalizer (3.3.2) doesn't match a supported version! warnings.warn( /usr/local/lib/python3.11/dist-packages/torch/cuda/__init__.py:58: FutureWarning: The pynvml package is deprecated. Please install nvidia-ml-py instead. If you did not install pynvml directly, please report this to the maintainers of the package that installed pynvml for you. import pynvml # type: ignore[import] /workspace/FLAME-MoE/.venv/lib/python3.11/site-packages/transformer_engine/pytorch/attention.py:108: UserWarning: To use flash-attn v3, please use the following commands to install: (1) pip install "git+https://github.com/Dao-AILab/flash-attention.git#egg=flashattn-hopper&subdirectory=hopper" (2) python_path=`python -c "import site; print(site.getsitepackages()[0])"` (3) mkdir -p $python_path/flashattn_hopper (4) wget -P $python_path/flashattn_hopper https://raw.githubusercontent.com/Dao-AILab/flash-attention/main/hopper/flash_attn_interface.py warnings.warn( /workspace/FLAME-MoE/Megatron-LM/megatron/core/models/gpt/gpt_layer_specs.py:54: UserWarning: Apex is not installed. Falling back to Torch Norm warnings.warn('Apex is not installed. Falling back to Torch Norm') /workspace/FLAME-MoE/Megatron-LM/megatron/core/models/retro/encoder_spec.py:47: UserWarning: Apex is not installed. Falling back to Torch Norm warnings.warn(f'Apex is not installed. Falling back to Torch Norm') /workspace/FLAME-MoE/Megatron-LM/megatron/core/models/retro/decoder_spec.py:39: UserWarning: Apex is not installed. Falling back to Torch Norm warnings.warn(f'Apex is not installed. Falling back to Torch Norm') 2026-07-23:04:34:52,825 INFO [__main__.py:279] Verbosity set to INFO 2026-07-23:04:35:25,223 INFO [__init__.py:459] The tag 'arc_ca' is already registered as a group, this tag will not be registered. This may affect tasks you want to call. 2026-07-23:04:35:25,240 INFO [__init__.py:459] The tag 'arc_ca' is already registered as a group, this tag will not be registered. This may affect tasks you want to call. 2026-07-23:04:35:40,724 INFO [__main__.py:376] Selected Tasks: ['arc_challenge', 'arc_easy', 'boolq', 'copa', 'hellaswag', 'lambada_openai', 'openbookqa', 'piqa', 'sciq', 'winogrande'] 2026-07-23:04:35:40,746 INFO [evaluator.py:164] Setting random seed to 42 | Setting numpy seed to 42 | Setting torch manual seed to 42 | Setting fewshot manual seed to 42 2026-07-23:04:35:40,746 INFO [evaluator.py:201] Initializing megatron_lm model, with arguments: {} using world size: 1, data-parallel size: 1, context-parallel size: 1, hierarchical context-parallel sizes: Nonetensor-model-parallel size: 1, encoder-tensor-model-parallel size: 0, pipeline-model-parallel size: 1, encoder-pipeline-model-parallel size: 0 WARNING: overriding default arguments for exit_on_missing_checkpoint:True with exit_on_missing_checkpoint:False setting global batch size to 32 Number of virtual stages per pipeline stage: None accumulate and all-reduce gradients in fp32 for bfloat16 data type. using torch.bfloat16 for parameters ... ------------------------ arguments ------------------------ account_for_embedding_in_pipeline_split ......... False account_for_loss_in_pipeline_split .............. False accumulate_allreduce_grads_in_fp32 .............. True adam_beta1 ...................................... 0.9 adam_beta2 ...................................... 0.999 adam_eps ........................................ 1e-08 add_bias_linear ................................. False add_position_embedding .......................... True add_qkv_bias .................................... False adlr_autoresume ................................. False adlr_autoresume_interval ........................ 1000 align_grad_reduce ............................... True align_param_gather .............................. False app_tag_run_name ................................ None app_tag_run_version ............................. 0.0.0 apply_layernorm_1p .............................. False apply_query_key_layer_scaling ................... False apply_residual_connection_post_layernorm ........ False apply_rope_fusion ............................... True async_save ...................................... None async_tensor_model_parallel_allreduce ........... True attention_backend ............................... AttnBackend.auto attention_dropout ............................... 0.0 attention_softmax_in_fp32 ....................... False auto_detect_ckpt_format ......................... False barrier_with_L1_time ............................ True bert_binary_head ................................ True bert_embedder_type .............................. megatron bert_load ....................................... None bf16 ............................................ True bias_dropout_fusion ............................. True bias_gelu_fusion ................................ False bias_swiglu_fusion .............................. True biencoder_projection_dim ........................ 0 biencoder_shared_query_context_model ............ False block_data_path ................................. None calc_ft_timeouts ................................ False calculate_per_token_loss ........................ False check_for_large_grads ........................... False check_for_nan_in_loss_and_grad .................. True check_for_spiky_loss ............................ False check_weight_hash_across_dp_replicas_interval ... None ckpt_assume_constant_structure .................. False ckpt_convert_format ............................. None ckpt_convert_save ............................... None ckpt_convert_update_legacy_dist_opt_format ...... False ckpt_format ..................................... torch_dist ckpt_fully_parallel_load ........................ False ckpt_fully_parallel_save ........................ True ckpt_fully_parallel_save_deprecated ............. False ckpt_step ....................................... None classes_fraction ................................ 1.0 clip_grad ....................................... 1.0 clone_scatter_output_in_embedding ............... True config_logger_dir ............................... consumed_train_samples .......................... 0 consumed_valid_samples .......................... 0 context_parallel_size ........................... 1 cp_comm_type .................................... ['p2p'] create_attention_mask_in_dataloader ............. True cross_entropy_loss_fusion ....................... False cuda_graph_warmup_steps ......................... 3 data_args_path .................................. None data_cache_path ................................. None data_parallel_random_init ....................... False data_parallel_sharding_strategy ................. no_shard data_parallel_size .............................. 1 data_path ....................................... None data_per_class_fraction ......................... 1.0 data_sharding ................................... True dataloader_type ................................. single ddp_average_in_collective ....................... False ddp_bucket_size ................................. None ddp_num_buckets ................................. None ddp_pad_buckets_for_high_nccl_busbw ............. False decoder_first_pipeline_num_layers ............... None decoder_last_pipeline_num_layers ................ None decoder_num_layers .............................. None decoder_seq_length .............................. None decoupled_lr .................................... None decoupled_min_lr ................................ None decrease_batch_size_if_needed ................... False defer_embedding_wgrad_compute ................... False deprecated_use_mcore_models ..................... False deterministic_mode .............................. False dino_bottleneck_size ............................ 256 dino_freeze_last_layer .......................... 1 dino_head_hidden_size ........................... 2048 dino_local_crops_number ......................... 10 dino_local_img_size ............................. 96 dino_norm_last_layer ............................ False dino_teacher_temp ............................... 0.07 dino_warmup_teacher_temp ........................ 0.04 dino_warmup_teacher_temp_epochs ................. 30 disable_straggler_on_startup .................... False dist_ckpt_format_deprecated ..................... None dist_ckpt_strictness ............................ assume_ok_unexpected distribute_saved_activations .................... False distributed_backend ............................. nccl distributed_timeout_minutes ..................... 10 embedding_path .................................. None empty_unused_memory_level ....................... 0 enable_cuda_graph ............................... False enable_ft_package ............................... False enable_gloo_process_groups ...................... True enable_one_logger ............................... True encoder_num_layers .............................. 9 encoder_pipeline_model_parallel_size ............ 0 encoder_seq_length .............................. 2048 encoder_tensor_model_parallel_size .............. 0 end_weight_decay ................................ 0.01 eod_mask_loss ................................... False error_injection_rate ............................ 0 error_injection_type ............................ transient_error eval_interval ................................... 1000 eval_iters ...................................... 100 evidence_data_path .............................. None exit_duration_in_mins ........................... None exit_interval ................................... None exit_on_missing_checkpoint ...................... False exit_signal_handler ............................. False exp_avg_dtype ................................... torch.float32 exp_avg_sq_dtype ................................ torch.float32 expert_model_parallel_size ...................... 1 expert_tensor_parallel_size ..................... 1 ffn_hidden_size ................................. 1026 finetune ........................................ False flash_decode .................................... False fp16 ............................................ False fp16_lm_cross_entropy ........................... False fp32_residual_connection ........................ False fp8 ............................................. None fp8_amax_compute_algo ........................... most_recent fp8_amax_history_len ............................ 1 fp8_interval .................................... 1 fp8_margin ...................................... 0 fp8_param_gather ................................ False fp8_wgrad ....................................... True global_batch_size ............................... 32 grad_reduce_in_bf16 ............................. False gradient_accumulation_fusion .................... False gradient_reduce_div_fusion ...................... True group_query_attention ........................... False head_lr_mult .................................... 1.0 hidden_dropout .................................. 0.0 hidden_size ..................................... 192 hierarchical_context_parallel_sizes ............. None hybrid_attention_ratio .......................... 0.0 hybrid_mlp_ratio ................................ 0.0 hybrid_override_pattern ......................... None hysteresis ...................................... 2 ict_head_size ................................... None ict_load ........................................ None img_h ........................................... 224 img_w ........................................... 224 indexer_batch_size .............................. 128 indexer_log_interval ............................ 1000 inference_batch_times_seqlen_threshold .......... -1 inference_max_batch_size ........................ 8 inference_max_seq_length ........................ 2560 inference_rng_tracker ........................... False init_method_std ................................. 0.02 init_method_xavier_uniform ...................... False init_model_with_meta_device ..................... False initial_loss_scale .............................. 4294967296 iter_per_epoch .................................. 1250 iterations_to_skip .............................. [] keep_fp8_transpose_cache_when_using_custom_fsdp . False kv_channels ..................................... 16 kv_lora_rank .................................... 32 lazy_mpu_init ................................... None load ............................................ /workspace/FLAME-MoE/results/phase0/runs/flame192_g3_moe/ckpt local_rank ...................................... 0 log_interval .................................... 100 log_loss_scale_to_tensorboard ................... True log_memory_to_tensorboard ....................... False log_num_zeros_in_grad ........................... False log_params_norm ................................. False log_progress .................................... False log_straggler ................................... False log_throughput .................................. False log_timers_to_tensorboard ....................... False log_validation_ppl_to_tensorboard ............... False log_world_size_to_tensorboard ................... False logging_level ................................... None loss_scale ...................................... None loss_scale_window ............................... 1000 lr .............................................. None lr_decay_iters .................................. None lr_decay_samples ................................ None lr_decay_style .................................. linear lr_warmup_fraction .............................. None lr_warmup_init .................................. 0.0 lr_warmup_iters ................................. 0 lr_warmup_samples ............................... 0 lr_wsd_decay_iters .............................. None lr_wsd_decay_samples ............................ None lr_wsd_decay_style .............................. exponential main_grads_dtype ................................ torch.float32 main_params_dtype ............................... torch.float32 make_vocab_size_divisible_by .................... 128 manual_gc ....................................... False manual_gc_eval .................................. True manual_gc_interval .............................. 0 mask_factor ..................................... 1.0 mask_prob ....................................... 0.15 mask_type ....................................... random masked_softmax_fusion ........................... True max_position_embeddings ......................... 2048 max_tokens_to_oom ............................... 10000000 memory_snapshot_path ............................ snapshot.pickle merge_file ...................................... None micro_batch_size ................................ 32 microbatch_group_size_per_vp_stage .............. None min_loss_scale .................................. 1.0 min_lr .......................................... 0.0 mmap_bin_files .................................. True mock_data ....................................... False moe_aux_loss_coeff .............................. 0.01 moe_enable_deepep ............................... False moe_expert_capacity_factor ...................... None moe_extended_tp ................................. False moe_ffn_hidden_size ............................. 44 moe_grouped_gemm ................................ True moe_input_jitter_eps ............................ None moe_layer_freq .................................. [0, 1, 1, 1, 1, 1, 1, 1, 1] moe_layer_recompute ............................. False moe_pad_expert_input_to_capacity ................ False moe_per_layer_logging ........................... False moe_permute_fusion .............................. False moe_router_bias_update_rate ..................... 0.001 moe_router_dtype ................................ fp32 moe_router_enable_expert_bias ................... False moe_router_group_topk ........................... None moe_router_load_balancing_type .................. aux_loss moe_router_num_groups ........................... None moe_router_pre_softmax .......................... True moe_router_score_function ....................... softmax moe_router_topk ................................. 18 moe_router_topk_scaling_factor .................. None moe_shared_expert_intermediate_size ............. 264 moe_shared_expert_overlap ....................... False moe_token_dispatcher_type ....................... alltoall moe_token_drop_policy ........................... probs moe_use_legacy_grouped_gemm ..................... False moe_use_upcycling ............................... False moe_z_loss_coeff ................................ 0.001 mscale .......................................... 1.0 mscale_all_dim .................................. 1.0 multi_latent_attention .......................... False nccl_communicator_config_path ................... None no_load_optim ................................... True no_load_rng ..................................... True no_persist_layer_norm ........................... False no_save_optim ................................... None no_save_rng ..................................... None non_persistent_ckpt_type ........................ None non_persistent_global_ckpt_dir .................. None non_persistent_local_ckpt_algo .................. fully_parallel non_persistent_local_ckpt_dir ................... None non_persistent_save_interval .................... None norm_epsilon .................................... 1e-06 normalization ................................... RMSNorm num_attention_heads ............................. 12 num_channels .................................... 3 num_classes ..................................... 1000 num_dataset_builder_threads ..................... 1 num_distributed_optimizer_instances ............. 1 num_experts ..................................... 192 num_layers ...................................... 9 num_layers_per_virtual_pipeline_stage ........... None num_query_groups ................................ 1 num_virtual_stages_per_pipeline_rank ............ None num_workers ..................................... 2 one_logger_async ................................ False one_logger_project .............................. megatron-lm one_logger_run_name ............................. None onnx_safe ....................................... None openai_gelu ..................................... False optimizer ....................................... adam optimizer_cpu_offload ........................... False optimizer_offload_fraction ...................... 1.0 output_bert_embeddings .......................... False overlap_cpu_optimizer_d2h_h2d ................... False overlap_grad_reduce ............................. False overlap_p2p_comm ................................ False overlap_p2p_comm_warmup_flush ................... False overlap_param_gather ............................ False overlap_param_gather_with_optimizer_step ........ False override_opt_param_scheduler .................... False params_dtype .................................... torch.bfloat16 patch_dim ....................................... 16 per_split_data_args_path ........................ None perform_initialization .......................... True pin_cpu_grads ................................... True pin_cpu_params .................................. True pipeline_model_parallel_comm_backend ............ None pipeline_model_parallel_size .................... 1 pipeline_model_parallel_split_rank .............. None position_embedding_type ......................... rope pretrained_checkpoint ........................... None profile ......................................... False profile_ranks ................................... [0] profile_step_end ................................ 12 profile_step_start .............................. 10 q_lora_rank ..................................... None qk_head_dim ..................................... 128 qk_layernorm .................................... False qk_pos_emb_head_dim ............................. 64 query_in_block_prob ............................. 0.1 rampup_batch_size ............................... None rank ............................................ 0 recompute_granularity ........................... None recompute_method ................................ None recompute_num_layers ............................ None record_memory_history ........................... False relative_attention_max_distance ................. 128 relative_attention_num_buckets .................. 32 replication ..................................... False replication_factor .............................. 2 replication_jump ................................ None rerun_mode ...................................... disabled reset_attention_mask ............................ False reset_position_ids .............................. False result_rejected_tracker_filename ................ None retriever_report_topk_accuracies ................ [] retriever_score_scaling ......................... False retriever_seq_length ............................ 256 retro_add_retriever ............................. False retro_attention_gate ............................ 1 retro_cyclic_train_iters ........................ None retro_encoder_attention_dropout ................. 0.1 retro_encoder_hidden_dropout .................... 0.1 retro_encoder_layers ............................ 2 retro_num_neighbors ............................. 2 retro_num_retrieved_chunks ...................... 2 retro_project_dir ............................... None retro_verify_neighbor_count ..................... True rope_scaling_factor ............................. 8.0 rotary_base ..................................... 10000 rotary_interleaved .............................. False rotary_percent .................................. 1.0 rotary_scaling_factor ........................... 1.0 rotary_seq_len_interpolation_factor ............. None s3_cache_path ................................... None sample_rate ..................................... 1.0 save ............................................ None save_interval ................................... None scatter_gather_tensors_in_pipeline .............. True seed ............................................ 42 seq_length ...................................... 2048 sequence_parallel ............................... False sgd_momentum .................................... 0.9 short_seq_prob .................................. 0.1 skip_train ...................................... False skipped_train_samples ........................... 0 spec ............................................ None split ........................................... None squared_relu .................................... False start_weight_decay .............................. 0.01 straggler_ctrlr_port ............................ 65535 straggler_minmax_count .......................... 1 suggested_communication_unit_size ............... 400000000 swiglu .......................................... True swin_backbone_type .............................. tiny te_rng_tracker .................................. False tensor_model_parallel_size ...................... 1 tensorboard_dir ................................. None tensorboard_log_interval ........................ 1 tensorboard_queue_size .......................... 1000 test_data_path .................................. None test_mode ....................................... False tiktoken_num_special_tokens ..................... 1000 tiktoken_pattern ................................ None tiktoken_special_tokens ......................... None timing_log_level ................................ 0 timing_log_option ............................... minmax titles_data_path ................................ None tokenizer_model ................................. EleutherAI/pythia-12b tokenizer_type .................................. HuggingFaceTokenizer tp_comm_bootstrap_backend ....................... nccl tp_comm_bulk_dgrad .............................. True tp_comm_bulk_wgrad .............................. True tp_comm_overlap ................................. False tp_comm_overlap_ag .............................. True tp_comm_overlap_cfg ............................. None tp_comm_overlap_rs .............................. True tp_comm_overlap_rs_dgrad ........................ False tp_comm_split_ag ................................ True tp_comm_split_rs ................................ True train_data_path ................................. None train_iters ..................................... None train_samples ................................... None train_sync_interval ............................. None transformer_impl ................................ transformer_engine transformer_pipeline_model_parallel_size ........ 1 untie_embeddings_and_output_weights ............. True use_checkpoint_args ............................. False use_checkpoint_opt_param_scheduler .............. False use_cpu_initialization .......................... None use_custom_fsdp ................................. False use_dist_ckpt ................................... True use_dist_ckpt_deprecated ........................ False use_distributed_optimizer ....................... True use_flash_attn .................................. False use_legacy_models ............................... False use_mp_args_from_checkpoint_args ................ False use_one_sent_docs ............................... False use_persistent_ckpt_worker ...................... False use_precision_aware_optimizer ................... False use_pytorch_profiler ............................ False use_ring_exchange_p2p ........................... False use_rope_scaling ................................ False use_rotary_position_embeddings .................. False use_tokenizer_model_from_checkpoint_args ........ True use_torch_fsdp2 ................................. False use_torch_optimizer_for_cpu_offload ............. False use_tp_pp_dp_mapping ............................ False v_head_dim ...................................... 128 valid_data_path ................................. None variable_seq_lengths ............................ False virtual_pipeline_model_parallel_size ............ None vision_backbone_type ............................ vit vision_pretraining .............................. False vision_pretraining_type ......................... classify vocab_extra_ids ................................. 0 vocab_file ...................................... None vocab_size ...................................... None wandb_exp_name .................................. wandb_project ................................... wandb_save_dir .................................. weight_decay .................................... 0.01 weight_decay_incr_style ......................... constant wgrad_deferral_limit ............................ 0 world_size ...................................... 1 yaml_cfg ........................................ None -------------------- end of arguments --------------------- 2026-07-23:04:35:40,762 INFO [num_microbatches_calculator.py:228] setting number of microbatches to constant 1 > building HuggingFaceTokenizer tokenizer ... 2026-07-23:04:35:40,960 INFO [_client.py:1038] HTTP Request: HEAD https://huggingface.co/EleutherAI/pythia-12b/resolve/main/config.json "HTTP/1.1 307 Temporary Redirect" 2026-07-23:04:35:40,972 INFO [_client.py:1038] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/EleutherAI/pythia-12b/bb1e3e710cdf6b524461d543cfb5ba773f0a81b6/config.json "HTTP/1.1 200 OK" 2026-07-23:04:35:41,096 INFO [_client.py:1038] HTTP Request: HEAD https://huggingface.co/EleutherAI/pythia-12b/resolve/main/tokenizer_config.json "HTTP/1.1 307 Temporary Redirect" 2026-07-23:04:35:41,108 INFO [_client.py:1038] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/EleutherAI/pythia-12b/bb1e3e710cdf6b524461d543cfb5ba773f0a81b6/tokenizer_config.json "HTTP/1.1 200 OK" 2026-07-23:04:35:41,217 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/models/EleutherAI/pythia-12b/tree/main/additional_chat_templates?recursive=false&expand=false "HTTP/1.1 404 Not Found" 2026-07-23:04:35:41,283 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/models/EleutherAI/pythia-12b/tree/main?recursive=true&expand=false "HTTP/1.1 200 OK" > padded vocab (size: 50277) with 27 dummy tokens (new size: 50304) WARNING: one_logger package is required to enable e2e metrics tracking. please go to https://confluence.nvidia.com/display/MLWFO/Package+Repositories for details to install it 2026-07-23:04:35:41,497 WARNING [rerun_state_machine.py:239] RerunStateMachine initialized in mode RerunMode.DISABLED > initializing torch distributed ... [W723 04:35:41.014936943 CUDAAllocatorConfig.h:28] Warning: expandable_segments not supported on this platform (function operator()) > initialized tensor model parallel with size 1 > initialized pipeline model parallel with size 1 > setting random seeds to 42 ... > compiling dataset index builder ... make: Entering directory '/workspace/FLAME-MoE/Megatron-LM/megatron/core/datasets' make: Nothing to be done for 'default'. make: Leaving directory '/workspace/FLAME-MoE/Megatron-LM/megatron/core/datasets' >>> done with dataset index builder. Compilation time: 0.808 seconds > compiling and loading fused kernels ... >>> done with compiling and loading fused kernels. Compilation time: 0.362 seconds building GPT model ... > number of parameters on (tensor, pipeline) model parallel rank (0, 0): 61678272 /workspace/FLAME-MoE/Megatron-LM/megatron/core/transformer/transformer_layer.py:339: UserWarning: TransformerLayer._get_layer_offset is deprecated.Please use get_transformer_layer_offset instead. warnings.warn( loading distributed checkpoint from /workspace/FLAME-MoE/results/phase0/runs/flame192_g3_moe/ckpt at iteration 3068 /workspace/FLAME-MoE/Megatron-LM/megatron/core/dist_checkpointing/strategies/torch.py:847: FutureWarning: `load_state_dict` is deprecated and will be removed in future versions. Please use `load` instead. checkpoint.load_state_dict( /usr/local/lib/python3.11/dist-packages/torch/distributed/checkpoint/planner_helpers.py:306: FutureWarning: Please use DTensor instead and we are deprecating ShardedTensor. device = getattr(value, "device", None) /workspace/FLAME-MoE/.venv/lib/python3.11/site-packages/transformer_engine/pytorch/module/base.py:566: FutureWarning: You are using `torch.load` with `weights_only=False` (the current default value), which uses the default pickle module implicitly. It is possible to construct malicious pickle data which will execute arbitrary code during unpickling (See https://github.com/pytorch/pytorch/blob/main/SECURITY.md#untrusted-models for more details). In a future release, the default value for `weights_only` will be flipped to `True`. This limits the functions that could be executed during unpickling. Arbitrary objects will no longer be allowed to be loaded via this mode unless they are explicitly allowlisted by the user via `torch.serialization.add_safe_globals`. We recommend you start setting `weights_only=True` for any use case where you don't have full control of the loaded file. Please open an issue on GitHub for any issues related to this experimental feature. state = torch.load(state, map_location="cuda") checkpoint version 3.0 2026-07-23:04:35:47,089 WARNING [rerun_state_machine.py:811] RerunStateMachine disabled via CLI, ignoring machine state saved in checkpoint successfully loaded checkpoint from /workspace/FLAME-MoE/results/phase0/runs/flame192_g3_moe/ckpt [ t 1/1, p 1/1 ] at iteration 3068 2026-07-23:04:35:47,237 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc "HTTP/1.1 200 OK" 2026-07-23:04:35:47,486 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc "HTTP/1.1 200 OK" 2026-07-23:04:35:47,553 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc/revision/210d026faf9955653af8916fad021475a3f00453 "HTTP/1.1 200 OK" 2026-07-23:04:35:47,644 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc/tree/210d026faf9955653af8916fad021475a3f00453?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:35:48,103 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc/tree/210d026faf9955653af8916fad021475a3f00453/ARC-Challenge?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:35:48,634 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc "HTTP/1.1 200 OK" 2026-07-23:04:35:48,865 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc "HTTP/1.1 200 OK" 2026-07-23:04:35:49,121 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc/tree/210d026faf9955653af8916fad021475a3f00453/ARC-Easy?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:35:49,264 WARNING [task.py:799] [Task: boolq] metric acc is defined, but aggregation is not. using default aggregation=mean 2026-07-23:04:35:49,264 WARNING [task.py:811] [Task: boolq] metric acc is defined, but higher_is_better is not. using default higher_is_better=True 2026-07-23:04:35:49,352 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue "HTTP/1.1 200 OK" 2026-07-23:04:35:49,583 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue "HTTP/1.1 200 OK" 2026-07-23:04:35:49,647 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/revision/3de24cf8022e94f4ee4b9d55a6f539891524d646 "HTTP/1.1 200 OK" 2026-07-23:04:35:49,706 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/tree/3de24cf8022e94f4ee4b9d55a6f539891524d646?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:35:50,257 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/tree/3de24cf8022e94f4ee4b9d55a6f539891524d646/axb?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:35:50,395 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/revision/3de24cf8022e94f4ee4b9d55a6f539891524d646 "HTTP/1.1 200 OK" 2026-07-23:04:35:50,472 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/tree/3de24cf8022e94f4ee4b9d55a6f539891524d646/boolq?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:35:50,723 WARNING [task.py:799] [Task: copa] metric acc is defined, but aggregation is not. using default aggregation=mean 2026-07-23:04:35:50,724 WARNING [task.py:811] [Task: copa] metric acc is defined, but higher_is_better is not. using default higher_is_better=True 2026-07-23:04:35:50,777 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue "HTTP/1.1 200 OK" 2026-07-23:04:35:51,028 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue "HTTP/1.1 200 OK" 2026-07-23:04:35:51,352 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/tree/3de24cf8022e94f4ee4b9d55a6f539891524d646/copa?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:35:51,494 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag "HTTP/1.1 200 OK" 2026-07-23:04:35:51,721 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag "HTTP/1.1 200 OK" 2026-07-23:04:35:51,785 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag/revision/218ec52e09a7e7462a5400043bb9a69a41d06b76 "HTTP/1.1 200 OK" 2026-07-23:04:35:51,844 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag/tree/218ec52e09a7e7462a5400043bb9a69a41d06b76?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:35:52,093 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag/tree/218ec52e09a7e7462a5400043bb9a69a41d06b76/data?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:35:52,153 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag/revision/218ec52e09a7e7462a5400043bb9a69a41d06b76 "HTTP/1.1 200 OK" 2026-07-23:04:35:54,196 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai "HTTP/1.1 200 OK" 2026-07-23:04:35:54,483 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai "HTTP/1.1 200 OK" 2026-07-23:04:35:54,543 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/revision/900124bf3b8235c6daf21033af9948b3f07346c4 "HTTP/1.1 200 OK" 2026-07-23:04:35:54,612 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/tree/900124bf3b8235c6daf21033af9948b3f07346c4?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:35:54,840 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/tree/900124bf3b8235c6daf21033af9948b3f07346c4/default%2Ftest?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:35:54,896 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/tree/900124bf3b8235c6daf21033af9948b3f07346c4/default?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:35:54,956 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/revision/900124bf3b8235c6daf21033af9948b3f07346c4 "HTTP/1.1 200 OK" 2026-07-23:04:35:54,968 WARNING [task.py:325] [Task: lambada_openai] has_training_docs and has_validation_docs are False, using test_docs as fewshot_docs but this is not recommended. 2026-07-23:04:35:54,968 WARNING [task.py:325] [Task: lambada_openai] has_training_docs and has_validation_docs are False, using test_docs as fewshot_docs but this is not recommended. 2026-07-23:04:35:55,087 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa "HTTP/1.1 200 OK" 2026-07-23:04:35:55,325 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa "HTTP/1.1 200 OK" 2026-07-23:04:35:55,386 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/revision/388097ea7776314e93a529163e0fea805b8a6454 "HTTP/1.1 200 OK" 2026-07-23:04:35:55,450 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/tree/388097ea7776314e93a529163e0fea805b8a6454?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:35:55,753 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/tree/388097ea7776314e93a529163e0fea805b8a6454/additional?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:35:55,812 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/revision/388097ea7776314e93a529163e0fea805b8a6454 "HTTP/1.1 200 OK" 2026-07-23:04:35:55,888 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/tree/388097ea7776314e93a529163e0fea805b8a6454/main?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:35:56,151 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa "HTTP/1.1 200 OK" 2026-07-23:04:35:56,209 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa "HTTP/1.1 200 OK" 2026-07-23:04:35:56,272 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa/tree/main?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:35:56,328 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/datasets/ybisk/piqa/resolve/main/piqa.py "HTTP/1.1 307 Temporary Redirect" 2026-07-23:04:35:56,339 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/resolve-cache/datasets/ybisk/piqa/2e8ac2dffd59bac8c3c6714948f4c551a0848bb0/piqa.py?%2Fdatasets%2Fybisk%2Fpiqa%2Fresolve%2Fmain%2Fpiqa.py=&etag=%220f0d96172f9bf7aa10069374b3236ef083cfeb8e%22 "HTTP/1.1 200 OK" 2026-07-23:04:35:56,884 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa/revision/main "HTTP/1.1 200 OK" 2026-07-23:04:35:56,943 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa/tree/main?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:35:57,365 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq "HTTP/1.1 200 OK" 2026-07-23:04:35:57,575 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq "HTTP/1.1 200 OK" 2026-07-23:04:35:57,635 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq/revision/2c94ad3e1aafab77146f384e23536f97a4849815 "HTTP/1.1 200 OK" 2026-07-23:04:35:57,696 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq/tree/2c94ad3e1aafab77146f384e23536f97a4849815?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:35:57,904 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq/tree/2c94ad3e1aafab77146f384e23536f97a4849815/data?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:35:58,008 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq/revision/2c94ad3e1aafab77146f384e23536f97a4849815 "HTTP/1.1 200 OK" 2026-07-23:04:35:58,429 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande "HTTP/1.1 200 OK" 2026-07-23:04:35:58,706 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande "HTTP/1.1 200 OK" 2026-07-23:04:35:58,765 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/revision/01e74176c63542e6b0bcb004dcdea22d94fb67b5 "HTTP/1.1 200 OK" 2026-07-23:04:35:58,833 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/tree/01e74176c63542e6b0bcb004dcdea22d94fb67b5?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:35:59,059 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/tree/01e74176c63542e6b0bcb004dcdea22d94fb67b5/winogrande_debiased?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:35:59,122 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/revision/01e74176c63542e6b0bcb004dcdea22d94fb67b5 "HTTP/1.1 200 OK" 2026-07-23:04:35:59,214 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/tree/01e74176c63542e6b0bcb004dcdea22d94fb67b5/winogrande_xl?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:35:59,861 WARNING [evaluator.py:270] Overwriting default num_fewshot of winogrande from None to 0 2026-07-23:04:35:59,861 WARNING [evaluator.py:270] Overwriting default num_fewshot of sciq from None to 0 2026-07-23:04:35:59,861 WARNING [evaluator.py:270] Overwriting default num_fewshot of piqa from None to 0 2026-07-23:04:35:59,861 WARNING [evaluator.py:270] Overwriting default num_fewshot of openbookqa from None to 0 2026-07-23:04:35:59,861 WARNING [evaluator.py:270] Overwriting default num_fewshot of lambada_openai from None to 0 2026-07-23:04:35:59,861 WARNING [evaluator.py:270] Overwriting default num_fewshot of hellaswag from None to 0 2026-07-23:04:35:59,861 WARNING [evaluator.py:270] Overwriting default num_fewshot of copa from None to 0 2026-07-23:04:35:59,861 WARNING [evaluator.py:270] Overwriting default num_fewshot of boolq from None to 0 2026-07-23:04:35:59,861 WARNING [evaluator.py:270] Overwriting default num_fewshot of arc_easy from None to 0 2026-07-23:04:35:59,861 WARNING [evaluator.py:270] Overwriting default num_fewshot of arc_challenge from None to 0 2026-07-23:04:35:59,896 INFO [task.py:415] Building contexts for winogrande on rank 0... 0%| | 0/1267 [00:00 /workspace/FLAME-MoE/results/ablations/flame1e18_downstream.csv ==================================================================== [run] flame192_g3_temporal mode=temporal TEMPORAL=1 H=192 heads=12 ffn=1026 experts=192 topk=18 moe_ffn=44 shared=264 [run] tasks=arc_challenge,arc_easy,hellaswag,openbookqa,piqa,winogrande,boolq,copa,lambada_openai,sciq out=/workspace/FLAME-MoE/results/phase0/runs/flame192_g3_temporal/lmeval_1e18_0shot ==================================================================== /usr/local/lib/python3.11/dist-packages/torch/cuda/__init__.py:58: FutureWarning: The pynvml package is deprecated. Please install nvidia-ml-py instead. If you did not install pynvml directly, please report this to the maintainers of the package that installed pynvml for you. import pynvml # type: ignore[import] [run_lmeval] stubbed VLM backends (hf_vlms, vllm_vlms) /usr/local/lib/python3.11/dist-packages/torch/cuda/__init__.py:58: FutureWarning: The pynvml package is deprecated. Please install nvidia-ml-py instead. If you did not install pynvml directly, please report this to the maintainers of the package that installed pynvml for you. import pynvml # type: ignore[import] /usr/local/lib/python3.11/dist-packages/torch/cuda/__init__.py:58: FutureWarning: The pynvml package is deprecated. Please install nvidia-ml-py instead. If you did not install pynvml directly, please report this to the maintainers of the package that installed pynvml for you. import pynvml # type: ignore[import] /workspace/FLAME-MoE/.venv/lib/python3.11/site-packages/transformer_engine/pytorch/attention.py:108: UserWarning: To use flash-attn v3, please use the following commands to install: (1) pip install "git+https://github.com/Dao-AILab/flash-attention.git#egg=flashattn-hopper&subdirectory=hopper" (2) python_path=`python -c "import site; print(site.getsitepackages()[0])"` (3) mkdir -p $python_path/flashattn_hopper (4) wget -P $python_path/flashattn_hopper https://raw.githubusercontent.com/Dao-AILab/flash-attention/main/hopper/flash_attn_interface.py warnings.warn( [temporal] rolling-residency router installed (evict=min_logit) [run_lmeval] installed temporal residency router (native regime) /usr/local/lib/python3.11/dist-packages/requests/__init__.py:113: RequestsDependencyWarning: urllib3 (2.2.3) or chardet (6.0.0.post1)/charset_normalizer (3.3.2) doesn't match a supported version! warnings.warn( 2026-07-23:04:45:45,471 INFO [__main__.py:279] Verbosity set to INFO 2026-07-23:04:46:16,896 INFO [__init__.py:459] The tag 'arc_ca' is already registered as a group, this tag will not be registered. This may affect tasks you want to call. 2026-07-23:04:46:16,910 INFO [__init__.py:459] The tag 'arc_ca' is already registered as a group, this tag will not be registered. This may affect tasks you want to call. 2026-07-23:04:46:34,021 INFO [__main__.py:376] Selected Tasks: ['arc_challenge', 'arc_easy', 'boolq', 'copa', 'hellaswag', 'lambada_openai', 'openbookqa', 'piqa', 'sciq', 'winogrande'] 2026-07-23:04:46:34,037 INFO [evaluator.py:164] Setting random seed to 42 | Setting numpy seed to 42 | Setting torch manual seed to 42 | Setting fewshot manual seed to 42 2026-07-23:04:46:34,037 INFO [evaluator.py:201] Initializing megatron_lm model, with arguments: {} using world size: 1, data-parallel size: 1, context-parallel size: 1, hierarchical context-parallel sizes: Nonetensor-model-parallel size: 1, encoder-tensor-model-parallel size: 0, pipeline-model-parallel size: 1, encoder-pipeline-model-parallel size: 0 WARNING: overriding default arguments for exit_on_missing_checkpoint:True with exit_on_missing_checkpoint:False setting global batch size to 32 Number of virtual stages per pipeline stage: None accumulate and all-reduce gradients in fp32 for bfloat16 data type. using torch.bfloat16 for parameters ... ------------------------ arguments ------------------------ account_for_embedding_in_pipeline_split ......... False account_for_loss_in_pipeline_split .............. False accumulate_allreduce_grads_in_fp32 .............. True adam_beta1 ...................................... 0.9 adam_beta2 ...................................... 0.999 adam_eps ........................................ 1e-08 add_bias_linear ................................. False add_position_embedding .......................... True add_qkv_bias .................................... False adlr_autoresume ................................. False adlr_autoresume_interval ........................ 1000 align_grad_reduce ............................... True align_param_gather .............................. False app_tag_run_name ................................ None app_tag_run_version ............................. 0.0.0 apply_layernorm_1p .............................. False apply_query_key_layer_scaling ................... False apply_residual_connection_post_layernorm ........ False apply_rope_fusion ............................... True async_save ...................................... None async_tensor_model_parallel_allreduce ........... True attention_backend ............................... AttnBackend.auto attention_dropout ............................... 0.0 attention_softmax_in_fp32 ....................... False auto_detect_ckpt_format ......................... False barrier_with_L1_time ............................ True bert_binary_head ................................ True bert_embedder_type .............................. megatron bert_load ....................................... None bf16 ............................................ True bias_dropout_fusion ............................. True bias_gelu_fusion ................................ False bias_swiglu_fusion .............................. True biencoder_projection_dim ........................ 0 biencoder_shared_query_context_model ............ False block_data_path ................................. None calc_ft_timeouts ................................ False calculate_per_token_loss ........................ False check_for_large_grads ........................... False check_for_nan_in_loss_and_grad .................. True check_for_spiky_loss ............................ False check_weight_hash_across_dp_replicas_interval ... None ckpt_assume_constant_structure .................. False ckpt_convert_format ............................. None ckpt_convert_save ............................... None ckpt_convert_update_legacy_dist_opt_format ...... False ckpt_format ..................................... torch_dist ckpt_fully_parallel_load ........................ False ckpt_fully_parallel_save ........................ True ckpt_fully_parallel_save_deprecated ............. False ckpt_step ....................................... None classes_fraction ................................ 1.0 clip_grad ....................................... 1.0 clone_scatter_output_in_embedding ............... True config_logger_dir ............................... consumed_train_samples .......................... 0 consumed_valid_samples .......................... 0 context_parallel_size ........................... 1 cp_comm_type .................................... ['p2p'] create_attention_mask_in_dataloader ............. True cross_entropy_loss_fusion ....................... False cuda_graph_warmup_steps ......................... 3 data_args_path .................................. None data_cache_path ................................. None data_parallel_random_init ....................... False data_parallel_sharding_strategy ................. no_shard data_parallel_size .............................. 1 data_path ....................................... None data_per_class_fraction ......................... 1.0 data_sharding ................................... True dataloader_type ................................. single ddp_average_in_collective ....................... False ddp_bucket_size ................................. None ddp_num_buckets ................................. None ddp_pad_buckets_for_high_nccl_busbw ............. False decoder_first_pipeline_num_layers ............... None decoder_last_pipeline_num_layers ................ None decoder_num_layers .............................. None decoder_seq_length .............................. None decoupled_lr .................................... None decoupled_min_lr ................................ None decrease_batch_size_if_needed ................... False defer_embedding_wgrad_compute ................... False deprecated_use_mcore_models ..................... False deterministic_mode .............................. False dino_bottleneck_size ............................ 256 dino_freeze_last_layer .......................... 1 dino_head_hidden_size ........................... 2048 dino_local_crops_number ......................... 10 dino_local_img_size ............................. 96 dino_norm_last_layer ............................ False dino_teacher_temp ............................... 0.07 dino_warmup_teacher_temp ........................ 0.04 dino_warmup_teacher_temp_epochs ................. 30 disable_straggler_on_startup .................... False dist_ckpt_format_deprecated ..................... None dist_ckpt_strictness ............................ assume_ok_unexpected distribute_saved_activations .................... False distributed_backend ............................. nccl distributed_timeout_minutes ..................... 10 embedding_path .................................. None empty_unused_memory_level ....................... 0 enable_cuda_graph ............................... False enable_ft_package ............................... False enable_gloo_process_groups ...................... True enable_one_logger ............................... True encoder_num_layers .............................. 9 encoder_pipeline_model_parallel_size ............ 0 encoder_seq_length .............................. 2048 encoder_tensor_model_parallel_size .............. 0 end_weight_decay ................................ 0.01 eod_mask_loss ................................... False error_injection_rate ............................ 0 error_injection_type ............................ transient_error eval_interval ................................... 1000 eval_iters ...................................... 100 evidence_data_path .............................. None exit_duration_in_mins ........................... None exit_interval ................................... None exit_on_missing_checkpoint ...................... False exit_signal_handler ............................. False exp_avg_dtype ................................... torch.float32 exp_avg_sq_dtype ................................ torch.float32 expert_model_parallel_size ...................... 1 expert_tensor_parallel_size ..................... 1 ffn_hidden_size ................................. 1026 finetune ........................................ False flash_decode .................................... False fp16 ............................................ False fp16_lm_cross_entropy ........................... False fp32_residual_connection ........................ False fp8 ............................................. None fp8_amax_compute_algo ........................... most_recent fp8_amax_history_len ............................ 1 fp8_interval .................................... 1 fp8_margin ...................................... 0 fp8_param_gather ................................ False fp8_wgrad ....................................... True global_batch_size ............................... 32 grad_reduce_in_bf16 ............................. False gradient_accumulation_fusion .................... False gradient_reduce_div_fusion ...................... True group_query_attention ........................... False head_lr_mult .................................... 1.0 hidden_dropout .................................. 0.0 hidden_size ..................................... 192 hierarchical_context_parallel_sizes ............. None hybrid_attention_ratio .......................... 0.0 hybrid_mlp_ratio ................................ 0.0 hybrid_override_pattern ......................... None hysteresis ...................................... 2 ict_head_size ................................... None ict_load ........................................ None img_h ........................................... 224 img_w ........................................... 224 indexer_batch_size .............................. 128 indexer_log_interval ............................ 1000 inference_batch_times_seqlen_threshold .......... -1 inference_max_batch_size ........................ 8 inference_max_seq_length ........................ 2560 inference_rng_tracker ........................... False init_method_std ................................. 0.02 init_method_xavier_uniform ...................... False init_model_with_meta_device ..................... False initial_loss_scale .............................. 4294967296 iter_per_epoch .................................. 1250 iterations_to_skip .............................. [] keep_fp8_transpose_cache_when_using_custom_fsdp . False kv_channels ..................................... 16 kv_lora_rank .................................... 32 lazy_mpu_init ................................... None load ............................................ /workspace/FLAME-MoE/results/phase0/runs/flame192_g3_temporal/ckpt local_rank ...................................... 0 log_interval .................................... 100 log_loss_scale_to_tensorboard ................... True log_memory_to_tensorboard ....................... False log_num_zeros_in_grad ........................... False log_params_norm ................................. False log_progress .................................... False log_straggler ................................... False log_throughput .................................. False log_timers_to_tensorboard ....................... False log_validation_ppl_to_tensorboard ............... False log_world_size_to_tensorboard ................... False logging_level ................................... None loss_scale ...................................... None loss_scale_window ............................... 1000 lr .............................................. None lr_decay_iters .................................. None lr_decay_samples ................................ None lr_decay_style .................................. linear lr_warmup_fraction .............................. None lr_warmup_init .................................. 0.0 lr_warmup_iters ................................. 0 lr_warmup_samples ............................... 0 lr_wsd_decay_iters .............................. None lr_wsd_decay_samples ............................ None lr_wsd_decay_style .............................. exponential main_grads_dtype ................................ torch.float32 main_params_dtype ............................... torch.float32 make_vocab_size_divisible_by .................... 128 manual_gc ....................................... False manual_gc_eval .................................. True manual_gc_interval .............................. 0 mask_factor ..................................... 1.0 mask_prob ....................................... 0.15 mask_type ....................................... random masked_softmax_fusion ........................... True max_position_embeddings ......................... 2048 max_tokens_to_oom ............................... 10000000 memory_snapshot_path ............................ snapshot.pickle merge_file ...................................... None micro_batch_size ................................ 32 microbatch_group_size_per_vp_stage .............. None min_loss_scale .................................. 1.0 min_lr .......................................... 0.0 mmap_bin_files .................................. True mock_data ....................................... False moe_aux_loss_coeff .............................. 0.01 moe_enable_deepep ............................... False moe_expert_capacity_factor ...................... None moe_extended_tp ................................. False moe_ffn_hidden_size ............................. 44 moe_grouped_gemm ................................ True moe_input_jitter_eps ............................ None moe_layer_freq .................................. [0, 1, 1, 1, 1, 1, 1, 1, 1] moe_layer_recompute ............................. False moe_pad_expert_input_to_capacity ................ False moe_per_layer_logging ........................... False moe_permute_fusion .............................. False moe_router_bias_update_rate ..................... 0.001 moe_router_dtype ................................ fp32 moe_router_enable_expert_bias ................... False moe_router_group_topk ........................... None moe_router_load_balancing_type .................. aux_loss moe_router_num_groups ........................... None moe_router_pre_softmax .......................... True moe_router_score_function ....................... softmax moe_router_topk ................................. 18 moe_router_topk_scaling_factor .................. None moe_shared_expert_intermediate_size ............. 264 moe_shared_expert_overlap ....................... False moe_token_dispatcher_type ....................... alltoall moe_token_drop_policy ........................... probs moe_use_legacy_grouped_gemm ..................... False moe_use_upcycling ............................... False moe_z_loss_coeff ................................ 0.001 mscale .......................................... 1.0 mscale_all_dim .................................. 1.0 multi_latent_attention .......................... False nccl_communicator_config_path ................... None no_load_optim ................................... True no_load_rng ..................................... True no_persist_layer_norm ........................... False no_save_optim ................................... None no_save_rng ..................................... None non_persistent_ckpt_type ........................ None non_persistent_global_ckpt_dir .................. None non_persistent_local_ckpt_algo .................. fully_parallel non_persistent_local_ckpt_dir ................... None non_persistent_save_interval .................... None norm_epsilon .................................... 1e-06 normalization ................................... RMSNorm num_attention_heads ............................. 12 num_channels .................................... 3 num_classes ..................................... 1000 num_dataset_builder_threads ..................... 1 num_distributed_optimizer_instances ............. 1 num_experts ..................................... 192 num_layers ...................................... 9 num_layers_per_virtual_pipeline_stage ........... None num_query_groups ................................ 1 num_virtual_stages_per_pipeline_rank ............ None num_workers ..................................... 2 one_logger_async ................................ False one_logger_project .............................. megatron-lm one_logger_run_name ............................. None onnx_safe ....................................... None openai_gelu ..................................... False optimizer ....................................... adam optimizer_cpu_offload ........................... False optimizer_offload_fraction ...................... 1.0 output_bert_embeddings .......................... False overlap_cpu_optimizer_d2h_h2d ................... False overlap_grad_reduce ............................. False overlap_p2p_comm ................................ False overlap_p2p_comm_warmup_flush ................... False overlap_param_gather ............................ False overlap_param_gather_with_optimizer_step ........ False override_opt_param_scheduler .................... False params_dtype .................................... torch.bfloat16 patch_dim ....................................... 16 per_split_data_args_path ........................ None perform_initialization .......................... True pin_cpu_grads ................................... True pin_cpu_params .................................. True pipeline_model_parallel_comm_backend ............ None pipeline_model_parallel_size .................... 1 pipeline_model_parallel_split_rank .............. None position_embedding_type ......................... rope pretrained_checkpoint ........................... None profile ......................................... False profile_ranks ................................... [0] profile_step_end ................................ 12 profile_step_start .............................. 10 q_lora_rank ..................................... None qk_head_dim ..................................... 128 qk_layernorm .................................... False qk_pos_emb_head_dim ............................. 64 query_in_block_prob ............................. 0.1 rampup_batch_size ............................... None rank ............................................ 0 recompute_granularity ........................... None recompute_method ................................ None recompute_num_layers ............................ None record_memory_history ........................... False relative_attention_max_distance ................. 128 relative_attention_num_buckets .................. 32 replication ..................................... False replication_factor .............................. 2 replication_jump ................................ None rerun_mode ...................................... disabled reset_attention_mask ............................ False reset_position_ids .............................. False result_rejected_tracker_filename ................ None retriever_report_topk_accuracies ................ [] retriever_score_scaling ......................... False retriever_seq_length ............................ 256 retro_add_retriever ............................. False retro_attention_gate ............................ 1 retro_cyclic_train_iters ........................ None retro_encoder_attention_dropout ................. 0.1 retro_encoder_hidden_dropout .................... 0.1 retro_encoder_layers ............................ 2 retro_num_neighbors ............................. 2 retro_num_retrieved_chunks ...................... 2 retro_project_dir ............................... None retro_verify_neighbor_count ..................... True rope_scaling_factor ............................. 8.0 rotary_base ..................................... 10000 rotary_interleaved .............................. False rotary_percent .................................. 1.0 rotary_scaling_factor ........................... 1.0 rotary_seq_len_interpolation_factor ............. None s3_cache_path ................................... None sample_rate ..................................... 1.0 save ............................................ None save_interval ................................... None scatter_gather_tensors_in_pipeline .............. True seed ............................................ 42 seq_length ...................................... 2048 sequence_parallel ............................... False sgd_momentum .................................... 0.9 short_seq_prob .................................. 0.1 skip_train ...................................... False skipped_train_samples ........................... 0 spec ............................................ None split ........................................... None squared_relu .................................... False start_weight_decay .............................. 0.01 straggler_ctrlr_port ............................ 65535 straggler_minmax_count .......................... 1 suggested_communication_unit_size ............... 400000000 swiglu .......................................... True swin_backbone_type .............................. tiny te_rng_tracker .................................. False tensor_model_parallel_size ...................... 1 tensorboard_dir ................................. None tensorboard_log_interval ........................ 1 tensorboard_queue_size .......................... 1000 test_data_path .................................. None test_mode ....................................... False tiktoken_num_special_tokens ..................... 1000 tiktoken_pattern ................................ None tiktoken_special_tokens ......................... None timing_log_level ................................ 0 timing_log_option ............................... minmax titles_data_path ................................ None tokenizer_model ................................. EleutherAI/pythia-12b tokenizer_type .................................. HuggingFaceTokenizer tp_comm_bootstrap_backend ....................... nccl tp_comm_bulk_dgrad .............................. True tp_comm_bulk_wgrad .............................. True tp_comm_overlap ................................. False tp_comm_overlap_ag .............................. True tp_comm_overlap_cfg ............................. None tp_comm_overlap_rs .............................. True tp_comm_overlap_rs_dgrad ........................ False tp_comm_split_ag ................................ True tp_comm_split_rs ................................ True train_data_path ................................. None train_iters ..................................... None train_samples ................................... None train_sync_interval ............................. None transformer_impl ................................ transformer_engine transformer_pipeline_model_parallel_size ........ 1 untie_embeddings_and_output_weights ............. True use_checkpoint_args ............................. False use_checkpoint_opt_param_scheduler .............. False use_cpu_initialization .......................... None use_custom_fsdp ................................. False use_dist_ckpt ................................... True use_dist_ckpt_deprecated ........................ False use_distributed_optimizer ....................... True use_flash_attn .................................. False use_legacy_models ............................... False use_mp_args_from_checkpoint_args ................ False use_one_sent_docs ............................... False use_persistent_ckpt_worker ...................... False use_precision_aware_optimizer ................... False use_pytorch_profiler ............................ False use_ring_exchange_p2p ........................... False use_rope_scaling ................................ False use_rotary_position_embeddings .................. False use_tokenizer_model_from_checkpoint_args ........ True use_torch_fsdp2 ................................. False use_torch_optimizer_for_cpu_offload ............. False use_tp_pp_dp_mapping ............................ False v_head_dim ...................................... 128 valid_data_path ................................. None variable_seq_lengths ............................ False virtual_pipeline_model_parallel_size ............ None vision_backbone_type ............................ vit vision_pretraining .............................. False vision_pretraining_type ......................... classify vocab_extra_ids ................................. 0 vocab_file ...................................... None vocab_size ...................................... None wandb_exp_name .................................. wandb_project ................................... wandb_save_dir .................................. weight_decay .................................... 0.01 weight_decay_incr_style ......................... constant wgrad_deferral_limit ............................ 0 world_size ...................................... 1 yaml_cfg ........................................ None -------------------- end of arguments --------------------- 2026-07-23:04:46:34,041 INFO [num_microbatches_calculator.py:228] setting number of microbatches to constant 1 > building HuggingFaceTokenizer tokenizer ... 2026-07-23:04:46:34,226 INFO [_client.py:1038] HTTP Request: HEAD https://huggingface.co/EleutherAI/pythia-12b/resolve/main/config.json "HTTP/1.1 307 Temporary Redirect" 2026-07-23:04:46:34,238 INFO [_client.py:1038] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/EleutherAI/pythia-12b/bb1e3e710cdf6b524461d543cfb5ba773f0a81b6/config.json "HTTP/1.1 200 OK" 2026-07-23:04:46:34,367 INFO [_client.py:1038] HTTP Request: HEAD https://huggingface.co/EleutherAI/pythia-12b/resolve/main/tokenizer_config.json "HTTP/1.1 307 Temporary Redirect" 2026-07-23:04:46:34,379 INFO [_client.py:1038] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/EleutherAI/pythia-12b/bb1e3e710cdf6b524461d543cfb5ba773f0a81b6/tokenizer_config.json "HTTP/1.1 200 OK" 2026-07-23:04:46:34,480 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/models/EleutherAI/pythia-12b/tree/main/additional_chat_templates?recursive=false&expand=false "HTTP/1.1 404 Not Found" 2026-07-23:04:46:34,543 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/models/EleutherAI/pythia-12b/tree/main?recursive=true&expand=false "HTTP/1.1 200 OK" > padded vocab (size: 50277) with 27 dummy tokens (new size: 50304) WARNING: one_logger package is required to enable e2e metrics tracking. please go to https://confluence.nvidia.com/display/MLWFO/Package+Repositories for details to install it 2026-07-23:04:46:34,750 WARNING [rerun_state_machine.py:239] RerunStateMachine initialized in mode RerunMode.DISABLED > initializing torch distributed ... [W723 04:46:34.267564285 CUDAAllocatorConfig.h:28] Warning: expandable_segments not supported on this platform (function operator()) > initialized tensor model parallel with size 1 > initialized pipeline model parallel with size 1 > setting random seeds to 42 ... > compiling dataset index builder ... make: Entering directory '/workspace/FLAME-MoE/Megatron-LM/megatron/core/datasets' make: Nothing to be done for 'default'. make: Leaving directory '/workspace/FLAME-MoE/Megatron-LM/megatron/core/datasets' >>> done with dataset index builder. Compilation time: 0.836 seconds > compiling and loading fused kernels ... >>> done with compiling and loading fused kernels. Compilation time: 0.327 seconds building GPT model ... > number of parameters on (tensor, pipeline) model parallel rank (0, 0): 61678272 /workspace/FLAME-MoE/Megatron-LM/megatron/core/transformer/transformer_layer.py:339: UserWarning: TransformerLayer._get_layer_offset is deprecated.Please use get_transformer_layer_offset instead. warnings.warn( loading distributed checkpoint from /workspace/FLAME-MoE/results/phase0/runs/flame192_g3_temporal/ckpt at iteration 3068 /workspace/FLAME-MoE/Megatron-LM/megatron/core/dist_checkpointing/strategies/torch.py:847: FutureWarning: `load_state_dict` is deprecated and will be removed in future versions. Please use `load` instead. checkpoint.load_state_dict( /usr/local/lib/python3.11/dist-packages/torch/distributed/checkpoint/planner_helpers.py:306: FutureWarning: Please use DTensor instead and we are deprecating ShardedTensor. device = getattr(value, "device", None) /workspace/FLAME-MoE/.venv/lib/python3.11/site-packages/transformer_engine/pytorch/module/base.py:566: FutureWarning: You are using `torch.load` with `weights_only=False` (the current default value), which uses the default pickle module implicitly. It is possible to construct malicious pickle data which will execute arbitrary code during unpickling (See https://github.com/pytorch/pytorch/blob/main/SECURITY.md#untrusted-models for more details). In a future release, the default value for `weights_only` will be flipped to `True`. This limits the functions that could be executed during unpickling. Arbitrary objects will no longer be allowed to be loaded via this mode unless they are explicitly allowlisted by the user via `torch.serialization.add_safe_globals`. We recommend you start setting `weights_only=True` for any use case where you don't have full control of the loaded file. Please open an issue on GitHub for any issues related to this experimental feature. state = torch.load(state, map_location="cuda") checkpoint version 3.0 2026-07-23:04:46:41,105 WARNING [rerun_state_machine.py:811] RerunStateMachine disabled via CLI, ignoring machine state saved in checkpoint successfully loaded checkpoint from /workspace/FLAME-MoE/results/phase0/runs/flame192_g3_temporal/ckpt [ t 1/1, p 1/1 ] at iteration 3068 2026-07-23:04:46:41,231 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc "HTTP/1.1 200 OK" 2026-07-23:04:46:41,469 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc "HTTP/1.1 200 OK" 2026-07-23:04:46:41,535 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc/revision/210d026faf9955653af8916fad021475a3f00453 "HTTP/1.1 200 OK" 2026-07-23:04:46:41,596 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc/tree/210d026faf9955653af8916fad021475a3f00453?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:46:41,885 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc/tree/210d026faf9955653af8916fad021475a3f00453/ARC-Challenge?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:46:42,384 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc "HTTP/1.1 200 OK" 2026-07-23:04:46:42,632 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc "HTTP/1.1 200 OK" 2026-07-23:04:46:42,868 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc/tree/210d026faf9955653af8916fad021475a3f00453/ARC-Easy?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:46:43,015 WARNING [task.py:799] [Task: boolq] metric acc is defined, but aggregation is not. using default aggregation=mean 2026-07-23:04:46:43,015 WARNING [task.py:811] [Task: boolq] metric acc is defined, but higher_is_better is not. using default higher_is_better=True 2026-07-23:04:46:43,080 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue "HTTP/1.1 200 OK" 2026-07-23:04:46:43,315 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue "HTTP/1.1 200 OK" 2026-07-23:04:46:43,380 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/revision/3de24cf8022e94f4ee4b9d55a6f539891524d646 "HTTP/1.1 200 OK" 2026-07-23:04:46:43,445 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/tree/3de24cf8022e94f4ee4b9d55a6f539891524d646?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:46:43,709 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/tree/3de24cf8022e94f4ee4b9d55a6f539891524d646/axb?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:46:43,764 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/revision/3de24cf8022e94f4ee4b9d55a6f539891524d646 "HTTP/1.1 200 OK" 2026-07-23:04:46:43,837 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/tree/3de24cf8022e94f4ee4b9d55a6f539891524d646/boolq?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:46:44,085 WARNING [task.py:799] [Task: copa] metric acc is defined, but aggregation is not. using default aggregation=mean 2026-07-23:04:46:44,085 WARNING [task.py:811] [Task: copa] metric acc is defined, but higher_is_better is not. using default higher_is_better=True 2026-07-23:04:46:44,148 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue "HTTP/1.1 200 OK" 2026-07-23:04:46:44,380 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue "HTTP/1.1 200 OK" 2026-07-23:04:46:44,718 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/tree/3de24cf8022e94f4ee4b9d55a6f539891524d646/copa?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:46:44,859 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag "HTTP/1.1 200 OK" 2026-07-23:04:46:45,098 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag "HTTP/1.1 200 OK" 2026-07-23:04:46:45,156 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag/revision/218ec52e09a7e7462a5400043bb9a69a41d06b76 "HTTP/1.1 200 OK" 2026-07-23:04:46:45,213 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag/tree/218ec52e09a7e7462a5400043bb9a69a41d06b76?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:46:45,441 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag/tree/218ec52e09a7e7462a5400043bb9a69a41d06b76/data?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:46:45,507 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag/revision/218ec52e09a7e7462a5400043bb9a69a41d06b76 "HTTP/1.1 200 OK" 2026-07-23:04:46:48,059 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai "HTTP/1.1 200 OK" 2026-07-23:04:46:48,304 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai "HTTP/1.1 200 OK" 2026-07-23:04:46:48,364 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/revision/900124bf3b8235c6daf21033af9948b3f07346c4 "HTTP/1.1 200 OK" 2026-07-23:04:46:48,431 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/tree/900124bf3b8235c6daf21033af9948b3f07346c4?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:46:48,723 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/tree/900124bf3b8235c6daf21033af9948b3f07346c4/default%2Ftest?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:46:48,850 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/tree/900124bf3b8235c6daf21033af9948b3f07346c4/default?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:46:48,911 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/revision/900124bf3b8235c6daf21033af9948b3f07346c4 "HTTP/1.1 200 OK" 2026-07-23:04:46:48,931 WARNING [task.py:325] [Task: lambada_openai] has_training_docs and has_validation_docs are False, using test_docs as fewshot_docs but this is not recommended. 2026-07-23:04:46:48,931 WARNING [task.py:325] [Task: lambada_openai] has_training_docs and has_validation_docs are False, using test_docs as fewshot_docs but this is not recommended. 2026-07-23:04:46:49,093 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa "HTTP/1.1 200 OK" 2026-07-23:04:46:49,351 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa "HTTP/1.1 200 OK" 2026-07-23:04:46:49,411 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/revision/388097ea7776314e93a529163e0fea805b8a6454 "HTTP/1.1 200 OK" 2026-07-23:04:46:49,472 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/tree/388097ea7776314e93a529163e0fea805b8a6454?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:46:49,687 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/tree/388097ea7776314e93a529163e0fea805b8a6454/additional?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:46:49,743 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/revision/388097ea7776314e93a529163e0fea805b8a6454 "HTTP/1.1 200 OK" 2026-07-23:04:46:49,832 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/tree/388097ea7776314e93a529163e0fea805b8a6454/main?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:46:50,097 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa "HTTP/1.1 200 OK" 2026-07-23:04:46:50,157 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa "HTTP/1.1 200 OK" 2026-07-23:04:46:50,218 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa/tree/main?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:46:50,280 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/datasets/ybisk/piqa/resolve/main/piqa.py "HTTP/1.1 307 Temporary Redirect" 2026-07-23:04:46:50,292 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/resolve-cache/datasets/ybisk/piqa/2e8ac2dffd59bac8c3c6714948f4c551a0848bb0/piqa.py?%2Fdatasets%2Fybisk%2Fpiqa%2Fresolve%2Fmain%2Fpiqa.py=&etag=%220f0d96172f9bf7aa10069374b3236ef083cfeb8e%22 "HTTP/1.1 200 OK" 2026-07-23:04:46:50,835 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa/revision/main "HTTP/1.1 200 OK" 2026-07-23:04:46:50,892 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa/tree/main?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:46:51,335 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq "HTTP/1.1 200 OK" 2026-07-23:04:46:51,543 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq "HTTP/1.1 200 OK" 2026-07-23:04:46:51,598 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq/revision/2c94ad3e1aafab77146f384e23536f97a4849815 "HTTP/1.1 200 OK" 2026-07-23:04:46:51,668 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq/tree/2c94ad3e1aafab77146f384e23536f97a4849815?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:46:51,904 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq/tree/2c94ad3e1aafab77146f384e23536f97a4849815/data?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:46:51,966 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq/revision/2c94ad3e1aafab77146f384e23536f97a4849815 "HTTP/1.1 200 OK" 2026-07-23:04:46:52,398 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande "HTTP/1.1 200 OK" 2026-07-23:04:46:52,617 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande "HTTP/1.1 200 OK" 2026-07-23:04:46:52,676 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/revision/01e74176c63542e6b0bcb004dcdea22d94fb67b5 "HTTP/1.1 200 OK" 2026-07-23:04:46:52,734 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/tree/01e74176c63542e6b0bcb004dcdea22d94fb67b5?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:46:52,967 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/tree/01e74176c63542e6b0bcb004dcdea22d94fb67b5/winogrande_debiased?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:46:53,029 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/revision/01e74176c63542e6b0bcb004dcdea22d94fb67b5 "HTTP/1.1 200 OK" 2026-07-23:04:46:53,111 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/tree/01e74176c63542e6b0bcb004dcdea22d94fb67b5/winogrande_xl?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:46:53,792 WARNING [evaluator.py:270] Overwriting default num_fewshot of winogrande from None to 0 2026-07-23:04:46:53,792 WARNING [evaluator.py:270] Overwriting default num_fewshot of sciq from None to 0 2026-07-23:04:46:53,792 WARNING [evaluator.py:270] Overwriting default num_fewshot of piqa from None to 0 2026-07-23:04:46:53,792 WARNING [evaluator.py:270] Overwriting default num_fewshot of openbookqa from None to 0 2026-07-23:04:46:53,792 WARNING [evaluator.py:270] Overwriting default num_fewshot of lambada_openai from None to 0 2026-07-23:04:46:53,792 WARNING [evaluator.py:270] Overwriting default num_fewshot of hellaswag from None to 0 2026-07-23:04:46:53,793 WARNING [evaluator.py:270] Overwriting default num_fewshot of copa from None to 0 2026-07-23:04:46:53,793 WARNING [evaluator.py:270] Overwriting default num_fewshot of boolq from None to 0 2026-07-23:04:46:53,793 WARNING [evaluator.py:270] Overwriting default num_fewshot of arc_easy from None to 0 2026-07-23:04:46:53,793 WARNING [evaluator.py:270] Overwriting default num_fewshot of arc_challenge from None to 0 2026-07-23:04:46:53,831 INFO [task.py:415] Building contexts for winogrande on rank 0... 0%| | 0/1267 [00:00 /workspace/FLAME-MoE/results/ablations/flame1e18_downstream.csv ==================================================================== [run] flame512_dense mode=dense TEMPORAL=0 H=512 heads=32 ffn=2826 experts=64 topk=6 moe_ffn=352 shared=704 [run] tasks=arc_challenge,arc_easy,hellaswag,openbookqa,piqa,winogrande,boolq,copa,lambada_openai,sciq out=/workspace/FLAME-MoE/results/phase0/runs/flame512_dense/lmeval_1e18_0shot ==================================================================== /usr/local/lib/python3.11/dist-packages/torch/cuda/__init__.py:58: FutureWarning: The pynvml package is deprecated. Please install nvidia-ml-py instead. If you did not install pynvml directly, please report this to the maintainers of the package that installed pynvml for you. import pynvml # type: ignore[import] [run_lmeval] stubbed VLM backends (hf_vlms, vllm_vlms) /usr/local/lib/python3.11/dist-packages/torch/cuda/__init__.py:58: FutureWarning: The pynvml package is deprecated. Please install nvidia-ml-py instead. If you did not install pynvml directly, please report this to the maintainers of the package that installed pynvml for you. import pynvml # type: ignore[import] /usr/local/lib/python3.11/dist-packages/requests/__init__.py:113: RequestsDependencyWarning: urllib3 (2.2.3) or chardet (6.0.0.post1)/charset_normalizer (3.3.2) doesn't match a supported version! warnings.warn( /usr/local/lib/python3.11/dist-packages/torch/cuda/__init__.py:58: FutureWarning: The pynvml package is deprecated. Please install nvidia-ml-py instead. If you did not install pynvml directly, please report this to the maintainers of the package that installed pynvml for you. import pynvml # type: ignore[import] /workspace/FLAME-MoE/.venv/lib/python3.11/site-packages/transformer_engine/pytorch/attention.py:108: UserWarning: To use flash-attn v3, please use the following commands to install: (1) pip install "git+https://github.com/Dao-AILab/flash-attention.git#egg=flashattn-hopper&subdirectory=hopper" (2) python_path=`python -c "import site; print(site.getsitepackages()[0])"` (3) mkdir -p $python_path/flashattn_hopper (4) wget -P $python_path/flashattn_hopper https://raw.githubusercontent.com/Dao-AILab/flash-attention/main/hopper/flash_attn_interface.py warnings.warn( /workspace/FLAME-MoE/Megatron-LM/megatron/core/models/gpt/gpt_layer_specs.py:54: UserWarning: Apex is not installed. Falling back to Torch Norm warnings.warn('Apex is not installed. Falling back to Torch Norm') /workspace/FLAME-MoE/Megatron-LM/megatron/core/models/retro/encoder_spec.py:47: UserWarning: Apex is not installed. Falling back to Torch Norm warnings.warn(f'Apex is not installed. Falling back to Torch Norm') /workspace/FLAME-MoE/Megatron-LM/megatron/core/models/retro/decoder_spec.py:39: UserWarning: Apex is not installed. Falling back to Torch Norm warnings.warn(f'Apex is not installed. Falling back to Torch Norm') 2026-07-23:04:56:16,525 INFO [__main__.py:279] Verbosity set to INFO 2026-07-23:04:56:48,810 INFO [__init__.py:459] The tag 'arc_ca' is already registered as a group, this tag will not be registered. This may affect tasks you want to call. 2026-07-23:04:56:48,826 INFO [__init__.py:459] The tag 'arc_ca' is already registered as a group, this tag will not be registered. This may affect tasks you want to call. 2026-07-23:04:57:04,926 INFO [__main__.py:376] Selected Tasks: ['arc_challenge', 'arc_easy', 'boolq', 'copa', 'hellaswag', 'lambada_openai', 'openbookqa', 'piqa', 'sciq', 'winogrande'] 2026-07-23:04:57:04,946 INFO [evaluator.py:164] Setting random seed to 42 | Setting numpy seed to 42 | Setting torch manual seed to 42 | Setting fewshot manual seed to 42 2026-07-23:04:57:04,946 INFO [evaluator.py:201] Initializing megatron_lm model, with arguments: {} using world size: 1, data-parallel size: 1, context-parallel size: 1, hierarchical context-parallel sizes: Nonetensor-model-parallel size: 1, encoder-tensor-model-parallel size: 0, pipeline-model-parallel size: 1, encoder-pipeline-model-parallel size: 0 WARNING: overriding default arguments for exit_on_missing_checkpoint:True with exit_on_missing_checkpoint:False setting global batch size to 32 Number of virtual stages per pipeline stage: None accumulate and all-reduce gradients in fp32 for bfloat16 data type. using torch.bfloat16 for parameters ... ------------------------ arguments ------------------------ account_for_embedding_in_pipeline_split ......... False account_for_loss_in_pipeline_split .............. False accumulate_allreduce_grads_in_fp32 .............. True adam_beta1 ...................................... 0.9 adam_beta2 ...................................... 0.999 adam_eps ........................................ 1e-08 add_bias_linear ................................. False add_position_embedding .......................... True add_qkv_bias .................................... False adlr_autoresume ................................. False adlr_autoresume_interval ........................ 1000 align_grad_reduce ............................... True align_param_gather .............................. False app_tag_run_name ................................ None app_tag_run_version ............................. 0.0.0 apply_layernorm_1p .............................. False apply_query_key_layer_scaling ................... False apply_residual_connection_post_layernorm ........ False apply_rope_fusion ............................... True async_save ...................................... None async_tensor_model_parallel_allreduce ........... True attention_backend ............................... AttnBackend.auto attention_dropout ............................... 0.0 attention_softmax_in_fp32 ....................... False auto_detect_ckpt_format ......................... False barrier_with_L1_time ............................ True bert_binary_head ................................ True bert_embedder_type .............................. megatron bert_load ....................................... None bf16 ............................................ True bias_dropout_fusion ............................. True bias_gelu_fusion ................................ False bias_swiglu_fusion .............................. True biencoder_projection_dim ........................ 0 biencoder_shared_query_context_model ............ False block_data_path ................................. None calc_ft_timeouts ................................ False calculate_per_token_loss ........................ False check_for_large_grads ........................... False check_for_nan_in_loss_and_grad .................. True check_for_spiky_loss ............................ False check_weight_hash_across_dp_replicas_interval ... None ckpt_assume_constant_structure .................. False ckpt_convert_format ............................. None ckpt_convert_save ............................... None ckpt_convert_update_legacy_dist_opt_format ...... False ckpt_format ..................................... torch_dist ckpt_fully_parallel_load ........................ False ckpt_fully_parallel_save ........................ True ckpt_fully_parallel_save_deprecated ............. False ckpt_step ....................................... None classes_fraction ................................ 1.0 clip_grad ....................................... 1.0 clone_scatter_output_in_embedding ............... True config_logger_dir ............................... consumed_train_samples .......................... 0 consumed_valid_samples .......................... 0 context_parallel_size ........................... 1 cp_comm_type .................................... ['p2p'] create_attention_mask_in_dataloader ............. True cross_entropy_loss_fusion ....................... False cuda_graph_warmup_steps ......................... 3 data_args_path .................................. None data_cache_path ................................. None data_parallel_random_init ....................... False data_parallel_sharding_strategy ................. no_shard data_parallel_size .............................. 1 data_path ....................................... None data_per_class_fraction ......................... 1.0 data_sharding ................................... True dataloader_type ................................. single ddp_average_in_collective ....................... False ddp_bucket_size ................................. None ddp_num_buckets ................................. None ddp_pad_buckets_for_high_nccl_busbw ............. False decoder_first_pipeline_num_layers ............... None decoder_last_pipeline_num_layers ................ None decoder_num_layers .............................. None decoder_seq_length .............................. None decoupled_lr .................................... None decoupled_min_lr ................................ None decrease_batch_size_if_needed ................... False defer_embedding_wgrad_compute ................... False deprecated_use_mcore_models ..................... False deterministic_mode .............................. False dino_bottleneck_size ............................ 256 dino_freeze_last_layer .......................... 1 dino_head_hidden_size ........................... 2048 dino_local_crops_number ......................... 10 dino_local_img_size ............................. 96 dino_norm_last_layer ............................ False dino_teacher_temp ............................... 0.07 dino_warmup_teacher_temp ........................ 0.04 dino_warmup_teacher_temp_epochs ................. 30 disable_straggler_on_startup .................... False dist_ckpt_format_deprecated ..................... None dist_ckpt_strictness ............................ assume_ok_unexpected distribute_saved_activations .................... False distributed_backend ............................. nccl distributed_timeout_minutes ..................... 10 embedding_path .................................. None empty_unused_memory_level ....................... 0 enable_cuda_graph ............................... False enable_ft_package ............................... False enable_gloo_process_groups ...................... True enable_one_logger ............................... True encoder_num_layers .............................. 9 encoder_pipeline_model_parallel_size ............ 0 encoder_seq_length .............................. 2048 encoder_tensor_model_parallel_size .............. 0 end_weight_decay ................................ 0.01 eod_mask_loss ................................... False error_injection_rate ............................ 0 error_injection_type ............................ transient_error eval_interval ................................... 1000 eval_iters ...................................... 100 evidence_data_path .............................. None exit_duration_in_mins ........................... None exit_interval ................................... None exit_on_missing_checkpoint ...................... False exit_signal_handler ............................. False exp_avg_dtype ................................... torch.float32 exp_avg_sq_dtype ................................ torch.float32 expert_model_parallel_size ...................... 1 expert_tensor_parallel_size ..................... 1 ffn_hidden_size ................................. 2826 finetune ........................................ False flash_decode .................................... False fp16 ............................................ False fp16_lm_cross_entropy ........................... False fp32_residual_connection ........................ False fp8 ............................................. None fp8_amax_compute_algo ........................... most_recent fp8_amax_history_len ............................ 1 fp8_interval .................................... 1 fp8_margin ...................................... 0 fp8_param_gather ................................ False fp8_wgrad ....................................... True global_batch_size ............................... 32 grad_reduce_in_bf16 ............................. False gradient_accumulation_fusion .................... False gradient_reduce_div_fusion ...................... True group_query_attention ........................... False head_lr_mult .................................... 1.0 hidden_dropout .................................. 0.0 hidden_size ..................................... 512 hierarchical_context_parallel_sizes ............. None hybrid_attention_ratio .......................... 0.0 hybrid_mlp_ratio ................................ 0.0 hybrid_override_pattern ......................... None hysteresis ...................................... 2 ict_head_size ................................... None ict_load ........................................ None img_h ........................................... 224 img_w ........................................... 224 indexer_batch_size .............................. 128 indexer_log_interval ............................ 1000 inference_batch_times_seqlen_threshold .......... -1 inference_max_batch_size ........................ 8 inference_max_seq_length ........................ 2560 inference_rng_tracker ........................... False init_method_std ................................. 0.02 init_method_xavier_uniform ...................... False init_model_with_meta_device ..................... False initial_loss_scale .............................. 4294967296 iter_per_epoch .................................. 1250 iterations_to_skip .............................. [] keep_fp8_transpose_cache_when_using_custom_fsdp . False kv_channels ..................................... 16 kv_lora_rank .................................... 32 lazy_mpu_init ................................... None load ............................................ /workspace/FLAME-MoE/results/phase0/runs/flame512_dense/ckpt local_rank ...................................... 0 log_interval .................................... 100 log_loss_scale_to_tensorboard ................... True log_memory_to_tensorboard ....................... False log_num_zeros_in_grad ........................... False log_params_norm ................................. False log_progress .................................... False log_straggler ................................... False log_throughput .................................. False log_timers_to_tensorboard ....................... False log_validation_ppl_to_tensorboard ............... False log_world_size_to_tensorboard ................... False logging_level ................................... None loss_scale ...................................... None loss_scale_window ............................... 1000 lr .............................................. None lr_decay_iters .................................. None lr_decay_samples ................................ None lr_decay_style .................................. linear lr_warmup_fraction .............................. None lr_warmup_init .................................. 0.0 lr_warmup_iters ................................. 0 lr_warmup_samples ............................... 0 lr_wsd_decay_iters .............................. None lr_wsd_decay_samples ............................ None lr_wsd_decay_style .............................. exponential main_grads_dtype ................................ torch.float32 main_params_dtype ............................... torch.float32 make_vocab_size_divisible_by .................... 128 manual_gc ....................................... False manual_gc_eval .................................. True manual_gc_interval .............................. 0 mask_factor ..................................... 1.0 mask_prob ....................................... 0.15 mask_type ....................................... random masked_softmax_fusion ........................... True max_position_embeddings ......................... 2048 max_tokens_to_oom ............................... 10000000 memory_snapshot_path ............................ snapshot.pickle merge_file ...................................... None micro_batch_size ................................ 32 microbatch_group_size_per_vp_stage .............. None min_loss_scale .................................. 1.0 min_lr .......................................... 0.0 mmap_bin_files .................................. True mock_data ....................................... False moe_aux_loss_coeff .............................. 0.0 moe_enable_deepep ............................... False moe_expert_capacity_factor ...................... None moe_extended_tp ................................. False moe_ffn_hidden_size ............................. 2826 moe_grouped_gemm ................................ False moe_input_jitter_eps ............................ None moe_layer_freq .................................. 1 moe_layer_recompute ............................. False moe_pad_expert_input_to_capacity ................ False moe_per_layer_logging ........................... False moe_permute_fusion .............................. False moe_router_bias_update_rate ..................... 0.001 moe_router_dtype ................................ None moe_router_enable_expert_bias ................... False moe_router_group_topk ........................... None moe_router_load_balancing_type .................. aux_loss moe_router_num_groups ........................... None moe_router_pre_softmax .......................... False moe_router_score_function ....................... softmax moe_router_topk ................................. 2 moe_router_topk_scaling_factor .................. None moe_shared_expert_intermediate_size ............. None moe_shared_expert_overlap ....................... False moe_token_dispatcher_type ....................... allgather moe_token_drop_policy ........................... probs moe_use_legacy_grouped_gemm ..................... False moe_use_upcycling ............................... False moe_z_loss_coeff ................................ None mscale .......................................... 1.0 mscale_all_dim .................................. 1.0 multi_latent_attention .......................... False nccl_communicator_config_path ................... None no_load_optim ................................... True no_load_rng ..................................... True no_persist_layer_norm ........................... False no_save_optim ................................... None no_save_rng ..................................... None non_persistent_ckpt_type ........................ None non_persistent_global_ckpt_dir .................. None non_persistent_local_ckpt_algo .................. fully_parallel non_persistent_local_ckpt_dir ................... None non_persistent_save_interval .................... None norm_epsilon .................................... 1e-06 normalization ................................... RMSNorm num_attention_heads ............................. 32 num_channels .................................... 3 num_classes ..................................... 1000 num_dataset_builder_threads ..................... 1 num_distributed_optimizer_instances ............. 1 num_experts ..................................... None num_layers ...................................... 9 num_layers_per_virtual_pipeline_stage ........... None num_query_groups ................................ 1 num_virtual_stages_per_pipeline_rank ............ None num_workers ..................................... 2 one_logger_async ................................ False one_logger_project .............................. megatron-lm one_logger_run_name ............................. None onnx_safe ....................................... None openai_gelu ..................................... False optimizer ....................................... adam optimizer_cpu_offload ........................... False optimizer_offload_fraction ...................... 1.0 output_bert_embeddings .......................... False overlap_cpu_optimizer_d2h_h2d ................... False overlap_grad_reduce ............................. False overlap_p2p_comm ................................ False overlap_p2p_comm_warmup_flush ................... False overlap_param_gather ............................ False overlap_param_gather_with_optimizer_step ........ False override_opt_param_scheduler .................... False params_dtype .................................... torch.bfloat16 patch_dim ....................................... 16 per_split_data_args_path ........................ None perform_initialization .......................... True pin_cpu_grads ................................... True pin_cpu_params .................................. True pipeline_model_parallel_comm_backend ............ None pipeline_model_parallel_size .................... 1 pipeline_model_parallel_split_rank .............. None position_embedding_type ......................... rope pretrained_checkpoint ........................... None profile ......................................... False profile_ranks ................................... [0] profile_step_end ................................ 12 profile_step_start .............................. 10 q_lora_rank ..................................... None qk_head_dim ..................................... 128 qk_layernorm .................................... False qk_pos_emb_head_dim ............................. 64 query_in_block_prob ............................. 0.1 rampup_batch_size ............................... None rank ............................................ 0 recompute_granularity ........................... None recompute_method ................................ None recompute_num_layers ............................ None record_memory_history ........................... False relative_attention_max_distance ................. 128 relative_attention_num_buckets .................. 32 replication ..................................... False replication_factor .............................. 2 replication_jump ................................ None rerun_mode ...................................... disabled reset_attention_mask ............................ False reset_position_ids .............................. False result_rejected_tracker_filename ................ None retriever_report_topk_accuracies ................ [] retriever_score_scaling ......................... False retriever_seq_length ............................ 256 retro_add_retriever ............................. False retro_attention_gate ............................ 1 retro_cyclic_train_iters ........................ None retro_encoder_attention_dropout ................. 0.1 retro_encoder_hidden_dropout .................... 0.1 retro_encoder_layers ............................ 2 retro_num_neighbors ............................. 2 retro_num_retrieved_chunks ...................... 2 retro_project_dir ............................... None retro_verify_neighbor_count ..................... True rope_scaling_factor ............................. 8.0 rotary_base ..................................... 10000 rotary_interleaved .............................. False rotary_percent .................................. 1.0 rotary_scaling_factor ........................... 1.0 rotary_seq_len_interpolation_factor ............. None s3_cache_path ................................... None sample_rate ..................................... 1.0 save ............................................ None save_interval ................................... None scatter_gather_tensors_in_pipeline .............. True seed ............................................ 42 seq_length ...................................... 2048 sequence_parallel ............................... False sgd_momentum .................................... 0.9 short_seq_prob .................................. 0.1 skip_train ...................................... False skipped_train_samples ........................... 0 spec ............................................ None split ........................................... None squared_relu .................................... False start_weight_decay .............................. 0.01 straggler_ctrlr_port ............................ 65535 straggler_minmax_count .......................... 1 suggested_communication_unit_size ............... 400000000 swiglu .......................................... True swin_backbone_type .............................. tiny te_rng_tracker .................................. False tensor_model_parallel_size ...................... 1 tensorboard_dir ................................. None tensorboard_log_interval ........................ 1 tensorboard_queue_size .......................... 1000 test_data_path .................................. None test_mode ....................................... False tiktoken_num_special_tokens ..................... 1000 tiktoken_pattern ................................ None tiktoken_special_tokens ......................... None timing_log_level ................................ 0 timing_log_option ............................... minmax titles_data_path ................................ None tokenizer_model ................................. EleutherAI/pythia-12b tokenizer_type .................................. HuggingFaceTokenizer tp_comm_bootstrap_backend ....................... nccl tp_comm_bulk_dgrad .............................. True tp_comm_bulk_wgrad .............................. True tp_comm_overlap ................................. False tp_comm_overlap_ag .............................. True tp_comm_overlap_cfg ............................. None tp_comm_overlap_rs .............................. True tp_comm_overlap_rs_dgrad ........................ False tp_comm_split_ag ................................ True tp_comm_split_rs ................................ True train_data_path ................................. None train_iters ..................................... None train_samples ................................... None train_sync_interval ............................. None transformer_impl ................................ transformer_engine transformer_pipeline_model_parallel_size ........ 1 untie_embeddings_and_output_weights ............. True use_checkpoint_args ............................. False use_checkpoint_opt_param_scheduler .............. False use_cpu_initialization .......................... None use_custom_fsdp ................................. False use_dist_ckpt ................................... True use_dist_ckpt_deprecated ........................ False use_distributed_optimizer ....................... True use_flash_attn .................................. False use_legacy_models ............................... False use_mp_args_from_checkpoint_args ................ False use_one_sent_docs ............................... False use_persistent_ckpt_worker ...................... False use_precision_aware_optimizer ................... False use_pytorch_profiler ............................ False use_ring_exchange_p2p ........................... False use_rope_scaling ................................ False use_rotary_position_embeddings .................. False use_tokenizer_model_from_checkpoint_args ........ True use_torch_fsdp2 ................................. False use_torch_optimizer_for_cpu_offload ............. False use_tp_pp_dp_mapping ............................ False v_head_dim ...................................... 128 valid_data_path ................................. None variable_seq_lengths ............................ False virtual_pipeline_model_parallel_size ............ None vision_backbone_type ............................ vit vision_pretraining .............................. False vision_pretraining_type ......................... classify vocab_extra_ids ................................. 0 vocab_file ...................................... None vocab_size ...................................... None wandb_exp_name .................................. wandb_project ................................... wandb_save_dir .................................. weight_decay .................................... 0.01 weight_decay_incr_style ......................... constant wgrad_deferral_limit ............................ 0 world_size ...................................... 1 yaml_cfg ........................................ None -------------------- end of arguments --------------------- 2026-07-23:04:57:04,959 INFO [num_microbatches_calculator.py:228] setting number of microbatches to constant 1 > building HuggingFaceTokenizer tokenizer ... 2026-07-23:04:57:05,148 INFO [_client.py:1038] HTTP Request: HEAD https://huggingface.co/EleutherAI/pythia-12b/resolve/main/config.json "HTTP/1.1 307 Temporary Redirect" 2026-07-23:04:57:05,159 INFO [_client.py:1038] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/EleutherAI/pythia-12b/bb1e3e710cdf6b524461d543cfb5ba773f0a81b6/config.json "HTTP/1.1 200 OK" 2026-07-23:04:57:05,285 INFO [_client.py:1038] HTTP Request: HEAD https://huggingface.co/EleutherAI/pythia-12b/resolve/main/tokenizer_config.json "HTTP/1.1 307 Temporary Redirect" 2026-07-23:04:57:05,298 INFO [_client.py:1038] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/EleutherAI/pythia-12b/bb1e3e710cdf6b524461d543cfb5ba773f0a81b6/tokenizer_config.json "HTTP/1.1 200 OK" 2026-07-23:04:57:05,390 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/models/EleutherAI/pythia-12b/tree/main/additional_chat_templates?recursive=false&expand=false "HTTP/1.1 404 Not Found" 2026-07-23:04:57:05,455 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/models/EleutherAI/pythia-12b/tree/main?recursive=true&expand=false "HTTP/1.1 200 OK" > padded vocab (size: 50277) with 27 dummy tokens (new size: 50304) WARNING: one_logger package is required to enable e2e metrics tracking. please go to https://confluence.nvidia.com/display/MLWFO/Package+Repositories for details to install it 2026-07-23:04:57:05,664 WARNING [rerun_state_machine.py:239] RerunStateMachine initialized in mode RerunMode.DISABLED > initializing torch distributed ... [W723 04:57:05.182218380 CUDAAllocatorConfig.h:28] Warning: expandable_segments not supported on this platform (function operator()) > initialized tensor model parallel with size 1 > initialized pipeline model parallel with size 1 > setting random seeds to 42 ... > compiling dataset index builder ... make: Entering directory '/workspace/FLAME-MoE/Megatron-LM/megatron/core/datasets' make: Nothing to be done for 'default'. make: Leaving directory '/workspace/FLAME-MoE/Megatron-LM/megatron/core/datasets' >>> done with dataset index builder. Compilation time: 0.766 seconds > compiling and loading fused kernels ... >>> done with compiling and loading fused kernels. Compilation time: 0.381 seconds building GPT model ... /workspace/FLAME-MoE/Megatron-LM/megatron/core/models/gpt/gpt_layer_specs.py:81: UserWarning: The fp8 argument in "get_gpt_layer_with_transformer_engine_spec" has been deprecated and will be removed soon. Please update your code accordingly. warnings.warn( > number of parameters on (tensor, pipeline) model parallel rank (0, 0): 100024832 /workspace/FLAME-MoE/Megatron-LM/megatron/core/transformer/transformer_layer.py:339: UserWarning: TransformerLayer._get_layer_offset is deprecated.Please use get_transformer_layer_offset instead. warnings.warn( loading distributed checkpoint from /workspace/FLAME-MoE/results/phase0/runs/flame512_dense/ckpt at iteration 802 /workspace/FLAME-MoE/Megatron-LM/megatron/core/dist_checkpointing/strategies/torch.py:847: FutureWarning: `load_state_dict` is deprecated and will be removed in future versions. Please use `load` instead. checkpoint.load_state_dict( /usr/local/lib/python3.11/dist-packages/torch/distributed/checkpoint/planner_helpers.py:306: FutureWarning: Please use DTensor instead and we are deprecating ShardedTensor. device = getattr(value, "device", None) /workspace/FLAME-MoE/.venv/lib/python3.11/site-packages/transformer_engine/pytorch/module/base.py:566: FutureWarning: You are using `torch.load` with `weights_only=False` (the current default value), which uses the default pickle module implicitly. It is possible to construct malicious pickle data which will execute arbitrary code during unpickling (See https://github.com/pytorch/pytorch/blob/main/SECURITY.md#untrusted-models for more details). In a future release, the default value for `weights_only` will be flipped to `True`. This limits the functions that could be executed during unpickling. Arbitrary objects will no longer be allowed to be loaded via this mode unless they are explicitly allowlisted by the user via `torch.serialization.add_safe_globals`. We recommend you start setting `weights_only=True` for any use case where you don't have full control of the loaded file. Please open an issue on GitHub for any issues related to this experimental feature. state = torch.load(state, map_location="cuda") checkpoint version 3.0 2026-07-23:04:57:09,043 WARNING [rerun_state_machine.py:811] RerunStateMachine disabled via CLI, ignoring machine state saved in checkpoint successfully loaded checkpoint from /workspace/FLAME-MoE/results/phase0/runs/flame512_dense/ckpt [ t 1/1, p 1/1 ] at iteration 802 2026-07-23:04:57:09,121 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc "HTTP/1.1 200 OK" 2026-07-23:04:57:09,377 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc "HTTP/1.1 200 OK" 2026-07-23:04:57:09,446 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc/revision/210d026faf9955653af8916fad021475a3f00453 "HTTP/1.1 200 OK" 2026-07-23:04:57:09,529 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc/tree/210d026faf9955653af8916fad021475a3f00453?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:57:09,798 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc/tree/210d026faf9955653af8916fad021475a3f00453/ARC-Challenge?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:57:10,376 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc "HTTP/1.1 200 OK" 2026-07-23:04:57:10,623 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc "HTTP/1.1 200 OK" 2026-07-23:04:57:10,859 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc/tree/210d026faf9955653af8916fad021475a3f00453/ARC-Easy?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:57:11,009 WARNING [task.py:799] [Task: boolq] metric acc is defined, but aggregation is not. using default aggregation=mean 2026-07-23:04:57:11,009 WARNING [task.py:811] [Task: boolq] metric acc is defined, but higher_is_better is not. using default higher_is_better=True 2026-07-23:04:57:11,073 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue "HTTP/1.1 200 OK" 2026-07-23:04:57:11,313 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue "HTTP/1.1 200 OK" 2026-07-23:04:57:11,370 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/revision/3de24cf8022e94f4ee4b9d55a6f539891524d646 "HTTP/1.1 200 OK" 2026-07-23:04:57:11,434 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/tree/3de24cf8022e94f4ee4b9d55a6f539891524d646?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:57:11,691 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/tree/3de24cf8022e94f4ee4b9d55a6f539891524d646/axb?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:57:11,789 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/revision/3de24cf8022e94f4ee4b9d55a6f539891524d646 "HTTP/1.1 200 OK" 2026-07-23:04:57:11,865 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/tree/3de24cf8022e94f4ee4b9d55a6f539891524d646/boolq?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:57:12,134 WARNING [task.py:799] [Task: copa] metric acc is defined, but aggregation is not. using default aggregation=mean 2026-07-23:04:57:12,135 WARNING [task.py:811] [Task: copa] metric acc is defined, but higher_is_better is not. using default higher_is_better=True 2026-07-23:04:57:12,246 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue "HTTP/1.1 200 OK" 2026-07-23:04:57:12,453 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue "HTTP/1.1 200 OK" 2026-07-23:04:57:12,755 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/tree/3de24cf8022e94f4ee4b9d55a6f539891524d646/copa?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:57:12,899 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag "HTTP/1.1 200 OK" 2026-07-23:04:57:13,141 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag "HTTP/1.1 200 OK" 2026-07-23:04:57:13,204 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag/revision/218ec52e09a7e7462a5400043bb9a69a41d06b76 "HTTP/1.1 200 OK" 2026-07-23:04:57:13,268 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag/tree/218ec52e09a7e7462a5400043bb9a69a41d06b76?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:57:13,605 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag/tree/218ec52e09a7e7462a5400043bb9a69a41d06b76/data?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:57:13,668 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag/revision/218ec52e09a7e7462a5400043bb9a69a41d06b76 "HTTP/1.1 200 OK" 2026-07-23:04:57:15,753 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai "HTTP/1.1 200 OK" 2026-07-23:04:57:15,975 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai "HTTP/1.1 200 OK" 2026-07-23:04:57:16,046 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/revision/900124bf3b8235c6daf21033af9948b3f07346c4 "HTTP/1.1 200 OK" 2026-07-23:04:57:16,110 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/tree/900124bf3b8235c6daf21033af9948b3f07346c4?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:57:16,336 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/tree/900124bf3b8235c6daf21033af9948b3f07346c4/default%2Ftest?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:57:16,400 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/tree/900124bf3b8235c6daf21033af9948b3f07346c4/default?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:57:16,467 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/revision/900124bf3b8235c6daf21033af9948b3f07346c4 "HTTP/1.1 200 OK" 2026-07-23:04:57:16,487 WARNING [task.py:325] [Task: lambada_openai] has_training_docs and has_validation_docs are False, using test_docs as fewshot_docs but this is not recommended. 2026-07-23:04:57:16,487 WARNING [task.py:325] [Task: lambada_openai] has_training_docs and has_validation_docs are False, using test_docs as fewshot_docs but this is not recommended. 2026-07-23:04:57:16,660 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa "HTTP/1.1 200 OK" 2026-07-23:04:57:16,881 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa "HTTP/1.1 200 OK" 2026-07-23:04:57:16,947 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/revision/388097ea7776314e93a529163e0fea805b8a6454 "HTTP/1.1 200 OK" 2026-07-23:04:57:17,002 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/tree/388097ea7776314e93a529163e0fea805b8a6454?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:57:17,219 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/tree/388097ea7776314e93a529163e0fea805b8a6454/additional?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:57:17,280 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/revision/388097ea7776314e93a529163e0fea805b8a6454 "HTTP/1.1 200 OK" 2026-07-23:04:57:17,368 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/tree/388097ea7776314e93a529163e0fea805b8a6454/main?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:57:17,954 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa "HTTP/1.1 200 OK" 2026-07-23:04:57:18,012 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa "HTTP/1.1 200 OK" 2026-07-23:04:57:18,076 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa/tree/main?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:57:18,137 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/datasets/ybisk/piqa/resolve/main/piqa.py "HTTP/1.1 307 Temporary Redirect" 2026-07-23:04:57:18,151 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/resolve-cache/datasets/ybisk/piqa/2e8ac2dffd59bac8c3c6714948f4c551a0848bb0/piqa.py?%2Fdatasets%2Fybisk%2Fpiqa%2Fresolve%2Fmain%2Fpiqa.py=&etag=%220f0d96172f9bf7aa10069374b3236ef083cfeb8e%22 "HTTP/1.1 200 OK" 2026-07-23:04:57:18,721 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa/revision/main "HTTP/1.1 200 OK" 2026-07-23:04:57:18,785 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa/tree/main?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:57:19,232 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq "HTTP/1.1 200 OK" 2026-07-23:04:57:19,443 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq "HTTP/1.1 200 OK" 2026-07-23:04:57:19,516 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq/revision/2c94ad3e1aafab77146f384e23536f97a4849815 "HTTP/1.1 200 OK" 2026-07-23:04:57:19,575 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq/tree/2c94ad3e1aafab77146f384e23536f97a4849815?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:57:19,760 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq/tree/2c94ad3e1aafab77146f384e23536f97a4849815/data?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:57:19,818 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq/revision/2c94ad3e1aafab77146f384e23536f97a4849815 "HTTP/1.1 200 OK" 2026-07-23:04:57:20,248 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande "HTTP/1.1 200 OK" 2026-07-23:04:57:20,479 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande "HTTP/1.1 200 OK" 2026-07-23:04:57:20,542 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/revision/01e74176c63542e6b0bcb004dcdea22d94fb67b5 "HTTP/1.1 200 OK" 2026-07-23:04:57:20,604 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/tree/01e74176c63542e6b0bcb004dcdea22d94fb67b5?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:57:20,935 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/tree/01e74176c63542e6b0bcb004dcdea22d94fb67b5/winogrande_debiased?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:57:20,992 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/revision/01e74176c63542e6b0bcb004dcdea22d94fb67b5 "HTTP/1.1 200 OK" 2026-07-23:04:57:21,075 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/tree/01e74176c63542e6b0bcb004dcdea22d94fb67b5/winogrande_xl?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:04:57:21,727 WARNING [evaluator.py:270] Overwriting default num_fewshot of winogrande from None to 0 2026-07-23:04:57:21,728 WARNING [evaluator.py:270] Overwriting default num_fewshot of sciq from None to 0 2026-07-23:04:57:21,728 WARNING [evaluator.py:270] Overwriting default num_fewshot of piqa from None to 0 2026-07-23:04:57:21,728 WARNING [evaluator.py:270] Overwriting default num_fewshot of openbookqa from None to 0 2026-07-23:04:57:21,728 WARNING [evaluator.py:270] Overwriting default num_fewshot of lambada_openai from None to 0 2026-07-23:04:57:21,728 WARNING [evaluator.py:270] Overwriting default num_fewshot of hellaswag from None to 0 2026-07-23:04:57:21,728 WARNING [evaluator.py:270] Overwriting default num_fewshot of copa from None to 0 2026-07-23:04:57:21,728 WARNING [evaluator.py:270] Overwriting default num_fewshot of boolq from None to 0 2026-07-23:04:57:21,728 WARNING [evaluator.py:270] Overwriting default num_fewshot of arc_easy from None to 0 2026-07-23:04:57:21,728 WARNING [evaluator.py:270] Overwriting default num_fewshot of arc_challenge from None to 0 2026-07-23:04:57:21,757 INFO [task.py:415] Building contexts for winogrande on rank 0... 0%| | 0/1267 [00:00 /workspace/FLAME-MoE/results/ablations/flame1e18_downstream.csv ==================================================================== [run] flame512_g1_moe mode=moe TEMPORAL=0 H=512 heads=32 ffn=2736 experts=64 topk=6 moe_ffn=352 shared=704 [run] tasks=arc_challenge,arc_easy,hellaswag,openbookqa,piqa,winogrande,boolq,copa,lambada_openai,sciq out=/workspace/FLAME-MoE/results/phase0/runs/flame512_g1_moe/lmeval_1e18_0shot ==================================================================== /usr/local/lib/python3.11/dist-packages/torch/cuda/__init__.py:58: FutureWarning: The pynvml package is deprecated. Please install nvidia-ml-py instead. If you did not install pynvml directly, please report this to the maintainers of the package that installed pynvml for you. import pynvml # type: ignore[import] [run_lmeval] stubbed VLM backends (hf_vlms, vllm_vlms) /usr/local/lib/python3.11/dist-packages/torch/cuda/__init__.py:58: FutureWarning: The pynvml package is deprecated. Please install nvidia-ml-py instead. If you did not install pynvml directly, please report this to the maintainers of the package that installed pynvml for you. import pynvml # type: ignore[import] /usr/local/lib/python3.11/dist-packages/requests/__init__.py:113: RequestsDependencyWarning: urllib3 (2.2.3) or chardet (6.0.0.post1)/charset_normalizer (3.3.2) doesn't match a supported version! warnings.warn( /usr/local/lib/python3.11/dist-packages/torch/cuda/__init__.py:58: FutureWarning: The pynvml package is deprecated. Please install nvidia-ml-py instead. If you did not install pynvml directly, please report this to the maintainers of the package that installed pynvml for you. import pynvml # type: ignore[import] /workspace/FLAME-MoE/.venv/lib/python3.11/site-packages/transformer_engine/pytorch/attention.py:108: UserWarning: To use flash-attn v3, please use the following commands to install: (1) pip install "git+https://github.com/Dao-AILab/flash-attention.git#egg=flashattn-hopper&subdirectory=hopper" (2) python_path=`python -c "import site; print(site.getsitepackages()[0])"` (3) mkdir -p $python_path/flashattn_hopper (4) wget -P $python_path/flashattn_hopper https://raw.githubusercontent.com/Dao-AILab/flash-attention/main/hopper/flash_attn_interface.py warnings.warn( /workspace/FLAME-MoE/Megatron-LM/megatron/core/models/gpt/gpt_layer_specs.py:54: UserWarning: Apex is not installed. Falling back to Torch Norm warnings.warn('Apex is not installed. Falling back to Torch Norm') /workspace/FLAME-MoE/Megatron-LM/megatron/core/models/retro/encoder_spec.py:47: UserWarning: Apex is not installed. Falling back to Torch Norm warnings.warn(f'Apex is not installed. Falling back to Torch Norm') /workspace/FLAME-MoE/Megatron-LM/megatron/core/models/retro/decoder_spec.py:39: UserWarning: Apex is not installed. Falling back to Torch Norm warnings.warn(f'Apex is not installed. Falling back to Torch Norm') 2026-07-23:05:03:08,320 INFO [__main__.py:279] Verbosity set to INFO 2026-07-23:05:03:39,865 INFO [__init__.py:459] The tag 'arc_ca' is already registered as a group, this tag will not be registered. This may affect tasks you want to call. 2026-07-23:05:03:39,878 INFO [__init__.py:459] The tag 'arc_ca' is already registered as a group, this tag will not be registered. This may affect tasks you want to call. 2026-07-23:05:03:54,292 INFO [__main__.py:376] Selected Tasks: ['arc_challenge', 'arc_easy', 'boolq', 'copa', 'hellaswag', 'lambada_openai', 'openbookqa', 'piqa', 'sciq', 'winogrande'] 2026-07-23:05:03:54,314 INFO [evaluator.py:164] Setting random seed to 42 | Setting numpy seed to 42 | Setting torch manual seed to 42 | Setting fewshot manual seed to 42 2026-07-23:05:03:54,315 INFO [evaluator.py:201] Initializing megatron_lm model, with arguments: {} using world size: 1, data-parallel size: 1, context-parallel size: 1, hierarchical context-parallel sizes: Nonetensor-model-parallel size: 1, encoder-tensor-model-parallel size: 0, pipeline-model-parallel size: 1, encoder-pipeline-model-parallel size: 0 WARNING: overriding default arguments for exit_on_missing_checkpoint:True with exit_on_missing_checkpoint:False setting global batch size to 32 Number of virtual stages per pipeline stage: None accumulate and all-reduce gradients in fp32 for bfloat16 data type. using torch.bfloat16 for parameters ... ------------------------ arguments ------------------------ account_for_embedding_in_pipeline_split ......... False account_for_loss_in_pipeline_split .............. False accumulate_allreduce_grads_in_fp32 .............. True adam_beta1 ...................................... 0.9 adam_beta2 ...................................... 0.999 adam_eps ........................................ 1e-08 add_bias_linear ................................. False add_position_embedding .......................... True add_qkv_bias .................................... False adlr_autoresume ................................. False adlr_autoresume_interval ........................ 1000 align_grad_reduce ............................... True align_param_gather .............................. False app_tag_run_name ................................ None app_tag_run_version ............................. 0.0.0 apply_layernorm_1p .............................. False apply_query_key_layer_scaling ................... False apply_residual_connection_post_layernorm ........ False apply_rope_fusion ............................... True async_save ...................................... None async_tensor_model_parallel_allreduce ........... True attention_backend ............................... AttnBackend.auto attention_dropout ............................... 0.0 attention_softmax_in_fp32 ....................... False auto_detect_ckpt_format ......................... False barrier_with_L1_time ............................ True bert_binary_head ................................ True bert_embedder_type .............................. megatron bert_load ....................................... None bf16 ............................................ True bias_dropout_fusion ............................. True bias_gelu_fusion ................................ False bias_swiglu_fusion .............................. True biencoder_projection_dim ........................ 0 biencoder_shared_query_context_model ............ False block_data_path ................................. None calc_ft_timeouts ................................ False calculate_per_token_loss ........................ False check_for_large_grads ........................... False check_for_nan_in_loss_and_grad .................. True check_for_spiky_loss ............................ False check_weight_hash_across_dp_replicas_interval ... None ckpt_assume_constant_structure .................. False ckpt_convert_format ............................. None ckpt_convert_save ............................... None ckpt_convert_update_legacy_dist_opt_format ...... False ckpt_format ..................................... torch_dist ckpt_fully_parallel_load ........................ False ckpt_fully_parallel_save ........................ True ckpt_fully_parallel_save_deprecated ............. False ckpt_step ....................................... None classes_fraction ................................ 1.0 clip_grad ....................................... 1.0 clone_scatter_output_in_embedding ............... True config_logger_dir ............................... consumed_train_samples .......................... 0 consumed_valid_samples .......................... 0 context_parallel_size ........................... 1 cp_comm_type .................................... ['p2p'] create_attention_mask_in_dataloader ............. True cross_entropy_loss_fusion ....................... False cuda_graph_warmup_steps ......................... 3 data_args_path .................................. None data_cache_path ................................. None data_parallel_random_init ....................... False data_parallel_sharding_strategy ................. no_shard data_parallel_size .............................. 1 data_path ....................................... None data_per_class_fraction ......................... 1.0 data_sharding ................................... True dataloader_type ................................. single ddp_average_in_collective ....................... False ddp_bucket_size ................................. None ddp_num_buckets ................................. None ddp_pad_buckets_for_high_nccl_busbw ............. False decoder_first_pipeline_num_layers ............... None decoder_last_pipeline_num_layers ................ None decoder_num_layers .............................. None decoder_seq_length .............................. None decoupled_lr .................................... None decoupled_min_lr ................................ None decrease_batch_size_if_needed ................... False defer_embedding_wgrad_compute ................... False deprecated_use_mcore_models ..................... False deterministic_mode .............................. False dino_bottleneck_size ............................ 256 dino_freeze_last_layer .......................... 1 dino_head_hidden_size ........................... 2048 dino_local_crops_number ......................... 10 dino_local_img_size ............................. 96 dino_norm_last_layer ............................ False dino_teacher_temp ............................... 0.07 dino_warmup_teacher_temp ........................ 0.04 dino_warmup_teacher_temp_epochs ................. 30 disable_straggler_on_startup .................... False dist_ckpt_format_deprecated ..................... None dist_ckpt_strictness ............................ assume_ok_unexpected distribute_saved_activations .................... False distributed_backend ............................. nccl distributed_timeout_minutes ..................... 10 embedding_path .................................. None empty_unused_memory_level ....................... 0 enable_cuda_graph ............................... False enable_ft_package ............................... False enable_gloo_process_groups ...................... True enable_one_logger ............................... True encoder_num_layers .............................. 9 encoder_pipeline_model_parallel_size ............ 0 encoder_seq_length .............................. 2048 encoder_tensor_model_parallel_size .............. 0 end_weight_decay ................................ 0.01 eod_mask_loss ................................... False error_injection_rate ............................ 0 error_injection_type ............................ transient_error eval_interval ................................... 1000 eval_iters ...................................... 100 evidence_data_path .............................. None exit_duration_in_mins ........................... None exit_interval ................................... None exit_on_missing_checkpoint ...................... False exit_signal_handler ............................. False exp_avg_dtype ................................... torch.float32 exp_avg_sq_dtype ................................ torch.float32 expert_model_parallel_size ...................... 1 expert_tensor_parallel_size ..................... 1 ffn_hidden_size ................................. 2736 finetune ........................................ False flash_decode .................................... False fp16 ............................................ False fp16_lm_cross_entropy ........................... False fp32_residual_connection ........................ False fp8 ............................................. None fp8_amax_compute_algo ........................... most_recent fp8_amax_history_len ............................ 1 fp8_interval .................................... 1 fp8_margin ...................................... 0 fp8_param_gather ................................ False fp8_wgrad ....................................... True global_batch_size ............................... 32 grad_reduce_in_bf16 ............................. False gradient_accumulation_fusion .................... False gradient_reduce_div_fusion ...................... True group_query_attention ........................... False head_lr_mult .................................... 1.0 hidden_dropout .................................. 0.0 hidden_size ..................................... 512 hierarchical_context_parallel_sizes ............. None hybrid_attention_ratio .......................... 0.0 hybrid_mlp_ratio ................................ 0.0 hybrid_override_pattern ......................... None hysteresis ...................................... 2 ict_head_size ................................... None ict_load ........................................ None img_h ........................................... 224 img_w ........................................... 224 indexer_batch_size .............................. 128 indexer_log_interval ............................ 1000 inference_batch_times_seqlen_threshold .......... -1 inference_max_batch_size ........................ 8 inference_max_seq_length ........................ 2560 inference_rng_tracker ........................... False init_method_std ................................. 0.02 init_method_xavier_uniform ...................... False init_model_with_meta_device ..................... False initial_loss_scale .............................. 4294967296 iter_per_epoch .................................. 1250 iterations_to_skip .............................. [] keep_fp8_transpose_cache_when_using_custom_fsdp . False kv_channels ..................................... 16 kv_lora_rank .................................... 32 lazy_mpu_init ................................... None load ............................................ /workspace/FLAME-MoE/results/phase0/runs/flame512_g1_moe/ckpt local_rank ...................................... 0 log_interval .................................... 100 log_loss_scale_to_tensorboard ................... True log_memory_to_tensorboard ....................... False log_num_zeros_in_grad ........................... False log_params_norm ................................. False log_progress .................................... False log_straggler ................................... False log_throughput .................................. False log_timers_to_tensorboard ....................... False log_validation_ppl_to_tensorboard ............... False log_world_size_to_tensorboard ................... False logging_level ................................... None loss_scale ...................................... None loss_scale_window ............................... 1000 lr .............................................. None lr_decay_iters .................................. None lr_decay_samples ................................ None lr_decay_style .................................. linear lr_warmup_fraction .............................. None lr_warmup_init .................................. 0.0 lr_warmup_iters ................................. 0 lr_warmup_samples ............................... 0 lr_wsd_decay_iters .............................. None lr_wsd_decay_samples ............................ None lr_wsd_decay_style .............................. exponential main_grads_dtype ................................ torch.float32 main_params_dtype ............................... torch.float32 make_vocab_size_divisible_by .................... 128 manual_gc ....................................... False manual_gc_eval .................................. True manual_gc_interval .............................. 0 mask_factor ..................................... 1.0 mask_prob ....................................... 0.15 mask_type ....................................... random masked_softmax_fusion ........................... True max_position_embeddings ......................... 2048 max_tokens_to_oom ............................... 10000000 memory_snapshot_path ............................ snapshot.pickle merge_file ...................................... None micro_batch_size ................................ 32 microbatch_group_size_per_vp_stage .............. None min_loss_scale .................................. 1.0 min_lr .......................................... 0.0 mmap_bin_files .................................. True mock_data ....................................... False moe_aux_loss_coeff .............................. 0.01 moe_enable_deepep ............................... False moe_expert_capacity_factor ...................... None moe_extended_tp ................................. False moe_ffn_hidden_size ............................. 352 moe_grouped_gemm ................................ True moe_input_jitter_eps ............................ None moe_layer_freq .................................. [0, 1, 1, 1, 1, 1, 1, 1, 1] moe_layer_recompute ............................. False moe_pad_expert_input_to_capacity ................ False moe_per_layer_logging ........................... False moe_permute_fusion .............................. False moe_router_bias_update_rate ..................... 0.001 moe_router_dtype ................................ fp32 moe_router_enable_expert_bias ................... False moe_router_group_topk ........................... None moe_router_load_balancing_type .................. aux_loss moe_router_num_groups ........................... None moe_router_pre_softmax .......................... True moe_router_score_function ....................... softmax moe_router_topk ................................. 6 moe_router_topk_scaling_factor .................. None moe_shared_expert_intermediate_size ............. 704 moe_shared_expert_overlap ....................... False moe_token_dispatcher_type ....................... alltoall moe_token_drop_policy ........................... probs moe_use_legacy_grouped_gemm ..................... False moe_use_upcycling ............................... False moe_z_loss_coeff ................................ 0.001 mscale .......................................... 1.0 mscale_all_dim .................................. 1.0 multi_latent_attention .......................... False nccl_communicator_config_path ................... None no_load_optim ................................... True no_load_rng ..................................... True no_persist_layer_norm ........................... False no_save_optim ................................... None no_save_rng ..................................... None non_persistent_ckpt_type ........................ None non_persistent_global_ckpt_dir .................. None non_persistent_local_ckpt_algo .................. fully_parallel non_persistent_local_ckpt_dir ................... None non_persistent_save_interval .................... None norm_epsilon .................................... 1e-06 normalization ................................... RMSNorm num_attention_heads ............................. 32 num_channels .................................... 3 num_classes ..................................... 1000 num_dataset_builder_threads ..................... 1 num_distributed_optimizer_instances ............. 1 num_experts ..................................... 64 num_layers ...................................... 9 num_layers_per_virtual_pipeline_stage ........... None num_query_groups ................................ 1 num_virtual_stages_per_pipeline_rank ............ None num_workers ..................................... 2 one_logger_async ................................ False one_logger_project .............................. megatron-lm one_logger_run_name ............................. None onnx_safe ....................................... None openai_gelu ..................................... False optimizer ....................................... adam optimizer_cpu_offload ........................... False optimizer_offload_fraction ...................... 1.0 output_bert_embeddings .......................... False overlap_cpu_optimizer_d2h_h2d ................... False overlap_grad_reduce ............................. False overlap_p2p_comm ................................ False overlap_p2p_comm_warmup_flush ................... False overlap_param_gather ............................ False overlap_param_gather_with_optimizer_step ........ False override_opt_param_scheduler .................... False params_dtype .................................... torch.bfloat16 patch_dim ....................................... 16 per_split_data_args_path ........................ None perform_initialization .......................... True pin_cpu_grads ................................... True pin_cpu_params .................................. True pipeline_model_parallel_comm_backend ............ None pipeline_model_parallel_size .................... 1 pipeline_model_parallel_split_rank .............. None position_embedding_type ......................... rope pretrained_checkpoint ........................... None profile ......................................... False profile_ranks ................................... [0] profile_step_end ................................ 12 profile_step_start .............................. 10 q_lora_rank ..................................... None qk_head_dim ..................................... 128 qk_layernorm .................................... False qk_pos_emb_head_dim ............................. 64 query_in_block_prob ............................. 0.1 rampup_batch_size ............................... None rank ............................................ 0 recompute_granularity ........................... None recompute_method ................................ None recompute_num_layers ............................ None record_memory_history ........................... False relative_attention_max_distance ................. 128 relative_attention_num_buckets .................. 32 replication ..................................... False replication_factor .............................. 2 replication_jump ................................ None rerun_mode ...................................... disabled reset_attention_mask ............................ False reset_position_ids .............................. False result_rejected_tracker_filename ................ None retriever_report_topk_accuracies ................ [] retriever_score_scaling ......................... False retriever_seq_length ............................ 256 retro_add_retriever ............................. False retro_attention_gate ............................ 1 retro_cyclic_train_iters ........................ None retro_encoder_attention_dropout ................. 0.1 retro_encoder_hidden_dropout .................... 0.1 retro_encoder_layers ............................ 2 retro_num_neighbors ............................. 2 retro_num_retrieved_chunks ...................... 2 retro_project_dir ............................... None retro_verify_neighbor_count ..................... True rope_scaling_factor ............................. 8.0 rotary_base ..................................... 10000 rotary_interleaved .............................. False rotary_percent .................................. 1.0 rotary_scaling_factor ........................... 1.0 rotary_seq_len_interpolation_factor ............. None s3_cache_path ................................... None sample_rate ..................................... 1.0 save ............................................ None save_interval ................................... None scatter_gather_tensors_in_pipeline .............. True seed ............................................ 42 seq_length ...................................... 2048 sequence_parallel ............................... False sgd_momentum .................................... 0.9 short_seq_prob .................................. 0.1 skip_train ...................................... False skipped_train_samples ........................... 0 spec ............................................ None split ........................................... None squared_relu .................................... False start_weight_decay .............................. 0.01 straggler_ctrlr_port ............................ 65535 straggler_minmax_count .......................... 1 suggested_communication_unit_size ............... 400000000 swiglu .......................................... True swin_backbone_type .............................. tiny te_rng_tracker .................................. False tensor_model_parallel_size ...................... 1 tensorboard_dir ................................. None tensorboard_log_interval ........................ 1 tensorboard_queue_size .......................... 1000 test_data_path .................................. None test_mode ....................................... False tiktoken_num_special_tokens ..................... 1000 tiktoken_pattern ................................ None tiktoken_special_tokens ......................... None timing_log_level ................................ 0 timing_log_option ............................... minmax titles_data_path ................................ None tokenizer_model ................................. EleutherAI/pythia-12b tokenizer_type .................................. HuggingFaceTokenizer tp_comm_bootstrap_backend ....................... nccl tp_comm_bulk_dgrad .............................. True tp_comm_bulk_wgrad .............................. True tp_comm_overlap ................................. False tp_comm_overlap_ag .............................. True tp_comm_overlap_cfg ............................. None tp_comm_overlap_rs .............................. True tp_comm_overlap_rs_dgrad ........................ False tp_comm_split_ag ................................ True tp_comm_split_rs ................................ True train_data_path ................................. None train_iters ..................................... None train_samples ................................... None train_sync_interval ............................. None transformer_impl ................................ transformer_engine transformer_pipeline_model_parallel_size ........ 1 untie_embeddings_and_output_weights ............. True use_checkpoint_args ............................. False use_checkpoint_opt_param_scheduler .............. False use_cpu_initialization .......................... None use_custom_fsdp ................................. False use_dist_ckpt ................................... True use_dist_ckpt_deprecated ........................ False use_distributed_optimizer ....................... True use_flash_attn .................................. False use_legacy_models ............................... False use_mp_args_from_checkpoint_args ................ False use_one_sent_docs ............................... False use_persistent_ckpt_worker ...................... False use_precision_aware_optimizer ................... False use_pytorch_profiler ............................ False use_ring_exchange_p2p ........................... False use_rope_scaling ................................ False use_rotary_position_embeddings .................. False use_tokenizer_model_from_checkpoint_args ........ True use_torch_fsdp2 ................................. False use_torch_optimizer_for_cpu_offload ............. False use_tp_pp_dp_mapping ............................ False v_head_dim ...................................... 128 valid_data_path ................................. None variable_seq_lengths ............................ False virtual_pipeline_model_parallel_size ............ None vision_backbone_type ............................ vit vision_pretraining .............................. False vision_pretraining_type ......................... classify vocab_extra_ids ................................. 0 vocab_file ...................................... None vocab_size ...................................... None wandb_exp_name .................................. wandb_project ................................... wandb_save_dir .................................. weight_decay .................................... 0.01 weight_decay_incr_style ......................... constant wgrad_deferral_limit ............................ 0 world_size ...................................... 1 yaml_cfg ........................................ None -------------------- end of arguments --------------------- 2026-07-23:05:03:54,330 INFO [num_microbatches_calculator.py:228] setting number of microbatches to constant 1 > building HuggingFaceTokenizer tokenizer ... 2026-07-23:05:03:54,527 INFO [_client.py:1038] HTTP Request: HEAD https://huggingface.co/EleutherAI/pythia-12b/resolve/main/config.json "HTTP/1.1 307 Temporary Redirect" 2026-07-23:05:03:54,540 INFO [_client.py:1038] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/EleutherAI/pythia-12b/bb1e3e710cdf6b524461d543cfb5ba773f0a81b6/config.json "HTTP/1.1 200 OK" 2026-07-23:05:03:54,662 INFO [_client.py:1038] HTTP Request: HEAD https://huggingface.co/EleutherAI/pythia-12b/resolve/main/tokenizer_config.json "HTTP/1.1 307 Temporary Redirect" 2026-07-23:05:03:54,675 INFO [_client.py:1038] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/EleutherAI/pythia-12b/bb1e3e710cdf6b524461d543cfb5ba773f0a81b6/tokenizer_config.json "HTTP/1.1 200 OK" 2026-07-23:05:03:54,773 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/models/EleutherAI/pythia-12b/tree/main/additional_chat_templates?recursive=false&expand=false "HTTP/1.1 404 Not Found" 2026-07-23:05:03:54,842 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/models/EleutherAI/pythia-12b/tree/main?recursive=true&expand=false "HTTP/1.1 200 OK" > padded vocab (size: 50277) with 27 dummy tokens (new size: 50304) WARNING: one_logger package is required to enable e2e metrics tracking. please go to https://confluence.nvidia.com/display/MLWFO/Package+Repositories for details to install it 2026-07-23:05:03:55,055 WARNING [rerun_state_machine.py:239] RerunStateMachine initialized in mode RerunMode.DISABLED > initializing torch distributed ... [W723 05:03:55.572863404 CUDAAllocatorConfig.h:28] Warning: expandable_segments not supported on this platform (function operator()) > initialized tensor model parallel with size 1 > initialized pipeline model parallel with size 1 > setting random seeds to 42 ... > compiling dataset index builder ... make: Entering directory '/workspace/FLAME-MoE/Megatron-LM/megatron/core/datasets' make: Nothing to be done for 'default'. make: Leaving directory '/workspace/FLAME-MoE/Megatron-LM/megatron/core/datasets' >>> done with dataset index builder. Compilation time: 0.652 seconds > compiling and loading fused kernels ... >>> done with compiling and loading fused kernels. Compilation time: 0.348 seconds building GPT model ... > number of parameters on (tensor, pipeline) model parallel rank (0, 0): 350897664 /workspace/FLAME-MoE/Megatron-LM/megatron/core/transformer/transformer_layer.py:339: UserWarning: TransformerLayer._get_layer_offset is deprecated.Please use get_transformer_layer_offset instead. warnings.warn( loading distributed checkpoint from /workspace/FLAME-MoE/results/phase0/runs/flame512_g1_moe/ckpt at iteration 802 /workspace/FLAME-MoE/Megatron-LM/megatron/core/dist_checkpointing/strategies/torch.py:847: FutureWarning: `load_state_dict` is deprecated and will be removed in future versions. Please use `load` instead. checkpoint.load_state_dict( /usr/local/lib/python3.11/dist-packages/torch/distributed/checkpoint/planner_helpers.py:306: FutureWarning: Please use DTensor instead and we are deprecating ShardedTensor. device = getattr(value, "device", None) /workspace/FLAME-MoE/.venv/lib/python3.11/site-packages/transformer_engine/pytorch/module/base.py:566: FutureWarning: You are using `torch.load` with `weights_only=False` (the current default value), which uses the default pickle module implicitly. It is possible to construct malicious pickle data which will execute arbitrary code during unpickling (See https://github.com/pytorch/pytorch/blob/main/SECURITY.md#untrusted-models for more details). In a future release, the default value for `weights_only` will be flipped to `True`. This limits the functions that could be executed during unpickling. Arbitrary objects will no longer be allowed to be loaded via this mode unless they are explicitly allowlisted by the user via `torch.serialization.add_safe_globals`. We recommend you start setting `weights_only=True` for any use case where you don't have full control of the loaded file. Please open an issue on GitHub for any issues related to this experimental feature. state = torch.load(state, map_location="cuda") checkpoint version 3.0 2026-07-23:05:04:02,664 WARNING [rerun_state_machine.py:811] RerunStateMachine disabled via CLI, ignoring machine state saved in checkpoint successfully loaded checkpoint from /workspace/FLAME-MoE/results/phase0/runs/flame512_g1_moe/ckpt [ t 1/1, p 1/1 ] at iteration 802 2026-07-23:05:04:02,815 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc "HTTP/1.1 200 OK" 2026-07-23:05:04:03,052 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc "HTTP/1.1 200 OK" 2026-07-23:05:04:03,115 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc/revision/210d026faf9955653af8916fad021475a3f00453 "HTTP/1.1 200 OK" 2026-07-23:05:04:03,175 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc/tree/210d026faf9955653af8916fad021475a3f00453?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:04:03,402 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc/tree/210d026faf9955653af8916fad021475a3f00453/ARC-Challenge?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:04:04,025 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc "HTTP/1.1 200 OK" 2026-07-23:05:04:04,245 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc "HTTP/1.1 200 OK" 2026-07-23:05:04:04,516 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc/tree/210d026faf9955653af8916fad021475a3f00453/ARC-Easy?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:04:04,662 WARNING [task.py:799] [Task: boolq] metric acc is defined, but aggregation is not. using default aggregation=mean 2026-07-23:05:04:04,662 WARNING [task.py:811] [Task: boolq] metric acc is defined, but higher_is_better is not. using default higher_is_better=True 2026-07-23:05:04:04,722 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue "HTTP/1.1 200 OK" 2026-07-23:05:04:04,954 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue "HTTP/1.1 200 OK" 2026-07-23:05:04:05,016 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/revision/3de24cf8022e94f4ee4b9d55a6f539891524d646 "HTTP/1.1 200 OK" 2026-07-23:05:04:05,083 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/tree/3de24cf8022e94f4ee4b9d55a6f539891524d646?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:04:05,382 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/tree/3de24cf8022e94f4ee4b9d55a6f539891524d646/axb?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:04:05,445 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/revision/3de24cf8022e94f4ee4b9d55a6f539891524d646 "HTTP/1.1 200 OK" 2026-07-23:05:04:05,514 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/tree/3de24cf8022e94f4ee4b9d55a6f539891524d646/boolq?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:04:05,771 WARNING [task.py:799] [Task: copa] metric acc is defined, but aggregation is not. using default aggregation=mean 2026-07-23:05:04:05,771 WARNING [task.py:811] [Task: copa] metric acc is defined, but higher_is_better is not. using default higher_is_better=True 2026-07-23:05:04:05,831 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue "HTTP/1.1 200 OK" 2026-07-23:05:04:06,079 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue "HTTP/1.1 200 OK" 2026-07-23:05:04:06,374 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/tree/3de24cf8022e94f4ee4b9d55a6f539891524d646/copa?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:04:06,513 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag "HTTP/1.1 200 OK" 2026-07-23:05:04:06,778 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag "HTTP/1.1 200 OK" 2026-07-23:05:04:06,836 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag/revision/218ec52e09a7e7462a5400043bb9a69a41d06b76 "HTTP/1.1 200 OK" 2026-07-23:05:04:06,903 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag/tree/218ec52e09a7e7462a5400043bb9a69a41d06b76?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:04:07,130 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag/tree/218ec52e09a7e7462a5400043bb9a69a41d06b76/data?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:04:07,190 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag/revision/218ec52e09a7e7462a5400043bb9a69a41d06b76 "HTTP/1.1 200 OK" 2026-07-23:05:04:09,576 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai "HTTP/1.1 200 OK" 2026-07-23:05:04:09,799 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai "HTTP/1.1 200 OK" 2026-07-23:05:04:09,859 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/revision/900124bf3b8235c6daf21033af9948b3f07346c4 "HTTP/1.1 200 OK" 2026-07-23:05:04:09,918 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/tree/900124bf3b8235c6daf21033af9948b3f07346c4?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:04:10,167 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/tree/900124bf3b8235c6daf21033af9948b3f07346c4/default%2Ftest?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:04:10,233 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/tree/900124bf3b8235c6daf21033af9948b3f07346c4/default?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:04:10,300 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/revision/900124bf3b8235c6daf21033af9948b3f07346c4 "HTTP/1.1 200 OK" 2026-07-23:05:04:10,317 WARNING [task.py:325] [Task: lambada_openai] has_training_docs and has_validation_docs are False, using test_docs as fewshot_docs but this is not recommended. 2026-07-23:05:04:10,318 WARNING [task.py:325] [Task: lambada_openai] has_training_docs and has_validation_docs are False, using test_docs as fewshot_docs but this is not recommended. 2026-07-23:05:04:10,482 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa "HTTP/1.1 200 OK" 2026-07-23:05:04:10,730 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa "HTTP/1.1 200 OK" 2026-07-23:05:04:10,804 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/revision/388097ea7776314e93a529163e0fea805b8a6454 "HTTP/1.1 200 OK" 2026-07-23:05:04:10,867 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/tree/388097ea7776314e93a529163e0fea805b8a6454?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:04:11,097 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/tree/388097ea7776314e93a529163e0fea805b8a6454/additional?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:04:11,156 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/revision/388097ea7776314e93a529163e0fea805b8a6454 "HTTP/1.1 200 OK" 2026-07-23:05:04:11,247 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/tree/388097ea7776314e93a529163e0fea805b8a6454/main?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:04:11,527 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa "HTTP/1.1 200 OK" 2026-07-23:05:04:11,582 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa "HTTP/1.1 200 OK" 2026-07-23:05:04:11,641 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa/tree/main?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:04:11,693 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/datasets/ybisk/piqa/resolve/main/piqa.py "HTTP/1.1 307 Temporary Redirect" 2026-07-23:05:04:11,704 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/resolve-cache/datasets/ybisk/piqa/2e8ac2dffd59bac8c3c6714948f4c551a0848bb0/piqa.py?%2Fdatasets%2Fybisk%2Fpiqa%2Fresolve%2Fmain%2Fpiqa.py=&etag=%220f0d96172f9bf7aa10069374b3236ef083cfeb8e%22 "HTTP/1.1 200 OK" 2026-07-23:05:04:12,257 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa/revision/main "HTTP/1.1 200 OK" 2026-07-23:05:04:12,323 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa/tree/main?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:04:12,751 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq "HTTP/1.1 200 OK" 2026-07-23:05:04:12,971 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq "HTTP/1.1 200 OK" 2026-07-23:05:04:13,025 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq/revision/2c94ad3e1aafab77146f384e23536f97a4849815 "HTTP/1.1 200 OK" 2026-07-23:05:04:13,082 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq/tree/2c94ad3e1aafab77146f384e23536f97a4849815?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:04:13,298 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq/tree/2c94ad3e1aafab77146f384e23536f97a4849815/data?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:04:13,362 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq/revision/2c94ad3e1aafab77146f384e23536f97a4849815 "HTTP/1.1 200 OK" 2026-07-23:05:04:13,833 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande "HTTP/1.1 200 OK" 2026-07-23:05:04:14,067 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande "HTTP/1.1 200 OK" 2026-07-23:05:04:14,127 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/revision/01e74176c63542e6b0bcb004dcdea22d94fb67b5 "HTTP/1.1 200 OK" 2026-07-23:05:04:14,183 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/tree/01e74176c63542e6b0bcb004dcdea22d94fb67b5?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:04:14,416 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/tree/01e74176c63542e6b0bcb004dcdea22d94fb67b5/winogrande_debiased?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:04:14,475 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/revision/01e74176c63542e6b0bcb004dcdea22d94fb67b5 "HTTP/1.1 200 OK" 2026-07-23:05:04:14,565 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/tree/01e74176c63542e6b0bcb004dcdea22d94fb67b5/winogrande_xl?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:04:15,215 WARNING [evaluator.py:270] Overwriting default num_fewshot of winogrande from None to 0 2026-07-23:05:04:15,215 WARNING [evaluator.py:270] Overwriting default num_fewshot of sciq from None to 0 2026-07-23:05:04:15,216 WARNING [evaluator.py:270] Overwriting default num_fewshot of piqa from None to 0 2026-07-23:05:04:15,216 WARNING [evaluator.py:270] Overwriting default num_fewshot of openbookqa from None to 0 2026-07-23:05:04:15,216 WARNING [evaluator.py:270] Overwriting default num_fewshot of lambada_openai from None to 0 2026-07-23:05:04:15,216 WARNING [evaluator.py:270] Overwriting default num_fewshot of hellaswag from None to 0 2026-07-23:05:04:15,216 WARNING [evaluator.py:270] Overwriting default num_fewshot of copa from None to 0 2026-07-23:05:04:15,216 WARNING [evaluator.py:270] Overwriting default num_fewshot of boolq from None to 0 2026-07-23:05:04:15,216 WARNING [evaluator.py:270] Overwriting default num_fewshot of arc_easy from None to 0 2026-07-23:05:04:15,216 WARNING [evaluator.py:270] Overwriting default num_fewshot of arc_challenge from None to 0 2026-07-23:05:04:15,249 INFO [task.py:415] Building contexts for winogrande on rank 0... 0%| | 0/1267 [00:00 /workspace/FLAME-MoE/results/ablations/flame1e18_downstream.csv ==================================================================== [run] flame512_g1_temporal mode=temporal TEMPORAL=1 H=512 heads=32 ffn=2736 experts=64 topk=6 moe_ffn=352 shared=704 [run] tasks=arc_challenge,arc_easy,hellaswag,openbookqa,piqa,winogrande,boolq,copa,lambada_openai,sciq out=/workspace/FLAME-MoE/results/phase0/runs/flame512_g1_temporal/lmeval_1e18_0shot ==================================================================== /usr/local/lib/python3.11/dist-packages/torch/cuda/__init__.py:58: FutureWarning: The pynvml package is deprecated. Please install nvidia-ml-py instead. If you did not install pynvml directly, please report this to the maintainers of the package that installed pynvml for you. import pynvml # type: ignore[import] [run_lmeval] stubbed VLM backends (hf_vlms, vllm_vlms) /usr/local/lib/python3.11/dist-packages/torch/cuda/__init__.py:58: FutureWarning: The pynvml package is deprecated. Please install nvidia-ml-py instead. If you did not install pynvml directly, please report this to the maintainers of the package that installed pynvml for you. import pynvml # type: ignore[import] /usr/local/lib/python3.11/dist-packages/torch/cuda/__init__.py:58: FutureWarning: The pynvml package is deprecated. Please install nvidia-ml-py instead. If you did not install pynvml directly, please report this to the maintainers of the package that installed pynvml for you. import pynvml # type: ignore[import] /workspace/FLAME-MoE/.venv/lib/python3.11/site-packages/transformer_engine/pytorch/attention.py:108: UserWarning: To use flash-attn v3, please use the following commands to install: (1) pip install "git+https://github.com/Dao-AILab/flash-attention.git#egg=flashattn-hopper&subdirectory=hopper" (2) python_path=`python -c "import site; print(site.getsitepackages()[0])"` (3) mkdir -p $python_path/flashattn_hopper (4) wget -P $python_path/flashattn_hopper https://raw.githubusercontent.com/Dao-AILab/flash-attention/main/hopper/flash_attn_interface.py warnings.warn( [temporal] rolling-residency router installed (evict=min_logit) [run_lmeval] installed temporal residency router (native regime) /usr/local/lib/python3.11/dist-packages/requests/__init__.py:113: RequestsDependencyWarning: urllib3 (2.2.3) or chardet (6.0.0.post1)/charset_normalizer (3.3.2) doesn't match a supported version! warnings.warn( 2026-07-23:05:11:58,971 INFO [__main__.py:279] Verbosity set to INFO 2026-07-23:05:12:29,605 INFO [__init__.py:459] The tag 'arc_ca' is already registered as a group, this tag will not be registered. This may affect tasks you want to call. 2026-07-23:05:12:29,620 INFO [__init__.py:459] The tag 'arc_ca' is already registered as a group, this tag will not be registered. This may affect tasks you want to call. 2026-07-23:05:12:46,442 INFO [__main__.py:376] Selected Tasks: ['arc_challenge', 'arc_easy', 'boolq', 'copa', 'hellaswag', 'lambada_openai', 'openbookqa', 'piqa', 'sciq', 'winogrande'] 2026-07-23:05:12:46,464 INFO [evaluator.py:164] Setting random seed to 42 | Setting numpy seed to 42 | Setting torch manual seed to 42 | Setting fewshot manual seed to 42 2026-07-23:05:12:46,464 INFO [evaluator.py:201] Initializing megatron_lm model, with arguments: {} using world size: 1, data-parallel size: 1, context-parallel size: 1, hierarchical context-parallel sizes: Nonetensor-model-parallel size: 1, encoder-tensor-model-parallel size: 0, pipeline-model-parallel size: 1, encoder-pipeline-model-parallel size: 0 WARNING: overriding default arguments for exit_on_missing_checkpoint:True with exit_on_missing_checkpoint:False setting global batch size to 32 Number of virtual stages per pipeline stage: None accumulate and all-reduce gradients in fp32 for bfloat16 data type. using torch.bfloat16 for parameters ... ------------------------ arguments ------------------------ account_for_embedding_in_pipeline_split ......... False account_for_loss_in_pipeline_split .............. False accumulate_allreduce_grads_in_fp32 .............. True adam_beta1 ...................................... 0.9 adam_beta2 ...................................... 0.999 adam_eps ........................................ 1e-08 add_bias_linear ................................. False add_position_embedding .......................... True add_qkv_bias .................................... False adlr_autoresume ................................. False adlr_autoresume_interval ........................ 1000 align_grad_reduce ............................... True align_param_gather .............................. False app_tag_run_name ................................ None app_tag_run_version ............................. 0.0.0 apply_layernorm_1p .............................. False apply_query_key_layer_scaling ................... False apply_residual_connection_post_layernorm ........ False apply_rope_fusion ............................... True async_save ...................................... None async_tensor_model_parallel_allreduce ........... True attention_backend ............................... AttnBackend.auto attention_dropout ............................... 0.0 attention_softmax_in_fp32 ....................... False auto_detect_ckpt_format ......................... False barrier_with_L1_time ............................ True bert_binary_head ................................ True bert_embedder_type .............................. megatron bert_load ....................................... None bf16 ............................................ True bias_dropout_fusion ............................. True bias_gelu_fusion ................................ False bias_swiglu_fusion .............................. True biencoder_projection_dim ........................ 0 biencoder_shared_query_context_model ............ False block_data_path ................................. None calc_ft_timeouts ................................ False calculate_per_token_loss ........................ False check_for_large_grads ........................... False check_for_nan_in_loss_and_grad .................. True check_for_spiky_loss ............................ False check_weight_hash_across_dp_replicas_interval ... None ckpt_assume_constant_structure .................. False ckpt_convert_format ............................. None ckpt_convert_save ............................... None ckpt_convert_update_legacy_dist_opt_format ...... False ckpt_format ..................................... torch_dist ckpt_fully_parallel_load ........................ False ckpt_fully_parallel_save ........................ True ckpt_fully_parallel_save_deprecated ............. False ckpt_step ....................................... None classes_fraction ................................ 1.0 clip_grad ....................................... 1.0 clone_scatter_output_in_embedding ............... True config_logger_dir ............................... consumed_train_samples .......................... 0 consumed_valid_samples .......................... 0 context_parallel_size ........................... 1 cp_comm_type .................................... ['p2p'] create_attention_mask_in_dataloader ............. True cross_entropy_loss_fusion ....................... False cuda_graph_warmup_steps ......................... 3 data_args_path .................................. None data_cache_path ................................. None data_parallel_random_init ....................... False data_parallel_sharding_strategy ................. no_shard data_parallel_size .............................. 1 data_path ....................................... None data_per_class_fraction ......................... 1.0 data_sharding ................................... True dataloader_type ................................. single ddp_average_in_collective ....................... False ddp_bucket_size ................................. None ddp_num_buckets ................................. None ddp_pad_buckets_for_high_nccl_busbw ............. False decoder_first_pipeline_num_layers ............... None decoder_last_pipeline_num_layers ................ None decoder_num_layers .............................. None decoder_seq_length .............................. None decoupled_lr .................................... None decoupled_min_lr ................................ None decrease_batch_size_if_needed ................... False defer_embedding_wgrad_compute ................... False deprecated_use_mcore_models ..................... False deterministic_mode .............................. False dino_bottleneck_size ............................ 256 dino_freeze_last_layer .......................... 1 dino_head_hidden_size ........................... 2048 dino_local_crops_number ......................... 10 dino_local_img_size ............................. 96 dino_norm_last_layer ............................ False dino_teacher_temp ............................... 0.07 dino_warmup_teacher_temp ........................ 0.04 dino_warmup_teacher_temp_epochs ................. 30 disable_straggler_on_startup .................... False dist_ckpt_format_deprecated ..................... None dist_ckpt_strictness ............................ assume_ok_unexpected distribute_saved_activations .................... False distributed_backend ............................. nccl distributed_timeout_minutes ..................... 10 embedding_path .................................. None empty_unused_memory_level ....................... 0 enable_cuda_graph ............................... False enable_ft_package ............................... False enable_gloo_process_groups ...................... True enable_one_logger ............................... True encoder_num_layers .............................. 9 encoder_pipeline_model_parallel_size ............ 0 encoder_seq_length .............................. 2048 encoder_tensor_model_parallel_size .............. 0 end_weight_decay ................................ 0.01 eod_mask_loss ................................... False error_injection_rate ............................ 0 error_injection_type ............................ transient_error eval_interval ................................... 1000 eval_iters ...................................... 100 evidence_data_path .............................. None exit_duration_in_mins ........................... None exit_interval ................................... None exit_on_missing_checkpoint ...................... False exit_signal_handler ............................. False exp_avg_dtype ................................... torch.float32 exp_avg_sq_dtype ................................ torch.float32 expert_model_parallel_size ...................... 1 expert_tensor_parallel_size ..................... 1 ffn_hidden_size ................................. 2736 finetune ........................................ False flash_decode .................................... False fp16 ............................................ False fp16_lm_cross_entropy ........................... False fp32_residual_connection ........................ False fp8 ............................................. None fp8_amax_compute_algo ........................... most_recent fp8_amax_history_len ............................ 1 fp8_interval .................................... 1 fp8_margin ...................................... 0 fp8_param_gather ................................ False fp8_wgrad ....................................... True global_batch_size ............................... 32 grad_reduce_in_bf16 ............................. False gradient_accumulation_fusion .................... False gradient_reduce_div_fusion ...................... True group_query_attention ........................... False head_lr_mult .................................... 1.0 hidden_dropout .................................. 0.0 hidden_size ..................................... 512 hierarchical_context_parallel_sizes ............. None hybrid_attention_ratio .......................... 0.0 hybrid_mlp_ratio ................................ 0.0 hybrid_override_pattern ......................... None hysteresis ...................................... 2 ict_head_size ................................... None ict_load ........................................ None img_h ........................................... 224 img_w ........................................... 224 indexer_batch_size .............................. 128 indexer_log_interval ............................ 1000 inference_batch_times_seqlen_threshold .......... -1 inference_max_batch_size ........................ 8 inference_max_seq_length ........................ 2560 inference_rng_tracker ........................... False init_method_std ................................. 0.02 init_method_xavier_uniform ...................... False init_model_with_meta_device ..................... False initial_loss_scale .............................. 4294967296 iter_per_epoch .................................. 1250 iterations_to_skip .............................. [] keep_fp8_transpose_cache_when_using_custom_fsdp . False kv_channels ..................................... 16 kv_lora_rank .................................... 32 lazy_mpu_init ................................... None load ............................................ /workspace/FLAME-MoE/results/phase0/runs/flame512_g1_temporal/ckpt local_rank ...................................... 0 log_interval .................................... 100 log_loss_scale_to_tensorboard ................... True log_memory_to_tensorboard ....................... False log_num_zeros_in_grad ........................... False log_params_norm ................................. False log_progress .................................... False log_straggler ................................... False log_throughput .................................. False log_timers_to_tensorboard ....................... False log_validation_ppl_to_tensorboard ............... False log_world_size_to_tensorboard ................... False logging_level ................................... None loss_scale ...................................... None loss_scale_window ............................... 1000 lr .............................................. None lr_decay_iters .................................. None lr_decay_samples ................................ None lr_decay_style .................................. linear lr_warmup_fraction .............................. None lr_warmup_init .................................. 0.0 lr_warmup_iters ................................. 0 lr_warmup_samples ............................... 0 lr_wsd_decay_iters .............................. None lr_wsd_decay_samples ............................ None lr_wsd_decay_style .............................. exponential main_grads_dtype ................................ torch.float32 main_params_dtype ............................... torch.float32 make_vocab_size_divisible_by .................... 128 manual_gc ....................................... False manual_gc_eval .................................. True manual_gc_interval .............................. 0 mask_factor ..................................... 1.0 mask_prob ....................................... 0.15 mask_type ....................................... random masked_softmax_fusion ........................... True max_position_embeddings ......................... 2048 max_tokens_to_oom ............................... 10000000 memory_snapshot_path ............................ snapshot.pickle merge_file ...................................... None micro_batch_size ................................ 32 microbatch_group_size_per_vp_stage .............. None min_loss_scale .................................. 1.0 min_lr .......................................... 0.0 mmap_bin_files .................................. True mock_data ....................................... False moe_aux_loss_coeff .............................. 0.01 moe_enable_deepep ............................... False moe_expert_capacity_factor ...................... None moe_extended_tp ................................. False moe_ffn_hidden_size ............................. 352 moe_grouped_gemm ................................ True moe_input_jitter_eps ............................ None moe_layer_freq .................................. [0, 1, 1, 1, 1, 1, 1, 1, 1] moe_layer_recompute ............................. False moe_pad_expert_input_to_capacity ................ False moe_per_layer_logging ........................... False moe_permute_fusion .............................. False moe_router_bias_update_rate ..................... 0.001 moe_router_dtype ................................ fp32 moe_router_enable_expert_bias ................... False moe_router_group_topk ........................... None moe_router_load_balancing_type .................. aux_loss moe_router_num_groups ........................... None moe_router_pre_softmax .......................... True moe_router_score_function ....................... softmax moe_router_topk ................................. 6 moe_router_topk_scaling_factor .................. None moe_shared_expert_intermediate_size ............. 704 moe_shared_expert_overlap ....................... False moe_token_dispatcher_type ....................... alltoall moe_token_drop_policy ........................... probs moe_use_legacy_grouped_gemm ..................... False moe_use_upcycling ............................... False moe_z_loss_coeff ................................ 0.001 mscale .......................................... 1.0 mscale_all_dim .................................. 1.0 multi_latent_attention .......................... False nccl_communicator_config_path ................... None no_load_optim ................................... True no_load_rng ..................................... True no_persist_layer_norm ........................... False no_save_optim ................................... None no_save_rng ..................................... None non_persistent_ckpt_type ........................ None non_persistent_global_ckpt_dir .................. None non_persistent_local_ckpt_algo .................. fully_parallel non_persistent_local_ckpt_dir ................... None non_persistent_save_interval .................... None norm_epsilon .................................... 1e-06 normalization ................................... RMSNorm num_attention_heads ............................. 32 num_channels .................................... 3 num_classes ..................................... 1000 num_dataset_builder_threads ..................... 1 num_distributed_optimizer_instances ............. 1 num_experts ..................................... 64 num_layers ...................................... 9 num_layers_per_virtual_pipeline_stage ........... None num_query_groups ................................ 1 num_virtual_stages_per_pipeline_rank ............ None num_workers ..................................... 2 one_logger_async ................................ False one_logger_project .............................. megatron-lm one_logger_run_name ............................. None onnx_safe ....................................... None openai_gelu ..................................... False optimizer ....................................... adam optimizer_cpu_offload ........................... False optimizer_offload_fraction ...................... 1.0 output_bert_embeddings .......................... False overlap_cpu_optimizer_d2h_h2d ................... False overlap_grad_reduce ............................. False overlap_p2p_comm ................................ False overlap_p2p_comm_warmup_flush ................... False overlap_param_gather ............................ False overlap_param_gather_with_optimizer_step ........ False override_opt_param_scheduler .................... False params_dtype .................................... torch.bfloat16 patch_dim ....................................... 16 per_split_data_args_path ........................ None perform_initialization .......................... True pin_cpu_grads ................................... True pin_cpu_params .................................. True pipeline_model_parallel_comm_backend ............ None pipeline_model_parallel_size .................... 1 pipeline_model_parallel_split_rank .............. None position_embedding_type ......................... rope pretrained_checkpoint ........................... None profile ......................................... False profile_ranks ................................... [0] profile_step_end ................................ 12 profile_step_start .............................. 10 q_lora_rank ..................................... None qk_head_dim ..................................... 128 qk_layernorm .................................... False qk_pos_emb_head_dim ............................. 64 query_in_block_prob ............................. 0.1 rampup_batch_size ............................... None rank ............................................ 0 recompute_granularity ........................... None recompute_method ................................ None recompute_num_layers ............................ None record_memory_history ........................... False relative_attention_max_distance ................. 128 relative_attention_num_buckets .................. 32 replication ..................................... False replication_factor .............................. 2 replication_jump ................................ None rerun_mode ...................................... disabled reset_attention_mask ............................ False reset_position_ids .............................. False result_rejected_tracker_filename ................ None retriever_report_topk_accuracies ................ [] retriever_score_scaling ......................... False retriever_seq_length ............................ 256 retro_add_retriever ............................. False retro_attention_gate ............................ 1 retro_cyclic_train_iters ........................ None retro_encoder_attention_dropout ................. 0.1 retro_encoder_hidden_dropout .................... 0.1 retro_encoder_layers ............................ 2 retro_num_neighbors ............................. 2 retro_num_retrieved_chunks ...................... 2 retro_project_dir ............................... None retro_verify_neighbor_count ..................... True rope_scaling_factor ............................. 8.0 rotary_base ..................................... 10000 rotary_interleaved .............................. False rotary_percent .................................. 1.0 rotary_scaling_factor ........................... 1.0 rotary_seq_len_interpolation_factor ............. None s3_cache_path ................................... None sample_rate ..................................... 1.0 save ............................................ None save_interval ................................... None scatter_gather_tensors_in_pipeline .............. True seed ............................................ 42 seq_length ...................................... 2048 sequence_parallel ............................... False sgd_momentum .................................... 0.9 short_seq_prob .................................. 0.1 skip_train ...................................... False skipped_train_samples ........................... 0 spec ............................................ None split ........................................... None squared_relu .................................... False start_weight_decay .............................. 0.01 straggler_ctrlr_port ............................ 65535 straggler_minmax_count .......................... 1 suggested_communication_unit_size ............... 400000000 swiglu .......................................... True swin_backbone_type .............................. tiny te_rng_tracker .................................. False tensor_model_parallel_size ...................... 1 tensorboard_dir ................................. None tensorboard_log_interval ........................ 1 tensorboard_queue_size .......................... 1000 test_data_path .................................. None test_mode ....................................... False tiktoken_num_special_tokens ..................... 1000 tiktoken_pattern ................................ None tiktoken_special_tokens ......................... None timing_log_level ................................ 0 timing_log_option ............................... minmax titles_data_path ................................ None tokenizer_model ................................. EleutherAI/pythia-12b tokenizer_type .................................. HuggingFaceTokenizer tp_comm_bootstrap_backend ....................... nccl tp_comm_bulk_dgrad .............................. True tp_comm_bulk_wgrad .............................. True tp_comm_overlap ................................. False tp_comm_overlap_ag .............................. True tp_comm_overlap_cfg ............................. None tp_comm_overlap_rs .............................. True tp_comm_overlap_rs_dgrad ........................ False tp_comm_split_ag ................................ True tp_comm_split_rs ................................ True train_data_path ................................. None train_iters ..................................... None train_samples ................................... None train_sync_interval ............................. None transformer_impl ................................ transformer_engine transformer_pipeline_model_parallel_size ........ 1 untie_embeddings_and_output_weights ............. True use_checkpoint_args ............................. False use_checkpoint_opt_param_scheduler .............. False use_cpu_initialization .......................... None use_custom_fsdp ................................. False use_dist_ckpt ................................... True use_dist_ckpt_deprecated ........................ False use_distributed_optimizer ....................... True use_flash_attn .................................. False use_legacy_models ............................... False use_mp_args_from_checkpoint_args ................ False use_one_sent_docs ............................... False use_persistent_ckpt_worker ...................... False use_precision_aware_optimizer ................... False use_pytorch_profiler ............................ False use_ring_exchange_p2p ........................... False use_rope_scaling ................................ False use_rotary_position_embeddings .................. False use_tokenizer_model_from_checkpoint_args ........ True use_torch_fsdp2 ................................. False use_torch_optimizer_for_cpu_offload ............. False use_tp_pp_dp_mapping ............................ False v_head_dim ...................................... 128 valid_data_path ................................. None variable_seq_lengths ............................ False virtual_pipeline_model_parallel_size ............ None vision_backbone_type ............................ vit vision_pretraining .............................. False vision_pretraining_type ......................... classify vocab_extra_ids ................................. 0 vocab_file ...................................... None vocab_size ...................................... None wandb_exp_name .................................. wandb_project ................................... wandb_save_dir .................................. weight_decay .................................... 0.01 weight_decay_incr_style ......................... constant wgrad_deferral_limit ............................ 0 world_size ...................................... 1 yaml_cfg ........................................ None -------------------- end of arguments --------------------- 2026-07-23:05:12:46,480 INFO [num_microbatches_calculator.py:228] setting number of microbatches to constant 1 > building HuggingFaceTokenizer tokenizer ... 2026-07-23:05:12:46,675 INFO [_client.py:1038] HTTP Request: HEAD https://huggingface.co/EleutherAI/pythia-12b/resolve/main/config.json "HTTP/1.1 307 Temporary Redirect" 2026-07-23:05:12:46,686 INFO [_client.py:1038] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/EleutherAI/pythia-12b/bb1e3e710cdf6b524461d543cfb5ba773f0a81b6/config.json "HTTP/1.1 200 OK" 2026-07-23:05:12:46,806 INFO [_client.py:1038] HTTP Request: HEAD https://huggingface.co/EleutherAI/pythia-12b/resolve/main/tokenizer_config.json "HTTP/1.1 307 Temporary Redirect" 2026-07-23:05:12:46,819 INFO [_client.py:1038] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/EleutherAI/pythia-12b/bb1e3e710cdf6b524461d543cfb5ba773f0a81b6/tokenizer_config.json "HTTP/1.1 200 OK" 2026-07-23:05:12:46,923 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/models/EleutherAI/pythia-12b/tree/main/additional_chat_templates?recursive=false&expand=false "HTTP/1.1 404 Not Found" 2026-07-23:05:12:46,988 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/models/EleutherAI/pythia-12b/tree/main?recursive=true&expand=false "HTTP/1.1 200 OK" > padded vocab (size: 50277) with 27 dummy tokens (new size: 50304) WARNING: one_logger package is required to enable e2e metrics tracking. please go to https://confluence.nvidia.com/display/MLWFO/Package+Repositories for details to install it 2026-07-23:05:12:47,202 WARNING [rerun_state_machine.py:239] RerunStateMachine initialized in mode RerunMode.DISABLED > initializing torch distributed ... [W723 05:12:47.719836875 CUDAAllocatorConfig.h:28] Warning: expandable_segments not supported on this platform (function operator()) > initialized tensor model parallel with size 1 > initialized pipeline model parallel with size 1 > setting random seeds to 42 ... > compiling dataset index builder ... make: Entering directory '/workspace/FLAME-MoE/Megatron-LM/megatron/core/datasets' make: Nothing to be done for 'default'. make: Leaving directory '/workspace/FLAME-MoE/Megatron-LM/megatron/core/datasets' >>> done with dataset index builder. Compilation time: 0.796 seconds > compiling and loading fused kernels ... >>> done with compiling and loading fused kernels. Compilation time: 0.351 seconds building GPT model ... > number of parameters on (tensor, pipeline) model parallel rank (0, 0): 350897664 /workspace/FLAME-MoE/Megatron-LM/megatron/core/transformer/transformer_layer.py:339: UserWarning: TransformerLayer._get_layer_offset is deprecated.Please use get_transformer_layer_offset instead. warnings.warn( loading distributed checkpoint from /workspace/FLAME-MoE/results/phase0/runs/flame512_g1_temporal/ckpt at iteration 802 /workspace/FLAME-MoE/Megatron-LM/megatron/core/dist_checkpointing/strategies/torch.py:847: FutureWarning: `load_state_dict` is deprecated and will be removed in future versions. Please use `load` instead. checkpoint.load_state_dict( /usr/local/lib/python3.11/dist-packages/torch/distributed/checkpoint/planner_helpers.py:306: FutureWarning: Please use DTensor instead and we are deprecating ShardedTensor. device = getattr(value, "device", None) /workspace/FLAME-MoE/.venv/lib/python3.11/site-packages/transformer_engine/pytorch/module/base.py:566: FutureWarning: You are using `torch.load` with `weights_only=False` (the current default value), which uses the default pickle module implicitly. It is possible to construct malicious pickle data which will execute arbitrary code during unpickling (See https://github.com/pytorch/pytorch/blob/main/SECURITY.md#untrusted-models for more details). In a future release, the default value for `weights_only` will be flipped to `True`. This limits the functions that could be executed during unpickling. Arbitrary objects will no longer be allowed to be loaded via this mode unless they are explicitly allowlisted by the user via `torch.serialization.add_safe_globals`. We recommend you start setting `weights_only=True` for any use case where you don't have full control of the loaded file. Please open an issue on GitHub for any issues related to this experimental feature. state = torch.load(state, map_location="cuda") checkpoint version 3.0 2026-07-23:05:12:54,957 WARNING [rerun_state_machine.py:811] RerunStateMachine disabled via CLI, ignoring machine state saved in checkpoint successfully loaded checkpoint from /workspace/FLAME-MoE/results/phase0/runs/flame512_g1_temporal/ckpt [ t 1/1, p 1/1 ] at iteration 802 2026-07-23:05:12:55,130 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc "HTTP/1.1 200 OK" 2026-07-23:05:12:55,375 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc "HTTP/1.1 200 OK" 2026-07-23:05:12:55,441 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc/revision/210d026faf9955653af8916fad021475a3f00453 "HTTP/1.1 200 OK" 2026-07-23:05:12:55,507 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc/tree/210d026faf9955653af8916fad021475a3f00453?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:12:56,026 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc/tree/210d026faf9955653af8916fad021475a3f00453/ARC-Challenge?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:12:56,532 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc "HTTP/1.1 200 OK" 2026-07-23:05:12:56,749 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc "HTTP/1.1 200 OK" 2026-07-23:05:12:56,992 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc/tree/210d026faf9955653af8916fad021475a3f00453/ARC-Easy?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:12:57,070 WARNING [task.py:799] [Task: boolq] metric acc is defined, but aggregation is not. using default aggregation=mean 2026-07-23:05:12:57,070 WARNING [task.py:811] [Task: boolq] metric acc is defined, but higher_is_better is not. using default higher_is_better=True 2026-07-23:05:12:57,132 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue "HTTP/1.1 200 OK" 2026-07-23:05:12:57,354 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue "HTTP/1.1 200 OK" 2026-07-23:05:12:57,415 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/revision/3de24cf8022e94f4ee4b9d55a6f539891524d646 "HTTP/1.1 200 OK" 2026-07-23:05:12:57,481 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/tree/3de24cf8022e94f4ee4b9d55a6f539891524d646?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:12:57,760 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/tree/3de24cf8022e94f4ee4b9d55a6f539891524d646/axb?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:12:57,822 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/revision/3de24cf8022e94f4ee4b9d55a6f539891524d646 "HTTP/1.1 200 OK" 2026-07-23:05:12:57,892 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/tree/3de24cf8022e94f4ee4b9d55a6f539891524d646/boolq?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:12:58,139 WARNING [task.py:799] [Task: copa] metric acc is defined, but aggregation is not. using default aggregation=mean 2026-07-23:05:12:58,139 WARNING [task.py:811] [Task: copa] metric acc is defined, but higher_is_better is not. using default higher_is_better=True 2026-07-23:05:12:58,201 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue "HTTP/1.1 200 OK" 2026-07-23:05:12:58,433 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue "HTTP/1.1 200 OK" 2026-07-23:05:12:58,688 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/tree/3de24cf8022e94f4ee4b9d55a6f539891524d646/copa?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:12:58,832 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag "HTTP/1.1 200 OK" 2026-07-23:05:12:59,081 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag "HTTP/1.1 200 OK" 2026-07-23:05:12:59,141 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag/revision/218ec52e09a7e7462a5400043bb9a69a41d06b76 "HTTP/1.1 200 OK" 2026-07-23:05:12:59,201 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag/tree/218ec52e09a7e7462a5400043bb9a69a41d06b76?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:12:59,383 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag/tree/218ec52e09a7e7462a5400043bb9a69a41d06b76/data?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:12:59,441 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag/revision/218ec52e09a7e7462a5400043bb9a69a41d06b76 "HTTP/1.1 200 OK" 2026-07-23:05:13:01,823 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai "HTTP/1.1 200 OK" 2026-07-23:05:13:02,174 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai "HTTP/1.1 200 OK" 2026-07-23:05:13:02,233 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/revision/900124bf3b8235c6daf21033af9948b3f07346c4 "HTTP/1.1 200 OK" 2026-07-23:05:13:02,291 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/tree/900124bf3b8235c6daf21033af9948b3f07346c4?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:13:02,551 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/tree/900124bf3b8235c6daf21033af9948b3f07346c4/default%2Ftest?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:13:02,615 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/tree/900124bf3b8235c6daf21033af9948b3f07346c4/default?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:13:02,681 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/revision/900124bf3b8235c6daf21033af9948b3f07346c4 "HTTP/1.1 200 OK" 2026-07-23:05:13:02,699 WARNING [task.py:325] [Task: lambada_openai] has_training_docs and has_validation_docs are False, using test_docs as fewshot_docs but this is not recommended. 2026-07-23:05:13:02,699 WARNING [task.py:325] [Task: lambada_openai] has_training_docs and has_validation_docs are False, using test_docs as fewshot_docs but this is not recommended. 2026-07-23:05:13:02,869 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa "HTTP/1.1 200 OK" 2026-07-23:05:13:03,109 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa "HTTP/1.1 200 OK" 2026-07-23:05:13:03,181 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/revision/388097ea7776314e93a529163e0fea805b8a6454 "HTTP/1.1 200 OK" 2026-07-23:05:13:03,242 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/tree/388097ea7776314e93a529163e0fea805b8a6454?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:13:03,471 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/tree/388097ea7776314e93a529163e0fea805b8a6454/additional?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:13:03,532 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/revision/388097ea7776314e93a529163e0fea805b8a6454 "HTTP/1.1 200 OK" 2026-07-23:05:13:03,614 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/tree/388097ea7776314e93a529163e0fea805b8a6454/main?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:13:03,882 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa "HTTP/1.1 200 OK" 2026-07-23:05:13:03,939 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa "HTTP/1.1 200 OK" 2026-07-23:05:13:04,000 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa/tree/main?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:13:04,058 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/datasets/ybisk/piqa/resolve/main/piqa.py "HTTP/1.1 307 Temporary Redirect" 2026-07-23:05:13:04,071 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/resolve-cache/datasets/ybisk/piqa/2e8ac2dffd59bac8c3c6714948f4c551a0848bb0/piqa.py?%2Fdatasets%2Fybisk%2Fpiqa%2Fresolve%2Fmain%2Fpiqa.py=&etag=%220f0d96172f9bf7aa10069374b3236ef083cfeb8e%22 "HTTP/1.1 200 OK" 2026-07-23:05:13:04,669 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa/revision/main "HTTP/1.1 200 OK" 2026-07-23:05:13:04,727 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa/tree/main?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:13:05,168 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq "HTTP/1.1 200 OK" 2026-07-23:05:13:05,441 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq "HTTP/1.1 200 OK" 2026-07-23:05:13:05,499 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq/revision/2c94ad3e1aafab77146f384e23536f97a4849815 "HTTP/1.1 200 OK" 2026-07-23:05:13:05,559 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq/tree/2c94ad3e1aafab77146f384e23536f97a4849815?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:13:05,825 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq/tree/2c94ad3e1aafab77146f384e23536f97a4849815/data?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:13:05,887 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq/revision/2c94ad3e1aafab77146f384e23536f97a4849815 "HTTP/1.1 200 OK" 2026-07-23:05:13:06,315 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande "HTTP/1.1 200 OK" 2026-07-23:05:13:06,617 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande "HTTP/1.1 200 OK" 2026-07-23:05:13:06,678 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/revision/01e74176c63542e6b0bcb004dcdea22d94fb67b5 "HTTP/1.1 200 OK" 2026-07-23:05:13:06,739 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/tree/01e74176c63542e6b0bcb004dcdea22d94fb67b5?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:13:06,975 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/tree/01e74176c63542e6b0bcb004dcdea22d94fb67b5/winogrande_debiased?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:13:07,035 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/revision/01e74176c63542e6b0bcb004dcdea22d94fb67b5 "HTTP/1.1 200 OK" 2026-07-23:05:13:07,118 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/tree/01e74176c63542e6b0bcb004dcdea22d94fb67b5/winogrande_xl?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:13:07,773 WARNING [evaluator.py:270] Overwriting default num_fewshot of winogrande from None to 0 2026-07-23:05:13:07,773 WARNING [evaluator.py:270] Overwriting default num_fewshot of sciq from None to 0 2026-07-23:05:13:07,773 WARNING [evaluator.py:270] Overwriting default num_fewshot of piqa from None to 0 2026-07-23:05:13:07,773 WARNING [evaluator.py:270] Overwriting default num_fewshot of openbookqa from None to 0 2026-07-23:05:13:07,773 WARNING [evaluator.py:270] Overwriting default num_fewshot of lambada_openai from None to 0 2026-07-23:05:13:07,773 WARNING [evaluator.py:270] Overwriting default num_fewshot of hellaswag from None to 0 2026-07-23:05:13:07,773 WARNING [evaluator.py:270] Overwriting default num_fewshot of copa from None to 0 2026-07-23:05:13:07,773 WARNING [evaluator.py:270] Overwriting default num_fewshot of boolq from None to 0 2026-07-23:05:13:07,773 WARNING [evaluator.py:270] Overwriting default num_fewshot of arc_easy from None to 0 2026-07-23:05:13:07,773 WARNING [evaluator.py:270] Overwriting default num_fewshot of arc_challenge from None to 0 2026-07-23:05:13:07,815 INFO [task.py:415] Building contexts for winogrande on rank 0... 0%| | 0/1267 [00:00 /workspace/FLAME-MoE/results/ablations/flame1e18_downstream.csv ==================================================================== [run] flame512_g3_moe mode=moe TEMPORAL=0 H=512 heads=32 ffn=2736 experts=192 topk=18 moe_ffn=118 shared=704 [run] tasks=arc_challenge,arc_easy,hellaswag,openbookqa,piqa,winogrande,boolq,copa,lambada_openai,sciq out=/workspace/FLAME-MoE/results/phase0/runs/flame512_g3_moe/lmeval_1e18_0shot ==================================================================== /usr/local/lib/python3.11/dist-packages/torch/cuda/__init__.py:58: FutureWarning: The pynvml package is deprecated. Please install nvidia-ml-py instead. If you did not install pynvml directly, please report this to the maintainers of the package that installed pynvml for you. import pynvml # type: ignore[import] [run_lmeval] stubbed VLM backends (hf_vlms, vllm_vlms) /usr/local/lib/python3.11/dist-packages/torch/cuda/__init__.py:58: FutureWarning: The pynvml package is deprecated. Please install nvidia-ml-py instead. If you did not install pynvml directly, please report this to the maintainers of the package that installed pynvml for you. import pynvml # type: ignore[import] /usr/local/lib/python3.11/dist-packages/requests/__init__.py:113: RequestsDependencyWarning: urllib3 (2.2.3) or chardet (6.0.0.post1)/charset_normalizer (3.3.2) doesn't match a supported version! warnings.warn( /usr/local/lib/python3.11/dist-packages/torch/cuda/__init__.py:58: FutureWarning: The pynvml package is deprecated. Please install nvidia-ml-py instead. If you did not install pynvml directly, please report this to the maintainers of the package that installed pynvml for you. import pynvml # type: ignore[import] /workspace/FLAME-MoE/.venv/lib/python3.11/site-packages/transformer_engine/pytorch/attention.py:108: UserWarning: To use flash-attn v3, please use the following commands to install: (1) pip install "git+https://github.com/Dao-AILab/flash-attention.git#egg=flashattn-hopper&subdirectory=hopper" (2) python_path=`python -c "import site; print(site.getsitepackages()[0])"` (3) mkdir -p $python_path/flashattn_hopper (4) wget -P $python_path/flashattn_hopper https://raw.githubusercontent.com/Dao-AILab/flash-attention/main/hopper/flash_attn_interface.py warnings.warn( /workspace/FLAME-MoE/Megatron-LM/megatron/core/models/gpt/gpt_layer_specs.py:54: UserWarning: Apex is not installed. Falling back to Torch Norm warnings.warn('Apex is not installed. Falling back to Torch Norm') /workspace/FLAME-MoE/Megatron-LM/megatron/core/models/retro/encoder_spec.py:47: UserWarning: Apex is not installed. Falling back to Torch Norm warnings.warn(f'Apex is not installed. Falling back to Torch Norm') /workspace/FLAME-MoE/Megatron-LM/megatron/core/models/retro/decoder_spec.py:39: UserWarning: Apex is not installed. Falling back to Torch Norm warnings.warn(f'Apex is not installed. Falling back to Torch Norm') 2026-07-23:05:20:57,088 INFO [__main__.py:279] Verbosity set to INFO 2026-07-23:05:21:26,734 INFO [__init__.py:459] The tag 'arc_ca' is already registered as a group, this tag will not be registered. This may affect tasks you want to call. 2026-07-23:05:21:26,749 INFO [__init__.py:459] The tag 'arc_ca' is already registered as a group, this tag will not be registered. This may affect tasks you want to call. 2026-07-23:05:21:41,417 INFO [__main__.py:376] Selected Tasks: ['arc_challenge', 'arc_easy', 'boolq', 'copa', 'hellaswag', 'lambada_openai', 'openbookqa', 'piqa', 'sciq', 'winogrande'] 2026-07-23:05:21:41,436 INFO [evaluator.py:164] Setting random seed to 42 | Setting numpy seed to 42 | Setting torch manual seed to 42 | Setting fewshot manual seed to 42 2026-07-23:05:21:41,437 INFO [evaluator.py:201] Initializing megatron_lm model, with arguments: {} using world size: 1, data-parallel size: 1, context-parallel size: 1, hierarchical context-parallel sizes: Nonetensor-model-parallel size: 1, encoder-tensor-model-parallel size: 0, pipeline-model-parallel size: 1, encoder-pipeline-model-parallel size: 0 WARNING: overriding default arguments for exit_on_missing_checkpoint:True with exit_on_missing_checkpoint:False setting global batch size to 32 Number of virtual stages per pipeline stage: None accumulate and all-reduce gradients in fp32 for bfloat16 data type. using torch.bfloat16 for parameters ... ------------------------ arguments ------------------------ account_for_embedding_in_pipeline_split ......... False account_for_loss_in_pipeline_split .............. False accumulate_allreduce_grads_in_fp32 .............. True adam_beta1 ...................................... 0.9 adam_beta2 ...................................... 0.999 adam_eps ........................................ 1e-08 add_bias_linear ................................. False add_position_embedding .......................... True add_qkv_bias .................................... False adlr_autoresume ................................. False adlr_autoresume_interval ........................ 1000 align_grad_reduce ............................... True align_param_gather .............................. False app_tag_run_name ................................ None app_tag_run_version ............................. 0.0.0 apply_layernorm_1p .............................. False apply_query_key_layer_scaling ................... False apply_residual_connection_post_layernorm ........ False apply_rope_fusion ............................... True async_save ...................................... None async_tensor_model_parallel_allreduce ........... True attention_backend ............................... AttnBackend.auto attention_dropout ............................... 0.0 attention_softmax_in_fp32 ....................... False auto_detect_ckpt_format ......................... False barrier_with_L1_time ............................ True bert_binary_head ................................ True bert_embedder_type .............................. megatron bert_load ....................................... None bf16 ............................................ True bias_dropout_fusion ............................. True bias_gelu_fusion ................................ False bias_swiglu_fusion .............................. True biencoder_projection_dim ........................ 0 biencoder_shared_query_context_model ............ False block_data_path ................................. None calc_ft_timeouts ................................ False calculate_per_token_loss ........................ False check_for_large_grads ........................... False check_for_nan_in_loss_and_grad .................. True check_for_spiky_loss ............................ False check_weight_hash_across_dp_replicas_interval ... None ckpt_assume_constant_structure .................. False ckpt_convert_format ............................. None ckpt_convert_save ............................... None ckpt_convert_update_legacy_dist_opt_format ...... False ckpt_format ..................................... torch_dist ckpt_fully_parallel_load ........................ False ckpt_fully_parallel_save ........................ True ckpt_fully_parallel_save_deprecated ............. False ckpt_step ....................................... None classes_fraction ................................ 1.0 clip_grad ....................................... 1.0 clone_scatter_output_in_embedding ............... True config_logger_dir ............................... consumed_train_samples .......................... 0 consumed_valid_samples .......................... 0 context_parallel_size ........................... 1 cp_comm_type .................................... ['p2p'] create_attention_mask_in_dataloader ............. True cross_entropy_loss_fusion ....................... False cuda_graph_warmup_steps ......................... 3 data_args_path .................................. None data_cache_path ................................. None data_parallel_random_init ....................... False data_parallel_sharding_strategy ................. no_shard data_parallel_size .............................. 1 data_path ....................................... None data_per_class_fraction ......................... 1.0 data_sharding ................................... True dataloader_type ................................. single ddp_average_in_collective ....................... False ddp_bucket_size ................................. None ddp_num_buckets ................................. None ddp_pad_buckets_for_high_nccl_busbw ............. False decoder_first_pipeline_num_layers ............... None decoder_last_pipeline_num_layers ................ None decoder_num_layers .............................. None decoder_seq_length .............................. None decoupled_lr .................................... None decoupled_min_lr ................................ None decrease_batch_size_if_needed ................... False defer_embedding_wgrad_compute ................... False deprecated_use_mcore_models ..................... False deterministic_mode .............................. False dino_bottleneck_size ............................ 256 dino_freeze_last_layer .......................... 1 dino_head_hidden_size ........................... 2048 dino_local_crops_number ......................... 10 dino_local_img_size ............................. 96 dino_norm_last_layer ............................ False dino_teacher_temp ............................... 0.07 dino_warmup_teacher_temp ........................ 0.04 dino_warmup_teacher_temp_epochs ................. 30 disable_straggler_on_startup .................... False dist_ckpt_format_deprecated ..................... None dist_ckpt_strictness ............................ assume_ok_unexpected distribute_saved_activations .................... False distributed_backend ............................. nccl distributed_timeout_minutes ..................... 10 embedding_path .................................. None empty_unused_memory_level ....................... 0 enable_cuda_graph ............................... False enable_ft_package ............................... False enable_gloo_process_groups ...................... True enable_one_logger ............................... True encoder_num_layers .............................. 9 encoder_pipeline_model_parallel_size ............ 0 encoder_seq_length .............................. 2048 encoder_tensor_model_parallel_size .............. 0 end_weight_decay ................................ 0.01 eod_mask_loss ................................... False error_injection_rate ............................ 0 error_injection_type ............................ transient_error eval_interval ................................... 1000 eval_iters ...................................... 100 evidence_data_path .............................. None exit_duration_in_mins ........................... None exit_interval ................................... None exit_on_missing_checkpoint ...................... False exit_signal_handler ............................. False exp_avg_dtype ................................... torch.float32 exp_avg_sq_dtype ................................ torch.float32 expert_model_parallel_size ...................... 1 expert_tensor_parallel_size ..................... 1 ffn_hidden_size ................................. 2736 finetune ........................................ False flash_decode .................................... False fp16 ............................................ False fp16_lm_cross_entropy ........................... False fp32_residual_connection ........................ False fp8 ............................................. None fp8_amax_compute_algo ........................... most_recent fp8_amax_history_len ............................ 1 fp8_interval .................................... 1 fp8_margin ...................................... 0 fp8_param_gather ................................ False fp8_wgrad ....................................... True global_batch_size ............................... 32 grad_reduce_in_bf16 ............................. False gradient_accumulation_fusion .................... False gradient_reduce_div_fusion ...................... True group_query_attention ........................... False head_lr_mult .................................... 1.0 hidden_dropout .................................. 0.0 hidden_size ..................................... 512 hierarchical_context_parallel_sizes ............. None hybrid_attention_ratio .......................... 0.0 hybrid_mlp_ratio ................................ 0.0 hybrid_override_pattern ......................... None hysteresis ...................................... 2 ict_head_size ................................... None ict_load ........................................ None img_h ........................................... 224 img_w ........................................... 224 indexer_batch_size .............................. 128 indexer_log_interval ............................ 1000 inference_batch_times_seqlen_threshold .......... -1 inference_max_batch_size ........................ 8 inference_max_seq_length ........................ 2560 inference_rng_tracker ........................... False init_method_std ................................. 0.02 init_method_xavier_uniform ...................... False init_model_with_meta_device ..................... False initial_loss_scale .............................. 4294967296 iter_per_epoch .................................. 1250 iterations_to_skip .............................. [] keep_fp8_transpose_cache_when_using_custom_fsdp . False kv_channels ..................................... 16 kv_lora_rank .................................... 32 lazy_mpu_init ................................... None load ............................................ /workspace/FLAME-MoE/results/phase0/runs/flame512_g3_moe/ckpt local_rank ...................................... 0 log_interval .................................... 100 log_loss_scale_to_tensorboard ................... True log_memory_to_tensorboard ....................... False log_num_zeros_in_grad ........................... False log_params_norm ................................. False log_progress .................................... False log_straggler ................................... False log_throughput .................................. False log_timers_to_tensorboard ....................... False log_validation_ppl_to_tensorboard ............... False log_world_size_to_tensorboard ................... False logging_level ................................... None loss_scale ...................................... None loss_scale_window ............................... 1000 lr .............................................. None lr_decay_iters .................................. None lr_decay_samples ................................ None lr_decay_style .................................. linear lr_warmup_fraction .............................. None lr_warmup_init .................................. 0.0 lr_warmup_iters ................................. 0 lr_warmup_samples ............................... 0 lr_wsd_decay_iters .............................. None lr_wsd_decay_samples ............................ None lr_wsd_decay_style .............................. exponential main_grads_dtype ................................ torch.float32 main_params_dtype ............................... torch.float32 make_vocab_size_divisible_by .................... 128 manual_gc ....................................... False manual_gc_eval .................................. True manual_gc_interval .............................. 0 mask_factor ..................................... 1.0 mask_prob ....................................... 0.15 mask_type ....................................... random masked_softmax_fusion ........................... True max_position_embeddings ......................... 2048 max_tokens_to_oom ............................... 10000000 memory_snapshot_path ............................ snapshot.pickle merge_file ...................................... None micro_batch_size ................................ 32 microbatch_group_size_per_vp_stage .............. None min_loss_scale .................................. 1.0 min_lr .......................................... 0.0 mmap_bin_files .................................. True mock_data ....................................... False moe_aux_loss_coeff .............................. 0.01 moe_enable_deepep ............................... False moe_expert_capacity_factor ...................... None moe_extended_tp ................................. False moe_ffn_hidden_size ............................. 118 moe_grouped_gemm ................................ True moe_input_jitter_eps ............................ None moe_layer_freq .................................. [0, 1, 1, 1, 1, 1, 1, 1, 1] moe_layer_recompute ............................. False moe_pad_expert_input_to_capacity ................ False moe_per_layer_logging ........................... False moe_permute_fusion .............................. False moe_router_bias_update_rate ..................... 0.001 moe_router_dtype ................................ fp32 moe_router_enable_expert_bias ................... False moe_router_group_topk ........................... None moe_router_load_balancing_type .................. aux_loss moe_router_num_groups ........................... None moe_router_pre_softmax .......................... True moe_router_score_function ....................... softmax moe_router_topk ................................. 18 moe_router_topk_scaling_factor .................. None moe_shared_expert_intermediate_size ............. 704 moe_shared_expert_overlap ....................... False moe_token_dispatcher_type ....................... alltoall moe_token_drop_policy ........................... probs moe_use_legacy_grouped_gemm ..................... False moe_use_upcycling ............................... False moe_z_loss_coeff ................................ 0.001 mscale .......................................... 1.0 mscale_all_dim .................................. 1.0 multi_latent_attention .......................... False nccl_communicator_config_path ................... None no_load_optim ................................... True no_load_rng ..................................... True no_persist_layer_norm ........................... False no_save_optim ................................... None no_save_rng ..................................... None non_persistent_ckpt_type ........................ None non_persistent_global_ckpt_dir .................. None non_persistent_local_ckpt_algo .................. fully_parallel non_persistent_local_ckpt_dir ................... None non_persistent_save_interval .................... None norm_epsilon .................................... 1e-06 normalization ................................... RMSNorm num_attention_heads ............................. 32 num_channels .................................... 3 num_classes ..................................... 1000 num_dataset_builder_threads ..................... 1 num_distributed_optimizer_instances ............. 1 num_experts ..................................... 192 num_layers ...................................... 9 num_layers_per_virtual_pipeline_stage ........... None num_query_groups ................................ 1 num_virtual_stages_per_pipeline_rank ............ None num_workers ..................................... 2 one_logger_async ................................ False one_logger_project .............................. megatron-lm one_logger_run_name ............................. None onnx_safe ....................................... None openai_gelu ..................................... False optimizer ....................................... adam optimizer_cpu_offload ........................... False optimizer_offload_fraction ...................... 1.0 output_bert_embeddings .......................... False overlap_cpu_optimizer_d2h_h2d ................... False overlap_grad_reduce ............................. False overlap_p2p_comm ................................ False overlap_p2p_comm_warmup_flush ................... False overlap_param_gather ............................ False overlap_param_gather_with_optimizer_step ........ False override_opt_param_scheduler .................... False params_dtype .................................... torch.bfloat16 patch_dim ....................................... 16 per_split_data_args_path ........................ None perform_initialization .......................... True pin_cpu_grads ................................... True pin_cpu_params .................................. True pipeline_model_parallel_comm_backend ............ None pipeline_model_parallel_size .................... 1 pipeline_model_parallel_split_rank .............. None position_embedding_type ......................... rope pretrained_checkpoint ........................... None profile ......................................... False profile_ranks ................................... [0] profile_step_end ................................ 12 profile_step_start .............................. 10 q_lora_rank ..................................... None qk_head_dim ..................................... 128 qk_layernorm .................................... False qk_pos_emb_head_dim ............................. 64 query_in_block_prob ............................. 0.1 rampup_batch_size ............................... None rank ............................................ 0 recompute_granularity ........................... None recompute_method ................................ None recompute_num_layers ............................ None record_memory_history ........................... False relative_attention_max_distance ................. 128 relative_attention_num_buckets .................. 32 replication ..................................... False replication_factor .............................. 2 replication_jump ................................ None rerun_mode ...................................... disabled reset_attention_mask ............................ False reset_position_ids .............................. False result_rejected_tracker_filename ................ None retriever_report_topk_accuracies ................ [] retriever_score_scaling ......................... False retriever_seq_length ............................ 256 retro_add_retriever ............................. False retro_attention_gate ............................ 1 retro_cyclic_train_iters ........................ None retro_encoder_attention_dropout ................. 0.1 retro_encoder_hidden_dropout .................... 0.1 retro_encoder_layers ............................ 2 retro_num_neighbors ............................. 2 retro_num_retrieved_chunks ...................... 2 retro_project_dir ............................... None retro_verify_neighbor_count ..................... True rope_scaling_factor ............................. 8.0 rotary_base ..................................... 10000 rotary_interleaved .............................. False rotary_percent .................................. 1.0 rotary_scaling_factor ........................... 1.0 rotary_seq_len_interpolation_factor ............. None s3_cache_path ................................... None sample_rate ..................................... 1.0 save ............................................ None save_interval ................................... None scatter_gather_tensors_in_pipeline .............. True seed ............................................ 42 seq_length ...................................... 2048 sequence_parallel ............................... False sgd_momentum .................................... 0.9 short_seq_prob .................................. 0.1 skip_train ...................................... False skipped_train_samples ........................... 0 spec ............................................ None split ........................................... None squared_relu .................................... False start_weight_decay .............................. 0.01 straggler_ctrlr_port ............................ 65535 straggler_minmax_count .......................... 1 suggested_communication_unit_size ............... 400000000 swiglu .......................................... True swin_backbone_type .............................. tiny te_rng_tracker .................................. False tensor_model_parallel_size ...................... 1 tensorboard_dir ................................. None tensorboard_log_interval ........................ 1 tensorboard_queue_size .......................... 1000 test_data_path .................................. None test_mode ....................................... False tiktoken_num_special_tokens ..................... 1000 tiktoken_pattern ................................ None tiktoken_special_tokens ......................... None timing_log_level ................................ 0 timing_log_option ............................... minmax titles_data_path ................................ None tokenizer_model ................................. EleutherAI/pythia-12b tokenizer_type .................................. HuggingFaceTokenizer tp_comm_bootstrap_backend ....................... nccl tp_comm_bulk_dgrad .............................. True tp_comm_bulk_wgrad .............................. True tp_comm_overlap ................................. False tp_comm_overlap_ag .............................. True tp_comm_overlap_cfg ............................. None tp_comm_overlap_rs .............................. True tp_comm_overlap_rs_dgrad ........................ False tp_comm_split_ag ................................ True tp_comm_split_rs ................................ True train_data_path ................................. None train_iters ..................................... None train_samples ................................... None train_sync_interval ............................. None transformer_impl ................................ transformer_engine transformer_pipeline_model_parallel_size ........ 1 untie_embeddings_and_output_weights ............. True use_checkpoint_args ............................. False use_checkpoint_opt_param_scheduler .............. False use_cpu_initialization .......................... None use_custom_fsdp ................................. False use_dist_ckpt ................................... True use_dist_ckpt_deprecated ........................ False use_distributed_optimizer ....................... True use_flash_attn .................................. False use_legacy_models ............................... False use_mp_args_from_checkpoint_args ................ False use_one_sent_docs ............................... False use_persistent_ckpt_worker ...................... False use_precision_aware_optimizer ................... False use_pytorch_profiler ............................ False use_ring_exchange_p2p ........................... False use_rope_scaling ................................ False use_rotary_position_embeddings .................. False use_tokenizer_model_from_checkpoint_args ........ True use_torch_fsdp2 ................................. False use_torch_optimizer_for_cpu_offload ............. False use_tp_pp_dp_mapping ............................ False v_head_dim ...................................... 128 valid_data_path ................................. None variable_seq_lengths ............................ False virtual_pipeline_model_parallel_size ............ None vision_backbone_type ............................ vit vision_pretraining .............................. False vision_pretraining_type ......................... classify vocab_extra_ids ................................. 0 vocab_file ...................................... None vocab_size ...................................... None wandb_exp_name .................................. wandb_project ................................... wandb_save_dir .................................. weight_decay .................................... 0.01 weight_decay_incr_style ......................... constant wgrad_deferral_limit ............................ 0 world_size ...................................... 1 yaml_cfg ........................................ None -------------------- end of arguments --------------------- 2026-07-23:05:21:41,451 INFO [num_microbatches_calculator.py:228] setting number of microbatches to constant 1 > building HuggingFaceTokenizer tokenizer ... 2026-07-23:05:21:41,664 INFO [_client.py:1038] HTTP Request: HEAD https://huggingface.co/EleutherAI/pythia-12b/resolve/main/config.json "HTTP/1.1 307 Temporary Redirect" 2026-07-23:05:21:41,677 INFO [_client.py:1038] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/EleutherAI/pythia-12b/bb1e3e710cdf6b524461d543cfb5ba773f0a81b6/config.json "HTTP/1.1 200 OK" 2026-07-23:05:21:41,795 INFO [_client.py:1038] HTTP Request: HEAD https://huggingface.co/EleutherAI/pythia-12b/resolve/main/tokenizer_config.json "HTTP/1.1 307 Temporary Redirect" 2026-07-23:05:21:41,808 INFO [_client.py:1038] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/EleutherAI/pythia-12b/bb1e3e710cdf6b524461d543cfb5ba773f0a81b6/tokenizer_config.json "HTTP/1.1 200 OK" 2026-07-23:05:21:41,902 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/models/EleutherAI/pythia-12b/tree/main/additional_chat_templates?recursive=false&expand=false "HTTP/1.1 404 Not Found" 2026-07-23:05:21:41,971 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/models/EleutherAI/pythia-12b/tree/main?recursive=true&expand=false "HTTP/1.1 200 OK" > padded vocab (size: 50277) with 27 dummy tokens (new size: 50304) WARNING: one_logger package is required to enable e2e metrics tracking. please go to https://confluence.nvidia.com/display/MLWFO/Package+Repositories for details to install it 2026-07-23:05:21:42,172 WARNING [rerun_state_machine.py:239] RerunStateMachine initialized in mode RerunMode.DISABLED > initializing torch distributed ... [W723 05:21:42.689962719 CUDAAllocatorConfig.h:28] Warning: expandable_segments not supported on this platform (function operator()) > initialized tensor model parallel with size 1 > initialized pipeline model parallel with size 1 > setting random seeds to 42 ... > compiling dataset index builder ... make: Entering directory '/workspace/FLAME-MoE/Megatron-LM/megatron/core/datasets' make: Nothing to be done for 'default'. make: Leaving directory '/workspace/FLAME-MoE/Megatron-LM/megatron/core/datasets' >>> done with dataset index builder. Compilation time: 0.755 seconds > compiling and loading fused kernels ... >>> done with compiling and loading fused kernels. Compilation time: 0.355 seconds building GPT model ... > number of parameters on (tensor, pipeline) model parallel rank (0, 0): 352994816 /workspace/FLAME-MoE/Megatron-LM/megatron/core/transformer/transformer_layer.py:339: UserWarning: TransformerLayer._get_layer_offset is deprecated.Please use get_transformer_layer_offset instead. warnings.warn( loading distributed checkpoint from /workspace/FLAME-MoE/results/phase0/runs/flame512_g3_moe/ckpt at iteration 802 /workspace/FLAME-MoE/Megatron-LM/megatron/core/dist_checkpointing/strategies/torch.py:847: FutureWarning: `load_state_dict` is deprecated and will be removed in future versions. Please use `load` instead. checkpoint.load_state_dict( /usr/local/lib/python3.11/dist-packages/torch/distributed/checkpoint/planner_helpers.py:306: FutureWarning: Please use DTensor instead and we are deprecating ShardedTensor. device = getattr(value, "device", None) /workspace/FLAME-MoE/.venv/lib/python3.11/site-packages/transformer_engine/pytorch/module/base.py:566: FutureWarning: You are using `torch.load` with `weights_only=False` (the current default value), which uses the default pickle module implicitly. It is possible to construct malicious pickle data which will execute arbitrary code during unpickling (See https://github.com/pytorch/pytorch/blob/main/SECURITY.md#untrusted-models for more details). In a future release, the default value for `weights_only` will be flipped to `True`. This limits the functions that could be executed during unpickling. Arbitrary objects will no longer be allowed to be loaded via this mode unless they are explicitly allowlisted by the user via `torch.serialization.add_safe_globals`. We recommend you start setting `weights_only=True` for any use case where you don't have full control of the loaded file. Please open an issue on GitHub for any issues related to this experimental feature. state = torch.load(state, map_location="cuda") checkpoint version 3.0 2026-07-23:05:21:53,001 WARNING [rerun_state_machine.py:811] RerunStateMachine disabled via CLI, ignoring machine state saved in checkpoint successfully loaded checkpoint from /workspace/FLAME-MoE/results/phase0/runs/flame512_g3_moe/ckpt [ t 1/1, p 1/1 ] at iteration 802 2026-07-23:05:21:53,159 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc "HTTP/1.1 200 OK" 2026-07-23:05:21:53,403 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc "HTTP/1.1 200 OK" 2026-07-23:05:21:53,470 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc/revision/210d026faf9955653af8916fad021475a3f00453 "HTTP/1.1 200 OK" 2026-07-23:05:21:53,528 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc/tree/210d026faf9955653af8916fad021475a3f00453?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:21:53,858 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc/tree/210d026faf9955653af8916fad021475a3f00453/ARC-Challenge?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:21:54,388 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc "HTTP/1.1 200 OK" 2026-07-23:05:21:54,613 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc "HTTP/1.1 200 OK" 2026-07-23:05:21:54,861 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc/tree/210d026faf9955653af8916fad021475a3f00453/ARC-Easy?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:21:55,003 WARNING [task.py:799] [Task: boolq] metric acc is defined, but aggregation is not. using default aggregation=mean 2026-07-23:05:21:55,004 WARNING [task.py:811] [Task: boolq] metric acc is defined, but higher_is_better is not. using default higher_is_better=True 2026-07-23:05:21:55,059 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue "HTTP/1.1 200 OK" 2026-07-23:05:21:55,311 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue "HTTP/1.1 200 OK" 2026-07-23:05:21:55,391 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/revision/3de24cf8022e94f4ee4b9d55a6f539891524d646 "HTTP/1.1 200 OK" 2026-07-23:05:21:55,456 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/tree/3de24cf8022e94f4ee4b9d55a6f539891524d646?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:21:56,065 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/tree/3de24cf8022e94f4ee4b9d55a6f539891524d646/axb?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:21:56,125 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/revision/3de24cf8022e94f4ee4b9d55a6f539891524d646 "HTTP/1.1 200 OK" 2026-07-23:05:21:56,195 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/tree/3de24cf8022e94f4ee4b9d55a6f539891524d646/boolq?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:21:56,393 WARNING [task.py:799] [Task: copa] metric acc is defined, but aggregation is not. using default aggregation=mean 2026-07-23:05:21:56,393 WARNING [task.py:811] [Task: copa] metric acc is defined, but higher_is_better is not. using default higher_is_better=True 2026-07-23:05:21:56,454 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue "HTTP/1.1 200 OK" 2026-07-23:05:21:56,680 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue "HTTP/1.1 200 OK" 2026-07-23:05:21:56,995 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/tree/3de24cf8022e94f4ee4b9d55a6f539891524d646/copa?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:21:57,139 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag "HTTP/1.1 200 OK" 2026-07-23:05:21:57,357 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag "HTTP/1.1 200 OK" 2026-07-23:05:21:57,414 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag/revision/218ec52e09a7e7462a5400043bb9a69a41d06b76 "HTTP/1.1 200 OK" 2026-07-23:05:21:57,475 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag/tree/218ec52e09a7e7462a5400043bb9a69a41d06b76?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:21:57,668 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag/tree/218ec52e09a7e7462a5400043bb9a69a41d06b76/data?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:21:57,727 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag/revision/218ec52e09a7e7462a5400043bb9a69a41d06b76 "HTTP/1.1 200 OK" 2026-07-23:05:21:59,787 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai "HTTP/1.1 200 OK" 2026-07-23:05:22:00,013 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai "HTTP/1.1 200 OK" 2026-07-23:05:22:00,069 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/revision/900124bf3b8235c6daf21033af9948b3f07346c4 "HTTP/1.1 200 OK" 2026-07-23:05:22:00,132 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/tree/900124bf3b8235c6daf21033af9948b3f07346c4?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:22:00,358 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/tree/900124bf3b8235c6daf21033af9948b3f07346c4/default%2Ftest?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:22:00,418 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/tree/900124bf3b8235c6daf21033af9948b3f07346c4/default?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:22:00,479 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/revision/900124bf3b8235c6daf21033af9948b3f07346c4 "HTTP/1.1 200 OK" 2026-07-23:05:22:00,498 WARNING [task.py:325] [Task: lambada_openai] has_training_docs and has_validation_docs are False, using test_docs as fewshot_docs but this is not recommended. 2026-07-23:05:22:00,499 WARNING [task.py:325] [Task: lambada_openai] has_training_docs and has_validation_docs are False, using test_docs as fewshot_docs but this is not recommended. 2026-07-23:05:22:00,661 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa "HTTP/1.1 200 OK" 2026-07-23:05:22:00,883 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa "HTTP/1.1 200 OK" 2026-07-23:05:22:00,955 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/revision/388097ea7776314e93a529163e0fea805b8a6454 "HTTP/1.1 200 OK" 2026-07-23:05:22:01,017 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/tree/388097ea7776314e93a529163e0fea805b8a6454?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:22:01,224 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/tree/388097ea7776314e93a529163e0fea805b8a6454/additional?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:22:01,402 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/revision/388097ea7776314e93a529163e0fea805b8a6454 "HTTP/1.1 200 OK" 2026-07-23:05:22:01,480 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/tree/388097ea7776314e93a529163e0fea805b8a6454/main?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:22:01,739 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa "HTTP/1.1 200 OK" 2026-07-23:05:22:01,792 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa "HTTP/1.1 200 OK" 2026-07-23:05:22:01,848 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa/tree/main?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:22:01,907 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/datasets/ybisk/piqa/resolve/main/piqa.py "HTTP/1.1 307 Temporary Redirect" 2026-07-23:05:22:01,919 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/resolve-cache/datasets/ybisk/piqa/2e8ac2dffd59bac8c3c6714948f4c551a0848bb0/piqa.py?%2Fdatasets%2Fybisk%2Fpiqa%2Fresolve%2Fmain%2Fpiqa.py=&etag=%220f0d96172f9bf7aa10069374b3236ef083cfeb8e%22 "HTTP/1.1 200 OK" 2026-07-23:05:22:02,447 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa/revision/main "HTTP/1.1 200 OK" 2026-07-23:05:22:02,504 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa/tree/main?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:22:02,918 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq "HTTP/1.1 200 OK" 2026-07-23:05:22:03,175 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq "HTTP/1.1 200 OK" 2026-07-23:05:22:03,233 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq/revision/2c94ad3e1aafab77146f384e23536f97a4849815 "HTTP/1.1 200 OK" 2026-07-23:05:22:03,293 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq/tree/2c94ad3e1aafab77146f384e23536f97a4849815?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:22:03,485 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq/tree/2c94ad3e1aafab77146f384e23536f97a4849815/data?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:22:03,551 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq/revision/2c94ad3e1aafab77146f384e23536f97a4849815 "HTTP/1.1 200 OK" 2026-07-23:05:22:03,964 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande "HTTP/1.1 200 OK" 2026-07-23:05:22:04,196 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande "HTTP/1.1 200 OK" 2026-07-23:05:22:04,254 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/revision/01e74176c63542e6b0bcb004dcdea22d94fb67b5 "HTTP/1.1 200 OK" 2026-07-23:05:22:04,366 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/tree/01e74176c63542e6b0bcb004dcdea22d94fb67b5?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:22:04,593 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/tree/01e74176c63542e6b0bcb004dcdea22d94fb67b5/winogrande_debiased?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:22:04,652 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/revision/01e74176c63542e6b0bcb004dcdea22d94fb67b5 "HTTP/1.1 200 OK" 2026-07-23:05:22:04,730 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/tree/01e74176c63542e6b0bcb004dcdea22d94fb67b5/winogrande_xl?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:22:05,368 WARNING [evaluator.py:270] Overwriting default num_fewshot of winogrande from None to 0 2026-07-23:05:22:05,369 WARNING [evaluator.py:270] Overwriting default num_fewshot of sciq from None to 0 2026-07-23:05:22:05,369 WARNING [evaluator.py:270] Overwriting default num_fewshot of piqa from None to 0 2026-07-23:05:22:05,369 WARNING [evaluator.py:270] Overwriting default num_fewshot of openbookqa from None to 0 2026-07-23:05:22:05,369 WARNING [evaluator.py:270] Overwriting default num_fewshot of lambada_openai from None to 0 2026-07-23:05:22:05,369 WARNING [evaluator.py:270] Overwriting default num_fewshot of hellaswag from None to 0 2026-07-23:05:22:05,369 WARNING [evaluator.py:270] Overwriting default num_fewshot of copa from None to 0 2026-07-23:05:22:05,369 WARNING [evaluator.py:270] Overwriting default num_fewshot of boolq from None to 0 2026-07-23:05:22:05,369 WARNING [evaluator.py:270] Overwriting default num_fewshot of arc_easy from None to 0 2026-07-23:05:22:05,369 WARNING [evaluator.py:270] Overwriting default num_fewshot of arc_challenge from None to 0 2026-07-23:05:22:05,399 INFO [task.py:415] Building contexts for winogrande on rank 0... 0%| | 0/1267 [00:00 /workspace/FLAME-MoE/results/ablations/flame1e18_downstream.csv ==================================================================== [run] flame512_g3_temporal mode=temporal TEMPORAL=1 H=512 heads=32 ffn=2736 experts=192 topk=18 moe_ffn=118 shared=704 [run] tasks=arc_challenge,arc_easy,hellaswag,openbookqa,piqa,winogrande,boolq,copa,lambada_openai,sciq out=/workspace/FLAME-MoE/results/phase0/runs/flame512_g3_temporal/lmeval_1e18_0shot ==================================================================== /usr/local/lib/python3.11/dist-packages/torch/cuda/__init__.py:58: FutureWarning: The pynvml package is deprecated. Please install nvidia-ml-py instead. If you did not install pynvml directly, please report this to the maintainers of the package that installed pynvml for you. import pynvml # type: ignore[import] [run_lmeval] stubbed VLM backends (hf_vlms, vllm_vlms) /usr/local/lib/python3.11/dist-packages/torch/cuda/__init__.py:58: FutureWarning: The pynvml package is deprecated. Please install nvidia-ml-py instead. If you did not install pynvml directly, please report this to the maintainers of the package that installed pynvml for you. import pynvml # type: ignore[import] /usr/local/lib/python3.11/dist-packages/torch/cuda/__init__.py:58: FutureWarning: The pynvml package is deprecated. Please install nvidia-ml-py instead. If you did not install pynvml directly, please report this to the maintainers of the package that installed pynvml for you. import pynvml # type: ignore[import] /workspace/FLAME-MoE/.venv/lib/python3.11/site-packages/transformer_engine/pytorch/attention.py:108: UserWarning: To use flash-attn v3, please use the following commands to install: (1) pip install "git+https://github.com/Dao-AILab/flash-attention.git#egg=flashattn-hopper&subdirectory=hopper" (2) python_path=`python -c "import site; print(site.getsitepackages()[0])"` (3) mkdir -p $python_path/flashattn_hopper (4) wget -P $python_path/flashattn_hopper https://raw.githubusercontent.com/Dao-AILab/flash-attention/main/hopper/flash_attn_interface.py warnings.warn( [temporal] rolling-residency router installed (evict=min_logit) [run_lmeval] installed temporal residency router (native regime) /usr/local/lib/python3.11/dist-packages/requests/__init__.py:113: RequestsDependencyWarning: urllib3 (2.2.3) or chardet (6.0.0.post1)/charset_normalizer (3.3.2) doesn't match a supported version! warnings.warn( 2026-07-23:05:31:46,492 INFO [__main__.py:279] Verbosity set to INFO 2026-07-23:05:32:16,965 INFO [__init__.py:459] The tag 'arc_ca' is already registered as a group, this tag will not be registered. This may affect tasks you want to call. 2026-07-23:05:32:16,980 INFO [__init__.py:459] The tag 'arc_ca' is already registered as a group, this tag will not be registered. This may affect tasks you want to call. 2026-07-23:05:32:33,156 INFO [__main__.py:376] Selected Tasks: ['arc_challenge', 'arc_easy', 'boolq', 'copa', 'hellaswag', 'lambada_openai', 'openbookqa', 'piqa', 'sciq', 'winogrande'] 2026-07-23:05:32:33,191 INFO [evaluator.py:164] Setting random seed to 42 | Setting numpy seed to 42 | Setting torch manual seed to 42 | Setting fewshot manual seed to 42 2026-07-23:05:32:33,192 INFO [evaluator.py:201] Initializing megatron_lm model, with arguments: {} using world size: 1, data-parallel size: 1, context-parallel size: 1, hierarchical context-parallel sizes: Nonetensor-model-parallel size: 1, encoder-tensor-model-parallel size: 0, pipeline-model-parallel size: 1, encoder-pipeline-model-parallel size: 0 WARNING: overriding default arguments for exit_on_missing_checkpoint:True with exit_on_missing_checkpoint:False setting global batch size to 32 Number of virtual stages per pipeline stage: None accumulate and all-reduce gradients in fp32 for bfloat16 data type. using torch.bfloat16 for parameters ... ------------------------ arguments ------------------------ account_for_embedding_in_pipeline_split ......... False account_for_loss_in_pipeline_split .............. False accumulate_allreduce_grads_in_fp32 .............. True adam_beta1 ...................................... 0.9 adam_beta2 ...................................... 0.999 adam_eps ........................................ 1e-08 add_bias_linear ................................. False add_position_embedding .......................... True add_qkv_bias .................................... False adlr_autoresume ................................. False adlr_autoresume_interval ........................ 1000 align_grad_reduce ............................... True align_param_gather .............................. False app_tag_run_name ................................ None app_tag_run_version ............................. 0.0.0 apply_layernorm_1p .............................. False apply_query_key_layer_scaling ................... False apply_residual_connection_post_layernorm ........ False apply_rope_fusion ............................... True async_save ...................................... None async_tensor_model_parallel_allreduce ........... True attention_backend ............................... AttnBackend.auto attention_dropout ............................... 0.0 attention_softmax_in_fp32 ....................... False auto_detect_ckpt_format ......................... False barrier_with_L1_time ............................ True bert_binary_head ................................ True bert_embedder_type .............................. megatron bert_load ....................................... None bf16 ............................................ True bias_dropout_fusion ............................. True bias_gelu_fusion ................................ False bias_swiglu_fusion .............................. True biencoder_projection_dim ........................ 0 biencoder_shared_query_context_model ............ False block_data_path ................................. None calc_ft_timeouts ................................ False calculate_per_token_loss ........................ False check_for_large_grads ........................... False check_for_nan_in_loss_and_grad .................. True check_for_spiky_loss ............................ False check_weight_hash_across_dp_replicas_interval ... None ckpt_assume_constant_structure .................. False ckpt_convert_format ............................. None ckpt_convert_save ............................... None ckpt_convert_update_legacy_dist_opt_format ...... False ckpt_format ..................................... torch_dist ckpt_fully_parallel_load ........................ False ckpt_fully_parallel_save ........................ True ckpt_fully_parallel_save_deprecated ............. False ckpt_step ....................................... None classes_fraction ................................ 1.0 clip_grad ....................................... 1.0 clone_scatter_output_in_embedding ............... True config_logger_dir ............................... consumed_train_samples .......................... 0 consumed_valid_samples .......................... 0 context_parallel_size ........................... 1 cp_comm_type .................................... ['p2p'] create_attention_mask_in_dataloader ............. True cross_entropy_loss_fusion ....................... False cuda_graph_warmup_steps ......................... 3 data_args_path .................................. None data_cache_path ................................. None data_parallel_random_init ....................... False data_parallel_sharding_strategy ................. no_shard data_parallel_size .............................. 1 data_path ....................................... None data_per_class_fraction ......................... 1.0 data_sharding ................................... True dataloader_type ................................. single ddp_average_in_collective ....................... False ddp_bucket_size ................................. None ddp_num_buckets ................................. None ddp_pad_buckets_for_high_nccl_busbw ............. False decoder_first_pipeline_num_layers ............... None decoder_last_pipeline_num_layers ................ None decoder_num_layers .............................. None decoder_seq_length .............................. None decoupled_lr .................................... None decoupled_min_lr ................................ None decrease_batch_size_if_needed ................... False defer_embedding_wgrad_compute ................... False deprecated_use_mcore_models ..................... False deterministic_mode .............................. False dino_bottleneck_size ............................ 256 dino_freeze_last_layer .......................... 1 dino_head_hidden_size ........................... 2048 dino_local_crops_number ......................... 10 dino_local_img_size ............................. 96 dino_norm_last_layer ............................ False dino_teacher_temp ............................... 0.07 dino_warmup_teacher_temp ........................ 0.04 dino_warmup_teacher_temp_epochs ................. 30 disable_straggler_on_startup .................... False dist_ckpt_format_deprecated ..................... None dist_ckpt_strictness ............................ assume_ok_unexpected distribute_saved_activations .................... False distributed_backend ............................. nccl distributed_timeout_minutes ..................... 10 embedding_path .................................. None empty_unused_memory_level ....................... 0 enable_cuda_graph ............................... False enable_ft_package ............................... False enable_gloo_process_groups ...................... True enable_one_logger ............................... True encoder_num_layers .............................. 9 encoder_pipeline_model_parallel_size ............ 0 encoder_seq_length .............................. 2048 encoder_tensor_model_parallel_size .............. 0 end_weight_decay ................................ 0.01 eod_mask_loss ................................... False error_injection_rate ............................ 0 error_injection_type ............................ transient_error eval_interval ................................... 1000 eval_iters ...................................... 100 evidence_data_path .............................. None exit_duration_in_mins ........................... None exit_interval ................................... None exit_on_missing_checkpoint ...................... False exit_signal_handler ............................. False exp_avg_dtype ................................... torch.float32 exp_avg_sq_dtype ................................ torch.float32 expert_model_parallel_size ...................... 1 expert_tensor_parallel_size ..................... 1 ffn_hidden_size ................................. 2736 finetune ........................................ False flash_decode .................................... False fp16 ............................................ False fp16_lm_cross_entropy ........................... False fp32_residual_connection ........................ False fp8 ............................................. None fp8_amax_compute_algo ........................... most_recent fp8_amax_history_len ............................ 1 fp8_interval .................................... 1 fp8_margin ...................................... 0 fp8_param_gather ................................ False fp8_wgrad ....................................... True global_batch_size ............................... 32 grad_reduce_in_bf16 ............................. False gradient_accumulation_fusion .................... False gradient_reduce_div_fusion ...................... True group_query_attention ........................... False head_lr_mult .................................... 1.0 hidden_dropout .................................. 0.0 hidden_size ..................................... 512 hierarchical_context_parallel_sizes ............. None hybrid_attention_ratio .......................... 0.0 hybrid_mlp_ratio ................................ 0.0 hybrid_override_pattern ......................... None hysteresis ...................................... 2 ict_head_size ................................... None ict_load ........................................ None img_h ........................................... 224 img_w ........................................... 224 indexer_batch_size .............................. 128 indexer_log_interval ............................ 1000 inference_batch_times_seqlen_threshold .......... -1 inference_max_batch_size ........................ 8 inference_max_seq_length ........................ 2560 inference_rng_tracker ........................... False init_method_std ................................. 0.02 init_method_xavier_uniform ...................... False init_model_with_meta_device ..................... False initial_loss_scale .............................. 4294967296 iter_per_epoch .................................. 1250 iterations_to_skip .............................. [] keep_fp8_transpose_cache_when_using_custom_fsdp . False kv_channels ..................................... 16 kv_lora_rank .................................... 32 lazy_mpu_init ................................... None load ............................................ /workspace/FLAME-MoE/results/phase0/runs/flame512_g3_temporal/ckpt local_rank ...................................... 0 log_interval .................................... 100 log_loss_scale_to_tensorboard ................... True log_memory_to_tensorboard ....................... False log_num_zeros_in_grad ........................... False log_params_norm ................................. False log_progress .................................... False log_straggler ................................... False log_throughput .................................. False log_timers_to_tensorboard ....................... False log_validation_ppl_to_tensorboard ............... False log_world_size_to_tensorboard ................... False logging_level ................................... None loss_scale ...................................... None loss_scale_window ............................... 1000 lr .............................................. None lr_decay_iters .................................. None lr_decay_samples ................................ None lr_decay_style .................................. linear lr_warmup_fraction .............................. None lr_warmup_init .................................. 0.0 lr_warmup_iters ................................. 0 lr_warmup_samples ............................... 0 lr_wsd_decay_iters .............................. None lr_wsd_decay_samples ............................ None lr_wsd_decay_style .............................. exponential main_grads_dtype ................................ torch.float32 main_params_dtype ............................... torch.float32 make_vocab_size_divisible_by .................... 128 manual_gc ....................................... False manual_gc_eval .................................. True manual_gc_interval .............................. 0 mask_factor ..................................... 1.0 mask_prob ....................................... 0.15 mask_type ....................................... random masked_softmax_fusion ........................... True max_position_embeddings ......................... 2048 max_tokens_to_oom ............................... 10000000 memory_snapshot_path ............................ snapshot.pickle merge_file ...................................... None micro_batch_size ................................ 32 microbatch_group_size_per_vp_stage .............. None min_loss_scale .................................. 1.0 min_lr .......................................... 0.0 mmap_bin_files .................................. True mock_data ....................................... False moe_aux_loss_coeff .............................. 0.01 moe_enable_deepep ............................... False moe_expert_capacity_factor ...................... None moe_extended_tp ................................. False moe_ffn_hidden_size ............................. 118 moe_grouped_gemm ................................ True moe_input_jitter_eps ............................ None moe_layer_freq .................................. [0, 1, 1, 1, 1, 1, 1, 1, 1] moe_layer_recompute ............................. False moe_pad_expert_input_to_capacity ................ False moe_per_layer_logging ........................... False moe_permute_fusion .............................. False moe_router_bias_update_rate ..................... 0.001 moe_router_dtype ................................ fp32 moe_router_enable_expert_bias ................... False moe_router_group_topk ........................... None moe_router_load_balancing_type .................. aux_loss moe_router_num_groups ........................... None moe_router_pre_softmax .......................... True moe_router_score_function ....................... softmax moe_router_topk ................................. 18 moe_router_topk_scaling_factor .................. None moe_shared_expert_intermediate_size ............. 704 moe_shared_expert_overlap ....................... False moe_token_dispatcher_type ....................... alltoall moe_token_drop_policy ........................... probs moe_use_legacy_grouped_gemm ..................... False moe_use_upcycling ............................... False moe_z_loss_coeff ................................ 0.001 mscale .......................................... 1.0 mscale_all_dim .................................. 1.0 multi_latent_attention .......................... False nccl_communicator_config_path ................... None no_load_optim ................................... True no_load_rng ..................................... True no_persist_layer_norm ........................... False no_save_optim ................................... None no_save_rng ..................................... None non_persistent_ckpt_type ........................ None non_persistent_global_ckpt_dir .................. None non_persistent_local_ckpt_algo .................. fully_parallel non_persistent_local_ckpt_dir ................... None non_persistent_save_interval .................... None norm_epsilon .................................... 1e-06 normalization ................................... RMSNorm num_attention_heads ............................. 32 num_channels .................................... 3 num_classes ..................................... 1000 num_dataset_builder_threads ..................... 1 num_distributed_optimizer_instances ............. 1 num_experts ..................................... 192 num_layers ...................................... 9 num_layers_per_virtual_pipeline_stage ........... None num_query_groups ................................ 1 num_virtual_stages_per_pipeline_rank ............ None num_workers ..................................... 2 one_logger_async ................................ False one_logger_project .............................. megatron-lm one_logger_run_name ............................. None onnx_safe ....................................... None openai_gelu ..................................... False optimizer ....................................... adam optimizer_cpu_offload ........................... False optimizer_offload_fraction ...................... 1.0 output_bert_embeddings .......................... False overlap_cpu_optimizer_d2h_h2d ................... False overlap_grad_reduce ............................. False overlap_p2p_comm ................................ False overlap_p2p_comm_warmup_flush ................... False overlap_param_gather ............................ False overlap_param_gather_with_optimizer_step ........ False override_opt_param_scheduler .................... False params_dtype .................................... torch.bfloat16 patch_dim ....................................... 16 per_split_data_args_path ........................ None perform_initialization .......................... True pin_cpu_grads ................................... True pin_cpu_params .................................. True pipeline_model_parallel_comm_backend ............ None pipeline_model_parallel_size .................... 1 pipeline_model_parallel_split_rank .............. None position_embedding_type ......................... rope pretrained_checkpoint ........................... None profile ......................................... False profile_ranks ................................... [0] profile_step_end ................................ 12 profile_step_start .............................. 10 q_lora_rank ..................................... None qk_head_dim ..................................... 128 qk_layernorm .................................... False qk_pos_emb_head_dim ............................. 64 query_in_block_prob ............................. 0.1 rampup_batch_size ............................... None rank ............................................ 0 recompute_granularity ........................... None recompute_method ................................ None recompute_num_layers ............................ None record_memory_history ........................... False relative_attention_max_distance ................. 128 relative_attention_num_buckets .................. 32 replication ..................................... False replication_factor .............................. 2 replication_jump ................................ None rerun_mode ...................................... disabled reset_attention_mask ............................ False reset_position_ids .............................. False result_rejected_tracker_filename ................ None retriever_report_topk_accuracies ................ [] retriever_score_scaling ......................... False retriever_seq_length ............................ 256 retro_add_retriever ............................. False retro_attention_gate ............................ 1 retro_cyclic_train_iters ........................ None retro_encoder_attention_dropout ................. 0.1 retro_encoder_hidden_dropout .................... 0.1 retro_encoder_layers ............................ 2 retro_num_neighbors ............................. 2 retro_num_retrieved_chunks ...................... 2 retro_project_dir ............................... None retro_verify_neighbor_count ..................... True rope_scaling_factor ............................. 8.0 rotary_base ..................................... 10000 rotary_interleaved .............................. False rotary_percent .................................. 1.0 rotary_scaling_factor ........................... 1.0 rotary_seq_len_interpolation_factor ............. None s3_cache_path ................................... None sample_rate ..................................... 1.0 save ............................................ None save_interval ................................... None scatter_gather_tensors_in_pipeline .............. True seed ............................................ 42 seq_length ...................................... 2048 sequence_parallel ............................... False sgd_momentum .................................... 0.9 short_seq_prob .................................. 0.1 skip_train ...................................... False skipped_train_samples ........................... 0 spec ............................................ None split ........................................... None squared_relu .................................... False start_weight_decay .............................. 0.01 straggler_ctrlr_port ............................ 65535 straggler_minmax_count .......................... 1 suggested_communication_unit_size ............... 400000000 swiglu .......................................... True swin_backbone_type .............................. tiny te_rng_tracker .................................. False tensor_model_parallel_size ...................... 1 tensorboard_dir ................................. None tensorboard_log_interval ........................ 1 tensorboard_queue_size .......................... 1000 test_data_path .................................. None test_mode ....................................... False tiktoken_num_special_tokens ..................... 1000 tiktoken_pattern ................................ None tiktoken_special_tokens ......................... None timing_log_level ................................ 0 timing_log_option ............................... minmax titles_data_path ................................ None tokenizer_model ................................. EleutherAI/pythia-12b tokenizer_type .................................. HuggingFaceTokenizer tp_comm_bootstrap_backend ....................... nccl tp_comm_bulk_dgrad .............................. True tp_comm_bulk_wgrad .............................. True tp_comm_overlap ................................. False tp_comm_overlap_ag .............................. True tp_comm_overlap_cfg ............................. None tp_comm_overlap_rs .............................. True tp_comm_overlap_rs_dgrad ........................ False tp_comm_split_ag ................................ True tp_comm_split_rs ................................ True train_data_path ................................. None train_iters ..................................... None train_samples ................................... None train_sync_interval ............................. None transformer_impl ................................ transformer_engine transformer_pipeline_model_parallel_size ........ 1 untie_embeddings_and_output_weights ............. True use_checkpoint_args ............................. False use_checkpoint_opt_param_scheduler .............. False use_cpu_initialization .......................... None use_custom_fsdp ................................. False use_dist_ckpt ................................... True use_dist_ckpt_deprecated ........................ False use_distributed_optimizer ....................... True use_flash_attn .................................. False use_legacy_models ............................... False use_mp_args_from_checkpoint_args ................ False use_one_sent_docs ............................... False use_persistent_ckpt_worker ...................... False use_precision_aware_optimizer ................... False use_pytorch_profiler ............................ False use_ring_exchange_p2p ........................... False use_rope_scaling ................................ False use_rotary_position_embeddings .................. False use_tokenizer_model_from_checkpoint_args ........ True use_torch_fsdp2 ................................. False use_torch_optimizer_for_cpu_offload ............. False use_tp_pp_dp_mapping ............................ False v_head_dim ...................................... 128 valid_data_path ................................. None variable_seq_lengths ............................ False virtual_pipeline_model_parallel_size ............ None vision_backbone_type ............................ vit vision_pretraining .............................. False vision_pretraining_type ......................... classify vocab_extra_ids ................................. 0 vocab_file ...................................... None vocab_size ...................................... None wandb_exp_name .................................. wandb_project ................................... wandb_save_dir .................................. weight_decay .................................... 0.01 weight_decay_incr_style ......................... constant wgrad_deferral_limit ............................ 0 world_size ...................................... 1 yaml_cfg ........................................ None -------------------- end of arguments --------------------- 2026-07-23:05:32:33,207 INFO [num_microbatches_calculator.py:228] setting number of microbatches to constant 1 > building HuggingFaceTokenizer tokenizer ... 2026-07-23:05:32:33,434 INFO [_client.py:1038] HTTP Request: HEAD https://huggingface.co/EleutherAI/pythia-12b/resolve/main/config.json "HTTP/1.1 307 Temporary Redirect" 2026-07-23:05:32:33,447 INFO [_client.py:1038] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/EleutherAI/pythia-12b/bb1e3e710cdf6b524461d543cfb5ba773f0a81b6/config.json "HTTP/1.1 200 OK" 2026-07-23:05:32:33,569 INFO [_client.py:1038] HTTP Request: HEAD https://huggingface.co/EleutherAI/pythia-12b/resolve/main/tokenizer_config.json "HTTP/1.1 307 Temporary Redirect" 2026-07-23:05:32:33,582 INFO [_client.py:1038] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/EleutherAI/pythia-12b/bb1e3e710cdf6b524461d543cfb5ba773f0a81b6/tokenizer_config.json "HTTP/1.1 200 OK" 2026-07-23:05:32:33,687 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/models/EleutherAI/pythia-12b/tree/main/additional_chat_templates?recursive=false&expand=false "HTTP/1.1 404 Not Found" 2026-07-23:05:32:33,749 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/models/EleutherAI/pythia-12b/tree/main?recursive=true&expand=false "HTTP/1.1 200 OK" > padded vocab (size: 50277) with 27 dummy tokens (new size: 50304) WARNING: one_logger package is required to enable e2e metrics tracking. please go to https://confluence.nvidia.com/display/MLWFO/Package+Repositories for details to install it 2026-07-23:05:32:33,961 WARNING [rerun_state_machine.py:239] RerunStateMachine initialized in mode RerunMode.DISABLED > initializing torch distributed ... [W723 05:32:33.479174937 CUDAAllocatorConfig.h:28] Warning: expandable_segments not supported on this platform (function operator()) > initialized tensor model parallel with size 1 > initialized pipeline model parallel with size 1 > setting random seeds to 42 ... > compiling dataset index builder ... make: Entering directory '/workspace/FLAME-MoE/Megatron-LM/megatron/core/datasets' make: Nothing to be done for 'default'. make: Leaving directory '/workspace/FLAME-MoE/Megatron-LM/megatron/core/datasets' >>> done with dataset index builder. Compilation time: 0.781 seconds > compiling and loading fused kernels ... >>> done with compiling and loading fused kernels. Compilation time: 0.361 seconds building GPT model ... > number of parameters on (tensor, pipeline) model parallel rank (0, 0): 352994816 /workspace/FLAME-MoE/Megatron-LM/megatron/core/transformer/transformer_layer.py:339: UserWarning: TransformerLayer._get_layer_offset is deprecated.Please use get_transformer_layer_offset instead. warnings.warn( loading distributed checkpoint from /workspace/FLAME-MoE/results/phase0/runs/flame512_g3_temporal/ckpt at iteration 802 /workspace/FLAME-MoE/Megatron-LM/megatron/core/dist_checkpointing/strategies/torch.py:847: FutureWarning: `load_state_dict` is deprecated and will be removed in future versions. Please use `load` instead. checkpoint.load_state_dict( /usr/local/lib/python3.11/dist-packages/torch/distributed/checkpoint/planner_helpers.py:306: FutureWarning: Please use DTensor instead and we are deprecating ShardedTensor. device = getattr(value, "device", None) /workspace/FLAME-MoE/.venv/lib/python3.11/site-packages/transformer_engine/pytorch/module/base.py:566: FutureWarning: You are using `torch.load` with `weights_only=False` (the current default value), which uses the default pickle module implicitly. It is possible to construct malicious pickle data which will execute arbitrary code during unpickling (See https://github.com/pytorch/pytorch/blob/main/SECURITY.md#untrusted-models for more details). In a future release, the default value for `weights_only` will be flipped to `True`. This limits the functions that could be executed during unpickling. Arbitrary objects will no longer be allowed to be loaded via this mode unless they are explicitly allowlisted by the user via `torch.serialization.add_safe_globals`. We recommend you start setting `weights_only=True` for any use case where you don't have full control of the loaded file. Please open an issue on GitHub for any issues related to this experimental feature. state = torch.load(state, map_location="cuda") checkpoint version 3.0 2026-07-23:05:32:44,901 WARNING [rerun_state_machine.py:811] RerunStateMachine disabled via CLI, ignoring machine state saved in checkpoint successfully loaded checkpoint from /workspace/FLAME-MoE/results/phase0/runs/flame512_g3_temporal/ckpt [ t 1/1, p 1/1 ] at iteration 802 2026-07-23:05:32:45,015 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc "HTTP/1.1 200 OK" 2026-07-23:05:32:45,261 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc "HTTP/1.1 200 OK" 2026-07-23:05:32:45,331 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc/revision/210d026faf9955653af8916fad021475a3f00453 "HTTP/1.1 200 OK" 2026-07-23:05:32:45,397 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc/tree/210d026faf9955653af8916fad021475a3f00453?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:32:45,660 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc/tree/210d026faf9955653af8916fad021475a3f00453/ARC-Challenge?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:32:46,134 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc "HTTP/1.1 200 OK" 2026-07-23:05:32:46,372 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc "HTTP/1.1 200 OK" 2026-07-23:05:32:46,700 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/ai2_arc/tree/210d026faf9955653af8916fad021475a3f00453/ARC-Easy?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:32:46,840 WARNING [task.py:799] [Task: boolq] metric acc is defined, but aggregation is not. using default aggregation=mean 2026-07-23:05:32:46,840 WARNING [task.py:811] [Task: boolq] metric acc is defined, but higher_is_better is not. using default higher_is_better=True 2026-07-23:05:32:46,906 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue "HTTP/1.1 200 OK" 2026-07-23:05:32:47,160 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue "HTTP/1.1 200 OK" 2026-07-23:05:32:47,228 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/revision/3de24cf8022e94f4ee4b9d55a6f539891524d646 "HTTP/1.1 200 OK" 2026-07-23:05:32:47,288 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/tree/3de24cf8022e94f4ee4b9d55a6f539891524d646?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:32:47,629 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/tree/3de24cf8022e94f4ee4b9d55a6f539891524d646/axb?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:32:47,691 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/revision/3de24cf8022e94f4ee4b9d55a6f539891524d646 "HTTP/1.1 200 OK" 2026-07-23:05:32:47,765 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/tree/3de24cf8022e94f4ee4b9d55a6f539891524d646/boolq?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:32:48,019 WARNING [task.py:799] [Task: copa] metric acc is defined, but aggregation is not. using default aggregation=mean 2026-07-23:05:32:48,019 WARNING [task.py:811] [Task: copa] metric acc is defined, but higher_is_better is not. using default higher_is_better=True 2026-07-23:05:32:48,103 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue "HTTP/1.1 200 OK" 2026-07-23:05:32:48,339 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue "HTTP/1.1 200 OK" 2026-07-23:05:32:48,608 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/aps/super_glue/tree/3de24cf8022e94f4ee4b9d55a6f539891524d646/copa?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:32:48,750 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag "HTTP/1.1 200 OK" 2026-07-23:05:32:48,973 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag "HTTP/1.1 200 OK" 2026-07-23:05:32:49,041 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag/revision/218ec52e09a7e7462a5400043bb9a69a41d06b76 "HTTP/1.1 200 OK" 2026-07-23:05:32:49,112 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag/tree/218ec52e09a7e7462a5400043bb9a69a41d06b76?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:32:49,339 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag/tree/218ec52e09a7e7462a5400043bb9a69a41d06b76/data?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:32:49,397 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/Rowan/hellaswag/revision/218ec52e09a7e7462a5400043bb9a69a41d06b76 "HTTP/1.1 200 OK" 2026-07-23:05:32:51,801 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai "HTTP/1.1 200 OK" 2026-07-23:05:32:52,040 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai "HTTP/1.1 200 OK" 2026-07-23:05:32:52,100 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/revision/900124bf3b8235c6daf21033af9948b3f07346c4 "HTTP/1.1 200 OK" 2026-07-23:05:32:52,162 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/tree/900124bf3b8235c6daf21033af9948b3f07346c4?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:32:52,382 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/tree/900124bf3b8235c6daf21033af9948b3f07346c4/default%2Ftest?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:32:52,448 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/tree/900124bf3b8235c6daf21033af9948b3f07346c4/default?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:32:52,509 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/EleutherAI/lambada_openai/revision/900124bf3b8235c6daf21033af9948b3f07346c4 "HTTP/1.1 200 OK" 2026-07-23:05:32:52,526 WARNING [task.py:325] [Task: lambada_openai] has_training_docs and has_validation_docs are False, using test_docs as fewshot_docs but this is not recommended. 2026-07-23:05:32:52,526 WARNING [task.py:325] [Task: lambada_openai] has_training_docs and has_validation_docs are False, using test_docs as fewshot_docs but this is not recommended. 2026-07-23:05:32:52,697 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa "HTTP/1.1 200 OK" 2026-07-23:05:32:52,938 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa "HTTP/1.1 200 OK" 2026-07-23:05:32:53,003 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/revision/388097ea7776314e93a529163e0fea805b8a6454 "HTTP/1.1 200 OK" 2026-07-23:05:32:53,062 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/tree/388097ea7776314e93a529163e0fea805b8a6454?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:32:53,278 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/tree/388097ea7776314e93a529163e0fea805b8a6454/additional?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:32:53,337 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/revision/388097ea7776314e93a529163e0fea805b8a6454 "HTTP/1.1 200 OK" 2026-07-23:05:32:53,414 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/openbookqa/tree/388097ea7776314e93a529163e0fea805b8a6454/main?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:32:53,705 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa "HTTP/1.1 200 OK" 2026-07-23:05:32:53,759 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa "HTTP/1.1 200 OK" 2026-07-23:05:32:53,820 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa/tree/main?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:32:53,884 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/datasets/ybisk/piqa/resolve/main/piqa.py "HTTP/1.1 307 Temporary Redirect" 2026-07-23:05:32:53,897 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/resolve-cache/datasets/ybisk/piqa/2e8ac2dffd59bac8c3c6714948f4c551a0848bb0/piqa.py?%2Fdatasets%2Fybisk%2Fpiqa%2Fresolve%2Fmain%2Fpiqa.py=&etag=%220f0d96172f9bf7aa10069374b3236ef083cfeb8e%22 "HTTP/1.1 200 OK" 2026-07-23:05:32:54,437 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa/revision/main "HTTP/1.1 200 OK" 2026-07-23:05:32:54,493 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/ybisk/piqa/tree/main?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:32:54,939 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq "HTTP/1.1 200 OK" 2026-07-23:05:32:55,146 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq "HTTP/1.1 200 OK" 2026-07-23:05:32:55,210 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq/revision/2c94ad3e1aafab77146f384e23536f97a4849815 "HTTP/1.1 200 OK" 2026-07-23:05:32:55,277 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq/tree/2c94ad3e1aafab77146f384e23536f97a4849815?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:32:55,491 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq/tree/2c94ad3e1aafab77146f384e23536f97a4849815/data?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:32:55,564 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/sciq/revision/2c94ad3e1aafab77146f384e23536f97a4849815 "HTTP/1.1 200 OK" 2026-07-23:05:32:55,994 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande "HTTP/1.1 200 OK" 2026-07-23:05:32:56,221 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande "HTTP/1.1 200 OK" 2026-07-23:05:32:56,342 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/revision/01e74176c63542e6b0bcb004dcdea22d94fb67b5 "HTTP/1.1 200 OK" 2026-07-23:05:32:56,413 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/tree/01e74176c63542e6b0bcb004dcdea22d94fb67b5?recursive=false&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:32:56,655 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/tree/01e74176c63542e6b0bcb004dcdea22d94fb67b5/winogrande_debiased?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:32:56,717 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/revision/01e74176c63542e6b0bcb004dcdea22d94fb67b5 "HTTP/1.1 200 OK" 2026-07-23:05:32:56,798 INFO [_client.py:1038] HTTP Request: GET https://huggingface.co/api/datasets/allenai/winogrande/tree/01e74176c63542e6b0bcb004dcdea22d94fb67b5/winogrande_xl?recursive=true&expand=false "HTTP/1.1 200 OK" 2026-07-23:05:32:57,453 WARNING [evaluator.py:270] Overwriting default num_fewshot of winogrande from None to 0 2026-07-23:05:32:57,453 WARNING [evaluator.py:270] Overwriting default num_fewshot of sciq from None to 0 2026-07-23:05:32:57,453 WARNING [evaluator.py:270] Overwriting default num_fewshot of piqa from None to 0 2026-07-23:05:32:57,453 WARNING [evaluator.py:270] Overwriting default num_fewshot of openbookqa from None to 0 2026-07-23:05:32:57,453 WARNING [evaluator.py:270] Overwriting default num_fewshot of lambada_openai from None to 0 2026-07-23:05:32:57,453 WARNING [evaluator.py:270] Overwriting default num_fewshot of hellaswag from None to 0 2026-07-23:05:32:57,453 WARNING [evaluator.py:270] Overwriting default num_fewshot of copa from None to 0 2026-07-23:05:32:57,453 WARNING [evaluator.py:270] Overwriting default num_fewshot of boolq from None to 0 2026-07-23:05:32:57,453 WARNING [evaluator.py:270] Overwriting default num_fewshot of arc_easy from None to 0 2026-07-23:05:32:57,453 WARNING [evaluator.py:270] Overwriting default num_fewshot of arc_challenge from None to 0 2026-07-23:05:32:57,478 INFO [task.py:415] Building contexts for winogrande on rank 0... 0%| | 0/1267 [00:00 /workspace/FLAME-MoE/results/ablations/flame1e18_downstream.csv [done] driver finished for: flame38m_dense_local flame38m_g1_moe flame38m_g1_moe_s2 flame38m_g1_moe_s3 flame38m_g1_temporal flame38m_g1_temporal_s2 flame38m_g1_temporal_s3 flame38m_g3_moe flame38m_g3_moe_s2 flame38m_g3_moe_s3 flame38m_g3_temporal flame38m_g3_temporal_s2 flame38m_g3_temporal_s3 flame192_g3_moe flame192_g3_temporal flame512_dense flame512_g1_moe flame512_g1_temporal flame512_g3_moe flame512_g3_temporal [done] all runs succeeded