hbfreed's picture
Add files using upload-large-folder tool
8717f59 verified
Raw
History Blame Contribute Delete
8.18 kB
(APIServer pid=1275172) INFO 07-24 12:12:54 [api_utils.py:339]
(APIServer pid=1275172) INFO 07-24 12:12:54 [api_utils.py:339] β–ˆ β–ˆ β–ˆβ–„ β–„β–ˆ
(APIServer pid=1275172) INFO 07-24 12:12:54 [api_utils.py:339] β–„β–„ β–„β–ˆ β–ˆ β–ˆ β–ˆ β–€β–„β–€ β–ˆ version 0.25.0
(APIServer pid=1275172) INFO 07-24 12:12:54 [api_utils.py:339] β–ˆβ–„β–ˆβ–€ β–ˆ β–ˆ β–ˆ β–ˆ model outputs/healed_dl/keep50_step0200
(APIServer pid=1275172) INFO 07-24 12:12:54 [api_utils.py:339] β–€β–€ β–€β–€β–€β–€β–€ β–€β–€β–€β–€β–€ β–€ β–€
(APIServer pid=1275172) INFO 07-24 12:12:54 [api_utils.py:339]
(APIServer pid=1275172) INFO 07-24 12:12:54 [api_utils.py:273] non-default args: {'model_tag': 'outputs/healed_dl/keep50_step0200', 'default_chat_template_kwargs': {'enable_thinking': True}, 'host': '127.0.0.1', 'port': 8399, 'model': 'outputs/healed_dl/keep50_step0200', 'max_model_len': 8192, 'enforce_eager': True, 'served_model_name': ['student'], 'pipeline_parallel_size': 2, 'gpu_memory_utilization': 0.9}
(APIServer pid=1275172) INFO 07-24 12:12:54 [model.py:619] Resolved architecture: PrunedQwen3_5MoeForCausalLM
(APIServer pid=1275172) INFO 07-24 12:12:54 [model.py:1776] Using max model len 8192
(APIServer pid=1275172) INFO 07-24 12:12:54 [vllm.py:1042] Asynchronous scheduling is enabled.
(APIServer pid=1275172) WARNING 07-24 12:12:54 [vllm.py:1096] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
(APIServer pid=1275172) WARNING 07-24 12:12:54 [vllm.py:1144] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
(APIServer pid=1275172) INFO 07-24 12:12:54 [kernel.py:292] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native'])
(APIServer pid=1275172) INFO 07-24 12:12:54 [vllm.py:1322] Cudagraph is disabled under eager mode
(APIServer pid=1275172) INFO 07-24 12:12:54 [compilation.py:312] Enabled custom fusions: norm_quant, act_quant
(EngineCore pid=1275369) INFO 07-24 12:13:06 [core.py:114] Initializing a V1 LLM engine (v0.25.0) with config: model='outputs/healed_dl/keep50_step0200', speculative_config=None, tokenizer='outputs/healed_dl/keep50_step0200', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=False, dtype=torch.bfloat16, max_seq_len=8192, download_dir=None, load_format=auto, tensor_parallel_size=1, pipeline_parallel_size=2, data_parallel_size=1, decode_context_parallel_size=1, dcp_comm_backend=ag_rs, disable_custom_all_reduce=False, quantization=None, quantization_config=None, enforce_eager=True, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False, jit_monitor_mode='warn', jit_monitor_verbose=False), seed=0, served_model_name=student, enable_prefix_caching=False, enable_chunked_prefill=True, pooler_config=None, compilation_config={'mode': <CompilationMode.NONE: 0>, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'ir_enable_torch_wrap': False, 'splitting_ops': [], 'compile_mm_encoder': False, 'cudagraph_mm_encoder': False, 'encoder_cudagraph_token_budgets': [], 'encoder_cudagraph_max_vision_items_per_batch': 0, 'encoder_cudagraph_max_frames_per_batch': None, 'compile_sizes': [], 'compile_ranges_endpoints': [2048], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'size_asserts': False, 'alignment_asserts': False, 'scalar_asserts': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': <CUDAGraphMode.NONE: 0>, 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': True, 'fuse_act_quant': True, 'fuse_attn_quant': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False, 'fuse_rope_kvcache_cat_mla': False, 'fuse_act_padding': False}, 'max_cudagraph_capture_size': 0, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': False, 'static_all_moe_layers': []}, kernel_config=KernelConfig(ir_op_priority=IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native']), enable_flashinfer_autotune=True, enable_cutedsl_warmup=True, moe_backend='auto', linear_backend='auto')
(EngineCore pid=1275369) WARNING 07-24 12:13:06 [multiproc_executor.py:1067] Reducing Torch parallelism from 24 threads to 1 to avoid unnecessary CPU contention. Set OMP_NUM_THREADS in the external environment to tune this value as needed.
(EngineCore pid=1275369) INFO 07-24 12:13:06 [multiproc_executor.py:140] DP group leader: node_rank=0, node_rank_within_dp=0, master_addr=127.0.0.1, mq_connect_ip=192.168.0.15 (local), world_size=2, local_world_size=2
(Worker pid=1275483) INFO 07-24 12:13:14 [parallel_state.py:1607] world_size=2 rank=0 local_rank=0 distributed_init_method=tcp://127.0.0.1:33603 backend=nccl
(Worker pid=1275484) INFO 07-24 12:13:14 [parallel_state.py:1607] world_size=2 rank=1 local_rank=1 distributed_init_method=tcp://127.0.0.1:33603 backend=nccl
(Worker pid=1275483) INFO 07-24 12:13:15 [pynccl.py:113] vLLM is using nccl==2.28.9
(Worker pid=1275483) INFO 07-24 12:13:16 [cuda_communicator.py:264] Using ['PYNCCL'] all-reduce backends (in dispatch order) for group 'pp:0' out of potential backends: ['NCCL_SYMM_MEM', 'QUICK_REDUCE', 'FLASHINFER', 'AITER_CUSTOM', 'CUSTOM', 'SYMM_MEM', 'PYNCCL'].
(Worker pid=1275483) INFO 07-24 12:13:16 [parallel_state.py:1942] rank 0 in world size 2 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank 0, EPLB rank N/A
(Worker pid=1275483) INFO 07-24 12:13:16 [topk_topp_sampler.py:55] Using FlashInfer for top-p & top-k sampling.
(Worker_PP0 pid=1275483) INFO 07-24 12:13:16 [gpu_model_runner.py:5209] Starting to load model outputs/healed_dl/keep50_step0200...
(Worker_PP0 pid=1275483) INFO 07-24 12:13:16 [qwen_gdn_linear_attn.py:228] Using Triton/FLA GDN prefill kernel (requested=auto, head_k_dim=128).
(Worker_PP0 pid=1275483) /home/henry/.cache/glean/megablocks-variable-93a1479bc15b/megablocks/grouped_gemm_util.py:10: UserWarning: Grouped GEMM not available.
(Worker_PP0 pid=1275483) warnings.warn('Grouped GEMM not available.')
(Worker_PP1 pid=1275484) /home/henry/.cache/glean/megablocks-variable-93a1479bc15b/megablocks/grouped_gemm_util.py:10: UserWarning: Grouped GEMM not available.
(Worker_PP1 pid=1275484) warnings.warn('Grouped GEMM not available.')
(Worker_PP0 pid=1275483) INFO 07-24 12:13:17 [cuda.py:476] Using FLASH_ATTN attention backend out of potential backends: ['FLASH_ATTN', 'FLASHINFER', 'TRITON_ATTN', 'FLEX_ATTENTION'].
(Worker_PP0 pid=1275483) INFO 07-24 12:13:17 [flash_attn.py:718] Using FlashAttention version 2
(Worker_PP0 pid=1275483) INFO 07-24 12:13:17 [weight_utils.py:849] Filesystem type for checkpoints: EXT4. Checkpoint size: 34.55 GiB. Available RAM: 106.11 GiB.
(Worker_PP0 pid=1275483) INFO 07-24 12:13:17 [weight_utils.py:872] Auto-prefetch is disabled because the filesystem (EXT4) is not a recognized network FS (NFS/Lustre). If you want to force prefetching, start vLLM with --safetensors-load-strategy=prefetch.
(Worker_PP0 pid=1275483) Loading safetensors checkpoint shards: 0% Completed | 0/1 [00:00<?, ?it/s]