Add files using upload-large-folder tool
Browse filesThis view is limited to 50 files because it contains too many changes. Β See raw diff
- evals/grid_math_unhealed/glean_keep25_unhealed.log +8 -0
- evals/grid_math_unhealed/glean_keep25_unhealed_chat.json +0 -0
- evals/grid_math_unhealed/glean_keep25_unhealed_chat.json.server.log +109 -0
- evals/grid_math_unhealed/glean_keep50_unhealed.log +8 -0
- evals/grid_math_unhealed/glean_keep50_unhealed_chat.json +0 -0
- evals/grid_math_unhealed/glean_keep50_unhealed_chat.json.server.log +102 -0
- evals/grid_math_unhealed/glean_keep75_unhealed.log +8 -0
- evals/grid_math_unhealed/glean_keep75_unhealed_chat.json +0 -0
- evals/grid_math_unhealed/glean_keep75_unhealed_chat.json.server.log +103 -0
- evals/grid_math_unhealed/reap_keep25_unhealed.log +8 -0
- evals/grid_math_unhealed/reap_keep25_unhealed_chat.json +0 -0
- evals/grid_math_unhealed/reap_keep25_unhealed_chat.json.server.log +117 -0
- evals/grid_math_unhealed/reap_keep50_unhealed.log +8 -0
- evals/grid_math_unhealed/reap_keep50_unhealed_chat.json +0 -0
- evals/grid_math_unhealed/reap_keep50_unhealed_chat.json.server.log +110 -0
- evals/grid_math_unhealed/reap_keep75_unhealed.log +8 -0
- evals/grid_math_unhealed/reap_keep75_unhealed_chat.json +0 -0
- evals/grid_math_unhealed/reap_keep75_unhealed_chat.json.server.log +105 -0
- evals/grid_math_unhealed/uniform_keep25_unhealed.log +8 -0
- evals/grid_math_unhealed/uniform_keep25_unhealed_chat.json +0 -0
- evals/grid_math_unhealed/uniform_keep25_unhealed_chat.json.server.log +115 -0
- evals/grid_math_unhealed/uniform_keep50_unhealed.log +8 -0
- evals/grid_math_unhealed/uniform_keep50_unhealed_chat.json +0 -0
- evals/grid_math_unhealed/uniform_keep50_unhealed_chat.json.server.log +101 -0
- evals/grid_math_unhealed/uniform_keep75_unhealed.log +8 -0
- evals/grid_math_unhealed/uniform_keep75_unhealed_chat.json +0 -0
- evals/grid_math_unhealed/uniform_keep75_unhealed_chat.json.server.log +103 -0
- evals/healing_breadth/glean_math_keep25_seed1224.json +0 -0
- evals/healing_breadth/glean_math_keep25_seed1224.json.server.log +139 -0
- evals/healing_breadth/glean_math_keep25_seed1224_long768_step0500_chat.json +0 -0
- evals/healing_breadth/glean_math_keep25_seed1224_long768_step0500_chat.json.server.log +255 -0
- evals/healing_breadth/glean_math_keep25_seed1224_long768_step0500_raw.json +0 -0
- evals/healing_breadth/glean_math_keep25_seed1224_long768_step0500_raw.json.server.log +377 -0
- evals/healing_breadth/glean_math_keep25_seed1224_step0050_chat.json +0 -0
- evals/healing_breadth/glean_math_keep25_seed1224_step0050_chat.json.server.log +138 -0
- evals/healing_breadth/glean_math_keep25_seed1224_step0050_raw.json +0 -0
- evals/healing_breadth/glean_math_keep25_seed1224_step0050_raw.json.server.log +260 -0
- evals/healing_breadth/glean_math_keep75_oneshot_chat.json +0 -0
- evals/healing_breadth/glean_math_keep75_oneshot_chat.json.server.log +130 -0
- evals/healing_breadth/glean_math_keep75_oneshot_raw.json +0 -0
- evals/healing_breadth/glean_math_keep75_oneshot_raw.json.server.log +129 -0
- evals/healing_breadth/reap_math_keep75_seed1224.json +0 -0
- evals/healing_breadth/reap_math_keep75_seed1224.json.server.log +128 -0
- evals/healing_breadth/uniform_math_keep50_seed1224.json +0 -0
- evals/healing_breadth/uniform_math_keep50_seed1224.json.server.log +130 -0
- evals/math_unhealed/glean_winnow-olmoe-math-keep25.eval.log +0 -0
- evals/math_unhealed/glean_winnow-olmoe-math-keep50.eval.log +70 -0
- evals/math_unhealed/glean_winnow-olmoe-math-keep75.eval.log +0 -0
- evals/math_unhealed/reap_reap-math-keep25.eval.log +0 -0
- evals/math_unhealed/reap_reap-math-keep50.eval.log +0 -0
evals/grid_math_unhealed/glean_keep25_unhealed.log
ADDED
|
@@ -0,0 +1,8 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"correct": 143,
|
| 3 |
+
"accuracy": 0.10841546626231995,
|
| 4 |
+
"finished": 588,
|
| 5 |
+
"finish_rate": 0.44579226686884005,
|
| 6 |
+
"mean_completion_tokens": 327.01440485216074
|
| 7 |
+
}
|
| 8 |
+
saved item-level results -> outputs/evals/grid_math_unhealed/glean_keep25_unhealed_chat.json
|
evals/grid_math_unhealed/glean_keep25_unhealed_chat.json
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
evals/grid_math_unhealed/glean_keep25_unhealed_chat.json.server.log
ADDED
|
@@ -0,0 +1,109 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
Skipping import of cpp extensions due to incompatible torch version 2.10.0+cu128 for torchao version 0.15.0 Please see https://github.com/pytorch/ao/issues/2919 for more info
|
| 2 |
+
WARNING 08-10 21:24:57 [registry.py:915] Model architecture OlmoeForCausalLM is already registered, and will be overwritten by the new model class glean_vllm.pruned_olmoe:PrunedOlmoeForCausalLM.
|
| 3 |
+
(APIServer pid=1153466) INFO 08-10 21:24:57 [utils.py:299]
|
| 4 |
+
(APIServer pid=1153466) INFO 08-10 21:24:57 [utils.py:299] β β ββ ββ
|
| 5 |
+
(APIServer pid=1153466) INFO 08-10 21:24:57 [utils.py:299] ββ ββ β β β βββ β version 0.19.0
|
| 6 |
+
(APIServer pid=1153466) INFO 08-10 21:24:57 [utils.py:299] ββββ β β β β model outputs/release/unhealed/winnow-olmoe-math-keep25
|
| 7 |
+
(APIServer pid=1153466) INFO 08-10 21:24:57 [utils.py:299] ββ βββββ βββββ β β
|
| 8 |
+
(APIServer pid=1153466) INFO 08-10 21:24:57 [utils.py:299]
|
| 9 |
+
(APIServer pid=1153466) INFO 08-10 21:24:57 [utils.py:233] non-default args: {'model_tag': 'outputs/release/unhealed/winnow-olmoe-math-keep25', 'host': '127.0.0.1', 'port': 8610, 'model': 'outputs/release/unhealed/winnow-olmoe-math-keep25', 'max_model_len': 2048, 'enforce_eager': True, 'served_model_name': ['student'], 'gpu_memory_utilization': 0.85}
|
| 10 |
+
(APIServer pid=1153466) INFO 08-10 21:25:05 [model.py:549] Resolved architecture: PrunedOlmoeForCausalLM
|
| 11 |
+
(APIServer pid=1153466) INFO 08-10 21:25:05 [model.py:1678] Using max model len 2048
|
| 12 |
+
(APIServer pid=1153466) INFO 08-10 21:25:05 [vllm.py:790] Asynchronous scheduling is enabled.
|
| 13 |
+
(APIServer pid=1153466) WARNING 08-10 21:25:05 [vllm.py:848] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
|
| 14 |
+
(APIServer pid=1153466) WARNING 08-10 21:25:05 [vllm.py:859] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
|
| 15 |
+
(APIServer pid=1153466) INFO 08-10 21:25:05 [vllm.py:1025] Cudagraph is disabled under eager mode
|
| 16 |
+
(APIServer pid=1153466) INFO 08-10 21:25:05 [compilation.py:290] Enabled custom fusions: norm_quant, act_quant
|
| 17 |
+
Skipping import of cpp extensions due to incompatible torch version 2.10.0+cu128 for torchao version 0.15.0 Please see https://github.com/pytorch/ao/issues/2919 for more info
|
| 18 |
+
(EngineCore pid=1154734) WARNING 08-10 21:25:14 [registry.py:915] Model architecture OlmoeForCausalLM is already registered, and will be overwritten by the new model class glean_vllm.pruned_olmoe:PrunedOlmoeForCausalLM.
|
| 19 |
+
(EngineCore pid=1154734) INFO 08-10 21:25:14 [core.py:105] Initializing a V1 LLM engine (v0.19.0) with config: model='outputs/release/unhealed/winnow-olmoe-math-keep25', speculative_config=None, tokenizer='outputs/release/unhealed/winnow-olmoe-math-keep25', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=False, dtype=torch.bfloat16, max_seq_len=2048, download_dir=None, load_format=auto, tensor_parallel_size=1, pipeline_parallel_size=1, data_parallel_size=1, decode_context_parallel_size=1, dcp_comm_backend=ag_rs, disable_custom_all_reduce=False, quantization=None, enforce_eager=True, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False), seed=0, served_model_name=student, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'mode': <CompilationMode.NONE: 0>, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'splitting_ops': [], 'compile_mm_encoder': False, 'cudagraph_mm_encoder': False, 'encoder_cudagraph_token_budgets': [], 'encoder_cudagraph_max_images_per_batch': 0, 'compile_sizes': [], 'compile_ranges_endpoints': [2048], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'size_asserts': False, 'alignment_asserts': False, 'scalar_asserts': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': <CUDAGraphMode.NONE: 0>, 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': True, 'fuse_act_quant': True, 'fuse_attn_quant': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False}, 'max_cudagraph_capture_size': 0, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': True, 'static_all_moe_layers': []}
|
| 20 |
+
(EngineCore pid=1154734) INFO 08-10 21:25:14 [parallel_state.py:1400] world_size=1 rank=0 local_rank=0 distributed_init_method=tcp://192.168.0.27:46169 backend=nccl
|
| 21 |
+
(EngineCore pid=1154734) INFO 08-10 21:25:14 [parallel_state.py:1716] rank 0 in world size 1 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank 0, EPLB rank N/A
|
| 22 |
+
(EngineCore pid=1154734) INFO 08-10 21:25:15 [gpu_model_runner.py:4735] Starting to load model outputs/release/unhealed/winnow-olmoe-math-keep25...
|
| 23 |
+
(EngineCore pid=1154734) INFO 08-10 21:25:16 [cuda.py:334] Using FLASH_ATTN attention backend out of potential backends: ['FLASH_ATTN', 'FLASHINFER', 'TRITON_ATTN', 'FLEX_ATTENTION'].
|
| 24 |
+
(EngineCore pid=1154734) INFO 08-10 21:25:16 [flash_attn.py:596] Using FlashAttention version 2
|
| 25 |
+
(EngineCore pid=1154734)
|
| 26 |
+
(EngineCore pid=1154734)
|
| 27 |
+
(EngineCore pid=1154734)
|
| 28 |
+
(EngineCore pid=1154734)
|
| 29 |
+
(EngineCore pid=1154734) INFO 08-10 21:25:22 [default_loader.py:384] Loading weights took 5.84 seconds
|
| 30 |
+
(EngineCore pid=1154734) INFO 08-10 21:25:22 [gpu_model_runner.py:4820] Model loading took 3.89 GiB memory and 6.430009 seconds
|
| 31 |
+
(EngineCore pid=1154734) INFO 08-10 21:25:24 [gpu_worker.py:436] Available KV cache memory: 15.86 GiB
|
| 32 |
+
(EngineCore pid=1154734) INFO 08-10 21:25:24 [kv_cache_utils.py:1319] GPU KV cache size: 129,920 tokens
|
| 33 |
+
(EngineCore pid=1154734) INFO 08-10 21:25:24 [kv_cache_utils.py:1324] Maximum concurrency for 2,048 tokens per request: 63.44x
|
| 34 |
+
(EngineCore pid=1154734) INFO 08-10 21:25:24 [core.py:283] init engine (profile, create kv cache, warmup model) took 1.67 seconds
|
| 35 |
+
(EngineCore pid=1154734) INFO 08-10 21:25:25 [vllm.py:790] Asynchronous scheduling is enabled.
|
| 36 |
+
(EngineCore pid=1154734) WARNING 08-10 21:25:25 [vllm.py:848] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
|
| 37 |
+
(EngineCore pid=1154734) WARNING 08-10 21:25:25 [vllm.py:859] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
|
| 38 |
+
(EngineCore pid=1154734) INFO 08-10 21:25:25 [vllm.py:1025] Cudagraph is disabled under eager mode
|
| 39 |
+
(EngineCore pid=1154734) INFO 08-10 21:25:25 [compilation.py:290] Enabled custom fusions: norm_quant, act_quant
|
| 40 |
+
(APIServer pid=1153466) INFO 08-10 21:25:25 [api_server.py:590] Supported tasks: ['generate']
|
| 41 |
+
(APIServer pid=1153466) WARNING 08-10 21:25:25 [__init__.py:14] SECURITY WARNING: Development endpoints are enabled! This should NOT be used in production!
|
| 42 |
+
(APIServer pid=1153466) INFO 08-10 21:25:25 [hf.py:314] Detected the chat template content format to be 'string'. You can set `--chat-template-content-format` to override this.
|
| 43 |
+
(APIServer pid=1153466) INFO 08-10 21:25:25 [api_server.py:594] Starting vLLM server on http://127.0.0.1:8610
|
| 44 |
+
(APIServer pid=1153466) INFO 08-10 21:25:25 [launcher.py:37] Available routes are:
|
| 45 |
+
(APIServer pid=1153466) INFO 08-10 21:25:25 [launcher.py:46] Route: /openapi.json, Methods: GET, HEAD
|
| 46 |
+
(APIServer pid=1153466) INFO 08-10 21:25:25 [launcher.py:46] Route: /docs, Methods: GET, HEAD
|
| 47 |
+
(APIServer pid=1153466) INFO 08-10 21:25:25 [launcher.py:46] Route: /docs/oauth2-redirect, Methods: GET, HEAD
|
| 48 |
+
(APIServer pid=1153466) INFO 08-10 21:25:25 [launcher.py:46] Route: /redoc, Methods: GET, HEAD
|
| 49 |
+
(APIServer pid=1153466) INFO 08-10 21:25:25 [launcher.py:46] Route: /sleep, Methods: POST
|
| 50 |
+
(APIServer pid=1153466) INFO 08-10 21:25:25 [launcher.py:46] Route: /wake_up, Methods: POST
|
| 51 |
+
(APIServer pid=1153466) INFO 08-10 21:25:25 [launcher.py:46] Route: /is_sleeping, Methods: GET
|
| 52 |
+
(APIServer pid=1153466) INFO 08-10 21:25:25 [launcher.py:46] Route: /collective_rpc, Methods: POST
|
| 53 |
+
(APIServer pid=1153466) INFO 08-10 21:25:25 [launcher.py:46] Route: /reset_prefix_cache, Methods: POST
|
| 54 |
+
(APIServer pid=1153466) INFO 08-10 21:25:25 [launcher.py:46] Route: /reset_mm_cache, Methods: POST
|
| 55 |
+
(APIServer pid=1153466) INFO 08-10 21:25:25 [launcher.py:46] Route: /reset_encoder_cache, Methods: POST
|
| 56 |
+
(APIServer pid=1153466) INFO 08-10 21:25:25 [launcher.py:46] Route: /tokenize, Methods: POST
|
| 57 |
+
(APIServer pid=1153466) INFO 08-10 21:25:25 [launcher.py:46] Route: /detokenize, Methods: POST
|
| 58 |
+
(APIServer pid=1153466) INFO 08-10 21:25:25 [launcher.py:46] Route: /load, Methods: GET
|
| 59 |
+
(APIServer pid=1153466) INFO 08-10 21:25:25 [launcher.py:46] Route: /version, Methods: GET
|
| 60 |
+
(APIServer pid=1153466) INFO 08-10 21:25:25 [launcher.py:46] Route: /health, Methods: GET
|
| 61 |
+
(APIServer pid=1153466) INFO 08-10 21:25:25 [launcher.py:46] Route: /metrics, Methods: GET
|
| 62 |
+
(APIServer pid=1153466) INFO 08-10 21:25:25 [launcher.py:46] Route: /server_info, Methods: GET
|
| 63 |
+
(APIServer pid=1153466) INFO 08-10 21:25:25 [launcher.py:46] Route: /v1/models, Methods: GET
|
| 64 |
+
(APIServer pid=1153466) INFO 08-10 21:25:25 [launcher.py:46] Route: /ping, Methods: GET
|
| 65 |
+
(APIServer pid=1153466) INFO 08-10 21:25:25 [launcher.py:46] Route: /ping, Methods: POST
|
| 66 |
+
(APIServer pid=1153466) INFO 08-10 21:25:25 [launcher.py:46] Route: /invocations, Methods: POST
|
| 67 |
+
(APIServer pid=1153466) INFO 08-10 21:25:25 [launcher.py:46] Route: /v1/chat/completions, Methods: POST
|
| 68 |
+
(APIServer pid=1153466) INFO 08-10 21:25:25 [launcher.py:46] Route: /v1/chat/completions/batch, Methods: POST
|
| 69 |
+
(APIServer pid=1153466) INFO 08-10 21:25:25 [launcher.py:46] Route: /v1/responses, Methods: POST
|
| 70 |
+
(APIServer pid=1153466) INFO 08-10 21:25:25 [launcher.py:46] Route: /v1/responses/{response_id}, Methods: GET
|
| 71 |
+
(APIServer pid=1153466) INFO 08-10 21:25:25 [launcher.py:46] Route: /v1/responses/{response_id}/cancel, Methods: POST
|
| 72 |
+
(APIServer pid=1153466) INFO 08-10 21:25:25 [launcher.py:46] Route: /v1/completions, Methods: POST
|
| 73 |
+
(APIServer pid=1153466) INFO 08-10 21:25:25 [launcher.py:46] Route: /v1/messages, Methods: POST
|
| 74 |
+
(APIServer pid=1153466) INFO 08-10 21:25:25 [launcher.py:46] Route: /v1/messages/count_tokens, Methods: POST
|
| 75 |
+
(APIServer pid=1153466) INFO 08-10 21:25:25 [launcher.py:46] Route: /inference/v1/generate, Methods: POST
|
| 76 |
+
(APIServer pid=1153466) INFO 08-10 21:25:25 [launcher.py:46] Route: /pause, Methods: POST
|
| 77 |
+
(APIServer pid=1153466) INFO 08-10 21:25:25 [launcher.py:46] Route: /resume, Methods: POST
|
| 78 |
+
(APIServer pid=1153466) INFO 08-10 21:25:25 [launcher.py:46] Route: /is_paused, Methods: GET
|
| 79 |
+
(APIServer pid=1153466) INFO 08-10 21:25:25 [launcher.py:46] Route: /init_weight_transfer_engine, Methods: POST
|
| 80 |
+
(APIServer pid=1153466) INFO 08-10 21:25:25 [launcher.py:46] Route: /update_weights, Methods: POST
|
| 81 |
+
(APIServer pid=1153466) INFO 08-10 21:25:25 [launcher.py:46] Route: /get_world_size, Methods: GET
|
| 82 |
+
(APIServer pid=1153466) INFO 08-10 21:25:25 [launcher.py:46] Route: /scale_elastic_ep, Methods: POST
|
| 83 |
+
(APIServer pid=1153466) INFO 08-10 21:25:25 [launcher.py:46] Route: /is_scaling_elastic_ep, Methods: POST
|
| 84 |
+
(APIServer pid=1153466) INFO 08-10 21:25:25 [launcher.py:46] Route: /v1/chat/completions/render, Methods: POST
|
| 85 |
+
(APIServer pid=1153466) INFO 08-10 21:25:25 [launcher.py:46] Route: /v1/completions/render, Methods: POST
|
| 86 |
+
(APIServer pid=1153466) INFO: Started server process [1153466]
|
| 87 |
+
(APIServer pid=1153466) INFO: Waiting for application startup.
|
| 88 |
+
(APIServer pid=1153466) INFO: Application startup complete.
|
| 89 |
+
(APIServer pid=1153466) INFO: 127.0.0.1:55122 - "GET /health HTTP/1.1" 200 OK
|
| 90 |
+
(APIServer pid=1153466) INFO 08-10 21:25:35 [loggers.py:259] Engine 000: Avg prompt throughput: 2593.9 tokens/s, Avg generation throughput: 2526.0 tokens/s, Running: 256 reqs, Waiting: 983 reqs, GPU KV cache usage: 35.2%, Prefix cache hit rate: 93.2%
|
| 91 |
+
(APIServer pid=1153466) INFO 08-10 21:25:45 [loggers.py:259] Engine 000: Avg prompt throughput: 519.2 tokens/s, Avg generation throughput: 3397.7 tokens/s, Running: 256 reqs, Waiting: 915 reqs, GPU KV cache usage: 56.2%, Prefix cache hit rate: 93.2%
|
| 92 |
+
(APIServer pid=1153466) INFO 08-10 21:25:55 [loggers.py:259] Engine 000: Avg prompt throughput: 239.9 tokens/s, Avg generation throughput: 3248.2 tokens/s, Running: 256 reqs, Waiting: 890 reqs, GPU KV cache usage: 78.8%, Prefix cache hit rate: 93.2%
|
| 93 |
+
(APIServer pid=1153466) INFO 08-10 21:26:05 [loggers.py:259] Engine 000: Avg prompt throughput: 127.4 tokens/s, Avg generation throughput: 3094.9 tokens/s, Running: 256 reqs, Waiting: 874 reqs, GPU KV cache usage: 100.0%, Prefix cache hit rate: 93.2%
|
| 94 |
+
(APIServer pid=1153466) INFO 08-10 21:26:15 [loggers.py:259] Engine 000: Avg prompt throughput: 1839.2 tokens/s, Avg generation throughput: 3085.6 tokens/s, Running: 254 reqs, Waiting: 632 reqs, GPU KV cache usage: 47.1%, Prefix cache hit rate: 93.3%
|
| 95 |
+
(APIServer pid=1153466) INFO 08-10 21:26:25 [loggers.py:259] Engine 000: Avg prompt throughput: 930.8 tokens/s, Avg generation throughput: 3341.7 tokens/s, Running: 256 reqs, Waiting: 517 reqs, GPU KV cache usage: 51.5%, Prefix cache hit rate: 93.3%
|
| 96 |
+
(APIServer pid=1153466) INFO 08-10 21:26:35 [loggers.py:259] Engine 000: Avg prompt throughput: 415.0 tokens/s, Avg generation throughput: 3296.3 tokens/s, Running: 256 reqs, Waiting: 466 reqs, GPU KV cache usage: 68.5%, Prefix cache hit rate: 93.3%
|
| 97 |
+
(APIServer pid=1153466) INFO 08-10 21:26:45 [loggers.py:259] Engine 000: Avg prompt throughput: 189.8 tokens/s, Avg generation throughput: 3169.8 tokens/s, Running: 256 reqs, Waiting: 441 reqs, GPU KV cache usage: 88.3%, Prefix cache hit rate: 93.3%
|
| 98 |
+
(APIServer pid=1153466) INFO 08-10 21:26:55 [loggers.py:259] Engine 000: Avg prompt throughput: 1409.8 tokens/s, Avg generation throughput: 3105.7 tokens/s, Running: 256 reqs, Waiting: 268 reqs, GPU KV cache usage: 57.2%, Prefix cache hit rate: 93.4%
|
| 99 |
+
(APIServer pid=1153466) INFO 08-10 21:27:05 [loggers.py:259] Engine 000: Avg prompt throughput: 992.5 tokens/s, Avg generation throughput: 3263.5 tokens/s, Running: 256 reqs, Waiting: 141 reqs, GPU KV cache usage: 52.8%, Prefix cache hit rate: 93.4%
|
| 100 |
+
(APIServer pid=1153466) INFO 08-10 21:27:15 [loggers.py:259] Engine 000: Avg prompt throughput: 697.8 tokens/s, Avg generation throughput: 3293.6 tokens/s, Running: 256 reqs, Waiting: 57 reqs, GPU KV cache usage: 62.2%, Prefix cache hit rate: 93.4%
|
| 101 |
+
(APIServer pid=1153466) INFO 08-10 21:27:25 [loggers.py:259] Engine 000: Avg prompt throughput: 419.7 tokens/s, Avg generation throughput: 3194.6 tokens/s, Running: 256 reqs, Waiting: 5 reqs, GPU KV cache usage: 78.3%, Prefix cache hit rate: 93.4%
|
| 102 |
+
(APIServer pid=1153466) INFO 08-10 21:27:35 [loggers.py:259] Engine 000: Avg prompt throughput: 37.3 tokens/s, Avg generation throughput: 3053.0 tokens/s, Running: 123 reqs, Waiting: 0 reqs, GPU KV cache usage: 42.5%, Prefix cache hit rate: 93.4%
|
| 103 |
+
(APIServer pid=1153466) INFO: 127.0.0.1:55138 - "POST /v1/completions HTTP/1.1" 200 OK
|
| 104 |
+
(EngineCore pid=1154734) INFO 08-10 21:27:45 [core.py:1210] Shutdown initiated (timeout=0)
|
| 105 |
+
(EngineCore pid=1154734) INFO 08-10 21:27:45 [core.py:1233] Shutdown complete
|
| 106 |
+
(APIServer pid=1153466) INFO: Shutting down
|
| 107 |
+
(APIServer pid=1153466) INFO: Waiting for application shutdown.
|
| 108 |
+
(APIServer pid=1153466) INFO: Application shutdown complete.
|
| 109 |
+
(APIServer pid=1153466) INFO: Finished server process [1153466]
|
evals/grid_math_unhealed/glean_keep50_unhealed.log
ADDED
|
@@ -0,0 +1,8 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"correct": 773,
|
| 3 |
+
"accuracy": 0.5860500379075056,
|
| 4 |
+
"finished": 1313,
|
| 5 |
+
"finish_rate": 0.9954510993176648,
|
| 6 |
+
"mean_completion_tokens": 115.38665655799848
|
| 7 |
+
}
|
| 8 |
+
saved item-level results -> outputs/evals/grid_math_unhealed/glean_keep50_unhealed_chat.json
|
evals/grid_math_unhealed/glean_keep50_unhealed_chat.json
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
evals/grid_math_unhealed/glean_keep50_unhealed_chat.json.server.log
ADDED
|
@@ -0,0 +1,102 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
Skipping import of cpp extensions due to incompatible torch version 2.10.0+cu128 for torchao version 0.15.0 Please see https://github.com/pytorch/ao/issues/2919 for more info
|
| 2 |
+
WARNING 08-10 21:22:07 [registry.py:915] Model architecture OlmoeForCausalLM is already registered, and will be overwritten by the new model class glean_vllm.pruned_olmoe:PrunedOlmoeForCausalLM.
|
| 3 |
+
(APIServer pid=1147686) INFO 08-10 21:22:07 [utils.py:299]
|
| 4 |
+
(APIServer pid=1147686) INFO 08-10 21:22:07 [utils.py:299] β β ββ ββ
|
| 5 |
+
(APIServer pid=1147686) INFO 08-10 21:22:07 [utils.py:299] ββ ββ β β β βββ β version 0.19.0
|
| 6 |
+
(APIServer pid=1147686) INFO 08-10 21:22:07 [utils.py:299] ββββ β β β β model outputs/release/unhealed/winnow-olmoe-math-keep50
|
| 7 |
+
(APIServer pid=1147686) INFO 08-10 21:22:07 [utils.py:299] ββ βββββ βββββ β β
|
| 8 |
+
(APIServer pid=1147686) INFO 08-10 21:22:07 [utils.py:299]
|
| 9 |
+
(APIServer pid=1147686) INFO 08-10 21:22:07 [utils.py:233] non-default args: {'model_tag': 'outputs/release/unhealed/winnow-olmoe-math-keep50', 'host': '127.0.0.1', 'port': 8610, 'model': 'outputs/release/unhealed/winnow-olmoe-math-keep50', 'max_model_len': 2048, 'enforce_eager': True, 'served_model_name': ['student'], 'gpu_memory_utilization': 0.85}
|
| 10 |
+
(APIServer pid=1147686) INFO 08-10 21:22:15 [model.py:549] Resolved architecture: PrunedOlmoeForCausalLM
|
| 11 |
+
(APIServer pid=1147686) INFO 08-10 21:22:15 [model.py:1678] Using max model len 2048
|
| 12 |
+
(APIServer pid=1147686) INFO 08-10 21:22:15 [vllm.py:790] Asynchronous scheduling is enabled.
|
| 13 |
+
(APIServer pid=1147686) WARNING 08-10 21:22:15 [vllm.py:848] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
|
| 14 |
+
(APIServer pid=1147686) WARNING 08-10 21:22:15 [vllm.py:859] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
|
| 15 |
+
(APIServer pid=1147686) INFO 08-10 21:22:15 [vllm.py:1025] Cudagraph is disabled under eager mode
|
| 16 |
+
(APIServer pid=1147686) INFO 08-10 21:22:15 [compilation.py:290] Enabled custom fusions: norm_quant, act_quant
|
| 17 |
+
Skipping import of cpp extensions due to incompatible torch version 2.10.0+cu128 for torchao version 0.15.0 Please see https://github.com/pytorch/ao/issues/2919 for more info
|
| 18 |
+
(EngineCore pid=1148963) WARNING 08-10 21:22:24 [registry.py:915] Model architecture OlmoeForCausalLM is already registered, and will be overwritten by the new model class glean_vllm.pruned_olmoe:PrunedOlmoeForCausalLM.
|
| 19 |
+
(EngineCore pid=1148963) INFO 08-10 21:22:24 [core.py:105] Initializing a V1 LLM engine (v0.19.0) with config: model='outputs/release/unhealed/winnow-olmoe-math-keep50', speculative_config=None, tokenizer='outputs/release/unhealed/winnow-olmoe-math-keep50', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=False, dtype=torch.bfloat16, max_seq_len=2048, download_dir=None, load_format=auto, tensor_parallel_size=1, pipeline_parallel_size=1, data_parallel_size=1, decode_context_parallel_size=1, dcp_comm_backend=ag_rs, disable_custom_all_reduce=False, quantization=None, enforce_eager=True, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False), seed=0, served_model_name=student, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'mode': <CompilationMode.NONE: 0>, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'splitting_ops': [], 'compile_mm_encoder': False, 'cudagraph_mm_encoder': False, 'encoder_cudagraph_token_budgets': [], 'encoder_cudagraph_max_images_per_batch': 0, 'compile_sizes': [], 'compile_ranges_endpoints': [2048], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'size_asserts': False, 'alignment_asserts': False, 'scalar_asserts': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': <CUDAGraphMode.NONE: 0>, 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': True, 'fuse_act_quant': True, 'fuse_attn_quant': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False}, 'max_cudagraph_capture_size': 0, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': True, 'static_all_moe_layers': []}
|
| 20 |
+
(EngineCore pid=1148963) INFO 08-10 21:22:24 [parallel_state.py:1400] world_size=1 rank=0 local_rank=0 distributed_init_method=tcp://192.168.0.27:43405 backend=nccl
|
| 21 |
+
(EngineCore pid=1148963) INFO 08-10 21:22:24 [parallel_state.py:1716] rank 0 in world size 1 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank 0, EPLB rank N/A
|
| 22 |
+
(EngineCore pid=1148963) INFO 08-10 21:22:25 [gpu_model_runner.py:4735] Starting to load model outputs/release/unhealed/winnow-olmoe-math-keep50...
|
| 23 |
+
(EngineCore pid=1148963) INFO 08-10 21:22:25 [cuda.py:334] Using FLASH_ATTN attention backend out of potential backends: ['FLASH_ATTN', 'FLASHINFER', 'TRITON_ATTN', 'FLEX_ATTENTION'].
|
| 24 |
+
(EngineCore pid=1148963) INFO 08-10 21:22:25 [flash_attn.py:596] Using FlashAttention version 2
|
| 25 |
+
(EngineCore pid=1148963)
|
| 26 |
+
(EngineCore pid=1148963)
|
| 27 |
+
(EngineCore pid=1148963)
|
| 28 |
+
(EngineCore pid=1148963)
|
| 29 |
+
(EngineCore pid=1148963)
|
| 30 |
+
(EngineCore pid=1148963) INFO 08-10 21:22:27 [default_loader.py:384] Loading weights took 1.42 seconds
|
| 31 |
+
(EngineCore pid=1148963) INFO 08-10 21:22:27 [gpu_model_runner.py:4820] Model loading took 6.89 GiB memory and 2.012000 seconds
|
| 32 |
+
(EngineCore pid=1148963) INFO 08-10 21:22:29 [gpu_worker.py:436] Available KV cache memory: 12.86 GiB
|
| 33 |
+
(EngineCore pid=1148963) INFO 08-10 21:22:29 [kv_cache_utils.py:1319] GPU KV cache size: 105,328 tokens
|
| 34 |
+
(EngineCore pid=1148963) INFO 08-10 21:22:29 [kv_cache_utils.py:1324] Maximum concurrency for 2,048 tokens per request: 51.43x
|
| 35 |
+
(EngineCore pid=1148963) INFO 08-10 21:22:29 [core.py:283] init engine (profile, create kv cache, warmup model) took 1.80 seconds
|
| 36 |
+
(EngineCore pid=1148963) INFO 08-10 21:22:30 [vllm.py:790] Asynchronous scheduling is enabled.
|
| 37 |
+
(EngineCore pid=1148963) WARNING 08-10 21:22:30 [vllm.py:848] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
|
| 38 |
+
(EngineCore pid=1148963) WARNING 08-10 21:22:30 [vllm.py:859] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
|
| 39 |
+
(EngineCore pid=1148963) INFO 08-10 21:22:30 [vllm.py:1025] Cudagraph is disabled under eager mode
|
| 40 |
+
(EngineCore pid=1148963) INFO 08-10 21:22:30 [compilation.py:290] Enabled custom fusions: norm_quant, act_quant
|
| 41 |
+
(APIServer pid=1147686) INFO 08-10 21:22:30 [api_server.py:590] Supported tasks: ['generate']
|
| 42 |
+
(APIServer pid=1147686) WARNING 08-10 21:22:30 [__init__.py:14] SECURITY WARNING: Development endpoints are enabled! This should NOT be used in production!
|
| 43 |
+
(APIServer pid=1147686) INFO 08-10 21:22:30 [hf.py:314] Detected the chat template content format to be 'string'. You can set `--chat-template-content-format` to override this.
|
| 44 |
+
(APIServer pid=1147686) INFO 08-10 21:22:30 [api_server.py:594] Starting vLLM server on http://127.0.0.1:8610
|
| 45 |
+
(APIServer pid=1147686) INFO 08-10 21:22:30 [launcher.py:37] Available routes are:
|
| 46 |
+
(APIServer pid=1147686) INFO 08-10 21:22:30 [launcher.py:46] Route: /openapi.json, Methods: HEAD, GET
|
| 47 |
+
(APIServer pid=1147686) INFO 08-10 21:22:30 [launcher.py:46] Route: /docs, Methods: HEAD, GET
|
| 48 |
+
(APIServer pid=1147686) INFO 08-10 21:22:30 [launcher.py:46] Route: /docs/oauth2-redirect, Methods: HEAD, GET
|
| 49 |
+
(APIServer pid=1147686) INFO 08-10 21:22:30 [launcher.py:46] Route: /redoc, Methods: HEAD, GET
|
| 50 |
+
(APIServer pid=1147686) INFO 08-10 21:22:30 [launcher.py:46] Route: /sleep, Methods: POST
|
| 51 |
+
(APIServer pid=1147686) INFO 08-10 21:22:30 [launcher.py:46] Route: /wake_up, Methods: POST
|
| 52 |
+
(APIServer pid=1147686) INFO 08-10 21:22:30 [launcher.py:46] Route: /is_sleeping, Methods: GET
|
| 53 |
+
(APIServer pid=1147686) INFO 08-10 21:22:30 [launcher.py:46] Route: /collective_rpc, Methods: POST
|
| 54 |
+
(APIServer pid=1147686) INFO 08-10 21:22:30 [launcher.py:46] Route: /reset_prefix_cache, Methods: POST
|
| 55 |
+
(APIServer pid=1147686) INFO 08-10 21:22:30 [launcher.py:46] Route: /reset_mm_cache, Methods: POST
|
| 56 |
+
(APIServer pid=1147686) INFO 08-10 21:22:30 [launcher.py:46] Route: /reset_encoder_cache, Methods: POST
|
| 57 |
+
(APIServer pid=1147686) INFO 08-10 21:22:30 [launcher.py:46] Route: /tokenize, Methods: POST
|
| 58 |
+
(APIServer pid=1147686) INFO 08-10 21:22:30 [launcher.py:46] Route: /detokenize, Methods: POST
|
| 59 |
+
(APIServer pid=1147686) INFO 08-10 21:22:30 [launcher.py:46] Route: /load, Methods: GET
|
| 60 |
+
(APIServer pid=1147686) INFO 08-10 21:22:30 [launcher.py:46] Route: /version, Methods: GET
|
| 61 |
+
(APIServer pid=1147686) INFO 08-10 21:22:30 [launcher.py:46] Route: /health, Methods: GET
|
| 62 |
+
(APIServer pid=1147686) INFO 08-10 21:22:30 [launcher.py:46] Route: /metrics, Methods: GET
|
| 63 |
+
(APIServer pid=1147686) INFO 08-10 21:22:30 [launcher.py:46] Route: /server_info, Methods: GET
|
| 64 |
+
(APIServer pid=1147686) INFO 08-10 21:22:30 [launcher.py:46] Route: /v1/models, Methods: GET
|
| 65 |
+
(APIServer pid=1147686) INFO 08-10 21:22:30 [launcher.py:46] Route: /ping, Methods: GET
|
| 66 |
+
(APIServer pid=1147686) INFO 08-10 21:22:30 [launcher.py:46] Route: /ping, Methods: POST
|
| 67 |
+
(APIServer pid=1147686) INFO 08-10 21:22:30 [launcher.py:46] Route: /invocations, Methods: POST
|
| 68 |
+
(APIServer pid=1147686) INFO 08-10 21:22:30 [launcher.py:46] Route: /v1/chat/completions, Methods: POST
|
| 69 |
+
(APIServer pid=1147686) INFO 08-10 21:22:30 [launcher.py:46] Route: /v1/chat/completions/batch, Methods: POST
|
| 70 |
+
(APIServer pid=1147686) INFO 08-10 21:22:30 [launcher.py:46] Route: /v1/responses, Methods: POST
|
| 71 |
+
(APIServer pid=1147686) INFO 08-10 21:22:30 [launcher.py:46] Route: /v1/responses/{response_id}, Methods: GET
|
| 72 |
+
(APIServer pid=1147686) INFO 08-10 21:22:30 [launcher.py:46] Route: /v1/responses/{response_id}/cancel, Methods: POST
|
| 73 |
+
(APIServer pid=1147686) INFO 08-10 21:22:30 [launcher.py:46] Route: /v1/completions, Methods: POST
|
| 74 |
+
(APIServer pid=1147686) INFO 08-10 21:22:30 [launcher.py:46] Route: /v1/messages, Methods: POST
|
| 75 |
+
(APIServer pid=1147686) INFO 08-10 21:22:30 [launcher.py:46] Route: /v1/messages/count_tokens, Methods: POST
|
| 76 |
+
(APIServer pid=1147686) INFO 08-10 21:22:30 [launcher.py:46] Route: /inference/v1/generate, Methods: POST
|
| 77 |
+
(APIServer pid=1147686) INFO 08-10 21:22:30 [launcher.py:46] Route: /pause, Methods: POST
|
| 78 |
+
(APIServer pid=1147686) INFO 08-10 21:22:30 [launcher.py:46] Route: /resume, Methods: POST
|
| 79 |
+
(APIServer pid=1147686) INFO 08-10 21:22:30 [launcher.py:46] Route: /is_paused, Methods: GET
|
| 80 |
+
(APIServer pid=1147686) INFO 08-10 21:22:30 [launcher.py:46] Route: /init_weight_transfer_engine, Methods: POST
|
| 81 |
+
(APIServer pid=1147686) INFO 08-10 21:22:30 [launcher.py:46] Route: /update_weights, Methods: POST
|
| 82 |
+
(APIServer pid=1147686) INFO 08-10 21:22:30 [launcher.py:46] Route: /get_world_size, Methods: GET
|
| 83 |
+
(APIServer pid=1147686) INFO 08-10 21:22:30 [launcher.py:46] Route: /scale_elastic_ep, Methods: POST
|
| 84 |
+
(APIServer pid=1147686) INFO 08-10 21:22:30 [launcher.py:46] Route: /is_scaling_elastic_ep, Methods: POST
|
| 85 |
+
(APIServer pid=1147686) INFO 08-10 21:22:30 [launcher.py:46] Route: /v1/chat/completions/render, Methods: POST
|
| 86 |
+
(APIServer pid=1147686) INFO 08-10 21:22:30 [launcher.py:46] Route: /v1/completions/render, Methods: POST
|
| 87 |
+
(APIServer pid=1147686) INFO: Started server process [1147686]
|
| 88 |
+
(APIServer pid=1147686) INFO: Waiting for application startup.
|
| 89 |
+
(APIServer pid=1147686) INFO: Application startup complete.
|
| 90 |
+
(APIServer pid=1147686) INFO: 127.0.0.1:37708 - "GET /health HTTP/1.1" 200 OK
|
| 91 |
+
(APIServer pid=1147686) INFO 08-10 21:22:40 [loggers.py:259] Engine 000: Avg prompt throughput: 2662.8 tokens/s, Avg generation throughput: 2173.5 tokens/s, Running: 253 reqs, Waiting: 972 reqs, GPU KV cache usage: 38.1%, Prefix cache hit rate: 93.2%
|
| 92 |
+
(APIServer pid=1147686) INFO 08-10 21:22:50 [loggers.py:259] Engine 000: Avg prompt throughput: 2160.5 tokens/s, Avg generation throughput: 3146.4 tokens/s, Running: 252 reqs, Waiting: 696 reqs, GPU KV cache usage: 39.2%, Prefix cache hit rate: 93.3%
|
| 93 |
+
(APIServer pid=1147686) INFO 08-10 21:23:00 [loggers.py:259] Engine 000: Avg prompt throughput: 2245.8 tokens/s, Avg generation throughput: 3093.6 tokens/s, Running: 255 reqs, Waiting: 409 reqs, GPU KV cache usage: 38.2%, Prefix cache hit rate: 93.4%
|
| 94 |
+
(APIServer pid=1147686) INFO 08-10 21:23:10 [loggers.py:259] Engine 000: Avg prompt throughput: 2095.3 tokens/s, Avg generation throughput: 3147.7 tokens/s, Running: 253 reqs, Waiting: 149 reqs, GPU KV cache usage: 40.2%, Prefix cache hit rate: 93.4%
|
| 95 |
+
(APIServer pid=1147686) INFO 08-10 21:23:20 [loggers.py:259] Engine 000: Avg prompt throughput: 1232.5 tokens/s, Avg generation throughput: 3092.0 tokens/s, Running: 88 reqs, Waiting: 0 reqs, GPU KV cache usage: 20.9%, Prefix cache hit rate: 93.4%
|
| 96 |
+
(APIServer pid=1147686) INFO: 127.0.0.1:37722 - "POST /v1/completions HTTP/1.1" 200 OK
|
| 97 |
+
(EngineCore pid=1148963) INFO 08-10 21:23:27 [core.py:1210] Shutdown initiated (timeout=0)
|
| 98 |
+
(EngineCore pid=1148963) INFO 08-10 21:23:27 [core.py:1233] Shutdown complete
|
| 99 |
+
(APIServer pid=1147686) INFO: Shutting down
|
| 100 |
+
(APIServer pid=1147686) INFO: Waiting for application shutdown.
|
| 101 |
+
(APIServer pid=1147686) INFO: Application shutdown complete.
|
| 102 |
+
(APIServer pid=1147686) INFO: Finished server process [1147686]
|
evals/grid_math_unhealed/glean_keep75_unhealed.log
ADDED
|
@@ -0,0 +1,8 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"correct": 898,
|
| 3 |
+
"accuracy": 0.6808188021228203,
|
| 4 |
+
"finished": 1317,
|
| 5 |
+
"finish_rate": 0.9984836997725549,
|
| 6 |
+
"mean_completion_tokens": 109.84154662623199
|
| 7 |
+
}
|
| 8 |
+
saved item-level results -> outputs/evals/grid_math_unhealed/glean_keep75_unhealed_chat.json
|
evals/grid_math_unhealed/glean_keep75_unhealed_chat.json
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
evals/grid_math_unhealed/glean_keep75_unhealed_chat.json.server.log
ADDED
|
@@ -0,0 +1,103 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
Skipping import of cpp extensions due to incompatible torch version 2.10.0+cu128 for torchao version 0.15.0 Please see https://github.com/pytorch/ao/issues/2919 for more info
|
| 2 |
+
WARNING 08-10 21:20:11 [registry.py:915] Model architecture OlmoeForCausalLM is already registered, and will be overwritten by the new model class glean_vllm.pruned_olmoe:PrunedOlmoeForCausalLM.
|
| 3 |
+
(APIServer pid=1142912) INFO 08-10 21:20:11 [utils.py:299]
|
| 4 |
+
(APIServer pid=1142912) INFO 08-10 21:20:11 [utils.py:299] β β ββ ββ
|
| 5 |
+
(APIServer pid=1142912) INFO 08-10 21:20:11 [utils.py:299] ββ ββ β β β βββ β version 0.19.0
|
| 6 |
+
(APIServer pid=1142912) INFO 08-10 21:20:11 [utils.py:299] ββββ β β β β model outputs/release/unhealed/winnow-olmoe-math-keep75
|
| 7 |
+
(APIServer pid=1142912) INFO 08-10 21:20:11 [utils.py:299] ββ βββββ βββββ β β
|
| 8 |
+
(APIServer pid=1142912) INFO 08-10 21:20:11 [utils.py:299]
|
| 9 |
+
(APIServer pid=1142912) INFO 08-10 21:20:11 [utils.py:233] non-default args: {'model_tag': 'outputs/release/unhealed/winnow-olmoe-math-keep75', 'host': '127.0.0.1', 'port': 8610, 'model': 'outputs/release/unhealed/winnow-olmoe-math-keep75', 'max_model_len': 2048, 'enforce_eager': True, 'served_model_name': ['student'], 'gpu_memory_utilization': 0.85}
|
| 10 |
+
(APIServer pid=1142912) INFO 08-10 21:20:21 [model.py:549] Resolved architecture: PrunedOlmoeForCausalLM
|
| 11 |
+
(APIServer pid=1142912) INFO 08-10 21:20:21 [model.py:1678] Using max model len 2048
|
| 12 |
+
(APIServer pid=1142912) INFO 08-10 21:20:22 [vllm.py:790] Asynchronous scheduling is enabled.
|
| 13 |
+
(APIServer pid=1142912) WARNING 08-10 21:20:22 [vllm.py:848] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
|
| 14 |
+
(APIServer pid=1142912) WARNING 08-10 21:20:22 [vllm.py:859] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
|
| 15 |
+
(APIServer pid=1142912) INFO 08-10 21:20:22 [vllm.py:1025] Cudagraph is disabled under eager mode
|
| 16 |
+
(APIServer pid=1142912) INFO 08-10 21:20:22 [compilation.py:290] Enabled custom fusions: norm_quant, act_quant
|
| 17 |
+
Skipping import of cpp extensions due to incompatible torch version 2.10.0+cu128 for torchao version 0.15.0 Please see https://github.com/pytorch/ao/issues/2919 for more info
|
| 18 |
+
(EngineCore pid=1144303) WARNING 08-10 21:20:30 [registry.py:915] Model architecture OlmoeForCausalLM is already registered, and will be overwritten by the new model class glean_vllm.pruned_olmoe:PrunedOlmoeForCausalLM.
|
| 19 |
+
(EngineCore pid=1144303) INFO 08-10 21:20:30 [core.py:105] Initializing a V1 LLM engine (v0.19.0) with config: model='outputs/release/unhealed/winnow-olmoe-math-keep75', speculative_config=None, tokenizer='outputs/release/unhealed/winnow-olmoe-math-keep75', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=False, dtype=torch.bfloat16, max_seq_len=2048, download_dir=None, load_format=auto, tensor_parallel_size=1, pipeline_parallel_size=1, data_parallel_size=1, decode_context_parallel_size=1, dcp_comm_backend=ag_rs, disable_custom_all_reduce=False, quantization=None, enforce_eager=True, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False), seed=0, served_model_name=student, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'mode': <CompilationMode.NONE: 0>, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'splitting_ops': [], 'compile_mm_encoder': False, 'cudagraph_mm_encoder': False, 'encoder_cudagraph_token_budgets': [], 'encoder_cudagraph_max_images_per_batch': 0, 'compile_sizes': [], 'compile_ranges_endpoints': [2048], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'size_asserts': False, 'alignment_asserts': False, 'scalar_asserts': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': <CUDAGraphMode.NONE: 0>, 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': True, 'fuse_act_quant': True, 'fuse_attn_quant': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False}, 'max_cudagraph_capture_size': 0, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': True, 'static_all_moe_layers': []}
|
| 20 |
+
(EngineCore pid=1144303) INFO 08-10 21:20:31 [parallel_state.py:1400] world_size=1 rank=0 local_rank=0 distributed_init_method=tcp://192.168.0.27:53323 backend=nccl
|
| 21 |
+
(EngineCore pid=1144303) INFO 08-10 21:20:31 [parallel_state.py:1716] rank 0 in world size 1 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank 0, EPLB rank N/A
|
| 22 |
+
(EngineCore pid=1144303) INFO 08-10 21:20:31 [gpu_model_runner.py:4735] Starting to load model outputs/release/unhealed/winnow-olmoe-math-keep75...
|
| 23 |
+
(EngineCore pid=1144303) INFO 08-10 21:20:33 [cuda.py:334] Using FLASH_ATTN attention backend out of potential backends: ['FLASH_ATTN', 'FLASHINFER', 'TRITON_ATTN', 'FLEX_ATTENTION'].
|
| 24 |
+
(EngineCore pid=1144303) INFO 08-10 21:20:33 [flash_attn.py:596] Using FlashAttention version 2
|
| 25 |
+
(EngineCore pid=1144303)
|
| 26 |
+
(EngineCore pid=1144303)
|
| 27 |
+
(EngineCore pid=1144303)
|
| 28 |
+
(EngineCore pid=1144303)
|
| 29 |
+
(EngineCore pid=1144303)
|
| 30 |
+
(EngineCore pid=1144303)
|
| 31 |
+
(EngineCore pid=1144303) INFO 08-10 21:20:35 [default_loader.py:384] Loading weights took 1.88 seconds
|
| 32 |
+
(EngineCore pid=1144303) INFO 08-10 21:20:35 [gpu_model_runner.py:4820] Model loading took 9.89 GiB memory and 3.012713 seconds
|
| 33 |
+
(EngineCore pid=1144303) INFO 08-10 21:20:39 [gpu_worker.py:436] Available KV cache memory: 9.86 GiB
|
| 34 |
+
(EngineCore pid=1144303) INFO 08-10 21:20:39 [kv_cache_utils.py:1319] GPU KV cache size: 80,752 tokens
|
| 35 |
+
(EngineCore pid=1144303) INFO 08-10 21:20:39 [kv_cache_utils.py:1324] Maximum concurrency for 2,048 tokens per request: 39.43x
|
| 36 |
+
(EngineCore pid=1144303) INFO 08-10 21:20:39 [core.py:283] init engine (profile, create kv cache, warmup model) took 3.97 seconds
|
| 37 |
+
(EngineCore pid=1144303) INFO 08-10 21:20:40 [vllm.py:790] Asynchronous scheduling is enabled.
|
| 38 |
+
(EngineCore pid=1144303) WARNING 08-10 21:20:40 [vllm.py:848] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
|
| 39 |
+
(EngineCore pid=1144303) WARNING 08-10 21:20:40 [vllm.py:859] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
|
| 40 |
+
(EngineCore pid=1144303) INFO 08-10 21:20:40 [vllm.py:1025] Cudagraph is disabled under eager mode
|
| 41 |
+
(EngineCore pid=1144303) INFO 08-10 21:20:40 [compilation.py:290] Enabled custom fusions: norm_quant, act_quant
|
| 42 |
+
(APIServer pid=1142912) INFO 08-10 21:20:40 [api_server.py:590] Supported tasks: ['generate']
|
| 43 |
+
(APIServer pid=1142912) WARNING 08-10 21:20:40 [__init__.py:14] SECURITY WARNING: Development endpoints are enabled! This should NOT be used in production!
|
| 44 |
+
(APIServer pid=1142912) INFO 08-10 21:20:40 [hf.py:314] Detected the chat template content format to be 'string'. You can set `--chat-template-content-format` to override this.
|
| 45 |
+
(APIServer pid=1142912) INFO 08-10 21:20:40 [api_server.py:594] Starting vLLM server on http://127.0.0.1:8610
|
| 46 |
+
(APIServer pid=1142912) INFO 08-10 21:20:40 [launcher.py:37] Available routes are:
|
| 47 |
+
(APIServer pid=1142912) INFO 08-10 21:20:40 [launcher.py:46] Route: /openapi.json, Methods: GET, HEAD
|
| 48 |
+
(APIServer pid=1142912) INFO 08-10 21:20:40 [launcher.py:46] Route: /docs, Methods: GET, HEAD
|
| 49 |
+
(APIServer pid=1142912) INFO 08-10 21:20:40 [launcher.py:46] Route: /docs/oauth2-redirect, Methods: GET, HEAD
|
| 50 |
+
(APIServer pid=1142912) INFO 08-10 21:20:40 [launcher.py:46] Route: /redoc, Methods: GET, HEAD
|
| 51 |
+
(APIServer pid=1142912) INFO 08-10 21:20:40 [launcher.py:46] Route: /sleep, Methods: POST
|
| 52 |
+
(APIServer pid=1142912) INFO 08-10 21:20:40 [launcher.py:46] Route: /wake_up, Methods: POST
|
| 53 |
+
(APIServer pid=1142912) INFO 08-10 21:20:40 [launcher.py:46] Route: /is_sleeping, Methods: GET
|
| 54 |
+
(APIServer pid=1142912) INFO 08-10 21:20:40 [launcher.py:46] Route: /collective_rpc, Methods: POST
|
| 55 |
+
(APIServer pid=1142912) INFO 08-10 21:20:40 [launcher.py:46] Route: /reset_prefix_cache, Methods: POST
|
| 56 |
+
(APIServer pid=1142912) INFO 08-10 21:20:40 [launcher.py:46] Route: /reset_mm_cache, Methods: POST
|
| 57 |
+
(APIServer pid=1142912) INFO 08-10 21:20:40 [launcher.py:46] Route: /reset_encoder_cache, Methods: POST
|
| 58 |
+
(APIServer pid=1142912) INFO 08-10 21:20:40 [launcher.py:46] Route: /tokenize, Methods: POST
|
| 59 |
+
(APIServer pid=1142912) INFO 08-10 21:20:40 [launcher.py:46] Route: /detokenize, Methods: POST
|
| 60 |
+
(APIServer pid=1142912) INFO 08-10 21:20:40 [launcher.py:46] Route: /load, Methods: GET
|
| 61 |
+
(APIServer pid=1142912) INFO 08-10 21:20:40 [launcher.py:46] Route: /version, Methods: GET
|
| 62 |
+
(APIServer pid=1142912) INFO 08-10 21:20:40 [launcher.py:46] Route: /health, Methods: GET
|
| 63 |
+
(APIServer pid=1142912) INFO 08-10 21:20:40 [launcher.py:46] Route: /metrics, Methods: GET
|
| 64 |
+
(APIServer pid=1142912) INFO 08-10 21:20:40 [launcher.py:46] Route: /server_info, Methods: GET
|
| 65 |
+
(APIServer pid=1142912) INFO 08-10 21:20:40 [launcher.py:46] Route: /v1/models, Methods: GET
|
| 66 |
+
(APIServer pid=1142912) INFO 08-10 21:20:40 [launcher.py:46] Route: /ping, Methods: GET
|
| 67 |
+
(APIServer pid=1142912) INFO 08-10 21:20:40 [launcher.py:46] Route: /ping, Methods: POST
|
| 68 |
+
(APIServer pid=1142912) INFO 08-10 21:20:40 [launcher.py:46] Route: /invocations, Methods: POST
|
| 69 |
+
(APIServer pid=1142912) INFO 08-10 21:20:40 [launcher.py:46] Route: /v1/chat/completions, Methods: POST
|
| 70 |
+
(APIServer pid=1142912) INFO 08-10 21:20:40 [launcher.py:46] Route: /v1/chat/completions/batch, Methods: POST
|
| 71 |
+
(APIServer pid=1142912) INFO 08-10 21:20:40 [launcher.py:46] Route: /v1/responses, Methods: POST
|
| 72 |
+
(APIServer pid=1142912) INFO 08-10 21:20:40 [launcher.py:46] Route: /v1/responses/{response_id}, Methods: GET
|
| 73 |
+
(APIServer pid=1142912) INFO 08-10 21:20:40 [launcher.py:46] Route: /v1/responses/{response_id}/cancel, Methods: POST
|
| 74 |
+
(APIServer pid=1142912) INFO 08-10 21:20:40 [launcher.py:46] Route: /v1/completions, Methods: POST
|
| 75 |
+
(APIServer pid=1142912) INFO 08-10 21:20:40 [launcher.py:46] Route: /v1/messages, Methods: POST
|
| 76 |
+
(APIServer pid=1142912) INFO 08-10 21:20:40 [launcher.py:46] Route: /v1/messages/count_tokens, Methods: POST
|
| 77 |
+
(APIServer pid=1142912) INFO 08-10 21:20:40 [launcher.py:46] Route: /inference/v1/generate, Methods: POST
|
| 78 |
+
(APIServer pid=1142912) INFO 08-10 21:20:40 [launcher.py:46] Route: /pause, Methods: POST
|
| 79 |
+
(APIServer pid=1142912) INFO 08-10 21:20:40 [launcher.py:46] Route: /resume, Methods: POST
|
| 80 |
+
(APIServer pid=1142912) INFO 08-10 21:20:40 [launcher.py:46] Route: /is_paused, Methods: GET
|
| 81 |
+
(APIServer pid=1142912) INFO 08-10 21:20:40 [launcher.py:46] Route: /init_weight_transfer_engine, Methods: POST
|
| 82 |
+
(APIServer pid=1142912) INFO 08-10 21:20:40 [launcher.py:46] Route: /update_weights, Methods: POST
|
| 83 |
+
(APIServer pid=1142912) INFO 08-10 21:20:40 [launcher.py:46] Route: /get_world_size, Methods: GET
|
| 84 |
+
(APIServer pid=1142912) INFO 08-10 21:20:40 [launcher.py:46] Route: /scale_elastic_ep, Methods: POST
|
| 85 |
+
(APIServer pid=1142912) INFO 08-10 21:20:40 [launcher.py:46] Route: /is_scaling_elastic_ep, Methods: POST
|
| 86 |
+
(APIServer pid=1142912) INFO 08-10 21:20:40 [launcher.py:46] Route: /v1/chat/completions/render, Methods: POST
|
| 87 |
+
(APIServer pid=1142912) INFO 08-10 21:20:40 [launcher.py:46] Route: /v1/completions/render, Methods: POST
|
| 88 |
+
(APIServer pid=1142912) INFO: Started server process [1142912]
|
| 89 |
+
(APIServer pid=1142912) INFO: Waiting for application startup.
|
| 90 |
+
(APIServer pid=1142912) INFO: Application startup complete.
|
| 91 |
+
(APIServer pid=1142912) INFO: 127.0.0.1:58890 - "GET /health HTTP/1.1" 200 OK
|
| 92 |
+
(APIServer pid=1142912) INFO 08-10 21:20:50 [loggers.py:259] Engine 000: Avg prompt throughput: 2625.2 tokens/s, Avg generation throughput: 1944.8 tokens/s, Running: 254 reqs, Waiting: 978 reqs, GPU KV cache usage: 48.0%, Prefix cache hit rate: 93.2%
|
| 93 |
+
(APIServer pid=1142912) INFO 08-10 21:21:00 [loggers.py:259] Engine 000: Avg prompt throughput: 2095.0 tokens/s, Avg generation throughput: 2917.1 tokens/s, Running: 255 reqs, Waiting: 710 reqs, GPU KV cache usage: 49.7%, Prefix cache hit rate: 93.3%
|
| 94 |
+
(APIServer pid=1142912) INFO 08-10 21:21:10 [loggers.py:259] Engine 000: Avg prompt throughput: 2228.9 tokens/s, Avg generation throughput: 2914.7 tokens/s, Running: 256 reqs, Waiting: 428 reqs, GPU KV cache usage: 49.1%, Prefix cache hit rate: 93.4%
|
| 95 |
+
(APIServer pid=1142912) INFO 08-10 21:21:20 [loggers.py:259] Engine 000: Avg prompt throughput: 2113.2 tokens/s, Avg generation throughput: 2942.7 tokens/s, Running: 255 reqs, Waiting: 164 reqs, GPU KV cache usage: 50.1%, Prefix cache hit rate: 93.4%
|
| 96 |
+
(APIServer pid=1142912) INFO 08-10 21:21:30 [loggers.py:259] Engine 000: Avg prompt throughput: 1354.3 tokens/s, Avg generation throughput: 2657.7 tokens/s, Running: 171 reqs, Waiting: 0 reqs, GPU KV cache usage: 40.4%, Prefix cache hit rate: 93.4%
|
| 97 |
+
(APIServer pid=1142912) INFO: 127.0.0.1:58892 - "POST /v1/completions HTTP/1.1" 200 OK
|
| 98 |
+
(EngineCore pid=1144303) INFO 08-10 21:21:39 [core.py:1210] Shutdown initiated (timeout=0)
|
| 99 |
+
(EngineCore pid=1144303) INFO 08-10 21:21:39 [core.py:1233] Shutdown complete
|
| 100 |
+
(APIServer pid=1142912) INFO: Shutting down
|
| 101 |
+
(APIServer pid=1142912) INFO: Waiting for application shutdown.
|
| 102 |
+
(APIServer pid=1142912) INFO: Application shutdown complete.
|
| 103 |
+
(APIServer pid=1142912) INFO: Finished server process [1142912]
|
evals/grid_math_unhealed/reap_keep25_unhealed.log
ADDED
|
@@ -0,0 +1,8 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"correct": 13,
|
| 3 |
+
"accuracy": 0.009855951478392721,
|
| 4 |
+
"finished": 254,
|
| 5 |
+
"finish_rate": 0.19257012888551933,
|
| 6 |
+
"mean_completion_tokens": 455.9560272934041
|
| 7 |
+
}
|
| 8 |
+
saved item-level results -> outputs/evals/grid_math_unhealed/reap_keep25_unhealed_chat.json
|
evals/grid_math_unhealed/reap_keep25_unhealed_chat.json
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
evals/grid_math_unhealed/reap_keep25_unhealed_chat.json.server.log
ADDED
|
@@ -0,0 +1,117 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
Skipping import of cpp extensions due to incompatible torch version 2.10.0+cu128 for torchao version 0.15.0 Please see https://github.com/pytorch/ao/issues/2919 for more info
|
| 2 |
+
WARNING 08-10 21:24:57 [registry.py:915] Model architecture OlmoeForCausalLM is already registered, and will be overwritten by the new model class glean_vllm.pruned_olmoe:PrunedOlmoeForCausalLM.
|
| 3 |
+
(APIServer pid=1153469) INFO 08-10 21:24:57 [utils.py:299]
|
| 4 |
+
(APIServer pid=1153469) INFO 08-10 21:24:57 [utils.py:299] β β ββ ββ
|
| 5 |
+
(APIServer pid=1153469) INFO 08-10 21:24:57 [utils.py:299] ββ ββ β β β βββ β version 0.19.0
|
| 6 |
+
(APIServer pid=1153469) INFO 08-10 21:24:57 [utils.py:299] ββββ β β β β model /media/henry/MoreFiles/winnow_release/math_arms/reap-math-keep25
|
| 7 |
+
(APIServer pid=1153469) INFO 08-10 21:24:57 [utils.py:299] ββ βββββ βββββ β β
|
| 8 |
+
(APIServer pid=1153469) INFO 08-10 21:24:57 [utils.py:299]
|
| 9 |
+
(APIServer pid=1153469) INFO 08-10 21:24:57 [utils.py:233] non-default args: {'model_tag': '/media/henry/MoreFiles/winnow_release/math_arms/reap-math-keep25', 'host': '127.0.0.1', 'port': 8611, 'model': '/media/henry/MoreFiles/winnow_release/math_arms/reap-math-keep25', 'max_model_len': 2048, 'enforce_eager': True, 'served_model_name': ['student'], 'gpu_memory_utilization': 0.85}
|
| 10 |
+
(APIServer pid=1153469) INFO 08-10 21:25:06 [model.py:549] Resolved architecture: OlmoeForCausalLM
|
| 11 |
+
(APIServer pid=1153469) INFO 08-10 21:25:06 [model.py:1678] Using max model len 2048
|
| 12 |
+
(APIServer pid=1153469) INFO 08-10 21:25:06 [vllm.py:790] Asynchronous scheduling is enabled.
|
| 13 |
+
(APIServer pid=1153469) WARNING 08-10 21:25:06 [vllm.py:848] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
|
| 14 |
+
(APIServer pid=1153469) WARNING 08-10 21:25:06 [vllm.py:859] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
|
| 15 |
+
(APIServer pid=1153469) INFO 08-10 21:25:06 [vllm.py:1025] Cudagraph is disabled under eager mode
|
| 16 |
+
(APIServer pid=1153469) INFO 08-10 21:25:06 [compilation.py:290] Enabled custom fusions: norm_quant, act_quant
|
| 17 |
+
Skipping import of cpp extensions due to incompatible torch version 2.10.0+cu128 for torchao version 0.15.0 Please see https://github.com/pytorch/ao/issues/2919 for more info
|
| 18 |
+
(EngineCore pid=1154752) WARNING 08-10 21:25:14 [registry.py:915] Model architecture OlmoeForCausalLM is already registered, and will be overwritten by the new model class glean_vllm.pruned_olmoe:PrunedOlmoeForCausalLM.
|
| 19 |
+
(EngineCore pid=1154752) INFO 08-10 21:25:14 [core.py:105] Initializing a V1 LLM engine (v0.19.0) with config: model='/media/henry/MoreFiles/winnow_release/math_arms/reap-math-keep25', speculative_config=None, tokenizer='/media/henry/MoreFiles/winnow_release/math_arms/reap-math-keep25', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=False, dtype=torch.bfloat16, max_seq_len=2048, download_dir=None, load_format=auto, tensor_parallel_size=1, pipeline_parallel_size=1, data_parallel_size=1, decode_context_parallel_size=1, dcp_comm_backend=ag_rs, disable_custom_all_reduce=False, quantization=None, enforce_eager=True, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False), seed=0, served_model_name=student, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'mode': <CompilationMode.NONE: 0>, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'splitting_ops': [], 'compile_mm_encoder': False, 'cudagraph_mm_encoder': False, 'encoder_cudagraph_token_budgets': [], 'encoder_cudagraph_max_images_per_batch': 0, 'compile_sizes': [], 'compile_ranges_endpoints': [2048], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'size_asserts': False, 'alignment_asserts': False, 'scalar_asserts': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': <CUDAGraphMode.NONE: 0>, 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': True, 'fuse_act_quant': True, 'fuse_attn_quant': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False}, 'max_cudagraph_capture_size': 0, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': True, 'static_all_moe_layers': []}
|
| 20 |
+
(EngineCore pid=1154752) INFO 08-10 21:25:14 [parallel_state.py:1400] world_size=1 rank=0 local_rank=0 distributed_init_method=tcp://192.168.0.27:39421 backend=nccl
|
| 21 |
+
(EngineCore pid=1154752) INFO 08-10 21:25:14 [parallel_state.py:1716] rank 0 in world size 1 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank 0, EPLB rank N/A
|
| 22 |
+
(EngineCore pid=1154752) INFO 08-10 21:25:15 [gpu_model_runner.py:4735] Starting to load model /media/henry/MoreFiles/winnow_release/math_arms/reap-math-keep25...
|
| 23 |
+
(EngineCore pid=1154752) INFO 08-10 21:25:16 [cuda.py:334] Using FLASH_ATTN attention backend out of potential backends: ['FLASH_ATTN', 'FLASHINFER', 'TRITON_ATTN', 'FLEX_ATTENTION'].
|
| 24 |
+
(EngineCore pid=1154752) INFO 08-10 21:25:16 [flash_attn.py:596] Using FlashAttention version 2
|
| 25 |
+
(EngineCore pid=1154752) INFO 08-10 21:25:16 [unquantized.py:186] Using TRITON backend for Unquantized MoE
|
| 26 |
+
(EngineCore pid=1154752)
|
| 27 |
+
(EngineCore pid=1154752)
|
| 28 |
+
(EngineCore pid=1154752)
|
| 29 |
+
(EngineCore pid=1154752)
|
| 30 |
+
(EngineCore pid=1154752) INFO 08-10 21:25:19 [default_loader.py:384] Loading weights took 3.52 seconds
|
| 31 |
+
(EngineCore pid=1154752) INFO 08-10 21:25:20 [gpu_model_runner.py:4820] Model loading took 3.89 GiB memory and 4.085796 seconds
|
| 32 |
+
(EngineCore pid=1154752) WARNING 08-10 21:25:21 [fused_moe.py:1090] Using default MoE config. Performance might be sub-optimal! Config file not found at /home/henry/.local/lib/python3.10/site-packages/vllm/model_executor/layers/fused_moe/configs/E=16,N=1024,device_name=NVIDIA_GeForce_RTX_3090.json
|
| 33 |
+
(EngineCore pid=1154752) INFO 08-10 21:25:23 [gpu_worker.py:436] Available KV cache memory: 15.73 GiB
|
| 34 |
+
(EngineCore pid=1154752) INFO 08-10 21:25:23 [kv_cache_utils.py:1319] GPU KV cache size: 128,880 tokens
|
| 35 |
+
(EngineCore pid=1154752) INFO 08-10 21:25:23 [kv_cache_utils.py:1324] Maximum concurrency for 2,048 tokens per request: 62.93x
|
| 36 |
+
(EngineCore pid=1154752) INFO 08-10 21:25:23 [core.py:283] init engine (profile, create kv cache, warmup model) took 2.70 seconds
|
| 37 |
+
(EngineCore pid=1154752) INFO 08-10 21:25:23 [vllm.py:790] Asynchronous scheduling is enabled.
|
| 38 |
+
(EngineCore pid=1154752) WARNING 08-10 21:25:23 [vllm.py:848] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
|
| 39 |
+
(EngineCore pid=1154752) WARNING 08-10 21:25:23 [vllm.py:859] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
|
| 40 |
+
(EngineCore pid=1154752) INFO 08-10 21:25:23 [vllm.py:1025] Cudagraph is disabled under eager mode
|
| 41 |
+
(EngineCore pid=1154752) INFO 08-10 21:25:23 [compilation.py:290] Enabled custom fusions: norm_quant, act_quant
|
| 42 |
+
(APIServer pid=1153469) INFO 08-10 21:25:23 [api_server.py:590] Supported tasks: ['generate']
|
| 43 |
+
(APIServer pid=1153469) WARNING 08-10 21:25:23 [__init__.py:14] SECURITY WARNING: Development endpoints are enabled! This should NOT be used in production!
|
| 44 |
+
(APIServer pid=1153469) INFO 08-10 21:25:23 [hf.py:314] Detected the chat template content format to be 'string'. You can set `--chat-template-content-format` to override this.
|
| 45 |
+
(APIServer pid=1153469) INFO 08-10 21:25:24 [api_server.py:594] Starting vLLM server on http://127.0.0.1:8611
|
| 46 |
+
(APIServer pid=1153469) INFO 08-10 21:25:24 [launcher.py:37] Available routes are:
|
| 47 |
+
(APIServer pid=1153469) INFO 08-10 21:25:24 [launcher.py:46] Route: /openapi.json, Methods: HEAD, GET
|
| 48 |
+
(APIServer pid=1153469) INFO 08-10 21:25:24 [launcher.py:46] Route: /docs, Methods: HEAD, GET
|
| 49 |
+
(APIServer pid=1153469) INFO 08-10 21:25:24 [launcher.py:46] Route: /docs/oauth2-redirect, Methods: HEAD, GET
|
| 50 |
+
(APIServer pid=1153469) INFO 08-10 21:25:24 [launcher.py:46] Route: /redoc, Methods: HEAD, GET
|
| 51 |
+
(APIServer pid=1153469) INFO 08-10 21:25:24 [launcher.py:46] Route: /sleep, Methods: POST
|
| 52 |
+
(APIServer pid=1153469) INFO 08-10 21:25:24 [launcher.py:46] Route: /wake_up, Methods: POST
|
| 53 |
+
(APIServer pid=1153469) INFO 08-10 21:25:24 [launcher.py:46] Route: /is_sleeping, Methods: GET
|
| 54 |
+
(APIServer pid=1153469) INFO 08-10 21:25:24 [launcher.py:46] Route: /collective_rpc, Methods: POST
|
| 55 |
+
(APIServer pid=1153469) INFO 08-10 21:25:24 [launcher.py:46] Route: /reset_prefix_cache, Methods: POST
|
| 56 |
+
(APIServer pid=1153469) INFO 08-10 21:25:24 [launcher.py:46] Route: /reset_mm_cache, Methods: POST
|
| 57 |
+
(APIServer pid=1153469) INFO 08-10 21:25:24 [launcher.py:46] Route: /reset_encoder_cache, Methods: POST
|
| 58 |
+
(APIServer pid=1153469) INFO 08-10 21:25:24 [launcher.py:46] Route: /tokenize, Methods: POST
|
| 59 |
+
(APIServer pid=1153469) INFO 08-10 21:25:24 [launcher.py:46] Route: /detokenize, Methods: POST
|
| 60 |
+
(APIServer pid=1153469) INFO 08-10 21:25:24 [launcher.py:46] Route: /load, Methods: GET
|
| 61 |
+
(APIServer pid=1153469) INFO 08-10 21:25:24 [launcher.py:46] Route: /version, Methods: GET
|
| 62 |
+
(APIServer pid=1153469) INFO 08-10 21:25:24 [launcher.py:46] Route: /health, Methods: GET
|
| 63 |
+
(APIServer pid=1153469) INFO 08-10 21:25:24 [launcher.py:46] Route: /metrics, Methods: GET
|
| 64 |
+
(APIServer pid=1153469) INFO 08-10 21:25:24 [launcher.py:46] Route: /server_info, Methods: GET
|
| 65 |
+
(APIServer pid=1153469) INFO 08-10 21:25:24 [launcher.py:46] Route: /v1/models, Methods: GET
|
| 66 |
+
(APIServer pid=1153469) INFO 08-10 21:25:24 [launcher.py:46] Route: /ping, Methods: GET
|
| 67 |
+
(APIServer pid=1153469) INFO 08-10 21:25:24 [launcher.py:46] Route: /ping, Methods: POST
|
| 68 |
+
(APIServer pid=1153469) INFO 08-10 21:25:24 [launcher.py:46] Route: /invocations, Methods: POST
|
| 69 |
+
(APIServer pid=1153469) INFO 08-10 21:25:24 [launcher.py:46] Route: /v1/chat/completions, Methods: POST
|
| 70 |
+
(APIServer pid=1153469) INFO 08-10 21:25:24 [launcher.py:46] Route: /v1/chat/completions/batch, Methods: POST
|
| 71 |
+
(APIServer pid=1153469) INFO 08-10 21:25:24 [launcher.py:46] Route: /v1/responses, Methods: POST
|
| 72 |
+
(APIServer pid=1153469) INFO 08-10 21:25:24 [launcher.py:46] Route: /v1/responses/{response_id}, Methods: GET
|
| 73 |
+
(APIServer pid=1153469) INFO 08-10 21:25:24 [launcher.py:46] Route: /v1/responses/{response_id}/cancel, Methods: POST
|
| 74 |
+
(APIServer pid=1153469) INFO 08-10 21:25:24 [launcher.py:46] Route: /v1/completions, Methods: POST
|
| 75 |
+
(APIServer pid=1153469) INFO 08-10 21:25:24 [launcher.py:46] Route: /v1/messages, Methods: POST
|
| 76 |
+
(APIServer pid=1153469) INFO 08-10 21:25:24 [launcher.py:46] Route: /v1/messages/count_tokens, Methods: POST
|
| 77 |
+
(APIServer pid=1153469) INFO 08-10 21:25:24 [launcher.py:46] Route: /inference/v1/generate, Methods: POST
|
| 78 |
+
(APIServer pid=1153469) INFO 08-10 21:25:24 [launcher.py:46] Route: /pause, Methods: POST
|
| 79 |
+
(APIServer pid=1153469) INFO 08-10 21:25:24 [launcher.py:46] Route: /resume, Methods: POST
|
| 80 |
+
(APIServer pid=1153469) INFO 08-10 21:25:24 [launcher.py:46] Route: /is_paused, Methods: GET
|
| 81 |
+
(APIServer pid=1153469) INFO 08-10 21:25:24 [launcher.py:46] Route: /init_weight_transfer_engine, Methods: POST
|
| 82 |
+
(APIServer pid=1153469) INFO 08-10 21:25:24 [launcher.py:46] Route: /update_weights, Methods: POST
|
| 83 |
+
(APIServer pid=1153469) INFO 08-10 21:25:24 [launcher.py:46] Route: /get_world_size, Methods: GET
|
| 84 |
+
(APIServer pid=1153469) INFO 08-10 21:25:24 [launcher.py:46] Route: /scale_elastic_ep, Methods: POST
|
| 85 |
+
(APIServer pid=1153469) INFO 08-10 21:25:24 [launcher.py:46] Route: /is_scaling_elastic_ep, Methods: POST
|
| 86 |
+
(APIServer pid=1153469) INFO 08-10 21:25:24 [launcher.py:46] Route: /v1/chat/completions/render, Methods: POST
|
| 87 |
+
(APIServer pid=1153469) INFO 08-10 21:25:24 [launcher.py:46] Route: /v1/completions/render, Methods: POST
|
| 88 |
+
(APIServer pid=1153469) INFO: Started server process [1153469]
|
| 89 |
+
(APIServer pid=1153469) INFO: Waiting for application startup.
|
| 90 |
+
(APIServer pid=1153469) INFO: Application startup complete.
|
| 91 |
+
(APIServer pid=1153469) INFO: 127.0.0.1:56296 - "GET /health HTTP/1.1" 200 OK
|
| 92 |
+
(APIServer pid=1153469) INFO 08-10 21:25:34 [loggers.py:259] Engine 000: Avg prompt throughput: 2117.1 tokens/s, Avg generation throughput: 2702.9 tokens/s, Running: 256 reqs, Waiting: 1049 reqs, GPU KV cache usage: 39.5%, Prefix cache hit rate: 93.1%
|
| 93 |
+
(APIServer pid=1153469) INFO 08-10 21:25:44 [loggers.py:259] Engine 000: Avg prompt throughput: 152.4 tokens/s, Avg generation throughput: 3427.3 tokens/s, Running: 256 reqs, Waiting: 1028 reqs, GPU KV cache usage: 62.9%, Prefix cache hit rate: 93.1%
|
| 94 |
+
(APIServer pid=1153469) INFO 08-10 21:25:54 [loggers.py:259] Engine 000: Avg prompt throughput: 85.5 tokens/s, Avg generation throughput: 3199.1 tokens/s, Running: 256 reqs, Waiting: 1019 reqs, GPU KV cache usage: 86.1%, Prefix cache hit rate: 93.1%
|
| 95 |
+
(APIServer pid=1153469) INFO 08-10 21:26:04 [loggers.py:259] Engine 000: Avg prompt throughput: 45.3 tokens/s, Avg generation throughput: 2986.6 tokens/s, Running: 221 reqs, Waiting: 1046 reqs, GPU KV cache usage: 99.3%, Prefix cache hit rate: 93.1%
|
| 96 |
+
(APIServer pid=1153469) INFO 08-10 21:26:14 [loggers.py:259] Engine 000: Avg prompt throughput: 1864.1 tokens/s, Avg generation throughput: 2962.2 tokens/s, Running: 256 reqs, Waiting: 774 reqs, GPU KV cache usage: 41.8%, Prefix cache hit rate: 93.3%
|
| 97 |
+
(APIServer pid=1153469) INFO 08-10 21:26:24 [loggers.py:259] Engine 000: Avg prompt throughput: 205.4 tokens/s, Avg generation throughput: 3375.5 tokens/s, Running: 256 reqs, Waiting: 747 reqs, GPU KV cache usage: 61.7%, Prefix cache hit rate: 93.3%
|
| 98 |
+
(APIServer pid=1153469) INFO 08-10 21:26:34 [loggers.py:259] Engine 000: Avg prompt throughput: 276.4 tokens/s, Avg generation throughput: 3170.8 tokens/s, Running: 256 reqs, Waiting: 713 reqs, GPU KV cache usage: 77.1%, Prefix cache hit rate: 93.3%
|
| 99 |
+
(APIServer pid=1153469) INFO 08-10 21:26:44 [loggers.py:259] Engine 000: Avg prompt throughput: 122.8 tokens/s, Avg generation throughput: 3044.6 tokens/s, Running: 256 reqs, Waiting: 697 reqs, GPU KV cache usage: 95.3%, Prefix cache hit rate: 93.3%
|
| 100 |
+
(APIServer pid=1153469) INFO 08-10 21:26:54 [loggers.py:259] Engine 000: Avg prompt throughput: 1477.9 tokens/s, Avg generation throughput: 3049.6 tokens/s, Running: 256 reqs, Waiting: 508 reqs, GPU KV cache usage: 47.0%, Prefix cache hit rate: 93.3%
|
| 101 |
+
(APIServer pid=1153469) INFO 08-10 21:27:04 [loggers.py:259] Engine 000: Avg prompt throughput: 372.5 tokens/s, Avg generation throughput: 3322.5 tokens/s, Running: 256 reqs, Waiting: 463 reqs, GPU KV cache usage: 60.0%, Prefix cache hit rate: 93.3%
|
| 102 |
+
(APIServer pid=1153469) INFO 08-10 21:27:14 [loggers.py:259] Engine 000: Avg prompt throughput: 270.9 tokens/s, Avg generation throughput: 3195.4 tokens/s, Running: 256 reqs, Waiting: 427 reqs, GPU KV cache usage: 72.5%, Prefix cache hit rate: 93.4%
|
| 103 |
+
(APIServer pid=1153469) INFO 08-10 21:27:24 [loggers.py:259] Engine 000: Avg prompt throughput: 196.2 tokens/s, Avg generation throughput: 3069.2 tokens/s, Running: 256 reqs, Waiting: 400 reqs, GPU KV cache usage: 87.6%, Prefix cache hit rate: 93.4%
|
| 104 |
+
(APIServer pid=1153469) INFO 08-10 21:27:34 [loggers.py:259] Engine 000: Avg prompt throughput: 1343.3 tokens/s, Avg generation throughput: 3004.6 tokens/s, Running: 256 reqs, Waiting: 240 reqs, GPU KV cache usage: 50.5%, Prefix cache hit rate: 93.4%
|
| 105 |
+
(APIServer pid=1153469) INFO 08-10 21:27:44 [loggers.py:259] Engine 000: Avg prompt throughput: 343.4 tokens/s, Avg generation throughput: 3272.2 tokens/s, Running: 256 reqs, Waiting: 195 reqs, GPU KV cache usage: 60.5%, Prefix cache hit rate: 93.4%
|
| 106 |
+
(APIServer pid=1153469) INFO 08-10 21:27:54 [loggers.py:259] Engine 000: Avg prompt throughput: 372.2 tokens/s, Avg generation throughput: 3193.9 tokens/s, Running: 256 reqs, Waiting: 146 reqs, GPU KV cache usage: 69.4%, Prefix cache hit rate: 93.4%
|
| 107 |
+
(APIServer pid=1153469) INFO 08-10 21:28:04 [loggers.py:259] Engine 000: Avg prompt throughput: 230.9 tokens/s, Avg generation throughput: 3094.1 tokens/s, Running: 256 reqs, Waiting: 119 reqs, GPU KV cache usage: 83.9%, Prefix cache hit rate: 93.4%
|
| 108 |
+
(APIServer pid=1153469) INFO 08-10 21:28:14 [loggers.py:259] Engine 000: Avg prompt throughput: 966.9 tokens/s, Avg generation throughput: 2975.0 tokens/s, Running: 228 reqs, Waiting: 0 reqs, GPU KV cache usage: 48.7%, Prefix cache hit rate: 93.4%
|
| 109 |
+
(APIServer pid=1153469) INFO 08-10 21:28:24 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 3140.1 tokens/s, Running: 172 reqs, Waiting: 0 reqs, GPU KV cache usage: 49.8%, Prefix cache hit rate: 93.4%
|
| 110 |
+
(APIServer pid=1153469) INFO 08-10 21:28:34 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 2828.4 tokens/s, Running: 99 reqs, Waiting: 0 reqs, GPU KV cache usage: 39.5%, Prefix cache hit rate: 93.4%
|
| 111 |
+
(APIServer pid=1153469) INFO: 127.0.0.1:56304 - "POST /v1/completions HTTP/1.1" 200 OK
|
| 112 |
+
(EngineCore pid=1154752) INFO 08-10 21:28:39 [core.py:1210] Shutdown initiated (timeout=0)
|
| 113 |
+
(EngineCore pid=1154752) INFO 08-10 21:28:39 [core.py:1233] Shutdown complete
|
| 114 |
+
(APIServer pid=1153469) INFO: Shutting down
|
| 115 |
+
(APIServer pid=1153469) INFO: Waiting for application shutdown.
|
| 116 |
+
(APIServer pid=1153469) INFO: Application shutdown complete.
|
| 117 |
+
(APIServer pid=1153469) INFO: Finished server process [1153469]
|
evals/grid_math_unhealed/reap_keep50_unhealed.log
ADDED
|
@@ -0,0 +1,8 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"correct": 195,
|
| 3 |
+
"accuracy": 0.14783927217589082,
|
| 4 |
+
"finished": 1054,
|
| 5 |
+
"finish_rate": 0.799090219863533,
|
| 6 |
+
"mean_completion_tokens": 268.20621683093253
|
| 7 |
+
}
|
| 8 |
+
saved item-level results -> outputs/evals/grid_math_unhealed/reap_keep50_unhealed_chat.json
|
evals/grid_math_unhealed/reap_keep50_unhealed_chat.json
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
evals/grid_math_unhealed/reap_keep50_unhealed_chat.json.server.log
ADDED
|
@@ -0,0 +1,110 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
Skipping import of cpp extensions due to incompatible torch version 2.10.0+cu128 for torchao version 0.15.0 Please see https://github.com/pytorch/ao/issues/2919 for more info
|
| 2 |
+
WARNING 08-10 21:22:07 [registry.py:915] Model architecture OlmoeForCausalLM is already registered, and will be overwritten by the new model class glean_vllm.pruned_olmoe:PrunedOlmoeForCausalLM.
|
| 3 |
+
(APIServer pid=1147688) INFO 08-10 21:22:07 [utils.py:299]
|
| 4 |
+
(APIServer pid=1147688) INFO 08-10 21:22:07 [utils.py:299] β β ββ ββ
|
| 5 |
+
(APIServer pid=1147688) INFO 08-10 21:22:07 [utils.py:299] ββ ββ β β β βββ β version 0.19.0
|
| 6 |
+
(APIServer pid=1147688) INFO 08-10 21:22:07 [utils.py:299] ββββ β β β β model /media/henry/MoreFiles/winnow_release/math_arms/reap-math-keep50
|
| 7 |
+
(APIServer pid=1147688) INFO 08-10 21:22:07 [utils.py:299] ββ βββββ βββββ β β
|
| 8 |
+
(APIServer pid=1147688) INFO 08-10 21:22:07 [utils.py:299]
|
| 9 |
+
(APIServer pid=1147688) INFO 08-10 21:22:07 [utils.py:233] non-default args: {'model_tag': '/media/henry/MoreFiles/winnow_release/math_arms/reap-math-keep50', 'host': '127.0.0.1', 'port': 8611, 'model': '/media/henry/MoreFiles/winnow_release/math_arms/reap-math-keep50', 'max_model_len': 2048, 'enforce_eager': True, 'served_model_name': ['student'], 'gpu_memory_utilization': 0.85}
|
| 10 |
+
(APIServer pid=1147688) INFO 08-10 21:22:15 [model.py:549] Resolved architecture: OlmoeForCausalLM
|
| 11 |
+
(APIServer pid=1147688) INFO 08-10 21:22:15 [model.py:1678] Using max model len 2048
|
| 12 |
+
(APIServer pid=1147688) INFO 08-10 21:22:15 [vllm.py:790] Asynchronous scheduling is enabled.
|
| 13 |
+
(APIServer pid=1147688) WARNING 08-10 21:22:15 [vllm.py:848] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
|
| 14 |
+
(APIServer pid=1147688) WARNING 08-10 21:22:15 [vllm.py:859] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
|
| 15 |
+
(APIServer pid=1147688) INFO 08-10 21:22:15 [vllm.py:1025] Cudagraph is disabled under eager mode
|
| 16 |
+
(APIServer pid=1147688) INFO 08-10 21:22:15 [compilation.py:290] Enabled custom fusions: norm_quant, act_quant
|
| 17 |
+
Skipping import of cpp extensions due to incompatible torch version 2.10.0+cu128 for torchao version 0.15.0 Please see https://github.com/pytorch/ao/issues/2919 for more info
|
| 18 |
+
(EngineCore pid=1148972) WARNING 08-10 21:22:24 [registry.py:915] Model architecture OlmoeForCausalLM is already registered, and will be overwritten by the new model class glean_vllm.pruned_olmoe:PrunedOlmoeForCausalLM.
|
| 19 |
+
(EngineCore pid=1148972) INFO 08-10 21:22:24 [core.py:105] Initializing a V1 LLM engine (v0.19.0) with config: model='/media/henry/MoreFiles/winnow_release/math_arms/reap-math-keep50', speculative_config=None, tokenizer='/media/henry/MoreFiles/winnow_release/math_arms/reap-math-keep50', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=False, dtype=torch.bfloat16, max_seq_len=2048, download_dir=None, load_format=auto, tensor_parallel_size=1, pipeline_parallel_size=1, data_parallel_size=1, decode_context_parallel_size=1, dcp_comm_backend=ag_rs, disable_custom_all_reduce=False, quantization=None, enforce_eager=True, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False), seed=0, served_model_name=student, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'mode': <CompilationMode.NONE: 0>, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'splitting_ops': [], 'compile_mm_encoder': False, 'cudagraph_mm_encoder': False, 'encoder_cudagraph_token_budgets': [], 'encoder_cudagraph_max_images_per_batch': 0, 'compile_sizes': [], 'compile_ranges_endpoints': [2048], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'size_asserts': False, 'alignment_asserts': False, 'scalar_asserts': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': <CUDAGraphMode.NONE: 0>, 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': True, 'fuse_act_quant': True, 'fuse_attn_quant': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False}, 'max_cudagraph_capture_size': 0, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': True, 'static_all_moe_layers': []}
|
| 20 |
+
(EngineCore pid=1148972) INFO 08-10 21:22:24 [parallel_state.py:1400] world_size=1 rank=0 local_rank=0 distributed_init_method=tcp://192.168.0.27:39471 backend=nccl
|
| 21 |
+
(EngineCore pid=1148972) INFO 08-10 21:22:24 [parallel_state.py:1716] rank 0 in world size 1 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank 0, EPLB rank N/A
|
| 22 |
+
(EngineCore pid=1148972) INFO 08-10 21:22:25 [gpu_model_runner.py:4735] Starting to load model /media/henry/MoreFiles/winnow_release/math_arms/reap-math-keep50...
|
| 23 |
+
(EngineCore pid=1148972) INFO 08-10 21:22:25 [cuda.py:334] Using FLASH_ATTN attention backend out of potential backends: ['FLASH_ATTN', 'FLASHINFER', 'TRITON_ATTN', 'FLEX_ATTENTION'].
|
| 24 |
+
(EngineCore pid=1148972) INFO 08-10 21:22:25 [flash_attn.py:596] Using FlashAttention version 2
|
| 25 |
+
(EngineCore pid=1148972) INFO 08-10 21:22:25 [unquantized.py:186] Using TRITON backend for Unquantized MoE
|
| 26 |
+
(EngineCore pid=1148972)
|
| 27 |
+
(EngineCore pid=1148972)
|
| 28 |
+
(EngineCore pid=1148972)
|
| 29 |
+
(EngineCore pid=1148972)
|
| 30 |
+
(EngineCore pid=1148972)
|
| 31 |
+
(EngineCore pid=1148972) INFO 08-10 21:22:28 [default_loader.py:384] Loading weights took 2.36 seconds
|
| 32 |
+
(EngineCore pid=1148972) INFO 08-10 21:22:28 [gpu_model_runner.py:4820] Model loading took 6.89 GiB memory and 2.924399 seconds
|
| 33 |
+
(EngineCore pid=1148972) WARNING 08-10 21:22:29 [fused_moe.py:1090] Using default MoE config. Performance might be sub-optimal! Config file not found at /home/henry/.local/lib/python3.10/site-packages/vllm/model_executor/layers/fused_moe/configs/E=32,N=1024,device_name=NVIDIA_GeForce_RTX_3090.json
|
| 34 |
+
(EngineCore pid=1148972) INFO 08-10 21:22:30 [gpu_worker.py:436] Available KV cache memory: 12.73 GiB
|
| 35 |
+
(EngineCore pid=1148972) INFO 08-10 21:22:30 [kv_cache_utils.py:1319] GPU KV cache size: 104,304 tokens
|
| 36 |
+
(EngineCore pid=1148972) INFO 08-10 21:22:30 [kv_cache_utils.py:1324] Maximum concurrency for 2,048 tokens per request: 50.93x
|
| 37 |
+
(EngineCore pid=1148972) INFO 08-10 21:22:30 [core.py:283] init engine (profile, create kv cache, warmup model) took 1.68 seconds
|
| 38 |
+
(EngineCore pid=1148972) INFO 08-10 21:22:31 [vllm.py:790] Asynchronous scheduling is enabled.
|
| 39 |
+
(EngineCore pid=1148972) WARNING 08-10 21:22:31 [vllm.py:848] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
|
| 40 |
+
(EngineCore pid=1148972) WARNING 08-10 21:22:31 [vllm.py:859] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
|
| 41 |
+
(EngineCore pid=1148972) INFO 08-10 21:22:31 [vllm.py:1025] Cudagraph is disabled under eager mode
|
| 42 |
+
(EngineCore pid=1148972) INFO 08-10 21:22:31 [compilation.py:290] Enabled custom fusions: norm_quant, act_quant
|
| 43 |
+
(APIServer pid=1147688) INFO 08-10 21:22:31 [api_server.py:590] Supported tasks: ['generate']
|
| 44 |
+
(APIServer pid=1147688) WARNING 08-10 21:22:31 [__init__.py:14] SECURITY WARNING: Development endpoints are enabled! This should NOT be used in production!
|
| 45 |
+
(APIServer pid=1147688) INFO 08-10 21:22:31 [hf.py:314] Detected the chat template content format to be 'string'. You can set `--chat-template-content-format` to override this.
|
| 46 |
+
(APIServer pid=1147688) INFO 08-10 21:22:31 [api_server.py:594] Starting vLLM server on http://127.0.0.1:8611
|
| 47 |
+
(APIServer pid=1147688) INFO 08-10 21:22:31 [launcher.py:37] Available routes are:
|
| 48 |
+
(APIServer pid=1147688) INFO 08-10 21:22:31 [launcher.py:46] Route: /openapi.json, Methods: GET, HEAD
|
| 49 |
+
(APIServer pid=1147688) INFO 08-10 21:22:31 [launcher.py:46] Route: /docs, Methods: GET, HEAD
|
| 50 |
+
(APIServer pid=1147688) INFO 08-10 21:22:31 [launcher.py:46] Route: /docs/oauth2-redirect, Methods: GET, HEAD
|
| 51 |
+
(APIServer pid=1147688) INFO 08-10 21:22:31 [launcher.py:46] Route: /redoc, Methods: GET, HEAD
|
| 52 |
+
(APIServer pid=1147688) INFO 08-10 21:22:31 [launcher.py:46] Route: /sleep, Methods: POST
|
| 53 |
+
(APIServer pid=1147688) INFO 08-10 21:22:31 [launcher.py:46] Route: /wake_up, Methods: POST
|
| 54 |
+
(APIServer pid=1147688) INFO 08-10 21:22:31 [launcher.py:46] Route: /is_sleeping, Methods: GET
|
| 55 |
+
(APIServer pid=1147688) INFO 08-10 21:22:31 [launcher.py:46] Route: /collective_rpc, Methods: POST
|
| 56 |
+
(APIServer pid=1147688) INFO 08-10 21:22:31 [launcher.py:46] Route: /reset_prefix_cache, Methods: POST
|
| 57 |
+
(APIServer pid=1147688) INFO 08-10 21:22:31 [launcher.py:46] Route: /reset_mm_cache, Methods: POST
|
| 58 |
+
(APIServer pid=1147688) INFO 08-10 21:22:31 [launcher.py:46] Route: /reset_encoder_cache, Methods: POST
|
| 59 |
+
(APIServer pid=1147688) INFO 08-10 21:22:31 [launcher.py:46] Route: /tokenize, Methods: POST
|
| 60 |
+
(APIServer pid=1147688) INFO 08-10 21:22:31 [launcher.py:46] Route: /detokenize, Methods: POST
|
| 61 |
+
(APIServer pid=1147688) INFO 08-10 21:22:31 [launcher.py:46] Route: /load, Methods: GET
|
| 62 |
+
(APIServer pid=1147688) INFO 08-10 21:22:31 [launcher.py:46] Route: /version, Methods: GET
|
| 63 |
+
(APIServer pid=1147688) INFO 08-10 21:22:31 [launcher.py:46] Route: /health, Methods: GET
|
| 64 |
+
(APIServer pid=1147688) INFO 08-10 21:22:31 [launcher.py:46] Route: /metrics, Methods: GET
|
| 65 |
+
(APIServer pid=1147688) INFO 08-10 21:22:31 [launcher.py:46] Route: /server_info, Methods: GET
|
| 66 |
+
(APIServer pid=1147688) INFO 08-10 21:22:31 [launcher.py:46] Route: /v1/models, Methods: GET
|
| 67 |
+
(APIServer pid=1147688) INFO 08-10 21:22:31 [launcher.py:46] Route: /ping, Methods: GET
|
| 68 |
+
(APIServer pid=1147688) INFO 08-10 21:22:31 [launcher.py:46] Route: /ping, Methods: POST
|
| 69 |
+
(APIServer pid=1147688) INFO 08-10 21:22:31 [launcher.py:46] Route: /invocations, Methods: POST
|
| 70 |
+
(APIServer pid=1147688) INFO 08-10 21:22:31 [launcher.py:46] Route: /v1/chat/completions, Methods: POST
|
| 71 |
+
(APIServer pid=1147688) INFO 08-10 21:22:31 [launcher.py:46] Route: /v1/chat/completions/batch, Methods: POST
|
| 72 |
+
(APIServer pid=1147688) INFO 08-10 21:22:31 [launcher.py:46] Route: /v1/responses, Methods: POST
|
| 73 |
+
(APIServer pid=1147688) INFO 08-10 21:22:31 [launcher.py:46] Route: /v1/responses/{response_id}, Methods: GET
|
| 74 |
+
(APIServer pid=1147688) INFO 08-10 21:22:31 [launcher.py:46] Route: /v1/responses/{response_id}/cancel, Methods: POST
|
| 75 |
+
(APIServer pid=1147688) INFO 08-10 21:22:31 [launcher.py:46] Route: /v1/completions, Methods: POST
|
| 76 |
+
(APIServer pid=1147688) INFO 08-10 21:22:31 [launcher.py:46] Route: /v1/messages, Methods: POST
|
| 77 |
+
(APIServer pid=1147688) INFO 08-10 21:22:31 [launcher.py:46] Route: /v1/messages/count_tokens, Methods: POST
|
| 78 |
+
(APIServer pid=1147688) INFO 08-10 21:22:31 [launcher.py:46] Route: /inference/v1/generate, Methods: POST
|
| 79 |
+
(APIServer pid=1147688) INFO 08-10 21:22:31 [launcher.py:46] Route: /pause, Methods: POST
|
| 80 |
+
(APIServer pid=1147688) INFO 08-10 21:22:31 [launcher.py:46] Route: /resume, Methods: POST
|
| 81 |
+
(APIServer pid=1147688) INFO 08-10 21:22:31 [launcher.py:46] Route: /is_paused, Methods: GET
|
| 82 |
+
(APIServer pid=1147688) INFO 08-10 21:22:31 [launcher.py:46] Route: /init_weight_transfer_engine, Methods: POST
|
| 83 |
+
(APIServer pid=1147688) INFO 08-10 21:22:31 [launcher.py:46] Route: /update_weights, Methods: POST
|
| 84 |
+
(APIServer pid=1147688) INFO 08-10 21:22:31 [launcher.py:46] Route: /get_world_size, Methods: GET
|
| 85 |
+
(APIServer pid=1147688) INFO 08-10 21:22:31 [launcher.py:46] Route: /scale_elastic_ep, Methods: POST
|
| 86 |
+
(APIServer pid=1147688) INFO 08-10 21:22:31 [launcher.py:46] Route: /is_scaling_elastic_ep, Methods: POST
|
| 87 |
+
(APIServer pid=1147688) INFO 08-10 21:22:31 [launcher.py:46] Route: /v1/chat/completions/render, Methods: POST
|
| 88 |
+
(APIServer pid=1147688) INFO 08-10 21:22:31 [launcher.py:46] Route: /v1/completions/render, Methods: POST
|
| 89 |
+
(APIServer pid=1147688) INFO: Started server process [1147688]
|
| 90 |
+
(APIServer pid=1147688) INFO: Waiting for application startup.
|
| 91 |
+
(APIServer pid=1147688) INFO: Application startup complete.
|
| 92 |
+
(APIServer pid=1147688) INFO: 127.0.0.1:46976 - "GET /health HTTP/1.1" 200 OK
|
| 93 |
+
(APIServer pid=1147688) INFO 08-10 21:22:41 [loggers.py:259] Engine 000: Avg prompt throughput: 2242.6 tokens/s, Avg generation throughput: 2459.7 tokens/s, Running: 255 reqs, Waiting: 1030 reqs, GPU KV cache usage: 44.8%, Prefix cache hit rate: 93.1%
|
| 94 |
+
(APIServer pid=1147688) INFO 08-10 21:22:51 [loggers.py:259] Engine 000: Avg prompt throughput: 904.4 tokens/s, Avg generation throughput: 3290.0 tokens/s, Running: 256 reqs, Waiting: 915 reqs, GPU KV cache usage: 60.6%, Prefix cache hit rate: 93.2%
|
| 95 |
+
(APIServer pid=1147688) INFO 08-10 21:23:01 [loggers.py:259] Engine 000: Avg prompt throughput: 884.6 tokens/s, Avg generation throughput: 3162.2 tokens/s, Running: 255 reqs, Waiting: 804 reqs, GPU KV cache usage: 68.7%, Prefix cache hit rate: 93.3%
|
| 96 |
+
(APIServer pid=1147688) INFO 08-10 21:23:11 [loggers.py:259] Engine 000: Avg prompt throughput: 723.3 tokens/s, Avg generation throughput: 3138.0 tokens/s, Running: 256 reqs, Waiting: 710 reqs, GPU KV cache usage: 78.8%, Prefix cache hit rate: 93.3%
|
| 97 |
+
(APIServer pid=1147688) INFO 08-10 21:23:21 [loggers.py:259] Engine 000: Avg prompt throughput: 1120.1 tokens/s, Avg generation throughput: 3108.8 tokens/s, Running: 256 reqs, Waiting: 566 reqs, GPU KV cache usage: 63.1%, Prefix cache hit rate: 93.3%
|
| 98 |
+
(APIServer pid=1147688) INFO 08-10 21:23:31 [loggers.py:259] Engine 000: Avg prompt throughput: 953.6 tokens/s, Avg generation throughput: 3135.7 tokens/s, Running: 256 reqs, Waiting: 447 reqs, GPU KV cache usage: 65.3%, Prefix cache hit rate: 93.3%
|
| 99 |
+
(APIServer pid=1147688) INFO 08-10 21:23:41 [loggers.py:259] Engine 000: Avg prompt throughput: 996.2 tokens/s, Avg generation throughput: 3135.3 tokens/s, Running: 256 reqs, Waiting: 321 reqs, GPU KV cache usage: 65.1%, Prefix cache hit rate: 93.3%
|
| 100 |
+
(APIServer pid=1147688) INFO 08-10 21:23:51 [loggers.py:259] Engine 000: Avg prompt throughput: 916.9 tokens/s, Avg generation throughput: 3111.0 tokens/s, Running: 256 reqs, Waiting: 210 reqs, GPU KV cache usage: 67.8%, Prefix cache hit rate: 93.4%
|
| 101 |
+
(APIServer pid=1147688) INFO 08-10 21:24:01 [loggers.py:259] Engine 000: Avg prompt throughput: 814.5 tokens/s, Avg generation throughput: 3138.3 tokens/s, Running: 255 reqs, Waiting: 107 reqs, GPU KV cache usage: 70.9%, Prefix cache hit rate: 93.4%
|
| 102 |
+
(APIServer pid=1147688) INFO 08-10 21:24:11 [loggers.py:259] Engine 000: Avg prompt throughput: 880.7 tokens/s, Avg generation throughput: 3076.6 tokens/s, Running: 243 reqs, Waiting: 0 reqs, GPU KV cache usage: 68.8%, Prefix cache hit rate: 93.4%
|
| 103 |
+
(APIServer pid=1147688) INFO 08-10 21:24:21 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 2944.0 tokens/s, Running: 120 reqs, Waiting: 0 reqs, GPU KV cache usage: 48.8%, Prefix cache hit rate: 93.4%
|
| 104 |
+
(APIServer pid=1147688) INFO: 127.0.0.1:46980 - "POST /v1/completions HTTP/1.1" 200 OK
|
| 105 |
+
(EngineCore pid=1148972) INFO 08-10 21:24:30 [core.py:1210] Shutdown initiated (timeout=0)
|
| 106 |
+
(EngineCore pid=1148972) INFO 08-10 21:24:30 [core.py:1233] Shutdown complete
|
| 107 |
+
(APIServer pid=1147688) INFO: Shutting down
|
| 108 |
+
(APIServer pid=1147688) INFO: Waiting for application shutdown.
|
| 109 |
+
(APIServer pid=1147688) INFO: Application shutdown complete.
|
| 110 |
+
(APIServer pid=1147688) INFO: Finished server process [1147688]
|
evals/grid_math_unhealed/reap_keep75_unhealed.log
ADDED
|
@@ -0,0 +1,8 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"correct": 833,
|
| 3 |
+
"accuracy": 0.6315390447308568,
|
| 4 |
+
"finished": 1316,
|
| 5 |
+
"finish_rate": 0.9977255496588324,
|
| 6 |
+
"mean_completion_tokens": 108.13646702047005
|
| 7 |
+
}
|
| 8 |
+
saved item-level results -> outputs/evals/grid_math_unhealed/reap_keep75_unhealed_chat.json
|
evals/grid_math_unhealed/reap_keep75_unhealed_chat.json
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
evals/grid_math_unhealed/reap_keep75_unhealed_chat.json.server.log
ADDED
|
@@ -0,0 +1,105 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
Skipping import of cpp extensions due to incompatible torch version 2.10.0+cu128 for torchao version 0.15.0 Please see https://github.com/pytorch/ao/issues/2919 for more info
|
| 2 |
+
WARNING 08-10 21:20:11 [registry.py:915] Model architecture OlmoeForCausalLM is already registered, and will be overwritten by the new model class glean_vllm.pruned_olmoe:PrunedOlmoeForCausalLM.
|
| 3 |
+
(APIServer pid=1142905) INFO 08-10 21:20:11 [utils.py:299]
|
| 4 |
+
(APIServer pid=1142905) INFO 08-10 21:20:11 [utils.py:299] β β ββ ββ
|
| 5 |
+
(APIServer pid=1142905) INFO 08-10 21:20:11 [utils.py:299] ββ ββ β β β βββ β version 0.19.0
|
| 6 |
+
(APIServer pid=1142905) INFO 08-10 21:20:11 [utils.py:299] ββββ β β β β model /media/henry/MoreFiles/winnow_release/math_arms/reap-math-keep75
|
| 7 |
+
(APIServer pid=1142905) INFO 08-10 21:20:11 [utils.py:299] ββ βββββ βββββ β β
|
| 8 |
+
(APIServer pid=1142905) INFO 08-10 21:20:11 [utils.py:299]
|
| 9 |
+
(APIServer pid=1142905) INFO 08-10 21:20:11 [utils.py:233] non-default args: {'model_tag': '/media/henry/MoreFiles/winnow_release/math_arms/reap-math-keep75', 'host': '127.0.0.1', 'port': 8611, 'model': '/media/henry/MoreFiles/winnow_release/math_arms/reap-math-keep75', 'max_model_len': 2048, 'enforce_eager': True, 'served_model_name': ['student'], 'gpu_memory_utilization': 0.85}
|
| 10 |
+
(APIServer pid=1142905) INFO 08-10 21:20:21 [model.py:549] Resolved architecture: OlmoeForCausalLM
|
| 11 |
+
(APIServer pid=1142905) INFO 08-10 21:20:21 [model.py:1678] Using max model len 2048
|
| 12 |
+
(APIServer pid=1142905) INFO 08-10 21:20:22 [vllm.py:790] Asynchronous scheduling is enabled.
|
| 13 |
+
(APIServer pid=1142905) WARNING 08-10 21:20:22 [vllm.py:848] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
|
| 14 |
+
(APIServer pid=1142905) WARNING 08-10 21:20:22 [vllm.py:859] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
|
| 15 |
+
(APIServer pid=1142905) INFO 08-10 21:20:22 [vllm.py:1025] Cudagraph is disabled under eager mode
|
| 16 |
+
(APIServer pid=1142905) INFO 08-10 21:20:22 [compilation.py:290] Enabled custom fusions: norm_quant, act_quant
|
| 17 |
+
Skipping import of cpp extensions due to incompatible torch version 2.10.0+cu128 for torchao version 0.15.0 Please see https://github.com/pytorch/ao/issues/2919 for more info
|
| 18 |
+
(EngineCore pid=1144304) WARNING 08-10 21:20:30 [registry.py:915] Model architecture OlmoeForCausalLM is already registered, and will be overwritten by the new model class glean_vllm.pruned_olmoe:PrunedOlmoeForCausalLM.
|
| 19 |
+
(EngineCore pid=1144304) INFO 08-10 21:20:30 [core.py:105] Initializing a V1 LLM engine (v0.19.0) with config: model='/media/henry/MoreFiles/winnow_release/math_arms/reap-math-keep75', speculative_config=None, tokenizer='/media/henry/MoreFiles/winnow_release/math_arms/reap-math-keep75', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=False, dtype=torch.bfloat16, max_seq_len=2048, download_dir=None, load_format=auto, tensor_parallel_size=1, pipeline_parallel_size=1, data_parallel_size=1, decode_context_parallel_size=1, dcp_comm_backend=ag_rs, disable_custom_all_reduce=False, quantization=None, enforce_eager=True, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False), seed=0, served_model_name=student, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'mode': <CompilationMode.NONE: 0>, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'splitting_ops': [], 'compile_mm_encoder': False, 'cudagraph_mm_encoder': False, 'encoder_cudagraph_token_budgets': [], 'encoder_cudagraph_max_images_per_batch': 0, 'compile_sizes': [], 'compile_ranges_endpoints': [2048], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'size_asserts': False, 'alignment_asserts': False, 'scalar_asserts': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': <CUDAGraphMode.NONE: 0>, 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': True, 'fuse_act_quant': True, 'fuse_attn_quant': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False}, 'max_cudagraph_capture_size': 0, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': True, 'static_all_moe_layers': []}
|
| 20 |
+
(EngineCore pid=1144304) INFO 08-10 21:20:31 [parallel_state.py:1400] world_size=1 rank=0 local_rank=0 distributed_init_method=tcp://192.168.0.27:40093 backend=nccl
|
| 21 |
+
(EngineCore pid=1144304) INFO 08-10 21:20:31 [parallel_state.py:1716] rank 0 in world size 1 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank 0, EPLB rank N/A
|
| 22 |
+
(EngineCore pid=1144304) INFO 08-10 21:20:31 [gpu_model_runner.py:4735] Starting to load model /media/henry/MoreFiles/winnow_release/math_arms/reap-math-keep75...
|
| 23 |
+
(EngineCore pid=1144304) INFO 08-10 21:20:33 [cuda.py:334] Using FLASH_ATTN attention backend out of potential backends: ['FLASH_ATTN', 'FLASHINFER', 'TRITON_ATTN', 'FLEX_ATTENTION'].
|
| 24 |
+
(EngineCore pid=1144304) INFO 08-10 21:20:33 [flash_attn.py:596] Using FlashAttention version 2
|
| 25 |
+
(EngineCore pid=1144304) INFO 08-10 21:20:33 [unquantized.py:186] Using TRITON backend for Unquantized MoE
|
| 26 |
+
(EngineCore pid=1144304)
|
| 27 |
+
(EngineCore pid=1144304)
|
| 28 |
+
(EngineCore pid=1144304)
|
| 29 |
+
(EngineCore pid=1144304)
|
| 30 |
+
(EngineCore pid=1144304)
|
| 31 |
+
(EngineCore pid=1144304)
|
| 32 |
+
(EngineCore pid=1144304) INFO 08-10 21:20:35 [default_loader.py:384] Loading weights took 2.25 seconds
|
| 33 |
+
(EngineCore pid=1144304) INFO 08-10 21:20:36 [gpu_model_runner.py:4820] Model loading took 9.89 GiB memory and 3.345540 seconds
|
| 34 |
+
(EngineCore pid=1144304) WARNING 08-10 21:20:36 [fused_moe.py:1090] Using default MoE config. Performance might be sub-optimal! Config file not found at /home/henry/.local/lib/python3.10/site-packages/vllm/model_executor/layers/fused_moe/configs/E=48,N=1024,device_name=NVIDIA_GeForce_RTX_3090.json
|
| 35 |
+
(EngineCore pid=1144304) INFO 08-10 21:20:37 [gpu_worker.py:436] Available KV cache memory: 9.73 GiB
|
| 36 |
+
(EngineCore pid=1144304) INFO 08-10 21:20:37 [kv_cache_utils.py:1319] GPU KV cache size: 79,712 tokens
|
| 37 |
+
(EngineCore pid=1144304) INFO 08-10 21:20:37 [kv_cache_utils.py:1324] Maximum concurrency for 2,048 tokens per request: 38.92x
|
| 38 |
+
(EngineCore pid=1144304) INFO 08-10 21:20:37 [core.py:283] init engine (profile, create kv cache, warmup model) took 1.75 seconds
|
| 39 |
+
(EngineCore pid=1144304) INFO 08-10 21:20:38 [vllm.py:790] Asynchronous scheduling is enabled.
|
| 40 |
+
(EngineCore pid=1144304) WARNING 08-10 21:20:38 [vllm.py:848] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
|
| 41 |
+
(EngineCore pid=1144304) WARNING 08-10 21:20:38 [vllm.py:859] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
|
| 42 |
+
(EngineCore pid=1144304) INFO 08-10 21:20:38 [vllm.py:1025] Cudagraph is disabled under eager mode
|
| 43 |
+
(EngineCore pid=1144304) INFO 08-10 21:20:38 [compilation.py:290] Enabled custom fusions: norm_quant, act_quant
|
| 44 |
+
(APIServer pid=1142905) INFO 08-10 21:20:38 [api_server.py:590] Supported tasks: ['generate']
|
| 45 |
+
(APIServer pid=1142905) WARNING 08-10 21:20:38 [__init__.py:14] SECURITY WARNING: Development endpoints are enabled! This should NOT be used in production!
|
| 46 |
+
(APIServer pid=1142905) INFO 08-10 21:20:38 [hf.py:314] Detected the chat template content format to be 'string'. You can set `--chat-template-content-format` to override this.
|
| 47 |
+
(APIServer pid=1142905) INFO 08-10 21:20:38 [api_server.py:594] Starting vLLM server on http://127.0.0.1:8611
|
| 48 |
+
(APIServer pid=1142905) INFO 08-10 21:20:38 [launcher.py:37] Available routes are:
|
| 49 |
+
(APIServer pid=1142905) INFO 08-10 21:20:38 [launcher.py:46] Route: /openapi.json, Methods: HEAD, GET
|
| 50 |
+
(APIServer pid=1142905) INFO 08-10 21:20:38 [launcher.py:46] Route: /docs, Methods: HEAD, GET
|
| 51 |
+
(APIServer pid=1142905) INFO 08-10 21:20:38 [launcher.py:46] Route: /docs/oauth2-redirect, Methods: HEAD, GET
|
| 52 |
+
(APIServer pid=1142905) INFO 08-10 21:20:38 [launcher.py:46] Route: /redoc, Methods: HEAD, GET
|
| 53 |
+
(APIServer pid=1142905) INFO 08-10 21:20:38 [launcher.py:46] Route: /sleep, Methods: POST
|
| 54 |
+
(APIServer pid=1142905) INFO 08-10 21:20:38 [launcher.py:46] Route: /wake_up, Methods: POST
|
| 55 |
+
(APIServer pid=1142905) INFO 08-10 21:20:38 [launcher.py:46] Route: /is_sleeping, Methods: GET
|
| 56 |
+
(APIServer pid=1142905) INFO 08-10 21:20:38 [launcher.py:46] Route: /collective_rpc, Methods: POST
|
| 57 |
+
(APIServer pid=1142905) INFO 08-10 21:20:38 [launcher.py:46] Route: /reset_prefix_cache, Methods: POST
|
| 58 |
+
(APIServer pid=1142905) INFO 08-10 21:20:38 [launcher.py:46] Route: /reset_mm_cache, Methods: POST
|
| 59 |
+
(APIServer pid=1142905) INFO 08-10 21:20:38 [launcher.py:46] Route: /reset_encoder_cache, Methods: POST
|
| 60 |
+
(APIServer pid=1142905) INFO 08-10 21:20:38 [launcher.py:46] Route: /tokenize, Methods: POST
|
| 61 |
+
(APIServer pid=1142905) INFO 08-10 21:20:38 [launcher.py:46] Route: /detokenize, Methods: POST
|
| 62 |
+
(APIServer pid=1142905) INFO 08-10 21:20:38 [launcher.py:46] Route: /load, Methods: GET
|
| 63 |
+
(APIServer pid=1142905) INFO 08-10 21:20:38 [launcher.py:46] Route: /version, Methods: GET
|
| 64 |
+
(APIServer pid=1142905) INFO 08-10 21:20:38 [launcher.py:46] Route: /health, Methods: GET
|
| 65 |
+
(APIServer pid=1142905) INFO 08-10 21:20:38 [launcher.py:46] Route: /metrics, Methods: GET
|
| 66 |
+
(APIServer pid=1142905) INFO 08-10 21:20:38 [launcher.py:46] Route: /server_info, Methods: GET
|
| 67 |
+
(APIServer pid=1142905) INFO 08-10 21:20:38 [launcher.py:46] Route: /v1/models, Methods: GET
|
| 68 |
+
(APIServer pid=1142905) INFO 08-10 21:20:38 [launcher.py:46] Route: /ping, Methods: GET
|
| 69 |
+
(APIServer pid=1142905) INFO 08-10 21:20:38 [launcher.py:46] Route: /ping, Methods: POST
|
| 70 |
+
(APIServer pid=1142905) INFO 08-10 21:20:38 [launcher.py:46] Route: /invocations, Methods: POST
|
| 71 |
+
(APIServer pid=1142905) INFO 08-10 21:20:38 [launcher.py:46] Route: /v1/chat/completions, Methods: POST
|
| 72 |
+
(APIServer pid=1142905) INFO 08-10 21:20:38 [launcher.py:46] Route: /v1/chat/completions/batch, Methods: POST
|
| 73 |
+
(APIServer pid=1142905) INFO 08-10 21:20:38 [launcher.py:46] Route: /v1/responses, Methods: POST
|
| 74 |
+
(APIServer pid=1142905) INFO 08-10 21:20:38 [launcher.py:46] Route: /v1/responses/{response_id}, Methods: GET
|
| 75 |
+
(APIServer pid=1142905) INFO 08-10 21:20:38 [launcher.py:46] Route: /v1/responses/{response_id}/cancel, Methods: POST
|
| 76 |
+
(APIServer pid=1142905) INFO 08-10 21:20:38 [launcher.py:46] Route: /v1/completions, Methods: POST
|
| 77 |
+
(APIServer pid=1142905) INFO 08-10 21:20:38 [launcher.py:46] Route: /v1/messages, Methods: POST
|
| 78 |
+
(APIServer pid=1142905) INFO 08-10 21:20:38 [launcher.py:46] Route: /v1/messages/count_tokens, Methods: POST
|
| 79 |
+
(APIServer pid=1142905) INFO 08-10 21:20:38 [launcher.py:46] Route: /inference/v1/generate, Methods: POST
|
| 80 |
+
(APIServer pid=1142905) INFO 08-10 21:20:38 [launcher.py:46] Route: /pause, Methods: POST
|
| 81 |
+
(APIServer pid=1142905) INFO 08-10 21:20:38 [launcher.py:46] Route: /resume, Methods: POST
|
| 82 |
+
(APIServer pid=1142905) INFO 08-10 21:20:38 [launcher.py:46] Route: /is_paused, Methods: GET
|
| 83 |
+
(APIServer pid=1142905) INFO 08-10 21:20:38 [launcher.py:46] Route: /init_weight_transfer_engine, Methods: POST
|
| 84 |
+
(APIServer pid=1142905) INFO 08-10 21:20:38 [launcher.py:46] Route: /update_weights, Methods: POST
|
| 85 |
+
(APIServer pid=1142905) INFO 08-10 21:20:38 [launcher.py:46] Route: /get_world_size, Methods: GET
|
| 86 |
+
(APIServer pid=1142905) INFO 08-10 21:20:38 [launcher.py:46] Route: /scale_elastic_ep, Methods: POST
|
| 87 |
+
(APIServer pid=1142905) INFO 08-10 21:20:38 [launcher.py:46] Route: /is_scaling_elastic_ep, Methods: POST
|
| 88 |
+
(APIServer pid=1142905) INFO 08-10 21:20:38 [launcher.py:46] Route: /v1/chat/completions/render, Methods: POST
|
| 89 |
+
(APIServer pid=1142905) INFO 08-10 21:20:38 [launcher.py:46] Route: /v1/completions/render, Methods: POST
|
| 90 |
+
(APIServer pid=1142905) INFO: Started server process [1142905]
|
| 91 |
+
(APIServer pid=1142905) INFO: Waiting for application startup.
|
| 92 |
+
(APIServer pid=1142905) INFO: Application startup complete.
|
| 93 |
+
(APIServer pid=1142905) INFO: 127.0.0.1:54554 - "GET /health HTTP/1.1" 200 OK
|
| 94 |
+
(APIServer pid=1142905) INFO 08-10 21:20:49 [loggers.py:259] Engine 000: Avg prompt throughput: 2776.8 tokens/s, Avg generation throughput: 1984.7 tokens/s, Running: 252 reqs, Waiting: 955 reqs, GPU KV cache usage: 47.5%, Prefix cache hit rate: 93.2%
|
| 95 |
+
(APIServer pid=1142905) INFO 08-10 21:20:59 [loggers.py:259] Engine 000: Avg prompt throughput: 2107.1 tokens/s, Avg generation throughput: 3019.1 tokens/s, Running: 255 reqs, Waiting: 685 reqs, GPU KV cache usage: 51.1%, Prefix cache hit rate: 93.3%
|
| 96 |
+
(APIServer pid=1142905) INFO 08-10 21:21:09 [loggers.py:259] Engine 000: Avg prompt throughput: 2170.0 tokens/s, Avg generation throughput: 2967.4 tokens/s, Running: 256 reqs, Waiting: 409 reqs, GPU KV cache usage: 51.3%, Prefix cache hit rate: 93.4%
|
| 97 |
+
(APIServer pid=1142905) INFO 08-10 21:21:19 [loggers.py:259] Engine 000: Avg prompt throughput: 2265.7 tokens/s, Avg generation throughput: 2966.5 tokens/s, Running: 255 reqs, Waiting: 128 reqs, GPU KV cache usage: 51.7%, Prefix cache hit rate: 93.4%
|
| 98 |
+
(APIServer pid=1142905) INFO 08-10 21:21:29 [loggers.py:259] Engine 000: Avg prompt throughput: 1072.9 tokens/s, Avg generation throughput: 2630.6 tokens/s, Running: 103 reqs, Waiting: 0 reqs, GPU KV cache usage: 31.6%, Prefix cache hit rate: 93.4%
|
| 99 |
+
(APIServer pid=1142905) INFO: 127.0.0.1:54568 - "POST /v1/completions HTTP/1.1" 200 OK
|
| 100 |
+
(EngineCore pid=1144304) INFO 08-10 21:21:38 [core.py:1210] Shutdown initiated (timeout=0)
|
| 101 |
+
(EngineCore pid=1144304) INFO 08-10 21:21:38 [core.py:1233] Shutdown complete
|
| 102 |
+
(APIServer pid=1142905) INFO: Shutting down
|
| 103 |
+
(APIServer pid=1142905) INFO: Waiting for application shutdown.
|
| 104 |
+
(APIServer pid=1142905) INFO: Application shutdown complete.
|
| 105 |
+
(APIServer pid=1142905) INFO: Finished server process [1142905]
|
evals/grid_math_unhealed/uniform_keep25_unhealed.log
ADDED
|
@@ -0,0 +1,8 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"correct": 118,
|
| 3 |
+
"accuracy": 0.08946171341925702,
|
| 4 |
+
"finished": 272,
|
| 5 |
+
"finish_rate": 0.20621683093252463,
|
| 6 |
+
"mean_completion_tokens": 477.21379833206976
|
| 7 |
+
}
|
| 8 |
+
saved item-level results -> outputs/evals/grid_math_unhealed/uniform_keep25_unhealed_chat.json
|
evals/grid_math_unhealed/uniform_keep25_unhealed_chat.json
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
evals/grid_math_unhealed/uniform_keep25_unhealed_chat.json.server.log
ADDED
|
@@ -0,0 +1,115 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
Skipping import of cpp extensions due to incompatible torch version 2.10.0+cu128 for torchao version 0.15.0 Please see https://github.com/pytorch/ao/issues/2919 for more info
|
| 2 |
+
WARNING 08-10 21:24:57 [registry.py:915] Model architecture OlmoeForCausalLM is already registered, and will be overwritten by the new model class glean_vllm.pruned_olmoe:PrunedOlmoeForCausalLM.
|
| 3 |
+
(APIServer pid=1153465) INFO 08-10 21:24:57 [utils.py:299]
|
| 4 |
+
(APIServer pid=1153465) INFO 08-10 21:24:57 [utils.py:299] β β ββ ββ
|
| 5 |
+
(APIServer pid=1153465) INFO 08-10 21:24:57 [utils.py:299] ββ ββ β β β βββ β version 0.19.0
|
| 6 |
+
(APIServer pid=1153465) INFO 08-10 21:24:57 [utils.py:299] ββββ β β β β model /media/henry/MoreFiles/winnow_release/math_arms/uniform_keep25
|
| 7 |
+
(APIServer pid=1153465) INFO 08-10 21:24:57 [utils.py:299] ββ βββββ βββββ β β
|
| 8 |
+
(APIServer pid=1153465) INFO 08-10 21:24:57 [utils.py:299]
|
| 9 |
+
(APIServer pid=1153465) INFO 08-10 21:24:57 [utils.py:233] non-default args: {'model_tag': '/media/henry/MoreFiles/winnow_release/math_arms/uniform_keep25', 'host': '127.0.0.1', 'port': 8612, 'model': '/media/henry/MoreFiles/winnow_release/math_arms/uniform_keep25', 'max_model_len': 2048, 'enforce_eager': True, 'served_model_name': ['student'], 'gpu_memory_utilization': 0.85}
|
| 10 |
+
(APIServer pid=1153465) INFO 08-10 21:25:05 [model.py:549] Resolved architecture: PrunedOlmoeForCausalLM
|
| 11 |
+
(APIServer pid=1153465) INFO 08-10 21:25:05 [model.py:1678] Using max model len 2048
|
| 12 |
+
(APIServer pid=1153465) INFO 08-10 21:25:06 [vllm.py:790] Asynchronous scheduling is enabled.
|
| 13 |
+
(APIServer pid=1153465) WARNING 08-10 21:25:06 [vllm.py:848] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
|
| 14 |
+
(APIServer pid=1153465) WARNING 08-10 21:25:06 [vllm.py:859] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
|
| 15 |
+
(APIServer pid=1153465) INFO 08-10 21:25:06 [vllm.py:1025] Cudagraph is disabled under eager mode
|
| 16 |
+
(APIServer pid=1153465) INFO 08-10 21:25:06 [compilation.py:290] Enabled custom fusions: norm_quant, act_quant
|
| 17 |
+
Skipping import of cpp extensions due to incompatible torch version 2.10.0+cu128 for torchao version 0.15.0 Please see https://github.com/pytorch/ao/issues/2919 for more info
|
| 18 |
+
(EngineCore pid=1154743) WARNING 08-10 21:25:14 [registry.py:915] Model architecture OlmoeForCausalLM is already registered, and will be overwritten by the new model class glean_vllm.pruned_olmoe:PrunedOlmoeForCausalLM.
|
| 19 |
+
(EngineCore pid=1154743) INFO 08-10 21:25:14 [core.py:105] Initializing a V1 LLM engine (v0.19.0) with config: model='/media/henry/MoreFiles/winnow_release/math_arms/uniform_keep25', speculative_config=None, tokenizer='/media/henry/MoreFiles/winnow_release/math_arms/uniform_keep25', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=False, dtype=torch.bfloat16, max_seq_len=2048, download_dir=None, load_format=auto, tensor_parallel_size=1, pipeline_parallel_size=1, data_parallel_size=1, decode_context_parallel_size=1, dcp_comm_backend=ag_rs, disable_custom_all_reduce=False, quantization=None, enforce_eager=True, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False), seed=0, served_model_name=student, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'mode': <CompilationMode.NONE: 0>, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'splitting_ops': [], 'compile_mm_encoder': False, 'cudagraph_mm_encoder': False, 'encoder_cudagraph_token_budgets': [], 'encoder_cudagraph_max_images_per_batch': 0, 'compile_sizes': [], 'compile_ranges_endpoints': [2048], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'size_asserts': False, 'alignment_asserts': False, 'scalar_asserts': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': <CUDAGraphMode.NONE: 0>, 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': True, 'fuse_act_quant': True, 'fuse_attn_quant': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False}, 'max_cudagraph_capture_size': 0, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': True, 'static_all_moe_layers': []}
|
| 20 |
+
(EngineCore pid=1154743) INFO 08-10 21:25:14 [parallel_state.py:1400] world_size=1 rank=0 local_rank=0 distributed_init_method=tcp://192.168.0.27:44057 backend=nccl
|
| 21 |
+
(EngineCore pid=1154743) INFO 08-10 21:25:14 [parallel_state.py:1716] rank 0 in world size 1 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank 0, EPLB rank N/A
|
| 22 |
+
(EngineCore pid=1154743) INFO 08-10 21:25:15 [gpu_model_runner.py:4735] Starting to load model /media/henry/MoreFiles/winnow_release/math_arms/uniform_keep25...
|
| 23 |
+
(EngineCore pid=1154743) INFO 08-10 21:25:16 [cuda.py:334] Using FLASH_ATTN attention backend out of potential backends: ['FLASH_ATTN', 'FLASHINFER', 'TRITON_ATTN', 'FLEX_ATTENTION'].
|
| 24 |
+
(EngineCore pid=1154743) INFO 08-10 21:25:16 [flash_attn.py:596] Using FlashAttention version 2
|
| 25 |
+
(EngineCore pid=1154743)
|
| 26 |
+
(EngineCore pid=1154743)
|
| 27 |
+
(EngineCore pid=1154743)
|
| 28 |
+
(EngineCore pid=1154743)
|
| 29 |
+
(EngineCore pid=1154743) INFO 08-10 21:25:20 [default_loader.py:384] Loading weights took 4.10 seconds
|
| 30 |
+
(EngineCore pid=1154743) INFO 08-10 21:25:21 [gpu_model_runner.py:4820] Model loading took 3.89 GiB memory and 4.696206 seconds
|
| 31 |
+
(EngineCore pid=1154743) INFO 08-10 21:25:23 [gpu_worker.py:436] Available KV cache memory: 15.86 GiB
|
| 32 |
+
(EngineCore pid=1154743) INFO 08-10 21:25:23 [kv_cache_utils.py:1319] GPU KV cache size: 129,888 tokens
|
| 33 |
+
(EngineCore pid=1154743) INFO 08-10 21:25:23 [kv_cache_utils.py:1324] Maximum concurrency for 2,048 tokens per request: 63.42x
|
| 34 |
+
(EngineCore pid=1154743) INFO 08-10 21:25:24 [core.py:283] init engine (profile, create kv cache, warmup model) took 2.86 seconds
|
| 35 |
+
(EngineCore pid=1154743) INFO 08-10 21:25:24 [vllm.py:790] Asynchronous scheduling is enabled.
|
| 36 |
+
(EngineCore pid=1154743) WARNING 08-10 21:25:24 [vllm.py:848] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
|
| 37 |
+
(EngineCore pid=1154743) WARNING 08-10 21:25:24 [vllm.py:859] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
|
| 38 |
+
(EngineCore pid=1154743) INFO 08-10 21:25:24 [vllm.py:1025] Cudagraph is disabled under eager mode
|
| 39 |
+
(EngineCore pid=1154743) INFO 08-10 21:25:24 [compilation.py:290] Enabled custom fusions: norm_quant, act_quant
|
| 40 |
+
(APIServer pid=1153465) INFO 08-10 21:25:24 [api_server.py:590] Supported tasks: ['generate']
|
| 41 |
+
(APIServer pid=1153465) WARNING 08-10 21:25:24 [__init__.py:14] SECURITY WARNING: Development endpoints are enabled! This should NOT be used in production!
|
| 42 |
+
(APIServer pid=1153465) INFO 08-10 21:25:24 [hf.py:314] Detected the chat template content format to be 'string'. You can set `--chat-template-content-format` to override this.
|
| 43 |
+
(APIServer pid=1153465) INFO 08-10 21:25:24 [api_server.py:594] Starting vLLM server on http://127.0.0.1:8612
|
| 44 |
+
(APIServer pid=1153465) INFO 08-10 21:25:24 [launcher.py:37] Available routes are:
|
| 45 |
+
(APIServer pid=1153465) INFO 08-10 21:25:24 [launcher.py:46] Route: /openapi.json, Methods: GET, HEAD
|
| 46 |
+
(APIServer pid=1153465) INFO 08-10 21:25:24 [launcher.py:46] Route: /docs, Methods: GET, HEAD
|
| 47 |
+
(APIServer pid=1153465) INFO 08-10 21:25:24 [launcher.py:46] Route: /docs/oauth2-redirect, Methods: GET, HEAD
|
| 48 |
+
(APIServer pid=1153465) INFO 08-10 21:25:24 [launcher.py:46] Route: /redoc, Methods: GET, HEAD
|
| 49 |
+
(APIServer pid=1153465) INFO 08-10 21:25:24 [launcher.py:46] Route: /sleep, Methods: POST
|
| 50 |
+
(APIServer pid=1153465) INFO 08-10 21:25:24 [launcher.py:46] Route: /wake_up, Methods: POST
|
| 51 |
+
(APIServer pid=1153465) INFO 08-10 21:25:24 [launcher.py:46] Route: /is_sleeping, Methods: GET
|
| 52 |
+
(APIServer pid=1153465) INFO 08-10 21:25:24 [launcher.py:46] Route: /collective_rpc, Methods: POST
|
| 53 |
+
(APIServer pid=1153465) INFO 08-10 21:25:24 [launcher.py:46] Route: /reset_prefix_cache, Methods: POST
|
| 54 |
+
(APIServer pid=1153465) INFO 08-10 21:25:24 [launcher.py:46] Route: /reset_mm_cache, Methods: POST
|
| 55 |
+
(APIServer pid=1153465) INFO 08-10 21:25:24 [launcher.py:46] Route: /reset_encoder_cache, Methods: POST
|
| 56 |
+
(APIServer pid=1153465) INFO 08-10 21:25:24 [launcher.py:46] Route: /tokenize, Methods: POST
|
| 57 |
+
(APIServer pid=1153465) INFO 08-10 21:25:24 [launcher.py:46] Route: /detokenize, Methods: POST
|
| 58 |
+
(APIServer pid=1153465) INFO 08-10 21:25:24 [launcher.py:46] Route: /load, Methods: GET
|
| 59 |
+
(APIServer pid=1153465) INFO 08-10 21:25:24 [launcher.py:46] Route: /version, Methods: GET
|
| 60 |
+
(APIServer pid=1153465) INFO 08-10 21:25:24 [launcher.py:46] Route: /health, Methods: GET
|
| 61 |
+
(APIServer pid=1153465) INFO 08-10 21:25:24 [launcher.py:46] Route: /metrics, Methods: GET
|
| 62 |
+
(APIServer pid=1153465) INFO 08-10 21:25:24 [launcher.py:46] Route: /server_info, Methods: GET
|
| 63 |
+
(APIServer pid=1153465) INFO 08-10 21:25:24 [launcher.py:46] Route: /v1/models, Methods: GET
|
| 64 |
+
(APIServer pid=1153465) INFO 08-10 21:25:24 [launcher.py:46] Route: /ping, Methods: GET
|
| 65 |
+
(APIServer pid=1153465) INFO 08-10 21:25:24 [launcher.py:46] Route: /ping, Methods: POST
|
| 66 |
+
(APIServer pid=1153465) INFO 08-10 21:25:24 [launcher.py:46] Route: /invocations, Methods: POST
|
| 67 |
+
(APIServer pid=1153465) INFO 08-10 21:25:24 [launcher.py:46] Route: /v1/chat/completions, Methods: POST
|
| 68 |
+
(APIServer pid=1153465) INFO 08-10 21:25:24 [launcher.py:46] Route: /v1/chat/completions/batch, Methods: POST
|
| 69 |
+
(APIServer pid=1153465) INFO 08-10 21:25:24 [launcher.py:46] Route: /v1/responses, Methods: POST
|
| 70 |
+
(APIServer pid=1153465) INFO 08-10 21:25:24 [launcher.py:46] Route: /v1/responses/{response_id}, Methods: GET
|
| 71 |
+
(APIServer pid=1153465) INFO 08-10 21:25:24 [launcher.py:46] Route: /v1/responses/{response_id}/cancel, Methods: POST
|
| 72 |
+
(APIServer pid=1153465) INFO 08-10 21:25:24 [launcher.py:46] Route: /v1/completions, Methods: POST
|
| 73 |
+
(APIServer pid=1153465) INFO 08-10 21:25:24 [launcher.py:46] Route: /v1/messages, Methods: POST
|
| 74 |
+
(APIServer pid=1153465) INFO 08-10 21:25:24 [launcher.py:46] Route: /v1/messages/count_tokens, Methods: POST
|
| 75 |
+
(APIServer pid=1153465) INFO 08-10 21:25:24 [launcher.py:46] Route: /inference/v1/generate, Methods: POST
|
| 76 |
+
(APIServer pid=1153465) INFO 08-10 21:25:24 [launcher.py:46] Route: /pause, Methods: POST
|
| 77 |
+
(APIServer pid=1153465) INFO 08-10 21:25:24 [launcher.py:46] Route: /resume, Methods: POST
|
| 78 |
+
(APIServer pid=1153465) INFO 08-10 21:25:24 [launcher.py:46] Route: /is_paused, Methods: GET
|
| 79 |
+
(APIServer pid=1153465) INFO 08-10 21:25:24 [launcher.py:46] Route: /init_weight_transfer_engine, Methods: POST
|
| 80 |
+
(APIServer pid=1153465) INFO 08-10 21:25:24 [launcher.py:46] Route: /update_weights, Methods: POST
|
| 81 |
+
(APIServer pid=1153465) INFO 08-10 21:25:24 [launcher.py:46] Route: /get_world_size, Methods: GET
|
| 82 |
+
(APIServer pid=1153465) INFO 08-10 21:25:24 [launcher.py:46] Route: /scale_elastic_ep, Methods: POST
|
| 83 |
+
(APIServer pid=1153465) INFO 08-10 21:25:24 [launcher.py:46] Route: /is_scaling_elastic_ep, Methods: POST
|
| 84 |
+
(APIServer pid=1153465) INFO 08-10 21:25:24 [launcher.py:46] Route: /v1/chat/completions/render, Methods: POST
|
| 85 |
+
(APIServer pid=1153465) INFO 08-10 21:25:24 [launcher.py:46] Route: /v1/completions/render, Methods: POST
|
| 86 |
+
(APIServer pid=1153465) INFO: Started server process [1153465]
|
| 87 |
+
(APIServer pid=1153465) INFO: Waiting for application startup.
|
| 88 |
+
(APIServer pid=1153465) INFO: Application startup complete.
|
| 89 |
+
(APIServer pid=1153465) INFO: 127.0.0.1:33884 - "GET /health HTTP/1.1" 200 OK
|
| 90 |
+
(APIServer pid=1153465) INFO 08-10 21:25:35 [loggers.py:259] Engine 000: Avg prompt throughput: 2041.2 tokens/s, Avg generation throughput: 2581.6 tokens/s, Running: 256 reqs, Waiting: 1060 reqs, GPU KV cache usage: 38.7%, Prefix cache hit rate: 93.1%
|
| 91 |
+
(APIServer pid=1153465) INFO 08-10 21:25:45 [loggers.py:259] Engine 000: Avg prompt throughput: 23.3 tokens/s, Avg generation throughput: 3582.5 tokens/s, Running: 256 reqs, Waiting: 1056 reqs, GPU KV cache usage: 65.7%, Prefix cache hit rate: 93.1%
|
| 92 |
+
(APIServer pid=1153465) INFO 08-10 21:25:55 [loggers.py:259] Engine 000: Avg prompt throughput: 77.7 tokens/s, Avg generation throughput: 3325.8 tokens/s, Running: 256 reqs, Waiting: 1045 reqs, GPU KV cache usage: 88.7%, Prefix cache hit rate: 93.1%
|
| 93 |
+
(APIServer pid=1153465) INFO 08-10 21:26:05 [loggers.py:259] Engine 000: Avg prompt throughput: 145.0 tokens/s, Avg generation throughput: 2930.4 tokens/s, Running: 229 reqs, Waiting: 1046 reqs, GPU KV cache usage: 99.9%, Prefix cache hit rate: 93.1%
|
| 94 |
+
(APIServer pid=1153465) INFO 08-10 21:26:15 [loggers.py:259] Engine 000: Avg prompt throughput: 1818.6 tokens/s, Avg generation throughput: 3485.2 tokens/s, Running: 256 reqs, Waiting: 796 reqs, GPU KV cache usage: 43.2%, Prefix cache hit rate: 93.3%
|
| 95 |
+
(APIServer pid=1153465) INFO 08-10 21:26:25 [loggers.py:259] Engine 000: Avg prompt throughput: 73.1 tokens/s, Avg generation throughput: 3530.9 tokens/s, Running: 256 reqs, Waiting: 787 reqs, GPU KV cache usage: 68.5%, Prefix cache hit rate: 93.3%
|
| 96 |
+
(APIServer pid=1153465) INFO 08-10 21:26:35 [loggers.py:259] Engine 000: Avg prompt throughput: 114.3 tokens/s, Avg generation throughput: 3299.9 tokens/s, Running: 255 reqs, Waiting: 770 reqs, GPU KV cache usage: 87.7%, Prefix cache hit rate: 93.3%
|
| 97 |
+
(APIServer pid=1153465) INFO 08-10 21:26:45 [loggers.py:259] Engine 000: Avg prompt throughput: 487.9 tokens/s, Avg generation throughput: 3165.6 tokens/s, Running: 239 reqs, Waiting: 690 reqs, GPU KV cache usage: 75.2%, Prefix cache hit rate: 93.3%
|
| 98 |
+
(APIServer pid=1153465) INFO 08-10 21:26:55 [loggers.py:259] Engine 000: Avg prompt throughput: 1483.6 tokens/s, Avg generation throughput: 3591.5 tokens/s, Running: 256 reqs, Waiting: 520 reqs, GPU KV cache usage: 49.4%, Prefix cache hit rate: 93.3%
|
| 99 |
+
(APIServer pid=1153465) INFO 08-10 21:27:05 [loggers.py:259] Engine 000: Avg prompt throughput: 153.1 tokens/s, Avg generation throughput: 3478.6 tokens/s, Running: 256 reqs, Waiting: 502 reqs, GPU KV cache usage: 70.8%, Prefix cache hit rate: 93.3%
|
| 100 |
+
(APIServer pid=1153465) INFO 08-10 21:27:15 [loggers.py:259] Engine 000: Avg prompt throughput: 247.5 tokens/s, Avg generation throughput: 3298.3 tokens/s, Running: 255 reqs, Waiting: 471 reqs, GPU KV cache usage: 85.2%, Prefix cache hit rate: 93.3%
|
| 101 |
+
(APIServer pid=1153465) INFO 08-10 21:27:25 [loggers.py:259] Engine 000: Avg prompt throughput: 1526.9 tokens/s, Avg generation throughput: 3154.4 tokens/s, Running: 256 reqs, Waiting: 283 reqs, GPU KV cache usage: 37.7%, Prefix cache hit rate: 93.4%
|
| 102 |
+
(APIServer pid=1153465) INFO 08-10 21:27:35 [loggers.py:259] Engine 000: Avg prompt throughput: 279.5 tokens/s, Avg generation throughput: 3605.4 tokens/s, Running: 256 reqs, Waiting: 246 reqs, GPU KV cache usage: 54.2%, Prefix cache hit rate: 93.4%
|
| 103 |
+
(APIServer pid=1153465) INFO 08-10 21:27:45 [loggers.py:259] Engine 000: Avg prompt throughput: 155.5 tokens/s, Avg generation throughput: 3453.2 tokens/s, Running: 256 reqs, Waiting: 228 reqs, GPU KV cache usage: 74.6%, Prefix cache hit rate: 93.4%
|
| 104 |
+
(APIServer pid=1153465) INFO 08-10 21:27:55 [loggers.py:259] Engine 000: Avg prompt throughput: 358.4 tokens/s, Avg generation throughput: 3271.9 tokens/s, Running: 256 reqs, Waiting: 181 reqs, GPU KV cache usage: 82.9%, Prefix cache hit rate: 93.4%
|
| 105 |
+
(APIServer pid=1153465) INFO 08-10 21:28:05 [loggers.py:259] Engine 000: Avg prompt throughput: 1343.6 tokens/s, Avg generation throughput: 3259.3 tokens/s, Running: 256 reqs, Waiting: 15 reqs, GPU KV cache usage: 44.9%, Prefix cache hit rate: 93.4%
|
| 106 |
+
(APIServer pid=1153465) INFO 08-10 21:28:15 [loggers.py:259] Engine 000: Avg prompt throughput: 116.1 tokens/s, Avg generation throughput: 3538.8 tokens/s, Running: 228 reqs, Waiting: 0 reqs, GPU KV cache usage: 55.7%, Prefix cache hit rate: 93.4%
|
| 107 |
+
(APIServer pid=1153465) INFO 08-10 21:28:25 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 3280.1 tokens/s, Running: 202 reqs, Waiting: 0 reqs, GPU KV cache usage: 69.6%, Prefix cache hit rate: 93.4%
|
| 108 |
+
(APIServer pid=1153465) INFO 08-10 21:28:35 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 2844.0 tokens/s, Running: 79 reqs, Waiting: 0 reqs, GPU KV cache usage: 36.5%, Prefix cache hit rate: 93.4%
|
| 109 |
+
(APIServer pid=1153465) INFO: 127.0.0.1:33886 - "POST /v1/completions HTTP/1.1" 200 OK
|
| 110 |
+
(EngineCore pid=1154743) INFO 08-10 21:28:37 [core.py:1210] Shutdown initiated (timeout=0)
|
| 111 |
+
(EngineCore pid=1154743) INFO 08-10 21:28:37 [core.py:1233] Shutdown complete
|
| 112 |
+
(APIServer pid=1153465) INFO: Shutting down
|
| 113 |
+
(APIServer pid=1153465) INFO: Waiting for application shutdown.
|
| 114 |
+
(APIServer pid=1153465) INFO: Application shutdown complete.
|
| 115 |
+
(APIServer pid=1153465) INFO: Finished server process [1153465]
|
evals/grid_math_unhealed/uniform_keep50_unhealed.log
ADDED
|
@@ -0,0 +1,8 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"correct": 548,
|
| 3 |
+
"accuracy": 0.41546626231993933,
|
| 4 |
+
"finished": 1315,
|
| 5 |
+
"finish_rate": 0.9969673995451099,
|
| 6 |
+
"mean_completion_tokens": 94.5352539802881
|
| 7 |
+
}
|
| 8 |
+
saved item-level results -> outputs/evals/grid_math_unhealed/uniform_keep50_unhealed_chat.json
|
evals/grid_math_unhealed/uniform_keep50_unhealed_chat.json
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
evals/grid_math_unhealed/uniform_keep50_unhealed_chat.json.server.log
ADDED
|
@@ -0,0 +1,101 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
Skipping import of cpp extensions due to incompatible torch version 2.10.0+cu128 for torchao version 0.15.0 Please see https://github.com/pytorch/ao/issues/2919 for more info
|
| 2 |
+
WARNING 08-10 21:22:07 [registry.py:915] Model architecture OlmoeForCausalLM is already registered, and will be overwritten by the new model class glean_vllm.pruned_olmoe:PrunedOlmoeForCausalLM.
|
| 3 |
+
(APIServer pid=1147687) INFO 08-10 21:22:07 [utils.py:299]
|
| 4 |
+
(APIServer pid=1147687) INFO 08-10 21:22:07 [utils.py:299] β β ββ ββ
|
| 5 |
+
(APIServer pid=1147687) INFO 08-10 21:22:07 [utils.py:299] ββ ββ β β β βββ β version 0.19.0
|
| 6 |
+
(APIServer pid=1147687) INFO 08-10 21:22:07 [utils.py:299] ββββ β β β β model /media/henry/MoreFiles/winnow_release/math_arms/uniform_keep50
|
| 7 |
+
(APIServer pid=1147687) INFO 08-10 21:22:07 [utils.py:299] ββ βββββ βββββ β β
|
| 8 |
+
(APIServer pid=1147687) INFO 08-10 21:22:07 [utils.py:299]
|
| 9 |
+
(APIServer pid=1147687) INFO 08-10 21:22:07 [utils.py:233] non-default args: {'model_tag': '/media/henry/MoreFiles/winnow_release/math_arms/uniform_keep50', 'host': '127.0.0.1', 'port': 8612, 'model': '/media/henry/MoreFiles/winnow_release/math_arms/uniform_keep50', 'max_model_len': 2048, 'enforce_eager': True, 'served_model_name': ['student'], 'gpu_memory_utilization': 0.85}
|
| 10 |
+
(APIServer pid=1147687) INFO 08-10 21:22:15 [model.py:549] Resolved architecture: PrunedOlmoeForCausalLM
|
| 11 |
+
(APIServer pid=1147687) INFO 08-10 21:22:15 [model.py:1678] Using max model len 2048
|
| 12 |
+
(APIServer pid=1147687) INFO 08-10 21:22:15 [vllm.py:790] Asynchronous scheduling is enabled.
|
| 13 |
+
(APIServer pid=1147687) WARNING 08-10 21:22:15 [vllm.py:848] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
|
| 14 |
+
(APIServer pid=1147687) WARNING 08-10 21:22:15 [vllm.py:859] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
|
| 15 |
+
(APIServer pid=1147687) INFO 08-10 21:22:15 [vllm.py:1025] Cudagraph is disabled under eager mode
|
| 16 |
+
(APIServer pid=1147687) INFO 08-10 21:22:15 [compilation.py:290] Enabled custom fusions: norm_quant, act_quant
|
| 17 |
+
Skipping import of cpp extensions due to incompatible torch version 2.10.0+cu128 for torchao version 0.15.0 Please see https://github.com/pytorch/ao/issues/2919 for more info
|
| 18 |
+
(EngineCore pid=1148946) WARNING 08-10 21:22:24 [registry.py:915] Model architecture OlmoeForCausalLM is already registered, and will be overwritten by the new model class glean_vllm.pruned_olmoe:PrunedOlmoeForCausalLM.
|
| 19 |
+
(EngineCore pid=1148946) INFO 08-10 21:22:24 [core.py:105] Initializing a V1 LLM engine (v0.19.0) with config: model='/media/henry/MoreFiles/winnow_release/math_arms/uniform_keep50', speculative_config=None, tokenizer='/media/henry/MoreFiles/winnow_release/math_arms/uniform_keep50', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=False, dtype=torch.bfloat16, max_seq_len=2048, download_dir=None, load_format=auto, tensor_parallel_size=1, pipeline_parallel_size=1, data_parallel_size=1, decode_context_parallel_size=1, dcp_comm_backend=ag_rs, disable_custom_all_reduce=False, quantization=None, enforce_eager=True, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False), seed=0, served_model_name=student, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'mode': <CompilationMode.NONE: 0>, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'splitting_ops': [], 'compile_mm_encoder': False, 'cudagraph_mm_encoder': False, 'encoder_cudagraph_token_budgets': [], 'encoder_cudagraph_max_images_per_batch': 0, 'compile_sizes': [], 'compile_ranges_endpoints': [2048], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'size_asserts': False, 'alignment_asserts': False, 'scalar_asserts': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': <CUDAGraphMode.NONE: 0>, 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': True, 'fuse_act_quant': True, 'fuse_attn_quant': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False}, 'max_cudagraph_capture_size': 0, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': True, 'static_all_moe_layers': []}
|
| 20 |
+
(EngineCore pid=1148946) INFO 08-10 21:22:24 [parallel_state.py:1400] world_size=1 rank=0 local_rank=0 distributed_init_method=tcp://192.168.0.27:47297 backend=nccl
|
| 21 |
+
(EngineCore pid=1148946) INFO 08-10 21:22:24 [parallel_state.py:1716] rank 0 in world size 1 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank 0, EPLB rank N/A
|
| 22 |
+
(EngineCore pid=1148946) INFO 08-10 21:22:24 [gpu_model_runner.py:4735] Starting to load model /media/henry/MoreFiles/winnow_release/math_arms/uniform_keep50...
|
| 23 |
+
(EngineCore pid=1148946) INFO 08-10 21:22:25 [cuda.py:334] Using FLASH_ATTN attention backend out of potential backends: ['FLASH_ATTN', 'FLASHINFER', 'TRITON_ATTN', 'FLEX_ATTENTION'].
|
| 24 |
+
(EngineCore pid=1148946) INFO 08-10 21:22:25 [flash_attn.py:596] Using FlashAttention version 2
|
| 25 |
+
(EngineCore pid=1148946)
|
| 26 |
+
(EngineCore pid=1148946)
|
| 27 |
+
(EngineCore pid=1148946)
|
| 28 |
+
(EngineCore pid=1148946)
|
| 29 |
+
(EngineCore pid=1148946)
|
| 30 |
+
(EngineCore pid=1148946) INFO 08-10 21:22:34 [default_loader.py:384] Loading weights took 8.47 seconds
|
| 31 |
+
(EngineCore pid=1148946) INFO 08-10 21:22:34 [gpu_model_runner.py:4820] Model loading took 6.89 GiB memory and 9.059107 seconds
|
| 32 |
+
(EngineCore pid=1148946) INFO 08-10 21:22:37 [gpu_worker.py:436] Available KV cache memory: 12.86 GiB
|
| 33 |
+
(EngineCore pid=1148946) INFO 08-10 21:22:37 [kv_cache_utils.py:1319] GPU KV cache size: 105,312 tokens
|
| 34 |
+
(EngineCore pid=1148946) INFO 08-10 21:22:37 [kv_cache_utils.py:1324] Maximum concurrency for 2,048 tokens per request: 51.42x
|
| 35 |
+
(EngineCore pid=1148946) INFO 08-10 21:22:37 [core.py:283] init engine (profile, create kv cache, warmup model) took 2.42 seconds
|
| 36 |
+
(EngineCore pid=1148946) INFO 08-10 21:22:37 [vllm.py:790] Asynchronous scheduling is enabled.
|
| 37 |
+
(EngineCore pid=1148946) WARNING 08-10 21:22:37 [vllm.py:848] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
|
| 38 |
+
(EngineCore pid=1148946) WARNING 08-10 21:22:37 [vllm.py:859] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
|
| 39 |
+
(EngineCore pid=1148946) INFO 08-10 21:22:37 [vllm.py:1025] Cudagraph is disabled under eager mode
|
| 40 |
+
(EngineCore pid=1148946) INFO 08-10 21:22:37 [compilation.py:290] Enabled custom fusions: norm_quant, act_quant
|
| 41 |
+
(APIServer pid=1147687) INFO 08-10 21:22:37 [api_server.py:590] Supported tasks: ['generate']
|
| 42 |
+
(APIServer pid=1147687) WARNING 08-10 21:22:37 [__init__.py:14] SECURITY WARNING: Development endpoints are enabled! This should NOT be used in production!
|
| 43 |
+
(APIServer pid=1147687) INFO 08-10 21:22:37 [hf.py:314] Detected the chat template content format to be 'string'. You can set `--chat-template-content-format` to override this.
|
| 44 |
+
(APIServer pid=1147687) INFO 08-10 21:22:37 [api_server.py:594] Starting vLLM server on http://127.0.0.1:8612
|
| 45 |
+
(APIServer pid=1147687) INFO 08-10 21:22:37 [launcher.py:37] Available routes are:
|
| 46 |
+
(APIServer pid=1147687) INFO 08-10 21:22:37 [launcher.py:46] Route: /openapi.json, Methods: GET, HEAD
|
| 47 |
+
(APIServer pid=1147687) INFO 08-10 21:22:37 [launcher.py:46] Route: /docs, Methods: GET, HEAD
|
| 48 |
+
(APIServer pid=1147687) INFO 08-10 21:22:37 [launcher.py:46] Route: /docs/oauth2-redirect, Methods: GET, HEAD
|
| 49 |
+
(APIServer pid=1147687) INFO 08-10 21:22:37 [launcher.py:46] Route: /redoc, Methods: GET, HEAD
|
| 50 |
+
(APIServer pid=1147687) INFO 08-10 21:22:37 [launcher.py:46] Route: /sleep, Methods: POST
|
| 51 |
+
(APIServer pid=1147687) INFO 08-10 21:22:37 [launcher.py:46] Route: /wake_up, Methods: POST
|
| 52 |
+
(APIServer pid=1147687) INFO 08-10 21:22:37 [launcher.py:46] Route: /is_sleeping, Methods: GET
|
| 53 |
+
(APIServer pid=1147687) INFO 08-10 21:22:37 [launcher.py:46] Route: /collective_rpc, Methods: POST
|
| 54 |
+
(APIServer pid=1147687) INFO 08-10 21:22:37 [launcher.py:46] Route: /reset_prefix_cache, Methods: POST
|
| 55 |
+
(APIServer pid=1147687) INFO 08-10 21:22:37 [launcher.py:46] Route: /reset_mm_cache, Methods: POST
|
| 56 |
+
(APIServer pid=1147687) INFO 08-10 21:22:37 [launcher.py:46] Route: /reset_encoder_cache, Methods: POST
|
| 57 |
+
(APIServer pid=1147687) INFO 08-10 21:22:37 [launcher.py:46] Route: /tokenize, Methods: POST
|
| 58 |
+
(APIServer pid=1147687) INFO 08-10 21:22:37 [launcher.py:46] Route: /detokenize, Methods: POST
|
| 59 |
+
(APIServer pid=1147687) INFO 08-10 21:22:37 [launcher.py:46] Route: /load, Methods: GET
|
| 60 |
+
(APIServer pid=1147687) INFO 08-10 21:22:37 [launcher.py:46] Route: /version, Methods: GET
|
| 61 |
+
(APIServer pid=1147687) INFO 08-10 21:22:37 [launcher.py:46] Route: /health, Methods: GET
|
| 62 |
+
(APIServer pid=1147687) INFO 08-10 21:22:37 [launcher.py:46] Route: /metrics, Methods: GET
|
| 63 |
+
(APIServer pid=1147687) INFO 08-10 21:22:37 [launcher.py:46] Route: /server_info, Methods: GET
|
| 64 |
+
(APIServer pid=1147687) INFO 08-10 21:22:37 [launcher.py:46] Route: /v1/models, Methods: GET
|
| 65 |
+
(APIServer pid=1147687) INFO 08-10 21:22:37 [launcher.py:46] Route: /ping, Methods: GET
|
| 66 |
+
(APIServer pid=1147687) INFO 08-10 21:22:37 [launcher.py:46] Route: /ping, Methods: POST
|
| 67 |
+
(APIServer pid=1147687) INFO 08-10 21:22:37 [launcher.py:46] Route: /invocations, Methods: POST
|
| 68 |
+
(APIServer pid=1147687) INFO 08-10 21:22:37 [launcher.py:46] Route: /v1/chat/completions, Methods: POST
|
| 69 |
+
(APIServer pid=1147687) INFO 08-10 21:22:37 [launcher.py:46] Route: /v1/chat/completions/batch, Methods: POST
|
| 70 |
+
(APIServer pid=1147687) INFO 08-10 21:22:37 [launcher.py:46] Route: /v1/responses, Methods: POST
|
| 71 |
+
(APIServer pid=1147687) INFO 08-10 21:22:37 [launcher.py:46] Route: /v1/responses/{response_id}, Methods: GET
|
| 72 |
+
(APIServer pid=1147687) INFO 08-10 21:22:37 [launcher.py:46] Route: /v1/responses/{response_id}/cancel, Methods: POST
|
| 73 |
+
(APIServer pid=1147687) INFO 08-10 21:22:37 [launcher.py:46] Route: /v1/completions, Methods: POST
|
| 74 |
+
(APIServer pid=1147687) INFO 08-10 21:22:37 [launcher.py:46] Route: /v1/messages, Methods: POST
|
| 75 |
+
(APIServer pid=1147687) INFO 08-10 21:22:37 [launcher.py:46] Route: /v1/messages/count_tokens, Methods: POST
|
| 76 |
+
(APIServer pid=1147687) INFO 08-10 21:22:37 [launcher.py:46] Route: /inference/v1/generate, Methods: POST
|
| 77 |
+
(APIServer pid=1147687) INFO 08-10 21:22:37 [launcher.py:46] Route: /pause, Methods: POST
|
| 78 |
+
(APIServer pid=1147687) INFO 08-10 21:22:37 [launcher.py:46] Route: /resume, Methods: POST
|
| 79 |
+
(APIServer pid=1147687) INFO 08-10 21:22:37 [launcher.py:46] Route: /is_paused, Methods: GET
|
| 80 |
+
(APIServer pid=1147687) INFO 08-10 21:22:37 [launcher.py:46] Route: /init_weight_transfer_engine, Methods: POST
|
| 81 |
+
(APIServer pid=1147687) INFO 08-10 21:22:37 [launcher.py:46] Route: /update_weights, Methods: POST
|
| 82 |
+
(APIServer pid=1147687) INFO 08-10 21:22:37 [launcher.py:46] Route: /get_world_size, Methods: GET
|
| 83 |
+
(APIServer pid=1147687) INFO 08-10 21:22:37 [launcher.py:46] Route: /scale_elastic_ep, Methods: POST
|
| 84 |
+
(APIServer pid=1147687) INFO 08-10 21:22:37 [launcher.py:46] Route: /is_scaling_elastic_ep, Methods: POST
|
| 85 |
+
(APIServer pid=1147687) INFO 08-10 21:22:37 [launcher.py:46] Route: /v1/chat/completions/render, Methods: POST
|
| 86 |
+
(APIServer pid=1147687) INFO 08-10 21:22:37 [launcher.py:46] Route: /v1/completions/render, Methods: POST
|
| 87 |
+
(APIServer pid=1147687) INFO: Started server process [1147687]
|
| 88 |
+
(APIServer pid=1147687) INFO: Waiting for application startup.
|
| 89 |
+
(APIServer pid=1147687) INFO: Application startup complete.
|
| 90 |
+
(APIServer pid=1147687) INFO: 127.0.0.1:36040 - "GET /health HTTP/1.1" 200 OK
|
| 91 |
+
(APIServer pid=1147687) INFO 08-10 21:22:48 [loggers.py:259] Engine 000: Avg prompt throughput: 3508.3 tokens/s, Avg generation throughput: 2603.9 tokens/s, Running: 254 reqs, Waiting: 860 reqs, GPU KV cache usage: 35.0%, Prefix cache hit rate: 93.2%
|
| 92 |
+
(APIServer pid=1147687) INFO 08-10 21:22:58 [loggers.py:259] Engine 000: Avg prompt throughput: 2530.2 tokens/s, Avg generation throughput: 3217.5 tokens/s, Running: 255 reqs, Waiting: 532 reqs, GPU KV cache usage: 35.7%, Prefix cache hit rate: 93.3%
|
| 93 |
+
(APIServer pid=1147687) INFO 08-10 21:23:08 [loggers.py:259] Engine 000: Avg prompt throughput: 2784.6 tokens/s, Avg generation throughput: 3164.9 tokens/s, Running: 254 reqs, Waiting: 189 reqs, GPU KV cache usage: 36.6%, Prefix cache hit rate: 93.4%
|
| 94 |
+
(APIServer pid=1147687) INFO 08-10 21:23:18 [loggers.py:259] Engine 000: Avg prompt throughput: 1531.1 tokens/s, Avg generation throughput: 2918.6 tokens/s, Running: 102 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.7%, Prefix cache hit rate: 93.4%
|
| 95 |
+
(APIServer pid=1147687) INFO: 127.0.0.1:43906 - "POST /v1/completions HTTP/1.1" 200 OK
|
| 96 |
+
(EngineCore pid=1148946) INFO 08-10 21:23:22 [core.py:1210] Shutdown initiated (timeout=0)
|
| 97 |
+
(EngineCore pid=1148946) INFO 08-10 21:23:22 [core.py:1233] Shutdown complete
|
| 98 |
+
(APIServer pid=1147687) INFO: Shutting down
|
| 99 |
+
(APIServer pid=1147687) INFO: Waiting for application shutdown.
|
| 100 |
+
(APIServer pid=1147687) INFO: Application shutdown complete.
|
| 101 |
+
(APIServer pid=1147687) INFO: Finished server process [1147687]
|
evals/grid_math_unhealed/uniform_keep75_unhealed.log
ADDED
|
@@ -0,0 +1,8 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"correct": 816,
|
| 3 |
+
"accuracy": 0.6186504927975739,
|
| 4 |
+
"finished": 1316,
|
| 5 |
+
"finish_rate": 0.9977255496588324,
|
| 6 |
+
"mean_completion_tokens": 107.12736921910539
|
| 7 |
+
}
|
| 8 |
+
saved item-level results -> outputs/evals/grid_math_unhealed/uniform_keep75_unhealed_chat.json
|
evals/grid_math_unhealed/uniform_keep75_unhealed_chat.json
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
evals/grid_math_unhealed/uniform_keep75_unhealed_chat.json.server.log
ADDED
|
@@ -0,0 +1,103 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
Skipping import of cpp extensions due to incompatible torch version 2.10.0+cu128 for torchao version 0.15.0 Please see https://github.com/pytorch/ao/issues/2919 for more info
|
| 2 |
+
WARNING 08-10 21:20:11 [registry.py:915] Model architecture OlmoeForCausalLM is already registered, and will be overwritten by the new model class glean_vllm.pruned_olmoe:PrunedOlmoeForCausalLM.
|
| 3 |
+
(APIServer pid=1142904) INFO 08-10 21:20:11 [utils.py:299]
|
| 4 |
+
(APIServer pid=1142904) INFO 08-10 21:20:11 [utils.py:299] β β ββ ββ
|
| 5 |
+
(APIServer pid=1142904) INFO 08-10 21:20:11 [utils.py:299] ββ ββ β β β βββ β version 0.19.0
|
| 6 |
+
(APIServer pid=1142904) INFO 08-10 21:20:11 [utils.py:299] ββββ β β β β model /media/henry/MoreFiles/winnow_release/math_arms/uniform_keep75
|
| 7 |
+
(APIServer pid=1142904) INFO 08-10 21:20:11 [utils.py:299] ββ βββββ βββββ β β
|
| 8 |
+
(APIServer pid=1142904) INFO 08-10 21:20:11 [utils.py:299]
|
| 9 |
+
(APIServer pid=1142904) INFO 08-10 21:20:11 [utils.py:233] non-default args: {'model_tag': '/media/henry/MoreFiles/winnow_release/math_arms/uniform_keep75', 'host': '127.0.0.1', 'port': 8612, 'model': '/media/henry/MoreFiles/winnow_release/math_arms/uniform_keep75', 'max_model_len': 2048, 'enforce_eager': True, 'served_model_name': ['student'], 'gpu_memory_utilization': 0.85}
|
| 10 |
+
(APIServer pid=1142904) INFO 08-10 21:20:21 [model.py:549] Resolved architecture: PrunedOlmoeForCausalLM
|
| 11 |
+
(APIServer pid=1142904) INFO 08-10 21:20:21 [model.py:1678] Using max model len 2048
|
| 12 |
+
(APIServer pid=1142904) INFO 08-10 21:20:22 [vllm.py:790] Asynchronous scheduling is enabled.
|
| 13 |
+
(APIServer pid=1142904) WARNING 08-10 21:20:22 [vllm.py:848] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
|
| 14 |
+
(APIServer pid=1142904) WARNING 08-10 21:20:22 [vllm.py:859] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
|
| 15 |
+
(APIServer pid=1142904) INFO 08-10 21:20:22 [vllm.py:1025] Cudagraph is disabled under eager mode
|
| 16 |
+
(APIServer pid=1142904) INFO 08-10 21:20:22 [compilation.py:290] Enabled custom fusions: norm_quant, act_quant
|
| 17 |
+
Skipping import of cpp extensions due to incompatible torch version 2.10.0+cu128 for torchao version 0.15.0 Please see https://github.com/pytorch/ao/issues/2919 for more info
|
| 18 |
+
(EngineCore pid=1144302) WARNING 08-10 21:20:30 [registry.py:915] Model architecture OlmoeForCausalLM is already registered, and will be overwritten by the new model class glean_vllm.pruned_olmoe:PrunedOlmoeForCausalLM.
|
| 19 |
+
(EngineCore pid=1144302) INFO 08-10 21:20:30 [core.py:105] Initializing a V1 LLM engine (v0.19.0) with config: model='/media/henry/MoreFiles/winnow_release/math_arms/uniform_keep75', speculative_config=None, tokenizer='/media/henry/MoreFiles/winnow_release/math_arms/uniform_keep75', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=False, dtype=torch.bfloat16, max_seq_len=2048, download_dir=None, load_format=auto, tensor_parallel_size=1, pipeline_parallel_size=1, data_parallel_size=1, decode_context_parallel_size=1, dcp_comm_backend=ag_rs, disable_custom_all_reduce=False, quantization=None, enforce_eager=True, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False), seed=0, served_model_name=student, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'mode': <CompilationMode.NONE: 0>, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'splitting_ops': [], 'compile_mm_encoder': False, 'cudagraph_mm_encoder': False, 'encoder_cudagraph_token_budgets': [], 'encoder_cudagraph_max_images_per_batch': 0, 'compile_sizes': [], 'compile_ranges_endpoints': [2048], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'size_asserts': False, 'alignment_asserts': False, 'scalar_asserts': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': <CUDAGraphMode.NONE: 0>, 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': True, 'fuse_act_quant': True, 'fuse_attn_quant': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False}, 'max_cudagraph_capture_size': 0, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': True, 'static_all_moe_layers': []}
|
| 20 |
+
(EngineCore pid=1144302) INFO 08-10 21:20:31 [parallel_state.py:1400] world_size=1 rank=0 local_rank=0 distributed_init_method=tcp://192.168.0.27:57179 backend=nccl
|
| 21 |
+
(EngineCore pid=1144302) INFO 08-10 21:20:31 [parallel_state.py:1716] rank 0 in world size 1 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank 0, EPLB rank N/A
|
| 22 |
+
(EngineCore pid=1144302) INFO 08-10 21:20:32 [gpu_model_runner.py:4735] Starting to load model /media/henry/MoreFiles/winnow_release/math_arms/uniform_keep75...
|
| 23 |
+
(EngineCore pid=1144302) INFO 08-10 21:20:33 [cuda.py:334] Using FLASH_ATTN attention backend out of potential backends: ['FLASH_ATTN', 'FLASHINFER', 'TRITON_ATTN', 'FLEX_ATTENTION'].
|
| 24 |
+
(EngineCore pid=1144302) INFO 08-10 21:20:33 [flash_attn.py:596] Using FlashAttention version 2
|
| 25 |
+
(EngineCore pid=1144302)
|
| 26 |
+
(EngineCore pid=1144302)
|
| 27 |
+
(EngineCore pid=1144302)
|
| 28 |
+
(EngineCore pid=1144302)
|
| 29 |
+
(EngineCore pid=1144302)
|
| 30 |
+
(EngineCore pid=1144302)
|
| 31 |
+
(EngineCore pid=1144302) INFO 08-10 21:20:35 [default_loader.py:384] Loading weights took 2.05 seconds
|
| 32 |
+
(EngineCore pid=1144302) INFO 08-10 21:20:36 [gpu_model_runner.py:4820] Model loading took 9.89 GiB memory and 3.191814 seconds
|
| 33 |
+
(EngineCore pid=1144302) INFO 08-10 21:20:38 [gpu_worker.py:436] Available KV cache memory: 9.86 GiB
|
| 34 |
+
(EngineCore pid=1144302) INFO 08-10 21:20:38 [kv_cache_utils.py:1319] GPU KV cache size: 80,736 tokens
|
| 35 |
+
(EngineCore pid=1144302) INFO 08-10 21:20:38 [kv_cache_utils.py:1324] Maximum concurrency for 2,048 tokens per request: 39.42x
|
| 36 |
+
(EngineCore pid=1144302) INFO 08-10 21:20:38 [core.py:283] init engine (profile, create kv cache, warmup model) took 2.64 seconds
|
| 37 |
+
(EngineCore pid=1144302) INFO 08-10 21:20:39 [vllm.py:790] Asynchronous scheduling is enabled.
|
| 38 |
+
(EngineCore pid=1144302) WARNING 08-10 21:20:39 [vllm.py:848] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
|
| 39 |
+
(EngineCore pid=1144302) WARNING 08-10 21:20:39 [vllm.py:859] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
|
| 40 |
+
(EngineCore pid=1144302) INFO 08-10 21:20:39 [vllm.py:1025] Cudagraph is disabled under eager mode
|
| 41 |
+
(EngineCore pid=1144302) INFO 08-10 21:20:39 [compilation.py:290] Enabled custom fusions: norm_quant, act_quant
|
| 42 |
+
(APIServer pid=1142904) INFO 08-10 21:20:39 [api_server.py:590] Supported tasks: ['generate']
|
| 43 |
+
(APIServer pid=1142904) WARNING 08-10 21:20:39 [__init__.py:14] SECURITY WARNING: Development endpoints are enabled! This should NOT be used in production!
|
| 44 |
+
(APIServer pid=1142904) INFO 08-10 21:20:39 [hf.py:314] Detected the chat template content format to be 'string'. You can set `--chat-template-content-format` to override this.
|
| 45 |
+
(APIServer pid=1142904) INFO 08-10 21:20:39 [api_server.py:594] Starting vLLM server on http://127.0.0.1:8612
|
| 46 |
+
(APIServer pid=1142904) INFO 08-10 21:20:39 [launcher.py:37] Available routes are:
|
| 47 |
+
(APIServer pid=1142904) INFO 08-10 21:20:39 [launcher.py:46] Route: /openapi.json, Methods: GET, HEAD
|
| 48 |
+
(APIServer pid=1142904) INFO 08-10 21:20:39 [launcher.py:46] Route: /docs, Methods: GET, HEAD
|
| 49 |
+
(APIServer pid=1142904) INFO 08-10 21:20:39 [launcher.py:46] Route: /docs/oauth2-redirect, Methods: GET, HEAD
|
| 50 |
+
(APIServer pid=1142904) INFO 08-10 21:20:39 [launcher.py:46] Route: /redoc, Methods: GET, HEAD
|
| 51 |
+
(APIServer pid=1142904) INFO 08-10 21:20:39 [launcher.py:46] Route: /sleep, Methods: POST
|
| 52 |
+
(APIServer pid=1142904) INFO 08-10 21:20:39 [launcher.py:46] Route: /wake_up, Methods: POST
|
| 53 |
+
(APIServer pid=1142904) INFO 08-10 21:20:39 [launcher.py:46] Route: /is_sleeping, Methods: GET
|
| 54 |
+
(APIServer pid=1142904) INFO 08-10 21:20:39 [launcher.py:46] Route: /collective_rpc, Methods: POST
|
| 55 |
+
(APIServer pid=1142904) INFO 08-10 21:20:39 [launcher.py:46] Route: /reset_prefix_cache, Methods: POST
|
| 56 |
+
(APIServer pid=1142904) INFO 08-10 21:20:39 [launcher.py:46] Route: /reset_mm_cache, Methods: POST
|
| 57 |
+
(APIServer pid=1142904) INFO 08-10 21:20:39 [launcher.py:46] Route: /reset_encoder_cache, Methods: POST
|
| 58 |
+
(APIServer pid=1142904) INFO 08-10 21:20:39 [launcher.py:46] Route: /tokenize, Methods: POST
|
| 59 |
+
(APIServer pid=1142904) INFO 08-10 21:20:39 [launcher.py:46] Route: /detokenize, Methods: POST
|
| 60 |
+
(APIServer pid=1142904) INFO 08-10 21:20:39 [launcher.py:46] Route: /load, Methods: GET
|
| 61 |
+
(APIServer pid=1142904) INFO 08-10 21:20:39 [launcher.py:46] Route: /version, Methods: GET
|
| 62 |
+
(APIServer pid=1142904) INFO 08-10 21:20:39 [launcher.py:46] Route: /health, Methods: GET
|
| 63 |
+
(APIServer pid=1142904) INFO 08-10 21:20:39 [launcher.py:46] Route: /metrics, Methods: GET
|
| 64 |
+
(APIServer pid=1142904) INFO 08-10 21:20:39 [launcher.py:46] Route: /server_info, Methods: GET
|
| 65 |
+
(APIServer pid=1142904) INFO 08-10 21:20:39 [launcher.py:46] Route: /v1/models, Methods: GET
|
| 66 |
+
(APIServer pid=1142904) INFO 08-10 21:20:39 [launcher.py:46] Route: /ping, Methods: GET
|
| 67 |
+
(APIServer pid=1142904) INFO 08-10 21:20:39 [launcher.py:46] Route: /ping, Methods: POST
|
| 68 |
+
(APIServer pid=1142904) INFO 08-10 21:20:39 [launcher.py:46] Route: /invocations, Methods: POST
|
| 69 |
+
(APIServer pid=1142904) INFO 08-10 21:20:39 [launcher.py:46] Route: /v1/chat/completions, Methods: POST
|
| 70 |
+
(APIServer pid=1142904) INFO 08-10 21:20:39 [launcher.py:46] Route: /v1/chat/completions/batch, Methods: POST
|
| 71 |
+
(APIServer pid=1142904) INFO 08-10 21:20:39 [launcher.py:46] Route: /v1/responses, Methods: POST
|
| 72 |
+
(APIServer pid=1142904) INFO 08-10 21:20:39 [launcher.py:46] Route: /v1/responses/{response_id}, Methods: GET
|
| 73 |
+
(APIServer pid=1142904) INFO 08-10 21:20:39 [launcher.py:46] Route: /v1/responses/{response_id}/cancel, Methods: POST
|
| 74 |
+
(APIServer pid=1142904) INFO 08-10 21:20:39 [launcher.py:46] Route: /v1/completions, Methods: POST
|
| 75 |
+
(APIServer pid=1142904) INFO 08-10 21:20:39 [launcher.py:46] Route: /v1/messages, Methods: POST
|
| 76 |
+
(APIServer pid=1142904) INFO 08-10 21:20:39 [launcher.py:46] Route: /v1/messages/count_tokens, Methods: POST
|
| 77 |
+
(APIServer pid=1142904) INFO 08-10 21:20:39 [launcher.py:46] Route: /inference/v1/generate, Methods: POST
|
| 78 |
+
(APIServer pid=1142904) INFO 08-10 21:20:39 [launcher.py:46] Route: /pause, Methods: POST
|
| 79 |
+
(APIServer pid=1142904) INFO 08-10 21:20:39 [launcher.py:46] Route: /resume, Methods: POST
|
| 80 |
+
(APIServer pid=1142904) INFO 08-10 21:20:39 [launcher.py:46] Route: /is_paused, Methods: GET
|
| 81 |
+
(APIServer pid=1142904) INFO 08-10 21:20:39 [launcher.py:46] Route: /init_weight_transfer_engine, Methods: POST
|
| 82 |
+
(APIServer pid=1142904) INFO 08-10 21:20:39 [launcher.py:46] Route: /update_weights, Methods: POST
|
| 83 |
+
(APIServer pid=1142904) INFO 08-10 21:20:39 [launcher.py:46] Route: /get_world_size, Methods: GET
|
| 84 |
+
(APIServer pid=1142904) INFO 08-10 21:20:39 [launcher.py:46] Route: /scale_elastic_ep, Methods: POST
|
| 85 |
+
(APIServer pid=1142904) INFO 08-10 21:20:39 [launcher.py:46] Route: /is_scaling_elastic_ep, Methods: POST
|
| 86 |
+
(APIServer pid=1142904) INFO 08-10 21:20:39 [launcher.py:46] Route: /v1/chat/completions/render, Methods: POST
|
| 87 |
+
(APIServer pid=1142904) INFO 08-10 21:20:39 [launcher.py:46] Route: /v1/completions/render, Methods: POST
|
| 88 |
+
(APIServer pid=1142904) INFO: Started server process [1142904]
|
| 89 |
+
(APIServer pid=1142904) INFO: Waiting for application startup.
|
| 90 |
+
(APIServer pid=1142904) INFO: Application startup complete.
|
| 91 |
+
(APIServer pid=1142904) INFO: 127.0.0.1:46028 - "GET /health HTTP/1.1" 200 OK
|
| 92 |
+
(APIServer pid=1142904) INFO 08-10 21:20:49 [loggers.py:259] Engine 000: Avg prompt throughput: 2616.4 tokens/s, Avg generation throughput: 1962.2 tokens/s, Running: 251 reqs, Waiting: 976 reqs, GPU KV cache usage: 48.0%, Prefix cache hit rate: 93.2%
|
| 93 |
+
(APIServer pid=1142904) INFO 08-10 21:20:59 [loggers.py:259] Engine 000: Avg prompt throughput: 2274.7 tokens/s, Avg generation throughput: 2965.4 tokens/s, Running: 252 reqs, Waiting: 686 reqs, GPU KV cache usage: 47.8%, Prefix cache hit rate: 93.3%
|
| 94 |
+
(APIServer pid=1142904) INFO 08-10 21:21:09 [loggers.py:259] Engine 000: Avg prompt throughput: 2289.4 tokens/s, Avg generation throughput: 2965.3 tokens/s, Running: 255 reqs, Waiting: 396 reqs, GPU KV cache usage: 47.5%, Prefix cache hit rate: 93.4%
|
| 95 |
+
(APIServer pid=1142904) INFO 08-10 21:21:19 [loggers.py:259] Engine 000: Avg prompt throughput: 2179.1 tokens/s, Avg generation throughput: 2993.4 tokens/s, Running: 255 reqs, Waiting: 125 reqs, GPU KV cache usage: 50.0%, Prefix cache hit rate: 93.4%
|
| 96 |
+
(APIServer pid=1142904) INFO 08-10 21:21:29 [loggers.py:259] Engine 000: Avg prompt throughput: 1048.0 tokens/s, Avg generation throughput: 2675.2 tokens/s, Running: 95 reqs, Waiting: 0 reqs, GPU KV cache usage: 28.3%, Prefix cache hit rate: 93.4%
|
| 97 |
+
(APIServer pid=1142904) INFO: 127.0.0.1:46038 - "POST /v1/completions HTTP/1.1" 200 OK
|
| 98 |
+
(EngineCore pid=1144302) INFO 08-10 21:21:36 [core.py:1210] Shutdown initiated (timeout=0)
|
| 99 |
+
(EngineCore pid=1144302) INFO 08-10 21:21:36 [core.py:1233] Shutdown complete
|
| 100 |
+
(APIServer pid=1142904) INFO: Shutting down
|
| 101 |
+
(APIServer pid=1142904) INFO: Waiting for application shutdown.
|
| 102 |
+
(APIServer pid=1142904) INFO: Application shutdown complete.
|
| 103 |
+
(APIServer pid=1142904) INFO: Finished server process [1142904]
|
evals/healing_breadth/glean_math_keep25_seed1224.json
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
evals/healing_breadth/glean_math_keep25_seed1224.json.server.log
ADDED
|
@@ -0,0 +1,139 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
(APIServer pid=199309) INFO 07-14 21:55:42 [api_utils.py:339]
|
| 2 |
+
(APIServer pid=199309) INFO 07-14 21:55:42 [api_utils.py:339] β β ββ ββ
|
| 3 |
+
(APIServer pid=199309) INFO 07-14 21:55:42 [api_utils.py:339] ββ ββ β β β βββ β version 0.25.0
|
| 4 |
+
(APIServer pid=199309) INFO 07-14 21:55:42 [api_utils.py:339] ββββ β β β β model outputs/healed/healing_breadth/glean_math_keep25_seed1224/step0050
|
| 5 |
+
(APIServer pid=199309) INFO 07-14 21:55:42 [api_utils.py:339] ββ βββββ βββββ β β
|
| 6 |
+
(APIServer pid=199309) INFO 07-14 21:55:42 [api_utils.py:339]
|
| 7 |
+
(APIServer pid=199309) INFO 07-14 21:55:42 [api_utils.py:273] non-default args: {'model_tag': 'outputs/healed/healing_breadth/glean_math_keep25_seed1224/step0050', 'host': '127.0.0.1', 'port': 8377, 'model': 'outputs/healed/healing_breadth/glean_math_keep25_seed1224/step0050', 'max_model_len': 2048, 'enforce_eager': True, 'served_model_name': ['student'], 'gpu_memory_utilization': 0.85}
|
| 8 |
+
(APIServer pid=199309) INFO 07-14 21:55:42 [model.py:619] Resolved architecture: PrunedOlmoeForCausalLM
|
| 9 |
+
(APIServer pid=199309) INFO 07-14 21:55:42 [model.py:1776] Using max model len 2048
|
| 10 |
+
(APIServer pid=199309) INFO 07-14 21:55:42 [vllm.py:1042] Asynchronous scheduling is enabled.
|
| 11 |
+
(APIServer pid=199309) WARNING 07-14 21:55:42 [vllm.py:1096] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
|
| 12 |
+
(APIServer pid=199309) WARNING 07-14 21:55:42 [vllm.py:1144] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
|
| 13 |
+
(APIServer pid=199309) INFO 07-14 21:55:42 [kernel.py:292] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native'])
|
| 14 |
+
(APIServer pid=199309) INFO 07-14 21:55:42 [vllm.py:1322] Cudagraph is disabled under eager mode
|
| 15 |
+
(APIServer pid=199309) INFO 07-14 21:55:42 [compilation.py:312] Enabled custom fusions: norm_quant, act_quant
|
| 16 |
+
(EngineCore pid=199426) INFO 07-14 21:55:49 [core.py:114] Initializing a V1 LLM engine (v0.25.0) with config: model='outputs/healed/healing_breadth/glean_math_keep25_seed1224/step0050', speculative_config=None, tokenizer='outputs/healed/healing_breadth/glean_math_keep25_seed1224/step0050', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=False, dtype=torch.bfloat16, max_seq_len=2048, download_dir=None, load_format=auto, tensor_parallel_size=1, pipeline_parallel_size=1, data_parallel_size=1, decode_context_parallel_size=1, dcp_comm_backend=ag_rs, disable_custom_all_reduce=False, quantization=None, quantization_config=None, enforce_eager=True, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False, jit_monitor_mode='warn', jit_monitor_verbose=False), seed=0, served_model_name=student, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'mode': <CompilationMode.NONE: 0>, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'ir_enable_torch_wrap': False, 'splitting_ops': [], 'compile_mm_encoder': False, 'cudagraph_mm_encoder': False, 'encoder_cudagraph_token_budgets': [], 'encoder_cudagraph_max_vision_items_per_batch': 0, 'encoder_cudagraph_max_frames_per_batch': None, 'compile_sizes': [], 'compile_ranges_endpoints': [2048], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'size_asserts': False, 'alignment_asserts': False, 'scalar_asserts': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': <CUDAGraphMode.NONE: 0>, 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': True, 'fuse_act_quant': True, 'fuse_attn_quant': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False, 'fuse_rope_kvcache_cat_mla': False, 'fuse_act_padding': False}, 'max_cudagraph_capture_size': 0, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': False, 'static_all_moe_layers': []}, kernel_config=KernelConfig(ir_op_priority=IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native']), enable_flashinfer_autotune=True, enable_cutedsl_warmup=True, moe_backend='auto', linear_backend='auto')
|
| 17 |
+
(EngineCore pid=199426) INFO 07-14 21:55:50 [parallel_state.py:1607] world_size=1 rank=0 local_rank=0 distributed_init_method=tcp://192.168.0.15:42095 backend=nccl
|
| 18 |
+
(EngineCore pid=199426) INFO 07-14 21:55:50 [parallel_state.py:1942] rank 0 in world size 1 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank 0, EPLB rank N/A
|
| 19 |
+
(EngineCore pid=199426) INFO 07-14 21:55:51 [topk_topp_sampler.py:55] Using FlashInfer for top-p & top-k sampling.
|
| 20 |
+
(EngineCore pid=199426) INFO 07-14 21:55:51 [gpu_model_runner.py:5209] Starting to load model outputs/healed/healing_breadth/glean_math_keep25_seed1224/step0050...
|
| 21 |
+
(EngineCore pid=199426) INFO 07-14 21:55:51 [cuda.py:476] Using FLASH_ATTN attention backend out of potential backends: ['FLASH_ATTN', 'FLASHINFER', 'TRITON_ATTN', 'FLEX_ATTENTION'].
|
| 22 |
+
(EngineCore pid=199426) INFO 07-14 21:55:51 [flash_attn.py:718] Using FlashAttention version 2
|
| 23 |
+
(EngineCore pid=199426) /home/henry/.cache/glean/megablocks-variable-93a1479bc15b/megablocks/grouped_gemm_util.py:10: UserWarning: Grouped GEMM not available.
|
| 24 |
+
(EngineCore pid=199426) warnings.warn('Grouped GEMM not available.')
|
| 25 |
+
(EngineCore pid=199426) INFO 07-14 21:55:51 [weight_utils.py:849] Filesystem type for checkpoints: EXT4. Checkpoint size: 3.89 GiB. Available RAM: 77.80 GiB.
|
| 26 |
+
(EngineCore pid=199426) INFO 07-14 21:55:51 [weight_utils.py:872] Auto-prefetch is disabled because the filesystem (EXT4) is not a recognized network FS (NFS/Lustre). If you want to force prefetching, start vLLM with --safetensors-load-strategy=prefetch.
|
| 27 |
+
(EngineCore pid=199426)
|
| 28 |
+
(EngineCore pid=199426)
|
| 29 |
+
(EngineCore pid=199426)
|
| 30 |
+
(EngineCore pid=199426)
|
| 31 |
+
(EngineCore pid=199426) INFO 07-14 21:55:53 [default_loader.py:430] Loading weights took 2.44 seconds
|
| 32 |
+
(EngineCore pid=199426) INFO 07-14 21:55:54 [gpu_model_runner.py:5306] Model loading took 3.89 GiB memory and 2.620414 seconds
|
| 33 |
+
(EngineCore pid=199426) INFO 07-14 21:55:56 [gpu_worker.py:538] Available KV cache memory: 15.82 GiB
|
| 34 |
+
(EngineCore pid=199426) INFO 07-14 21:55:56 [kv_cache_utils.py:2146] GPU KV cache size: 129,584 tokens
|
| 35 |
+
(EngineCore pid=199426) INFO 07-14 21:55:56 [kv_cache_utils.py:2147] Maximum concurrency for 2,048 tokens per request: 63.27x
|
| 36 |
+
(EngineCore pid=199426) INFO 07-14 21:55:56 [cutedsl_warmup.py:97] Skipping CuTeDSL warmup because no compile units were requested.
|
| 37 |
+
(EngineCore pid=199426) INFO 07-14 21:55:56 [jit_monitor.py:73] Kernel JIT monitor activated; monitored JIT compilations during inference will use mode=warn.
|
| 38 |
+
(EngineCore pid=199426) INFO 07-14 21:55:56 [core.py:344] init engine (profile, create kv cache, warmup model) took 2.05 s
|
| 39 |
+
(EngineCore pid=199426) INFO 07-14 21:55:56 [vllm.py:1042] Asynchronous scheduling is enabled.
|
| 40 |
+
(EngineCore pid=199426) WARNING 07-14 21:55:56 [vllm.py:1096] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
|
| 41 |
+
(EngineCore pid=199426) WARNING 07-14 21:55:56 [vllm.py:1144] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
|
| 42 |
+
(EngineCore pid=199426) INFO 07-14 21:55:56 [kernel.py:292] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native'])
|
| 43 |
+
(EngineCore pid=199426) INFO 07-14 21:55:56 [vllm.py:1322] Cudagraph is disabled under eager mode
|
| 44 |
+
(EngineCore pid=199426) INFO 07-14 21:55:56 [compilation.py:312] Enabled custom fusions: norm_quant, act_quant
|
| 45 |
+
(APIServer pid=199309) INFO 07-14 21:55:56 [api_server.py:612] Supported tasks: ['generate']
|
| 46 |
+
(APIServer pid=199309) WARNING 07-14 21:55:56 [__init__.py:36] SECURITY WARNING: Development endpoints are enabled! This should NOT be used in production!
|
| 47 |
+
(APIServer pid=199309) INFO 07-14 21:55:56 [hf.py:548] Detected the chat template content format to be 'string'. You can set `--chat-template-content-format` to override this.
|
| 48 |
+
(APIServer pid=199309) INFO 07-14 21:55:56 [api_server.py:616] Starting vLLM server on http://127.0.0.1:8377
|
| 49 |
+
(APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:37] Available routes are:
|
| 50 |
+
(APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /openapi.json, Methods: GET, HEAD
|
| 51 |
+
(APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /docs, Methods: GET, HEAD
|
| 52 |
+
(APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /docs/oauth2-redirect, Methods: GET, HEAD
|
| 53 |
+
(APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /redoc, Methods: GET, HEAD
|
| 54 |
+
(APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /load, Methods: GET
|
| 55 |
+
(APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /version, Methods: GET
|
| 56 |
+
(APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /health, Methods: GET
|
| 57 |
+
(APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /metrics, Methods: GET
|
| 58 |
+
(APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /tokenize, Methods: POST
|
| 59 |
+
(APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /detokenize, Methods: POST
|
| 60 |
+
(APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /v1/models, Methods: GET
|
| 61 |
+
(APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /ping, Methods: GET
|
| 62 |
+
(APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /ping, Methods: POST
|
| 63 |
+
(APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /invocations, Methods: POST
|
| 64 |
+
(APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /reset_prefix_cache, Methods: POST
|
| 65 |
+
(APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /reset_mm_cache, Methods: POST
|
| 66 |
+
(APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /reset_encoder_cache, Methods: POST
|
| 67 |
+
(APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /pause, Methods: POST
|
| 68 |
+
(APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /resume, Methods: POST
|
| 69 |
+
(APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /is_paused, Methods: GET
|
| 70 |
+
(APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /init_weight_transfer_engine, Methods: POST
|
| 71 |
+
(APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /start_weight_update, Methods: POST
|
| 72 |
+
(APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /update_weights, Methods: POST
|
| 73 |
+
(APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /finish_weight_update, Methods: POST
|
| 74 |
+
(APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /get_world_size, Methods: GET
|
| 75 |
+
(APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /collective_rpc, Methods: POST
|
| 76 |
+
(APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /server_info, Methods: GET
|
| 77 |
+
(APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /sleep, Methods: POST
|
| 78 |
+
(APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /wake_up, Methods: POST
|
| 79 |
+
(APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /is_sleeping, Methods: GET
|
| 80 |
+
(APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /v1/chat/completions, Methods: POST
|
| 81 |
+
(APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /v1/chat/completions/batch, Methods: POST
|
| 82 |
+
(APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /v1/responses, Methods: POST
|
| 83 |
+
(APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /v1/responses/{response_id}, Methods: GET
|
| 84 |
+
(APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /v1/responses/{response_id}/cancel, Methods: POST
|
| 85 |
+
(APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /v1/completions, Methods: POST
|
| 86 |
+
(APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /v1/messages, Methods: POST
|
| 87 |
+
(APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /v1/messages/count_tokens, Methods: POST
|
| 88 |
+
(APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /generative_scoring, Methods: POST
|
| 89 |
+
(APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /scale_elastic_ep, Methods: POST
|
| 90 |
+
(APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /is_scaling_elastic_ep, Methods: POST
|
| 91 |
+
(APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /v1/chat/completions/render, Methods: POST
|
| 92 |
+
(APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /v1/completions/render, Methods: POST
|
| 93 |
+
(APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /v1/chat/completions/derender, Methods: POST
|
| 94 |
+
(APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /v1/completions/derender, Methods: POST
|
| 95 |
+
(APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /inference/v1/generate, Methods: POST
|
| 96 |
+
(APIServer pid=199309) INFO: Started server process [199309]
|
| 97 |
+
(APIServer pid=199309) INFO: Waiting for application startup.
|
| 98 |
+
(APIServer pid=199309) INFO: Application startup complete.
|
| 99 |
+
(APIServer pid=199309) INFO: 127.0.0.1:60678 - "GET /health HTTP/1.1" 200 OK
|
| 100 |
+
(EngineCore pid=199426) WARNING 07-14 21:55:58 [jit_monitor.py:129] Triton kernel JIT compilation during inference: _build_route_rows. This causes a latency spike; consider extending warmup to cover this shape/config.
|
| 101 |
+
(APIServer pid=199309) INFO 07-14 21:56:07 [loggers.py:273] Engine 000: Avg prompt throughput: 2049.1 tokens/s, Avg generation throughput: 2315.6 tokens/s, Running: 256 reqs, Waiting: 1059 reqs, GPU KV cache usage: 36.6%, Prefix cache hit rate: 93.1%
|
| 102 |
+
(APIServer pid=199309) INFO 07-14 21:56:17 [loggers.py:273] Engine 000: Avg prompt throughput: 281.0 tokens/s, Avg generation throughput: 2862.8 tokens/s, Running: 256 reqs, Waiting: 1022 reqs, GPU KV cache usage: 54.2%, Prefix cache hit rate: 93.1%
|
| 103 |
+
(APIServer pid=199309) INFO 07-14 21:56:27 [loggers.py:273] Engine 000: Avg prompt throughput: 467.0 tokens/s, Avg generation throughput: 2732.6 tokens/s, Running: 256 reqs, Waiting: 961 reqs, GPU KV cache usage: 63.6%, Prefix cache hit rate: 93.2%
|
| 104 |
+
(APIServer pid=199309) INFO 07-14 21:56:37 [loggers.py:273] Engine 000: Avg prompt throughput: 426.2 tokens/s, Avg generation throughput: 2733.7 tokens/s, Running: 255 reqs, Waiting: 908 reqs, GPU KV cache usage: 72.4%, Prefix cache hit rate: 93.2%
|
| 105 |
+
(APIServer pid=199309) INFO 07-14 21:56:47 [loggers.py:273] Engine 000: Avg prompt throughput: 1128.2 tokens/s, Avg generation throughput: 2494.4 tokens/s, Running: 256 reqs, Waiting: 764 reqs, GPU KV cache usage: 40.0%, Prefix cache hit rate: 93.3%
|
| 106 |
+
(APIServer pid=199309) INFO 07-14 21:56:57 [loggers.py:273] Engine 000: Avg prompt throughput: 319.6 tokens/s, Avg generation throughput: 2888.4 tokens/s, Running: 256 reqs, Waiting: 722 reqs, GPU KV cache usage: 53.6%, Prefix cache hit rate: 93.3%
|
| 107 |
+
(APIServer pid=199309) INFO 07-14 21:57:07 [loggers.py:273] Engine 000: Avg prompt throughput: 530.8 tokens/s, Avg generation throughput: 2834.6 tokens/s, Running: 255 reqs, Waiting: 654 reqs, GPU KV cache usage: 59.5%, Prefix cache hit rate: 93.3%
|
| 108 |
+
(APIServer pid=199309) INFO 07-14 21:57:17 [loggers.py:273] Engine 000: Avg prompt throughput: 643.1 tokens/s, Avg generation throughput: 2781.9 tokens/s, Running: 255 reqs, Waiting: 572 reqs, GPU KV cache usage: 58.6%, Prefix cache hit rate: 93.3%
|
| 109 |
+
(APIServer pid=199309) INFO 07-14 21:57:27 [loggers.py:273] Engine 000: Avg prompt throughput: 604.2 tokens/s, Avg generation throughput: 2731.0 tokens/s, Running: 256 reqs, Waiting: 498 reqs, GPU KV cache usage: 62.2%, Prefix cache hit rate: 93.3%
|
| 110 |
+
(APIServer pid=199309) INFO 07-14 21:57:37 [loggers.py:273] Engine 000: Avg prompt throughput: 780.9 tokens/s, Avg generation throughput: 2728.8 tokens/s, Running: 256 reqs, Waiting: 395 reqs, GPU KV cache usage: 50.9%, Prefix cache hit rate: 93.4%
|
| 111 |
+
(APIServer pid=199309) INFO 07-14 21:57:47 [loggers.py:273] Engine 000: Avg prompt throughput: 467.6 tokens/s, Avg generation throughput: 2835.7 tokens/s, Running: 255 reqs, Waiting: 339 reqs, GPU KV cache usage: 60.1%, Prefix cache hit rate: 93.3%
|
| 112 |
+
(APIServer pid=199309) INFO 07-14 21:57:57 [loggers.py:273] Engine 000: Avg prompt throughput: 614.5 tokens/s, Avg generation throughput: 2706.4 tokens/s, Running: 254 reqs, Waiting: 266 reqs, GPU KV cache usage: 61.5%, Prefix cache hit rate: 93.4%
|
| 113 |
+
(APIServer pid=199309) INFO 07-14 21:58:07 [loggers.py:273] Engine 000: Avg prompt throughput: 541.1 tokens/s, Avg generation throughput: 2783.3 tokens/s, Running: 255 reqs, Waiting: 198 reqs, GPU KV cache usage: 63.2%, Prefix cache hit rate: 93.4%
|
| 114 |
+
(APIServer pid=199309) INFO 07-14 21:58:17 [loggers.py:273] Engine 000: Avg prompt throughput: 587.5 tokens/s, Avg generation throughput: 2834.1 tokens/s, Running: 256 reqs, Waiting: 123 reqs, GPU KV cache usage: 62.1%, Prefix cache hit rate: 93.4%
|
| 115 |
+
(APIServer pid=199309) INFO 07-14 21:58:27 [loggers.py:273] Engine 000: Avg prompt throughput: 721.8 tokens/s, Avg generation throughput: 2704.7 tokens/s, Running: 255 reqs, Waiting: 35 reqs, GPU KV cache usage: 56.7%, Prefix cache hit rate: 93.4%
|
| 116 |
+
(APIServer pid=199309) INFO 07-14 21:58:37 [loggers.py:273] Engine 000: Avg prompt throughput: 285.7 tokens/s, Avg generation throughput: 2830.4 tokens/s, Running: 226 reqs, Waiting: 0 reqs, GPU KV cache usage: 58.2%, Prefix cache hit rate: 93.4%
|
| 117 |
+
(APIServer pid=199309) INFO 07-14 21:58:47 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 2559.2 tokens/s, Running: 143 reqs, Waiting: 0 reqs, GPU KV cache usage: 47.3%, Prefix cache hit rate: 93.4%
|
| 118 |
+
(APIServer pid=199309) INFO 07-14 21:58:57 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 1863.9 tokens/s, Running: 33 reqs, Waiting: 0 reqs, GPU KV cache usage: 14.7%, Prefix cache hit rate: 93.4%
|
| 119 |
+
(APIServer pid=199309) INFO: 127.0.0.1:60680 - "POST /v1/completions HTTP/1.1" 200 OK
|
| 120 |
+
(EngineCore pid=199426) INFO 07-14 21:59:01 [core.py:1214] [shutdown] EngineCore: trigger received signal=SIGTERM
|
| 121 |
+
(APIServer pid=199309) INFO 07-14 21:59:01 [launcher.py:100] [shutdown] API server: shutdown triggered
|
| 122 |
+
(APIServer pid=199309) INFO 07-14 21:59:01 [launcher.py:116] [shutdown] API server: stopping engine client mode=abort timeout=0s
|
| 123 |
+
(EngineCore pid=199426) INFO 07-14 21:59:01 [core.py:1333] [shutdown] EngineCore: start mode=abort timeout=0s
|
| 124 |
+
(EngineCore pid=199426) INFO 07-14 21:59:01 [core.py:1364] [shutdown] EngineCore: request processing complete; starting resource teardown
|
| 125 |
+
(EngineCore pid=199426) INFO 07-14 21:59:01 [core.py:1227] [shutdown] EngineCore: exiting busy loop
|
| 126 |
+
(APIServer pid=199309) INFO 07-14 21:59:01 [core_client.py:655] [shutdown] MPClient: start timeout=0s
|
| 127 |
+
(APIServer pid=199309) INFO 07-14 21:59:01 [core_client.py:657] [shutdown] MPClient: stopping engine manager
|
| 128 |
+
(APIServer pid=199309) WARNING 07-14 21:59:01 [utils.py:626] [shutdown] Process manager: force killing remaining processes count=1
|
| 129 |
+
(APIServer pid=199309) INFO 07-14 21:59:01 [core_client.py:659] [shutdown] MPClient: engine manager stopped
|
| 130 |
+
(APIServer pid=199309) INFO 07-14 21:59:01 [core_client.py:660] [shutdown] MPClient: cleaning up background resources
|
| 131 |
+
(APIServer pid=199309) INFO 07-14 21:59:01 [core_client.py:662] [shutdown] MPClient: complete
|
| 132 |
+
(APIServer pid=199309) INFO 07-14 21:59:01 [launcher.py:125] [shutdown] API server: engine client stopped
|
| 133 |
+
(APIServer pid=199309) INFO 07-14 21:59:01 [launcher.py:128] [shutdown] API server: signalling HTTP server shutdown
|
| 134 |
+
(APIServer pid=199309) INFO 07-14 21:59:01 [launcher.py:149] [shutdown] API server: shutting down FastAPI HTTP server
|
| 135 |
+
(APIServer pid=199309) INFO: Shutting down
|
| 136 |
+
(APIServer pid=199309) INFO: Waiting for application shutdown.
|
| 137 |
+
(APIServer pid=199309) INFO: Application shutdown complete.
|
| 138 |
+
/home/henry/.local/share/uv/python/cpython-3.12.12-linux-x86_64-gnu/lib/python3.12/multiprocessing/resource_tracker.py:279: UserWarning: resource_tracker: There appear to be 1 leaked semaphore objects to clean up at shutdown
|
| 139 |
+
warnings.warn('resource_tracker: There appear to be %d '
|
evals/healing_breadth/glean_math_keep25_seed1224_long768_step0500_chat.json
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
evals/healing_breadth/glean_math_keep25_seed1224_long768_step0500_chat.json.server.log
ADDED
|
@@ -0,0 +1,255 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
(APIServer pid=304120) INFO 07-15 18:42:16 [api_utils.py:339]
|
| 2 |
+
(APIServer pid=304120) INFO 07-15 18:42:16 [api_utils.py:339] β β ββ ββ
|
| 3 |
+
(APIServer pid=304120) INFO 07-15 18:42:16 [api_utils.py:339] ββ ββ β β β βββ β version 0.25.0
|
| 4 |
+
(APIServer pid=304120) INFO 07-15 18:42:16 [api_utils.py:339] ββββ β β β β model outputs/healed/healing_breadth/glean_math_keep25_seed1224_long768/step0500
|
| 5 |
+
(APIServer pid=304120) INFO 07-15 18:42:16 [api_utils.py:339] ββ βββββ βββββ β β
|
| 6 |
+
(APIServer pid=304120) INFO 07-15 18:42:16 [api_utils.py:339]
|
| 7 |
+
(APIServer pid=304120) INFO 07-15 18:42:16 [api_utils.py:273] non-default args: {'model_tag': 'outputs/healed/healing_breadth/glean_math_keep25_seed1224_long768/step0500', 'host': '127.0.0.1', 'port': 8377, 'model': 'outputs/healed/healing_breadth/glean_math_keep25_seed1224_long768/step0500', 'max_model_len': 2048, 'enforce_eager': True, 'served_model_name': ['student'], 'gpu_memory_utilization': 0.85}
|
| 8 |
+
(APIServer pid=304120) INFO 07-15 18:42:16 [model.py:619] Resolved architecture: PrunedOlmoeForCausalLM
|
| 9 |
+
(APIServer pid=304120) INFO 07-15 18:42:16 [model.py:1776] Using max model len 2048
|
| 10 |
+
(APIServer pid=304120) INFO 07-15 18:42:16 [vllm.py:1042] Asynchronous scheduling is enabled.
|
| 11 |
+
(APIServer pid=304120) WARNING 07-15 18:42:16 [vllm.py:1096] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
|
| 12 |
+
(APIServer pid=304120) WARNING 07-15 18:42:16 [vllm.py:1144] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
|
| 13 |
+
(APIServer pid=304120) INFO 07-15 18:42:16 [kernel.py:292] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native'])
|
| 14 |
+
(APIServer pid=304120) INFO 07-15 18:42:16 [vllm.py:1322] Cudagraph is disabled under eager mode
|
| 15 |
+
(APIServer pid=304120) INFO 07-15 18:42:16 [compilation.py:312] Enabled custom fusions: norm_quant, act_quant
|
| 16 |
+
(EngineCore pid=304237) INFO 07-15 18:42:24 [core.py:114] Initializing a V1 LLM engine (v0.25.0) with config: model='outputs/healed/healing_breadth/glean_math_keep25_seed1224_long768/step0500', speculative_config=None, tokenizer='outputs/healed/healing_breadth/glean_math_keep25_seed1224_long768/step0500', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=False, dtype=torch.bfloat16, max_seq_len=2048, download_dir=None, load_format=auto, tensor_parallel_size=1, pipeline_parallel_size=1, data_parallel_size=1, decode_context_parallel_size=1, dcp_comm_backend=ag_rs, disable_custom_all_reduce=False, quantization=None, quantization_config=None, enforce_eager=True, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False, jit_monitor_mode='warn', jit_monitor_verbose=False), seed=0, served_model_name=student, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'mode': <CompilationMode.NONE: 0>, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'ir_enable_torch_wrap': False, 'splitting_ops': [], 'compile_mm_encoder': False, 'cudagraph_mm_encoder': False, 'encoder_cudagraph_token_budgets': [], 'encoder_cudagraph_max_vision_items_per_batch': 0, 'encoder_cudagraph_max_frames_per_batch': None, 'compile_sizes': [], 'compile_ranges_endpoints': [2048], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'size_asserts': False, 'alignment_asserts': False, 'scalar_asserts': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': <CUDAGraphMode.NONE: 0>, 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': True, 'fuse_act_quant': True, 'fuse_attn_quant': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False, 'fuse_rope_kvcache_cat_mla': False, 'fuse_act_padding': False}, 'max_cudagraph_capture_size': 0, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': False, 'static_all_moe_layers': []}, kernel_config=KernelConfig(ir_op_priority=IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native']), enable_flashinfer_autotune=True, enable_cutedsl_warmup=True, moe_backend='auto', linear_backend='auto')
|
| 17 |
+
(EngineCore pid=304237) INFO 07-15 18:42:24 [parallel_state.py:1607] world_size=1 rank=0 local_rank=0 distributed_init_method=tcp://192.168.0.15:38123 backend=nccl
|
| 18 |
+
(EngineCore pid=304237) INFO 07-15 18:42:24 [parallel_state.py:1942] rank 0 in world size 1 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank 0, EPLB rank N/A
|
| 19 |
+
(EngineCore pid=304237) INFO 07-15 18:42:25 [topk_topp_sampler.py:55] Using FlashInfer for top-p & top-k sampling.
|
| 20 |
+
(EngineCore pid=304237) INFO 07-15 18:42:25 [gpu_model_runner.py:5209] Starting to load model outputs/healed/healing_breadth/glean_math_keep25_seed1224_long768/step0500...
|
| 21 |
+
(EngineCore pid=304237) INFO 07-15 18:42:26 [cuda.py:476] Using FLASH_ATTN attention backend out of potential backends: ['FLASH_ATTN', 'FLASHINFER', 'TRITON_ATTN', 'FLEX_ATTENTION'].
|
| 22 |
+
(EngineCore pid=304237) INFO 07-15 18:42:26 [flash_attn.py:718] Using FlashAttention version 2
|
| 23 |
+
(EngineCore pid=304237) /home/henry/.cache/glean/megablocks-variable-93a1479bc15b/megablocks/grouped_gemm_util.py:10: UserWarning: Grouped GEMM not available.
|
| 24 |
+
(EngineCore pid=304237) warnings.warn('Grouped GEMM not available.')
|
| 25 |
+
(EngineCore pid=304237) INFO 07-15 18:42:26 [weight_utils.py:849] Filesystem type for checkpoints: EXT4. Checkpoint size: 3.89 GiB. Available RAM: 109.34 GiB.
|
| 26 |
+
(EngineCore pid=304237) INFO 07-15 18:42:26 [weight_utils.py:872] Auto-prefetch is disabled because the filesystem (EXT4) is not a recognized network FS (NFS/Lustre). If you want to force prefetching, start vLLM with --safetensors-load-strategy=prefetch.
|
| 27 |
+
(EngineCore pid=304237)
|
| 28 |
+
(EngineCore pid=304237)
|
| 29 |
+
(EngineCore pid=304237)
|
| 30 |
+
(EngineCore pid=304237)
|
| 31 |
+
(EngineCore pid=304237) INFO 07-15 18:42:28 [default_loader.py:430] Loading weights took 2.36 seconds
|
| 32 |
+
(EngineCore pid=304237) INFO 07-15 18:42:28 [gpu_model_runner.py:5306] Model loading took 3.89 GiB memory and 2.546971 seconds
|
| 33 |
+
(EngineCore pid=304237) INFO 07-15 18:42:30 [gpu_worker.py:538] Available KV cache memory: 15.82 GiB
|
| 34 |
+
(EngineCore pid=304237) INFO 07-15 18:42:30 [kv_cache_utils.py:2146] GPU KV cache size: 129,584 tokens
|
| 35 |
+
(EngineCore pid=304237) INFO 07-15 18:42:30 [kv_cache_utils.py:2147] Maximum concurrency for 2,048 tokens per request: 63.27x
|
| 36 |
+
(EngineCore pid=304237) INFO 07-15 18:42:30 [cutedsl_warmup.py:97] Skipping CuTeDSL warmup because no compile units were requested.
|
| 37 |
+
(EngineCore pid=304237) INFO 07-15 18:42:30 [jit_monitor.py:73] Kernel JIT monitor activated; monitored JIT compilations during inference will use mode=warn.
|
| 38 |
+
(EngineCore pid=304237) INFO 07-15 18:42:31 [core.py:344] init engine (profile, create kv cache, warmup model) took 2.09 s
|
| 39 |
+
(EngineCore pid=304237) INFO 07-15 18:42:31 [vllm.py:1042] Asynchronous scheduling is enabled.
|
| 40 |
+
(EngineCore pid=304237) WARNING 07-15 18:42:31 [vllm.py:1096] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
|
| 41 |
+
(EngineCore pid=304237) WARNING 07-15 18:42:31 [vllm.py:1144] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
|
| 42 |
+
(EngineCore pid=304237) INFO 07-15 18:42:31 [kernel.py:292] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native'])
|
| 43 |
+
(EngineCore pid=304237) INFO 07-15 18:42:31 [vllm.py:1322] Cudagraph is disabled under eager mode
|
| 44 |
+
(EngineCore pid=304237) INFO 07-15 18:42:31 [compilation.py:312] Enabled custom fusions: norm_quant, act_quant
|
| 45 |
+
(APIServer pid=304120) INFO 07-15 18:42:31 [api_server.py:612] Supported tasks: ['generate']
|
| 46 |
+
(APIServer pid=304120) WARNING 07-15 18:42:31 [__init__.py:36] SECURITY WARNING: Development endpoints are enabled! This should NOT be used in production!
|
| 47 |
+
(APIServer pid=304120) INFO 07-15 18:42:31 [hf.py:548] Detected the chat template content format to be 'string'. You can set `--chat-template-content-format` to override this.
|
| 48 |
+
(APIServer pid=304120) INFO 07-15 18:42:31 [api_server.py:616] Starting vLLM server on http://127.0.0.1:8377
|
| 49 |
+
(APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:37] Available routes are:
|
| 50 |
+
(APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /openapi.json, Methods: HEAD, GET
|
| 51 |
+
(APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /docs, Methods: HEAD, GET
|
| 52 |
+
(APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /docs/oauth2-redirect, Methods: HEAD, GET
|
| 53 |
+
(APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /redoc, Methods: HEAD, GET
|
| 54 |
+
(APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /load, Methods: GET
|
| 55 |
+
(APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /version, Methods: GET
|
| 56 |
+
(APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /health, Methods: GET
|
| 57 |
+
(APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /metrics, Methods: GET
|
| 58 |
+
(APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /tokenize, Methods: POST
|
| 59 |
+
(APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /detokenize, Methods: POST
|
| 60 |
+
(APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /v1/models, Methods: GET
|
| 61 |
+
(APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /ping, Methods: GET
|
| 62 |
+
(APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /ping, Methods: POST
|
| 63 |
+
(APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /invocations, Methods: POST
|
| 64 |
+
(APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /reset_prefix_cache, Methods: POST
|
| 65 |
+
(APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /reset_mm_cache, Methods: POST
|
| 66 |
+
(APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /reset_encoder_cache, Methods: POST
|
| 67 |
+
(APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /pause, Methods: POST
|
| 68 |
+
(APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /resume, Methods: POST
|
| 69 |
+
(APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /is_paused, Methods: GET
|
| 70 |
+
(APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /init_weight_transfer_engine, Methods: POST
|
| 71 |
+
(APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /start_weight_update, Methods: POST
|
| 72 |
+
(APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /update_weights, Methods: POST
|
| 73 |
+
(APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /finish_weight_update, Methods: POST
|
| 74 |
+
(APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /get_world_size, Methods: GET
|
| 75 |
+
(APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /collective_rpc, Methods: POST
|
| 76 |
+
(APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /server_info, Methods: GET
|
| 77 |
+
(APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /sleep, Methods: POST
|
| 78 |
+
(APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /wake_up, Methods: POST
|
| 79 |
+
(APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /is_sleeping, Methods: GET
|
| 80 |
+
(APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /v1/chat/completions, Methods: POST
|
| 81 |
+
(APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /v1/chat/completions/batch, Methods: POST
|
| 82 |
+
(APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /v1/responses, Methods: POST
|
| 83 |
+
(APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /v1/responses/{response_id}, Methods: GET
|
| 84 |
+
(APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /v1/responses/{response_id}/cancel, Methods: POST
|
| 85 |
+
(APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /v1/completions, Methods: POST
|
| 86 |
+
(APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /v1/messages, Methods: POST
|
| 87 |
+
(APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /v1/messages/count_tokens, Methods: POST
|
| 88 |
+
(APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /generative_scoring, Methods: POST
|
| 89 |
+
(APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /scale_elastic_ep, Methods: POST
|
| 90 |
+
(APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /is_scaling_elastic_ep, Methods: POST
|
| 91 |
+
(APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /v1/chat/completions/render, Methods: POST
|
| 92 |
+
(APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /v1/completions/render, Methods: POST
|
| 93 |
+
(APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /v1/chat/completions/derender, Methods: POST
|
| 94 |
+
(APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /v1/completions/derender, Methods: POST
|
| 95 |
+
(APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /inference/v1/generate, Methods: POST
|
| 96 |
+
(APIServer pid=304120) INFO: Started server process [304120]
|
| 97 |
+
(APIServer pid=304120) INFO: Waiting for application startup.
|
| 98 |
+
(APIServer pid=304120) INFO: Application startup complete.
|
| 99 |
+
(APIServer pid=304120) INFO: 127.0.0.1:52674 - "GET /health HTTP/1.1" 200 OK
|
| 100 |
+
(APIServer pid=304120) ERROR 07-15 18:42:33 [server_utils.py:384] Exception caught. Request id: None
|
| 101 |
+
(APIServer pid=304120) INFO: 127.0.0.1:52682 - "POST /v1/completions HTTP/1.1" 400 Bad Request
|
| 102 |
+
(APIServer pid=304120) INFO 07-15 18:42:33 [launcher.py:100] [shutdown] API server: shutdown triggered
|
| 103 |
+
(APIServer pid=304120) INFO 07-15 18:42:33 [launcher.py:116] [shutdown] API server: stopping engine client mode=abort timeout=0s
|
| 104 |
+
(EngineCore pid=304237) INFO 07-15 18:42:33 [core.py:1214] [shutdown] EngineCore: trigger received signal=SIGTERM
|
| 105 |
+
(EngineCore pid=304237) INFO 07-15 18:42:33 [core.py:1333] [shutdown] EngineCore: start mode=abort timeout=0s
|
| 106 |
+
(EngineCore pid=304237) INFO 07-15 18:42:33 [core.py:1364] [shutdown] EngineCore: request processing complete; starting resource teardown
|
| 107 |
+
(EngineCore pid=304237) INFO 07-15 18:42:33 [core.py:1227] [shutdown] EngineCore: exiting busy loop
|
| 108 |
+
(APIServer pid=304120) INFO 07-15 18:42:33 [core_client.py:655] [shutdown] MPClient: start timeout=0s
|
| 109 |
+
(APIServer pid=304120) INFO 07-15 18:42:33 [core_client.py:657] [shutdown] MPClient: stopping engine manager
|
| 110 |
+
(APIServer pid=304120) WARNING 07-15 18:42:33 [utils.py:626] [shutdown] Process manager: force killing remaining processes count=1
|
| 111 |
+
(APIServer pid=304120) INFO 07-15 18:42:33 [core_client.py:659] [shutdown] MPClient: engine manager stopped
|
| 112 |
+
(APIServer pid=304120) INFO 07-15 18:42:33 [core_client.py:660] [shutdown] MPClient: cleaning up background resources
|
| 113 |
+
(APIServer pid=304120) INFO 07-15 18:42:33 [core_client.py:662] [shutdown] MPClient: complete
|
| 114 |
+
(APIServer pid=304120) INFO 07-15 18:42:33 [launcher.py:125] [shutdown] API server: engine client stopped
|
| 115 |
+
(APIServer pid=304120) INFO 07-15 18:42:33 [launcher.py:128] [shutdown] API server: signalling HTTP server shutdown
|
| 116 |
+
(APIServer pid=304120) INFO 07-15 18:42:33 [launcher.py:149] [shutdown] API server: shutting down FastAPI HTTP server
|
| 117 |
+
(APIServer pid=304120) INFO: Shutting down
|
| 118 |
+
(APIServer pid=304120) INFO: Waiting for application shutdown.
|
| 119 |
+
(APIServer pid=304120) INFO: Application shutdown complete.
|
| 120 |
+
/home/henry/.local/share/uv/python/cpython-3.12.12-linux-x86_64-gnu/lib/python3.12/multiprocessing/resource_tracker.py:279: UserWarning: resource_tracker: There appear to be 1 leaked semaphore objects to clean up at shutdown
|
| 121 |
+
warnings.warn('resource_tracker: There appear to be %d '
|
| 122 |
+
(APIServer pid=306333) INFO 07-15 18:47:04 [api_utils.py:339]
|
| 123 |
+
(APIServer pid=306333) INFO 07-15 18:47:04 [api_utils.py:339] β β ββ ββ
|
| 124 |
+
(APIServer pid=306333) INFO 07-15 18:47:04 [api_utils.py:339] ββ ββ β β β βββ β version 0.25.0
|
| 125 |
+
(APIServer pid=306333) INFO 07-15 18:47:04 [api_utils.py:339] ββββ β β β β model outputs/healed/healing_breadth/glean_math_keep25_seed1224_long768/step0500
|
| 126 |
+
(APIServer pid=306333) INFO 07-15 18:47:04 [api_utils.py:339] ββ βββββ βββββ β β
|
| 127 |
+
(APIServer pid=306333) INFO 07-15 18:47:04 [api_utils.py:339]
|
| 128 |
+
(APIServer pid=306333) INFO 07-15 18:47:04 [api_utils.py:273] non-default args: {'model_tag': 'outputs/healed/healing_breadth/glean_math_keep25_seed1224_long768/step0500', 'host': '127.0.0.1', 'port': 8377, 'model': 'outputs/healed/healing_breadth/glean_math_keep25_seed1224_long768/step0500', 'max_model_len': 2048, 'enforce_eager': True, 'served_model_name': ['student'], 'gpu_memory_utilization': 0.85}
|
| 129 |
+
(APIServer pid=306333) INFO 07-15 18:47:04 [model.py:619] Resolved architecture: PrunedOlmoeForCausalLM
|
| 130 |
+
(APIServer pid=306333) INFO 07-15 18:47:04 [model.py:1776] Using max model len 2048
|
| 131 |
+
(APIServer pid=306333) INFO 07-15 18:47:04 [vllm.py:1042] Asynchronous scheduling is enabled.
|
| 132 |
+
(APIServer pid=306333) WARNING 07-15 18:47:04 [vllm.py:1096] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
|
| 133 |
+
(APIServer pid=306333) WARNING 07-15 18:47:04 [vllm.py:1144] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
|
| 134 |
+
(APIServer pid=306333) INFO 07-15 18:47:04 [kernel.py:292] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native'])
|
| 135 |
+
(APIServer pid=306333) INFO 07-15 18:47:04 [vllm.py:1322] Cudagraph is disabled under eager mode
|
| 136 |
+
(APIServer pid=306333) INFO 07-15 18:47:04 [compilation.py:312] Enabled custom fusions: norm_quant, act_quant
|
| 137 |
+
(EngineCore pid=306450) INFO 07-15 18:47:11 [core.py:114] Initializing a V1 LLM engine (v0.25.0) with config: model='outputs/healed/healing_breadth/glean_math_keep25_seed1224_long768/step0500', speculative_config=None, tokenizer='outputs/healed/healing_breadth/glean_math_keep25_seed1224_long768/step0500', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=False, dtype=torch.bfloat16, max_seq_len=2048, download_dir=None, load_format=auto, tensor_parallel_size=1, pipeline_parallel_size=1, data_parallel_size=1, decode_context_parallel_size=1, dcp_comm_backend=ag_rs, disable_custom_all_reduce=False, quantization=None, quantization_config=None, enforce_eager=True, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False, jit_monitor_mode='warn', jit_monitor_verbose=False), seed=0, served_model_name=student, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'mode': <CompilationMode.NONE: 0>, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'ir_enable_torch_wrap': False, 'splitting_ops': [], 'compile_mm_encoder': False, 'cudagraph_mm_encoder': False, 'encoder_cudagraph_token_budgets': [], 'encoder_cudagraph_max_vision_items_per_batch': 0, 'encoder_cudagraph_max_frames_per_batch': None, 'compile_sizes': [], 'compile_ranges_endpoints': [2048], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'size_asserts': False, 'alignment_asserts': False, 'scalar_asserts': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': <CUDAGraphMode.NONE: 0>, 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': True, 'fuse_act_quant': True, 'fuse_attn_quant': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False, 'fuse_rope_kvcache_cat_mla': False, 'fuse_act_padding': False}, 'max_cudagraph_capture_size': 0, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': False, 'static_all_moe_layers': []}, kernel_config=KernelConfig(ir_op_priority=IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native']), enable_flashinfer_autotune=True, enable_cutedsl_warmup=True, moe_backend='auto', linear_backend='auto')
|
| 138 |
+
(EngineCore pid=306450) INFO 07-15 18:47:12 [parallel_state.py:1607] world_size=1 rank=0 local_rank=0 distributed_init_method=tcp://192.168.0.15:44905 backend=nccl
|
| 139 |
+
(EngineCore pid=306450) INFO 07-15 18:47:12 [parallel_state.py:1942] rank 0 in world size 1 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank 0, EPLB rank N/A
|
| 140 |
+
(EngineCore pid=306450) INFO 07-15 18:47:13 [topk_topp_sampler.py:55] Using FlashInfer for top-p & top-k sampling.
|
| 141 |
+
(EngineCore pid=306450) INFO 07-15 18:47:13 [gpu_model_runner.py:5209] Starting to load model outputs/healed/healing_breadth/glean_math_keep25_seed1224_long768/step0500...
|
| 142 |
+
(EngineCore pid=306450) INFO 07-15 18:47:13 [cuda.py:476] Using FLASH_ATTN attention backend out of potential backends: ['FLASH_ATTN', 'FLASHINFER', 'TRITON_ATTN', 'FLEX_ATTENTION'].
|
| 143 |
+
(EngineCore pid=306450) INFO 07-15 18:47:13 [flash_attn.py:718] Using FlashAttention version 2
|
| 144 |
+
(EngineCore pid=306450) /home/henry/.cache/glean/megablocks-variable-93a1479bc15b/megablocks/grouped_gemm_util.py:10: UserWarning: Grouped GEMM not available.
|
| 145 |
+
(EngineCore pid=306450) warnings.warn('Grouped GEMM not available.')
|
| 146 |
+
(EngineCore pid=306450) INFO 07-15 18:47:13 [weight_utils.py:849] Filesystem type for checkpoints: EXT4. Checkpoint size: 3.89 GiB. Available RAM: 109.49 GiB.
|
| 147 |
+
(EngineCore pid=306450) INFO 07-15 18:47:13 [weight_utils.py:872] Auto-prefetch is disabled because the filesystem (EXT4) is not a recognized network FS (NFS/Lustre). If you want to force prefetching, start vLLM with --safetensors-load-strategy=prefetch.
|
| 148 |
+
(EngineCore pid=306450)
|
| 149 |
+
(EngineCore pid=306450)
|
| 150 |
+
(EngineCore pid=306450)
|
| 151 |
+
(EngineCore pid=306450)
|
| 152 |
+
(EngineCore pid=306450) INFO 07-15 18:47:16 [default_loader.py:430] Loading weights took 2.45 seconds
|
| 153 |
+
(EngineCore pid=306450) INFO 07-15 18:47:16 [gpu_model_runner.py:5306] Model loading took 3.89 GiB memory and 2.628739 seconds
|
| 154 |
+
(EngineCore pid=306450) INFO 07-15 18:47:18 [gpu_worker.py:538] Available KV cache memory: 15.82 GiB
|
| 155 |
+
(EngineCore pid=306450) INFO 07-15 18:47:18 [kv_cache_utils.py:2146] GPU KV cache size: 129,584 tokens
|
| 156 |
+
(EngineCore pid=306450) INFO 07-15 18:47:18 [kv_cache_utils.py:2147] Maximum concurrency for 2,048 tokens per request: 63.27x
|
| 157 |
+
(EngineCore pid=306450) INFO 07-15 18:47:18 [cutedsl_warmup.py:97] Skipping CuTeDSL warmup because no compile units were requested.
|
| 158 |
+
(EngineCore pid=306450) INFO 07-15 18:47:18 [jit_monitor.py:73] Kernel JIT monitor activated; monitored JIT compilations during inference will use mode=warn.
|
| 159 |
+
(EngineCore pid=306450) INFO 07-15 18:47:18 [core.py:344] init engine (profile, create kv cache, warmup model) took 2.07 s
|
| 160 |
+
(EngineCore pid=306450) INFO 07-15 18:47:19 [vllm.py:1042] Asynchronous scheduling is enabled.
|
| 161 |
+
(EngineCore pid=306450) WARNING 07-15 18:47:19 [vllm.py:1096] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
|
| 162 |
+
(EngineCore pid=306450) WARNING 07-15 18:47:19 [vllm.py:1144] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
|
| 163 |
+
(EngineCore pid=306450) INFO 07-15 18:47:19 [kernel.py:292] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native'])
|
| 164 |
+
(EngineCore pid=306450) INFO 07-15 18:47:19 [vllm.py:1322] Cudagraph is disabled under eager mode
|
| 165 |
+
(EngineCore pid=306450) INFO 07-15 18:47:19 [compilation.py:312] Enabled custom fusions: norm_quant, act_quant
|
| 166 |
+
(APIServer pid=306333) INFO 07-15 18:47:19 [api_server.py:612] Supported tasks: ['generate']
|
| 167 |
+
(APIServer pid=306333) WARNING 07-15 18:47:19 [__init__.py:36] SECURITY WARNING: Development endpoints are enabled! This should NOT be used in production!
|
| 168 |
+
(APIServer pid=306333) INFO 07-15 18:47:19 [hf.py:548] Detected the chat template content format to be 'string'. You can set `--chat-template-content-format` to override this.
|
| 169 |
+
(APIServer pid=306333) INFO 07-15 18:47:19 [api_server.py:616] Starting vLLM server on http://127.0.0.1:8377
|
| 170 |
+
(APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:37] Available routes are:
|
| 171 |
+
(APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /openapi.json, Methods: GET, HEAD
|
| 172 |
+
(APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /docs, Methods: GET, HEAD
|
| 173 |
+
(APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /docs/oauth2-redirect, Methods: GET, HEAD
|
| 174 |
+
(APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /redoc, Methods: GET, HEAD
|
| 175 |
+
(APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /load, Methods: GET
|
| 176 |
+
(APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /version, Methods: GET
|
| 177 |
+
(APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /health, Methods: GET
|
| 178 |
+
(APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /metrics, Methods: GET
|
| 179 |
+
(APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /tokenize, Methods: POST
|
| 180 |
+
(APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /detokenize, Methods: POST
|
| 181 |
+
(APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /v1/models, Methods: GET
|
| 182 |
+
(APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /ping, Methods: GET
|
| 183 |
+
(APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /ping, Methods: POST
|
| 184 |
+
(APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /invocations, Methods: POST
|
| 185 |
+
(APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /reset_prefix_cache, Methods: POST
|
| 186 |
+
(APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /reset_mm_cache, Methods: POST
|
| 187 |
+
(APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /reset_encoder_cache, Methods: POST
|
| 188 |
+
(APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /pause, Methods: POST
|
| 189 |
+
(APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /resume, Methods: POST
|
| 190 |
+
(APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /is_paused, Methods: GET
|
| 191 |
+
(APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /init_weight_transfer_engine, Methods: POST
|
| 192 |
+
(APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /start_weight_update, Methods: POST
|
| 193 |
+
(APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /update_weights, Methods: POST
|
| 194 |
+
(APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /finish_weight_update, Methods: POST
|
| 195 |
+
(APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /get_world_size, Methods: GET
|
| 196 |
+
(APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /collective_rpc, Methods: POST
|
| 197 |
+
(APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /server_info, Methods: GET
|
| 198 |
+
(APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /sleep, Methods: POST
|
| 199 |
+
(APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /wake_up, Methods: POST
|
| 200 |
+
(APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /is_sleeping, Methods: GET
|
| 201 |
+
(APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /v1/chat/completions, Methods: POST
|
| 202 |
+
(APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /v1/chat/completions/batch, Methods: POST
|
| 203 |
+
(APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /v1/responses, Methods: POST
|
| 204 |
+
(APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /v1/responses/{response_id}, Methods: GET
|
| 205 |
+
(APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /v1/responses/{response_id}/cancel, Methods: POST
|
| 206 |
+
(APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /v1/completions, Methods: POST
|
| 207 |
+
(APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /v1/messages, Methods: POST
|
| 208 |
+
(APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /v1/messages/count_tokens, Methods: POST
|
| 209 |
+
(APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /generative_scoring, Methods: POST
|
| 210 |
+
(APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /scale_elastic_ep, Methods: POST
|
| 211 |
+
(APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /is_scaling_elastic_ep, Methods: POST
|
| 212 |
+
(APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /v1/chat/completions/render, Methods: POST
|
| 213 |
+
(APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /v1/completions/render, Methods: POST
|
| 214 |
+
(APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /v1/chat/completions/derender, Methods: POST
|
| 215 |
+
(APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /v1/completions/derender, Methods: POST
|
| 216 |
+
(APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /inference/v1/generate, Methods: POST
|
| 217 |
+
(APIServer pid=306333) INFO: Started server process [306333]
|
| 218 |
+
(APIServer pid=306333) INFO: Waiting for application startup.
|
| 219 |
+
(APIServer pid=306333) INFO: Application startup complete.
|
| 220 |
+
(APIServer pid=306333) INFO: 127.0.0.1:55846 - "GET /health HTTP/1.1" 200 OK
|
| 221 |
+
(EngineCore pid=306450) WARNING 07-15 18:47:21 [jit_monitor.py:129] Triton kernel JIT compilation during inference: _build_route_rows. This causes a latency spike; consider extending warmup to cover this shape/config.
|
| 222 |
+
(APIServer pid=306333) INFO 07-15 18:47:29 [loggers.py:273] Engine 000: Avg prompt throughput: 2022.5 tokens/s, Avg generation throughput: 2244.1 tokens/s, Running: 256 reqs, Waiting: 1063 reqs, GPU KV cache usage: 36.3%, Prefix cache hit rate: 93.0%
|
| 223 |
+
(APIServer pid=306333) INFO 07-15 18:47:39 [loggers.py:273] Engine 000: Avg prompt throughput: 662.3 tokens/s, Avg generation throughput: 2934.7 tokens/s, Running: 255 reqs, Waiting: 976 reqs, GPU KV cache usage: 48.2%, Prefix cache hit rate: 93.2%
|
| 224 |
+
(APIServer pid=306333) INFO 07-15 18:47:49 [loggers.py:273] Engine 000: Avg prompt throughput: 1015.2 tokens/s, Avg generation throughput: 2879.5 tokens/s, Running: 256 reqs, Waiting: 850 reqs, GPU KV cache usage: 47.1%, Prefix cache hit rate: 93.2%
|
| 225 |
+
(APIServer pid=306333) INFO 07-15 18:47:59 [loggers.py:273] Engine 000: Avg prompt throughput: 600.1 tokens/s, Avg generation throughput: 2935.5 tokens/s, Running: 253 reqs, Waiting: 769 reqs, GPU KV cache usage: 52.1%, Prefix cache hit rate: 93.3%
|
| 226 |
+
(APIServer pid=306333) INFO 07-15 18:48:09 [loggers.py:273] Engine 000: Avg prompt throughput: 1022.8 tokens/s, Avg generation throughput: 2879.4 tokens/s, Running: 254 reqs, Waiting: 637 reqs, GPU KV cache usage: 44.7%, Prefix cache hit rate: 93.3%
|
| 227 |
+
(APIServer pid=306333) INFO 07-15 18:48:19 [loggers.py:273] Engine 000: Avg prompt throughput: 858.2 tokens/s, Avg generation throughput: 2932.9 tokens/s, Running: 254 reqs, Waiting: 530 reqs, GPU KV cache usage: 46.3%, Prefix cache hit rate: 93.3%
|
| 228 |
+
(APIServer pid=306333) INFO 07-15 18:48:29 [loggers.py:273] Engine 000: Avg prompt throughput: 884.5 tokens/s, Avg generation throughput: 2932.2 tokens/s, Running: 255 reqs, Waiting: 418 reqs, GPU KV cache usage: 47.5%, Prefix cache hit rate: 93.4%
|
| 229 |
+
(APIServer pid=306333) INFO 07-15 18:48:39 [loggers.py:273] Engine 000: Avg prompt throughput: 1016.5 tokens/s, Avg generation throughput: 2930.8 tokens/s, Running: 256 reqs, Waiting: 295 reqs, GPU KV cache usage: 46.8%, Prefix cache hit rate: 93.4%
|
| 230 |
+
(APIServer pid=306333) INFO 07-15 18:48:49 [loggers.py:273] Engine 000: Avg prompt throughput: 882.8 tokens/s, Avg generation throughput: 2932.5 tokens/s, Running: 255 reqs, Waiting: 184 reqs, GPU KV cache usage: 46.7%, Prefix cache hit rate: 93.4%
|
| 231 |
+
(APIServer pid=306333) INFO 07-15 18:48:59 [loggers.py:273] Engine 000: Avg prompt throughput: 813.3 tokens/s, Avg generation throughput: 2933.6 tokens/s, Running: 254 reqs, Waiting: 82 reqs, GPU KV cache usage: 49.3%, Prefix cache hit rate: 93.4%
|
| 232 |
+
(APIServer pid=306333) INFO 07-15 18:49:09 [loggers.py:273] Engine 000: Avg prompt throughput: 669.5 tokens/s, Avg generation throughput: 2900.8 tokens/s, Running: 231 reqs, Waiting: 0 reqs, GPU KV cache usage: 48.0%, Prefix cache hit rate: 93.4%
|
| 233 |
+
(APIServer pid=306333) INFO 07-15 18:49:19 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 2483.2 tokens/s, Running: 84 reqs, Waiting: 0 reqs, GPU KV cache usage: 25.9%, Prefix cache hit rate: 93.4%
|
| 234 |
+
(APIServer pid=306333) INFO: 127.0.0.1:55862 - "POST /v1/completions HTTP/1.1" 200 OK
|
| 235 |
+
(APIServer pid=306333) INFO 07-15 18:49:29 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 682.5 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 93.4%
|
| 236 |
+
(EngineCore pid=306450) INFO 07-15 18:49:29 [core.py:1214] [shutdown] EngineCore: trigger received signal=SIGTERM
|
| 237 |
+
(APIServer pid=306333) INFO 07-15 18:49:29 [launcher.py:100] [shutdown] API server: shutdown triggered
|
| 238 |
+
(APIServer pid=306333) INFO 07-15 18:49:29 [launcher.py:116] [shutdown] API server: stopping engine client mode=abort timeout=0s
|
| 239 |
+
(EngineCore pid=306450) INFO 07-15 18:49:29 [core.py:1333] [shutdown] EngineCore: start mode=abort timeout=0s
|
| 240 |
+
(EngineCore pid=306450) INFO 07-15 18:49:29 [core.py:1364] [shutdown] EngineCore: request processing complete; starting resource teardown
|
| 241 |
+
(EngineCore pid=306450) INFO 07-15 18:49:29 [core.py:1227] [shutdown] EngineCore: exiting busy loop
|
| 242 |
+
(APIServer pid=306333) INFO 07-15 18:49:29 [core_client.py:655] [shutdown] MPClient: start timeout=0s
|
| 243 |
+
(APIServer pid=306333) INFO 07-15 18:49:29 [core_client.py:657] [shutdown] MPClient: stopping engine manager
|
| 244 |
+
(APIServer pid=306333) WARNING 07-15 18:49:29 [utils.py:626] [shutdown] Process manager: force killing remaining processes count=1
|
| 245 |
+
(APIServer pid=306333) INFO 07-15 18:49:29 [core_client.py:659] [shutdown] MPClient: engine manager stopped
|
| 246 |
+
(APIServer pid=306333) INFO 07-15 18:49:29 [core_client.py:660] [shutdown] MPClient: cleaning up background resources
|
| 247 |
+
(APIServer pid=306333) INFO 07-15 18:49:29 [core_client.py:662] [shutdown] MPClient: complete
|
| 248 |
+
(APIServer pid=306333) INFO 07-15 18:49:29 [launcher.py:125] [shutdown] API server: engine client stopped
|
| 249 |
+
(APIServer pid=306333) INFO 07-15 18:49:29 [launcher.py:128] [shutdown] API server: signalling HTTP server shutdown
|
| 250 |
+
(APIServer pid=306333) INFO 07-15 18:49:29 [launcher.py:149] [shutdown] API server: shutting down FastAPI HTTP server
|
| 251 |
+
(APIServer pid=306333) INFO: Shutting down
|
| 252 |
+
(APIServer pid=306333) INFO: Waiting for application shutdown.
|
| 253 |
+
(APIServer pid=306333) INFO: Application shutdown complete.
|
| 254 |
+
/home/henry/.local/share/uv/python/cpython-3.12.12-linux-x86_64-gnu/lib/python3.12/multiprocessing/resource_tracker.py:279: UserWarning: resource_tracker: There appear to be 1 leaked semaphore objects to clean up at shutdown
|
| 255 |
+
warnings.warn('resource_tracker: There appear to be %d '
|
evals/healing_breadth/glean_math_keep25_seed1224_long768_step0500_raw.json
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
evals/healing_breadth/glean_math_keep25_seed1224_long768_step0500_raw.json.server.log
ADDED
|
@@ -0,0 +1,377 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
(APIServer pid=302497) INFO 07-15 18:37:53 [api_utils.py:339]
|
| 2 |
+
(APIServer pid=302497) INFO 07-15 18:37:53 [api_utils.py:339] β β ββ ββ
|
| 3 |
+
(APIServer pid=302497) INFO 07-15 18:37:53 [api_utils.py:339] ββ ββ β β β βββ β version 0.25.0
|
| 4 |
+
(APIServer pid=302497) INFO 07-15 18:37:53 [api_utils.py:339] ββββ β β β β model outputs/healed/healing_breadth/glean_math_keep25_seed1224_long768/step0500
|
| 5 |
+
(APIServer pid=302497) INFO 07-15 18:37:53 [api_utils.py:339] ββ βββββ βββββ β β
|
| 6 |
+
(APIServer pid=302497) INFO 07-15 18:37:53 [api_utils.py:339]
|
| 7 |
+
(APIServer pid=302497) INFO 07-15 18:37:53 [api_utils.py:273] non-default args: {'model_tag': 'outputs/healed/healing_breadth/glean_math_keep25_seed1224_long768/step0500', 'host': '127.0.0.1', 'port': 8377, 'model': 'outputs/healed/healing_breadth/glean_math_keep25_seed1224_long768/step0500', 'max_model_len': 2048, 'enforce_eager': True, 'served_model_name': ['student'], 'gpu_memory_utilization': 0.85}
|
| 8 |
+
(APIServer pid=302497) INFO 07-15 18:37:53 [model.py:619] Resolved architecture: PrunedOlmoeForCausalLM
|
| 9 |
+
(APIServer pid=302497) INFO 07-15 18:37:53 [model.py:1776] Using max model len 2048
|
| 10 |
+
(APIServer pid=302497) INFO 07-15 18:37:53 [vllm.py:1042] Asynchronous scheduling is enabled.
|
| 11 |
+
(APIServer pid=302497) WARNING 07-15 18:37:53 [vllm.py:1096] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
|
| 12 |
+
(APIServer pid=302497) WARNING 07-15 18:37:53 [vllm.py:1144] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
|
| 13 |
+
(APIServer pid=302497) INFO 07-15 18:37:53 [kernel.py:292] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native'])
|
| 14 |
+
(APIServer pid=302497) INFO 07-15 18:37:53 [vllm.py:1322] Cudagraph is disabled under eager mode
|
| 15 |
+
(APIServer pid=302497) INFO 07-15 18:37:53 [compilation.py:312] Enabled custom fusions: norm_quant, act_quant
|
| 16 |
+
(EngineCore pid=302614) INFO 07-15 18:38:00 [core.py:114] Initializing a V1 LLM engine (v0.25.0) with config: model='outputs/healed/healing_breadth/glean_math_keep25_seed1224_long768/step0500', speculative_config=None, tokenizer='outputs/healed/healing_breadth/glean_math_keep25_seed1224_long768/step0500', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=False, dtype=torch.bfloat16, max_seq_len=2048, download_dir=None, load_format=auto, tensor_parallel_size=1, pipeline_parallel_size=1, data_parallel_size=1, decode_context_parallel_size=1, dcp_comm_backend=ag_rs, disable_custom_all_reduce=False, quantization=None, quantization_config=None, enforce_eager=True, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False, jit_monitor_mode='warn', jit_monitor_verbose=False), seed=0, served_model_name=student, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'mode': <CompilationMode.NONE: 0>, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'ir_enable_torch_wrap': False, 'splitting_ops': [], 'compile_mm_encoder': False, 'cudagraph_mm_encoder': False, 'encoder_cudagraph_token_budgets': [], 'encoder_cudagraph_max_vision_items_per_batch': 0, 'encoder_cudagraph_max_frames_per_batch': None, 'compile_sizes': [], 'compile_ranges_endpoints': [2048], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'size_asserts': False, 'alignment_asserts': False, 'scalar_asserts': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': <CUDAGraphMode.NONE: 0>, 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': True, 'fuse_act_quant': True, 'fuse_attn_quant': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False, 'fuse_rope_kvcache_cat_mla': False, 'fuse_act_padding': False}, 'max_cudagraph_capture_size': 0, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': False, 'static_all_moe_layers': []}, kernel_config=KernelConfig(ir_op_priority=IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native']), enable_flashinfer_autotune=True, enable_cutedsl_warmup=True, moe_backend='auto', linear_backend='auto')
|
| 17 |
+
(EngineCore pid=302614) INFO 07-15 18:38:01 [parallel_state.py:1607] world_size=1 rank=0 local_rank=0 distributed_init_method=tcp://192.168.0.15:34145 backend=nccl
|
| 18 |
+
(EngineCore pid=302614) INFO 07-15 18:38:01 [parallel_state.py:1942] rank 0 in world size 1 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank 0, EPLB rank N/A
|
| 19 |
+
(EngineCore pid=302614) INFO 07-15 18:38:02 [topk_topp_sampler.py:55] Using FlashInfer for top-p & top-k sampling.
|
| 20 |
+
(EngineCore pid=302614) INFO 07-15 18:38:02 [gpu_model_runner.py:5209] Starting to load model outputs/healed/healing_breadth/glean_math_keep25_seed1224_long768/step0500...
|
| 21 |
+
(EngineCore pid=302614) INFO 07-15 18:38:02 [cuda.py:476] Using FLASH_ATTN attention backend out of potential backends: ['FLASH_ATTN', 'FLASHINFER', 'TRITON_ATTN', 'FLEX_ATTENTION'].
|
| 22 |
+
(EngineCore pid=302614) INFO 07-15 18:38:02 [flash_attn.py:718] Using FlashAttention version 2
|
| 23 |
+
(EngineCore pid=302614) /home/henry/.cache/glean/megablocks-variable-93a1479bc15b/megablocks/grouped_gemm_util.py:10: UserWarning: Grouped GEMM not available.
|
| 24 |
+
(EngineCore pid=302614) warnings.warn('Grouped GEMM not available.')
|
| 25 |
+
(EngineCore pid=302614) INFO 07-15 18:38:02 [weight_utils.py:849] Filesystem type for checkpoints: EXT4. Checkpoint size: 3.89 GiB. Available RAM: 109.79 GiB.
|
| 26 |
+
(EngineCore pid=302614) INFO 07-15 18:38:02 [weight_utils.py:872] Auto-prefetch is disabled because the filesystem (EXT4) is not a recognized network FS (NFS/Lustre). If you want to force prefetching, start vLLM with --safetensors-load-strategy=prefetch.
|
| 27 |
+
(EngineCore pid=302614)
|
| 28 |
+
(EngineCore pid=302614)
|
| 29 |
+
(EngineCore pid=302614)
|
| 30 |
+
(EngineCore pid=302614)
|
| 31 |
+
(EngineCore pid=302614) INFO 07-15 18:38:05 [default_loader.py:430] Loading weights took 2.50 seconds
|
| 32 |
+
(EngineCore pid=302614) INFO 07-15 18:38:05 [gpu_model_runner.py:5306] Model loading took 3.89 GiB memory and 2.681781 seconds
|
| 33 |
+
(EngineCore pid=302614) INFO 07-15 18:38:07 [gpu_worker.py:538] Available KV cache memory: 15.82 GiB
|
| 34 |
+
(EngineCore pid=302614) INFO 07-15 18:38:07 [kv_cache_utils.py:2146] GPU KV cache size: 129,584 tokens
|
| 35 |
+
(EngineCore pid=302614) INFO 07-15 18:38:07 [kv_cache_utils.py:2147] Maximum concurrency for 2,048 tokens per request: 63.27x
|
| 36 |
+
(EngineCore pid=302614) INFO 07-15 18:38:07 [cutedsl_warmup.py:97] Skipping CuTeDSL warmup because no compile units were requested.
|
| 37 |
+
(EngineCore pid=302614) INFO 07-15 18:38:07 [jit_monitor.py:73] Kernel JIT monitor activated; monitored JIT compilations during inference will use mode=warn.
|
| 38 |
+
(EngineCore pid=302614) INFO 07-15 18:38:07 [core.py:344] init engine (profile, create kv cache, warmup model) took 2.10 s
|
| 39 |
+
(EngineCore pid=302614) INFO 07-15 18:38:08 [vllm.py:1042] Asynchronous scheduling is enabled.
|
| 40 |
+
(EngineCore pid=302614) WARNING 07-15 18:38:08 [vllm.py:1096] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
|
| 41 |
+
(EngineCore pid=302614) WARNING 07-15 18:38:08 [vllm.py:1144] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
|
| 42 |
+
(EngineCore pid=302614) INFO 07-15 18:38:08 [kernel.py:292] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native'])
|
| 43 |
+
(EngineCore pid=302614) INFO 07-15 18:38:08 [vllm.py:1322] Cudagraph is disabled under eager mode
|
| 44 |
+
(EngineCore pid=302614) INFO 07-15 18:38:08 [compilation.py:312] Enabled custom fusions: norm_quant, act_quant
|
| 45 |
+
(APIServer pid=302497) INFO 07-15 18:38:08 [api_server.py:612] Supported tasks: ['generate']
|
| 46 |
+
(APIServer pid=302497) WARNING 07-15 18:38:08 [__init__.py:36] SECURITY WARNING: Development endpoints are enabled! This should NOT be used in production!
|
| 47 |
+
(APIServer pid=302497) INFO 07-15 18:38:08 [hf.py:548] Detected the chat template content format to be 'string'. You can set `--chat-template-content-format` to override this.
|
| 48 |
+
(APIServer pid=302497) INFO 07-15 18:38:08 [api_server.py:616] Starting vLLM server on http://127.0.0.1:8377
|
| 49 |
+
(APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:37] Available routes are:
|
| 50 |
+
(APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /openapi.json, Methods: GET, HEAD
|
| 51 |
+
(APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /docs, Methods: GET, HEAD
|
| 52 |
+
(APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /docs/oauth2-redirect, Methods: GET, HEAD
|
| 53 |
+
(APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /redoc, Methods: GET, HEAD
|
| 54 |
+
(APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /load, Methods: GET
|
| 55 |
+
(APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /version, Methods: GET
|
| 56 |
+
(APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /health, Methods: GET
|
| 57 |
+
(APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /metrics, Methods: GET
|
| 58 |
+
(APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /tokenize, Methods: POST
|
| 59 |
+
(APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /detokenize, Methods: POST
|
| 60 |
+
(APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /v1/models, Methods: GET
|
| 61 |
+
(APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /ping, Methods: GET
|
| 62 |
+
(APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /ping, Methods: POST
|
| 63 |
+
(APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /invocations, Methods: POST
|
| 64 |
+
(APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /reset_prefix_cache, Methods: POST
|
| 65 |
+
(APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /reset_mm_cache, Methods: POST
|
| 66 |
+
(APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /reset_encoder_cache, Methods: POST
|
| 67 |
+
(APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /pause, Methods: POST
|
| 68 |
+
(APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /resume, Methods: POST
|
| 69 |
+
(APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /is_paused, Methods: GET
|
| 70 |
+
(APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /init_weight_transfer_engine, Methods: POST
|
| 71 |
+
(APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /start_weight_update, Methods: POST
|
| 72 |
+
(APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /update_weights, Methods: POST
|
| 73 |
+
(APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /finish_weight_update, Methods: POST
|
| 74 |
+
(APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /get_world_size, Methods: GET
|
| 75 |
+
(APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /collective_rpc, Methods: POST
|
| 76 |
+
(APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /server_info, Methods: GET
|
| 77 |
+
(APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /sleep, Methods: POST
|
| 78 |
+
(APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /wake_up, Methods: POST
|
| 79 |
+
(APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /is_sleeping, Methods: GET
|
| 80 |
+
(APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /v1/chat/completions, Methods: POST
|
| 81 |
+
(APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /v1/chat/completions/batch, Methods: POST
|
| 82 |
+
(APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /v1/responses, Methods: POST
|
| 83 |
+
(APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /v1/responses/{response_id}, Methods: GET
|
| 84 |
+
(APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /v1/responses/{response_id}/cancel, Methods: POST
|
| 85 |
+
(APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /v1/completions, Methods: POST
|
| 86 |
+
(APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /v1/messages, Methods: POST
|
| 87 |
+
(APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /v1/messages/count_tokens, Methods: POST
|
| 88 |
+
(APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /generative_scoring, Methods: POST
|
| 89 |
+
(APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /scale_elastic_ep, Methods: POST
|
| 90 |
+
(APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /is_scaling_elastic_ep, Methods: POST
|
| 91 |
+
(APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /v1/chat/completions/render, Methods: POST
|
| 92 |
+
(APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /v1/completions/render, Methods: POST
|
| 93 |
+
(APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /v1/chat/completions/derender, Methods: POST
|
| 94 |
+
(APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /v1/completions/derender, Methods: POST
|
| 95 |
+
(APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /inference/v1/generate, Methods: POST
|
| 96 |
+
(APIServer pid=302497) INFO: Started server process [302497]
|
| 97 |
+
(APIServer pid=302497) INFO: Waiting for application startup.
|
| 98 |
+
(APIServer pid=302497) INFO: Application startup complete.
|
| 99 |
+
(APIServer pid=302497) INFO: 127.0.0.1:40196 - "GET /health HTTP/1.1" 200 OK
|
| 100 |
+
(APIServer pid=302497) ERROR 07-15 18:38:09 [server_utils.py:384] Exception caught. Request id: None
|
| 101 |
+
(APIServer pid=302497) INFO: 127.0.0.1:40210 - "POST /v1/completions HTTP/1.1" 400 Bad Request
|
| 102 |
+
(APIServer pid=302497) INFO 07-15 18:38:09 [launcher.py:100] [shutdown] API server: shutdown triggered
|
| 103 |
+
(APIServer pid=302497) INFO 07-15 18:38:09 [launcher.py:116] [shutdown] API server: stopping engine client mode=abort timeout=0s
|
| 104 |
+
(EngineCore pid=302614) INFO 07-15 18:38:09 [core.py:1214] [shutdown] EngineCore: trigger received signal=SIGTERM
|
| 105 |
+
(EngineCore pid=302614) INFO 07-15 18:38:09 [core.py:1333] [shutdown] EngineCore: start mode=abort timeout=0s
|
| 106 |
+
(EngineCore pid=302614) INFO 07-15 18:38:09 [core.py:1364] [shutdown] EngineCore: request processing complete; starting resource teardown
|
| 107 |
+
(EngineCore pid=302614) INFO 07-15 18:38:09 [core.py:1227] [shutdown] EngineCore: exiting busy loop
|
| 108 |
+
(APIServer pid=302497) INFO 07-15 18:38:09 [core_client.py:655] [shutdown] MPClient: start timeout=0s
|
| 109 |
+
(APIServer pid=302497) INFO 07-15 18:38:09 [core_client.py:657] [shutdown] MPClient: stopping engine manager
|
| 110 |
+
(APIServer pid=302497) WARNING 07-15 18:38:09 [utils.py:626] [shutdown] Process manager: force killing remaining processes count=1
|
| 111 |
+
(APIServer pid=302497) INFO 07-15 18:38:09 [core_client.py:659] [shutdown] MPClient: engine manager stopped
|
| 112 |
+
(APIServer pid=302497) INFO 07-15 18:38:09 [core_client.py:660] [shutdown] MPClient: cleaning up background resources
|
| 113 |
+
(APIServer pid=302497) INFO 07-15 18:38:09 [core_client.py:662] [shutdown] MPClient: complete
|
| 114 |
+
(APIServer pid=302497) INFO 07-15 18:38:09 [launcher.py:125] [shutdown] API server: engine client stopped
|
| 115 |
+
(APIServer pid=302497) INFO 07-15 18:38:09 [launcher.py:128] [shutdown] API server: signalling HTTP server shutdown
|
| 116 |
+
(APIServer pid=302497) INFO 07-15 18:38:09 [launcher.py:149] [shutdown] API server: shutting down FastAPI HTTP server
|
| 117 |
+
(APIServer pid=302497) INFO: Shutting down
|
| 118 |
+
(APIServer pid=302497) INFO: Waiting for application shutdown.
|
| 119 |
+
(APIServer pid=302497) INFO: Application shutdown complete.
|
| 120 |
+
/home/henry/.local/share/uv/python/cpython-3.12.12-linux-x86_64-gnu/lib/python3.12/multiprocessing/resource_tracker.py:279: UserWarning: resource_tracker: There appear to be 1 leaked semaphore objects to clean up at shutdown
|
| 121 |
+
warnings.warn('resource_tracker: There appear to be %d '
|
| 122 |
+
(APIServer pid=303449) INFO 07-15 18:41:35 [api_utils.py:339]
|
| 123 |
+
(APIServer pid=303449) INFO 07-15 18:41:35 [api_utils.py:339] β β ββ ββ
|
| 124 |
+
(APIServer pid=303449) INFO 07-15 18:41:35 [api_utils.py:339] ββ ββ β β β βββ β version 0.25.0
|
| 125 |
+
(APIServer pid=303449) INFO 07-15 18:41:35 [api_utils.py:339] ββββ β β β β model outputs/healed/healing_breadth/glean_math_keep25_seed1224_long768/step0500
|
| 126 |
+
(APIServer pid=303449) INFO 07-15 18:41:35 [api_utils.py:339] ββ βββββ βββββ β β
|
| 127 |
+
(APIServer pid=303449) INFO 07-15 18:41:35 [api_utils.py:339]
|
| 128 |
+
(APIServer pid=303449) INFO 07-15 18:41:35 [api_utils.py:273] non-default args: {'model_tag': 'outputs/healed/healing_breadth/glean_math_keep25_seed1224_long768/step0500', 'host': '127.0.0.1', 'port': 8377, 'model': 'outputs/healed/healing_breadth/glean_math_keep25_seed1224_long768/step0500', 'max_model_len': 2048, 'enforce_eager': True, 'served_model_name': ['student'], 'gpu_memory_utilization': 0.85}
|
| 129 |
+
(APIServer pid=303449) INFO 07-15 18:41:35 [model.py:619] Resolved architecture: PrunedOlmoeForCausalLM
|
| 130 |
+
(APIServer pid=303449) INFO 07-15 18:41:35 [model.py:1776] Using max model len 2048
|
| 131 |
+
(APIServer pid=303449) INFO 07-15 18:41:35 [vllm.py:1042] Asynchronous scheduling is enabled.
|
| 132 |
+
(APIServer pid=303449) WARNING 07-15 18:41:35 [vllm.py:1096] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
|
| 133 |
+
(APIServer pid=303449) WARNING 07-15 18:41:35 [vllm.py:1144] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
|
| 134 |
+
(APIServer pid=303449) INFO 07-15 18:41:35 [kernel.py:292] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native'])
|
| 135 |
+
(APIServer pid=303449) INFO 07-15 18:41:36 [vllm.py:1322] Cudagraph is disabled under eager mode
|
| 136 |
+
(APIServer pid=303449) INFO 07-15 18:41:36 [compilation.py:312] Enabled custom fusions: norm_quant, act_quant
|
| 137 |
+
(EngineCore pid=303564) INFO 07-15 18:41:43 [core.py:114] Initializing a V1 LLM engine (v0.25.0) with config: model='outputs/healed/healing_breadth/glean_math_keep25_seed1224_long768/step0500', speculative_config=None, tokenizer='outputs/healed/healing_breadth/glean_math_keep25_seed1224_long768/step0500', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=False, dtype=torch.bfloat16, max_seq_len=2048, download_dir=None, load_format=auto, tensor_parallel_size=1, pipeline_parallel_size=1, data_parallel_size=1, decode_context_parallel_size=1, dcp_comm_backend=ag_rs, disable_custom_all_reduce=False, quantization=None, quantization_config=None, enforce_eager=True, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False, jit_monitor_mode='warn', jit_monitor_verbose=False), seed=0, served_model_name=student, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'mode': <CompilationMode.NONE: 0>, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'ir_enable_torch_wrap': False, 'splitting_ops': [], 'compile_mm_encoder': False, 'cudagraph_mm_encoder': False, 'encoder_cudagraph_token_budgets': [], 'encoder_cudagraph_max_vision_items_per_batch': 0, 'encoder_cudagraph_max_frames_per_batch': None, 'compile_sizes': [], 'compile_ranges_endpoints': [2048], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'size_asserts': False, 'alignment_asserts': False, 'scalar_asserts': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': <CUDAGraphMode.NONE: 0>, 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': True, 'fuse_act_quant': True, 'fuse_attn_quant': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False, 'fuse_rope_kvcache_cat_mla': False, 'fuse_act_padding': False}, 'max_cudagraph_capture_size': 0, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': False, 'static_all_moe_layers': []}, kernel_config=KernelConfig(ir_op_priority=IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native']), enable_flashinfer_autotune=True, enable_cutedsl_warmup=True, moe_backend='auto', linear_backend='auto')
|
| 138 |
+
(EngineCore pid=303564) INFO 07-15 18:41:43 [parallel_state.py:1607] world_size=1 rank=0 local_rank=0 distributed_init_method=tcp://192.168.0.15:45641 backend=nccl
|
| 139 |
+
(EngineCore pid=303564) INFO 07-15 18:41:43 [parallel_state.py:1942] rank 0 in world size 1 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank 0, EPLB rank N/A
|
| 140 |
+
(EngineCore pid=303564) INFO 07-15 18:41:44 [topk_topp_sampler.py:55] Using FlashInfer for top-p & top-k sampling.
|
| 141 |
+
(EngineCore pid=303564) INFO 07-15 18:41:44 [gpu_model_runner.py:5209] Starting to load model outputs/healed/healing_breadth/glean_math_keep25_seed1224_long768/step0500...
|
| 142 |
+
(EngineCore pid=303564) INFO 07-15 18:41:45 [cuda.py:476] Using FLASH_ATTN attention backend out of potential backends: ['FLASH_ATTN', 'FLASHINFER', 'TRITON_ATTN', 'FLEX_ATTENTION'].
|
| 143 |
+
(EngineCore pid=303564) INFO 07-15 18:41:45 [flash_attn.py:718] Using FlashAttention version 2
|
| 144 |
+
(EngineCore pid=303564) /home/henry/.cache/glean/megablocks-variable-93a1479bc15b/megablocks/grouped_gemm_util.py:10: UserWarning: Grouped GEMM not available.
|
| 145 |
+
(EngineCore pid=303564) warnings.warn('Grouped GEMM not available.')
|
| 146 |
+
(EngineCore pid=303564) INFO 07-15 18:41:45 [weight_utils.py:849] Filesystem type for checkpoints: EXT4. Checkpoint size: 3.89 GiB. Available RAM: 109.58 GiB.
|
| 147 |
+
(EngineCore pid=303564) INFO 07-15 18:41:45 [weight_utils.py:872] Auto-prefetch is disabled because the filesystem (EXT4) is not a recognized network FS (NFS/Lustre). If you want to force prefetching, start vLLM with --safetensors-load-strategy=prefetch.
|
| 148 |
+
(EngineCore pid=303564)
|
| 149 |
+
(EngineCore pid=303564)
|
| 150 |
+
(EngineCore pid=303564)
|
| 151 |
+
(EngineCore pid=303564)
|
| 152 |
+
(EngineCore pid=303564) INFO 07-15 18:41:47 [default_loader.py:430] Loading weights took 2.36 seconds
|
| 153 |
+
(EngineCore pid=303564) INFO 07-15 18:41:47 [gpu_model_runner.py:5306] Model loading took 3.89 GiB memory and 2.540095 seconds
|
| 154 |
+
(EngineCore pid=303564) INFO 07-15 18:41:49 [gpu_worker.py:538] Available KV cache memory: 15.82 GiB
|
| 155 |
+
(EngineCore pid=303564) INFO 07-15 18:41:49 [kv_cache_utils.py:2146] GPU KV cache size: 129,584 tokens
|
| 156 |
+
(EngineCore pid=303564) INFO 07-15 18:41:49 [kv_cache_utils.py:2147] Maximum concurrency for 2,048 tokens per request: 63.27x
|
| 157 |
+
(EngineCore pid=303564) INFO 07-15 18:41:49 [cutedsl_warmup.py:97] Skipping CuTeDSL warmup because no compile units were requested.
|
| 158 |
+
(EngineCore pid=303564) INFO 07-15 18:41:49 [jit_monitor.py:73] Kernel JIT monitor activated; monitored JIT compilations during inference will use mode=warn.
|
| 159 |
+
(EngineCore pid=303564) INFO 07-15 18:41:50 [core.py:344] init engine (profile, create kv cache, warmup model) took 2.08 s
|
| 160 |
+
(EngineCore pid=303564) INFO 07-15 18:41:50 [vllm.py:1042] Asynchronous scheduling is enabled.
|
| 161 |
+
(EngineCore pid=303564) WARNING 07-15 18:41:50 [vllm.py:1096] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
|
| 162 |
+
(EngineCore pid=303564) WARNING 07-15 18:41:50 [vllm.py:1144] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
|
| 163 |
+
(EngineCore pid=303564) INFO 07-15 18:41:50 [kernel.py:292] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native'])
|
| 164 |
+
(EngineCore pid=303564) INFO 07-15 18:41:50 [vllm.py:1322] Cudagraph is disabled under eager mode
|
| 165 |
+
(EngineCore pid=303564) INFO 07-15 18:41:50 [compilation.py:312] Enabled custom fusions: norm_quant, act_quant
|
| 166 |
+
(APIServer pid=303449) INFO 07-15 18:41:50 [api_server.py:612] Supported tasks: ['generate']
|
| 167 |
+
(APIServer pid=303449) WARNING 07-15 18:41:50 [__init__.py:36] SECURITY WARNING: Development endpoints are enabled! This should NOT be used in production!
|
| 168 |
+
(APIServer pid=303449) INFO 07-15 18:41:50 [hf.py:548] Detected the chat template content format to be 'string'. You can set `--chat-template-content-format` to override this.
|
| 169 |
+
(APIServer pid=303449) INFO 07-15 18:41:50 [api_server.py:616] Starting vLLM server on http://127.0.0.1:8377
|
| 170 |
+
(APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:37] Available routes are:
|
| 171 |
+
(APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /openapi.json, Methods: GET, HEAD
|
| 172 |
+
(APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /docs, Methods: GET, HEAD
|
| 173 |
+
(APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /docs/oauth2-redirect, Methods: GET, HEAD
|
| 174 |
+
(APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /redoc, Methods: GET, HEAD
|
| 175 |
+
(APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /load, Methods: GET
|
| 176 |
+
(APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /version, Methods: GET
|
| 177 |
+
(APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /health, Methods: GET
|
| 178 |
+
(APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /metrics, Methods: GET
|
| 179 |
+
(APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /tokenize, Methods: POST
|
| 180 |
+
(APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /detokenize, Methods: POST
|
| 181 |
+
(APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /v1/models, Methods: GET
|
| 182 |
+
(APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /ping, Methods: GET
|
| 183 |
+
(APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /ping, Methods: POST
|
| 184 |
+
(APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /invocations, Methods: POST
|
| 185 |
+
(APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /reset_prefix_cache, Methods: POST
|
| 186 |
+
(APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /reset_mm_cache, Methods: POST
|
| 187 |
+
(APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /reset_encoder_cache, Methods: POST
|
| 188 |
+
(APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /pause, Methods: POST
|
| 189 |
+
(APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /resume, Methods: POST
|
| 190 |
+
(APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /is_paused, Methods: GET
|
| 191 |
+
(APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /init_weight_transfer_engine, Methods: POST
|
| 192 |
+
(APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /start_weight_update, Methods: POST
|
| 193 |
+
(APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /update_weights, Methods: POST
|
| 194 |
+
(APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /finish_weight_update, Methods: POST
|
| 195 |
+
(APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /get_world_size, Methods: GET
|
| 196 |
+
(APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /collective_rpc, Methods: POST
|
| 197 |
+
(APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /server_info, Methods: GET
|
| 198 |
+
(APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /sleep, Methods: POST
|
| 199 |
+
(APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /wake_up, Methods: POST
|
| 200 |
+
(APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /is_sleeping, Methods: GET
|
| 201 |
+
(APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /v1/chat/completions, Methods: POST
|
| 202 |
+
(APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /v1/chat/completions/batch, Methods: POST
|
| 203 |
+
(APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /v1/responses, Methods: POST
|
| 204 |
+
(APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /v1/responses/{response_id}, Methods: GET
|
| 205 |
+
(APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /v1/responses/{response_id}/cancel, Methods: POST
|
| 206 |
+
(APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /v1/completions, Methods: POST
|
| 207 |
+
(APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /v1/messages, Methods: POST
|
| 208 |
+
(APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /v1/messages/count_tokens, Methods: POST
|
| 209 |
+
(APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /generative_scoring, Methods: POST
|
| 210 |
+
(APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /scale_elastic_ep, Methods: POST
|
| 211 |
+
(APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /is_scaling_elastic_ep, Methods: POST
|
| 212 |
+
(APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /v1/chat/completions/render, Methods: POST
|
| 213 |
+
(APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /v1/completions/render, Methods: POST
|
| 214 |
+
(APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /v1/chat/completions/derender, Methods: POST
|
| 215 |
+
(APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /v1/completions/derender, Methods: POST
|
| 216 |
+
(APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /inference/v1/generate, Methods: POST
|
| 217 |
+
(APIServer pid=303449) INFO: Started server process [303449]
|
| 218 |
+
(APIServer pid=303449) INFO: Waiting for application startup.
|
| 219 |
+
(APIServer pid=303449) INFO: Application startup complete.
|
| 220 |
+
(APIServer pid=303449) INFO: 127.0.0.1:41890 - "GET /health HTTP/1.1" 200 OK
|
| 221 |
+
(APIServer pid=303449) ERROR 07-15 18:41:52 [server_utils.py:384] Exception caught. Request id: None
|
| 222 |
+
(APIServer pid=303449) INFO: 127.0.0.1:41906 - "POST /v1/completions HTTP/1.1" 400 Bad Request
|
| 223 |
+
(APIServer pid=303449) INFO 07-15 18:41:52 [launcher.py:100] [shutdown] API server: shutdown triggered
|
| 224 |
+
(APIServer pid=303449) INFO 07-15 18:41:52 [launcher.py:116] [shutdown] API server: stopping engine client mode=abort timeout=0s
|
| 225 |
+
(EngineCore pid=303564) INFO 07-15 18:41:52 [core.py:1214] [shutdown] EngineCore: trigger received signal=SIGTERM
|
| 226 |
+
(EngineCore pid=303564) INFO 07-15 18:41:52 [core.py:1333] [shutdown] EngineCore: start mode=abort timeout=0s
|
| 227 |
+
(EngineCore pid=303564) INFO 07-15 18:41:52 [core.py:1364] [shutdown] EngineCore: request processing complete; starting resource teardown
|
| 228 |
+
(EngineCore pid=303564) INFO 07-15 18:41:52 [core.py:1227] [shutdown] EngineCore: exiting busy loop
|
| 229 |
+
(APIServer pid=303449) INFO 07-15 18:41:52 [core_client.py:655] [shutdown] MPClient: start timeout=0s
|
| 230 |
+
(APIServer pid=303449) INFO 07-15 18:41:52 [core_client.py:657] [shutdown] MPClient: stopping engine manager
|
| 231 |
+
(APIServer pid=303449) WARNING 07-15 18:41:52 [utils.py:626] [shutdown] Process manager: force killing remaining processes count=1
|
| 232 |
+
(APIServer pid=303449) INFO 07-15 18:41:52 [core_client.py:659] [shutdown] MPClient: engine manager stopped
|
| 233 |
+
(APIServer pid=303449) INFO 07-15 18:41:52 [core_client.py:660] [shutdown] MPClient: cleaning up background resources
|
| 234 |
+
(APIServer pid=303449) INFO 07-15 18:41:52 [core_client.py:662] [shutdown] MPClient: complete
|
| 235 |
+
(APIServer pid=303449) INFO 07-15 18:41:52 [launcher.py:125] [shutdown] API server: engine client stopped
|
| 236 |
+
(APIServer pid=303449) INFO 07-15 18:41:52 [launcher.py:128] [shutdown] API server: signalling HTTP server shutdown
|
| 237 |
+
(APIServer pid=303449) INFO 07-15 18:41:52 [launcher.py:149] [shutdown] API server: shutting down FastAPI HTTP server
|
| 238 |
+
(APIServer pid=303449) INFO: Shutting down
|
| 239 |
+
(APIServer pid=303449) INFO: Waiting for application shutdown.
|
| 240 |
+
(APIServer pid=303449) INFO: Application shutdown complete.
|
| 241 |
+
/home/henry/.local/share/uv/python/cpython-3.12.12-linux-x86_64-gnu/lib/python3.12/multiprocessing/resource_tracker.py:279: UserWarning: resource_tracker: There appear to be 1 leaked semaphore objects to clean up at shutdown
|
| 242 |
+
warnings.warn('resource_tracker: There appear to be %d '
|
| 243 |
+
(APIServer pid=305508) INFO 07-15 18:43:59 [api_utils.py:339]
|
| 244 |
+
(APIServer pid=305508) INFO 07-15 18:43:59 [api_utils.py:339] β β ββ ββ
|
| 245 |
+
(APIServer pid=305508) INFO 07-15 18:43:59 [api_utils.py:339] ββ ββ β β β βββ β version 0.25.0
|
| 246 |
+
(APIServer pid=305508) INFO 07-15 18:43:59 [api_utils.py:339] ββββ β β β β model outputs/healed/healing_breadth/glean_math_keep25_seed1224_long768/step0500
|
| 247 |
+
(APIServer pid=305508) INFO 07-15 18:43:59 [api_utils.py:339] ββ βββββ βββββ β β
|
| 248 |
+
(APIServer pid=305508) INFO 07-15 18:43:59 [api_utils.py:339]
|
| 249 |
+
(APIServer pid=305508) INFO 07-15 18:43:59 [api_utils.py:273] non-default args: {'model_tag': 'outputs/healed/healing_breadth/glean_math_keep25_seed1224_long768/step0500', 'host': '127.0.0.1', 'port': 8377, 'model': 'outputs/healed/healing_breadth/glean_math_keep25_seed1224_long768/step0500', 'max_model_len': 2048, 'enforce_eager': True, 'served_model_name': ['student'], 'gpu_memory_utilization': 0.85}
|
| 250 |
+
(APIServer pid=305508) INFO 07-15 18:43:59 [model.py:619] Resolved architecture: PrunedOlmoeForCausalLM
|
| 251 |
+
(APIServer pid=305508) INFO 07-15 18:43:59 [model.py:1776] Using max model len 2048
|
| 252 |
+
(APIServer pid=305508) INFO 07-15 18:43:59 [vllm.py:1042] Asynchronous scheduling is enabled.
|
| 253 |
+
(APIServer pid=305508) WARNING 07-15 18:43:59 [vllm.py:1096] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
|
| 254 |
+
(APIServer pid=305508) WARNING 07-15 18:43:59 [vllm.py:1144] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
|
| 255 |
+
(APIServer pid=305508) INFO 07-15 18:43:59 [kernel.py:292] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native'])
|
| 256 |
+
(APIServer pid=305508) INFO 07-15 18:43:59 [vllm.py:1322] Cudagraph is disabled under eager mode
|
| 257 |
+
(APIServer pid=305508) INFO 07-15 18:43:59 [compilation.py:312] Enabled custom fusions: norm_quant, act_quant
|
| 258 |
+
(EngineCore pid=305625) INFO 07-15 18:44:06 [core.py:114] Initializing a V1 LLM engine (v0.25.0) with config: model='outputs/healed/healing_breadth/glean_math_keep25_seed1224_long768/step0500', speculative_config=None, tokenizer='outputs/healed/healing_breadth/glean_math_keep25_seed1224_long768/step0500', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=False, dtype=torch.bfloat16, max_seq_len=2048, download_dir=None, load_format=auto, tensor_parallel_size=1, pipeline_parallel_size=1, data_parallel_size=1, decode_context_parallel_size=1, dcp_comm_backend=ag_rs, disable_custom_all_reduce=False, quantization=None, quantization_config=None, enforce_eager=True, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False, jit_monitor_mode='warn', jit_monitor_verbose=False), seed=0, served_model_name=student, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'mode': <CompilationMode.NONE: 0>, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'ir_enable_torch_wrap': False, 'splitting_ops': [], 'compile_mm_encoder': False, 'cudagraph_mm_encoder': False, 'encoder_cudagraph_token_budgets': [], 'encoder_cudagraph_max_vision_items_per_batch': 0, 'encoder_cudagraph_max_frames_per_batch': None, 'compile_sizes': [], 'compile_ranges_endpoints': [2048], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'size_asserts': False, 'alignment_asserts': False, 'scalar_asserts': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': <CUDAGraphMode.NONE: 0>, 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': True, 'fuse_act_quant': True, 'fuse_attn_quant': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False, 'fuse_rope_kvcache_cat_mla': False, 'fuse_act_padding': False}, 'max_cudagraph_capture_size': 0, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': False, 'static_all_moe_layers': []}, kernel_config=KernelConfig(ir_op_priority=IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native']), enable_flashinfer_autotune=True, enable_cutedsl_warmup=True, moe_backend='auto', linear_backend='auto')
|
| 259 |
+
(EngineCore pid=305625) INFO 07-15 18:44:06 [parallel_state.py:1607] world_size=1 rank=0 local_rank=0 distributed_init_method=tcp://192.168.0.15:53125 backend=nccl
|
| 260 |
+
(EngineCore pid=305625) INFO 07-15 18:44:07 [parallel_state.py:1942] rank 0 in world size 1 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank 0, EPLB rank N/A
|
| 261 |
+
(EngineCore pid=305625) INFO 07-15 18:44:07 [topk_topp_sampler.py:55] Using FlashInfer for top-p & top-k sampling.
|
| 262 |
+
(EngineCore pid=305625) INFO 07-15 18:44:07 [gpu_model_runner.py:5209] Starting to load model outputs/healed/healing_breadth/glean_math_keep25_seed1224_long768/step0500...
|
| 263 |
+
(EngineCore pid=305625) INFO 07-15 18:44:08 [cuda.py:476] Using FLASH_ATTN attention backend out of potential backends: ['FLASH_ATTN', 'FLASHINFER', 'TRITON_ATTN', 'FLEX_ATTENTION'].
|
| 264 |
+
(EngineCore pid=305625) INFO 07-15 18:44:08 [flash_attn.py:718] Using FlashAttention version 2
|
| 265 |
+
(EngineCore pid=305625) /home/henry/.cache/glean/megablocks-variable-93a1479bc15b/megablocks/grouped_gemm_util.py:10: UserWarning: Grouped GEMM not available.
|
| 266 |
+
(EngineCore pid=305625) warnings.warn('Grouped GEMM not available.')
|
| 267 |
+
(EngineCore pid=305625) INFO 07-15 18:44:08 [weight_utils.py:849] Filesystem type for checkpoints: EXT4. Checkpoint size: 3.89 GiB. Available RAM: 109.52 GiB.
|
| 268 |
+
(EngineCore pid=305625) INFO 07-15 18:44:08 [weight_utils.py:872] Auto-prefetch is disabled because the filesystem (EXT4) is not a recognized network FS (NFS/Lustre). If you want to force prefetching, start vLLM with --safetensors-load-strategy=prefetch.
|
| 269 |
+
(EngineCore pid=305625)
|
| 270 |
+
(EngineCore pid=305625)
|
| 271 |
+
(EngineCore pid=305625)
|
| 272 |
+
(EngineCore pid=305625)
|
| 273 |
+
(EngineCore pid=305625) INFO 07-15 18:44:10 [default_loader.py:430] Loading weights took 2.39 seconds
|
| 274 |
+
(EngineCore pid=305625) INFO 07-15 18:44:11 [gpu_model_runner.py:5306] Model loading took 3.89 GiB memory and 2.570320 seconds
|
| 275 |
+
(EngineCore pid=305625) INFO 07-15 18:44:12 [gpu_worker.py:538] Available KV cache memory: 15.82 GiB
|
| 276 |
+
(EngineCore pid=305625) INFO 07-15 18:44:12 [kv_cache_utils.py:2146] GPU KV cache size: 129,584 tokens
|
| 277 |
+
(EngineCore pid=305625) INFO 07-15 18:44:12 [kv_cache_utils.py:2147] Maximum concurrency for 2,048 tokens per request: 63.27x
|
| 278 |
+
(EngineCore pid=305625) INFO 07-15 18:44:12 [cutedsl_warmup.py:97] Skipping CuTeDSL warmup because no compile units were requested.
|
| 279 |
+
(EngineCore pid=305625) INFO 07-15 18:44:13 [jit_monitor.py:73] Kernel JIT monitor activated; monitored JIT compilations during inference will use mode=warn.
|
| 280 |
+
(EngineCore pid=305625) INFO 07-15 18:44:13 [core.py:344] init engine (profile, create kv cache, warmup model) took 2.05 s
|
| 281 |
+
(EngineCore pid=305625) INFO 07-15 18:44:13 [vllm.py:1042] Asynchronous scheduling is enabled.
|
| 282 |
+
(EngineCore pid=305625) WARNING 07-15 18:44:13 [vllm.py:1096] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
|
| 283 |
+
(EngineCore pid=305625) WARNING 07-15 18:44:13 [vllm.py:1144] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
|
| 284 |
+
(EngineCore pid=305625) INFO 07-15 18:44:13 [kernel.py:292] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native'])
|
| 285 |
+
(EngineCore pid=305625) INFO 07-15 18:44:13 [vllm.py:1322] Cudagraph is disabled under eager mode
|
| 286 |
+
(EngineCore pid=305625) INFO 07-15 18:44:13 [compilation.py:312] Enabled custom fusions: norm_quant, act_quant
|
| 287 |
+
(APIServer pid=305508) INFO 07-15 18:44:13 [api_server.py:612] Supported tasks: ['generate']
|
| 288 |
+
(APIServer pid=305508) WARNING 07-15 18:44:13 [__init__.py:36] SECURITY WARNING: Development endpoints are enabled! This should NOT be used in production!
|
| 289 |
+
(APIServer pid=305508) INFO 07-15 18:44:13 [hf.py:548] Detected the chat template content format to be 'string'. You can set `--chat-template-content-format` to override this.
|
| 290 |
+
(APIServer pid=305508) INFO 07-15 18:44:13 [api_server.py:616] Starting vLLM server on http://127.0.0.1:8377
|
| 291 |
+
(APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:37] Available routes are:
|
| 292 |
+
(APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /openapi.json, Methods: GET, HEAD
|
| 293 |
+
(APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /docs, Methods: GET, HEAD
|
| 294 |
+
(APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /docs/oauth2-redirect, Methods: GET, HEAD
|
| 295 |
+
(APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /redoc, Methods: GET, HEAD
|
| 296 |
+
(APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /load, Methods: GET
|
| 297 |
+
(APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /version, Methods: GET
|
| 298 |
+
(APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /health, Methods: GET
|
| 299 |
+
(APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /metrics, Methods: GET
|
| 300 |
+
(APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /tokenize, Methods: POST
|
| 301 |
+
(APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /detokenize, Methods: POST
|
| 302 |
+
(APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /v1/models, Methods: GET
|
| 303 |
+
(APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /ping, Methods: GET
|
| 304 |
+
(APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /ping, Methods: POST
|
| 305 |
+
(APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /invocations, Methods: POST
|
| 306 |
+
(APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /reset_prefix_cache, Methods: POST
|
| 307 |
+
(APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /reset_mm_cache, Methods: POST
|
| 308 |
+
(APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /reset_encoder_cache, Methods: POST
|
| 309 |
+
(APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /pause, Methods: POST
|
| 310 |
+
(APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /resume, Methods: POST
|
| 311 |
+
(APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /is_paused, Methods: GET
|
| 312 |
+
(APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /init_weight_transfer_engine, Methods: POST
|
| 313 |
+
(APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /start_weight_update, Methods: POST
|
| 314 |
+
(APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /update_weights, Methods: POST
|
| 315 |
+
(APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /finish_weight_update, Methods: POST
|
| 316 |
+
(APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /get_world_size, Methods: GET
|
| 317 |
+
(APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /collective_rpc, Methods: POST
|
| 318 |
+
(APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /server_info, Methods: GET
|
| 319 |
+
(APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /sleep, Methods: POST
|
| 320 |
+
(APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /wake_up, Methods: POST
|
| 321 |
+
(APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /is_sleeping, Methods: GET
|
| 322 |
+
(APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /v1/chat/completions, Methods: POST
|
| 323 |
+
(APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /v1/chat/completions/batch, Methods: POST
|
| 324 |
+
(APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /v1/responses, Methods: POST
|
| 325 |
+
(APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /v1/responses/{response_id}, Methods: GET
|
| 326 |
+
(APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /v1/responses/{response_id}/cancel, Methods: POST
|
| 327 |
+
(APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /v1/completions, Methods: POST
|
| 328 |
+
(APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /v1/messages, Methods: POST
|
| 329 |
+
(APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /v1/messages/count_tokens, Methods: POST
|
| 330 |
+
(APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /generative_scoring, Methods: POST
|
| 331 |
+
(APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /scale_elastic_ep, Methods: POST
|
| 332 |
+
(APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /is_scaling_elastic_ep, Methods: POST
|
| 333 |
+
(APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /v1/chat/completions/render, Methods: POST
|
| 334 |
+
(APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /v1/completions/render, Methods: POST
|
| 335 |
+
(APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /v1/chat/completions/derender, Methods: POST
|
| 336 |
+
(APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /v1/completions/derender, Methods: POST
|
| 337 |
+
(APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /inference/v1/generate, Methods: POST
|
| 338 |
+
(APIServer pid=305508) INFO: Started server process [305508]
|
| 339 |
+
(APIServer pid=305508) INFO: Waiting for application startup.
|
| 340 |
+
(APIServer pid=305508) INFO: Application startup complete.
|
| 341 |
+
(APIServer pid=305508) INFO: 127.0.0.1:39768 - "GET /health HTTP/1.1" 200 OK
|
| 342 |
+
(EngineCore pid=305625) WARNING 07-15 18:44:15 [jit_monitor.py:129] Triton kernel JIT compilation during inference: _build_route_rows. This causes a latency spike; consider extending warmup to cover this shape/config.
|
| 343 |
+
(APIServer pid=305508) INFO 07-15 18:44:23 [loggers.py:273] Engine 000: Avg prompt throughput: 1848.7 tokens/s, Avg generation throughput: 2142.8 tokens/s, Running: 253 reqs, Waiting: 1038 reqs, GPU KV cache usage: 31.1%, Prefix cache hit rate: 94.1%
|
| 344 |
+
(APIServer pid=305508) INFO 07-15 18:44:33 [loggers.py:273] Engine 000: Avg prompt throughput: 796.1 tokens/s, Avg generation throughput: 2956.5 tokens/s, Running: 256 reqs, Waiting: 915 reqs, GPU KV cache usage: 43.0%, Prefix cache hit rate: 94.3%
|
| 345 |
+
(APIServer pid=305508) INFO 07-15 18:44:43 [loggers.py:273] Engine 000: Avg prompt throughput: 452.7 tokens/s, Avg generation throughput: 2936.8 tokens/s, Running: 256 reqs, Waiting: 850 reqs, GPU KV cache usage: 58.3%, Prefix cache hit rate: 94.2%
|
| 346 |
+
(APIServer pid=305508) INFO 07-15 18:44:53 [loggers.py:273] Engine 000: Avg prompt throughput: 219.0 tokens/s, Avg generation throughput: 2863.2 tokens/s, Running: 255 reqs, Waiting: 815 reqs, GPU KV cache usage: 75.4%, Prefix cache hit rate: 94.3%
|
| 347 |
+
(APIServer pid=305508) INFO 07-15 18:45:03 [loggers.py:273] Engine 000: Avg prompt throughput: 960.0 tokens/s, Avg generation throughput: 2749.6 tokens/s, Running: 256 reqs, Waiting: 665 reqs, GPU KV cache usage: 45.6%, Prefix cache hit rate: 94.3%
|
| 348 |
+
(APIServer pid=305508) INFO 07-15 18:45:13 [loggers.py:273] Engine 000: Avg prompt throughput: 647.5 tokens/s, Avg generation throughput: 2908.7 tokens/s, Running: 255 reqs, Waiting: 567 reqs, GPU KV cache usage: 49.9%, Prefix cache hit rate: 94.4%
|
| 349 |
+
(APIServer pid=305508) INFO 07-15 18:45:23 [loggers.py:273] Engine 000: Avg prompt throughput: 662.9 tokens/s, Avg generation throughput: 2908.2 tokens/s, Running: 255 reqs, Waiting: 470 reqs, GPU KV cache usage: 53.1%, Prefix cache hit rate: 94.4%
|
| 350 |
+
(APIServer pid=305508) INFO 07-15 18:45:33 [loggers.py:273] Engine 000: Avg prompt throughput: 589.0 tokens/s, Avg generation throughput: 2909.2 tokens/s, Running: 255 reqs, Waiting: 379 reqs, GPU KV cache usage: 57.6%, Prefix cache hit rate: 94.4%
|
| 351 |
+
(APIServer pid=305508) INFO 07-15 18:45:43 [loggers.py:273] Engine 000: Avg prompt throughput: 457.3 tokens/s, Avg generation throughput: 2860.3 tokens/s, Running: 255 reqs, Waiting: 314 reqs, GPU KV cache usage: 68.8%, Prefix cache hit rate: 94.4%
|
| 352 |
+
(APIServer pid=305508) INFO 07-15 18:45:53 [loggers.py:273] Engine 000: Avg prompt throughput: 707.3 tokens/s, Avg generation throughput: 2856.2 tokens/s, Running: 253 reqs, Waiting: 209 reqs, GPU KV cache usage: 58.3%, Prefix cache hit rate: 94.4%
|
| 353 |
+
(APIServer pid=305508) INFO 07-15 18:46:03 [loggers.py:273] Engine 000: Avg prompt throughput: 818.5 tokens/s, Avg generation throughput: 2854.9 tokens/s, Running: 255 reqs, Waiting: 87 reqs, GPU KV cache usage: 53.2%, Prefix cache hit rate: 94.4%
|
| 354 |
+
(APIServer pid=305508) INFO 07-15 18:46:14 [loggers.py:273] Engine 000: Avg prompt throughput: 592.2 tokens/s, Avg generation throughput: 2896.5 tokens/s, Running: 241 reqs, Waiting: 0 reqs, GPU KV cache usage: 54.5%, Prefix cache hit rate: 94.4%
|
| 355 |
+
(APIServer pid=305508) INFO 07-15 18:46:24 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 2555.5 tokens/s, Running: 144 reqs, Waiting: 0 reqs, GPU KV cache usage: 44.9%, Prefix cache hit rate: 94.4%
|
| 356 |
+
(APIServer pid=305508) INFO 07-15 18:46:34 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 2061.1 tokens/s, Running: 64 reqs, Waiting: 0 reqs, GPU KV cache usage: 25.6%, Prefix cache hit rate: 94.4%
|
| 357 |
+
(APIServer pid=305508) INFO: 127.0.0.1:45906 - "POST /v1/completions HTTP/1.1" 200 OK
|
| 358 |
+
(EngineCore pid=305625) INFO 07-15 18:46:40 [core.py:1214] [shutdown] EngineCore: trigger received signal=SIGTERM
|
| 359 |
+
(APIServer pid=305508) INFO 07-15 18:46:40 [launcher.py:100] [shutdown] API server: shutdown triggered
|
| 360 |
+
(APIServer pid=305508) INFO 07-15 18:46:40 [launcher.py:116] [shutdown] API server: stopping engine client mode=abort timeout=0s
|
| 361 |
+
(EngineCore pid=305625) INFO 07-15 18:46:40 [core.py:1333] [shutdown] EngineCore: start mode=abort timeout=0s
|
| 362 |
+
(EngineCore pid=305625) INFO 07-15 18:46:40 [core.py:1364] [shutdown] EngineCore: request processing complete; starting resource teardown
|
| 363 |
+
(EngineCore pid=305625) INFO 07-15 18:46:40 [core.py:1227] [shutdown] EngineCore: exiting busy loop
|
| 364 |
+
(APIServer pid=305508) INFO 07-15 18:46:40 [core_client.py:655] [shutdown] MPClient: start timeout=0s
|
| 365 |
+
(APIServer pid=305508) INFO 07-15 18:46:40 [core_client.py:657] [shutdown] MPClient: stopping engine manager
|
| 366 |
+
(APIServer pid=305508) WARNING 07-15 18:46:40 [utils.py:626] [shutdown] Process manager: force killing remaining processes count=1
|
| 367 |
+
(APIServer pid=305508) INFO 07-15 18:46:40 [core_client.py:659] [shutdown] MPClient: engine manager stopped
|
| 368 |
+
(APIServer pid=305508) INFO 07-15 18:46:40 [core_client.py:660] [shutdown] MPClient: cleaning up background resources
|
| 369 |
+
(APIServer pid=305508) INFO 07-15 18:46:40 [core_client.py:662] [shutdown] MPClient: complete
|
| 370 |
+
(APIServer pid=305508) INFO 07-15 18:46:40 [launcher.py:125] [shutdown] API server: engine client stopped
|
| 371 |
+
(APIServer pid=305508) INFO 07-15 18:46:40 [launcher.py:128] [shutdown] API server: signalling HTTP server shutdown
|
| 372 |
+
(APIServer pid=305508) INFO 07-15 18:46:40 [launcher.py:149] [shutdown] API server: shutting down FastAPI HTTP server
|
| 373 |
+
(APIServer pid=305508) INFO: Shutting down
|
| 374 |
+
(APIServer pid=305508) INFO: Waiting for application shutdown.
|
| 375 |
+
(APIServer pid=305508) INFO: Application shutdown complete.
|
| 376 |
+
/home/henry/.local/share/uv/python/cpython-3.12.12-linux-x86_64-gnu/lib/python3.12/multiprocessing/resource_tracker.py:279: UserWarning: resource_tracker: There appear to be 1 leaked semaphore objects to clean up at shutdown
|
| 377 |
+
warnings.warn('resource_tracker: There appear to be %d '
|
evals/healing_breadth/glean_math_keep25_seed1224_step0050_chat.json
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
evals/healing_breadth/glean_math_keep25_seed1224_step0050_chat.json.server.log
ADDED
|
@@ -0,0 +1,138 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
(APIServer pid=308068) INFO 07-15 18:54:12 [api_utils.py:339]
|
| 2 |
+
(APIServer pid=308068) INFO 07-15 18:54:12 [api_utils.py:339] β β ββ ββ
|
| 3 |
+
(APIServer pid=308068) INFO 07-15 18:54:12 [api_utils.py:339] ββ ββ β β β βββ β version 0.25.0
|
| 4 |
+
(APIServer pid=308068) INFO 07-15 18:54:12 [api_utils.py:339] ββββ β β β β model outputs/healed/healing_breadth/glean_math_keep25_seed1224/step0050
|
| 5 |
+
(APIServer pid=308068) INFO 07-15 18:54:12 [api_utils.py:339] ββ βββββ βββββ β β
|
| 6 |
+
(APIServer pid=308068) INFO 07-15 18:54:12 [api_utils.py:339]
|
| 7 |
+
(APIServer pid=308068) INFO 07-15 18:54:12 [api_utils.py:273] non-default args: {'model_tag': 'outputs/healed/healing_breadth/glean_math_keep25_seed1224/step0050', 'host': '127.0.0.1', 'port': 8377, 'model': 'outputs/healed/healing_breadth/glean_math_keep25_seed1224/step0050', 'max_model_len': 2048, 'enforce_eager': True, 'served_model_name': ['student'], 'gpu_memory_utilization': 0.85}
|
| 8 |
+
(APIServer pid=308068) INFO 07-15 18:54:12 [model.py:619] Resolved architecture: PrunedOlmoeForCausalLM
|
| 9 |
+
(APIServer pid=308068) INFO 07-15 18:54:12 [model.py:1776] Using max model len 2048
|
| 10 |
+
(APIServer pid=308068) INFO 07-15 18:54:12 [vllm.py:1042] Asynchronous scheduling is enabled.
|
| 11 |
+
(APIServer pid=308068) WARNING 07-15 18:54:12 [vllm.py:1096] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
|
| 12 |
+
(APIServer pid=308068) WARNING 07-15 18:54:12 [vllm.py:1144] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
|
| 13 |
+
(APIServer pid=308068) INFO 07-15 18:54:12 [kernel.py:292] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native'])
|
| 14 |
+
(APIServer pid=308068) INFO 07-15 18:54:12 [vllm.py:1322] Cudagraph is disabled under eager mode
|
| 15 |
+
(APIServer pid=308068) INFO 07-15 18:54:12 [compilation.py:312] Enabled custom fusions: norm_quant, act_quant
|
| 16 |
+
(EngineCore pid=308191) INFO 07-15 18:54:19 [core.py:114] Initializing a V1 LLM engine (v0.25.0) with config: model='outputs/healed/healing_breadth/glean_math_keep25_seed1224/step0050', speculative_config=None, tokenizer='outputs/healed/healing_breadth/glean_math_keep25_seed1224/step0050', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=False, dtype=torch.bfloat16, max_seq_len=2048, download_dir=None, load_format=auto, tensor_parallel_size=1, pipeline_parallel_size=1, data_parallel_size=1, decode_context_parallel_size=1, dcp_comm_backend=ag_rs, disable_custom_all_reduce=False, quantization=None, quantization_config=None, enforce_eager=True, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False, jit_monitor_mode='warn', jit_monitor_verbose=False), seed=0, served_model_name=student, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'mode': <CompilationMode.NONE: 0>, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'ir_enable_torch_wrap': False, 'splitting_ops': [], 'compile_mm_encoder': False, 'cudagraph_mm_encoder': False, 'encoder_cudagraph_token_budgets': [], 'encoder_cudagraph_max_vision_items_per_batch': 0, 'encoder_cudagraph_max_frames_per_batch': None, 'compile_sizes': [], 'compile_ranges_endpoints': [2048], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'size_asserts': False, 'alignment_asserts': False, 'scalar_asserts': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': <CUDAGraphMode.NONE: 0>, 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': True, 'fuse_act_quant': True, 'fuse_attn_quant': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False, 'fuse_rope_kvcache_cat_mla': False, 'fuse_act_padding': False}, 'max_cudagraph_capture_size': 0, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': False, 'static_all_moe_layers': []}, kernel_config=KernelConfig(ir_op_priority=IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native']), enable_flashinfer_autotune=True, enable_cutedsl_warmup=True, moe_backend='auto', linear_backend='auto')
|
| 17 |
+
(EngineCore pid=308191) INFO 07-15 18:54:20 [parallel_state.py:1607] world_size=1 rank=0 local_rank=0 distributed_init_method=tcp://192.168.0.15:60767 backend=nccl
|
| 18 |
+
(EngineCore pid=308191) INFO 07-15 18:54:20 [parallel_state.py:1942] rank 0 in world size 1 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank 0, EPLB rank N/A
|
| 19 |
+
(EngineCore pid=308191) INFO 07-15 18:54:20 [topk_topp_sampler.py:55] Using FlashInfer for top-p & top-k sampling.
|
| 20 |
+
(EngineCore pid=308191) INFO 07-15 18:54:20 [gpu_model_runner.py:5209] Starting to load model outputs/healed/healing_breadth/glean_math_keep25_seed1224/step0050...
|
| 21 |
+
(EngineCore pid=308191) INFO 07-15 18:54:21 [cuda.py:476] Using FLASH_ATTN attention backend out of potential backends: ['FLASH_ATTN', 'FLASHINFER', 'TRITON_ATTN', 'FLEX_ATTENTION'].
|
| 22 |
+
(EngineCore pid=308191) INFO 07-15 18:54:21 [flash_attn.py:718] Using FlashAttention version 2
|
| 23 |
+
(EngineCore pid=308191) /home/henry/.cache/glean/megablocks-variable-93a1479bc15b/megablocks/grouped_gemm_util.py:10: UserWarning: Grouped GEMM not available.
|
| 24 |
+
(EngineCore pid=308191) warnings.warn('Grouped GEMM not available.')
|
| 25 |
+
(EngineCore pid=308191) INFO 07-15 18:54:21 [weight_utils.py:849] Filesystem type for checkpoints: EXT4. Checkpoint size: 3.89 GiB. Available RAM: 109.29 GiB.
|
| 26 |
+
(EngineCore pid=308191) INFO 07-15 18:54:21 [weight_utils.py:872] Auto-prefetch is disabled because the filesystem (EXT4) is not a recognized network FS (NFS/Lustre). If you want to force prefetching, start vLLM with --safetensors-load-strategy=prefetch.
|
| 27 |
+
(EngineCore pid=308191)
|
| 28 |
+
(EngineCore pid=308191)
|
| 29 |
+
(EngineCore pid=308191)
|
| 30 |
+
(EngineCore pid=308191)
|
| 31 |
+
(EngineCore pid=308191) INFO 07-15 18:54:24 [default_loader.py:430] Loading weights took 2.72 seconds
|
| 32 |
+
(EngineCore pid=308191) INFO 07-15 18:54:24 [gpu_model_runner.py:5306] Model loading took 3.89 GiB memory and 2.901110 seconds
|
| 33 |
+
(EngineCore pid=308191) INFO 07-15 18:54:26 [gpu_worker.py:538] Available KV cache memory: 15.82 GiB
|
| 34 |
+
(EngineCore pid=308191) INFO 07-15 18:54:26 [kv_cache_utils.py:2146] GPU KV cache size: 129,584 tokens
|
| 35 |
+
(EngineCore pid=308191) INFO 07-15 18:54:26 [kv_cache_utils.py:2147] Maximum concurrency for 2,048 tokens per request: 63.27x
|
| 36 |
+
(EngineCore pid=308191) INFO 07-15 18:54:26 [cutedsl_warmup.py:97] Skipping CuTeDSL warmup because no compile units were requested.
|
| 37 |
+
(EngineCore pid=308191) INFO 07-15 18:54:26 [jit_monitor.py:73] Kernel JIT monitor activated; monitored JIT compilations during inference will use mode=warn.
|
| 38 |
+
(EngineCore pid=308191) INFO 07-15 18:54:26 [core.py:344] init engine (profile, create kv cache, warmup model) took 2.06 s
|
| 39 |
+
(EngineCore pid=308191) INFO 07-15 18:54:26 [vllm.py:1042] Asynchronous scheduling is enabled.
|
| 40 |
+
(EngineCore pid=308191) WARNING 07-15 18:54:26 [vllm.py:1096] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
|
| 41 |
+
(EngineCore pid=308191) WARNING 07-15 18:54:26 [vllm.py:1144] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
|
| 42 |
+
(EngineCore pid=308191) INFO 07-15 18:54:26 [kernel.py:292] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native'])
|
| 43 |
+
(EngineCore pid=308191) INFO 07-15 18:54:26 [vllm.py:1322] Cudagraph is disabled under eager mode
|
| 44 |
+
(EngineCore pid=308191) INFO 07-15 18:54:26 [compilation.py:312] Enabled custom fusions: norm_quant, act_quant
|
| 45 |
+
(APIServer pid=308068) INFO 07-15 18:54:26 [api_server.py:612] Supported tasks: ['generate']
|
| 46 |
+
(APIServer pid=308068) WARNING 07-15 18:54:26 [__init__.py:36] SECURITY WARNING: Development endpoints are enabled! This should NOT be used in production!
|
| 47 |
+
(APIServer pid=308068) INFO 07-15 18:54:27 [hf.py:548] Detected the chat template content format to be 'string'. You can set `--chat-template-content-format` to override this.
|
| 48 |
+
(APIServer pid=308068) INFO 07-15 18:54:27 [api_server.py:616] Starting vLLM server on http://127.0.0.1:8377
|
| 49 |
+
(APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:37] Available routes are:
|
| 50 |
+
(APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /openapi.json, Methods: HEAD, GET
|
| 51 |
+
(APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /docs, Methods: HEAD, GET
|
| 52 |
+
(APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /docs/oauth2-redirect, Methods: HEAD, GET
|
| 53 |
+
(APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /redoc, Methods: HEAD, GET
|
| 54 |
+
(APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /load, Methods: GET
|
| 55 |
+
(APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /version, Methods: GET
|
| 56 |
+
(APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /health, Methods: GET
|
| 57 |
+
(APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /metrics, Methods: GET
|
| 58 |
+
(APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /tokenize, Methods: POST
|
| 59 |
+
(APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /detokenize, Methods: POST
|
| 60 |
+
(APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /v1/models, Methods: GET
|
| 61 |
+
(APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /ping, Methods: GET
|
| 62 |
+
(APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /ping, Methods: POST
|
| 63 |
+
(APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /invocations, Methods: POST
|
| 64 |
+
(APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /reset_prefix_cache, Methods: POST
|
| 65 |
+
(APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /reset_mm_cache, Methods: POST
|
| 66 |
+
(APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /reset_encoder_cache, Methods: POST
|
| 67 |
+
(APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /pause, Methods: POST
|
| 68 |
+
(APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /resume, Methods: POST
|
| 69 |
+
(APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /is_paused, Methods: GET
|
| 70 |
+
(APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /init_weight_transfer_engine, Methods: POST
|
| 71 |
+
(APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /start_weight_update, Methods: POST
|
| 72 |
+
(APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /update_weights, Methods: POST
|
| 73 |
+
(APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /finish_weight_update, Methods: POST
|
| 74 |
+
(APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /get_world_size, Methods: GET
|
| 75 |
+
(APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /collective_rpc, Methods: POST
|
| 76 |
+
(APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /server_info, Methods: GET
|
| 77 |
+
(APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /sleep, Methods: POST
|
| 78 |
+
(APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /wake_up, Methods: POST
|
| 79 |
+
(APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /is_sleeping, Methods: GET
|
| 80 |
+
(APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /v1/chat/completions, Methods: POST
|
| 81 |
+
(APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /v1/chat/completions/batch, Methods: POST
|
| 82 |
+
(APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /v1/responses, Methods: POST
|
| 83 |
+
(APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /v1/responses/{response_id}, Methods: GET
|
| 84 |
+
(APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /v1/responses/{response_id}/cancel, Methods: POST
|
| 85 |
+
(APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /v1/completions, Methods: POST
|
| 86 |
+
(APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /v1/messages, Methods: POST
|
| 87 |
+
(APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /v1/messages/count_tokens, Methods: POST
|
| 88 |
+
(APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /generative_scoring, Methods: POST
|
| 89 |
+
(APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /scale_elastic_ep, Methods: POST
|
| 90 |
+
(APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /is_scaling_elastic_ep, Methods: POST
|
| 91 |
+
(APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /v1/chat/completions/render, Methods: POST
|
| 92 |
+
(APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /v1/completions/render, Methods: POST
|
| 93 |
+
(APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /v1/chat/completions/derender, Methods: POST
|
| 94 |
+
(APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /v1/completions/derender, Methods: POST
|
| 95 |
+
(APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /inference/v1/generate, Methods: POST
|
| 96 |
+
(APIServer pid=308068) INFO: Started server process [308068]
|
| 97 |
+
(APIServer pid=308068) INFO: Waiting for application startup.
|
| 98 |
+
(APIServer pid=308068) INFO: Application startup complete.
|
| 99 |
+
(APIServer pid=308068) INFO: 127.0.0.1:34220 - "GET /health HTTP/1.1" 200 OK
|
| 100 |
+
(EngineCore pid=308191) WARNING 07-15 18:54:28 [jit_monitor.py:129] Triton kernel JIT compilation during inference: _build_route_rows. This causes a latency spike; consider extending warmup to cover this shape/config.
|
| 101 |
+
(APIServer pid=308068) INFO 07-15 18:54:37 [loggers.py:273] Engine 000: Avg prompt throughput: 2059.2 tokens/s, Avg generation throughput: 2322.0 tokens/s, Running: 256 reqs, Waiting: 1057 reqs, GPU KV cache usage: 36.5%, Prefix cache hit rate: 93.1%
|
| 102 |
+
(APIServer pid=308068) INFO 07-15 18:54:47 [loggers.py:273] Engine 000: Avg prompt throughput: 316.3 tokens/s, Avg generation throughput: 2990.5 tokens/s, Running: 256 reqs, Waiting: 1017 reqs, GPU KV cache usage: 54.9%, Prefix cache hit rate: 93.1%
|
| 103 |
+
(APIServer pid=308068) INFO 07-15 18:54:57 [loggers.py:273] Engine 000: Avg prompt throughput: 377.3 tokens/s, Avg generation throughput: 2887.0 tokens/s, Running: 256 reqs, Waiting: 968 reqs, GPU KV cache usage: 67.2%, Prefix cache hit rate: 93.2%
|
| 104 |
+
(APIServer pid=308068) INFO 07-15 18:55:07 [loggers.py:273] Engine 000: Avg prompt throughput: 374.5 tokens/s, Avg generation throughput: 2810.4 tokens/s, Running: 254 reqs, Waiting: 919 reqs, GPU KV cache usage: 76.5%, Prefix cache hit rate: 93.2%
|
| 105 |
+
(APIServer pid=308068) INFO 07-15 18:55:17 [loggers.py:273] Engine 000: Avg prompt throughput: 1206.0 tokens/s, Avg generation throughput: 2749.2 tokens/s, Running: 256 reqs, Waiting: 766 reqs, GPU KV cache usage: 42.7%, Prefix cache hit rate: 93.3%
|
| 106 |
+
(APIServer pid=308068) INFO 07-15 18:55:27 [loggers.py:273] Engine 000: Avg prompt throughput: 472.2 tokens/s, Avg generation throughput: 2962.8 tokens/s, Running: 253 reqs, Waiting: 705 reqs, GPU KV cache usage: 51.4%, Prefix cache hit rate: 93.3%
|
| 107 |
+
(APIServer pid=308068) INFO 07-15 18:55:37 [loggers.py:273] Engine 000: Avg prompt throughput: 409.7 tokens/s, Avg generation throughput: 2912.6 tokens/s, Running: 254 reqs, Waiting: 652 reqs, GPU KV cache usage: 61.6%, Prefix cache hit rate: 93.3%
|
| 108 |
+
(APIServer pid=308068) INFO 07-15 18:55:47 [loggers.py:273] Engine 000: Avg prompt throughput: 655.1 tokens/s, Avg generation throughput: 2832.9 tokens/s, Running: 256 reqs, Waiting: 567 reqs, GPU KV cache usage: 62.0%, Prefix cache hit rate: 93.3%
|
| 109 |
+
(APIServer pid=308068) INFO 07-15 18:55:57 [loggers.py:273] Engine 000: Avg prompt throughput: 561.6 tokens/s, Avg generation throughput: 2860.4 tokens/s, Running: 255 reqs, Waiting: 499 reqs, GPU KV cache usage: 66.1%, Prefix cache hit rate: 93.3%
|
| 110 |
+
(APIServer pid=308068) INFO 07-15 18:56:07 [loggers.py:273] Engine 000: Avg prompt throughput: 781.7 tokens/s, Avg generation throughput: 2882.3 tokens/s, Running: 254 reqs, Waiting: 397 reqs, GPU KV cache usage: 54.7%, Prefix cache hit rate: 93.4%
|
| 111 |
+
(APIServer pid=308068) INFO 07-15 18:56:17 [loggers.py:273] Engine 000: Avg prompt throughput: 661.3 tokens/s, Avg generation throughput: 2884.0 tokens/s, Running: 256 reqs, Waiting: 319 reqs, GPU KV cache usage: 59.2%, Prefix cache hit rate: 93.3%
|
| 112 |
+
(APIServer pid=308068) INFO 07-15 18:56:27 [loggers.py:273] Engine 000: Avg prompt throughput: 591.5 tokens/s, Avg generation throughput: 2859.2 tokens/s, Running: 256 reqs, Waiting: 246 reqs, GPU KV cache usage: 61.7%, Prefix cache hit rate: 93.4%
|
| 113 |
+
(APIServer pid=308068) INFO 07-15 18:56:37 [loggers.py:273] Engine 000: Avg prompt throughput: 639.7 tokens/s, Avg generation throughput: 2884.4 tokens/s, Running: 255 reqs, Waiting: 163 reqs, GPU KV cache usage: 58.6%, Prefix cache hit rate: 93.4%
|
| 114 |
+
(APIServer pid=308068) INFO 07-15 18:56:47 [loggers.py:273] Engine 000: Avg prompt throughput: 662.0 tokens/s, Avg generation throughput: 2858.9 tokens/s, Running: 254 reqs, Waiting: 83 reqs, GPU KV cache usage: 57.5%, Prefix cache hit rate: 93.4%
|
| 115 |
+
(APIServer pid=308068) INFO 07-15 18:56:57 [loggers.py:273] Engine 000: Avg prompt throughput: 677.8 tokens/s, Avg generation throughput: 2908.3 tokens/s, Running: 253 reqs, Waiting: 0 reqs, GPU KV cache usage: 57.3%, Prefix cache hit rate: 93.4%
|
| 116 |
+
(APIServer pid=308068) INFO 07-15 18:57:07 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 2752.8 tokens/s, Running: 182 reqs, Waiting: 0 reqs, GPU KV cache usage: 53.3%, Prefix cache hit rate: 93.4%
|
| 117 |
+
(APIServer pid=308068) INFO 07-15 18:57:17 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 2179.0 tokens/s, Running: 74 reqs, Waiting: 0 reqs, GPU KV cache usage: 29.0%, Prefix cache hit rate: 93.4%
|
| 118 |
+
(APIServer pid=308068) INFO: 127.0.0.1:34236 - "POST /v1/completions HTTP/1.1" 200 OK
|
| 119 |
+
(EngineCore pid=308191) INFO 07-15 18:57:24 [core.py:1214] [shutdown] EngineCore: trigger received signal=SIGTERM
|
| 120 |
+
(APIServer pid=308068) INFO 07-15 18:57:24 [launcher.py:100] [shutdown] API server: shutdown triggered
|
| 121 |
+
(APIServer pid=308068) INFO 07-15 18:57:24 [launcher.py:116] [shutdown] API server: stopping engine client mode=abort timeout=0s
|
| 122 |
+
(EngineCore pid=308191) INFO 07-15 18:57:24 [core.py:1333] [shutdown] EngineCore: start mode=abort timeout=0s
|
| 123 |
+
(EngineCore pid=308191) INFO 07-15 18:57:24 [core.py:1364] [shutdown] EngineCore: request processing complete; starting resource teardown
|
| 124 |
+
(EngineCore pid=308191) INFO 07-15 18:57:24 [core.py:1227] [shutdown] EngineCore: exiting busy loop
|
| 125 |
+
(APIServer pid=308068) INFO 07-15 18:57:24 [core_client.py:655] [shutdown] MPClient: start timeout=0s
|
| 126 |
+
(APIServer pid=308068) INFO 07-15 18:57:24 [core_client.py:657] [shutdown] MPClient: stopping engine manager
|
| 127 |
+
(APIServer pid=308068) WARNING 07-15 18:57:24 [utils.py:626] [shutdown] Process manager: force killing remaining processes count=1
|
| 128 |
+
(APIServer pid=308068) INFO 07-15 18:57:24 [core_client.py:659] [shutdown] MPClient: engine manager stopped
|
| 129 |
+
(APIServer pid=308068) INFO 07-15 18:57:24 [core_client.py:660] [shutdown] MPClient: cleaning up background resources
|
| 130 |
+
(APIServer pid=308068) INFO 07-15 18:57:24 [core_client.py:662] [shutdown] MPClient: complete
|
| 131 |
+
(APIServer pid=308068) INFO 07-15 18:57:24 [launcher.py:125] [shutdown] API server: engine client stopped
|
| 132 |
+
(APIServer pid=308068) INFO 07-15 18:57:24 [launcher.py:128] [shutdown] API server: signalling HTTP server shutdown
|
| 133 |
+
(APIServer pid=308068) INFO 07-15 18:57:24 [launcher.py:149] [shutdown] API server: shutting down FastAPI HTTP server
|
| 134 |
+
(APIServer pid=308068) INFO: Shutting down
|
| 135 |
+
(APIServer pid=308068) INFO: Waiting for application shutdown.
|
| 136 |
+
(APIServer pid=308068) INFO: Application shutdown complete.
|
| 137 |
+
/home/henry/.local/share/uv/python/cpython-3.12.12-linux-x86_64-gnu/lib/python3.12/multiprocessing/resource_tracker.py:279: UserWarning: resource_tracker: There appear to be 1 leaked semaphore objects to clean up at shutdown
|
| 138 |
+
warnings.warn('resource_tracker: There appear to be %d '
|
evals/healing_breadth/glean_math_keep25_seed1224_step0050_raw.json
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
evals/healing_breadth/glean_math_keep25_seed1224_step0050_raw.json.server.log
ADDED
|
@@ -0,0 +1,260 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
(APIServer pid=304785) INFO 07-15 18:42:58 [api_utils.py:339]
|
| 2 |
+
(APIServer pid=304785) INFO 07-15 18:42:58 [api_utils.py:339] β β ββ ββ
|
| 3 |
+
(APIServer pid=304785) INFO 07-15 18:42:58 [api_utils.py:339] ββ ββ β β β βββ β version 0.25.0
|
| 4 |
+
(APIServer pid=304785) INFO 07-15 18:42:58 [api_utils.py:339] ββββ β β β β model outputs/healed/healing_breadth/glean_math_keep25_seed1224/step0050
|
| 5 |
+
(APIServer pid=304785) INFO 07-15 18:42:58 [api_utils.py:339] ββ βββββ βββββ β β
|
| 6 |
+
(APIServer pid=304785) INFO 07-15 18:42:58 [api_utils.py:339]
|
| 7 |
+
(APIServer pid=304785) INFO 07-15 18:42:58 [api_utils.py:273] non-default args: {'model_tag': 'outputs/healed/healing_breadth/glean_math_keep25_seed1224/step0050', 'host': '127.0.0.1', 'port': 8377, 'model': 'outputs/healed/healing_breadth/glean_math_keep25_seed1224/step0050', 'max_model_len': 2048, 'enforce_eager': True, 'served_model_name': ['student'], 'gpu_memory_utilization': 0.85}
|
| 8 |
+
(APIServer pid=304785) INFO 07-15 18:42:58 [model.py:619] Resolved architecture: PrunedOlmoeForCausalLM
|
| 9 |
+
(APIServer pid=304785) INFO 07-15 18:42:58 [model.py:1776] Using max model len 2048
|
| 10 |
+
(APIServer pid=304785) INFO 07-15 18:42:58 [vllm.py:1042] Asynchronous scheduling is enabled.
|
| 11 |
+
(APIServer pid=304785) WARNING 07-15 18:42:58 [vllm.py:1096] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
|
| 12 |
+
(APIServer pid=304785) WARNING 07-15 18:42:58 [vllm.py:1144] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
|
| 13 |
+
(APIServer pid=304785) INFO 07-15 18:42:58 [kernel.py:292] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native'])
|
| 14 |
+
(APIServer pid=304785) INFO 07-15 18:42:58 [vllm.py:1322] Cudagraph is disabled under eager mode
|
| 15 |
+
(APIServer pid=304785) INFO 07-15 18:42:58 [compilation.py:312] Enabled custom fusions: norm_quant, act_quant
|
| 16 |
+
(EngineCore pid=304916) INFO 07-15 18:43:05 [core.py:114] Initializing a V1 LLM engine (v0.25.0) with config: model='outputs/healed/healing_breadth/glean_math_keep25_seed1224/step0050', speculative_config=None, tokenizer='outputs/healed/healing_breadth/glean_math_keep25_seed1224/step0050', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=False, dtype=torch.bfloat16, max_seq_len=2048, download_dir=None, load_format=auto, tensor_parallel_size=1, pipeline_parallel_size=1, data_parallel_size=1, decode_context_parallel_size=1, dcp_comm_backend=ag_rs, disable_custom_all_reduce=False, quantization=None, quantization_config=None, enforce_eager=True, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False, jit_monitor_mode='warn', jit_monitor_verbose=False), seed=0, served_model_name=student, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'mode': <CompilationMode.NONE: 0>, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'ir_enable_torch_wrap': False, 'splitting_ops': [], 'compile_mm_encoder': False, 'cudagraph_mm_encoder': False, 'encoder_cudagraph_token_budgets': [], 'encoder_cudagraph_max_vision_items_per_batch': 0, 'encoder_cudagraph_max_frames_per_batch': None, 'compile_sizes': [], 'compile_ranges_endpoints': [2048], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'size_asserts': False, 'alignment_asserts': False, 'scalar_asserts': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': <CUDAGraphMode.NONE: 0>, 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': True, 'fuse_act_quant': True, 'fuse_attn_quant': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False, 'fuse_rope_kvcache_cat_mla': False, 'fuse_act_padding': False}, 'max_cudagraph_capture_size': 0, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': False, 'static_all_moe_layers': []}, kernel_config=KernelConfig(ir_op_priority=IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native']), enable_flashinfer_autotune=True, enable_cutedsl_warmup=True, moe_backend='auto', linear_backend='auto')
|
| 17 |
+
(EngineCore pid=304916) INFO 07-15 18:43:05 [parallel_state.py:1607] world_size=1 rank=0 local_rank=0 distributed_init_method=tcp://192.168.0.15:45135 backend=nccl
|
| 18 |
+
(EngineCore pid=304916) INFO 07-15 18:43:06 [parallel_state.py:1942] rank 0 in world size 1 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank 0, EPLB rank N/A
|
| 19 |
+
(EngineCore pid=304916) INFO 07-15 18:43:06 [topk_topp_sampler.py:55] Using FlashInfer for top-p & top-k sampling.
|
| 20 |
+
(EngineCore pid=304916) INFO 07-15 18:43:06 [gpu_model_runner.py:5209] Starting to load model outputs/healed/healing_breadth/glean_math_keep25_seed1224/step0050...
|
| 21 |
+
(EngineCore pid=304916) INFO 07-15 18:43:07 [cuda.py:476] Using FLASH_ATTN attention backend out of potential backends: ['FLASH_ATTN', 'FLASHINFER', 'TRITON_ATTN', 'FLEX_ATTENTION'].
|
| 22 |
+
(EngineCore pid=304916) INFO 07-15 18:43:07 [flash_attn.py:718] Using FlashAttention version 2
|
| 23 |
+
(EngineCore pid=304916) /home/henry/.cache/glean/megablocks-variable-93a1479bc15b/megablocks/grouped_gemm_util.py:10: UserWarning: Grouped GEMM not available.
|
| 24 |
+
(EngineCore pid=304916) warnings.warn('Grouped GEMM not available.')
|
| 25 |
+
(EngineCore pid=304916) INFO 07-15 18:43:07 [weight_utils.py:849] Filesystem type for checkpoints: EXT4. Checkpoint size: 3.89 GiB. Available RAM: 109.76 GiB.
|
| 26 |
+
(EngineCore pid=304916) INFO 07-15 18:43:07 [weight_utils.py:872] Auto-prefetch is disabled because the filesystem (EXT4) is not a recognized network FS (NFS/Lustre). If you want to force prefetching, start vLLM with --safetensors-load-strategy=prefetch.
|
| 27 |
+
(EngineCore pid=304916)
|
| 28 |
+
(EngineCore pid=304916)
|
| 29 |
+
(EngineCore pid=304916)
|
| 30 |
+
(EngineCore pid=304916)
|
| 31 |
+
(EngineCore pid=304916) INFO 07-15 18:43:09 [default_loader.py:430] Loading weights took 2.61 seconds
|
| 32 |
+
(EngineCore pid=304916) INFO 07-15 18:43:10 [gpu_model_runner.py:5306] Model loading took 3.89 GiB memory and 2.790794 seconds
|
| 33 |
+
(EngineCore pid=304916) INFO 07-15 18:43:11 [gpu_worker.py:538] Available KV cache memory: 15.82 GiB
|
| 34 |
+
(EngineCore pid=304916) INFO 07-15 18:43:11 [kv_cache_utils.py:2146] GPU KV cache size: 129,584 tokens
|
| 35 |
+
(EngineCore pid=304916) INFO 07-15 18:43:11 [kv_cache_utils.py:2147] Maximum concurrency for 2,048 tokens per request: 63.27x
|
| 36 |
+
(EngineCore pid=304916) INFO 07-15 18:43:12 [cutedsl_warmup.py:97] Skipping CuTeDSL warmup because no compile units were requested.
|
| 37 |
+
(EngineCore pid=304916) INFO 07-15 18:43:12 [jit_monitor.py:73] Kernel JIT monitor activated; monitored JIT compilations during inference will use mode=warn.
|
| 38 |
+
(EngineCore pid=304916) INFO 07-15 18:43:12 [core.py:344] init engine (profile, create kv cache, warmup model) took 2.08 s
|
| 39 |
+
(EngineCore pid=304916) INFO 07-15 18:43:12 [vllm.py:1042] Asynchronous scheduling is enabled.
|
| 40 |
+
(EngineCore pid=304916) WARNING 07-15 18:43:12 [vllm.py:1096] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
|
| 41 |
+
(EngineCore pid=304916) WARNING 07-15 18:43:12 [vllm.py:1144] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
|
| 42 |
+
(EngineCore pid=304916) INFO 07-15 18:43:12 [kernel.py:292] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native'])
|
| 43 |
+
(EngineCore pid=304916) INFO 07-15 18:43:12 [vllm.py:1322] Cudagraph is disabled under eager mode
|
| 44 |
+
(EngineCore pid=304916) INFO 07-15 18:43:12 [compilation.py:312] Enabled custom fusions: norm_quant, act_quant
|
| 45 |
+
(APIServer pid=304785) INFO 07-15 18:43:12 [api_server.py:612] Supported tasks: ['generate']
|
| 46 |
+
(APIServer pid=304785) WARNING 07-15 18:43:12 [__init__.py:36] SECURITY WARNING: Development endpoints are enabled! This should NOT be used in production!
|
| 47 |
+
(APIServer pid=304785) INFO 07-15 18:43:12 [hf.py:548] Detected the chat template content format to be 'string'. You can set `--chat-template-content-format` to override this.
|
| 48 |
+
(APIServer pid=304785) INFO 07-15 18:43:12 [api_server.py:616] Starting vLLM server on http://127.0.0.1:8377
|
| 49 |
+
(APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:37] Available routes are:
|
| 50 |
+
(APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /openapi.json, Methods: HEAD, GET
|
| 51 |
+
(APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /docs, Methods: HEAD, GET
|
| 52 |
+
(APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /docs/oauth2-redirect, Methods: HEAD, GET
|
| 53 |
+
(APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /redoc, Methods: HEAD, GET
|
| 54 |
+
(APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /load, Methods: GET
|
| 55 |
+
(APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /version, Methods: GET
|
| 56 |
+
(APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /health, Methods: GET
|
| 57 |
+
(APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /metrics, Methods: GET
|
| 58 |
+
(APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /tokenize, Methods: POST
|
| 59 |
+
(APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /detokenize, Methods: POST
|
| 60 |
+
(APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /v1/models, Methods: GET
|
| 61 |
+
(APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /ping, Methods: GET
|
| 62 |
+
(APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /ping, Methods: POST
|
| 63 |
+
(APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /invocations, Methods: POST
|
| 64 |
+
(APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /reset_prefix_cache, Methods: POST
|
| 65 |
+
(APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /reset_mm_cache, Methods: POST
|
| 66 |
+
(APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /reset_encoder_cache, Methods: POST
|
| 67 |
+
(APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /pause, Methods: POST
|
| 68 |
+
(APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /resume, Methods: POST
|
| 69 |
+
(APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /is_paused, Methods: GET
|
| 70 |
+
(APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /init_weight_transfer_engine, Methods: POST
|
| 71 |
+
(APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /start_weight_update, Methods: POST
|
| 72 |
+
(APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /update_weights, Methods: POST
|
| 73 |
+
(APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /finish_weight_update, Methods: POST
|
| 74 |
+
(APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /get_world_size, Methods: GET
|
| 75 |
+
(APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /collective_rpc, Methods: POST
|
| 76 |
+
(APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /server_info, Methods: GET
|
| 77 |
+
(APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /sleep, Methods: POST
|
| 78 |
+
(APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /wake_up, Methods: POST
|
| 79 |
+
(APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /is_sleeping, Methods: GET
|
| 80 |
+
(APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /v1/chat/completions, Methods: POST
|
| 81 |
+
(APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /v1/chat/completions/batch, Methods: POST
|
| 82 |
+
(APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /v1/responses, Methods: POST
|
| 83 |
+
(APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /v1/responses/{response_id}, Methods: GET
|
| 84 |
+
(APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /v1/responses/{response_id}/cancel, Methods: POST
|
| 85 |
+
(APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /v1/completions, Methods: POST
|
| 86 |
+
(APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /v1/messages, Methods: POST
|
| 87 |
+
(APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /v1/messages/count_tokens, Methods: POST
|
| 88 |
+
(APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /generative_scoring, Methods: POST
|
| 89 |
+
(APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /scale_elastic_ep, Methods: POST
|
| 90 |
+
(APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /is_scaling_elastic_ep, Methods: POST
|
| 91 |
+
(APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /v1/chat/completions/render, Methods: POST
|
| 92 |
+
(APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /v1/completions/render, Methods: POST
|
| 93 |
+
(APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /v1/chat/completions/derender, Methods: POST
|
| 94 |
+
(APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /v1/completions/derender, Methods: POST
|
| 95 |
+
(APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /inference/v1/generate, Methods: POST
|
| 96 |
+
(APIServer pid=304785) INFO: Started server process [304785]
|
| 97 |
+
(APIServer pid=304785) INFO: Waiting for application startup.
|
| 98 |
+
(APIServer pid=304785) INFO: Application startup complete.
|
| 99 |
+
(APIServer pid=304785) INFO 07-15 18:43:22 [launcher.py:100] [shutdown] API server: shutdown triggered
|
| 100 |
+
(APIServer pid=304785) INFO 07-15 18:43:22 [launcher.py:116] [shutdown] API server: stopping engine client mode=abort timeout=0s
|
| 101 |
+
(APIServer pid=304785) INFO 07-15 18:43:22 [core_client.py:655] [shutdown] MPClient: start timeout=0s
|
| 102 |
+
(APIServer pid=304785) INFO 07-15 18:43:22 [core_client.py:657] [shutdown] MPClient: stopping engine manager
|
| 103 |
+
(APIServer pid=304785) WARNING 07-15 18:43:22 [utils.py:626] [shutdown] Process manager: force killing remaining processes count=1
|
| 104 |
+
(EngineCore pid=304916) INFO 07-15 18:43:22 [core.py:1214] [shutdown] EngineCore: trigger received signal=SIGTERM
|
| 105 |
+
(EngineCore pid=304916) INFO 07-15 18:43:22 [core.py:1333] [shutdown] EngineCore: start mode=abort timeout=0s
|
| 106 |
+
(EngineCore pid=304916) INFO 07-15 18:43:22 [core.py:1364] [shutdown] EngineCore: request processing complete; starting resource teardown
|
| 107 |
+
(EngineCore pid=304916) INFO 07-15 18:43:22 [core.py:1227] [shutdown] EngineCore: exiting busy loop
|
| 108 |
+
(APIServer pid=304785) INFO 07-15 18:43:22 [core_client.py:659] [shutdown] MPClient: engine manager stopped
|
| 109 |
+
(APIServer pid=304785) INFO 07-15 18:43:22 [core_client.py:660] [shutdown] MPClient: cleaning up background resources
|
| 110 |
+
(APIServer pid=304785) INFO 07-15 18:43:22 [core_client.py:662] [shutdown] MPClient: complete
|
| 111 |
+
(APIServer pid=304785) INFO 07-15 18:43:22 [launcher.py:125] [shutdown] API server: engine client stopped
|
| 112 |
+
(APIServer pid=304785) INFO 07-15 18:43:22 [launcher.py:128] [shutdown] API server: signalling HTTP server shutdown
|
| 113 |
+
(APIServer pid=304785) INFO 07-15 18:43:22 [launcher.py:149] [shutdown] API server: shutting down FastAPI HTTP server
|
| 114 |
+
(APIServer pid=304785) INFO: Shutting down
|
| 115 |
+
(APIServer pid=304785) INFO: Waiting for application shutdown.
|
| 116 |
+
(APIServer pid=304785) INFO: Application shutdown complete.
|
| 117 |
+
/home/henry/.local/share/uv/python/cpython-3.12.12-linux-x86_64-gnu/lib/python3.12/multiprocessing/resource_tracker.py:279: UserWarning: resource_tracker: There appear to be 1 leaked semaphore objects to clean up at shutdown
|
| 118 |
+
warnings.warn('resource_tracker: There appear to be %d '
|
| 119 |
+
(APIServer pid=307152) INFO 07-15 18:49:55 [api_utils.py:339]
|
| 120 |
+
(APIServer pid=307152) INFO 07-15 18:49:55 [api_utils.py:339] β β ββ ββ
|
| 121 |
+
(APIServer pid=307152) INFO 07-15 18:49:55 [api_utils.py:339] ββ ββ β β β βββ β version 0.25.0
|
| 122 |
+
(APIServer pid=307152) INFO 07-15 18:49:55 [api_utils.py:339] ββββ β β β β model outputs/healed/healing_breadth/glean_math_keep25_seed1224/step0050
|
| 123 |
+
(APIServer pid=307152) INFO 07-15 18:49:55 [api_utils.py:339] ββ βββββ βββββ β β
|
| 124 |
+
(APIServer pid=307152) INFO 07-15 18:49:55 [api_utils.py:339]
|
| 125 |
+
(APIServer pid=307152) INFO 07-15 18:49:55 [api_utils.py:273] non-default args: {'model_tag': 'outputs/healed/healing_breadth/glean_math_keep25_seed1224/step0050', 'host': '127.0.0.1', 'port': 8377, 'model': 'outputs/healed/healing_breadth/glean_math_keep25_seed1224/step0050', 'max_model_len': 2048, 'enforce_eager': True, 'served_model_name': ['student'], 'gpu_memory_utilization': 0.85}
|
| 126 |
+
(APIServer pid=307152) INFO 07-15 18:49:55 [model.py:619] Resolved architecture: PrunedOlmoeForCausalLM
|
| 127 |
+
(APIServer pid=307152) INFO 07-15 18:49:55 [model.py:1776] Using max model len 2048
|
| 128 |
+
(APIServer pid=307152) INFO 07-15 18:49:55 [vllm.py:1042] Asynchronous scheduling is enabled.
|
| 129 |
+
(APIServer pid=307152) WARNING 07-15 18:49:55 [vllm.py:1096] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
|
| 130 |
+
(APIServer pid=307152) WARNING 07-15 18:49:55 [vllm.py:1144] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
|
| 131 |
+
(APIServer pid=307152) INFO 07-15 18:49:55 [kernel.py:292] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native'])
|
| 132 |
+
(APIServer pid=307152) INFO 07-15 18:49:55 [vllm.py:1322] Cudagraph is disabled under eager mode
|
| 133 |
+
(APIServer pid=307152) INFO 07-15 18:49:55 [compilation.py:312] Enabled custom fusions: norm_quant, act_quant
|
| 134 |
+
(EngineCore pid=307273) INFO 07-15 18:50:02 [core.py:114] Initializing a V1 LLM engine (v0.25.0) with config: model='outputs/healed/healing_breadth/glean_math_keep25_seed1224/step0050', speculative_config=None, tokenizer='outputs/healed/healing_breadth/glean_math_keep25_seed1224/step0050', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=False, dtype=torch.bfloat16, max_seq_len=2048, download_dir=None, load_format=auto, tensor_parallel_size=1, pipeline_parallel_size=1, data_parallel_size=1, decode_context_parallel_size=1, dcp_comm_backend=ag_rs, disable_custom_all_reduce=False, quantization=None, quantization_config=None, enforce_eager=True, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False, jit_monitor_mode='warn', jit_monitor_verbose=False), seed=0, served_model_name=student, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'mode': <CompilationMode.NONE: 0>, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'ir_enable_torch_wrap': False, 'splitting_ops': [], 'compile_mm_encoder': False, 'cudagraph_mm_encoder': False, 'encoder_cudagraph_token_budgets': [], 'encoder_cudagraph_max_vision_items_per_batch': 0, 'encoder_cudagraph_max_frames_per_batch': None, 'compile_sizes': [], 'compile_ranges_endpoints': [2048], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'size_asserts': False, 'alignment_asserts': False, 'scalar_asserts': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': <CUDAGraphMode.NONE: 0>, 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': True, 'fuse_act_quant': True, 'fuse_attn_quant': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False, 'fuse_rope_kvcache_cat_mla': False, 'fuse_act_padding': False}, 'max_cudagraph_capture_size': 0, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': False, 'static_all_moe_layers': []}, kernel_config=KernelConfig(ir_op_priority=IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native']), enable_flashinfer_autotune=True, enable_cutedsl_warmup=True, moe_backend='auto', linear_backend='auto')
|
| 135 |
+
(EngineCore pid=307273) INFO 07-15 18:50:03 [parallel_state.py:1607] world_size=1 rank=0 local_rank=0 distributed_init_method=tcp://192.168.0.15:56251 backend=nccl
|
| 136 |
+
(EngineCore pid=307273) INFO 07-15 18:50:03 [parallel_state.py:1942] rank 0 in world size 1 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank 0, EPLB rank N/A
|
| 137 |
+
(EngineCore pid=307273) INFO 07-15 18:50:03 [topk_topp_sampler.py:55] Using FlashInfer for top-p & top-k sampling.
|
| 138 |
+
(EngineCore pid=307273) INFO 07-15 18:50:03 [gpu_model_runner.py:5209] Starting to load model outputs/healed/healing_breadth/glean_math_keep25_seed1224/step0050...
|
| 139 |
+
(EngineCore pid=307273) INFO 07-15 18:50:04 [cuda.py:476] Using FLASH_ATTN attention backend out of potential backends: ['FLASH_ATTN', 'FLASHINFER', 'TRITON_ATTN', 'FLEX_ATTENTION'].
|
| 140 |
+
(EngineCore pid=307273) INFO 07-15 18:50:04 [flash_attn.py:718] Using FlashAttention version 2
|
| 141 |
+
(EngineCore pid=307273) /home/henry/.cache/glean/megablocks-variable-93a1479bc15b/megablocks/grouped_gemm_util.py:10: UserWarning: Grouped GEMM not available.
|
| 142 |
+
(EngineCore pid=307273) warnings.warn('Grouped GEMM not available.')
|
| 143 |
+
(EngineCore pid=307273) INFO 07-15 18:50:04 [weight_utils.py:849] Filesystem type for checkpoints: EXT4. Checkpoint size: 3.89 GiB. Available RAM: 109.43 GiB.
|
| 144 |
+
(EngineCore pid=307273) INFO 07-15 18:50:04 [weight_utils.py:872] Auto-prefetch is disabled because the filesystem (EXT4) is not a recognized network FS (NFS/Lustre). If you want to force prefetching, start vLLM with --safetensors-load-strategy=prefetch.
|
| 145 |
+
(EngineCore pid=307273)
|
| 146 |
+
(EngineCore pid=307273)
|
| 147 |
+
(EngineCore pid=307273)
|
| 148 |
+
(EngineCore pid=307273)
|
| 149 |
+
(EngineCore pid=307273) INFO 07-15 18:50:06 [default_loader.py:430] Loading weights took 2.49 seconds
|
| 150 |
+
(EngineCore pid=307273) INFO 07-15 18:50:07 [gpu_model_runner.py:5306] Model loading took 3.89 GiB memory and 2.678073 seconds
|
| 151 |
+
(EngineCore pid=307273) INFO 07-15 18:50:08 [gpu_worker.py:538] Available KV cache memory: 15.82 GiB
|
| 152 |
+
(EngineCore pid=307273) INFO 07-15 18:50:08 [kv_cache_utils.py:2146] GPU KV cache size: 129,584 tokens
|
| 153 |
+
(EngineCore pid=307273) INFO 07-15 18:50:08 [kv_cache_utils.py:2147] Maximum concurrency for 2,048 tokens per request: 63.27x
|
| 154 |
+
(EngineCore pid=307273) INFO 07-15 18:50:09 [cutedsl_warmup.py:97] Skipping CuTeDSL warmup because no compile units were requested.
|
| 155 |
+
(EngineCore pid=307273) INFO 07-15 18:50:09 [jit_monitor.py:73] Kernel JIT monitor activated; monitored JIT compilations during inference will use mode=warn.
|
| 156 |
+
(EngineCore pid=307273) INFO 07-15 18:50:09 [core.py:344] init engine (profile, create kv cache, warmup model) took 2.07 s
|
| 157 |
+
(EngineCore pid=307273) INFO 07-15 18:50:09 [vllm.py:1042] Asynchronous scheduling is enabled.
|
| 158 |
+
(EngineCore pid=307273) WARNING 07-15 18:50:09 [vllm.py:1096] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
|
| 159 |
+
(EngineCore pid=307273) WARNING 07-15 18:50:09 [vllm.py:1144] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
|
| 160 |
+
(EngineCore pid=307273) INFO 07-15 18:50:09 [kernel.py:292] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native'])
|
| 161 |
+
(EngineCore pid=307273) INFO 07-15 18:50:09 [vllm.py:1322] Cudagraph is disabled under eager mode
|
| 162 |
+
(EngineCore pid=307273) INFO 07-15 18:50:09 [compilation.py:312] Enabled custom fusions: norm_quant, act_quant
|
| 163 |
+
(APIServer pid=307152) INFO 07-15 18:50:09 [api_server.py:612] Supported tasks: ['generate']
|
| 164 |
+
(APIServer pid=307152) WARNING 07-15 18:50:09 [__init__.py:36] SECURITY WARNING: Development endpoints are enabled! This should NOT be used in production!
|
| 165 |
+
(APIServer pid=307152) INFO 07-15 18:50:09 [hf.py:548] Detected the chat template content format to be 'string'. You can set `--chat-template-content-format` to override this.
|
| 166 |
+
(APIServer pid=307152) INFO 07-15 18:50:09 [api_server.py:616] Starting vLLM server on http://127.0.0.1:8377
|
| 167 |
+
(APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:37] Available routes are:
|
| 168 |
+
(APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /openapi.json, Methods: HEAD, GET
|
| 169 |
+
(APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /docs, Methods: HEAD, GET
|
| 170 |
+
(APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /docs/oauth2-redirect, Methods: HEAD, GET
|
| 171 |
+
(APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /redoc, Methods: HEAD, GET
|
| 172 |
+
(APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /load, Methods: GET
|
| 173 |
+
(APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /version, Methods: GET
|
| 174 |
+
(APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /health, Methods: GET
|
| 175 |
+
(APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /metrics, Methods: GET
|
| 176 |
+
(APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /tokenize, Methods: POST
|
| 177 |
+
(APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /detokenize, Methods: POST
|
| 178 |
+
(APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /v1/models, Methods: GET
|
| 179 |
+
(APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /ping, Methods: GET
|
| 180 |
+
(APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /ping, Methods: POST
|
| 181 |
+
(APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /invocations, Methods: POST
|
| 182 |
+
(APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /reset_prefix_cache, Methods: POST
|
| 183 |
+
(APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /reset_mm_cache, Methods: POST
|
| 184 |
+
(APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /reset_encoder_cache, Methods: POST
|
| 185 |
+
(APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /pause, Methods: POST
|
| 186 |
+
(APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /resume, Methods: POST
|
| 187 |
+
(APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /is_paused, Methods: GET
|
| 188 |
+
(APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /init_weight_transfer_engine, Methods: POST
|
| 189 |
+
(APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /start_weight_update, Methods: POST
|
| 190 |
+
(APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /update_weights, Methods: POST
|
| 191 |
+
(APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /finish_weight_update, Methods: POST
|
| 192 |
+
(APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /get_world_size, Methods: GET
|
| 193 |
+
(APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /collective_rpc, Methods: POST
|
| 194 |
+
(APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /server_info, Methods: GET
|
| 195 |
+
(APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /sleep, Methods: POST
|
| 196 |
+
(APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /wake_up, Methods: POST
|
| 197 |
+
(APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /is_sleeping, Methods: GET
|
| 198 |
+
(APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /v1/chat/completions, Methods: POST
|
| 199 |
+
(APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /v1/chat/completions/batch, Methods: POST
|
| 200 |
+
(APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /v1/responses, Methods: POST
|
| 201 |
+
(APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /v1/responses/{response_id}, Methods: GET
|
| 202 |
+
(APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /v1/responses/{response_id}/cancel, Methods: POST
|
| 203 |
+
(APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /v1/completions, Methods: POST
|
| 204 |
+
(APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /v1/messages, Methods: POST
|
| 205 |
+
(APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /v1/messages/count_tokens, Methods: POST
|
| 206 |
+
(APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /generative_scoring, Methods: POST
|
| 207 |
+
(APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /scale_elastic_ep, Methods: POST
|
| 208 |
+
(APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /is_scaling_elastic_ep, Methods: POST
|
| 209 |
+
(APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /v1/chat/completions/render, Methods: POST
|
| 210 |
+
(APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /v1/completions/render, Methods: POST
|
| 211 |
+
(APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /v1/chat/completions/derender, Methods: POST
|
| 212 |
+
(APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /v1/completions/derender, Methods: POST
|
| 213 |
+
(APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /inference/v1/generate, Methods: POST
|
| 214 |
+
(APIServer pid=307152) INFO: Started server process [307152]
|
| 215 |
+
(APIServer pid=307152) INFO: Waiting for application startup.
|
| 216 |
+
(APIServer pid=307152) INFO: Application startup complete.
|
| 217 |
+
(APIServer pid=307152) INFO: 127.0.0.1:37086 - "GET /health HTTP/1.1" 200 OK
|
| 218 |
+
(EngineCore pid=307273) WARNING 07-15 18:50:11 [jit_monitor.py:129] Triton kernel JIT compilation during inference: _build_route_rows. This causes a latency spike; consider extending warmup to cover this shape/config.
|
| 219 |
+
(APIServer pid=307152) INFO 07-15 18:50:20 [loggers.py:273] Engine 000: Avg prompt throughput: 1728.4 tokens/s, Avg generation throughput: 2268.5 tokens/s, Running: 256 reqs, Waiting: 1059 reqs, GPU KV cache usage: 33.5%, Prefix cache hit rate: 94.1%
|
| 220 |
+
(APIServer pid=307152) INFO 07-15 18:50:30 [loggers.py:273] Engine 000: Avg prompt throughput: 129.1 tokens/s, Avg generation throughput: 3043.4 tokens/s, Running: 256 reqs, Waiting: 1038 reqs, GPU KV cache usage: 54.3%, Prefix cache hit rate: 94.1%
|
| 221 |
+
(APIServer pid=307152) INFO 07-15 18:50:40 [loggers.py:273] Engine 000: Avg prompt throughput: 112.1 tokens/s, Avg generation throughput: 2916.1 tokens/s, Running: 256 reqs, Waiting: 1021 reqs, GPU KV cache usage: 73.4%, Prefix cache hit rate: 94.2%
|
| 222 |
+
(APIServer pid=307152) INFO 07-15 18:50:50 [loggers.py:273] Engine 000: Avg prompt throughput: 72.7 tokens/s, Avg generation throughput: 2788.6 tokens/s, Running: 256 reqs, Waiting: 1010 reqs, GPU KV cache usage: 92.2%, Prefix cache hit rate: 94.2%
|
| 223 |
+
(APIServer pid=307152) INFO 07-15 18:51:00 [loggers.py:273] Engine 000: Avg prompt throughput: 1368.9 tokens/s, Avg generation throughput: 2629.0 tokens/s, Running: 256 reqs, Waiting: 799 reqs, GPU KV cache usage: 30.0%, Prefix cache hit rate: 94.3%
|
| 224 |
+
(APIServer pid=307152) INFO 07-15 18:51:10 [loggers.py:273] Engine 000: Avg prompt throughput: 142.4 tokens/s, Avg generation throughput: 3069.3 tokens/s, Running: 256 reqs, Waiting: 779 reqs, GPU KV cache usage: 48.9%, Prefix cache hit rate: 94.3%
|
| 225 |
+
(APIServer pid=307152) INFO 07-15 18:51:20 [loggers.py:273] Engine 000: Avg prompt throughput: 174.0 tokens/s, Avg generation throughput: 2966.1 tokens/s, Running: 256 reqs, Waiting: 750 reqs, GPU KV cache usage: 63.9%, Prefix cache hit rate: 94.3%
|
| 226 |
+
(APIServer pid=307152) INFO 07-15 18:51:30 [loggers.py:273] Engine 000: Avg prompt throughput: 197.2 tokens/s, Avg generation throughput: 2838.7 tokens/s, Running: 256 reqs, Waiting: 721 reqs, GPU KV cache usage: 76.8%, Prefix cache hit rate: 94.3%
|
| 227 |
+
(APIServer pid=307152) INFO 07-15 18:51:40 [loggers.py:273] Engine 000: Avg prompt throughput: 134.4 tokens/s, Avg generation throughput: 2787.9 tokens/s, Running: 255 reqs, Waiting: 701 reqs, GPU KV cache usage: 91.1%, Prefix cache hit rate: 94.3%
|
| 228 |
+
(APIServer pid=307152) INFO 07-15 18:51:50 [loggers.py:273] Engine 000: Avg prompt throughput: 1103.7 tokens/s, Avg generation throughput: 2850.0 tokens/s, Running: 255 reqs, Waiting: 531 reqs, GPU KV cache usage: 47.4%, Prefix cache hit rate: 94.4%
|
| 229 |
+
(APIServer pid=307152) INFO 07-15 18:52:00 [loggers.py:273] Engine 000: Avg prompt throughput: 322.1 tokens/s, Avg generation throughput: 2939.5 tokens/s, Running: 256 reqs, Waiting: 487 reqs, GPU KV cache usage: 57.3%, Prefix cache hit rate: 94.3%
|
| 230 |
+
(APIServer pid=307152) INFO 07-15 18:52:10 [loggers.py:273] Engine 000: Avg prompt throughput: 242.7 tokens/s, Avg generation throughput: 2914.0 tokens/s, Running: 255 reqs, Waiting: 448 reqs, GPU KV cache usage: 66.5%, Prefix cache hit rate: 94.4%
|
| 231 |
+
(APIServer pid=307152) INFO 07-15 18:52:20 [loggers.py:273] Engine 000: Avg prompt throughput: 270.3 tokens/s, Avg generation throughput: 2836.3 tokens/s, Running: 256 reqs, Waiting: 403 reqs, GPU KV cache usage: 75.4%, Prefix cache hit rate: 94.4%
|
| 232 |
+
(APIServer pid=307152) INFO 07-15 18:52:30 [loggers.py:273] Engine 000: Avg prompt throughput: 1025.0 tokens/s, Avg generation throughput: 2647.8 tokens/s, Running: 256 reqs, Waiting: 257 reqs, GPU KV cache usage: 40.7%, Prefix cache hit rate: 94.4%
|
| 233 |
+
(APIServer pid=307152) INFO 07-15 18:52:40 [loggers.py:273] Engine 000: Avg prompt throughput: 196.8 tokens/s, Avg generation throughput: 2992.2 tokens/s, Running: 255 reqs, Waiting: 229 reqs, GPU KV cache usage: 55.4%, Prefix cache hit rate: 94.4%
|
| 234 |
+
(APIServer pid=307152) INFO 07-15 18:52:50 [loggers.py:273] Engine 000: Avg prompt throughput: 296.9 tokens/s, Avg generation throughput: 2938.7 tokens/s, Running: 256 reqs, Waiting: 182 reqs, GPU KV cache usage: 63.4%, Prefix cache hit rate: 94.4%
|
| 235 |
+
(APIServer pid=307152) INFO 07-15 18:53:00 [loggers.py:273] Engine 000: Avg prompt throughput: 268.7 tokens/s, Avg generation throughput: 2836.8 tokens/s, Running: 256 reqs, Waiting: 140 reqs, GPU KV cache usage: 71.0%, Prefix cache hit rate: 94.4%
|
| 236 |
+
(APIServer pid=307152) INFO 07-15 18:53:10 [loggers.py:273] Engine 000: Avg prompt throughput: 354.0 tokens/s, Avg generation throughput: 2836.6 tokens/s, Running: 255 reqs, Waiting: 90 reqs, GPU KV cache usage: 75.4%, Prefix cache hit rate: 94.4%
|
| 237 |
+
(APIServer pid=307152) INFO 07-15 18:53:20 [loggers.py:273] Engine 000: Avg prompt throughput: 616.3 tokens/s, Avg generation throughput: 2818.2 tokens/s, Running: 226 reqs, Waiting: 0 reqs, GPU KV cache usage: 51.0%, Prefix cache hit rate: 94.4%
|
| 238 |
+
(APIServer pid=307152) INFO 07-15 18:53:30 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 2701.6 tokens/s, Running: 164 reqs, Waiting: 0 reqs, GPU KV cache usage: 47.0%, Prefix cache hit rate: 94.4%
|
| 239 |
+
(APIServer pid=307152) INFO 07-15 18:53:40 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 2325.4 tokens/s, Running: 107 reqs, Waiting: 0 reqs, GPU KV cache usage: 40.3%, Prefix cache hit rate: 94.4%
|
| 240 |
+
(APIServer pid=307152) INFO: 127.0.0.1:37094 - "POST /v1/completions HTTP/1.1" 200 OK
|
| 241 |
+
(EngineCore pid=307273) INFO 07-15 18:53:47 [core.py:1214] [shutdown] EngineCore: trigger received signal=SIGTERM
|
| 242 |
+
(APIServer pid=307152) INFO 07-15 18:53:47 [launcher.py:100] [shutdown] API server: shutdown triggered
|
| 243 |
+
(APIServer pid=307152) INFO 07-15 18:53:47 [launcher.py:116] [shutdown] API server: stopping engine client mode=abort timeout=0s
|
| 244 |
+
(EngineCore pid=307273) INFO 07-15 18:53:47 [core.py:1333] [shutdown] EngineCore: start mode=abort timeout=0s
|
| 245 |
+
(EngineCore pid=307273) INFO 07-15 18:53:47 [core.py:1364] [shutdown] EngineCore: request processing complete; starting resource teardown
|
| 246 |
+
(EngineCore pid=307273) INFO 07-15 18:53:47 [core.py:1227] [shutdown] EngineCore: exiting busy loop
|
| 247 |
+
(APIServer pid=307152) INFO 07-15 18:53:47 [core_client.py:655] [shutdown] MPClient: start timeout=0s
|
| 248 |
+
(APIServer pid=307152) INFO 07-15 18:53:47 [core_client.py:657] [shutdown] MPClient: stopping engine manager
|
| 249 |
+
(APIServer pid=307152) WARNING 07-15 18:53:47 [utils.py:626] [shutdown] Process manager: force killing remaining processes count=1
|
| 250 |
+
(APIServer pid=307152) INFO 07-15 18:53:47 [core_client.py:659] [shutdown] MPClient: engine manager stopped
|
| 251 |
+
(APIServer pid=307152) INFO 07-15 18:53:47 [core_client.py:660] [shutdown] MPClient: cleaning up background resources
|
| 252 |
+
(APIServer pid=307152) INFO 07-15 18:53:47 [core_client.py:662] [shutdown] MPClient: complete
|
| 253 |
+
(APIServer pid=307152) INFO 07-15 18:53:47 [launcher.py:125] [shutdown] API server: engine client stopped
|
| 254 |
+
(APIServer pid=307152) INFO 07-15 18:53:47 [launcher.py:128] [shutdown] API server: signalling HTTP server shutdown
|
| 255 |
+
(APIServer pid=307152) INFO 07-15 18:53:47 [launcher.py:149] [shutdown] API server: shutting down FastAPI HTTP server
|
| 256 |
+
(APIServer pid=307152) INFO: Shutting down
|
| 257 |
+
(APIServer pid=307152) INFO: Waiting for application shutdown.
|
| 258 |
+
(APIServer pid=307152) INFO: Application shutdown complete.
|
| 259 |
+
/home/henry/.local/share/uv/python/cpython-3.12.12-linux-x86_64-gnu/lib/python3.12/multiprocessing/resource_tracker.py:279: UserWarning: resource_tracker: There appear to be 1 leaked semaphore objects to clean up at shutdown
|
| 260 |
+
warnings.warn('resource_tracker: There appear to be %d '
|
evals/healing_breadth/glean_math_keep75_oneshot_chat.json
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
evals/healing_breadth/glean_math_keep75_oneshot_chat.json.server.log
ADDED
|
@@ -0,0 +1,130 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
(APIServer pid=239542) INFO 07-15 06:24:31 [api_utils.py:339]
|
| 2 |
+
(APIServer pid=239542) INFO 07-15 06:24:31 [api_utils.py:339] β β ββ ββ
|
| 3 |
+
(APIServer pid=239542) INFO 07-15 06:24:31 [api_utils.py:339] ββ ββ β β β βββ β version 0.25.0
|
| 4 |
+
(APIServer pid=239542) INFO 07-15 06:24:31 [api_utils.py:339] ββββ β β β β model outputs/pruned/glean-0125inst-math-keep75
|
| 5 |
+
(APIServer pid=239542) INFO 07-15 06:24:31 [api_utils.py:339] ββ βββββ βββββ β β
|
| 6 |
+
(APIServer pid=239542) INFO 07-15 06:24:31 [api_utils.py:339]
|
| 7 |
+
(APIServer pid=239542) INFO 07-15 06:24:31 [api_utils.py:273] non-default args: {'model_tag': 'outputs/pruned/glean-0125inst-math-keep75', 'host': '127.0.0.1', 'port': 8377, 'model': 'outputs/pruned/glean-0125inst-math-keep75', 'max_model_len': 2048, 'enforce_eager': True, 'served_model_name': ['student'], 'gpu_memory_utilization': 0.85}
|
| 8 |
+
(APIServer pid=239542) INFO 07-15 06:24:31 [model.py:619] Resolved architecture: OlmoeForCausalLM
|
| 9 |
+
(APIServer pid=239542) INFO 07-15 06:24:31 [model.py:1776] Using max model len 2048
|
| 10 |
+
(APIServer pid=239542) INFO 07-15 06:24:31 [vllm.py:1042] Asynchronous scheduling is enabled.
|
| 11 |
+
(APIServer pid=239542) WARNING 07-15 06:24:31 [vllm.py:1096] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
|
| 12 |
+
(APIServer pid=239542) WARNING 07-15 06:24:31 [vllm.py:1144] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
|
| 13 |
+
(APIServer pid=239542) INFO 07-15 06:24:31 [kernel.py:292] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native'])
|
| 14 |
+
(APIServer pid=239542) INFO 07-15 06:24:32 [vllm.py:1322] Cudagraph is disabled under eager mode
|
| 15 |
+
(APIServer pid=239542) INFO 07-15 06:24:32 [compilation.py:312] Enabled custom fusions: norm_quant, act_quant
|
| 16 |
+
(EngineCore pid=239661) INFO 07-15 06:24:39 [core.py:114] Initializing a V1 LLM engine (v0.25.0) with config: model='outputs/pruned/glean-0125inst-math-keep75', speculative_config=None, tokenizer='outputs/pruned/glean-0125inst-math-keep75', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=False, dtype=torch.bfloat16, max_seq_len=2048, download_dir=None, load_format=auto, tensor_parallel_size=1, pipeline_parallel_size=1, data_parallel_size=1, decode_context_parallel_size=1, dcp_comm_backend=ag_rs, disable_custom_all_reduce=False, quantization=None, quantization_config=None, enforce_eager=True, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False, jit_monitor_mode='warn', jit_monitor_verbose=False), seed=0, served_model_name=student, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'mode': <CompilationMode.NONE: 0>, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'ir_enable_torch_wrap': False, 'splitting_ops': [], 'compile_mm_encoder': False, 'cudagraph_mm_encoder': False, 'encoder_cudagraph_token_budgets': [], 'encoder_cudagraph_max_vision_items_per_batch': 0, 'encoder_cudagraph_max_frames_per_batch': None, 'compile_sizes': [], 'compile_ranges_endpoints': [2048], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'size_asserts': False, 'alignment_asserts': False, 'scalar_asserts': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': <CUDAGraphMode.NONE: 0>, 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': True, 'fuse_act_quant': True, 'fuse_attn_quant': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False, 'fuse_rope_kvcache_cat_mla': False, 'fuse_act_padding': False}, 'max_cudagraph_capture_size': 0, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': False, 'static_all_moe_layers': []}, kernel_config=KernelConfig(ir_op_priority=IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native']), enable_flashinfer_autotune=True, enable_cutedsl_warmup=True, moe_backend='auto', linear_backend='auto')
|
| 17 |
+
(EngineCore pid=239661) INFO 07-15 06:24:39 [parallel_state.py:1607] world_size=1 rank=0 local_rank=0 distributed_init_method=tcp://192.168.0.15:38341 backend=nccl
|
| 18 |
+
(EngineCore pid=239661) INFO 07-15 06:24:39 [parallel_state.py:1942] rank 0 in world size 1 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank 0, EPLB rank N/A
|
| 19 |
+
(EngineCore pid=239661) INFO 07-15 06:24:40 [topk_topp_sampler.py:55] Using FlashInfer for top-p & top-k sampling.
|
| 20 |
+
(EngineCore pid=239661) INFO 07-15 06:24:40 [gpu_model_runner.py:5209] Starting to load model outputs/pruned/glean-0125inst-math-keep75...
|
| 21 |
+
(EngineCore pid=239661) INFO 07-15 06:24:41 [cuda.py:476] Using FLASH_ATTN attention backend out of potential backends: ['FLASH_ATTN', 'FLASHINFER', 'TRITON_ATTN', 'FLEX_ATTENTION'].
|
| 22 |
+
(EngineCore pid=239661) INFO 07-15 06:24:41 [flash_attn.py:718] Using FlashAttention version 2
|
| 23 |
+
(EngineCore pid=239661) /home/henry/.cache/glean/megablocks-variable-93a1479bc15b/megablocks/grouped_gemm_util.py:10: UserWarning: Grouped GEMM not available.
|
| 24 |
+
(EngineCore pid=239661) warnings.warn('Grouped GEMM not available.')
|
| 25 |
+
(EngineCore pid=239661) INFO 07-15 06:24:41 [weight_utils.py:849] Filesystem type for checkpoints: EXT4. Checkpoint size: 9.89 GiB. Available RAM: 61.30 GiB.
|
| 26 |
+
(EngineCore pid=239661) INFO 07-15 06:24:41 [weight_utils.py:872] Auto-prefetch is disabled because the filesystem (EXT4) is not a recognized network FS (NFS/Lustre). If you want to force prefetching, start vLLM with --safetensors-load-strategy=prefetch.
|
| 27 |
+
(EngineCore pid=239661)
|
| 28 |
+
(EngineCore pid=239661)
|
| 29 |
+
(EngineCore pid=239661)
|
| 30 |
+
(EngineCore pid=239661)
|
| 31 |
+
(EngineCore pid=239661)
|
| 32 |
+
(EngineCore pid=239661)
|
| 33 |
+
(EngineCore pid=239661) INFO 07-15 06:24:48 [default_loader.py:430] Loading weights took 7.29 seconds
|
| 34 |
+
(EngineCore pid=239661) INFO 07-15 06:24:48 [gpu_model_runner.py:5306] Model loading took 9.89 GiB memory and 7.475187 seconds
|
| 35 |
+
(EngineCore pid=239661) INFO 07-15 06:24:50 [gpu_worker.py:538] Available KV cache memory: 9.8 GiB
|
| 36 |
+
(EngineCore pid=239661) INFO 07-15 06:24:50 [kv_cache_utils.py:2146] GPU KV cache size: 80,272 tokens
|
| 37 |
+
(EngineCore pid=239661) INFO 07-15 06:24:50 [kv_cache_utils.py:2147] Maximum concurrency for 2,048 tokens per request: 39.20x
|
| 38 |
+
(EngineCore pid=239661) INFO 07-15 06:24:50 [cutedsl_warmup.py:97] Skipping CuTeDSL warmup because no compile units were requested.
|
| 39 |
+
(EngineCore pid=239661) INFO 07-15 06:24:50 [jit_monitor.py:73] Kernel JIT monitor activated; monitored JIT compilations during inference will use mode=warn.
|
| 40 |
+
(EngineCore pid=239661) INFO 07-15 06:24:50 [core.py:344] init engine (profile, create kv cache, warmup model) took 2.09 s
|
| 41 |
+
(EngineCore pid=239661) INFO 07-15 06:24:51 [vllm.py:1042] Asynchronous scheduling is enabled.
|
| 42 |
+
(EngineCore pid=239661) WARNING 07-15 06:24:51 [vllm.py:1096] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
|
| 43 |
+
(EngineCore pid=239661) WARNING 07-15 06:24:51 [vllm.py:1144] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
|
| 44 |
+
(EngineCore pid=239661) INFO 07-15 06:24:51 [kernel.py:292] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native'])
|
| 45 |
+
(EngineCore pid=239661) INFO 07-15 06:24:51 [vllm.py:1322] Cudagraph is disabled under eager mode
|
| 46 |
+
(EngineCore pid=239661) INFO 07-15 06:24:51 [compilation.py:312] Enabled custom fusions: norm_quant, act_quant
|
| 47 |
+
(APIServer pid=239542) INFO 07-15 06:24:51 [api_server.py:612] Supported tasks: ['generate']
|
| 48 |
+
(APIServer pid=239542) WARNING 07-15 06:24:51 [__init__.py:36] SECURITY WARNING: Development endpoints are enabled! This should NOT be used in production!
|
| 49 |
+
(APIServer pid=239542) INFO 07-15 06:24:51 [hf.py:548] Detected the chat template content format to be 'string'. You can set `--chat-template-content-format` to override this.
|
| 50 |
+
(APIServer pid=239542) INFO 07-15 06:24:51 [api_server.py:616] Starting vLLM server on http://127.0.0.1:8377
|
| 51 |
+
(APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:37] Available routes are:
|
| 52 |
+
(APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /openapi.json, Methods: HEAD, GET
|
| 53 |
+
(APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /docs, Methods: HEAD, GET
|
| 54 |
+
(APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /docs/oauth2-redirect, Methods: HEAD, GET
|
| 55 |
+
(APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /redoc, Methods: HEAD, GET
|
| 56 |
+
(APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /load, Methods: GET
|
| 57 |
+
(APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /version, Methods: GET
|
| 58 |
+
(APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /health, Methods: GET
|
| 59 |
+
(APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /metrics, Methods: GET
|
| 60 |
+
(APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /tokenize, Methods: POST
|
| 61 |
+
(APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /detokenize, Methods: POST
|
| 62 |
+
(APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /v1/models, Methods: GET
|
| 63 |
+
(APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /ping, Methods: GET
|
| 64 |
+
(APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /ping, Methods: POST
|
| 65 |
+
(APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /invocations, Methods: POST
|
| 66 |
+
(APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /reset_prefix_cache, Methods: POST
|
| 67 |
+
(APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /reset_mm_cache, Methods: POST
|
| 68 |
+
(APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /reset_encoder_cache, Methods: POST
|
| 69 |
+
(APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /pause, Methods: POST
|
| 70 |
+
(APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /resume, Methods: POST
|
| 71 |
+
(APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /is_paused, Methods: GET
|
| 72 |
+
(APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /init_weight_transfer_engine, Methods: POST
|
| 73 |
+
(APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /start_weight_update, Methods: POST
|
| 74 |
+
(APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /update_weights, Methods: POST
|
| 75 |
+
(APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /finish_weight_update, Methods: POST
|
| 76 |
+
(APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /get_world_size, Methods: GET
|
| 77 |
+
(APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /collective_rpc, Methods: POST
|
| 78 |
+
(APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /server_info, Methods: GET
|
| 79 |
+
(APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /sleep, Methods: POST
|
| 80 |
+
(APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /wake_up, Methods: POST
|
| 81 |
+
(APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /is_sleeping, Methods: GET
|
| 82 |
+
(APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /v1/chat/completions, Methods: POST
|
| 83 |
+
(APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /v1/chat/completions/batch, Methods: POST
|
| 84 |
+
(APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /v1/responses, Methods: POST
|
| 85 |
+
(APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /v1/responses/{response_id}, Methods: GET
|
| 86 |
+
(APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /v1/responses/{response_id}/cancel, Methods: POST
|
| 87 |
+
(APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /v1/completions, Methods: POST
|
| 88 |
+
(APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /v1/messages, Methods: POST
|
| 89 |
+
(APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /v1/messages/count_tokens, Methods: POST
|
| 90 |
+
(APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /generative_scoring, Methods: POST
|
| 91 |
+
(APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /scale_elastic_ep, Methods: POST
|
| 92 |
+
(APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /is_scaling_elastic_ep, Methods: POST
|
| 93 |
+
(APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /v1/chat/completions/render, Methods: POST
|
| 94 |
+
(APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /v1/completions/render, Methods: POST
|
| 95 |
+
(APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /v1/chat/completions/derender, Methods: POST
|
| 96 |
+
(APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /v1/completions/derender, Methods: POST
|
| 97 |
+
(APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /inference/v1/generate, Methods: POST
|
| 98 |
+
(APIServer pid=239542) INFO: Started server process [239542]
|
| 99 |
+
(APIServer pid=239542) INFO: Waiting for application startup.
|
| 100 |
+
(APIServer pid=239542) INFO: Application startup complete.
|
| 101 |
+
(APIServer pid=239542) INFO: 127.0.0.1:51610 - "GET /health HTTP/1.1" 200 OK
|
| 102 |
+
(EngineCore pid=239661) WARNING 07-15 06:24:52 [jit_monitor.py:129] Triton kernel JIT compilation during inference: _build_route_rows. This causes a latency spike; consider extending warmup to cover this shape/config.
|
| 103 |
+
(APIServer pid=239542) INFO 07-15 06:25:01 [loggers.py:273] Engine 000: Avg prompt throughput: 2621.7 tokens/s, Avg generation throughput: 1945.2 tokens/s, Running: 254 reqs, Waiting: 978 reqs, GPU KV cache usage: 48.5%, Prefix cache hit rate: 93.2%
|
| 104 |
+
(APIServer pid=239542) INFO 07-15 06:25:11 [loggers.py:273] Engine 000: Avg prompt throughput: 1793.6 tokens/s, Avg generation throughput: 2485.0 tokens/s, Running: 251 reqs, Waiting: 747 reqs, GPU KV cache usage: 49.0%, Prefix cache hit rate: 93.3%
|
| 105 |
+
(APIServer pid=239542) INFO 07-15 06:25:21 [loggers.py:273] Engine 000: Avg prompt throughput: 1816.8 tokens/s, Avg generation throughput: 2460.1 tokens/s, Running: 254 reqs, Waiting: 517 reqs, GPU KV cache usage: 50.1%, Prefix cache hit rate: 93.3%
|
| 106 |
+
(APIServer pid=239542) INFO 07-15 06:25:31 [loggers.py:273] Engine 000: Avg prompt throughput: 1842.7 tokens/s, Avg generation throughput: 2460.4 tokens/s, Running: 252 reqs, Waiting: 291 reqs, GPU KV cache usage: 50.3%, Prefix cache hit rate: 93.4%
|
| 107 |
+
(APIServer pid=239542) INFO 07-15 06:25:41 [loggers.py:273] Engine 000: Avg prompt throughput: 1720.0 tokens/s, Avg generation throughput: 2485.7 tokens/s, Running: 256 reqs, Waiting: 71 reqs, GPU KV cache usage: 52.5%, Prefix cache hit rate: 93.4%
|
| 108 |
+
(APIServer pid=239542) INFO 07-15 06:25:51 [loggers.py:273] Engine 000: Avg prompt throughput: 618.3 tokens/s, Avg generation throughput: 2265.6 tokens/s, Running: 64 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.5%, Prefix cache hit rate: 93.4%
|
| 109 |
+
(APIServer pid=239542) INFO 07-15 06:26:01 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 254.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.2%, Prefix cache hit rate: 93.4%
|
| 110 |
+
(APIServer pid=239542) INFO: 127.0.0.1:51612 - "POST /v1/completions HTTP/1.1" 200 OK
|
| 111 |
+
(EngineCore pid=239661) INFO 07-15 06:26:02 [core.py:1214] [shutdown] EngineCore: trigger received signal=SIGTERM
|
| 112 |
+
(APIServer pid=239542) INFO 07-15 06:26:02 [launcher.py:100] [shutdown] API server: shutdown triggered
|
| 113 |
+
(APIServer pid=239542) INFO 07-15 06:26:02 [launcher.py:116] [shutdown] API server: stopping engine client mode=abort timeout=0s
|
| 114 |
+
(EngineCore pid=239661) INFO 07-15 06:26:02 [core.py:1333] [shutdown] EngineCore: start mode=abort timeout=0s
|
| 115 |
+
(EngineCore pid=239661) INFO 07-15 06:26:02 [core.py:1364] [shutdown] EngineCore: request processing complete; starting resource teardown
|
| 116 |
+
(EngineCore pid=239661) INFO 07-15 06:26:02 [core.py:1227] [shutdown] EngineCore: exiting busy loop
|
| 117 |
+
(APIServer pid=239542) INFO 07-15 06:26:02 [core_client.py:655] [shutdown] MPClient: start timeout=0s
|
| 118 |
+
(APIServer pid=239542) INFO 07-15 06:26:02 [core_client.py:657] [shutdown] MPClient: stopping engine manager
|
| 119 |
+
(APIServer pid=239542) WARNING 07-15 06:26:02 [utils.py:626] [shutdown] Process manager: force killing remaining processes count=1
|
| 120 |
+
(APIServer pid=239542) INFO 07-15 06:26:02 [core_client.py:659] [shutdown] MPClient: engine manager stopped
|
| 121 |
+
(APIServer pid=239542) INFO 07-15 06:26:02 [core_client.py:660] [shutdown] MPClient: cleaning up background resources
|
| 122 |
+
(APIServer pid=239542) INFO 07-15 06:26:02 [core_client.py:662] [shutdown] MPClient: complete
|
| 123 |
+
(APIServer pid=239542) INFO 07-15 06:26:02 [launcher.py:125] [shutdown] API server: engine client stopped
|
| 124 |
+
(APIServer pid=239542) INFO 07-15 06:26:02 [launcher.py:128] [shutdown] API server: signalling HTTP server shutdown
|
| 125 |
+
(APIServer pid=239542) INFO 07-15 06:26:02 [launcher.py:149] [shutdown] API server: shutting down FastAPI HTTP server
|
| 126 |
+
(APIServer pid=239542) INFO: Shutting down
|
| 127 |
+
(APIServer pid=239542) INFO: Waiting for application shutdown.
|
| 128 |
+
(APIServer pid=239542) INFO: Application shutdown complete.
|
| 129 |
+
/home/henry/.local/share/uv/python/cpython-3.12.12-linux-x86_64-gnu/lib/python3.12/multiprocessing/resource_tracker.py:279: UserWarning: resource_tracker: There appear to be 1 leaked semaphore objects to clean up at shutdown
|
| 130 |
+
warnings.warn('resource_tracker: There appear to be %d '
|
evals/healing_breadth/glean_math_keep75_oneshot_raw.json
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
evals/healing_breadth/glean_math_keep75_oneshot_raw.json.server.log
ADDED
|
@@ -0,0 +1,129 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
(APIServer pid=238699) INFO 07-15 06:22:30 [api_utils.py:339]
|
| 2 |
+
(APIServer pid=238699) INFO 07-15 06:22:30 [api_utils.py:339] β β ββ ββ
|
| 3 |
+
(APIServer pid=238699) INFO 07-15 06:22:30 [api_utils.py:339] ββ ββ β β β βββ β version 0.25.0
|
| 4 |
+
(APIServer pid=238699) INFO 07-15 06:22:30 [api_utils.py:339] ββββ β β β β model outputs/pruned/glean-0125inst-math-keep75
|
| 5 |
+
(APIServer pid=238699) INFO 07-15 06:22:30 [api_utils.py:339] ββ βββββ βββββ β β
|
| 6 |
+
(APIServer pid=238699) INFO 07-15 06:22:30 [api_utils.py:339]
|
| 7 |
+
(APIServer pid=238699) INFO 07-15 06:22:30 [api_utils.py:273] non-default args: {'model_tag': 'outputs/pruned/glean-0125inst-math-keep75', 'host': '127.0.0.1', 'port': 8377, 'model': 'outputs/pruned/glean-0125inst-math-keep75', 'max_model_len': 2048, 'enforce_eager': True, 'served_model_name': ['student'], 'gpu_memory_utilization': 0.85}
|
| 8 |
+
(APIServer pid=238699) INFO 07-15 06:22:30 [model.py:619] Resolved architecture: OlmoeForCausalLM
|
| 9 |
+
(APIServer pid=238699) INFO 07-15 06:22:30 [model.py:1776] Using max model len 2048
|
| 10 |
+
(APIServer pid=238699) INFO 07-15 06:22:30 [vllm.py:1042] Asynchronous scheduling is enabled.
|
| 11 |
+
(APIServer pid=238699) WARNING 07-15 06:22:30 [vllm.py:1096] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
|
| 12 |
+
(APIServer pid=238699) WARNING 07-15 06:22:30 [vllm.py:1144] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
|
| 13 |
+
(APIServer pid=238699) INFO 07-15 06:22:30 [kernel.py:292] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native'])
|
| 14 |
+
(APIServer pid=238699) INFO 07-15 06:22:30 [vllm.py:1322] Cudagraph is disabled under eager mode
|
| 15 |
+
(APIServer pid=238699) INFO 07-15 06:22:30 [compilation.py:312] Enabled custom fusions: norm_quant, act_quant
|
| 16 |
+
(EngineCore pid=238829) INFO 07-15 06:22:38 [core.py:114] Initializing a V1 LLM engine (v0.25.0) with config: model='outputs/pruned/glean-0125inst-math-keep75', speculative_config=None, tokenizer='outputs/pruned/glean-0125inst-math-keep75', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=False, dtype=torch.bfloat16, max_seq_len=2048, download_dir=None, load_format=auto, tensor_parallel_size=1, pipeline_parallel_size=1, data_parallel_size=1, decode_context_parallel_size=1, dcp_comm_backend=ag_rs, disable_custom_all_reduce=False, quantization=None, quantization_config=None, enforce_eager=True, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False, jit_monitor_mode='warn', jit_monitor_verbose=False), seed=0, served_model_name=student, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'mode': <CompilationMode.NONE: 0>, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'ir_enable_torch_wrap': False, 'splitting_ops': [], 'compile_mm_encoder': False, 'cudagraph_mm_encoder': False, 'encoder_cudagraph_token_budgets': [], 'encoder_cudagraph_max_vision_items_per_batch': 0, 'encoder_cudagraph_max_frames_per_batch': None, 'compile_sizes': [], 'compile_ranges_endpoints': [2048], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'size_asserts': False, 'alignment_asserts': False, 'scalar_asserts': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': <CUDAGraphMode.NONE: 0>, 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': True, 'fuse_act_quant': True, 'fuse_attn_quant': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False, 'fuse_rope_kvcache_cat_mla': False, 'fuse_act_padding': False}, 'max_cudagraph_capture_size': 0, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': False, 'static_all_moe_layers': []}, kernel_config=KernelConfig(ir_op_priority=IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native']), enable_flashinfer_autotune=True, enable_cutedsl_warmup=True, moe_backend='auto', linear_backend='auto')
|
| 17 |
+
(EngineCore pid=238829) INFO 07-15 06:22:38 [parallel_state.py:1607] world_size=1 rank=0 local_rank=0 distributed_init_method=tcp://192.168.0.15:59451 backend=nccl
|
| 18 |
+
(EngineCore pid=238829) INFO 07-15 06:22:38 [parallel_state.py:1942] rank 0 in world size 1 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank 0, EPLB rank N/A
|
| 19 |
+
(EngineCore pid=238829) INFO 07-15 06:22:39 [topk_topp_sampler.py:55] Using FlashInfer for top-p & top-k sampling.
|
| 20 |
+
(EngineCore pid=238829) INFO 07-15 06:22:39 [gpu_model_runner.py:5209] Starting to load model outputs/pruned/glean-0125inst-math-keep75...
|
| 21 |
+
(EngineCore pid=238829) INFO 07-15 06:22:40 [cuda.py:476] Using FLASH_ATTN attention backend out of potential backends: ['FLASH_ATTN', 'FLASHINFER', 'TRITON_ATTN', 'FLEX_ATTENTION'].
|
| 22 |
+
(EngineCore pid=238829) INFO 07-15 06:22:40 [flash_attn.py:718] Using FlashAttention version 2
|
| 23 |
+
(EngineCore pid=238829) /home/henry/.cache/glean/megablocks-variable-93a1479bc15b/megablocks/grouped_gemm_util.py:10: UserWarning: Grouped GEMM not available.
|
| 24 |
+
(EngineCore pid=238829) warnings.warn('Grouped GEMM not available.')
|
| 25 |
+
(EngineCore pid=238829) INFO 07-15 06:22:40 [weight_utils.py:849] Filesystem type for checkpoints: EXT4. Checkpoint size: 9.89 GiB. Available RAM: 61.21 GiB.
|
| 26 |
+
(EngineCore pid=238829) INFO 07-15 06:22:40 [weight_utils.py:872] Auto-prefetch is disabled because the filesystem (EXT4) is not a recognized network FS (NFS/Lustre). If you want to force prefetching, start vLLM with --safetensors-load-strategy=prefetch.
|
| 27 |
+
(EngineCore pid=238829)
|
| 28 |
+
(EngineCore pid=238829)
|
| 29 |
+
(EngineCore pid=238829)
|
| 30 |
+
(EngineCore pid=238829)
|
| 31 |
+
(EngineCore pid=238829)
|
| 32 |
+
(EngineCore pid=238829)
|
| 33 |
+
(EngineCore pid=238829) INFO 07-15 06:22:55 [default_loader.py:430] Loading weights took 15.19 seconds
|
| 34 |
+
(EngineCore pid=238829) INFO 07-15 06:22:55 [gpu_model_runner.py:5306] Model loading took 9.89 GiB memory and 15.389861 seconds
|
| 35 |
+
(EngineCore pid=238829) INFO 07-15 06:22:57 [gpu_worker.py:538] Available KV cache memory: 9.8 GiB
|
| 36 |
+
(EngineCore pid=238829) INFO 07-15 06:22:57 [kv_cache_utils.py:2146] GPU KV cache size: 80,272 tokens
|
| 37 |
+
(EngineCore pid=238829) INFO 07-15 06:22:57 [kv_cache_utils.py:2147] Maximum concurrency for 2,048 tokens per request: 39.20x
|
| 38 |
+
(EngineCore pid=238829) INFO 07-15 06:22:57 [cutedsl_warmup.py:97] Skipping CuTeDSL warmup because no compile units were requested.
|
| 39 |
+
(EngineCore pid=238829) INFO 07-15 06:22:57 [jit_monitor.py:73] Kernel JIT monitor activated; monitored JIT compilations during inference will use mode=warn.
|
| 40 |
+
(EngineCore pid=238829) INFO 07-15 06:22:57 [core.py:344] init engine (profile, create kv cache, warmup model) took 2.13 s
|
| 41 |
+
(EngineCore pid=238829) INFO 07-15 06:22:58 [vllm.py:1042] Asynchronous scheduling is enabled.
|
| 42 |
+
(EngineCore pid=238829) WARNING 07-15 06:22:58 [vllm.py:1096] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
|
| 43 |
+
(EngineCore pid=238829) WARNING 07-15 06:22:58 [vllm.py:1144] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
|
| 44 |
+
(EngineCore pid=238829) INFO 07-15 06:22:58 [kernel.py:292] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native'])
|
| 45 |
+
(EngineCore pid=238829) INFO 07-15 06:22:58 [vllm.py:1322] Cudagraph is disabled under eager mode
|
| 46 |
+
(EngineCore pid=238829) INFO 07-15 06:22:58 [compilation.py:312] Enabled custom fusions: norm_quant, act_quant
|
| 47 |
+
(APIServer pid=238699) INFO 07-15 06:22:58 [api_server.py:612] Supported tasks: ['generate']
|
| 48 |
+
(APIServer pid=238699) WARNING 07-15 06:22:58 [__init__.py:36] SECURITY WARNING: Development endpoints are enabled! This should NOT be used in production!
|
| 49 |
+
(APIServer pid=238699) INFO 07-15 06:22:58 [hf.py:548] Detected the chat template content format to be 'string'. You can set `--chat-template-content-format` to override this.
|
| 50 |
+
(APIServer pid=238699) INFO 07-15 06:22:58 [api_server.py:616] Starting vLLM server on http://127.0.0.1:8377
|
| 51 |
+
(APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:37] Available routes are:
|
| 52 |
+
(APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /openapi.json, Methods: HEAD, GET
|
| 53 |
+
(APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /docs, Methods: HEAD, GET
|
| 54 |
+
(APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /docs/oauth2-redirect, Methods: HEAD, GET
|
| 55 |
+
(APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /redoc, Methods: HEAD, GET
|
| 56 |
+
(APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /load, Methods: GET
|
| 57 |
+
(APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /version, Methods: GET
|
| 58 |
+
(APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /health, Methods: GET
|
| 59 |
+
(APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /metrics, Methods: GET
|
| 60 |
+
(APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /tokenize, Methods: POST
|
| 61 |
+
(APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /detokenize, Methods: POST
|
| 62 |
+
(APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /v1/models, Methods: GET
|
| 63 |
+
(APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /ping, Methods: GET
|
| 64 |
+
(APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /ping, Methods: POST
|
| 65 |
+
(APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /invocations, Methods: POST
|
| 66 |
+
(APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /reset_prefix_cache, Methods: POST
|
| 67 |
+
(APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /reset_mm_cache, Methods: POST
|
| 68 |
+
(APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /reset_encoder_cache, Methods: POST
|
| 69 |
+
(APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /pause, Methods: POST
|
| 70 |
+
(APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /resume, Methods: POST
|
| 71 |
+
(APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /is_paused, Methods: GET
|
| 72 |
+
(APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /init_weight_transfer_engine, Methods: POST
|
| 73 |
+
(APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /start_weight_update, Methods: POST
|
| 74 |
+
(APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /update_weights, Methods: POST
|
| 75 |
+
(APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /finish_weight_update, Methods: POST
|
| 76 |
+
(APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /get_world_size, Methods: GET
|
| 77 |
+
(APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /collective_rpc, Methods: POST
|
| 78 |
+
(APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /server_info, Methods: GET
|
| 79 |
+
(APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /sleep, Methods: POST
|
| 80 |
+
(APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /wake_up, Methods: POST
|
| 81 |
+
(APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /is_sleeping, Methods: GET
|
| 82 |
+
(APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /v1/chat/completions, Methods: POST
|
| 83 |
+
(APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /v1/chat/completions/batch, Methods: POST
|
| 84 |
+
(APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /v1/responses, Methods: POST
|
| 85 |
+
(APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /v1/responses/{response_id}, Methods: GET
|
| 86 |
+
(APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /v1/responses/{response_id}/cancel, Methods: POST
|
| 87 |
+
(APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /v1/completions, Methods: POST
|
| 88 |
+
(APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /v1/messages, Methods: POST
|
| 89 |
+
(APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /v1/messages/count_tokens, Methods: POST
|
| 90 |
+
(APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /generative_scoring, Methods: POST
|
| 91 |
+
(APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /scale_elastic_ep, Methods: POST
|
| 92 |
+
(APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /is_scaling_elastic_ep, Methods: POST
|
| 93 |
+
(APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /v1/chat/completions/render, Methods: POST
|
| 94 |
+
(APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /v1/completions/render, Methods: POST
|
| 95 |
+
(APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /v1/chat/completions/derender, Methods: POST
|
| 96 |
+
(APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /v1/completions/derender, Methods: POST
|
| 97 |
+
(APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /inference/v1/generate, Methods: POST
|
| 98 |
+
(APIServer pid=238699) INFO: Started server process [238699]
|
| 99 |
+
(APIServer pid=238699) INFO: Waiting for application startup.
|
| 100 |
+
(APIServer pid=238699) INFO: Application startup complete.
|
| 101 |
+
(APIServer pid=238699) INFO: 127.0.0.1:52570 - "GET /health HTTP/1.1" 200 OK
|
| 102 |
+
(EngineCore pid=238829) WARNING 07-15 06:22:59 [jit_monitor.py:129] Triton kernel JIT compilation during inference: _build_route_rows. This causes a latency spike; consider extending warmup to cover this shape/config.
|
| 103 |
+
(APIServer pid=238699) INFO 07-15 06:23:08 [loggers.py:273] Engine 000: Avg prompt throughput: 2250.2 tokens/s, Avg generation throughput: 1992.1 tokens/s, Running: 251 reqs, Waiting: 968 reqs, GPU KV cache usage: 43.6%, Prefix cache hit rate: 94.2%
|
| 104 |
+
(APIServer pid=238699) INFO 07-15 06:23:18 [loggers.py:273] Engine 000: Avg prompt throughput: 1600.3 tokens/s, Avg generation throughput: 2509.7 tokens/s, Running: 254 reqs, Waiting: 725 reqs, GPU KV cache usage: 45.6%, Prefix cache hit rate: 94.3%
|
| 105 |
+
(APIServer pid=238699) INFO 07-15 06:23:28 [loggers.py:273] Engine 000: Avg prompt throughput: 1587.9 tokens/s, Avg generation throughput: 2459.3 tokens/s, Running: 255 reqs, Waiting: 488 reqs, GPU KV cache usage: 45.1%, Prefix cache hit rate: 94.3%
|
| 106 |
+
(APIServer pid=238699) INFO 07-15 06:23:38 [loggers.py:273] Engine 000: Avg prompt throughput: 1514.4 tokens/s, Avg generation throughput: 2510.7 tokens/s, Running: 256 reqs, Waiting: 264 reqs, GPU KV cache usage: 47.5%, Prefix cache hit rate: 94.4%
|
| 107 |
+
(APIServer pid=238699) INFO 07-15 06:23:48 [loggers.py:273] Engine 000: Avg prompt throughput: 1560.7 tokens/s, Avg generation throughput: 2536.6 tokens/s, Running: 251 reqs, Waiting: 31 reqs, GPU KV cache usage: 48.7%, Prefix cache hit rate: 94.4%
|
| 108 |
+
(APIServer pid=238699) INFO 07-15 06:23:58 [loggers.py:273] Engine 000: Avg prompt throughput: 211.0 tokens/s, Avg generation throughput: 1868.5 tokens/s, Running: 24 reqs, Waiting: 0 reqs, GPU KV cache usage: 10.8%, Prefix cache hit rate: 94.4%
|
| 109 |
+
(APIServer pid=238699) INFO: 127.0.0.1:52586 - "POST /v1/completions HTTP/1.1" 200 OK
|
| 110 |
+
(APIServer pid=238699) INFO 07-15 06:24:07 [launcher.py:100] [shutdown] API server: shutdown triggered
|
| 111 |
+
(APIServer pid=238699) INFO 07-15 06:24:07 [launcher.py:116] [shutdown] API server: stopping engine client mode=abort timeout=0s
|
| 112 |
+
(EngineCore pid=238829) INFO 07-15 06:24:07 [core.py:1214] [shutdown] EngineCore: trigger received signal=SIGTERM
|
| 113 |
+
(EngineCore pid=238829) INFO 07-15 06:24:07 [core.py:1333] [shutdown] EngineCore: start mode=abort timeout=0s
|
| 114 |
+
(EngineCore pid=238829) INFO 07-15 06:24:07 [core.py:1364] [shutdown] EngineCore: request processing complete; starting resource teardown
|
| 115 |
+
(EngineCore pid=238829) INFO 07-15 06:24:07 [core.py:1227] [shutdown] EngineCore: exiting busy loop
|
| 116 |
+
(APIServer pid=238699) INFO 07-15 06:24:07 [core_client.py:655] [shutdown] MPClient: start timeout=0s
|
| 117 |
+
(APIServer pid=238699) INFO 07-15 06:24:07 [core_client.py:657] [shutdown] MPClient: stopping engine manager
|
| 118 |
+
(APIServer pid=238699) WARNING 07-15 06:24:07 [utils.py:626] [shutdown] Process manager: force killing remaining processes count=1
|
| 119 |
+
(APIServer pid=238699) INFO 07-15 06:24:07 [core_client.py:659] [shutdown] MPClient: engine manager stopped
|
| 120 |
+
(APIServer pid=238699) INFO 07-15 06:24:07 [core_client.py:660] [shutdown] MPClient: cleaning up background resources
|
| 121 |
+
(APIServer pid=238699) INFO 07-15 06:24:07 [core_client.py:662] [shutdown] MPClient: complete
|
| 122 |
+
(APIServer pid=238699) INFO 07-15 06:24:07 [launcher.py:125] [shutdown] API server: engine client stopped
|
| 123 |
+
(APIServer pid=238699) INFO 07-15 06:24:07 [launcher.py:128] [shutdown] API server: signalling HTTP server shutdown
|
| 124 |
+
(APIServer pid=238699) INFO 07-15 06:24:07 [launcher.py:149] [shutdown] API server: shutting down FastAPI HTTP server
|
| 125 |
+
(APIServer pid=238699) INFO: Shutting down
|
| 126 |
+
(APIServer pid=238699) INFO: Waiting for application shutdown.
|
| 127 |
+
(APIServer pid=238699) INFO: Application shutdown complete.
|
| 128 |
+
/home/henry/.local/share/uv/python/cpython-3.12.12-linux-x86_64-gnu/lib/python3.12/multiprocessing/resource_tracker.py:279: UserWarning: resource_tracker: There appear to be 1 leaked semaphore objects to clean up at shutdown
|
| 129 |
+
warnings.warn('resource_tracker: There appear to be %d '
|
evals/healing_breadth/reap_math_keep75_seed1224.json
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
evals/healing_breadth/reap_math_keep75_seed1224.json.server.log
ADDED
|
@@ -0,0 +1,128 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
(APIServer pid=207721) INFO 07-14 23:28:07 [api_utils.py:339]
|
| 2 |
+
(APIServer pid=207721) INFO 07-14 23:28:07 [api_utils.py:339] β β ββ ββ
|
| 3 |
+
(APIServer pid=207721) INFO 07-14 23:28:07 [api_utils.py:339] ββ ββ β β β βββ β version 0.25.0
|
| 4 |
+
(APIServer pid=207721) INFO 07-14 23:28:07 [api_utils.py:339] ββββ β β β β model outputs/healed/healing_breadth/reap_math_keep75_seed1224/step0050
|
| 5 |
+
(APIServer pid=207721) INFO 07-14 23:28:07 [api_utils.py:339] ββ βββββ βββββ β β
|
| 6 |
+
(APIServer pid=207721) INFO 07-14 23:28:07 [api_utils.py:339]
|
| 7 |
+
(APIServer pid=207721) INFO 07-14 23:28:07 [api_utils.py:273] non-default args: {'model_tag': 'outputs/healed/healing_breadth/reap_math_keep75_seed1224/step0050', 'host': '127.0.0.1', 'port': 8377, 'model': 'outputs/healed/healing_breadth/reap_math_keep75_seed1224/step0050', 'max_model_len': 2048, 'enforce_eager': True, 'served_model_name': ['student'], 'gpu_memory_utilization': 0.85}
|
| 8 |
+
(APIServer pid=207721) INFO 07-14 23:28:07 [model.py:619] Resolved architecture: OlmoeForCausalLM
|
| 9 |
+
(APIServer pid=207721) INFO 07-14 23:28:07 [model.py:1776] Using max model len 2048
|
| 10 |
+
(APIServer pid=207721) INFO 07-14 23:28:07 [vllm.py:1042] Asynchronous scheduling is enabled.
|
| 11 |
+
(APIServer pid=207721) WARNING 07-14 23:28:07 [vllm.py:1096] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
|
| 12 |
+
(APIServer pid=207721) WARNING 07-14 23:28:07 [vllm.py:1144] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
|
| 13 |
+
(APIServer pid=207721) INFO 07-14 23:28:07 [kernel.py:292] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native'])
|
| 14 |
+
(APIServer pid=207721) INFO 07-14 23:28:07 [vllm.py:1322] Cudagraph is disabled under eager mode
|
| 15 |
+
(APIServer pid=207721) INFO 07-14 23:28:07 [compilation.py:312] Enabled custom fusions: norm_quant, act_quant
|
| 16 |
+
(EngineCore pid=207838) INFO 07-14 23:28:15 [core.py:114] Initializing a V1 LLM engine (v0.25.0) with config: model='outputs/healed/healing_breadth/reap_math_keep75_seed1224/step0050', speculative_config=None, tokenizer='outputs/healed/healing_breadth/reap_math_keep75_seed1224/step0050', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=False, dtype=torch.bfloat16, max_seq_len=2048, download_dir=None, load_format=auto, tensor_parallel_size=1, pipeline_parallel_size=1, data_parallel_size=1, decode_context_parallel_size=1, dcp_comm_backend=ag_rs, disable_custom_all_reduce=False, quantization=None, quantization_config=None, enforce_eager=True, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False, jit_monitor_mode='warn', jit_monitor_verbose=False), seed=0, served_model_name=student, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'mode': <CompilationMode.NONE: 0>, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'ir_enable_torch_wrap': False, 'splitting_ops': [], 'compile_mm_encoder': False, 'cudagraph_mm_encoder': False, 'encoder_cudagraph_token_budgets': [], 'encoder_cudagraph_max_vision_items_per_batch': 0, 'encoder_cudagraph_max_frames_per_batch': None, 'compile_sizes': [], 'compile_ranges_endpoints': [2048], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'size_asserts': False, 'alignment_asserts': False, 'scalar_asserts': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': <CUDAGraphMode.NONE: 0>, 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': True, 'fuse_act_quant': True, 'fuse_attn_quant': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False, 'fuse_rope_kvcache_cat_mla': False, 'fuse_act_padding': False}, 'max_cudagraph_capture_size': 0, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': False, 'static_all_moe_layers': []}, kernel_config=KernelConfig(ir_op_priority=IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native']), enable_flashinfer_autotune=True, enable_cutedsl_warmup=True, moe_backend='auto', linear_backend='auto')
|
| 17 |
+
(EngineCore pid=207838) INFO 07-14 23:28:15 [parallel_state.py:1607] world_size=1 rank=0 local_rank=0 distributed_init_method=tcp://192.168.0.15:33565 backend=nccl
|
| 18 |
+
(EngineCore pid=207838) INFO 07-14 23:28:15 [parallel_state.py:1942] rank 0 in world size 1 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank 0, EPLB rank N/A
|
| 19 |
+
(EngineCore pid=207838) INFO 07-14 23:28:16 [topk_topp_sampler.py:55] Using FlashInfer for top-p & top-k sampling.
|
| 20 |
+
(EngineCore pid=207838) INFO 07-14 23:28:16 [gpu_model_runner.py:5209] Starting to load model outputs/healed/healing_breadth/reap_math_keep75_seed1224/step0050...
|
| 21 |
+
(EngineCore pid=207838) INFO 07-14 23:28:17 [cuda.py:476] Using FLASH_ATTN attention backend out of potential backends: ['FLASH_ATTN', 'FLASHINFER', 'TRITON_ATTN', 'FLEX_ATTENTION'].
|
| 22 |
+
(EngineCore pid=207838) INFO 07-14 23:28:17 [flash_attn.py:718] Using FlashAttention version 2
|
| 23 |
+
(EngineCore pid=207838) INFO 07-14 23:28:17 [unquantized.py:262] Using TRITON Unquantized MoE backend out of potential backends: ['FlashInfer TRTLLM', 'FlashInfer CUTLASS', 'TRITON', 'BATCHED_TRITON'].
|
| 24 |
+
(EngineCore pid=207838) INFO 07-14 23:28:17 [weight_utils.py:849] Filesystem type for checkpoints: EXT4. Checkpoint size: 9.89 GiB. Available RAM: 68.46 GiB.
|
| 25 |
+
(EngineCore pid=207838) INFO 07-14 23:28:17 [weight_utils.py:872] Auto-prefetch is disabled because the filesystem (EXT4) is not a recognized network FS (NFS/Lustre). If you want to force prefetching, start vLLM with --safetensors-load-strategy=prefetch.
|
| 26 |
+
(EngineCore pid=207838)
|
| 27 |
+
(EngineCore pid=207838)
|
| 28 |
+
(EngineCore pid=207838)
|
| 29 |
+
(EngineCore pid=207838)
|
| 30 |
+
(EngineCore pid=207838)
|
| 31 |
+
(EngineCore pid=207838) INFO 07-14 23:28:22 [default_loader.py:430] Loading weights took 5.26 seconds
|
| 32 |
+
(EngineCore pid=207838) INFO 07-14 23:28:22 [unquantized.py:334] Using MoEPrepareAndFinalizeNoDPEPModular
|
| 33 |
+
(EngineCore pid=207838) INFO 07-14 23:28:23 [gpu_model_runner.py:5306] Model loading took 9.89 GiB memory and 5.436781 seconds
|
| 34 |
+
(EngineCore pid=207838) WARNING 07-14 23:28:23 [fused_moe.py:1106] Using default MoE config. Performance might be sub-optimal! Config file not found at /home/henry/Documents/PythonProjects/variable-reap/vllm-plugin/.venv25/lib/python3.12/site-packages/vllm/model_executor/layers/fused_moe/configs/E=48,N=1024,device_name=NVIDIA_GeForce_RTX_3090.json
|
| 35 |
+
(EngineCore pid=207838) INFO 07-14 23:28:24 [gpu_worker.py:538] Available KV cache memory: 9.73 GiB
|
| 36 |
+
(EngineCore pid=207838) INFO 07-14 23:28:24 [kv_cache_utils.py:2146] GPU KV cache size: 79,680 tokens
|
| 37 |
+
(EngineCore pid=207838) INFO 07-14 23:28:24 [kv_cache_utils.py:2147] Maximum concurrency for 2,048 tokens per request: 38.91x
|
| 38 |
+
(EngineCore pid=207838) INFO 07-14 23:28:24 [cutedsl_warmup.py:97] Skipping CuTeDSL warmup because no compile units were requested.
|
| 39 |
+
(EngineCore pid=207838) INFO 07-14 23:28:24 [jit_monitor.py:73] Kernel JIT monitor activated; monitored JIT compilations during inference will use mode=warn.
|
| 40 |
+
(EngineCore pid=207838) INFO 07-14 23:28:25 [core.py:344] init engine (profile, create kv cache, warmup model) took 2.11 s
|
| 41 |
+
(EngineCore pid=207838) INFO 07-14 23:28:25 [vllm.py:1042] Asynchronous scheduling is enabled.
|
| 42 |
+
(EngineCore pid=207838) WARNING 07-14 23:28:25 [vllm.py:1096] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
|
| 43 |
+
(EngineCore pid=207838) WARNING 07-14 23:28:25 [vllm.py:1144] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
|
| 44 |
+
(EngineCore pid=207838) INFO 07-14 23:28:25 [kernel.py:292] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native'])
|
| 45 |
+
(EngineCore pid=207838) INFO 07-14 23:28:25 [vllm.py:1322] Cudagraph is disabled under eager mode
|
| 46 |
+
(EngineCore pid=207838) INFO 07-14 23:28:25 [compilation.py:312] Enabled custom fusions: norm_quant, act_quant
|
| 47 |
+
(APIServer pid=207721) INFO 07-14 23:28:25 [api_server.py:612] Supported tasks: ['generate']
|
| 48 |
+
(APIServer pid=207721) WARNING 07-14 23:28:25 [__init__.py:36] SECURITY WARNING: Development endpoints are enabled! This should NOT be used in production!
|
| 49 |
+
(APIServer pid=207721) INFO 07-14 23:28:25 [hf.py:548] Detected the chat template content format to be 'string'. You can set `--chat-template-content-format` to override this.
|
| 50 |
+
(APIServer pid=207721) INFO 07-14 23:28:25 [api_server.py:616] Starting vLLM server on http://127.0.0.1:8377
|
| 51 |
+
(APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:37] Available routes are:
|
| 52 |
+
(APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /openapi.json, Methods: HEAD, GET
|
| 53 |
+
(APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /docs, Methods: HEAD, GET
|
| 54 |
+
(APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /docs/oauth2-redirect, Methods: HEAD, GET
|
| 55 |
+
(APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /redoc, Methods: HEAD, GET
|
| 56 |
+
(APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /load, Methods: GET
|
| 57 |
+
(APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /version, Methods: GET
|
| 58 |
+
(APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /health, Methods: GET
|
| 59 |
+
(APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /metrics, Methods: GET
|
| 60 |
+
(APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /tokenize, Methods: POST
|
| 61 |
+
(APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /detokenize, Methods: POST
|
| 62 |
+
(APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /v1/models, Methods: GET
|
| 63 |
+
(APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /ping, Methods: GET
|
| 64 |
+
(APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /ping, Methods: POST
|
| 65 |
+
(APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /invocations, Methods: POST
|
| 66 |
+
(APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /reset_prefix_cache, Methods: POST
|
| 67 |
+
(APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /reset_mm_cache, Methods: POST
|
| 68 |
+
(APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /reset_encoder_cache, Methods: POST
|
| 69 |
+
(APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /pause, Methods: POST
|
| 70 |
+
(APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /resume, Methods: POST
|
| 71 |
+
(APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /is_paused, Methods: GET
|
| 72 |
+
(APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /init_weight_transfer_engine, Methods: POST
|
| 73 |
+
(APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /start_weight_update, Methods: POST
|
| 74 |
+
(APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /update_weights, Methods: POST
|
| 75 |
+
(APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /finish_weight_update, Methods: POST
|
| 76 |
+
(APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /get_world_size, Methods: GET
|
| 77 |
+
(APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /collective_rpc, Methods: POST
|
| 78 |
+
(APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /server_info, Methods: GET
|
| 79 |
+
(APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /sleep, Methods: POST
|
| 80 |
+
(APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /wake_up, Methods: POST
|
| 81 |
+
(APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /is_sleeping, Methods: GET
|
| 82 |
+
(APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /v1/chat/completions, Methods: POST
|
| 83 |
+
(APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /v1/chat/completions/batch, Methods: POST
|
| 84 |
+
(APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /v1/responses, Methods: POST
|
| 85 |
+
(APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /v1/responses/{response_id}, Methods: GET
|
| 86 |
+
(APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /v1/responses/{response_id}/cancel, Methods: POST
|
| 87 |
+
(APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /v1/completions, Methods: POST
|
| 88 |
+
(APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /v1/messages, Methods: POST
|
| 89 |
+
(APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /v1/messages/count_tokens, Methods: POST
|
| 90 |
+
(APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /generative_scoring, Methods: POST
|
| 91 |
+
(APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /scale_elastic_ep, Methods: POST
|
| 92 |
+
(APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /is_scaling_elastic_ep, Methods: POST
|
| 93 |
+
(APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /v1/chat/completions/render, Methods: POST
|
| 94 |
+
(APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /v1/completions/render, Methods: POST
|
| 95 |
+
(APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /v1/chat/completions/derender, Methods: POST
|
| 96 |
+
(APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /v1/completions/derender, Methods: POST
|
| 97 |
+
(APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /inference/v1/generate, Methods: POST
|
| 98 |
+
(APIServer pid=207721) INFO: Started server process [207721]
|
| 99 |
+
(APIServer pid=207721) INFO: Waiting for application startup.
|
| 100 |
+
(APIServer pid=207721) INFO: Application startup complete.
|
| 101 |
+
(APIServer pid=207721) INFO: 127.0.0.1:52832 - "GET /health HTTP/1.1" 200 OK
|
| 102 |
+
(EngineCore pid=207838) WARNING 07-14 23:28:27 [jit_monitor.py:129] Triton kernel JIT compilation during inference: fused_moe_kernel. This causes a latency spike; consider extending warmup to cover this shape/config.
|
| 103 |
+
(APIServer pid=207721) INFO 07-14 23:28:35 [loggers.py:273] Engine 000: Avg prompt throughput: 2896.9 tokens/s, Avg generation throughput: 2135.2 tokens/s, Running: 252 reqs, Waiting: 942 reqs, GPU KV cache usage: 47.8%, Prefix cache hit rate: 93.2%
|
| 104 |
+
(APIServer pid=207721) INFO 07-14 23:28:45 [loggers.py:273] Engine 000: Avg prompt throughput: 2050.3 tokens/s, Avg generation throughput: 3019.8 tokens/s, Running: 254 reqs, Waiting: 678 reqs, GPU KV cache usage: 51.2%, Prefix cache hit rate: 93.3%
|
| 105 |
+
(APIServer pid=207721) INFO 07-14 23:28:55 [loggers.py:273] Engine 000: Avg prompt throughput: 2279.4 tokens/s, Avg generation throughput: 2991.5 tokens/s, Running: 253 reqs, Waiting: 388 reqs, GPU KV cache usage: 50.9%, Prefix cache hit rate: 93.4%
|
| 106 |
+
(APIServer pid=207721) INFO 07-14 23:29:05 [loggers.py:273] Engine 000: Avg prompt throughput: 2386.0 tokens/s, Avg generation throughput: 2965.4 tokens/s, Running: 255 reqs, Waiting: 94 reqs, GPU KV cache usage: 51.0%, Prefix cache hit rate: 93.4%
|
| 107 |
+
(APIServer pid=207721) INFO 07-14 23:29:15 [loggers.py:273] Engine 000: Avg prompt throughput: 787.1 tokens/s, Avg generation throughput: 2775.3 tokens/s, Running: 14 reqs, Waiting: 0 reqs, GPU KV cache usage: 7.4%, Prefix cache hit rate: 93.4%
|
| 108 |
+
(APIServer pid=207721) INFO: 127.0.0.1:52848 - "POST /v1/completions HTTP/1.1" 200 OK
|
| 109 |
+
(EngineCore pid=207838) INFO 07-14 23:29:20 [core.py:1214] [shutdown] EngineCore: trigger received signal=SIGTERM
|
| 110 |
+
(APIServer pid=207721) INFO 07-14 23:29:20 [launcher.py:100] [shutdown] API server: shutdown triggered
|
| 111 |
+
(APIServer pid=207721) INFO 07-14 23:29:20 [launcher.py:116] [shutdown] API server: stopping engine client mode=abort timeout=0s
|
| 112 |
+
(EngineCore pid=207838) INFO 07-14 23:29:20 [core.py:1333] [shutdown] EngineCore: start mode=abort timeout=0s
|
| 113 |
+
(EngineCore pid=207838) INFO 07-14 23:29:20 [core.py:1364] [shutdown] EngineCore: request processing complete; starting resource teardown
|
| 114 |
+
(EngineCore pid=207838) INFO 07-14 23:29:20 [core.py:1227] [shutdown] EngineCore: exiting busy loop
|
| 115 |
+
(APIServer pid=207721) INFO 07-14 23:29:20 [core_client.py:655] [shutdown] MPClient: start timeout=0s
|
| 116 |
+
(APIServer pid=207721) INFO 07-14 23:29:20 [core_client.py:657] [shutdown] MPClient: stopping engine manager
|
| 117 |
+
(APIServer pid=207721) WARNING 07-14 23:29:20 [utils.py:626] [shutdown] Process manager: force killing remaining processes count=1
|
| 118 |
+
(APIServer pid=207721) INFO 07-14 23:29:20 [core_client.py:659] [shutdown] MPClient: engine manager stopped
|
| 119 |
+
(APIServer pid=207721) INFO 07-14 23:29:20 [core_client.py:660] [shutdown] MPClient: cleaning up background resources
|
| 120 |
+
(APIServer pid=207721) INFO 07-14 23:29:20 [core_client.py:662] [shutdown] MPClient: complete
|
| 121 |
+
(APIServer pid=207721) INFO 07-14 23:29:20 [launcher.py:125] [shutdown] API server: engine client stopped
|
| 122 |
+
(APIServer pid=207721) INFO 07-14 23:29:20 [launcher.py:128] [shutdown] API server: signalling HTTP server shutdown
|
| 123 |
+
(APIServer pid=207721) INFO 07-14 23:29:20 [launcher.py:149] [shutdown] API server: shutting down FastAPI HTTP server
|
| 124 |
+
(APIServer pid=207721) INFO: Shutting down
|
| 125 |
+
(APIServer pid=207721) INFO: Waiting for application shutdown.
|
| 126 |
+
(APIServer pid=207721) INFO: Application shutdown complete.
|
| 127 |
+
/home/henry/.local/share/uv/python/cpython-3.12.12-linux-x86_64-gnu/lib/python3.12/multiprocessing/resource_tracker.py:279: UserWarning: resource_tracker: There appear to be 1 leaked semaphore objects to clean up at shutdown
|
| 128 |
+
warnings.warn('resource_tracker: There appear to be %d '
|
evals/healing_breadth/uniform_math_keep50_seed1224.json
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
evals/healing_breadth/uniform_math_keep50_seed1224.json.server.log
ADDED
|
@@ -0,0 +1,130 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
(APIServer pid=215091) INFO 07-15 00:43:36 [api_utils.py:339]
|
| 2 |
+
(APIServer pid=215091) INFO 07-15 00:43:36 [api_utils.py:339] β β ββ ββ
|
| 3 |
+
(APIServer pid=215091) INFO 07-15 00:43:36 [api_utils.py:339] ββ ββ β β β βββ β version 0.25.0
|
| 4 |
+
(APIServer pid=215091) INFO 07-15 00:43:36 [api_utils.py:339] ββββ β β β β model outputs/healed/healing_breadth/uniform_math_keep50_seed1224/step0050
|
| 5 |
+
(APIServer pid=215091) INFO 07-15 00:43:36 [api_utils.py:339] ββ βββββ βββββ β β
|
| 6 |
+
(APIServer pid=215091) INFO 07-15 00:43:36 [api_utils.py:339]
|
| 7 |
+
(APIServer pid=215091) INFO 07-15 00:43:36 [api_utils.py:273] non-default args: {'model_tag': 'outputs/healed/healing_breadth/uniform_math_keep50_seed1224/step0050', 'host': '127.0.0.1', 'port': 8377, 'model': 'outputs/healed/healing_breadth/uniform_math_keep50_seed1224/step0050', 'max_model_len': 2048, 'enforce_eager': True, 'served_model_name': ['student'], 'gpu_memory_utilization': 0.85}
|
| 8 |
+
(APIServer pid=215091) INFO 07-15 00:43:36 [model.py:619] Resolved architecture: PrunedOlmoeForCausalLM
|
| 9 |
+
(APIServer pid=215091) INFO 07-15 00:43:36 [model.py:1776] Using max model len 2048
|
| 10 |
+
(APIServer pid=215091) INFO 07-15 00:43:36 [vllm.py:1042] Asynchronous scheduling is enabled.
|
| 11 |
+
(APIServer pid=215091) WARNING 07-15 00:43:36 [vllm.py:1096] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
|
| 12 |
+
(APIServer pid=215091) WARNING 07-15 00:43:36 [vllm.py:1144] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
|
| 13 |
+
(APIServer pid=215091) INFO 07-15 00:43:36 [kernel.py:292] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native'])
|
| 14 |
+
(APIServer pid=215091) INFO 07-15 00:43:36 [vllm.py:1322] Cudagraph is disabled under eager mode
|
| 15 |
+
(APIServer pid=215091) INFO 07-15 00:43:36 [compilation.py:312] Enabled custom fusions: norm_quant, act_quant
|
| 16 |
+
(EngineCore pid=215209) INFO 07-15 00:43:44 [core.py:114] Initializing a V1 LLM engine (v0.25.0) with config: model='outputs/healed/healing_breadth/uniform_math_keep50_seed1224/step0050', speculative_config=None, tokenizer='outputs/healed/healing_breadth/uniform_math_keep50_seed1224/step0050', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=False, dtype=torch.bfloat16, max_seq_len=2048, download_dir=None, load_format=auto, tensor_parallel_size=1, pipeline_parallel_size=1, data_parallel_size=1, decode_context_parallel_size=1, dcp_comm_backend=ag_rs, disable_custom_all_reduce=False, quantization=None, quantization_config=None, enforce_eager=True, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False, jit_monitor_mode='warn', jit_monitor_verbose=False), seed=0, served_model_name=student, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'mode': <CompilationMode.NONE: 0>, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'ir_enable_torch_wrap': False, 'splitting_ops': [], 'compile_mm_encoder': False, 'cudagraph_mm_encoder': False, 'encoder_cudagraph_token_budgets': [], 'encoder_cudagraph_max_vision_items_per_batch': 0, 'encoder_cudagraph_max_frames_per_batch': None, 'compile_sizes': [], 'compile_ranges_endpoints': [2048], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'size_asserts': False, 'alignment_asserts': False, 'scalar_asserts': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': <CUDAGraphMode.NONE: 0>, 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': True, 'fuse_act_quant': True, 'fuse_attn_quant': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False, 'fuse_rope_kvcache_cat_mla': False, 'fuse_act_padding': False}, 'max_cudagraph_capture_size': 0, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': False, 'static_all_moe_layers': []}, kernel_config=KernelConfig(ir_op_priority=IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native']), enable_flashinfer_autotune=True, enable_cutedsl_warmup=True, moe_backend='auto', linear_backend='auto')
|
| 17 |
+
(EngineCore pid=215209) INFO 07-15 00:43:44 [parallel_state.py:1607] world_size=1 rank=0 local_rank=0 distributed_init_method=tcp://192.168.0.15:36073 backend=nccl
|
| 18 |
+
(EngineCore pid=215209) INFO 07-15 00:43:44 [parallel_state.py:1942] rank 0 in world size 1 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank 0, EPLB rank N/A
|
| 19 |
+
(EngineCore pid=215209) INFO 07-15 00:43:45 [topk_topp_sampler.py:55] Using FlashInfer for top-p & top-k sampling.
|
| 20 |
+
(EngineCore pid=215209) INFO 07-15 00:43:45 [gpu_model_runner.py:5209] Starting to load model outputs/healed/healing_breadth/uniform_math_keep50_seed1224/step0050...
|
| 21 |
+
(EngineCore pid=215209) INFO 07-15 00:43:46 [cuda.py:476] Using FLASH_ATTN attention backend out of potential backends: ['FLASH_ATTN', 'FLASHINFER', 'TRITON_ATTN', 'FLEX_ATTENTION'].
|
| 22 |
+
(EngineCore pid=215209) INFO 07-15 00:43:46 [flash_attn.py:718] Using FlashAttention version 2
|
| 23 |
+
(EngineCore pid=215209) /home/henry/.cache/glean/megablocks-variable-93a1479bc15b/megablocks/grouped_gemm_util.py:10: UserWarning: Grouped GEMM not available.
|
| 24 |
+
(EngineCore pid=215209) warnings.warn('Grouped GEMM not available.')
|
| 25 |
+
(EngineCore pid=215209) INFO 07-15 00:43:46 [weight_utils.py:849] Filesystem type for checkpoints: EXT4. Checkpoint size: 6.89 GiB. Available RAM: 61.47 GiB.
|
| 26 |
+
(EngineCore pid=215209) INFO 07-15 00:43:46 [weight_utils.py:872] Auto-prefetch is disabled because the filesystem (EXT4) is not a recognized network FS (NFS/Lustre). If you want to force prefetching, start vLLM with --safetensors-load-strategy=prefetch.
|
| 27 |
+
(EngineCore pid=215209)
|
| 28 |
+
(EngineCore pid=215209)
|
| 29 |
+
(EngineCore pid=215209)
|
| 30 |
+
(EngineCore pid=215209)
|
| 31 |
+
(EngineCore pid=215209)
|
| 32 |
+
(EngineCore pid=215209) INFO 07-15 00:43:50 [default_loader.py:430] Loading weights took 4.53 seconds
|
| 33 |
+
(EngineCore pid=215209) INFO 07-15 00:43:51 [gpu_model_runner.py:5306] Model loading took 6.89 GiB memory and 4.727233 seconds
|
| 34 |
+
(EngineCore pid=215209) INFO 07-15 00:43:52 [gpu_worker.py:538] Available KV cache memory: 12.84 GiB
|
| 35 |
+
(EngineCore pid=215209) INFO 07-15 00:43:52 [kv_cache_utils.py:2146] GPU KV cache size: 105,216 tokens
|
| 36 |
+
(EngineCore pid=215209) INFO 07-15 00:43:52 [kv_cache_utils.py:2147] Maximum concurrency for 2,048 tokens per request: 51.38x
|
| 37 |
+
(EngineCore pid=215209) INFO 07-15 00:43:52 [cutedsl_warmup.py:97] Skipping CuTeDSL warmup because no compile units were requested.
|
| 38 |
+
(EngineCore pid=215209) INFO 07-15 00:43:53 [jit_monitor.py:73] Kernel JIT monitor activated; monitored JIT compilations during inference will use mode=warn.
|
| 39 |
+
(EngineCore pid=215209) INFO 07-15 00:43:53 [core.py:344] init engine (profile, create kv cache, warmup model) took 2.12 s
|
| 40 |
+
(EngineCore pid=215209) INFO 07-15 00:43:53 [vllm.py:1042] Asynchronous scheduling is enabled.
|
| 41 |
+
(EngineCore pid=215209) WARNING 07-15 00:43:53 [vllm.py:1096] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
|
| 42 |
+
(EngineCore pid=215209) WARNING 07-15 00:43:53 [vllm.py:1144] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
|
| 43 |
+
(EngineCore pid=215209) INFO 07-15 00:43:53 [kernel.py:292] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native'])
|
| 44 |
+
(EngineCore pid=215209) INFO 07-15 00:43:53 [vllm.py:1322] Cudagraph is disabled under eager mode
|
| 45 |
+
(EngineCore pid=215209) INFO 07-15 00:43:53 [compilation.py:312] Enabled custom fusions: norm_quant, act_quant
|
| 46 |
+
(APIServer pid=215091) INFO 07-15 00:43:53 [api_server.py:612] Supported tasks: ['generate']
|
| 47 |
+
(APIServer pid=215091) WARNING 07-15 00:43:53 [__init__.py:36] SECURITY WARNING: Development endpoints are enabled! This should NOT be used in production!
|
| 48 |
+
(APIServer pid=215091) INFO 07-15 00:43:53 [hf.py:548] Detected the chat template content format to be 'string'. You can set `--chat-template-content-format` to override this.
|
| 49 |
+
(APIServer pid=215091) INFO 07-15 00:43:53 [api_server.py:616] Starting vLLM server on http://127.0.0.1:8377
|
| 50 |
+
(APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:37] Available routes are:
|
| 51 |
+
(APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /openapi.json, Methods: GET, HEAD
|
| 52 |
+
(APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /docs, Methods: GET, HEAD
|
| 53 |
+
(APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /docs/oauth2-redirect, Methods: GET, HEAD
|
| 54 |
+
(APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /redoc, Methods: GET, HEAD
|
| 55 |
+
(APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /load, Methods: GET
|
| 56 |
+
(APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /version, Methods: GET
|
| 57 |
+
(APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /health, Methods: GET
|
| 58 |
+
(APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /metrics, Methods: GET
|
| 59 |
+
(APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /tokenize, Methods: POST
|
| 60 |
+
(APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /detokenize, Methods: POST
|
| 61 |
+
(APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /v1/models, Methods: GET
|
| 62 |
+
(APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /ping, Methods: GET
|
| 63 |
+
(APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /ping, Methods: POST
|
| 64 |
+
(APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /invocations, Methods: POST
|
| 65 |
+
(APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /reset_prefix_cache, Methods: POST
|
| 66 |
+
(APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /reset_mm_cache, Methods: POST
|
| 67 |
+
(APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /reset_encoder_cache, Methods: POST
|
| 68 |
+
(APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /pause, Methods: POST
|
| 69 |
+
(APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /resume, Methods: POST
|
| 70 |
+
(APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /is_paused, Methods: GET
|
| 71 |
+
(APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /init_weight_transfer_engine, Methods: POST
|
| 72 |
+
(APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /start_weight_update, Methods: POST
|
| 73 |
+
(APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /update_weights, Methods: POST
|
| 74 |
+
(APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /finish_weight_update, Methods: POST
|
| 75 |
+
(APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /get_world_size, Methods: GET
|
| 76 |
+
(APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /collective_rpc, Methods: POST
|
| 77 |
+
(APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /server_info, Methods: GET
|
| 78 |
+
(APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /sleep, Methods: POST
|
| 79 |
+
(APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /wake_up, Methods: POST
|
| 80 |
+
(APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /is_sleeping, Methods: GET
|
| 81 |
+
(APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /v1/chat/completions, Methods: POST
|
| 82 |
+
(APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /v1/chat/completions/batch, Methods: POST
|
| 83 |
+
(APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /v1/responses, Methods: POST
|
| 84 |
+
(APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /v1/responses/{response_id}, Methods: GET
|
| 85 |
+
(APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /v1/responses/{response_id}/cancel, Methods: POST
|
| 86 |
+
(APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /v1/completions, Methods: POST
|
| 87 |
+
(APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /v1/messages, Methods: POST
|
| 88 |
+
(APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /v1/messages/count_tokens, Methods: POST
|
| 89 |
+
(APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /generative_scoring, Methods: POST
|
| 90 |
+
(APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /scale_elastic_ep, Methods: POST
|
| 91 |
+
(APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /is_scaling_elastic_ep, Methods: POST
|
| 92 |
+
(APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /v1/chat/completions/render, Methods: POST
|
| 93 |
+
(APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /v1/completions/render, Methods: POST
|
| 94 |
+
(APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /v1/chat/completions/derender, Methods: POST
|
| 95 |
+
(APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /v1/completions/derender, Methods: POST
|
| 96 |
+
(APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /inference/v1/generate, Methods: POST
|
| 97 |
+
(APIServer pid=215091) INFO: Started server process [215091]
|
| 98 |
+
(APIServer pid=215091) INFO: Waiting for application startup.
|
| 99 |
+
(APIServer pid=215091) INFO: Application startup complete.
|
| 100 |
+
(APIServer pid=215091) INFO: 127.0.0.1:42130 - "GET /health HTTP/1.1" 200 OK
|
| 101 |
+
(EngineCore pid=215209) WARNING 07-15 00:43:55 [jit_monitor.py:129] Triton kernel JIT compilation during inference: _build_route_rows. This causes a latency spike; consider extending warmup to cover this shape/config.
|
| 102 |
+
(APIServer pid=215091) INFO 07-15 00:44:04 [loggers.py:273] Engine 000: Avg prompt throughput: 2793.7 tokens/s, Avg generation throughput: 2084.8 tokens/s, Running: 253 reqs, Waiting: 957 reqs, GPU KV cache usage: 37.0%, Prefix cache hit rate: 93.2%
|
| 103 |
+
(APIServer pid=215091) INFO 07-15 00:44:14 [loggers.py:273] Engine 000: Avg prompt throughput: 1725.4 tokens/s, Avg generation throughput: 2588.6 tokens/s, Running: 256 reqs, Waiting: 733 reqs, GPU KV cache usage: 40.4%, Prefix cache hit rate: 93.3%
|
| 104 |
+
(APIServer pid=215091) INFO 07-15 00:44:24 [loggers.py:273] Engine 000: Avg prompt throughput: 1870.4 tokens/s, Avg generation throughput: 2587.0 tokens/s, Running: 253 reqs, Waiting: 496 reqs, GPU KV cache usage: 40.3%, Prefix cache hit rate: 93.3%
|
| 105 |
+
(APIServer pid=215091) INFO 07-15 00:44:34 [loggers.py:273] Engine 000: Avg prompt throughput: 1878.2 tokens/s, Avg generation throughput: 2561.9 tokens/s, Running: 253 reqs, Waiting: 263 reqs, GPU KV cache usage: 41.0%, Prefix cache hit rate: 93.4%
|
| 106 |
+
(APIServer pid=215091) INFO 07-15 00:44:44 [loggers.py:273] Engine 000: Avg prompt throughput: 1818.6 tokens/s, Avg generation throughput: 2639.5 tokens/s, Running: 255 reqs, Waiting: 35 reqs, GPU KV cache usage: 44.2%, Prefix cache hit rate: 93.4%
|
| 107 |
+
(APIServer pid=215091) INFO 07-15 00:44:54 [loggers.py:273] Engine 000: Avg prompt throughput: 315.9 tokens/s, Avg generation throughput: 2202.9 tokens/s, Running: 29 reqs, Waiting: 0 reqs, GPU KV cache usage: 10.0%, Prefix cache hit rate: 93.4%
|
| 108 |
+
(APIServer pid=215091) INFO: 127.0.0.1:42138 - "POST /v1/completions HTTP/1.1" 200 OK
|
| 109 |
+
(APIServer pid=215091) INFO 07-15 00:45:04 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 389.4 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 93.4%
|
| 110 |
+
(APIServer pid=215091) INFO 07-15 00:45:04 [launcher.py:100] [shutdown] API server: shutdown triggered
|
| 111 |
+
(EngineCore pid=215209) INFO 07-15 00:45:04 [core.py:1214] [shutdown] EngineCore: trigger received signal=SIGTERM
|
| 112 |
+
(APIServer pid=215091) INFO 07-15 00:45:04 [launcher.py:116] [shutdown] API server: stopping engine client mode=abort timeout=0s
|
| 113 |
+
(EngineCore pid=215209) INFO 07-15 00:45:04 [core.py:1333] [shutdown] EngineCore: start mode=abort timeout=0s
|
| 114 |
+
(EngineCore pid=215209) INFO 07-15 00:45:04 [core.py:1364] [shutdown] EngineCore: request processing complete; starting resource teardown
|
| 115 |
+
(EngineCore pid=215209) INFO 07-15 00:45:04 [core.py:1227] [shutdown] EngineCore: exiting busy loop
|
| 116 |
+
(APIServer pid=215091) INFO 07-15 00:45:04 [core_client.py:655] [shutdown] MPClient: start timeout=0s
|
| 117 |
+
(APIServer pid=215091) INFO 07-15 00:45:04 [core_client.py:657] [shutdown] MPClient: stopping engine manager
|
| 118 |
+
(APIServer pid=215091) WARNING 07-15 00:45:04 [utils.py:626] [shutdown] Process manager: force killing remaining processes count=1
|
| 119 |
+
(APIServer pid=215091) INFO: Shutting down
|
| 120 |
+
(APIServer pid=215091) INFO 07-15 00:45:04 [core_client.py:659] [shutdown] MPClient: engine manager stopped
|
| 121 |
+
(APIServer pid=215091) INFO 07-15 00:45:04 [core_client.py:660] [shutdown] MPClient: cleaning up background resources
|
| 122 |
+
(APIServer pid=215091) INFO 07-15 00:45:04 [core_client.py:662] [shutdown] MPClient: complete
|
| 123 |
+
(APIServer pid=215091) INFO 07-15 00:45:04 [launcher.py:125] [shutdown] API server: engine client stopped
|
| 124 |
+
(APIServer pid=215091) INFO 07-15 00:45:04 [launcher.py:128] [shutdown] API server: signalling HTTP server shutdown
|
| 125 |
+
(APIServer pid=215091) INFO 07-15 00:45:04 [launcher.py:149] [shutdown] API server: shutting down FastAPI HTTP server
|
| 126 |
+
(APIServer pid=215091) INFO: Shutting down
|
| 127 |
+
(APIServer pid=215091) INFO: Waiting for application shutdown.
|
| 128 |
+
(APIServer pid=215091) INFO: Application shutdown complete.
|
| 129 |
+
/home/henry/.local/share/uv/python/cpython-3.12.12-linux-x86_64-gnu/lib/python3.12/multiprocessing/resource_tracker.py:279: UserWarning: resource_tracker: There appear to be 1 leaked semaphore objects to clean up at shutdown
|
| 130 |
+
warnings.warn('resource_tracker: There appear to be %d '
|
evals/math_unhealed/glean_winnow-olmoe-math-keep25.eval.log
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
evals/math_unhealed/glean_winnow-olmoe-math-keep50.eval.log
ADDED
|
@@ -0,0 +1,70 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 0 |
0%| | 0/1319 [00:00<?, ?it/s]
|
| 1 |
14%|ββ | 191/1319 [00:00<00:00, 1902.21it/s]
|
| 2 |
29%|βββ | 385/1319 [00:00<00:00, 1923.43it/s]
|
| 3 |
44%|βββββ | 581/1319 [00:00<00:00, 1936.99it/s]
|
| 4 |
59%|ββββββ | 777/1319 [00:00<00:00, 1942.54it/s]
|
| 5 |
74%|ββββββββ | 973/1319 [00:00<00:00, 1945.81it/s]
|
| 6 |
89%|βββββββββ | 1168/1319 [00:00<00:00, 1939.97it/s]
|
|
|
|
|
|
|
| 7 |
0%| | 0/500 [00:00<?, ?it/s]
|
| 8 |
8%|β | 42/500 [00:00<00:01, 411.34it/s]
|
| 9 |
17%|ββ | 84/500 [00:00<00:00, 416.06it/s]
|
| 10 |
25%|βββ | 127/500 [00:00<00:00, 418.71it/s]
|
| 11 |
34%|ββββ | 170/500 [00:00<00:00, 421.00it/s]
|
| 12 |
43%|βββββ | 213/500 [00:00<00:00, 422.35it/s]
|
| 13 |
51%|βββββ | 256/500 [00:00<00:00, 424.21it/s]
|
| 14 |
60%|ββββββ | 299/500 [00:00<00:00, 424.80it/s]
|
| 15 |
68%|βββββββ | 342/500 [00:00<00:00, 424.12it/s]
|
| 16 |
77%|ββββββββ | 385/500 [00:00<00:00, 424.83it/s]
|
| 17 |
86%|βββββββββ | 428/500 [00:01<00:00, 424.33it/s]
|
| 18 |
94%|ββββββββββ| 471/500 [00:01<00:00, 425.60it/s]
|
|
|
|
|
|
|
| 19 |
0%| | 0/541 [00:00<?, ?it/s]
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 20 |
0%| | 0/164 [00:00<?, ?it/s]
|
| 21 |
87%|βββββββββ | 142/164 [00:00<00:00, 1419.11it/s]
|
|
|
|
|
|
|
| 22 |
0%| | 0/500 [00:00<?, ?it/s]
|
| 23 |
3%|β | 17/500 [00:00<00:02, 161.93it/s]
|
| 24 |
7%|β | 34/500 [00:00<00:02, 162.61it/s]
|
| 25 |
10%|β | 51/500 [00:00<00:02, 163.06it/s]
|
| 26 |
14%|ββ | 68/500 [00:00<00:02, 163.43it/s]
|
| 27 |
17%|ββ | 85/500 [00:00<00:02, 163.79it/s]
|
| 28 |
20%|ββ | 102/500 [00:00<00:02, 164.02it/s]
|
| 29 |
24%|βββ | 119/500 [00:00<00:02, 164.32it/s]
|
| 30 |
27%|βββ | 136/500 [00:00<00:02, 164.49it/s]
|
| 31 |
31%|βββ | 153/500 [00:00<00:02, 164.68it/s]
|
| 32 |
34%|ββββ | 170/500 [00:01<00:02, 164.78it/s]
|
| 33 |
37%|ββββ | 187/500 [00:01<00:01, 165.03it/s]
|
| 34 |
41%|ββββ | 204/500 [00:01<00:01, 165.27it/s]
|
| 35 |
44%|βββββ | 221/500 [00:01<00:01, 165.43it/s]
|
| 36 |
48%|βββββ | 238/500 [00:01<00:01, 165.60it/s]
|
| 37 |
51%|βββββ | 255/500 [00:01<00:01, 165.58it/s]
|
| 38 |
54%|ββββββ | 272/500 [00:01<00:01, 165.63it/s]
|
| 39 |
58%|ββββββ | 289/500 [00:01<00:01, 165.63it/s]
|
| 40 |
61%|ββββββ | 306/500 [00:01<00:01, 165.70it/s]
|
| 41 |
65%|βββββββ | 323/500 [00:01<00:01, 165.80it/s]
|
| 42 |
68%|βββββββ | 340/500 [00:02<00:00, 165.89it/s]
|
| 43 |
71%|ββββββββ | 357/500 [00:02<00:00, 165.80it/s]
|
| 44 |
75%|ββββββββ | 374/500 [00:02<00:00, 165.88it/s]
|
| 45 |
78%|ββββββββ | 391/500 [00:02<00:00, 165.74it/s]
|
| 46 |
82%|βββββββββ | 408/500 [00:02<00:00, 165.84it/s]
|
| 47 |
85%|βββββββββ | 425/500 [00:02<00:00, 165.82it/s]
|
| 48 |
88%|βββββββββ | 442/500 [00:02<00:00, 166.05it/s]
|
| 49 |
92%|ββββββββββ| 459/500 [00:02<00:00, 165.99it/s]
|
| 50 |
95%|ββββββββββ| 476/500 [00:02<00:00, 165.98it/s]
|
| 51 |
99%|ββββββββββ| 493/500 [00:02<00:00, 166.07it/s]
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
2026-08-10T14:55:12-07:00 serving outputs/release/unhealed/winnow-olmoe-math-keep50 on GPU 1 port 8521 (pp=1 think=template-default)
|
| 2 |
+
2026-08-10T14:55:12-07:00 waiting for server /health ...
|
| 3 |
+
2026-08-10T14:55:42-07:00 server up; chat pass [gsm8k_cot_zeroshot,minerva_math500,ifeval]
|
| 4 |
+
2026-08-10:14:55:50 INFO [_cli.run:388] Selected Tasks: ['gsm8k_cot_zeroshot', 'minerva_math500', 'ifeval']
|
| 5 |
+
2026-08-10:14:55:51 INFO [evaluator:214] Setting random seed to 0 | Setting numpy seed to 1234 | Setting torch manual seed to 1234 | Setting fewshot manual seed to 1234
|
| 6 |
+
2026-08-10:14:55:51 WARNING [evaluator:226] generation_kwargs: {'max_gen_toks': 1280} specified through cli, these settings will update set parameters in yaml tasks. Ensure 'do_sample=True' for non-greedy decoding!
|
| 7 |
+
2026-08-10:14:55:51 INFO [evaluator:239] Initializing local-chat-completions model, with arguments: {'model': 'student', 'base_url': 'http://127.0.0.1:8521/v1/chat/completions', 'num_concurrent': 48, 'tokenized_requests': False, 'max_retries': 3}
|
| 8 |
+
2026-08-10:14:55:51 INFO [models.api_models:179] Using max length 2048 - 1
|
| 9 |
+
2026-08-10:14:55:51 INFO [models.api_models:200] Using tokenizer None
|
| 10 |
+
2026-08-10:14:55:57 INFO [evaluator_utils:446] Selected tasks:
|
| 11 |
+
2026-08-10:14:55:57 INFO [evaluator_utils:480] Task: gsm8k_cot_zeroshot (gsm8k/gsm8k-cot-zeroshot.yaml)
|
| 12 |
+
2026-08-10:14:55:57 INFO [evaluator_utils:480] Task: ifeval (ifeval/ifeval.yaml)
|
| 13 |
+
2026-08-10:14:55:57 INFO [evaluator_utils:480] Task: minerva_math500 (minerva_math/minerva_math500.yaml)
|
| 14 |
+
2026-08-10:14:55:57 INFO [evaluator:314] gsm8k_cot_zeroshot: Using gen_kwargs: {'until': ['Q:', '</s>', '<|im_end|>'], 'do_sample': False, 'max_gen_toks': 1280}
|
| 15 |
+
2026-08-10:14:55:57 INFO [evaluator:314] minerva_math500: Using gen_kwargs: {'until': ['Problem:'], 'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 1280}
|
| 16 |
+
2026-08-10:14:55:57 INFO [evaluator:314] ifeval: Using gen_kwargs: {'until': [], 'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 1280}
|
| 17 |
+
2026-08-10:14:55:57 INFO [api.task:312] Building contexts for gsm8k_cot_zeroshot on rank 0...
|
| 18 |
+
|
| 19 |
0%| | 0/1319 [00:00<?, ?it/s]
|
| 20 |
14%|ββ | 191/1319 [00:00<00:00, 1902.21it/s]
|
| 21 |
29%|βββ | 385/1319 [00:00<00:00, 1923.43it/s]
|
| 22 |
44%|βββββ | 581/1319 [00:00<00:00, 1936.99it/s]
|
| 23 |
59%|ββββββ | 777/1319 [00:00<00:00, 1942.54it/s]
|
| 24 |
74%|ββββββββ | 973/1319 [00:00<00:00, 1945.81it/s]
|
| 25 |
89%|βββββββββ | 1168/1319 [00:00<00:00, 1939.97it/s]
|
| 26 |
+
2026-08-10:14:55:58 INFO [api.task:312] Building contexts for minerva_math500 on rank 0...
|
| 27 |
+
|
| 28 |
0%| | 0/500 [00:00<?, ?it/s]
|
| 29 |
8%|β | 42/500 [00:00<00:01, 411.34it/s]
|
| 30 |
17%|ββ | 84/500 [00:00<00:00, 416.06it/s]
|
| 31 |
25%|βββ | 127/500 [00:00<00:00, 418.71it/s]
|
| 32 |
34%|ββββ | 170/500 [00:00<00:00, 421.00it/s]
|
| 33 |
43%|βββββ | 213/500 [00:00<00:00, 422.35it/s]
|
| 34 |
51%|βββββ | 256/500 [00:00<00:00, 424.21it/s]
|
| 35 |
60%|ββββββ | 299/500 [00:00<00:00, 424.80it/s]
|
| 36 |
68%|βββββββ | 342/500 [00:00<00:00, 424.12it/s]
|
| 37 |
77%|ββββββββ | 385/500 [00:00<00:00, 424.83it/s]
|
| 38 |
86%|βββββββββ | 428/500 [00:01<00:00, 424.33it/s]
|
| 39 |
94%|ββββββββββ| 471/500 [00:01<00:00, 425.60it/s]
|
| 40 |
+
2026-08-10:14:55:59 INFO [api.task:312] Building contexts for ifeval on rank 0...
|
| 41 |
+
|
| 42 |
0%| | 0/541 [00:00<?, ?it/s]
|
| 43 |
+
2026-08-10:14:55:59 INFO [evaluator:585] Running generate_until requests
|
| 44 |
+
2026-08-10:14:55:59 INFO [models.api_models:747] Tokenized requests are disabled. Context + generation length is not checked.
|
| 45 |
+
|
| 46 |
+
|
| 47 |
+
|
| 48 |
+
2026-08-10:15:00:52 INFO [loggers.evaluation_tracker:247] Saving results aggregated
|
| 49 |
+
2026-08-10:15:00:52 INFO [loggers.evaluation_tracker:119] Saving per-task samples to outputs/evals/math_unhealed/glean_winnow-olmoe-math-keep50/student/*.jsonl
|
| 50 |
+
local-chat-completions ({'model': 'student', 'base_url': 'http://127.0.0.1:8521/v1/chat/completions', 'num_concurrent': 48, 'tokenized_requests': False, 'max_retries': 3}), gen_kwargs: ({'max_gen_toks': 1280}), limit: None, num_fewshot: None, batch_size: 1
|
| 51 |
+
| Tasks |Version| Filter |n-shot| Metric | |Value | |Stderr|
|
| 52 |
+
|------------------|------:|----------------|-----:|-----------------------|---|-----:|---|------|
|
| 53 |
+
|gsm8k_cot_zeroshot| 3|flexible-extract| 0|exact_match |β |0.4405|Β± |0.0137|
|
| 54 |
+
| | |strict-match | 0|exact_match |β |0.0000|Β± | 0|
|
| 55 |
+
|ifeval | 4|none | 0|inst_level_loose_acc |β |0.5420|Β± | N/A|
|
| 56 |
+
| | |none | 0|inst_level_strict_acc |β |0.5048|Β± | N/A|
|
| 57 |
+
| | |none | 0|prompt_level_loose_acc |β |0.4177|Β± |0.0212|
|
| 58 |
+
| | |none | 0|prompt_level_strict_acc|β |0.3789|Β± |0.0209|
|
| 59 |
+
|minerva_math500 | 3|none | 4|exact_match |β |0.1760|Β± |0.0170|
|
| 60 |
+
| | |none | 4|math_verify |β |0.1960|Β± |0.0178|
|
| 61 |
+
|
| 62 |
+
2026-08-10T15:00:54-07:00 code pass [humaneval,mbpp] via /v1/completions (function-continuation)
|
| 63 |
+
2026-08-10:15:01:01 INFO [_cli.run:388] Selected Tasks: ['humaneval', 'mbpp']
|
| 64 |
+
2026-08-10:15:01:02 INFO [evaluator:214] Setting random seed to 0 | Setting numpy seed to 1234 | Setting torch manual seed to 1234 | Setting fewshot manual seed to 1234
|
| 65 |
+
2026-08-10:15:01:02 INFO [evaluator:239] Initializing local-completions model, with arguments: {'model': 'student', 'base_url': 'http://127.0.0.1:8521/v1/completions', 'tokenizer': 'outputs/release/unhealed/winnow-olmoe-math-keep50', 'num_concurrent': 48, 'tokenized_requests': False, 'max_retries': 3}
|
| 66 |
+
2026-08-10:15:01:02 INFO [models.openai_completions:42] Remote tokenizer not supported. Using huggingface tokenizer backend.
|
| 67 |
+
2026-08-10:15:01:02 INFO [models.api_models:179] Using max length 2048 - 1
|
| 68 |
+
2026-08-10:15:01:02 INFO [models.api_models:200] Using tokenizer huggingface
|
| 69 |
+
2026-08-10:15:01:10 INFO [evaluator_utils:446] Selected tasks:
|
| 70 |
+
2026-08-10:15:01:10 INFO [evaluator_utils:480] Task: humaneval (humaneval/humaneval.yaml)
|
| 71 |
+
2026-08-10:15:01:10 INFO [evaluator_utils:480] Task: mbpp (mbpp/mbpp.yaml)
|
| 72 |
+
2026-08-10:15:01:10 INFO [evaluator:314] humaneval: Using gen_kwargs: {'until': ['\nclass', '\ndef', '\n#', '\nif', '\nprint'], 'max_gen_toks': 1024, 'do_sample': False}
|
| 73 |
+
2026-08-10:15:01:10 INFO [evaluator:314] mbpp: Using gen_kwargs: {'until': ['[DONE]'], 'do_sample': False}
|
| 74 |
+
2026-08-10:15:01:10 INFO [api.task:312] Building contexts for humaneval on rank 0...
|
| 75 |
+
|
| 76 |
0%| | 0/164 [00:00<?, ?it/s]
|
| 77 |
87%|βββββββββ | 142/164 [00:00<00:00, 1419.11it/s]
|
| 78 |
+
2026-08-10:15:01:10 INFO [api.task:312] Building contexts for mbpp on rank 0...
|
| 79 |
+
|
| 80 |
0%| | 0/500 [00:00<?, ?it/s]
|
| 81 |
3%|β | 17/500 [00:00<00:02, 161.93it/s]
|
| 82 |
7%|β | 34/500 [00:00<00:02, 162.61it/s]
|
| 83 |
10%|β | 51/500 [00:00<00:02, 163.06it/s]
|
| 84 |
14%|ββ | 68/500 [00:00<00:02, 163.43it/s]
|
| 85 |
17%|ββ | 85/500 [00:00<00:02, 163.79it/s]
|
| 86 |
20%|ββ | 102/500 [00:00<00:02, 164.02it/s]
|
| 87 |
24%|βββ | 119/500 [00:00<00:02, 164.32it/s]
|
| 88 |
27%|βββ | 136/500 [00:00<00:02, 164.49it/s]
|
| 89 |
31%|βββ | 153/500 [00:00<00:02, 164.68it/s]
|
| 90 |
34%|ββββ | 170/500 [00:01<00:02, 164.78it/s]
|
| 91 |
37%|ββββ | 187/500 [00:01<00:01, 165.03it/s]
|
| 92 |
41%|ββββ | 204/500 [00:01<00:01, 165.27it/s]
|
| 93 |
44%|βββββ | 221/500 [00:01<00:01, 165.43it/s]
|
| 94 |
48%|βββββ | 238/500 [00:01<00:01, 165.60it/s]
|
| 95 |
51%|βββββ | 255/500 [00:01<00:01, 165.58it/s]
|
| 96 |
54%|ββββββ | 272/500 [00:01<00:01, 165.63it/s]
|
| 97 |
58%|ββββββ | 289/500 [00:01<00:01, 165.63it/s]
|
| 98 |
61%|ββββββ | 306/500 [00:01<00:01, 165.70it/s]
|
| 99 |
65%|βββββββ | 323/500 [00:01<00:01, 165.80it/s]
|
| 100 |
68%|βββββββ | 340/500 [00:02<00:00, 165.89it/s]
|
| 101 |
71%|ββββββββ | 357/500 [00:02<00:00, 165.80it/s]
|
| 102 |
75%|ββββββββ | 374/500 [00:02<00:00, 165.88it/s]
|
| 103 |
78%|ββββββββ | 391/500 [00:02<00:00, 165.74it/s]
|
| 104 |
82%|βββββββββ | 408/500 [00:02<00:00, 165.84it/s]
|
| 105 |
85%|βββββββββ | 425/500 [00:02<00:00, 165.82it/s]
|
| 106 |
88%|βββββββββ | 442/500 [00:02<00:00, 166.05it/s]
|
| 107 |
92%|ββββββββββ| 459/500 [00:02<00:00, 165.99it/s]
|
| 108 |
95%|ββββββββββ| 476/500 [00:02<00:00, 165.98it/s]
|
| 109 |
99%|ββββββββββ| 493/500 [00:02<00:00, 166.07it/s]
|
| 110 |
+
2026-08-10:15:01:13 INFO [evaluator:585] Running generate_until requests
|
| 111 |
+
2026-08-10:15:01:13 INFO [models.api_models:747] Tokenized requests are disabled. Context + generation length is not checked.
|
| 112 |
+
|
| 113 |
+
|
| 114 |
+
2026-08-10:15:05:32 INFO [loggers.evaluation_tracker:247] Saving results aggregated
|
| 115 |
+
2026-08-10:15:05:32 INFO [loggers.evaluation_tracker:119] Saving per-task samples to outputs/evals/math_unhealed/glean_winnow-olmoe-math-keep50/student/*.jsonl
|
| 116 |
+
local-completions ({'model': 'student', 'base_url': 'http://127.0.0.1:8521/v1/completions', 'tokenizer': 'outputs/release/unhealed/winnow-olmoe-math-keep50', 'num_concurrent': 48, 'tokenized_requests': False, 'max_retries': 3}), gen_kwargs: ({}), limit: None, num_fewshot: None, batch_size: 1
|
| 117 |
+
| Tasks |Version| Filter |n-shot| Metric | |Value | |Stderr|
|
| 118 |
+
|---------|------:|-----------|-----:|---------|---|-----:|---|-----:|
|
| 119 |
+
|humaneval| 1|create_test| 0|pass@1 |β |0.0732|Β± |0.0204|
|
| 120 |
+
|mbpp | 1|none | 3|pass_at_1|β |0.0900|Β± |0.0128|
|
| 121 |
+
|
| 122 |
+
2026-08-10T15:05:33-07:00 lm_eval exit=0 -> outputs/evals/math_unhealed/glean_winnow-olmoe-math-keep50
|
evals/math_unhealed/glean_winnow-olmoe-math-keep75.eval.log
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
evals/math_unhealed/reap_reap-math-keep25.eval.log
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
evals/math_unhealed/reap_reap-math-keep50.eval.log
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|