hbfreed commited on
Commit
8717f59
Β·
verified Β·
1 Parent(s): c7b14e3

Add files using upload-large-folder tool

Browse files
This view is limited to 50 files because it contains too many changes. Β  See raw diff
Files changed (50) hide show
  1. evals/grid_math_unhealed/glean_keep25_unhealed.log +8 -0
  2. evals/grid_math_unhealed/glean_keep25_unhealed_chat.json +0 -0
  3. evals/grid_math_unhealed/glean_keep25_unhealed_chat.json.server.log +109 -0
  4. evals/grid_math_unhealed/glean_keep50_unhealed.log +8 -0
  5. evals/grid_math_unhealed/glean_keep50_unhealed_chat.json +0 -0
  6. evals/grid_math_unhealed/glean_keep50_unhealed_chat.json.server.log +102 -0
  7. evals/grid_math_unhealed/glean_keep75_unhealed.log +8 -0
  8. evals/grid_math_unhealed/glean_keep75_unhealed_chat.json +0 -0
  9. evals/grid_math_unhealed/glean_keep75_unhealed_chat.json.server.log +103 -0
  10. evals/grid_math_unhealed/reap_keep25_unhealed.log +8 -0
  11. evals/grid_math_unhealed/reap_keep25_unhealed_chat.json +0 -0
  12. evals/grid_math_unhealed/reap_keep25_unhealed_chat.json.server.log +117 -0
  13. evals/grid_math_unhealed/reap_keep50_unhealed.log +8 -0
  14. evals/grid_math_unhealed/reap_keep50_unhealed_chat.json +0 -0
  15. evals/grid_math_unhealed/reap_keep50_unhealed_chat.json.server.log +110 -0
  16. evals/grid_math_unhealed/reap_keep75_unhealed.log +8 -0
  17. evals/grid_math_unhealed/reap_keep75_unhealed_chat.json +0 -0
  18. evals/grid_math_unhealed/reap_keep75_unhealed_chat.json.server.log +105 -0
  19. evals/grid_math_unhealed/uniform_keep25_unhealed.log +8 -0
  20. evals/grid_math_unhealed/uniform_keep25_unhealed_chat.json +0 -0
  21. evals/grid_math_unhealed/uniform_keep25_unhealed_chat.json.server.log +115 -0
  22. evals/grid_math_unhealed/uniform_keep50_unhealed.log +8 -0
  23. evals/grid_math_unhealed/uniform_keep50_unhealed_chat.json +0 -0
  24. evals/grid_math_unhealed/uniform_keep50_unhealed_chat.json.server.log +101 -0
  25. evals/grid_math_unhealed/uniform_keep75_unhealed.log +8 -0
  26. evals/grid_math_unhealed/uniform_keep75_unhealed_chat.json +0 -0
  27. evals/grid_math_unhealed/uniform_keep75_unhealed_chat.json.server.log +103 -0
  28. evals/healing_breadth/glean_math_keep25_seed1224.json +0 -0
  29. evals/healing_breadth/glean_math_keep25_seed1224.json.server.log +139 -0
  30. evals/healing_breadth/glean_math_keep25_seed1224_long768_step0500_chat.json +0 -0
  31. evals/healing_breadth/glean_math_keep25_seed1224_long768_step0500_chat.json.server.log +255 -0
  32. evals/healing_breadth/glean_math_keep25_seed1224_long768_step0500_raw.json +0 -0
  33. evals/healing_breadth/glean_math_keep25_seed1224_long768_step0500_raw.json.server.log +377 -0
  34. evals/healing_breadth/glean_math_keep25_seed1224_step0050_chat.json +0 -0
  35. evals/healing_breadth/glean_math_keep25_seed1224_step0050_chat.json.server.log +138 -0
  36. evals/healing_breadth/glean_math_keep25_seed1224_step0050_raw.json +0 -0
  37. evals/healing_breadth/glean_math_keep25_seed1224_step0050_raw.json.server.log +260 -0
  38. evals/healing_breadth/glean_math_keep75_oneshot_chat.json +0 -0
  39. evals/healing_breadth/glean_math_keep75_oneshot_chat.json.server.log +130 -0
  40. evals/healing_breadth/glean_math_keep75_oneshot_raw.json +0 -0
  41. evals/healing_breadth/glean_math_keep75_oneshot_raw.json.server.log +129 -0
  42. evals/healing_breadth/reap_math_keep75_seed1224.json +0 -0
  43. evals/healing_breadth/reap_math_keep75_seed1224.json.server.log +128 -0
  44. evals/healing_breadth/uniform_math_keep50_seed1224.json +0 -0
  45. evals/healing_breadth/uniform_math_keep50_seed1224.json.server.log +130 -0
  46. evals/math_unhealed/glean_winnow-olmoe-math-keep25.eval.log +0 -0
  47. evals/math_unhealed/glean_winnow-olmoe-math-keep50.eval.log +70 -0
  48. evals/math_unhealed/glean_winnow-olmoe-math-keep75.eval.log +0 -0
  49. evals/math_unhealed/reap_reap-math-keep25.eval.log +0 -0
  50. evals/math_unhealed/reap_reap-math-keep50.eval.log +0 -0
evals/grid_math_unhealed/glean_keep25_unhealed.log ADDED
@@ -0,0 +1,8 @@
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "correct": 143,
3
+ "accuracy": 0.10841546626231995,
4
+ "finished": 588,
5
+ "finish_rate": 0.44579226686884005,
6
+ "mean_completion_tokens": 327.01440485216074
7
+ }
8
+ saved item-level results -> outputs/evals/grid_math_unhealed/glean_keep25_unhealed_chat.json
evals/grid_math_unhealed/glean_keep25_unhealed_chat.json ADDED
The diff for this file is too large to render. See raw diff
 
evals/grid_math_unhealed/glean_keep25_unhealed_chat.json.server.log ADDED
@@ -0,0 +1,109 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ Skipping import of cpp extensions due to incompatible torch version 2.10.0+cu128 for torchao version 0.15.0 Please see https://github.com/pytorch/ao/issues/2919 for more info
2
+ WARNING 08-10 21:24:57 [registry.py:915] Model architecture OlmoeForCausalLM is already registered, and will be overwritten by the new model class glean_vllm.pruned_olmoe:PrunedOlmoeForCausalLM.
3
+ (APIServer pid=1153466) INFO 08-10 21:24:57 [utils.py:299]
4
+ (APIServer pid=1153466) INFO 08-10 21:24:57 [utils.py:299] β–ˆ β–ˆ β–ˆβ–„ β–„β–ˆ
5
+ (APIServer pid=1153466) INFO 08-10 21:24:57 [utils.py:299] β–„β–„ β–„β–ˆ β–ˆ β–ˆ β–ˆ β–€β–„β–€ β–ˆ version 0.19.0
6
+ (APIServer pid=1153466) INFO 08-10 21:24:57 [utils.py:299] β–ˆβ–„β–ˆβ–€ β–ˆ β–ˆ β–ˆ β–ˆ model outputs/release/unhealed/winnow-olmoe-math-keep25
7
+ (APIServer pid=1153466) INFO 08-10 21:24:57 [utils.py:299] β–€β–€ β–€β–€β–€β–€β–€ β–€β–€β–€β–€β–€ β–€ β–€
8
+ (APIServer pid=1153466) INFO 08-10 21:24:57 [utils.py:299]
9
+ (APIServer pid=1153466) INFO 08-10 21:24:57 [utils.py:233] non-default args: {'model_tag': 'outputs/release/unhealed/winnow-olmoe-math-keep25', 'host': '127.0.0.1', 'port': 8610, 'model': 'outputs/release/unhealed/winnow-olmoe-math-keep25', 'max_model_len': 2048, 'enforce_eager': True, 'served_model_name': ['student'], 'gpu_memory_utilization': 0.85}
10
+ (APIServer pid=1153466) INFO 08-10 21:25:05 [model.py:549] Resolved architecture: PrunedOlmoeForCausalLM
11
+ (APIServer pid=1153466) INFO 08-10 21:25:05 [model.py:1678] Using max model len 2048
12
+ (APIServer pid=1153466) INFO 08-10 21:25:05 [vllm.py:790] Asynchronous scheduling is enabled.
13
+ (APIServer pid=1153466) WARNING 08-10 21:25:05 [vllm.py:848] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
14
+ (APIServer pid=1153466) WARNING 08-10 21:25:05 [vllm.py:859] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
15
+ (APIServer pid=1153466) INFO 08-10 21:25:05 [vllm.py:1025] Cudagraph is disabled under eager mode
16
+ (APIServer pid=1153466) INFO 08-10 21:25:05 [compilation.py:290] Enabled custom fusions: norm_quant, act_quant
17
+ Skipping import of cpp extensions due to incompatible torch version 2.10.0+cu128 for torchao version 0.15.0 Please see https://github.com/pytorch/ao/issues/2919 for more info
18
+ (EngineCore pid=1154734) WARNING 08-10 21:25:14 [registry.py:915] Model architecture OlmoeForCausalLM is already registered, and will be overwritten by the new model class glean_vllm.pruned_olmoe:PrunedOlmoeForCausalLM.
19
+ (EngineCore pid=1154734) INFO 08-10 21:25:14 [core.py:105] Initializing a V1 LLM engine (v0.19.0) with config: model='outputs/release/unhealed/winnow-olmoe-math-keep25', speculative_config=None, tokenizer='outputs/release/unhealed/winnow-olmoe-math-keep25', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=False, dtype=torch.bfloat16, max_seq_len=2048, download_dir=None, load_format=auto, tensor_parallel_size=1, pipeline_parallel_size=1, data_parallel_size=1, decode_context_parallel_size=1, dcp_comm_backend=ag_rs, disable_custom_all_reduce=False, quantization=None, enforce_eager=True, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False), seed=0, served_model_name=student, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'mode': <CompilationMode.NONE: 0>, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'splitting_ops': [], 'compile_mm_encoder': False, 'cudagraph_mm_encoder': False, 'encoder_cudagraph_token_budgets': [], 'encoder_cudagraph_max_images_per_batch': 0, 'compile_sizes': [], 'compile_ranges_endpoints': [2048], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'size_asserts': False, 'alignment_asserts': False, 'scalar_asserts': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': <CUDAGraphMode.NONE: 0>, 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': True, 'fuse_act_quant': True, 'fuse_attn_quant': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False}, 'max_cudagraph_capture_size': 0, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': True, 'static_all_moe_layers': []}
20
+ (EngineCore pid=1154734) INFO 08-10 21:25:14 [parallel_state.py:1400] world_size=1 rank=0 local_rank=0 distributed_init_method=tcp://192.168.0.27:46169 backend=nccl
21
+ (EngineCore pid=1154734) INFO 08-10 21:25:14 [parallel_state.py:1716] rank 0 in world size 1 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank 0, EPLB rank N/A
22
+ (EngineCore pid=1154734) INFO 08-10 21:25:15 [gpu_model_runner.py:4735] Starting to load model outputs/release/unhealed/winnow-olmoe-math-keep25...
23
+ (EngineCore pid=1154734) INFO 08-10 21:25:16 [cuda.py:334] Using FLASH_ATTN attention backend out of potential backends: ['FLASH_ATTN', 'FLASHINFER', 'TRITON_ATTN', 'FLEX_ATTENTION'].
24
+ (EngineCore pid=1154734) INFO 08-10 21:25:16 [flash_attn.py:596] Using FlashAttention version 2
25
+ (EngineCore pid=1154734)
26
+ (EngineCore pid=1154734)
27
+ (EngineCore pid=1154734)
28
+ (EngineCore pid=1154734)
29
+ (EngineCore pid=1154734) INFO 08-10 21:25:22 [default_loader.py:384] Loading weights took 5.84 seconds
30
+ (EngineCore pid=1154734) INFO 08-10 21:25:22 [gpu_model_runner.py:4820] Model loading took 3.89 GiB memory and 6.430009 seconds
31
+ (EngineCore pid=1154734) INFO 08-10 21:25:24 [gpu_worker.py:436] Available KV cache memory: 15.86 GiB
32
+ (EngineCore pid=1154734) INFO 08-10 21:25:24 [kv_cache_utils.py:1319] GPU KV cache size: 129,920 tokens
33
+ (EngineCore pid=1154734) INFO 08-10 21:25:24 [kv_cache_utils.py:1324] Maximum concurrency for 2,048 tokens per request: 63.44x
34
+ (EngineCore pid=1154734) INFO 08-10 21:25:24 [core.py:283] init engine (profile, create kv cache, warmup model) took 1.67 seconds
35
+ (EngineCore pid=1154734) INFO 08-10 21:25:25 [vllm.py:790] Asynchronous scheduling is enabled.
36
+ (EngineCore pid=1154734) WARNING 08-10 21:25:25 [vllm.py:848] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
37
+ (EngineCore pid=1154734) WARNING 08-10 21:25:25 [vllm.py:859] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
38
+ (EngineCore pid=1154734) INFO 08-10 21:25:25 [vllm.py:1025] Cudagraph is disabled under eager mode
39
+ (EngineCore pid=1154734) INFO 08-10 21:25:25 [compilation.py:290] Enabled custom fusions: norm_quant, act_quant
40
+ (APIServer pid=1153466) INFO 08-10 21:25:25 [api_server.py:590] Supported tasks: ['generate']
41
+ (APIServer pid=1153466) WARNING 08-10 21:25:25 [__init__.py:14] SECURITY WARNING: Development endpoints are enabled! This should NOT be used in production!
42
+ (APIServer pid=1153466) INFO 08-10 21:25:25 [hf.py:314] Detected the chat template content format to be 'string'. You can set `--chat-template-content-format` to override this.
43
+ (APIServer pid=1153466) INFO 08-10 21:25:25 [api_server.py:594] Starting vLLM server on http://127.0.0.1:8610
44
+ (APIServer pid=1153466) INFO 08-10 21:25:25 [launcher.py:37] Available routes are:
45
+ (APIServer pid=1153466) INFO 08-10 21:25:25 [launcher.py:46] Route: /openapi.json, Methods: GET, HEAD
46
+ (APIServer pid=1153466) INFO 08-10 21:25:25 [launcher.py:46] Route: /docs, Methods: GET, HEAD
47
+ (APIServer pid=1153466) INFO 08-10 21:25:25 [launcher.py:46] Route: /docs/oauth2-redirect, Methods: GET, HEAD
48
+ (APIServer pid=1153466) INFO 08-10 21:25:25 [launcher.py:46] Route: /redoc, Methods: GET, HEAD
49
+ (APIServer pid=1153466) INFO 08-10 21:25:25 [launcher.py:46] Route: /sleep, Methods: POST
50
+ (APIServer pid=1153466) INFO 08-10 21:25:25 [launcher.py:46] Route: /wake_up, Methods: POST
51
+ (APIServer pid=1153466) INFO 08-10 21:25:25 [launcher.py:46] Route: /is_sleeping, Methods: GET
52
+ (APIServer pid=1153466) INFO 08-10 21:25:25 [launcher.py:46] Route: /collective_rpc, Methods: POST
53
+ (APIServer pid=1153466) INFO 08-10 21:25:25 [launcher.py:46] Route: /reset_prefix_cache, Methods: POST
54
+ (APIServer pid=1153466) INFO 08-10 21:25:25 [launcher.py:46] Route: /reset_mm_cache, Methods: POST
55
+ (APIServer pid=1153466) INFO 08-10 21:25:25 [launcher.py:46] Route: /reset_encoder_cache, Methods: POST
56
+ (APIServer pid=1153466) INFO 08-10 21:25:25 [launcher.py:46] Route: /tokenize, Methods: POST
57
+ (APIServer pid=1153466) INFO 08-10 21:25:25 [launcher.py:46] Route: /detokenize, Methods: POST
58
+ (APIServer pid=1153466) INFO 08-10 21:25:25 [launcher.py:46] Route: /load, Methods: GET
59
+ (APIServer pid=1153466) INFO 08-10 21:25:25 [launcher.py:46] Route: /version, Methods: GET
60
+ (APIServer pid=1153466) INFO 08-10 21:25:25 [launcher.py:46] Route: /health, Methods: GET
61
+ (APIServer pid=1153466) INFO 08-10 21:25:25 [launcher.py:46] Route: /metrics, Methods: GET
62
+ (APIServer pid=1153466) INFO 08-10 21:25:25 [launcher.py:46] Route: /server_info, Methods: GET
63
+ (APIServer pid=1153466) INFO 08-10 21:25:25 [launcher.py:46] Route: /v1/models, Methods: GET
64
+ (APIServer pid=1153466) INFO 08-10 21:25:25 [launcher.py:46] Route: /ping, Methods: GET
65
+ (APIServer pid=1153466) INFO 08-10 21:25:25 [launcher.py:46] Route: /ping, Methods: POST
66
+ (APIServer pid=1153466) INFO 08-10 21:25:25 [launcher.py:46] Route: /invocations, Methods: POST
67
+ (APIServer pid=1153466) INFO 08-10 21:25:25 [launcher.py:46] Route: /v1/chat/completions, Methods: POST
68
+ (APIServer pid=1153466) INFO 08-10 21:25:25 [launcher.py:46] Route: /v1/chat/completions/batch, Methods: POST
69
+ (APIServer pid=1153466) INFO 08-10 21:25:25 [launcher.py:46] Route: /v1/responses, Methods: POST
70
+ (APIServer pid=1153466) INFO 08-10 21:25:25 [launcher.py:46] Route: /v1/responses/{response_id}, Methods: GET
71
+ (APIServer pid=1153466) INFO 08-10 21:25:25 [launcher.py:46] Route: /v1/responses/{response_id}/cancel, Methods: POST
72
+ (APIServer pid=1153466) INFO 08-10 21:25:25 [launcher.py:46] Route: /v1/completions, Methods: POST
73
+ (APIServer pid=1153466) INFO 08-10 21:25:25 [launcher.py:46] Route: /v1/messages, Methods: POST
74
+ (APIServer pid=1153466) INFO 08-10 21:25:25 [launcher.py:46] Route: /v1/messages/count_tokens, Methods: POST
75
+ (APIServer pid=1153466) INFO 08-10 21:25:25 [launcher.py:46] Route: /inference/v1/generate, Methods: POST
76
+ (APIServer pid=1153466) INFO 08-10 21:25:25 [launcher.py:46] Route: /pause, Methods: POST
77
+ (APIServer pid=1153466) INFO 08-10 21:25:25 [launcher.py:46] Route: /resume, Methods: POST
78
+ (APIServer pid=1153466) INFO 08-10 21:25:25 [launcher.py:46] Route: /is_paused, Methods: GET
79
+ (APIServer pid=1153466) INFO 08-10 21:25:25 [launcher.py:46] Route: /init_weight_transfer_engine, Methods: POST
80
+ (APIServer pid=1153466) INFO 08-10 21:25:25 [launcher.py:46] Route: /update_weights, Methods: POST
81
+ (APIServer pid=1153466) INFO 08-10 21:25:25 [launcher.py:46] Route: /get_world_size, Methods: GET
82
+ (APIServer pid=1153466) INFO 08-10 21:25:25 [launcher.py:46] Route: /scale_elastic_ep, Methods: POST
83
+ (APIServer pid=1153466) INFO 08-10 21:25:25 [launcher.py:46] Route: /is_scaling_elastic_ep, Methods: POST
84
+ (APIServer pid=1153466) INFO 08-10 21:25:25 [launcher.py:46] Route: /v1/chat/completions/render, Methods: POST
85
+ (APIServer pid=1153466) INFO 08-10 21:25:25 [launcher.py:46] Route: /v1/completions/render, Methods: POST
86
+ (APIServer pid=1153466) INFO: Started server process [1153466]
87
+ (APIServer pid=1153466) INFO: Waiting for application startup.
88
+ (APIServer pid=1153466) INFO: Application startup complete.
89
+ (APIServer pid=1153466) INFO: 127.0.0.1:55122 - "GET /health HTTP/1.1" 200 OK
90
+ (APIServer pid=1153466) INFO 08-10 21:25:35 [loggers.py:259] Engine 000: Avg prompt throughput: 2593.9 tokens/s, Avg generation throughput: 2526.0 tokens/s, Running: 256 reqs, Waiting: 983 reqs, GPU KV cache usage: 35.2%, Prefix cache hit rate: 93.2%
91
+ (APIServer pid=1153466) INFO 08-10 21:25:45 [loggers.py:259] Engine 000: Avg prompt throughput: 519.2 tokens/s, Avg generation throughput: 3397.7 tokens/s, Running: 256 reqs, Waiting: 915 reqs, GPU KV cache usage: 56.2%, Prefix cache hit rate: 93.2%
92
+ (APIServer pid=1153466) INFO 08-10 21:25:55 [loggers.py:259] Engine 000: Avg prompt throughput: 239.9 tokens/s, Avg generation throughput: 3248.2 tokens/s, Running: 256 reqs, Waiting: 890 reqs, GPU KV cache usage: 78.8%, Prefix cache hit rate: 93.2%
93
+ (APIServer pid=1153466) INFO 08-10 21:26:05 [loggers.py:259] Engine 000: Avg prompt throughput: 127.4 tokens/s, Avg generation throughput: 3094.9 tokens/s, Running: 256 reqs, Waiting: 874 reqs, GPU KV cache usage: 100.0%, Prefix cache hit rate: 93.2%
94
+ (APIServer pid=1153466) INFO 08-10 21:26:15 [loggers.py:259] Engine 000: Avg prompt throughput: 1839.2 tokens/s, Avg generation throughput: 3085.6 tokens/s, Running: 254 reqs, Waiting: 632 reqs, GPU KV cache usage: 47.1%, Prefix cache hit rate: 93.3%
95
+ (APIServer pid=1153466) INFO 08-10 21:26:25 [loggers.py:259] Engine 000: Avg prompt throughput: 930.8 tokens/s, Avg generation throughput: 3341.7 tokens/s, Running: 256 reqs, Waiting: 517 reqs, GPU KV cache usage: 51.5%, Prefix cache hit rate: 93.3%
96
+ (APIServer pid=1153466) INFO 08-10 21:26:35 [loggers.py:259] Engine 000: Avg prompt throughput: 415.0 tokens/s, Avg generation throughput: 3296.3 tokens/s, Running: 256 reqs, Waiting: 466 reqs, GPU KV cache usage: 68.5%, Prefix cache hit rate: 93.3%
97
+ (APIServer pid=1153466) INFO 08-10 21:26:45 [loggers.py:259] Engine 000: Avg prompt throughput: 189.8 tokens/s, Avg generation throughput: 3169.8 tokens/s, Running: 256 reqs, Waiting: 441 reqs, GPU KV cache usage: 88.3%, Prefix cache hit rate: 93.3%
98
+ (APIServer pid=1153466) INFO 08-10 21:26:55 [loggers.py:259] Engine 000: Avg prompt throughput: 1409.8 tokens/s, Avg generation throughput: 3105.7 tokens/s, Running: 256 reqs, Waiting: 268 reqs, GPU KV cache usage: 57.2%, Prefix cache hit rate: 93.4%
99
+ (APIServer pid=1153466) INFO 08-10 21:27:05 [loggers.py:259] Engine 000: Avg prompt throughput: 992.5 tokens/s, Avg generation throughput: 3263.5 tokens/s, Running: 256 reqs, Waiting: 141 reqs, GPU KV cache usage: 52.8%, Prefix cache hit rate: 93.4%
100
+ (APIServer pid=1153466) INFO 08-10 21:27:15 [loggers.py:259] Engine 000: Avg prompt throughput: 697.8 tokens/s, Avg generation throughput: 3293.6 tokens/s, Running: 256 reqs, Waiting: 57 reqs, GPU KV cache usage: 62.2%, Prefix cache hit rate: 93.4%
101
+ (APIServer pid=1153466) INFO 08-10 21:27:25 [loggers.py:259] Engine 000: Avg prompt throughput: 419.7 tokens/s, Avg generation throughput: 3194.6 tokens/s, Running: 256 reqs, Waiting: 5 reqs, GPU KV cache usage: 78.3%, Prefix cache hit rate: 93.4%
102
+ (APIServer pid=1153466) INFO 08-10 21:27:35 [loggers.py:259] Engine 000: Avg prompt throughput: 37.3 tokens/s, Avg generation throughput: 3053.0 tokens/s, Running: 123 reqs, Waiting: 0 reqs, GPU KV cache usage: 42.5%, Prefix cache hit rate: 93.4%
103
+ (APIServer pid=1153466) INFO: 127.0.0.1:55138 - "POST /v1/completions HTTP/1.1" 200 OK
104
+ (EngineCore pid=1154734) INFO 08-10 21:27:45 [core.py:1210] Shutdown initiated (timeout=0)
105
+ (EngineCore pid=1154734) INFO 08-10 21:27:45 [core.py:1233] Shutdown complete
106
+ (APIServer pid=1153466) INFO: Shutting down
107
+ (APIServer pid=1153466) INFO: Waiting for application shutdown.
108
+ (APIServer pid=1153466) INFO: Application shutdown complete.
109
+ (APIServer pid=1153466) INFO: Finished server process [1153466]
evals/grid_math_unhealed/glean_keep50_unhealed.log ADDED
@@ -0,0 +1,8 @@
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "correct": 773,
3
+ "accuracy": 0.5860500379075056,
4
+ "finished": 1313,
5
+ "finish_rate": 0.9954510993176648,
6
+ "mean_completion_tokens": 115.38665655799848
7
+ }
8
+ saved item-level results -> outputs/evals/grid_math_unhealed/glean_keep50_unhealed_chat.json
evals/grid_math_unhealed/glean_keep50_unhealed_chat.json ADDED
The diff for this file is too large to render. See raw diff
 
evals/grid_math_unhealed/glean_keep50_unhealed_chat.json.server.log ADDED
@@ -0,0 +1,102 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ Skipping import of cpp extensions due to incompatible torch version 2.10.0+cu128 for torchao version 0.15.0 Please see https://github.com/pytorch/ao/issues/2919 for more info
2
+ WARNING 08-10 21:22:07 [registry.py:915] Model architecture OlmoeForCausalLM is already registered, and will be overwritten by the new model class glean_vllm.pruned_olmoe:PrunedOlmoeForCausalLM.
3
+ (APIServer pid=1147686) INFO 08-10 21:22:07 [utils.py:299]
4
+ (APIServer pid=1147686) INFO 08-10 21:22:07 [utils.py:299] β–ˆ β–ˆ β–ˆβ–„ β–„β–ˆ
5
+ (APIServer pid=1147686) INFO 08-10 21:22:07 [utils.py:299] β–„β–„ β–„β–ˆ β–ˆ β–ˆ β–ˆ β–€β–„β–€ β–ˆ version 0.19.0
6
+ (APIServer pid=1147686) INFO 08-10 21:22:07 [utils.py:299] β–ˆβ–„β–ˆβ–€ β–ˆ β–ˆ β–ˆ β–ˆ model outputs/release/unhealed/winnow-olmoe-math-keep50
7
+ (APIServer pid=1147686) INFO 08-10 21:22:07 [utils.py:299] β–€β–€ β–€β–€β–€β–€β–€ β–€β–€β–€β–€β–€ β–€ β–€
8
+ (APIServer pid=1147686) INFO 08-10 21:22:07 [utils.py:299]
9
+ (APIServer pid=1147686) INFO 08-10 21:22:07 [utils.py:233] non-default args: {'model_tag': 'outputs/release/unhealed/winnow-olmoe-math-keep50', 'host': '127.0.0.1', 'port': 8610, 'model': 'outputs/release/unhealed/winnow-olmoe-math-keep50', 'max_model_len': 2048, 'enforce_eager': True, 'served_model_name': ['student'], 'gpu_memory_utilization': 0.85}
10
+ (APIServer pid=1147686) INFO 08-10 21:22:15 [model.py:549] Resolved architecture: PrunedOlmoeForCausalLM
11
+ (APIServer pid=1147686) INFO 08-10 21:22:15 [model.py:1678] Using max model len 2048
12
+ (APIServer pid=1147686) INFO 08-10 21:22:15 [vllm.py:790] Asynchronous scheduling is enabled.
13
+ (APIServer pid=1147686) WARNING 08-10 21:22:15 [vllm.py:848] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
14
+ (APIServer pid=1147686) WARNING 08-10 21:22:15 [vllm.py:859] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
15
+ (APIServer pid=1147686) INFO 08-10 21:22:15 [vllm.py:1025] Cudagraph is disabled under eager mode
16
+ (APIServer pid=1147686) INFO 08-10 21:22:15 [compilation.py:290] Enabled custom fusions: norm_quant, act_quant
17
+ Skipping import of cpp extensions due to incompatible torch version 2.10.0+cu128 for torchao version 0.15.0 Please see https://github.com/pytorch/ao/issues/2919 for more info
18
+ (EngineCore pid=1148963) WARNING 08-10 21:22:24 [registry.py:915] Model architecture OlmoeForCausalLM is already registered, and will be overwritten by the new model class glean_vllm.pruned_olmoe:PrunedOlmoeForCausalLM.
19
+ (EngineCore pid=1148963) INFO 08-10 21:22:24 [core.py:105] Initializing a V1 LLM engine (v0.19.0) with config: model='outputs/release/unhealed/winnow-olmoe-math-keep50', speculative_config=None, tokenizer='outputs/release/unhealed/winnow-olmoe-math-keep50', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=False, dtype=torch.bfloat16, max_seq_len=2048, download_dir=None, load_format=auto, tensor_parallel_size=1, pipeline_parallel_size=1, data_parallel_size=1, decode_context_parallel_size=1, dcp_comm_backend=ag_rs, disable_custom_all_reduce=False, quantization=None, enforce_eager=True, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False), seed=0, served_model_name=student, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'mode': <CompilationMode.NONE: 0>, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'splitting_ops': [], 'compile_mm_encoder': False, 'cudagraph_mm_encoder': False, 'encoder_cudagraph_token_budgets': [], 'encoder_cudagraph_max_images_per_batch': 0, 'compile_sizes': [], 'compile_ranges_endpoints': [2048], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'size_asserts': False, 'alignment_asserts': False, 'scalar_asserts': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': <CUDAGraphMode.NONE: 0>, 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': True, 'fuse_act_quant': True, 'fuse_attn_quant': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False}, 'max_cudagraph_capture_size': 0, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': True, 'static_all_moe_layers': []}
20
+ (EngineCore pid=1148963) INFO 08-10 21:22:24 [parallel_state.py:1400] world_size=1 rank=0 local_rank=0 distributed_init_method=tcp://192.168.0.27:43405 backend=nccl
21
+ (EngineCore pid=1148963) INFO 08-10 21:22:24 [parallel_state.py:1716] rank 0 in world size 1 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank 0, EPLB rank N/A
22
+ (EngineCore pid=1148963) INFO 08-10 21:22:25 [gpu_model_runner.py:4735] Starting to load model outputs/release/unhealed/winnow-olmoe-math-keep50...
23
+ (EngineCore pid=1148963) INFO 08-10 21:22:25 [cuda.py:334] Using FLASH_ATTN attention backend out of potential backends: ['FLASH_ATTN', 'FLASHINFER', 'TRITON_ATTN', 'FLEX_ATTENTION'].
24
+ (EngineCore pid=1148963) INFO 08-10 21:22:25 [flash_attn.py:596] Using FlashAttention version 2
25
+ (EngineCore pid=1148963)
26
+ (EngineCore pid=1148963)
27
+ (EngineCore pid=1148963)
28
+ (EngineCore pid=1148963)
29
+ (EngineCore pid=1148963)
30
+ (EngineCore pid=1148963) INFO 08-10 21:22:27 [default_loader.py:384] Loading weights took 1.42 seconds
31
+ (EngineCore pid=1148963) INFO 08-10 21:22:27 [gpu_model_runner.py:4820] Model loading took 6.89 GiB memory and 2.012000 seconds
32
+ (EngineCore pid=1148963) INFO 08-10 21:22:29 [gpu_worker.py:436] Available KV cache memory: 12.86 GiB
33
+ (EngineCore pid=1148963) INFO 08-10 21:22:29 [kv_cache_utils.py:1319] GPU KV cache size: 105,328 tokens
34
+ (EngineCore pid=1148963) INFO 08-10 21:22:29 [kv_cache_utils.py:1324] Maximum concurrency for 2,048 tokens per request: 51.43x
35
+ (EngineCore pid=1148963) INFO 08-10 21:22:29 [core.py:283] init engine (profile, create kv cache, warmup model) took 1.80 seconds
36
+ (EngineCore pid=1148963) INFO 08-10 21:22:30 [vllm.py:790] Asynchronous scheduling is enabled.
37
+ (EngineCore pid=1148963) WARNING 08-10 21:22:30 [vllm.py:848] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
38
+ (EngineCore pid=1148963) WARNING 08-10 21:22:30 [vllm.py:859] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
39
+ (EngineCore pid=1148963) INFO 08-10 21:22:30 [vllm.py:1025] Cudagraph is disabled under eager mode
40
+ (EngineCore pid=1148963) INFO 08-10 21:22:30 [compilation.py:290] Enabled custom fusions: norm_quant, act_quant
41
+ (APIServer pid=1147686) INFO 08-10 21:22:30 [api_server.py:590] Supported tasks: ['generate']
42
+ (APIServer pid=1147686) WARNING 08-10 21:22:30 [__init__.py:14] SECURITY WARNING: Development endpoints are enabled! This should NOT be used in production!
43
+ (APIServer pid=1147686) INFO 08-10 21:22:30 [hf.py:314] Detected the chat template content format to be 'string'. You can set `--chat-template-content-format` to override this.
44
+ (APIServer pid=1147686) INFO 08-10 21:22:30 [api_server.py:594] Starting vLLM server on http://127.0.0.1:8610
45
+ (APIServer pid=1147686) INFO 08-10 21:22:30 [launcher.py:37] Available routes are:
46
+ (APIServer pid=1147686) INFO 08-10 21:22:30 [launcher.py:46] Route: /openapi.json, Methods: HEAD, GET
47
+ (APIServer pid=1147686) INFO 08-10 21:22:30 [launcher.py:46] Route: /docs, Methods: HEAD, GET
48
+ (APIServer pid=1147686) INFO 08-10 21:22:30 [launcher.py:46] Route: /docs/oauth2-redirect, Methods: HEAD, GET
49
+ (APIServer pid=1147686) INFO 08-10 21:22:30 [launcher.py:46] Route: /redoc, Methods: HEAD, GET
50
+ (APIServer pid=1147686) INFO 08-10 21:22:30 [launcher.py:46] Route: /sleep, Methods: POST
51
+ (APIServer pid=1147686) INFO 08-10 21:22:30 [launcher.py:46] Route: /wake_up, Methods: POST
52
+ (APIServer pid=1147686) INFO 08-10 21:22:30 [launcher.py:46] Route: /is_sleeping, Methods: GET
53
+ (APIServer pid=1147686) INFO 08-10 21:22:30 [launcher.py:46] Route: /collective_rpc, Methods: POST
54
+ (APIServer pid=1147686) INFO 08-10 21:22:30 [launcher.py:46] Route: /reset_prefix_cache, Methods: POST
55
+ (APIServer pid=1147686) INFO 08-10 21:22:30 [launcher.py:46] Route: /reset_mm_cache, Methods: POST
56
+ (APIServer pid=1147686) INFO 08-10 21:22:30 [launcher.py:46] Route: /reset_encoder_cache, Methods: POST
57
+ (APIServer pid=1147686) INFO 08-10 21:22:30 [launcher.py:46] Route: /tokenize, Methods: POST
58
+ (APIServer pid=1147686) INFO 08-10 21:22:30 [launcher.py:46] Route: /detokenize, Methods: POST
59
+ (APIServer pid=1147686) INFO 08-10 21:22:30 [launcher.py:46] Route: /load, Methods: GET
60
+ (APIServer pid=1147686) INFO 08-10 21:22:30 [launcher.py:46] Route: /version, Methods: GET
61
+ (APIServer pid=1147686) INFO 08-10 21:22:30 [launcher.py:46] Route: /health, Methods: GET
62
+ (APIServer pid=1147686) INFO 08-10 21:22:30 [launcher.py:46] Route: /metrics, Methods: GET
63
+ (APIServer pid=1147686) INFO 08-10 21:22:30 [launcher.py:46] Route: /server_info, Methods: GET
64
+ (APIServer pid=1147686) INFO 08-10 21:22:30 [launcher.py:46] Route: /v1/models, Methods: GET
65
+ (APIServer pid=1147686) INFO 08-10 21:22:30 [launcher.py:46] Route: /ping, Methods: GET
66
+ (APIServer pid=1147686) INFO 08-10 21:22:30 [launcher.py:46] Route: /ping, Methods: POST
67
+ (APIServer pid=1147686) INFO 08-10 21:22:30 [launcher.py:46] Route: /invocations, Methods: POST
68
+ (APIServer pid=1147686) INFO 08-10 21:22:30 [launcher.py:46] Route: /v1/chat/completions, Methods: POST
69
+ (APIServer pid=1147686) INFO 08-10 21:22:30 [launcher.py:46] Route: /v1/chat/completions/batch, Methods: POST
70
+ (APIServer pid=1147686) INFO 08-10 21:22:30 [launcher.py:46] Route: /v1/responses, Methods: POST
71
+ (APIServer pid=1147686) INFO 08-10 21:22:30 [launcher.py:46] Route: /v1/responses/{response_id}, Methods: GET
72
+ (APIServer pid=1147686) INFO 08-10 21:22:30 [launcher.py:46] Route: /v1/responses/{response_id}/cancel, Methods: POST
73
+ (APIServer pid=1147686) INFO 08-10 21:22:30 [launcher.py:46] Route: /v1/completions, Methods: POST
74
+ (APIServer pid=1147686) INFO 08-10 21:22:30 [launcher.py:46] Route: /v1/messages, Methods: POST
75
+ (APIServer pid=1147686) INFO 08-10 21:22:30 [launcher.py:46] Route: /v1/messages/count_tokens, Methods: POST
76
+ (APIServer pid=1147686) INFO 08-10 21:22:30 [launcher.py:46] Route: /inference/v1/generate, Methods: POST
77
+ (APIServer pid=1147686) INFO 08-10 21:22:30 [launcher.py:46] Route: /pause, Methods: POST
78
+ (APIServer pid=1147686) INFO 08-10 21:22:30 [launcher.py:46] Route: /resume, Methods: POST
79
+ (APIServer pid=1147686) INFO 08-10 21:22:30 [launcher.py:46] Route: /is_paused, Methods: GET
80
+ (APIServer pid=1147686) INFO 08-10 21:22:30 [launcher.py:46] Route: /init_weight_transfer_engine, Methods: POST
81
+ (APIServer pid=1147686) INFO 08-10 21:22:30 [launcher.py:46] Route: /update_weights, Methods: POST
82
+ (APIServer pid=1147686) INFO 08-10 21:22:30 [launcher.py:46] Route: /get_world_size, Methods: GET
83
+ (APIServer pid=1147686) INFO 08-10 21:22:30 [launcher.py:46] Route: /scale_elastic_ep, Methods: POST
84
+ (APIServer pid=1147686) INFO 08-10 21:22:30 [launcher.py:46] Route: /is_scaling_elastic_ep, Methods: POST
85
+ (APIServer pid=1147686) INFO 08-10 21:22:30 [launcher.py:46] Route: /v1/chat/completions/render, Methods: POST
86
+ (APIServer pid=1147686) INFO 08-10 21:22:30 [launcher.py:46] Route: /v1/completions/render, Methods: POST
87
+ (APIServer pid=1147686) INFO: Started server process [1147686]
88
+ (APIServer pid=1147686) INFO: Waiting for application startup.
89
+ (APIServer pid=1147686) INFO: Application startup complete.
90
+ (APIServer pid=1147686) INFO: 127.0.0.1:37708 - "GET /health HTTP/1.1" 200 OK
91
+ (APIServer pid=1147686) INFO 08-10 21:22:40 [loggers.py:259] Engine 000: Avg prompt throughput: 2662.8 tokens/s, Avg generation throughput: 2173.5 tokens/s, Running: 253 reqs, Waiting: 972 reqs, GPU KV cache usage: 38.1%, Prefix cache hit rate: 93.2%
92
+ (APIServer pid=1147686) INFO 08-10 21:22:50 [loggers.py:259] Engine 000: Avg prompt throughput: 2160.5 tokens/s, Avg generation throughput: 3146.4 tokens/s, Running: 252 reqs, Waiting: 696 reqs, GPU KV cache usage: 39.2%, Prefix cache hit rate: 93.3%
93
+ (APIServer pid=1147686) INFO 08-10 21:23:00 [loggers.py:259] Engine 000: Avg prompt throughput: 2245.8 tokens/s, Avg generation throughput: 3093.6 tokens/s, Running: 255 reqs, Waiting: 409 reqs, GPU KV cache usage: 38.2%, Prefix cache hit rate: 93.4%
94
+ (APIServer pid=1147686) INFO 08-10 21:23:10 [loggers.py:259] Engine 000: Avg prompt throughput: 2095.3 tokens/s, Avg generation throughput: 3147.7 tokens/s, Running: 253 reqs, Waiting: 149 reqs, GPU KV cache usage: 40.2%, Prefix cache hit rate: 93.4%
95
+ (APIServer pid=1147686) INFO 08-10 21:23:20 [loggers.py:259] Engine 000: Avg prompt throughput: 1232.5 tokens/s, Avg generation throughput: 3092.0 tokens/s, Running: 88 reqs, Waiting: 0 reqs, GPU KV cache usage: 20.9%, Prefix cache hit rate: 93.4%
96
+ (APIServer pid=1147686) INFO: 127.0.0.1:37722 - "POST /v1/completions HTTP/1.1" 200 OK
97
+ (EngineCore pid=1148963) INFO 08-10 21:23:27 [core.py:1210] Shutdown initiated (timeout=0)
98
+ (EngineCore pid=1148963) INFO 08-10 21:23:27 [core.py:1233] Shutdown complete
99
+ (APIServer pid=1147686) INFO: Shutting down
100
+ (APIServer pid=1147686) INFO: Waiting for application shutdown.
101
+ (APIServer pid=1147686) INFO: Application shutdown complete.
102
+ (APIServer pid=1147686) INFO: Finished server process [1147686]
evals/grid_math_unhealed/glean_keep75_unhealed.log ADDED
@@ -0,0 +1,8 @@
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "correct": 898,
3
+ "accuracy": 0.6808188021228203,
4
+ "finished": 1317,
5
+ "finish_rate": 0.9984836997725549,
6
+ "mean_completion_tokens": 109.84154662623199
7
+ }
8
+ saved item-level results -> outputs/evals/grid_math_unhealed/glean_keep75_unhealed_chat.json
evals/grid_math_unhealed/glean_keep75_unhealed_chat.json ADDED
The diff for this file is too large to render. See raw diff
 
evals/grid_math_unhealed/glean_keep75_unhealed_chat.json.server.log ADDED
@@ -0,0 +1,103 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ Skipping import of cpp extensions due to incompatible torch version 2.10.0+cu128 for torchao version 0.15.0 Please see https://github.com/pytorch/ao/issues/2919 for more info
2
+ WARNING 08-10 21:20:11 [registry.py:915] Model architecture OlmoeForCausalLM is already registered, and will be overwritten by the new model class glean_vllm.pruned_olmoe:PrunedOlmoeForCausalLM.
3
+ (APIServer pid=1142912) INFO 08-10 21:20:11 [utils.py:299]
4
+ (APIServer pid=1142912) INFO 08-10 21:20:11 [utils.py:299] β–ˆ β–ˆ β–ˆβ–„ β–„β–ˆ
5
+ (APIServer pid=1142912) INFO 08-10 21:20:11 [utils.py:299] β–„β–„ β–„β–ˆ β–ˆ β–ˆ β–ˆ β–€β–„β–€ β–ˆ version 0.19.0
6
+ (APIServer pid=1142912) INFO 08-10 21:20:11 [utils.py:299] β–ˆβ–„β–ˆβ–€ β–ˆ β–ˆ β–ˆ β–ˆ model outputs/release/unhealed/winnow-olmoe-math-keep75
7
+ (APIServer pid=1142912) INFO 08-10 21:20:11 [utils.py:299] β–€β–€ β–€β–€β–€β–€β–€ β–€β–€β–€β–€β–€ β–€ β–€
8
+ (APIServer pid=1142912) INFO 08-10 21:20:11 [utils.py:299]
9
+ (APIServer pid=1142912) INFO 08-10 21:20:11 [utils.py:233] non-default args: {'model_tag': 'outputs/release/unhealed/winnow-olmoe-math-keep75', 'host': '127.0.0.1', 'port': 8610, 'model': 'outputs/release/unhealed/winnow-olmoe-math-keep75', 'max_model_len': 2048, 'enforce_eager': True, 'served_model_name': ['student'], 'gpu_memory_utilization': 0.85}
10
+ (APIServer pid=1142912) INFO 08-10 21:20:21 [model.py:549] Resolved architecture: PrunedOlmoeForCausalLM
11
+ (APIServer pid=1142912) INFO 08-10 21:20:21 [model.py:1678] Using max model len 2048
12
+ (APIServer pid=1142912) INFO 08-10 21:20:22 [vllm.py:790] Asynchronous scheduling is enabled.
13
+ (APIServer pid=1142912) WARNING 08-10 21:20:22 [vllm.py:848] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
14
+ (APIServer pid=1142912) WARNING 08-10 21:20:22 [vllm.py:859] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
15
+ (APIServer pid=1142912) INFO 08-10 21:20:22 [vllm.py:1025] Cudagraph is disabled under eager mode
16
+ (APIServer pid=1142912) INFO 08-10 21:20:22 [compilation.py:290] Enabled custom fusions: norm_quant, act_quant
17
+ Skipping import of cpp extensions due to incompatible torch version 2.10.0+cu128 for torchao version 0.15.0 Please see https://github.com/pytorch/ao/issues/2919 for more info
18
+ (EngineCore pid=1144303) WARNING 08-10 21:20:30 [registry.py:915] Model architecture OlmoeForCausalLM is already registered, and will be overwritten by the new model class glean_vllm.pruned_olmoe:PrunedOlmoeForCausalLM.
19
+ (EngineCore pid=1144303) INFO 08-10 21:20:30 [core.py:105] Initializing a V1 LLM engine (v0.19.0) with config: model='outputs/release/unhealed/winnow-olmoe-math-keep75', speculative_config=None, tokenizer='outputs/release/unhealed/winnow-olmoe-math-keep75', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=False, dtype=torch.bfloat16, max_seq_len=2048, download_dir=None, load_format=auto, tensor_parallel_size=1, pipeline_parallel_size=1, data_parallel_size=1, decode_context_parallel_size=1, dcp_comm_backend=ag_rs, disable_custom_all_reduce=False, quantization=None, enforce_eager=True, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False), seed=0, served_model_name=student, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'mode': <CompilationMode.NONE: 0>, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'splitting_ops': [], 'compile_mm_encoder': False, 'cudagraph_mm_encoder': False, 'encoder_cudagraph_token_budgets': [], 'encoder_cudagraph_max_images_per_batch': 0, 'compile_sizes': [], 'compile_ranges_endpoints': [2048], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'size_asserts': False, 'alignment_asserts': False, 'scalar_asserts': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': <CUDAGraphMode.NONE: 0>, 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': True, 'fuse_act_quant': True, 'fuse_attn_quant': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False}, 'max_cudagraph_capture_size': 0, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': True, 'static_all_moe_layers': []}
20
+ (EngineCore pid=1144303) INFO 08-10 21:20:31 [parallel_state.py:1400] world_size=1 rank=0 local_rank=0 distributed_init_method=tcp://192.168.0.27:53323 backend=nccl
21
+ (EngineCore pid=1144303) INFO 08-10 21:20:31 [parallel_state.py:1716] rank 0 in world size 1 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank 0, EPLB rank N/A
22
+ (EngineCore pid=1144303) INFO 08-10 21:20:31 [gpu_model_runner.py:4735] Starting to load model outputs/release/unhealed/winnow-olmoe-math-keep75...
23
+ (EngineCore pid=1144303) INFO 08-10 21:20:33 [cuda.py:334] Using FLASH_ATTN attention backend out of potential backends: ['FLASH_ATTN', 'FLASHINFER', 'TRITON_ATTN', 'FLEX_ATTENTION'].
24
+ (EngineCore pid=1144303) INFO 08-10 21:20:33 [flash_attn.py:596] Using FlashAttention version 2
25
+ (EngineCore pid=1144303)
26
+ (EngineCore pid=1144303)
27
+ (EngineCore pid=1144303)
28
+ (EngineCore pid=1144303)
29
+ (EngineCore pid=1144303)
30
+ (EngineCore pid=1144303)
31
+ (EngineCore pid=1144303) INFO 08-10 21:20:35 [default_loader.py:384] Loading weights took 1.88 seconds
32
+ (EngineCore pid=1144303) INFO 08-10 21:20:35 [gpu_model_runner.py:4820] Model loading took 9.89 GiB memory and 3.012713 seconds
33
+ (EngineCore pid=1144303) INFO 08-10 21:20:39 [gpu_worker.py:436] Available KV cache memory: 9.86 GiB
34
+ (EngineCore pid=1144303) INFO 08-10 21:20:39 [kv_cache_utils.py:1319] GPU KV cache size: 80,752 tokens
35
+ (EngineCore pid=1144303) INFO 08-10 21:20:39 [kv_cache_utils.py:1324] Maximum concurrency for 2,048 tokens per request: 39.43x
36
+ (EngineCore pid=1144303) INFO 08-10 21:20:39 [core.py:283] init engine (profile, create kv cache, warmup model) took 3.97 seconds
37
+ (EngineCore pid=1144303) INFO 08-10 21:20:40 [vllm.py:790] Asynchronous scheduling is enabled.
38
+ (EngineCore pid=1144303) WARNING 08-10 21:20:40 [vllm.py:848] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
39
+ (EngineCore pid=1144303) WARNING 08-10 21:20:40 [vllm.py:859] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
40
+ (EngineCore pid=1144303) INFO 08-10 21:20:40 [vllm.py:1025] Cudagraph is disabled under eager mode
41
+ (EngineCore pid=1144303) INFO 08-10 21:20:40 [compilation.py:290] Enabled custom fusions: norm_quant, act_quant
42
+ (APIServer pid=1142912) INFO 08-10 21:20:40 [api_server.py:590] Supported tasks: ['generate']
43
+ (APIServer pid=1142912) WARNING 08-10 21:20:40 [__init__.py:14] SECURITY WARNING: Development endpoints are enabled! This should NOT be used in production!
44
+ (APIServer pid=1142912) INFO 08-10 21:20:40 [hf.py:314] Detected the chat template content format to be 'string'. You can set `--chat-template-content-format` to override this.
45
+ (APIServer pid=1142912) INFO 08-10 21:20:40 [api_server.py:594] Starting vLLM server on http://127.0.0.1:8610
46
+ (APIServer pid=1142912) INFO 08-10 21:20:40 [launcher.py:37] Available routes are:
47
+ (APIServer pid=1142912) INFO 08-10 21:20:40 [launcher.py:46] Route: /openapi.json, Methods: GET, HEAD
48
+ (APIServer pid=1142912) INFO 08-10 21:20:40 [launcher.py:46] Route: /docs, Methods: GET, HEAD
49
+ (APIServer pid=1142912) INFO 08-10 21:20:40 [launcher.py:46] Route: /docs/oauth2-redirect, Methods: GET, HEAD
50
+ (APIServer pid=1142912) INFO 08-10 21:20:40 [launcher.py:46] Route: /redoc, Methods: GET, HEAD
51
+ (APIServer pid=1142912) INFO 08-10 21:20:40 [launcher.py:46] Route: /sleep, Methods: POST
52
+ (APIServer pid=1142912) INFO 08-10 21:20:40 [launcher.py:46] Route: /wake_up, Methods: POST
53
+ (APIServer pid=1142912) INFO 08-10 21:20:40 [launcher.py:46] Route: /is_sleeping, Methods: GET
54
+ (APIServer pid=1142912) INFO 08-10 21:20:40 [launcher.py:46] Route: /collective_rpc, Methods: POST
55
+ (APIServer pid=1142912) INFO 08-10 21:20:40 [launcher.py:46] Route: /reset_prefix_cache, Methods: POST
56
+ (APIServer pid=1142912) INFO 08-10 21:20:40 [launcher.py:46] Route: /reset_mm_cache, Methods: POST
57
+ (APIServer pid=1142912) INFO 08-10 21:20:40 [launcher.py:46] Route: /reset_encoder_cache, Methods: POST
58
+ (APIServer pid=1142912) INFO 08-10 21:20:40 [launcher.py:46] Route: /tokenize, Methods: POST
59
+ (APIServer pid=1142912) INFO 08-10 21:20:40 [launcher.py:46] Route: /detokenize, Methods: POST
60
+ (APIServer pid=1142912) INFO 08-10 21:20:40 [launcher.py:46] Route: /load, Methods: GET
61
+ (APIServer pid=1142912) INFO 08-10 21:20:40 [launcher.py:46] Route: /version, Methods: GET
62
+ (APIServer pid=1142912) INFO 08-10 21:20:40 [launcher.py:46] Route: /health, Methods: GET
63
+ (APIServer pid=1142912) INFO 08-10 21:20:40 [launcher.py:46] Route: /metrics, Methods: GET
64
+ (APIServer pid=1142912) INFO 08-10 21:20:40 [launcher.py:46] Route: /server_info, Methods: GET
65
+ (APIServer pid=1142912) INFO 08-10 21:20:40 [launcher.py:46] Route: /v1/models, Methods: GET
66
+ (APIServer pid=1142912) INFO 08-10 21:20:40 [launcher.py:46] Route: /ping, Methods: GET
67
+ (APIServer pid=1142912) INFO 08-10 21:20:40 [launcher.py:46] Route: /ping, Methods: POST
68
+ (APIServer pid=1142912) INFO 08-10 21:20:40 [launcher.py:46] Route: /invocations, Methods: POST
69
+ (APIServer pid=1142912) INFO 08-10 21:20:40 [launcher.py:46] Route: /v1/chat/completions, Methods: POST
70
+ (APIServer pid=1142912) INFO 08-10 21:20:40 [launcher.py:46] Route: /v1/chat/completions/batch, Methods: POST
71
+ (APIServer pid=1142912) INFO 08-10 21:20:40 [launcher.py:46] Route: /v1/responses, Methods: POST
72
+ (APIServer pid=1142912) INFO 08-10 21:20:40 [launcher.py:46] Route: /v1/responses/{response_id}, Methods: GET
73
+ (APIServer pid=1142912) INFO 08-10 21:20:40 [launcher.py:46] Route: /v1/responses/{response_id}/cancel, Methods: POST
74
+ (APIServer pid=1142912) INFO 08-10 21:20:40 [launcher.py:46] Route: /v1/completions, Methods: POST
75
+ (APIServer pid=1142912) INFO 08-10 21:20:40 [launcher.py:46] Route: /v1/messages, Methods: POST
76
+ (APIServer pid=1142912) INFO 08-10 21:20:40 [launcher.py:46] Route: /v1/messages/count_tokens, Methods: POST
77
+ (APIServer pid=1142912) INFO 08-10 21:20:40 [launcher.py:46] Route: /inference/v1/generate, Methods: POST
78
+ (APIServer pid=1142912) INFO 08-10 21:20:40 [launcher.py:46] Route: /pause, Methods: POST
79
+ (APIServer pid=1142912) INFO 08-10 21:20:40 [launcher.py:46] Route: /resume, Methods: POST
80
+ (APIServer pid=1142912) INFO 08-10 21:20:40 [launcher.py:46] Route: /is_paused, Methods: GET
81
+ (APIServer pid=1142912) INFO 08-10 21:20:40 [launcher.py:46] Route: /init_weight_transfer_engine, Methods: POST
82
+ (APIServer pid=1142912) INFO 08-10 21:20:40 [launcher.py:46] Route: /update_weights, Methods: POST
83
+ (APIServer pid=1142912) INFO 08-10 21:20:40 [launcher.py:46] Route: /get_world_size, Methods: GET
84
+ (APIServer pid=1142912) INFO 08-10 21:20:40 [launcher.py:46] Route: /scale_elastic_ep, Methods: POST
85
+ (APIServer pid=1142912) INFO 08-10 21:20:40 [launcher.py:46] Route: /is_scaling_elastic_ep, Methods: POST
86
+ (APIServer pid=1142912) INFO 08-10 21:20:40 [launcher.py:46] Route: /v1/chat/completions/render, Methods: POST
87
+ (APIServer pid=1142912) INFO 08-10 21:20:40 [launcher.py:46] Route: /v1/completions/render, Methods: POST
88
+ (APIServer pid=1142912) INFO: Started server process [1142912]
89
+ (APIServer pid=1142912) INFO: Waiting for application startup.
90
+ (APIServer pid=1142912) INFO: Application startup complete.
91
+ (APIServer pid=1142912) INFO: 127.0.0.1:58890 - "GET /health HTTP/1.1" 200 OK
92
+ (APIServer pid=1142912) INFO 08-10 21:20:50 [loggers.py:259] Engine 000: Avg prompt throughput: 2625.2 tokens/s, Avg generation throughput: 1944.8 tokens/s, Running: 254 reqs, Waiting: 978 reqs, GPU KV cache usage: 48.0%, Prefix cache hit rate: 93.2%
93
+ (APIServer pid=1142912) INFO 08-10 21:21:00 [loggers.py:259] Engine 000: Avg prompt throughput: 2095.0 tokens/s, Avg generation throughput: 2917.1 tokens/s, Running: 255 reqs, Waiting: 710 reqs, GPU KV cache usage: 49.7%, Prefix cache hit rate: 93.3%
94
+ (APIServer pid=1142912) INFO 08-10 21:21:10 [loggers.py:259] Engine 000: Avg prompt throughput: 2228.9 tokens/s, Avg generation throughput: 2914.7 tokens/s, Running: 256 reqs, Waiting: 428 reqs, GPU KV cache usage: 49.1%, Prefix cache hit rate: 93.4%
95
+ (APIServer pid=1142912) INFO 08-10 21:21:20 [loggers.py:259] Engine 000: Avg prompt throughput: 2113.2 tokens/s, Avg generation throughput: 2942.7 tokens/s, Running: 255 reqs, Waiting: 164 reqs, GPU KV cache usage: 50.1%, Prefix cache hit rate: 93.4%
96
+ (APIServer pid=1142912) INFO 08-10 21:21:30 [loggers.py:259] Engine 000: Avg prompt throughput: 1354.3 tokens/s, Avg generation throughput: 2657.7 tokens/s, Running: 171 reqs, Waiting: 0 reqs, GPU KV cache usage: 40.4%, Prefix cache hit rate: 93.4%
97
+ (APIServer pid=1142912) INFO: 127.0.0.1:58892 - "POST /v1/completions HTTP/1.1" 200 OK
98
+ (EngineCore pid=1144303) INFO 08-10 21:21:39 [core.py:1210] Shutdown initiated (timeout=0)
99
+ (EngineCore pid=1144303) INFO 08-10 21:21:39 [core.py:1233] Shutdown complete
100
+ (APIServer pid=1142912) INFO: Shutting down
101
+ (APIServer pid=1142912) INFO: Waiting for application shutdown.
102
+ (APIServer pid=1142912) INFO: Application shutdown complete.
103
+ (APIServer pid=1142912) INFO: Finished server process [1142912]
evals/grid_math_unhealed/reap_keep25_unhealed.log ADDED
@@ -0,0 +1,8 @@
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "correct": 13,
3
+ "accuracy": 0.009855951478392721,
4
+ "finished": 254,
5
+ "finish_rate": 0.19257012888551933,
6
+ "mean_completion_tokens": 455.9560272934041
7
+ }
8
+ saved item-level results -> outputs/evals/grid_math_unhealed/reap_keep25_unhealed_chat.json
evals/grid_math_unhealed/reap_keep25_unhealed_chat.json ADDED
The diff for this file is too large to render. See raw diff
 
evals/grid_math_unhealed/reap_keep25_unhealed_chat.json.server.log ADDED
@@ -0,0 +1,117 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ Skipping import of cpp extensions due to incompatible torch version 2.10.0+cu128 for torchao version 0.15.0 Please see https://github.com/pytorch/ao/issues/2919 for more info
2
+ WARNING 08-10 21:24:57 [registry.py:915] Model architecture OlmoeForCausalLM is already registered, and will be overwritten by the new model class glean_vllm.pruned_olmoe:PrunedOlmoeForCausalLM.
3
+ (APIServer pid=1153469) INFO 08-10 21:24:57 [utils.py:299]
4
+ (APIServer pid=1153469) INFO 08-10 21:24:57 [utils.py:299] β–ˆ β–ˆ β–ˆβ–„ β–„β–ˆ
5
+ (APIServer pid=1153469) INFO 08-10 21:24:57 [utils.py:299] β–„β–„ β–„β–ˆ β–ˆ β–ˆ β–ˆ β–€β–„β–€ β–ˆ version 0.19.0
6
+ (APIServer pid=1153469) INFO 08-10 21:24:57 [utils.py:299] β–ˆβ–„β–ˆβ–€ β–ˆ β–ˆ β–ˆ β–ˆ model /media/henry/MoreFiles/winnow_release/math_arms/reap-math-keep25
7
+ (APIServer pid=1153469) INFO 08-10 21:24:57 [utils.py:299] β–€β–€ β–€β–€β–€β–€β–€ β–€β–€β–€β–€β–€ β–€ β–€
8
+ (APIServer pid=1153469) INFO 08-10 21:24:57 [utils.py:299]
9
+ (APIServer pid=1153469) INFO 08-10 21:24:57 [utils.py:233] non-default args: {'model_tag': '/media/henry/MoreFiles/winnow_release/math_arms/reap-math-keep25', 'host': '127.0.0.1', 'port': 8611, 'model': '/media/henry/MoreFiles/winnow_release/math_arms/reap-math-keep25', 'max_model_len': 2048, 'enforce_eager': True, 'served_model_name': ['student'], 'gpu_memory_utilization': 0.85}
10
+ (APIServer pid=1153469) INFO 08-10 21:25:06 [model.py:549] Resolved architecture: OlmoeForCausalLM
11
+ (APIServer pid=1153469) INFO 08-10 21:25:06 [model.py:1678] Using max model len 2048
12
+ (APIServer pid=1153469) INFO 08-10 21:25:06 [vllm.py:790] Asynchronous scheduling is enabled.
13
+ (APIServer pid=1153469) WARNING 08-10 21:25:06 [vllm.py:848] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
14
+ (APIServer pid=1153469) WARNING 08-10 21:25:06 [vllm.py:859] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
15
+ (APIServer pid=1153469) INFO 08-10 21:25:06 [vllm.py:1025] Cudagraph is disabled under eager mode
16
+ (APIServer pid=1153469) INFO 08-10 21:25:06 [compilation.py:290] Enabled custom fusions: norm_quant, act_quant
17
+ Skipping import of cpp extensions due to incompatible torch version 2.10.0+cu128 for torchao version 0.15.0 Please see https://github.com/pytorch/ao/issues/2919 for more info
18
+ (EngineCore pid=1154752) WARNING 08-10 21:25:14 [registry.py:915] Model architecture OlmoeForCausalLM is already registered, and will be overwritten by the new model class glean_vllm.pruned_olmoe:PrunedOlmoeForCausalLM.
19
+ (EngineCore pid=1154752) INFO 08-10 21:25:14 [core.py:105] Initializing a V1 LLM engine (v0.19.0) with config: model='/media/henry/MoreFiles/winnow_release/math_arms/reap-math-keep25', speculative_config=None, tokenizer='/media/henry/MoreFiles/winnow_release/math_arms/reap-math-keep25', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=False, dtype=torch.bfloat16, max_seq_len=2048, download_dir=None, load_format=auto, tensor_parallel_size=1, pipeline_parallel_size=1, data_parallel_size=1, decode_context_parallel_size=1, dcp_comm_backend=ag_rs, disable_custom_all_reduce=False, quantization=None, enforce_eager=True, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False), seed=0, served_model_name=student, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'mode': <CompilationMode.NONE: 0>, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'splitting_ops': [], 'compile_mm_encoder': False, 'cudagraph_mm_encoder': False, 'encoder_cudagraph_token_budgets': [], 'encoder_cudagraph_max_images_per_batch': 0, 'compile_sizes': [], 'compile_ranges_endpoints': [2048], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'size_asserts': False, 'alignment_asserts': False, 'scalar_asserts': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': <CUDAGraphMode.NONE: 0>, 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': True, 'fuse_act_quant': True, 'fuse_attn_quant': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False}, 'max_cudagraph_capture_size': 0, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': True, 'static_all_moe_layers': []}
20
+ (EngineCore pid=1154752) INFO 08-10 21:25:14 [parallel_state.py:1400] world_size=1 rank=0 local_rank=0 distributed_init_method=tcp://192.168.0.27:39421 backend=nccl
21
+ (EngineCore pid=1154752) INFO 08-10 21:25:14 [parallel_state.py:1716] rank 0 in world size 1 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank 0, EPLB rank N/A
22
+ (EngineCore pid=1154752) INFO 08-10 21:25:15 [gpu_model_runner.py:4735] Starting to load model /media/henry/MoreFiles/winnow_release/math_arms/reap-math-keep25...
23
+ (EngineCore pid=1154752) INFO 08-10 21:25:16 [cuda.py:334] Using FLASH_ATTN attention backend out of potential backends: ['FLASH_ATTN', 'FLASHINFER', 'TRITON_ATTN', 'FLEX_ATTENTION'].
24
+ (EngineCore pid=1154752) INFO 08-10 21:25:16 [flash_attn.py:596] Using FlashAttention version 2
25
+ (EngineCore pid=1154752) INFO 08-10 21:25:16 [unquantized.py:186] Using TRITON backend for Unquantized MoE
26
+ (EngineCore pid=1154752)
27
+ (EngineCore pid=1154752)
28
+ (EngineCore pid=1154752)
29
+ (EngineCore pid=1154752)
30
+ (EngineCore pid=1154752) INFO 08-10 21:25:19 [default_loader.py:384] Loading weights took 3.52 seconds
31
+ (EngineCore pid=1154752) INFO 08-10 21:25:20 [gpu_model_runner.py:4820] Model loading took 3.89 GiB memory and 4.085796 seconds
32
+ (EngineCore pid=1154752) WARNING 08-10 21:25:21 [fused_moe.py:1090] Using default MoE config. Performance might be sub-optimal! Config file not found at /home/henry/.local/lib/python3.10/site-packages/vllm/model_executor/layers/fused_moe/configs/E=16,N=1024,device_name=NVIDIA_GeForce_RTX_3090.json
33
+ (EngineCore pid=1154752) INFO 08-10 21:25:23 [gpu_worker.py:436] Available KV cache memory: 15.73 GiB
34
+ (EngineCore pid=1154752) INFO 08-10 21:25:23 [kv_cache_utils.py:1319] GPU KV cache size: 128,880 tokens
35
+ (EngineCore pid=1154752) INFO 08-10 21:25:23 [kv_cache_utils.py:1324] Maximum concurrency for 2,048 tokens per request: 62.93x
36
+ (EngineCore pid=1154752) INFO 08-10 21:25:23 [core.py:283] init engine (profile, create kv cache, warmup model) took 2.70 seconds
37
+ (EngineCore pid=1154752) INFO 08-10 21:25:23 [vllm.py:790] Asynchronous scheduling is enabled.
38
+ (EngineCore pid=1154752) WARNING 08-10 21:25:23 [vllm.py:848] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
39
+ (EngineCore pid=1154752) WARNING 08-10 21:25:23 [vllm.py:859] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
40
+ (EngineCore pid=1154752) INFO 08-10 21:25:23 [vllm.py:1025] Cudagraph is disabled under eager mode
41
+ (EngineCore pid=1154752) INFO 08-10 21:25:23 [compilation.py:290] Enabled custom fusions: norm_quant, act_quant
42
+ (APIServer pid=1153469) INFO 08-10 21:25:23 [api_server.py:590] Supported tasks: ['generate']
43
+ (APIServer pid=1153469) WARNING 08-10 21:25:23 [__init__.py:14] SECURITY WARNING: Development endpoints are enabled! This should NOT be used in production!
44
+ (APIServer pid=1153469) INFO 08-10 21:25:23 [hf.py:314] Detected the chat template content format to be 'string'. You can set `--chat-template-content-format` to override this.
45
+ (APIServer pid=1153469) INFO 08-10 21:25:24 [api_server.py:594] Starting vLLM server on http://127.0.0.1:8611
46
+ (APIServer pid=1153469) INFO 08-10 21:25:24 [launcher.py:37] Available routes are:
47
+ (APIServer pid=1153469) INFO 08-10 21:25:24 [launcher.py:46] Route: /openapi.json, Methods: HEAD, GET
48
+ (APIServer pid=1153469) INFO 08-10 21:25:24 [launcher.py:46] Route: /docs, Methods: HEAD, GET
49
+ (APIServer pid=1153469) INFO 08-10 21:25:24 [launcher.py:46] Route: /docs/oauth2-redirect, Methods: HEAD, GET
50
+ (APIServer pid=1153469) INFO 08-10 21:25:24 [launcher.py:46] Route: /redoc, Methods: HEAD, GET
51
+ (APIServer pid=1153469) INFO 08-10 21:25:24 [launcher.py:46] Route: /sleep, Methods: POST
52
+ (APIServer pid=1153469) INFO 08-10 21:25:24 [launcher.py:46] Route: /wake_up, Methods: POST
53
+ (APIServer pid=1153469) INFO 08-10 21:25:24 [launcher.py:46] Route: /is_sleeping, Methods: GET
54
+ (APIServer pid=1153469) INFO 08-10 21:25:24 [launcher.py:46] Route: /collective_rpc, Methods: POST
55
+ (APIServer pid=1153469) INFO 08-10 21:25:24 [launcher.py:46] Route: /reset_prefix_cache, Methods: POST
56
+ (APIServer pid=1153469) INFO 08-10 21:25:24 [launcher.py:46] Route: /reset_mm_cache, Methods: POST
57
+ (APIServer pid=1153469) INFO 08-10 21:25:24 [launcher.py:46] Route: /reset_encoder_cache, Methods: POST
58
+ (APIServer pid=1153469) INFO 08-10 21:25:24 [launcher.py:46] Route: /tokenize, Methods: POST
59
+ (APIServer pid=1153469) INFO 08-10 21:25:24 [launcher.py:46] Route: /detokenize, Methods: POST
60
+ (APIServer pid=1153469) INFO 08-10 21:25:24 [launcher.py:46] Route: /load, Methods: GET
61
+ (APIServer pid=1153469) INFO 08-10 21:25:24 [launcher.py:46] Route: /version, Methods: GET
62
+ (APIServer pid=1153469) INFO 08-10 21:25:24 [launcher.py:46] Route: /health, Methods: GET
63
+ (APIServer pid=1153469) INFO 08-10 21:25:24 [launcher.py:46] Route: /metrics, Methods: GET
64
+ (APIServer pid=1153469) INFO 08-10 21:25:24 [launcher.py:46] Route: /server_info, Methods: GET
65
+ (APIServer pid=1153469) INFO 08-10 21:25:24 [launcher.py:46] Route: /v1/models, Methods: GET
66
+ (APIServer pid=1153469) INFO 08-10 21:25:24 [launcher.py:46] Route: /ping, Methods: GET
67
+ (APIServer pid=1153469) INFO 08-10 21:25:24 [launcher.py:46] Route: /ping, Methods: POST
68
+ (APIServer pid=1153469) INFO 08-10 21:25:24 [launcher.py:46] Route: /invocations, Methods: POST
69
+ (APIServer pid=1153469) INFO 08-10 21:25:24 [launcher.py:46] Route: /v1/chat/completions, Methods: POST
70
+ (APIServer pid=1153469) INFO 08-10 21:25:24 [launcher.py:46] Route: /v1/chat/completions/batch, Methods: POST
71
+ (APIServer pid=1153469) INFO 08-10 21:25:24 [launcher.py:46] Route: /v1/responses, Methods: POST
72
+ (APIServer pid=1153469) INFO 08-10 21:25:24 [launcher.py:46] Route: /v1/responses/{response_id}, Methods: GET
73
+ (APIServer pid=1153469) INFO 08-10 21:25:24 [launcher.py:46] Route: /v1/responses/{response_id}/cancel, Methods: POST
74
+ (APIServer pid=1153469) INFO 08-10 21:25:24 [launcher.py:46] Route: /v1/completions, Methods: POST
75
+ (APIServer pid=1153469) INFO 08-10 21:25:24 [launcher.py:46] Route: /v1/messages, Methods: POST
76
+ (APIServer pid=1153469) INFO 08-10 21:25:24 [launcher.py:46] Route: /v1/messages/count_tokens, Methods: POST
77
+ (APIServer pid=1153469) INFO 08-10 21:25:24 [launcher.py:46] Route: /inference/v1/generate, Methods: POST
78
+ (APIServer pid=1153469) INFO 08-10 21:25:24 [launcher.py:46] Route: /pause, Methods: POST
79
+ (APIServer pid=1153469) INFO 08-10 21:25:24 [launcher.py:46] Route: /resume, Methods: POST
80
+ (APIServer pid=1153469) INFO 08-10 21:25:24 [launcher.py:46] Route: /is_paused, Methods: GET
81
+ (APIServer pid=1153469) INFO 08-10 21:25:24 [launcher.py:46] Route: /init_weight_transfer_engine, Methods: POST
82
+ (APIServer pid=1153469) INFO 08-10 21:25:24 [launcher.py:46] Route: /update_weights, Methods: POST
83
+ (APIServer pid=1153469) INFO 08-10 21:25:24 [launcher.py:46] Route: /get_world_size, Methods: GET
84
+ (APIServer pid=1153469) INFO 08-10 21:25:24 [launcher.py:46] Route: /scale_elastic_ep, Methods: POST
85
+ (APIServer pid=1153469) INFO 08-10 21:25:24 [launcher.py:46] Route: /is_scaling_elastic_ep, Methods: POST
86
+ (APIServer pid=1153469) INFO 08-10 21:25:24 [launcher.py:46] Route: /v1/chat/completions/render, Methods: POST
87
+ (APIServer pid=1153469) INFO 08-10 21:25:24 [launcher.py:46] Route: /v1/completions/render, Methods: POST
88
+ (APIServer pid=1153469) INFO: Started server process [1153469]
89
+ (APIServer pid=1153469) INFO: Waiting for application startup.
90
+ (APIServer pid=1153469) INFO: Application startup complete.
91
+ (APIServer pid=1153469) INFO: 127.0.0.1:56296 - "GET /health HTTP/1.1" 200 OK
92
+ (APIServer pid=1153469) INFO 08-10 21:25:34 [loggers.py:259] Engine 000: Avg prompt throughput: 2117.1 tokens/s, Avg generation throughput: 2702.9 tokens/s, Running: 256 reqs, Waiting: 1049 reqs, GPU KV cache usage: 39.5%, Prefix cache hit rate: 93.1%
93
+ (APIServer pid=1153469) INFO 08-10 21:25:44 [loggers.py:259] Engine 000: Avg prompt throughput: 152.4 tokens/s, Avg generation throughput: 3427.3 tokens/s, Running: 256 reqs, Waiting: 1028 reqs, GPU KV cache usage: 62.9%, Prefix cache hit rate: 93.1%
94
+ (APIServer pid=1153469) INFO 08-10 21:25:54 [loggers.py:259] Engine 000: Avg prompt throughput: 85.5 tokens/s, Avg generation throughput: 3199.1 tokens/s, Running: 256 reqs, Waiting: 1019 reqs, GPU KV cache usage: 86.1%, Prefix cache hit rate: 93.1%
95
+ (APIServer pid=1153469) INFO 08-10 21:26:04 [loggers.py:259] Engine 000: Avg prompt throughput: 45.3 tokens/s, Avg generation throughput: 2986.6 tokens/s, Running: 221 reqs, Waiting: 1046 reqs, GPU KV cache usage: 99.3%, Prefix cache hit rate: 93.1%
96
+ (APIServer pid=1153469) INFO 08-10 21:26:14 [loggers.py:259] Engine 000: Avg prompt throughput: 1864.1 tokens/s, Avg generation throughput: 2962.2 tokens/s, Running: 256 reqs, Waiting: 774 reqs, GPU KV cache usage: 41.8%, Prefix cache hit rate: 93.3%
97
+ (APIServer pid=1153469) INFO 08-10 21:26:24 [loggers.py:259] Engine 000: Avg prompt throughput: 205.4 tokens/s, Avg generation throughput: 3375.5 tokens/s, Running: 256 reqs, Waiting: 747 reqs, GPU KV cache usage: 61.7%, Prefix cache hit rate: 93.3%
98
+ (APIServer pid=1153469) INFO 08-10 21:26:34 [loggers.py:259] Engine 000: Avg prompt throughput: 276.4 tokens/s, Avg generation throughput: 3170.8 tokens/s, Running: 256 reqs, Waiting: 713 reqs, GPU KV cache usage: 77.1%, Prefix cache hit rate: 93.3%
99
+ (APIServer pid=1153469) INFO 08-10 21:26:44 [loggers.py:259] Engine 000: Avg prompt throughput: 122.8 tokens/s, Avg generation throughput: 3044.6 tokens/s, Running: 256 reqs, Waiting: 697 reqs, GPU KV cache usage: 95.3%, Prefix cache hit rate: 93.3%
100
+ (APIServer pid=1153469) INFO 08-10 21:26:54 [loggers.py:259] Engine 000: Avg prompt throughput: 1477.9 tokens/s, Avg generation throughput: 3049.6 tokens/s, Running: 256 reqs, Waiting: 508 reqs, GPU KV cache usage: 47.0%, Prefix cache hit rate: 93.3%
101
+ (APIServer pid=1153469) INFO 08-10 21:27:04 [loggers.py:259] Engine 000: Avg prompt throughput: 372.5 tokens/s, Avg generation throughput: 3322.5 tokens/s, Running: 256 reqs, Waiting: 463 reqs, GPU KV cache usage: 60.0%, Prefix cache hit rate: 93.3%
102
+ (APIServer pid=1153469) INFO 08-10 21:27:14 [loggers.py:259] Engine 000: Avg prompt throughput: 270.9 tokens/s, Avg generation throughput: 3195.4 tokens/s, Running: 256 reqs, Waiting: 427 reqs, GPU KV cache usage: 72.5%, Prefix cache hit rate: 93.4%
103
+ (APIServer pid=1153469) INFO 08-10 21:27:24 [loggers.py:259] Engine 000: Avg prompt throughput: 196.2 tokens/s, Avg generation throughput: 3069.2 tokens/s, Running: 256 reqs, Waiting: 400 reqs, GPU KV cache usage: 87.6%, Prefix cache hit rate: 93.4%
104
+ (APIServer pid=1153469) INFO 08-10 21:27:34 [loggers.py:259] Engine 000: Avg prompt throughput: 1343.3 tokens/s, Avg generation throughput: 3004.6 tokens/s, Running: 256 reqs, Waiting: 240 reqs, GPU KV cache usage: 50.5%, Prefix cache hit rate: 93.4%
105
+ (APIServer pid=1153469) INFO 08-10 21:27:44 [loggers.py:259] Engine 000: Avg prompt throughput: 343.4 tokens/s, Avg generation throughput: 3272.2 tokens/s, Running: 256 reqs, Waiting: 195 reqs, GPU KV cache usage: 60.5%, Prefix cache hit rate: 93.4%
106
+ (APIServer pid=1153469) INFO 08-10 21:27:54 [loggers.py:259] Engine 000: Avg prompt throughput: 372.2 tokens/s, Avg generation throughput: 3193.9 tokens/s, Running: 256 reqs, Waiting: 146 reqs, GPU KV cache usage: 69.4%, Prefix cache hit rate: 93.4%
107
+ (APIServer pid=1153469) INFO 08-10 21:28:04 [loggers.py:259] Engine 000: Avg prompt throughput: 230.9 tokens/s, Avg generation throughput: 3094.1 tokens/s, Running: 256 reqs, Waiting: 119 reqs, GPU KV cache usage: 83.9%, Prefix cache hit rate: 93.4%
108
+ (APIServer pid=1153469) INFO 08-10 21:28:14 [loggers.py:259] Engine 000: Avg prompt throughput: 966.9 tokens/s, Avg generation throughput: 2975.0 tokens/s, Running: 228 reqs, Waiting: 0 reqs, GPU KV cache usage: 48.7%, Prefix cache hit rate: 93.4%
109
+ (APIServer pid=1153469) INFO 08-10 21:28:24 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 3140.1 tokens/s, Running: 172 reqs, Waiting: 0 reqs, GPU KV cache usage: 49.8%, Prefix cache hit rate: 93.4%
110
+ (APIServer pid=1153469) INFO 08-10 21:28:34 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 2828.4 tokens/s, Running: 99 reqs, Waiting: 0 reqs, GPU KV cache usage: 39.5%, Prefix cache hit rate: 93.4%
111
+ (APIServer pid=1153469) INFO: 127.0.0.1:56304 - "POST /v1/completions HTTP/1.1" 200 OK
112
+ (EngineCore pid=1154752) INFO 08-10 21:28:39 [core.py:1210] Shutdown initiated (timeout=0)
113
+ (EngineCore pid=1154752) INFO 08-10 21:28:39 [core.py:1233] Shutdown complete
114
+ (APIServer pid=1153469) INFO: Shutting down
115
+ (APIServer pid=1153469) INFO: Waiting for application shutdown.
116
+ (APIServer pid=1153469) INFO: Application shutdown complete.
117
+ (APIServer pid=1153469) INFO: Finished server process [1153469]
evals/grid_math_unhealed/reap_keep50_unhealed.log ADDED
@@ -0,0 +1,8 @@
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "correct": 195,
3
+ "accuracy": 0.14783927217589082,
4
+ "finished": 1054,
5
+ "finish_rate": 0.799090219863533,
6
+ "mean_completion_tokens": 268.20621683093253
7
+ }
8
+ saved item-level results -> outputs/evals/grid_math_unhealed/reap_keep50_unhealed_chat.json
evals/grid_math_unhealed/reap_keep50_unhealed_chat.json ADDED
The diff for this file is too large to render. See raw diff
 
evals/grid_math_unhealed/reap_keep50_unhealed_chat.json.server.log ADDED
@@ -0,0 +1,110 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ Skipping import of cpp extensions due to incompatible torch version 2.10.0+cu128 for torchao version 0.15.0 Please see https://github.com/pytorch/ao/issues/2919 for more info
2
+ WARNING 08-10 21:22:07 [registry.py:915] Model architecture OlmoeForCausalLM is already registered, and will be overwritten by the new model class glean_vllm.pruned_olmoe:PrunedOlmoeForCausalLM.
3
+ (APIServer pid=1147688) INFO 08-10 21:22:07 [utils.py:299]
4
+ (APIServer pid=1147688) INFO 08-10 21:22:07 [utils.py:299] β–ˆ β–ˆ β–ˆβ–„ β–„β–ˆ
5
+ (APIServer pid=1147688) INFO 08-10 21:22:07 [utils.py:299] β–„β–„ β–„β–ˆ β–ˆ β–ˆ β–ˆ β–€β–„β–€ β–ˆ version 0.19.0
6
+ (APIServer pid=1147688) INFO 08-10 21:22:07 [utils.py:299] β–ˆβ–„β–ˆβ–€ β–ˆ β–ˆ β–ˆ β–ˆ model /media/henry/MoreFiles/winnow_release/math_arms/reap-math-keep50
7
+ (APIServer pid=1147688) INFO 08-10 21:22:07 [utils.py:299] β–€β–€ β–€β–€β–€β–€β–€ β–€β–€β–€β–€β–€ β–€ β–€
8
+ (APIServer pid=1147688) INFO 08-10 21:22:07 [utils.py:299]
9
+ (APIServer pid=1147688) INFO 08-10 21:22:07 [utils.py:233] non-default args: {'model_tag': '/media/henry/MoreFiles/winnow_release/math_arms/reap-math-keep50', 'host': '127.0.0.1', 'port': 8611, 'model': '/media/henry/MoreFiles/winnow_release/math_arms/reap-math-keep50', 'max_model_len': 2048, 'enforce_eager': True, 'served_model_name': ['student'], 'gpu_memory_utilization': 0.85}
10
+ (APIServer pid=1147688) INFO 08-10 21:22:15 [model.py:549] Resolved architecture: OlmoeForCausalLM
11
+ (APIServer pid=1147688) INFO 08-10 21:22:15 [model.py:1678] Using max model len 2048
12
+ (APIServer pid=1147688) INFO 08-10 21:22:15 [vllm.py:790] Asynchronous scheduling is enabled.
13
+ (APIServer pid=1147688) WARNING 08-10 21:22:15 [vllm.py:848] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
14
+ (APIServer pid=1147688) WARNING 08-10 21:22:15 [vllm.py:859] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
15
+ (APIServer pid=1147688) INFO 08-10 21:22:15 [vllm.py:1025] Cudagraph is disabled under eager mode
16
+ (APIServer pid=1147688) INFO 08-10 21:22:15 [compilation.py:290] Enabled custom fusions: norm_quant, act_quant
17
+ Skipping import of cpp extensions due to incompatible torch version 2.10.0+cu128 for torchao version 0.15.0 Please see https://github.com/pytorch/ao/issues/2919 for more info
18
+ (EngineCore pid=1148972) WARNING 08-10 21:22:24 [registry.py:915] Model architecture OlmoeForCausalLM is already registered, and will be overwritten by the new model class glean_vllm.pruned_olmoe:PrunedOlmoeForCausalLM.
19
+ (EngineCore pid=1148972) INFO 08-10 21:22:24 [core.py:105] Initializing a V1 LLM engine (v0.19.0) with config: model='/media/henry/MoreFiles/winnow_release/math_arms/reap-math-keep50', speculative_config=None, tokenizer='/media/henry/MoreFiles/winnow_release/math_arms/reap-math-keep50', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=False, dtype=torch.bfloat16, max_seq_len=2048, download_dir=None, load_format=auto, tensor_parallel_size=1, pipeline_parallel_size=1, data_parallel_size=1, decode_context_parallel_size=1, dcp_comm_backend=ag_rs, disable_custom_all_reduce=False, quantization=None, enforce_eager=True, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False), seed=0, served_model_name=student, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'mode': <CompilationMode.NONE: 0>, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'splitting_ops': [], 'compile_mm_encoder': False, 'cudagraph_mm_encoder': False, 'encoder_cudagraph_token_budgets': [], 'encoder_cudagraph_max_images_per_batch': 0, 'compile_sizes': [], 'compile_ranges_endpoints': [2048], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'size_asserts': False, 'alignment_asserts': False, 'scalar_asserts': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': <CUDAGraphMode.NONE: 0>, 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': True, 'fuse_act_quant': True, 'fuse_attn_quant': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False}, 'max_cudagraph_capture_size': 0, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': True, 'static_all_moe_layers': []}
20
+ (EngineCore pid=1148972) INFO 08-10 21:22:24 [parallel_state.py:1400] world_size=1 rank=0 local_rank=0 distributed_init_method=tcp://192.168.0.27:39471 backend=nccl
21
+ (EngineCore pid=1148972) INFO 08-10 21:22:24 [parallel_state.py:1716] rank 0 in world size 1 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank 0, EPLB rank N/A
22
+ (EngineCore pid=1148972) INFO 08-10 21:22:25 [gpu_model_runner.py:4735] Starting to load model /media/henry/MoreFiles/winnow_release/math_arms/reap-math-keep50...
23
+ (EngineCore pid=1148972) INFO 08-10 21:22:25 [cuda.py:334] Using FLASH_ATTN attention backend out of potential backends: ['FLASH_ATTN', 'FLASHINFER', 'TRITON_ATTN', 'FLEX_ATTENTION'].
24
+ (EngineCore pid=1148972) INFO 08-10 21:22:25 [flash_attn.py:596] Using FlashAttention version 2
25
+ (EngineCore pid=1148972) INFO 08-10 21:22:25 [unquantized.py:186] Using TRITON backend for Unquantized MoE
26
+ (EngineCore pid=1148972)
27
+ (EngineCore pid=1148972)
28
+ (EngineCore pid=1148972)
29
+ (EngineCore pid=1148972)
30
+ (EngineCore pid=1148972)
31
+ (EngineCore pid=1148972) INFO 08-10 21:22:28 [default_loader.py:384] Loading weights took 2.36 seconds
32
+ (EngineCore pid=1148972) INFO 08-10 21:22:28 [gpu_model_runner.py:4820] Model loading took 6.89 GiB memory and 2.924399 seconds
33
+ (EngineCore pid=1148972) WARNING 08-10 21:22:29 [fused_moe.py:1090] Using default MoE config. Performance might be sub-optimal! Config file not found at /home/henry/.local/lib/python3.10/site-packages/vllm/model_executor/layers/fused_moe/configs/E=32,N=1024,device_name=NVIDIA_GeForce_RTX_3090.json
34
+ (EngineCore pid=1148972) INFO 08-10 21:22:30 [gpu_worker.py:436] Available KV cache memory: 12.73 GiB
35
+ (EngineCore pid=1148972) INFO 08-10 21:22:30 [kv_cache_utils.py:1319] GPU KV cache size: 104,304 tokens
36
+ (EngineCore pid=1148972) INFO 08-10 21:22:30 [kv_cache_utils.py:1324] Maximum concurrency for 2,048 tokens per request: 50.93x
37
+ (EngineCore pid=1148972) INFO 08-10 21:22:30 [core.py:283] init engine (profile, create kv cache, warmup model) took 1.68 seconds
38
+ (EngineCore pid=1148972) INFO 08-10 21:22:31 [vllm.py:790] Asynchronous scheduling is enabled.
39
+ (EngineCore pid=1148972) WARNING 08-10 21:22:31 [vllm.py:848] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
40
+ (EngineCore pid=1148972) WARNING 08-10 21:22:31 [vllm.py:859] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
41
+ (EngineCore pid=1148972) INFO 08-10 21:22:31 [vllm.py:1025] Cudagraph is disabled under eager mode
42
+ (EngineCore pid=1148972) INFO 08-10 21:22:31 [compilation.py:290] Enabled custom fusions: norm_quant, act_quant
43
+ (APIServer pid=1147688) INFO 08-10 21:22:31 [api_server.py:590] Supported tasks: ['generate']
44
+ (APIServer pid=1147688) WARNING 08-10 21:22:31 [__init__.py:14] SECURITY WARNING: Development endpoints are enabled! This should NOT be used in production!
45
+ (APIServer pid=1147688) INFO 08-10 21:22:31 [hf.py:314] Detected the chat template content format to be 'string'. You can set `--chat-template-content-format` to override this.
46
+ (APIServer pid=1147688) INFO 08-10 21:22:31 [api_server.py:594] Starting vLLM server on http://127.0.0.1:8611
47
+ (APIServer pid=1147688) INFO 08-10 21:22:31 [launcher.py:37] Available routes are:
48
+ (APIServer pid=1147688) INFO 08-10 21:22:31 [launcher.py:46] Route: /openapi.json, Methods: GET, HEAD
49
+ (APIServer pid=1147688) INFO 08-10 21:22:31 [launcher.py:46] Route: /docs, Methods: GET, HEAD
50
+ (APIServer pid=1147688) INFO 08-10 21:22:31 [launcher.py:46] Route: /docs/oauth2-redirect, Methods: GET, HEAD
51
+ (APIServer pid=1147688) INFO 08-10 21:22:31 [launcher.py:46] Route: /redoc, Methods: GET, HEAD
52
+ (APIServer pid=1147688) INFO 08-10 21:22:31 [launcher.py:46] Route: /sleep, Methods: POST
53
+ (APIServer pid=1147688) INFO 08-10 21:22:31 [launcher.py:46] Route: /wake_up, Methods: POST
54
+ (APIServer pid=1147688) INFO 08-10 21:22:31 [launcher.py:46] Route: /is_sleeping, Methods: GET
55
+ (APIServer pid=1147688) INFO 08-10 21:22:31 [launcher.py:46] Route: /collective_rpc, Methods: POST
56
+ (APIServer pid=1147688) INFO 08-10 21:22:31 [launcher.py:46] Route: /reset_prefix_cache, Methods: POST
57
+ (APIServer pid=1147688) INFO 08-10 21:22:31 [launcher.py:46] Route: /reset_mm_cache, Methods: POST
58
+ (APIServer pid=1147688) INFO 08-10 21:22:31 [launcher.py:46] Route: /reset_encoder_cache, Methods: POST
59
+ (APIServer pid=1147688) INFO 08-10 21:22:31 [launcher.py:46] Route: /tokenize, Methods: POST
60
+ (APIServer pid=1147688) INFO 08-10 21:22:31 [launcher.py:46] Route: /detokenize, Methods: POST
61
+ (APIServer pid=1147688) INFO 08-10 21:22:31 [launcher.py:46] Route: /load, Methods: GET
62
+ (APIServer pid=1147688) INFO 08-10 21:22:31 [launcher.py:46] Route: /version, Methods: GET
63
+ (APIServer pid=1147688) INFO 08-10 21:22:31 [launcher.py:46] Route: /health, Methods: GET
64
+ (APIServer pid=1147688) INFO 08-10 21:22:31 [launcher.py:46] Route: /metrics, Methods: GET
65
+ (APIServer pid=1147688) INFO 08-10 21:22:31 [launcher.py:46] Route: /server_info, Methods: GET
66
+ (APIServer pid=1147688) INFO 08-10 21:22:31 [launcher.py:46] Route: /v1/models, Methods: GET
67
+ (APIServer pid=1147688) INFO 08-10 21:22:31 [launcher.py:46] Route: /ping, Methods: GET
68
+ (APIServer pid=1147688) INFO 08-10 21:22:31 [launcher.py:46] Route: /ping, Methods: POST
69
+ (APIServer pid=1147688) INFO 08-10 21:22:31 [launcher.py:46] Route: /invocations, Methods: POST
70
+ (APIServer pid=1147688) INFO 08-10 21:22:31 [launcher.py:46] Route: /v1/chat/completions, Methods: POST
71
+ (APIServer pid=1147688) INFO 08-10 21:22:31 [launcher.py:46] Route: /v1/chat/completions/batch, Methods: POST
72
+ (APIServer pid=1147688) INFO 08-10 21:22:31 [launcher.py:46] Route: /v1/responses, Methods: POST
73
+ (APIServer pid=1147688) INFO 08-10 21:22:31 [launcher.py:46] Route: /v1/responses/{response_id}, Methods: GET
74
+ (APIServer pid=1147688) INFO 08-10 21:22:31 [launcher.py:46] Route: /v1/responses/{response_id}/cancel, Methods: POST
75
+ (APIServer pid=1147688) INFO 08-10 21:22:31 [launcher.py:46] Route: /v1/completions, Methods: POST
76
+ (APIServer pid=1147688) INFO 08-10 21:22:31 [launcher.py:46] Route: /v1/messages, Methods: POST
77
+ (APIServer pid=1147688) INFO 08-10 21:22:31 [launcher.py:46] Route: /v1/messages/count_tokens, Methods: POST
78
+ (APIServer pid=1147688) INFO 08-10 21:22:31 [launcher.py:46] Route: /inference/v1/generate, Methods: POST
79
+ (APIServer pid=1147688) INFO 08-10 21:22:31 [launcher.py:46] Route: /pause, Methods: POST
80
+ (APIServer pid=1147688) INFO 08-10 21:22:31 [launcher.py:46] Route: /resume, Methods: POST
81
+ (APIServer pid=1147688) INFO 08-10 21:22:31 [launcher.py:46] Route: /is_paused, Methods: GET
82
+ (APIServer pid=1147688) INFO 08-10 21:22:31 [launcher.py:46] Route: /init_weight_transfer_engine, Methods: POST
83
+ (APIServer pid=1147688) INFO 08-10 21:22:31 [launcher.py:46] Route: /update_weights, Methods: POST
84
+ (APIServer pid=1147688) INFO 08-10 21:22:31 [launcher.py:46] Route: /get_world_size, Methods: GET
85
+ (APIServer pid=1147688) INFO 08-10 21:22:31 [launcher.py:46] Route: /scale_elastic_ep, Methods: POST
86
+ (APIServer pid=1147688) INFO 08-10 21:22:31 [launcher.py:46] Route: /is_scaling_elastic_ep, Methods: POST
87
+ (APIServer pid=1147688) INFO 08-10 21:22:31 [launcher.py:46] Route: /v1/chat/completions/render, Methods: POST
88
+ (APIServer pid=1147688) INFO 08-10 21:22:31 [launcher.py:46] Route: /v1/completions/render, Methods: POST
89
+ (APIServer pid=1147688) INFO: Started server process [1147688]
90
+ (APIServer pid=1147688) INFO: Waiting for application startup.
91
+ (APIServer pid=1147688) INFO: Application startup complete.
92
+ (APIServer pid=1147688) INFO: 127.0.0.1:46976 - "GET /health HTTP/1.1" 200 OK
93
+ (APIServer pid=1147688) INFO 08-10 21:22:41 [loggers.py:259] Engine 000: Avg prompt throughput: 2242.6 tokens/s, Avg generation throughput: 2459.7 tokens/s, Running: 255 reqs, Waiting: 1030 reqs, GPU KV cache usage: 44.8%, Prefix cache hit rate: 93.1%
94
+ (APIServer pid=1147688) INFO 08-10 21:22:51 [loggers.py:259] Engine 000: Avg prompt throughput: 904.4 tokens/s, Avg generation throughput: 3290.0 tokens/s, Running: 256 reqs, Waiting: 915 reqs, GPU KV cache usage: 60.6%, Prefix cache hit rate: 93.2%
95
+ (APIServer pid=1147688) INFO 08-10 21:23:01 [loggers.py:259] Engine 000: Avg prompt throughput: 884.6 tokens/s, Avg generation throughput: 3162.2 tokens/s, Running: 255 reqs, Waiting: 804 reqs, GPU KV cache usage: 68.7%, Prefix cache hit rate: 93.3%
96
+ (APIServer pid=1147688) INFO 08-10 21:23:11 [loggers.py:259] Engine 000: Avg prompt throughput: 723.3 tokens/s, Avg generation throughput: 3138.0 tokens/s, Running: 256 reqs, Waiting: 710 reqs, GPU KV cache usage: 78.8%, Prefix cache hit rate: 93.3%
97
+ (APIServer pid=1147688) INFO 08-10 21:23:21 [loggers.py:259] Engine 000: Avg prompt throughput: 1120.1 tokens/s, Avg generation throughput: 3108.8 tokens/s, Running: 256 reqs, Waiting: 566 reqs, GPU KV cache usage: 63.1%, Prefix cache hit rate: 93.3%
98
+ (APIServer pid=1147688) INFO 08-10 21:23:31 [loggers.py:259] Engine 000: Avg prompt throughput: 953.6 tokens/s, Avg generation throughput: 3135.7 tokens/s, Running: 256 reqs, Waiting: 447 reqs, GPU KV cache usage: 65.3%, Prefix cache hit rate: 93.3%
99
+ (APIServer pid=1147688) INFO 08-10 21:23:41 [loggers.py:259] Engine 000: Avg prompt throughput: 996.2 tokens/s, Avg generation throughput: 3135.3 tokens/s, Running: 256 reqs, Waiting: 321 reqs, GPU KV cache usage: 65.1%, Prefix cache hit rate: 93.3%
100
+ (APIServer pid=1147688) INFO 08-10 21:23:51 [loggers.py:259] Engine 000: Avg prompt throughput: 916.9 tokens/s, Avg generation throughput: 3111.0 tokens/s, Running: 256 reqs, Waiting: 210 reqs, GPU KV cache usage: 67.8%, Prefix cache hit rate: 93.4%
101
+ (APIServer pid=1147688) INFO 08-10 21:24:01 [loggers.py:259] Engine 000: Avg prompt throughput: 814.5 tokens/s, Avg generation throughput: 3138.3 tokens/s, Running: 255 reqs, Waiting: 107 reqs, GPU KV cache usage: 70.9%, Prefix cache hit rate: 93.4%
102
+ (APIServer pid=1147688) INFO 08-10 21:24:11 [loggers.py:259] Engine 000: Avg prompt throughput: 880.7 tokens/s, Avg generation throughput: 3076.6 tokens/s, Running: 243 reqs, Waiting: 0 reqs, GPU KV cache usage: 68.8%, Prefix cache hit rate: 93.4%
103
+ (APIServer pid=1147688) INFO 08-10 21:24:21 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 2944.0 tokens/s, Running: 120 reqs, Waiting: 0 reqs, GPU KV cache usage: 48.8%, Prefix cache hit rate: 93.4%
104
+ (APIServer pid=1147688) INFO: 127.0.0.1:46980 - "POST /v1/completions HTTP/1.1" 200 OK
105
+ (EngineCore pid=1148972) INFO 08-10 21:24:30 [core.py:1210] Shutdown initiated (timeout=0)
106
+ (EngineCore pid=1148972) INFO 08-10 21:24:30 [core.py:1233] Shutdown complete
107
+ (APIServer pid=1147688) INFO: Shutting down
108
+ (APIServer pid=1147688) INFO: Waiting for application shutdown.
109
+ (APIServer pid=1147688) INFO: Application shutdown complete.
110
+ (APIServer pid=1147688) INFO: Finished server process [1147688]
evals/grid_math_unhealed/reap_keep75_unhealed.log ADDED
@@ -0,0 +1,8 @@
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "correct": 833,
3
+ "accuracy": 0.6315390447308568,
4
+ "finished": 1316,
5
+ "finish_rate": 0.9977255496588324,
6
+ "mean_completion_tokens": 108.13646702047005
7
+ }
8
+ saved item-level results -> outputs/evals/grid_math_unhealed/reap_keep75_unhealed_chat.json
evals/grid_math_unhealed/reap_keep75_unhealed_chat.json ADDED
The diff for this file is too large to render. See raw diff
 
evals/grid_math_unhealed/reap_keep75_unhealed_chat.json.server.log ADDED
@@ -0,0 +1,105 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ Skipping import of cpp extensions due to incompatible torch version 2.10.0+cu128 for torchao version 0.15.0 Please see https://github.com/pytorch/ao/issues/2919 for more info
2
+ WARNING 08-10 21:20:11 [registry.py:915] Model architecture OlmoeForCausalLM is already registered, and will be overwritten by the new model class glean_vllm.pruned_olmoe:PrunedOlmoeForCausalLM.
3
+ (APIServer pid=1142905) INFO 08-10 21:20:11 [utils.py:299]
4
+ (APIServer pid=1142905) INFO 08-10 21:20:11 [utils.py:299] β–ˆ β–ˆ β–ˆβ–„ β–„β–ˆ
5
+ (APIServer pid=1142905) INFO 08-10 21:20:11 [utils.py:299] β–„β–„ β–„β–ˆ β–ˆ β–ˆ β–ˆ β–€β–„β–€ β–ˆ version 0.19.0
6
+ (APIServer pid=1142905) INFO 08-10 21:20:11 [utils.py:299] β–ˆβ–„β–ˆβ–€ β–ˆ β–ˆ β–ˆ β–ˆ model /media/henry/MoreFiles/winnow_release/math_arms/reap-math-keep75
7
+ (APIServer pid=1142905) INFO 08-10 21:20:11 [utils.py:299] β–€β–€ β–€β–€β–€β–€β–€ β–€β–€β–€β–€β–€ β–€ β–€
8
+ (APIServer pid=1142905) INFO 08-10 21:20:11 [utils.py:299]
9
+ (APIServer pid=1142905) INFO 08-10 21:20:11 [utils.py:233] non-default args: {'model_tag': '/media/henry/MoreFiles/winnow_release/math_arms/reap-math-keep75', 'host': '127.0.0.1', 'port': 8611, 'model': '/media/henry/MoreFiles/winnow_release/math_arms/reap-math-keep75', 'max_model_len': 2048, 'enforce_eager': True, 'served_model_name': ['student'], 'gpu_memory_utilization': 0.85}
10
+ (APIServer pid=1142905) INFO 08-10 21:20:21 [model.py:549] Resolved architecture: OlmoeForCausalLM
11
+ (APIServer pid=1142905) INFO 08-10 21:20:21 [model.py:1678] Using max model len 2048
12
+ (APIServer pid=1142905) INFO 08-10 21:20:22 [vllm.py:790] Asynchronous scheduling is enabled.
13
+ (APIServer pid=1142905) WARNING 08-10 21:20:22 [vllm.py:848] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
14
+ (APIServer pid=1142905) WARNING 08-10 21:20:22 [vllm.py:859] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
15
+ (APIServer pid=1142905) INFO 08-10 21:20:22 [vllm.py:1025] Cudagraph is disabled under eager mode
16
+ (APIServer pid=1142905) INFO 08-10 21:20:22 [compilation.py:290] Enabled custom fusions: norm_quant, act_quant
17
+ Skipping import of cpp extensions due to incompatible torch version 2.10.0+cu128 for torchao version 0.15.0 Please see https://github.com/pytorch/ao/issues/2919 for more info
18
+ (EngineCore pid=1144304) WARNING 08-10 21:20:30 [registry.py:915] Model architecture OlmoeForCausalLM is already registered, and will be overwritten by the new model class glean_vllm.pruned_olmoe:PrunedOlmoeForCausalLM.
19
+ (EngineCore pid=1144304) INFO 08-10 21:20:30 [core.py:105] Initializing a V1 LLM engine (v0.19.0) with config: model='/media/henry/MoreFiles/winnow_release/math_arms/reap-math-keep75', speculative_config=None, tokenizer='/media/henry/MoreFiles/winnow_release/math_arms/reap-math-keep75', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=False, dtype=torch.bfloat16, max_seq_len=2048, download_dir=None, load_format=auto, tensor_parallel_size=1, pipeline_parallel_size=1, data_parallel_size=1, decode_context_parallel_size=1, dcp_comm_backend=ag_rs, disable_custom_all_reduce=False, quantization=None, enforce_eager=True, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False), seed=0, served_model_name=student, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'mode': <CompilationMode.NONE: 0>, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'splitting_ops': [], 'compile_mm_encoder': False, 'cudagraph_mm_encoder': False, 'encoder_cudagraph_token_budgets': [], 'encoder_cudagraph_max_images_per_batch': 0, 'compile_sizes': [], 'compile_ranges_endpoints': [2048], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'size_asserts': False, 'alignment_asserts': False, 'scalar_asserts': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': <CUDAGraphMode.NONE: 0>, 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': True, 'fuse_act_quant': True, 'fuse_attn_quant': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False}, 'max_cudagraph_capture_size': 0, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': True, 'static_all_moe_layers': []}
20
+ (EngineCore pid=1144304) INFO 08-10 21:20:31 [parallel_state.py:1400] world_size=1 rank=0 local_rank=0 distributed_init_method=tcp://192.168.0.27:40093 backend=nccl
21
+ (EngineCore pid=1144304) INFO 08-10 21:20:31 [parallel_state.py:1716] rank 0 in world size 1 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank 0, EPLB rank N/A
22
+ (EngineCore pid=1144304) INFO 08-10 21:20:31 [gpu_model_runner.py:4735] Starting to load model /media/henry/MoreFiles/winnow_release/math_arms/reap-math-keep75...
23
+ (EngineCore pid=1144304) INFO 08-10 21:20:33 [cuda.py:334] Using FLASH_ATTN attention backend out of potential backends: ['FLASH_ATTN', 'FLASHINFER', 'TRITON_ATTN', 'FLEX_ATTENTION'].
24
+ (EngineCore pid=1144304) INFO 08-10 21:20:33 [flash_attn.py:596] Using FlashAttention version 2
25
+ (EngineCore pid=1144304) INFO 08-10 21:20:33 [unquantized.py:186] Using TRITON backend for Unquantized MoE
26
+ (EngineCore pid=1144304)
27
+ (EngineCore pid=1144304)
28
+ (EngineCore pid=1144304)
29
+ (EngineCore pid=1144304)
30
+ (EngineCore pid=1144304)
31
+ (EngineCore pid=1144304)
32
+ (EngineCore pid=1144304) INFO 08-10 21:20:35 [default_loader.py:384] Loading weights took 2.25 seconds
33
+ (EngineCore pid=1144304) INFO 08-10 21:20:36 [gpu_model_runner.py:4820] Model loading took 9.89 GiB memory and 3.345540 seconds
34
+ (EngineCore pid=1144304) WARNING 08-10 21:20:36 [fused_moe.py:1090] Using default MoE config. Performance might be sub-optimal! Config file not found at /home/henry/.local/lib/python3.10/site-packages/vllm/model_executor/layers/fused_moe/configs/E=48,N=1024,device_name=NVIDIA_GeForce_RTX_3090.json
35
+ (EngineCore pid=1144304) INFO 08-10 21:20:37 [gpu_worker.py:436] Available KV cache memory: 9.73 GiB
36
+ (EngineCore pid=1144304) INFO 08-10 21:20:37 [kv_cache_utils.py:1319] GPU KV cache size: 79,712 tokens
37
+ (EngineCore pid=1144304) INFO 08-10 21:20:37 [kv_cache_utils.py:1324] Maximum concurrency for 2,048 tokens per request: 38.92x
38
+ (EngineCore pid=1144304) INFO 08-10 21:20:37 [core.py:283] init engine (profile, create kv cache, warmup model) took 1.75 seconds
39
+ (EngineCore pid=1144304) INFO 08-10 21:20:38 [vllm.py:790] Asynchronous scheduling is enabled.
40
+ (EngineCore pid=1144304) WARNING 08-10 21:20:38 [vllm.py:848] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
41
+ (EngineCore pid=1144304) WARNING 08-10 21:20:38 [vllm.py:859] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
42
+ (EngineCore pid=1144304) INFO 08-10 21:20:38 [vllm.py:1025] Cudagraph is disabled under eager mode
43
+ (EngineCore pid=1144304) INFO 08-10 21:20:38 [compilation.py:290] Enabled custom fusions: norm_quant, act_quant
44
+ (APIServer pid=1142905) INFO 08-10 21:20:38 [api_server.py:590] Supported tasks: ['generate']
45
+ (APIServer pid=1142905) WARNING 08-10 21:20:38 [__init__.py:14] SECURITY WARNING: Development endpoints are enabled! This should NOT be used in production!
46
+ (APIServer pid=1142905) INFO 08-10 21:20:38 [hf.py:314] Detected the chat template content format to be 'string'. You can set `--chat-template-content-format` to override this.
47
+ (APIServer pid=1142905) INFO 08-10 21:20:38 [api_server.py:594] Starting vLLM server on http://127.0.0.1:8611
48
+ (APIServer pid=1142905) INFO 08-10 21:20:38 [launcher.py:37] Available routes are:
49
+ (APIServer pid=1142905) INFO 08-10 21:20:38 [launcher.py:46] Route: /openapi.json, Methods: HEAD, GET
50
+ (APIServer pid=1142905) INFO 08-10 21:20:38 [launcher.py:46] Route: /docs, Methods: HEAD, GET
51
+ (APIServer pid=1142905) INFO 08-10 21:20:38 [launcher.py:46] Route: /docs/oauth2-redirect, Methods: HEAD, GET
52
+ (APIServer pid=1142905) INFO 08-10 21:20:38 [launcher.py:46] Route: /redoc, Methods: HEAD, GET
53
+ (APIServer pid=1142905) INFO 08-10 21:20:38 [launcher.py:46] Route: /sleep, Methods: POST
54
+ (APIServer pid=1142905) INFO 08-10 21:20:38 [launcher.py:46] Route: /wake_up, Methods: POST
55
+ (APIServer pid=1142905) INFO 08-10 21:20:38 [launcher.py:46] Route: /is_sleeping, Methods: GET
56
+ (APIServer pid=1142905) INFO 08-10 21:20:38 [launcher.py:46] Route: /collective_rpc, Methods: POST
57
+ (APIServer pid=1142905) INFO 08-10 21:20:38 [launcher.py:46] Route: /reset_prefix_cache, Methods: POST
58
+ (APIServer pid=1142905) INFO 08-10 21:20:38 [launcher.py:46] Route: /reset_mm_cache, Methods: POST
59
+ (APIServer pid=1142905) INFO 08-10 21:20:38 [launcher.py:46] Route: /reset_encoder_cache, Methods: POST
60
+ (APIServer pid=1142905) INFO 08-10 21:20:38 [launcher.py:46] Route: /tokenize, Methods: POST
61
+ (APIServer pid=1142905) INFO 08-10 21:20:38 [launcher.py:46] Route: /detokenize, Methods: POST
62
+ (APIServer pid=1142905) INFO 08-10 21:20:38 [launcher.py:46] Route: /load, Methods: GET
63
+ (APIServer pid=1142905) INFO 08-10 21:20:38 [launcher.py:46] Route: /version, Methods: GET
64
+ (APIServer pid=1142905) INFO 08-10 21:20:38 [launcher.py:46] Route: /health, Methods: GET
65
+ (APIServer pid=1142905) INFO 08-10 21:20:38 [launcher.py:46] Route: /metrics, Methods: GET
66
+ (APIServer pid=1142905) INFO 08-10 21:20:38 [launcher.py:46] Route: /server_info, Methods: GET
67
+ (APIServer pid=1142905) INFO 08-10 21:20:38 [launcher.py:46] Route: /v1/models, Methods: GET
68
+ (APIServer pid=1142905) INFO 08-10 21:20:38 [launcher.py:46] Route: /ping, Methods: GET
69
+ (APIServer pid=1142905) INFO 08-10 21:20:38 [launcher.py:46] Route: /ping, Methods: POST
70
+ (APIServer pid=1142905) INFO 08-10 21:20:38 [launcher.py:46] Route: /invocations, Methods: POST
71
+ (APIServer pid=1142905) INFO 08-10 21:20:38 [launcher.py:46] Route: /v1/chat/completions, Methods: POST
72
+ (APIServer pid=1142905) INFO 08-10 21:20:38 [launcher.py:46] Route: /v1/chat/completions/batch, Methods: POST
73
+ (APIServer pid=1142905) INFO 08-10 21:20:38 [launcher.py:46] Route: /v1/responses, Methods: POST
74
+ (APIServer pid=1142905) INFO 08-10 21:20:38 [launcher.py:46] Route: /v1/responses/{response_id}, Methods: GET
75
+ (APIServer pid=1142905) INFO 08-10 21:20:38 [launcher.py:46] Route: /v1/responses/{response_id}/cancel, Methods: POST
76
+ (APIServer pid=1142905) INFO 08-10 21:20:38 [launcher.py:46] Route: /v1/completions, Methods: POST
77
+ (APIServer pid=1142905) INFO 08-10 21:20:38 [launcher.py:46] Route: /v1/messages, Methods: POST
78
+ (APIServer pid=1142905) INFO 08-10 21:20:38 [launcher.py:46] Route: /v1/messages/count_tokens, Methods: POST
79
+ (APIServer pid=1142905) INFO 08-10 21:20:38 [launcher.py:46] Route: /inference/v1/generate, Methods: POST
80
+ (APIServer pid=1142905) INFO 08-10 21:20:38 [launcher.py:46] Route: /pause, Methods: POST
81
+ (APIServer pid=1142905) INFO 08-10 21:20:38 [launcher.py:46] Route: /resume, Methods: POST
82
+ (APIServer pid=1142905) INFO 08-10 21:20:38 [launcher.py:46] Route: /is_paused, Methods: GET
83
+ (APIServer pid=1142905) INFO 08-10 21:20:38 [launcher.py:46] Route: /init_weight_transfer_engine, Methods: POST
84
+ (APIServer pid=1142905) INFO 08-10 21:20:38 [launcher.py:46] Route: /update_weights, Methods: POST
85
+ (APIServer pid=1142905) INFO 08-10 21:20:38 [launcher.py:46] Route: /get_world_size, Methods: GET
86
+ (APIServer pid=1142905) INFO 08-10 21:20:38 [launcher.py:46] Route: /scale_elastic_ep, Methods: POST
87
+ (APIServer pid=1142905) INFO 08-10 21:20:38 [launcher.py:46] Route: /is_scaling_elastic_ep, Methods: POST
88
+ (APIServer pid=1142905) INFO 08-10 21:20:38 [launcher.py:46] Route: /v1/chat/completions/render, Methods: POST
89
+ (APIServer pid=1142905) INFO 08-10 21:20:38 [launcher.py:46] Route: /v1/completions/render, Methods: POST
90
+ (APIServer pid=1142905) INFO: Started server process [1142905]
91
+ (APIServer pid=1142905) INFO: Waiting for application startup.
92
+ (APIServer pid=1142905) INFO: Application startup complete.
93
+ (APIServer pid=1142905) INFO: 127.0.0.1:54554 - "GET /health HTTP/1.1" 200 OK
94
+ (APIServer pid=1142905) INFO 08-10 21:20:49 [loggers.py:259] Engine 000: Avg prompt throughput: 2776.8 tokens/s, Avg generation throughput: 1984.7 tokens/s, Running: 252 reqs, Waiting: 955 reqs, GPU KV cache usage: 47.5%, Prefix cache hit rate: 93.2%
95
+ (APIServer pid=1142905) INFO 08-10 21:20:59 [loggers.py:259] Engine 000: Avg prompt throughput: 2107.1 tokens/s, Avg generation throughput: 3019.1 tokens/s, Running: 255 reqs, Waiting: 685 reqs, GPU KV cache usage: 51.1%, Prefix cache hit rate: 93.3%
96
+ (APIServer pid=1142905) INFO 08-10 21:21:09 [loggers.py:259] Engine 000: Avg prompt throughput: 2170.0 tokens/s, Avg generation throughput: 2967.4 tokens/s, Running: 256 reqs, Waiting: 409 reqs, GPU KV cache usage: 51.3%, Prefix cache hit rate: 93.4%
97
+ (APIServer pid=1142905) INFO 08-10 21:21:19 [loggers.py:259] Engine 000: Avg prompt throughput: 2265.7 tokens/s, Avg generation throughput: 2966.5 tokens/s, Running: 255 reqs, Waiting: 128 reqs, GPU KV cache usage: 51.7%, Prefix cache hit rate: 93.4%
98
+ (APIServer pid=1142905) INFO 08-10 21:21:29 [loggers.py:259] Engine 000: Avg prompt throughput: 1072.9 tokens/s, Avg generation throughput: 2630.6 tokens/s, Running: 103 reqs, Waiting: 0 reqs, GPU KV cache usage: 31.6%, Prefix cache hit rate: 93.4%
99
+ (APIServer pid=1142905) INFO: 127.0.0.1:54568 - "POST /v1/completions HTTP/1.1" 200 OK
100
+ (EngineCore pid=1144304) INFO 08-10 21:21:38 [core.py:1210] Shutdown initiated (timeout=0)
101
+ (EngineCore pid=1144304) INFO 08-10 21:21:38 [core.py:1233] Shutdown complete
102
+ (APIServer pid=1142905) INFO: Shutting down
103
+ (APIServer pid=1142905) INFO: Waiting for application shutdown.
104
+ (APIServer pid=1142905) INFO: Application shutdown complete.
105
+ (APIServer pid=1142905) INFO: Finished server process [1142905]
evals/grid_math_unhealed/uniform_keep25_unhealed.log ADDED
@@ -0,0 +1,8 @@
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "correct": 118,
3
+ "accuracy": 0.08946171341925702,
4
+ "finished": 272,
5
+ "finish_rate": 0.20621683093252463,
6
+ "mean_completion_tokens": 477.21379833206976
7
+ }
8
+ saved item-level results -> outputs/evals/grid_math_unhealed/uniform_keep25_unhealed_chat.json
evals/grid_math_unhealed/uniform_keep25_unhealed_chat.json ADDED
The diff for this file is too large to render. See raw diff
 
evals/grid_math_unhealed/uniform_keep25_unhealed_chat.json.server.log ADDED
@@ -0,0 +1,115 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ Skipping import of cpp extensions due to incompatible torch version 2.10.0+cu128 for torchao version 0.15.0 Please see https://github.com/pytorch/ao/issues/2919 for more info
2
+ WARNING 08-10 21:24:57 [registry.py:915] Model architecture OlmoeForCausalLM is already registered, and will be overwritten by the new model class glean_vllm.pruned_olmoe:PrunedOlmoeForCausalLM.
3
+ (APIServer pid=1153465) INFO 08-10 21:24:57 [utils.py:299]
4
+ (APIServer pid=1153465) INFO 08-10 21:24:57 [utils.py:299] β–ˆ β–ˆ β–ˆβ–„ β–„β–ˆ
5
+ (APIServer pid=1153465) INFO 08-10 21:24:57 [utils.py:299] β–„β–„ β–„β–ˆ β–ˆ β–ˆ β–ˆ β–€β–„β–€ β–ˆ version 0.19.0
6
+ (APIServer pid=1153465) INFO 08-10 21:24:57 [utils.py:299] β–ˆβ–„β–ˆβ–€ β–ˆ β–ˆ β–ˆ β–ˆ model /media/henry/MoreFiles/winnow_release/math_arms/uniform_keep25
7
+ (APIServer pid=1153465) INFO 08-10 21:24:57 [utils.py:299] β–€β–€ β–€β–€β–€β–€β–€ β–€β–€β–€β–€β–€ β–€ β–€
8
+ (APIServer pid=1153465) INFO 08-10 21:24:57 [utils.py:299]
9
+ (APIServer pid=1153465) INFO 08-10 21:24:57 [utils.py:233] non-default args: {'model_tag': '/media/henry/MoreFiles/winnow_release/math_arms/uniform_keep25', 'host': '127.0.0.1', 'port': 8612, 'model': '/media/henry/MoreFiles/winnow_release/math_arms/uniform_keep25', 'max_model_len': 2048, 'enforce_eager': True, 'served_model_name': ['student'], 'gpu_memory_utilization': 0.85}
10
+ (APIServer pid=1153465) INFO 08-10 21:25:05 [model.py:549] Resolved architecture: PrunedOlmoeForCausalLM
11
+ (APIServer pid=1153465) INFO 08-10 21:25:05 [model.py:1678] Using max model len 2048
12
+ (APIServer pid=1153465) INFO 08-10 21:25:06 [vllm.py:790] Asynchronous scheduling is enabled.
13
+ (APIServer pid=1153465) WARNING 08-10 21:25:06 [vllm.py:848] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
14
+ (APIServer pid=1153465) WARNING 08-10 21:25:06 [vllm.py:859] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
15
+ (APIServer pid=1153465) INFO 08-10 21:25:06 [vllm.py:1025] Cudagraph is disabled under eager mode
16
+ (APIServer pid=1153465) INFO 08-10 21:25:06 [compilation.py:290] Enabled custom fusions: norm_quant, act_quant
17
+ Skipping import of cpp extensions due to incompatible torch version 2.10.0+cu128 for torchao version 0.15.0 Please see https://github.com/pytorch/ao/issues/2919 for more info
18
+ (EngineCore pid=1154743) WARNING 08-10 21:25:14 [registry.py:915] Model architecture OlmoeForCausalLM is already registered, and will be overwritten by the new model class glean_vllm.pruned_olmoe:PrunedOlmoeForCausalLM.
19
+ (EngineCore pid=1154743) INFO 08-10 21:25:14 [core.py:105] Initializing a V1 LLM engine (v0.19.0) with config: model='/media/henry/MoreFiles/winnow_release/math_arms/uniform_keep25', speculative_config=None, tokenizer='/media/henry/MoreFiles/winnow_release/math_arms/uniform_keep25', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=False, dtype=torch.bfloat16, max_seq_len=2048, download_dir=None, load_format=auto, tensor_parallel_size=1, pipeline_parallel_size=1, data_parallel_size=1, decode_context_parallel_size=1, dcp_comm_backend=ag_rs, disable_custom_all_reduce=False, quantization=None, enforce_eager=True, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False), seed=0, served_model_name=student, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'mode': <CompilationMode.NONE: 0>, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'splitting_ops': [], 'compile_mm_encoder': False, 'cudagraph_mm_encoder': False, 'encoder_cudagraph_token_budgets': [], 'encoder_cudagraph_max_images_per_batch': 0, 'compile_sizes': [], 'compile_ranges_endpoints': [2048], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'size_asserts': False, 'alignment_asserts': False, 'scalar_asserts': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': <CUDAGraphMode.NONE: 0>, 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': True, 'fuse_act_quant': True, 'fuse_attn_quant': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False}, 'max_cudagraph_capture_size': 0, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': True, 'static_all_moe_layers': []}
20
+ (EngineCore pid=1154743) INFO 08-10 21:25:14 [parallel_state.py:1400] world_size=1 rank=0 local_rank=0 distributed_init_method=tcp://192.168.0.27:44057 backend=nccl
21
+ (EngineCore pid=1154743) INFO 08-10 21:25:14 [parallel_state.py:1716] rank 0 in world size 1 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank 0, EPLB rank N/A
22
+ (EngineCore pid=1154743) INFO 08-10 21:25:15 [gpu_model_runner.py:4735] Starting to load model /media/henry/MoreFiles/winnow_release/math_arms/uniform_keep25...
23
+ (EngineCore pid=1154743) INFO 08-10 21:25:16 [cuda.py:334] Using FLASH_ATTN attention backend out of potential backends: ['FLASH_ATTN', 'FLASHINFER', 'TRITON_ATTN', 'FLEX_ATTENTION'].
24
+ (EngineCore pid=1154743) INFO 08-10 21:25:16 [flash_attn.py:596] Using FlashAttention version 2
25
+ (EngineCore pid=1154743)
26
+ (EngineCore pid=1154743)
27
+ (EngineCore pid=1154743)
28
+ (EngineCore pid=1154743)
29
+ (EngineCore pid=1154743) INFO 08-10 21:25:20 [default_loader.py:384] Loading weights took 4.10 seconds
30
+ (EngineCore pid=1154743) INFO 08-10 21:25:21 [gpu_model_runner.py:4820] Model loading took 3.89 GiB memory and 4.696206 seconds
31
+ (EngineCore pid=1154743) INFO 08-10 21:25:23 [gpu_worker.py:436] Available KV cache memory: 15.86 GiB
32
+ (EngineCore pid=1154743) INFO 08-10 21:25:23 [kv_cache_utils.py:1319] GPU KV cache size: 129,888 tokens
33
+ (EngineCore pid=1154743) INFO 08-10 21:25:23 [kv_cache_utils.py:1324] Maximum concurrency for 2,048 tokens per request: 63.42x
34
+ (EngineCore pid=1154743) INFO 08-10 21:25:24 [core.py:283] init engine (profile, create kv cache, warmup model) took 2.86 seconds
35
+ (EngineCore pid=1154743) INFO 08-10 21:25:24 [vllm.py:790] Asynchronous scheduling is enabled.
36
+ (EngineCore pid=1154743) WARNING 08-10 21:25:24 [vllm.py:848] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
37
+ (EngineCore pid=1154743) WARNING 08-10 21:25:24 [vllm.py:859] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
38
+ (EngineCore pid=1154743) INFO 08-10 21:25:24 [vllm.py:1025] Cudagraph is disabled under eager mode
39
+ (EngineCore pid=1154743) INFO 08-10 21:25:24 [compilation.py:290] Enabled custom fusions: norm_quant, act_quant
40
+ (APIServer pid=1153465) INFO 08-10 21:25:24 [api_server.py:590] Supported tasks: ['generate']
41
+ (APIServer pid=1153465) WARNING 08-10 21:25:24 [__init__.py:14] SECURITY WARNING: Development endpoints are enabled! This should NOT be used in production!
42
+ (APIServer pid=1153465) INFO 08-10 21:25:24 [hf.py:314] Detected the chat template content format to be 'string'. You can set `--chat-template-content-format` to override this.
43
+ (APIServer pid=1153465) INFO 08-10 21:25:24 [api_server.py:594] Starting vLLM server on http://127.0.0.1:8612
44
+ (APIServer pid=1153465) INFO 08-10 21:25:24 [launcher.py:37] Available routes are:
45
+ (APIServer pid=1153465) INFO 08-10 21:25:24 [launcher.py:46] Route: /openapi.json, Methods: GET, HEAD
46
+ (APIServer pid=1153465) INFO 08-10 21:25:24 [launcher.py:46] Route: /docs, Methods: GET, HEAD
47
+ (APIServer pid=1153465) INFO 08-10 21:25:24 [launcher.py:46] Route: /docs/oauth2-redirect, Methods: GET, HEAD
48
+ (APIServer pid=1153465) INFO 08-10 21:25:24 [launcher.py:46] Route: /redoc, Methods: GET, HEAD
49
+ (APIServer pid=1153465) INFO 08-10 21:25:24 [launcher.py:46] Route: /sleep, Methods: POST
50
+ (APIServer pid=1153465) INFO 08-10 21:25:24 [launcher.py:46] Route: /wake_up, Methods: POST
51
+ (APIServer pid=1153465) INFO 08-10 21:25:24 [launcher.py:46] Route: /is_sleeping, Methods: GET
52
+ (APIServer pid=1153465) INFO 08-10 21:25:24 [launcher.py:46] Route: /collective_rpc, Methods: POST
53
+ (APIServer pid=1153465) INFO 08-10 21:25:24 [launcher.py:46] Route: /reset_prefix_cache, Methods: POST
54
+ (APIServer pid=1153465) INFO 08-10 21:25:24 [launcher.py:46] Route: /reset_mm_cache, Methods: POST
55
+ (APIServer pid=1153465) INFO 08-10 21:25:24 [launcher.py:46] Route: /reset_encoder_cache, Methods: POST
56
+ (APIServer pid=1153465) INFO 08-10 21:25:24 [launcher.py:46] Route: /tokenize, Methods: POST
57
+ (APIServer pid=1153465) INFO 08-10 21:25:24 [launcher.py:46] Route: /detokenize, Methods: POST
58
+ (APIServer pid=1153465) INFO 08-10 21:25:24 [launcher.py:46] Route: /load, Methods: GET
59
+ (APIServer pid=1153465) INFO 08-10 21:25:24 [launcher.py:46] Route: /version, Methods: GET
60
+ (APIServer pid=1153465) INFO 08-10 21:25:24 [launcher.py:46] Route: /health, Methods: GET
61
+ (APIServer pid=1153465) INFO 08-10 21:25:24 [launcher.py:46] Route: /metrics, Methods: GET
62
+ (APIServer pid=1153465) INFO 08-10 21:25:24 [launcher.py:46] Route: /server_info, Methods: GET
63
+ (APIServer pid=1153465) INFO 08-10 21:25:24 [launcher.py:46] Route: /v1/models, Methods: GET
64
+ (APIServer pid=1153465) INFO 08-10 21:25:24 [launcher.py:46] Route: /ping, Methods: GET
65
+ (APIServer pid=1153465) INFO 08-10 21:25:24 [launcher.py:46] Route: /ping, Methods: POST
66
+ (APIServer pid=1153465) INFO 08-10 21:25:24 [launcher.py:46] Route: /invocations, Methods: POST
67
+ (APIServer pid=1153465) INFO 08-10 21:25:24 [launcher.py:46] Route: /v1/chat/completions, Methods: POST
68
+ (APIServer pid=1153465) INFO 08-10 21:25:24 [launcher.py:46] Route: /v1/chat/completions/batch, Methods: POST
69
+ (APIServer pid=1153465) INFO 08-10 21:25:24 [launcher.py:46] Route: /v1/responses, Methods: POST
70
+ (APIServer pid=1153465) INFO 08-10 21:25:24 [launcher.py:46] Route: /v1/responses/{response_id}, Methods: GET
71
+ (APIServer pid=1153465) INFO 08-10 21:25:24 [launcher.py:46] Route: /v1/responses/{response_id}/cancel, Methods: POST
72
+ (APIServer pid=1153465) INFO 08-10 21:25:24 [launcher.py:46] Route: /v1/completions, Methods: POST
73
+ (APIServer pid=1153465) INFO 08-10 21:25:24 [launcher.py:46] Route: /v1/messages, Methods: POST
74
+ (APIServer pid=1153465) INFO 08-10 21:25:24 [launcher.py:46] Route: /v1/messages/count_tokens, Methods: POST
75
+ (APIServer pid=1153465) INFO 08-10 21:25:24 [launcher.py:46] Route: /inference/v1/generate, Methods: POST
76
+ (APIServer pid=1153465) INFO 08-10 21:25:24 [launcher.py:46] Route: /pause, Methods: POST
77
+ (APIServer pid=1153465) INFO 08-10 21:25:24 [launcher.py:46] Route: /resume, Methods: POST
78
+ (APIServer pid=1153465) INFO 08-10 21:25:24 [launcher.py:46] Route: /is_paused, Methods: GET
79
+ (APIServer pid=1153465) INFO 08-10 21:25:24 [launcher.py:46] Route: /init_weight_transfer_engine, Methods: POST
80
+ (APIServer pid=1153465) INFO 08-10 21:25:24 [launcher.py:46] Route: /update_weights, Methods: POST
81
+ (APIServer pid=1153465) INFO 08-10 21:25:24 [launcher.py:46] Route: /get_world_size, Methods: GET
82
+ (APIServer pid=1153465) INFO 08-10 21:25:24 [launcher.py:46] Route: /scale_elastic_ep, Methods: POST
83
+ (APIServer pid=1153465) INFO 08-10 21:25:24 [launcher.py:46] Route: /is_scaling_elastic_ep, Methods: POST
84
+ (APIServer pid=1153465) INFO 08-10 21:25:24 [launcher.py:46] Route: /v1/chat/completions/render, Methods: POST
85
+ (APIServer pid=1153465) INFO 08-10 21:25:24 [launcher.py:46] Route: /v1/completions/render, Methods: POST
86
+ (APIServer pid=1153465) INFO: Started server process [1153465]
87
+ (APIServer pid=1153465) INFO: Waiting for application startup.
88
+ (APIServer pid=1153465) INFO: Application startup complete.
89
+ (APIServer pid=1153465) INFO: 127.0.0.1:33884 - "GET /health HTTP/1.1" 200 OK
90
+ (APIServer pid=1153465) INFO 08-10 21:25:35 [loggers.py:259] Engine 000: Avg prompt throughput: 2041.2 tokens/s, Avg generation throughput: 2581.6 tokens/s, Running: 256 reqs, Waiting: 1060 reqs, GPU KV cache usage: 38.7%, Prefix cache hit rate: 93.1%
91
+ (APIServer pid=1153465) INFO 08-10 21:25:45 [loggers.py:259] Engine 000: Avg prompt throughput: 23.3 tokens/s, Avg generation throughput: 3582.5 tokens/s, Running: 256 reqs, Waiting: 1056 reqs, GPU KV cache usage: 65.7%, Prefix cache hit rate: 93.1%
92
+ (APIServer pid=1153465) INFO 08-10 21:25:55 [loggers.py:259] Engine 000: Avg prompt throughput: 77.7 tokens/s, Avg generation throughput: 3325.8 tokens/s, Running: 256 reqs, Waiting: 1045 reqs, GPU KV cache usage: 88.7%, Prefix cache hit rate: 93.1%
93
+ (APIServer pid=1153465) INFO 08-10 21:26:05 [loggers.py:259] Engine 000: Avg prompt throughput: 145.0 tokens/s, Avg generation throughput: 2930.4 tokens/s, Running: 229 reqs, Waiting: 1046 reqs, GPU KV cache usage: 99.9%, Prefix cache hit rate: 93.1%
94
+ (APIServer pid=1153465) INFO 08-10 21:26:15 [loggers.py:259] Engine 000: Avg prompt throughput: 1818.6 tokens/s, Avg generation throughput: 3485.2 tokens/s, Running: 256 reqs, Waiting: 796 reqs, GPU KV cache usage: 43.2%, Prefix cache hit rate: 93.3%
95
+ (APIServer pid=1153465) INFO 08-10 21:26:25 [loggers.py:259] Engine 000: Avg prompt throughput: 73.1 tokens/s, Avg generation throughput: 3530.9 tokens/s, Running: 256 reqs, Waiting: 787 reqs, GPU KV cache usage: 68.5%, Prefix cache hit rate: 93.3%
96
+ (APIServer pid=1153465) INFO 08-10 21:26:35 [loggers.py:259] Engine 000: Avg prompt throughput: 114.3 tokens/s, Avg generation throughput: 3299.9 tokens/s, Running: 255 reqs, Waiting: 770 reqs, GPU KV cache usage: 87.7%, Prefix cache hit rate: 93.3%
97
+ (APIServer pid=1153465) INFO 08-10 21:26:45 [loggers.py:259] Engine 000: Avg prompt throughput: 487.9 tokens/s, Avg generation throughput: 3165.6 tokens/s, Running: 239 reqs, Waiting: 690 reqs, GPU KV cache usage: 75.2%, Prefix cache hit rate: 93.3%
98
+ (APIServer pid=1153465) INFO 08-10 21:26:55 [loggers.py:259] Engine 000: Avg prompt throughput: 1483.6 tokens/s, Avg generation throughput: 3591.5 tokens/s, Running: 256 reqs, Waiting: 520 reqs, GPU KV cache usage: 49.4%, Prefix cache hit rate: 93.3%
99
+ (APIServer pid=1153465) INFO 08-10 21:27:05 [loggers.py:259] Engine 000: Avg prompt throughput: 153.1 tokens/s, Avg generation throughput: 3478.6 tokens/s, Running: 256 reqs, Waiting: 502 reqs, GPU KV cache usage: 70.8%, Prefix cache hit rate: 93.3%
100
+ (APIServer pid=1153465) INFO 08-10 21:27:15 [loggers.py:259] Engine 000: Avg prompt throughput: 247.5 tokens/s, Avg generation throughput: 3298.3 tokens/s, Running: 255 reqs, Waiting: 471 reqs, GPU KV cache usage: 85.2%, Prefix cache hit rate: 93.3%
101
+ (APIServer pid=1153465) INFO 08-10 21:27:25 [loggers.py:259] Engine 000: Avg prompt throughput: 1526.9 tokens/s, Avg generation throughput: 3154.4 tokens/s, Running: 256 reqs, Waiting: 283 reqs, GPU KV cache usage: 37.7%, Prefix cache hit rate: 93.4%
102
+ (APIServer pid=1153465) INFO 08-10 21:27:35 [loggers.py:259] Engine 000: Avg prompt throughput: 279.5 tokens/s, Avg generation throughput: 3605.4 tokens/s, Running: 256 reqs, Waiting: 246 reqs, GPU KV cache usage: 54.2%, Prefix cache hit rate: 93.4%
103
+ (APIServer pid=1153465) INFO 08-10 21:27:45 [loggers.py:259] Engine 000: Avg prompt throughput: 155.5 tokens/s, Avg generation throughput: 3453.2 tokens/s, Running: 256 reqs, Waiting: 228 reqs, GPU KV cache usage: 74.6%, Prefix cache hit rate: 93.4%
104
+ (APIServer pid=1153465) INFO 08-10 21:27:55 [loggers.py:259] Engine 000: Avg prompt throughput: 358.4 tokens/s, Avg generation throughput: 3271.9 tokens/s, Running: 256 reqs, Waiting: 181 reqs, GPU KV cache usage: 82.9%, Prefix cache hit rate: 93.4%
105
+ (APIServer pid=1153465) INFO 08-10 21:28:05 [loggers.py:259] Engine 000: Avg prompt throughput: 1343.6 tokens/s, Avg generation throughput: 3259.3 tokens/s, Running: 256 reqs, Waiting: 15 reqs, GPU KV cache usage: 44.9%, Prefix cache hit rate: 93.4%
106
+ (APIServer pid=1153465) INFO 08-10 21:28:15 [loggers.py:259] Engine 000: Avg prompt throughput: 116.1 tokens/s, Avg generation throughput: 3538.8 tokens/s, Running: 228 reqs, Waiting: 0 reqs, GPU KV cache usage: 55.7%, Prefix cache hit rate: 93.4%
107
+ (APIServer pid=1153465) INFO 08-10 21:28:25 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 3280.1 tokens/s, Running: 202 reqs, Waiting: 0 reqs, GPU KV cache usage: 69.6%, Prefix cache hit rate: 93.4%
108
+ (APIServer pid=1153465) INFO 08-10 21:28:35 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 2844.0 tokens/s, Running: 79 reqs, Waiting: 0 reqs, GPU KV cache usage: 36.5%, Prefix cache hit rate: 93.4%
109
+ (APIServer pid=1153465) INFO: 127.0.0.1:33886 - "POST /v1/completions HTTP/1.1" 200 OK
110
+ (EngineCore pid=1154743) INFO 08-10 21:28:37 [core.py:1210] Shutdown initiated (timeout=0)
111
+ (EngineCore pid=1154743) INFO 08-10 21:28:37 [core.py:1233] Shutdown complete
112
+ (APIServer pid=1153465) INFO: Shutting down
113
+ (APIServer pid=1153465) INFO: Waiting for application shutdown.
114
+ (APIServer pid=1153465) INFO: Application shutdown complete.
115
+ (APIServer pid=1153465) INFO: Finished server process [1153465]
evals/grid_math_unhealed/uniform_keep50_unhealed.log ADDED
@@ -0,0 +1,8 @@
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "correct": 548,
3
+ "accuracy": 0.41546626231993933,
4
+ "finished": 1315,
5
+ "finish_rate": 0.9969673995451099,
6
+ "mean_completion_tokens": 94.5352539802881
7
+ }
8
+ saved item-level results -> outputs/evals/grid_math_unhealed/uniform_keep50_unhealed_chat.json
evals/grid_math_unhealed/uniform_keep50_unhealed_chat.json ADDED
The diff for this file is too large to render. See raw diff
 
evals/grid_math_unhealed/uniform_keep50_unhealed_chat.json.server.log ADDED
@@ -0,0 +1,101 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ Skipping import of cpp extensions due to incompatible torch version 2.10.0+cu128 for torchao version 0.15.0 Please see https://github.com/pytorch/ao/issues/2919 for more info
2
+ WARNING 08-10 21:22:07 [registry.py:915] Model architecture OlmoeForCausalLM is already registered, and will be overwritten by the new model class glean_vllm.pruned_olmoe:PrunedOlmoeForCausalLM.
3
+ (APIServer pid=1147687) INFO 08-10 21:22:07 [utils.py:299]
4
+ (APIServer pid=1147687) INFO 08-10 21:22:07 [utils.py:299] β–ˆ β–ˆ β–ˆβ–„ β–„β–ˆ
5
+ (APIServer pid=1147687) INFO 08-10 21:22:07 [utils.py:299] β–„β–„ β–„β–ˆ β–ˆ β–ˆ β–ˆ β–€β–„β–€ β–ˆ version 0.19.0
6
+ (APIServer pid=1147687) INFO 08-10 21:22:07 [utils.py:299] β–ˆβ–„β–ˆβ–€ β–ˆ β–ˆ β–ˆ β–ˆ model /media/henry/MoreFiles/winnow_release/math_arms/uniform_keep50
7
+ (APIServer pid=1147687) INFO 08-10 21:22:07 [utils.py:299] β–€β–€ β–€β–€β–€β–€β–€ β–€β–€β–€β–€β–€ β–€ β–€
8
+ (APIServer pid=1147687) INFO 08-10 21:22:07 [utils.py:299]
9
+ (APIServer pid=1147687) INFO 08-10 21:22:07 [utils.py:233] non-default args: {'model_tag': '/media/henry/MoreFiles/winnow_release/math_arms/uniform_keep50', 'host': '127.0.0.1', 'port': 8612, 'model': '/media/henry/MoreFiles/winnow_release/math_arms/uniform_keep50', 'max_model_len': 2048, 'enforce_eager': True, 'served_model_name': ['student'], 'gpu_memory_utilization': 0.85}
10
+ (APIServer pid=1147687) INFO 08-10 21:22:15 [model.py:549] Resolved architecture: PrunedOlmoeForCausalLM
11
+ (APIServer pid=1147687) INFO 08-10 21:22:15 [model.py:1678] Using max model len 2048
12
+ (APIServer pid=1147687) INFO 08-10 21:22:15 [vllm.py:790] Asynchronous scheduling is enabled.
13
+ (APIServer pid=1147687) WARNING 08-10 21:22:15 [vllm.py:848] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
14
+ (APIServer pid=1147687) WARNING 08-10 21:22:15 [vllm.py:859] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
15
+ (APIServer pid=1147687) INFO 08-10 21:22:15 [vllm.py:1025] Cudagraph is disabled under eager mode
16
+ (APIServer pid=1147687) INFO 08-10 21:22:15 [compilation.py:290] Enabled custom fusions: norm_quant, act_quant
17
+ Skipping import of cpp extensions due to incompatible torch version 2.10.0+cu128 for torchao version 0.15.0 Please see https://github.com/pytorch/ao/issues/2919 for more info
18
+ (EngineCore pid=1148946) WARNING 08-10 21:22:24 [registry.py:915] Model architecture OlmoeForCausalLM is already registered, and will be overwritten by the new model class glean_vllm.pruned_olmoe:PrunedOlmoeForCausalLM.
19
+ (EngineCore pid=1148946) INFO 08-10 21:22:24 [core.py:105] Initializing a V1 LLM engine (v0.19.0) with config: model='/media/henry/MoreFiles/winnow_release/math_arms/uniform_keep50', speculative_config=None, tokenizer='/media/henry/MoreFiles/winnow_release/math_arms/uniform_keep50', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=False, dtype=torch.bfloat16, max_seq_len=2048, download_dir=None, load_format=auto, tensor_parallel_size=1, pipeline_parallel_size=1, data_parallel_size=1, decode_context_parallel_size=1, dcp_comm_backend=ag_rs, disable_custom_all_reduce=False, quantization=None, enforce_eager=True, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False), seed=0, served_model_name=student, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'mode': <CompilationMode.NONE: 0>, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'splitting_ops': [], 'compile_mm_encoder': False, 'cudagraph_mm_encoder': False, 'encoder_cudagraph_token_budgets': [], 'encoder_cudagraph_max_images_per_batch': 0, 'compile_sizes': [], 'compile_ranges_endpoints': [2048], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'size_asserts': False, 'alignment_asserts': False, 'scalar_asserts': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': <CUDAGraphMode.NONE: 0>, 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': True, 'fuse_act_quant': True, 'fuse_attn_quant': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False}, 'max_cudagraph_capture_size': 0, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': True, 'static_all_moe_layers': []}
20
+ (EngineCore pid=1148946) INFO 08-10 21:22:24 [parallel_state.py:1400] world_size=1 rank=0 local_rank=0 distributed_init_method=tcp://192.168.0.27:47297 backend=nccl
21
+ (EngineCore pid=1148946) INFO 08-10 21:22:24 [parallel_state.py:1716] rank 0 in world size 1 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank 0, EPLB rank N/A
22
+ (EngineCore pid=1148946) INFO 08-10 21:22:24 [gpu_model_runner.py:4735] Starting to load model /media/henry/MoreFiles/winnow_release/math_arms/uniform_keep50...
23
+ (EngineCore pid=1148946) INFO 08-10 21:22:25 [cuda.py:334] Using FLASH_ATTN attention backend out of potential backends: ['FLASH_ATTN', 'FLASHINFER', 'TRITON_ATTN', 'FLEX_ATTENTION'].
24
+ (EngineCore pid=1148946) INFO 08-10 21:22:25 [flash_attn.py:596] Using FlashAttention version 2
25
+ (EngineCore pid=1148946)
26
+ (EngineCore pid=1148946)
27
+ (EngineCore pid=1148946)
28
+ (EngineCore pid=1148946)
29
+ (EngineCore pid=1148946)
30
+ (EngineCore pid=1148946) INFO 08-10 21:22:34 [default_loader.py:384] Loading weights took 8.47 seconds
31
+ (EngineCore pid=1148946) INFO 08-10 21:22:34 [gpu_model_runner.py:4820] Model loading took 6.89 GiB memory and 9.059107 seconds
32
+ (EngineCore pid=1148946) INFO 08-10 21:22:37 [gpu_worker.py:436] Available KV cache memory: 12.86 GiB
33
+ (EngineCore pid=1148946) INFO 08-10 21:22:37 [kv_cache_utils.py:1319] GPU KV cache size: 105,312 tokens
34
+ (EngineCore pid=1148946) INFO 08-10 21:22:37 [kv_cache_utils.py:1324] Maximum concurrency for 2,048 tokens per request: 51.42x
35
+ (EngineCore pid=1148946) INFO 08-10 21:22:37 [core.py:283] init engine (profile, create kv cache, warmup model) took 2.42 seconds
36
+ (EngineCore pid=1148946) INFO 08-10 21:22:37 [vllm.py:790] Asynchronous scheduling is enabled.
37
+ (EngineCore pid=1148946) WARNING 08-10 21:22:37 [vllm.py:848] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
38
+ (EngineCore pid=1148946) WARNING 08-10 21:22:37 [vllm.py:859] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
39
+ (EngineCore pid=1148946) INFO 08-10 21:22:37 [vllm.py:1025] Cudagraph is disabled under eager mode
40
+ (EngineCore pid=1148946) INFO 08-10 21:22:37 [compilation.py:290] Enabled custom fusions: norm_quant, act_quant
41
+ (APIServer pid=1147687) INFO 08-10 21:22:37 [api_server.py:590] Supported tasks: ['generate']
42
+ (APIServer pid=1147687) WARNING 08-10 21:22:37 [__init__.py:14] SECURITY WARNING: Development endpoints are enabled! This should NOT be used in production!
43
+ (APIServer pid=1147687) INFO 08-10 21:22:37 [hf.py:314] Detected the chat template content format to be 'string'. You can set `--chat-template-content-format` to override this.
44
+ (APIServer pid=1147687) INFO 08-10 21:22:37 [api_server.py:594] Starting vLLM server on http://127.0.0.1:8612
45
+ (APIServer pid=1147687) INFO 08-10 21:22:37 [launcher.py:37] Available routes are:
46
+ (APIServer pid=1147687) INFO 08-10 21:22:37 [launcher.py:46] Route: /openapi.json, Methods: GET, HEAD
47
+ (APIServer pid=1147687) INFO 08-10 21:22:37 [launcher.py:46] Route: /docs, Methods: GET, HEAD
48
+ (APIServer pid=1147687) INFO 08-10 21:22:37 [launcher.py:46] Route: /docs/oauth2-redirect, Methods: GET, HEAD
49
+ (APIServer pid=1147687) INFO 08-10 21:22:37 [launcher.py:46] Route: /redoc, Methods: GET, HEAD
50
+ (APIServer pid=1147687) INFO 08-10 21:22:37 [launcher.py:46] Route: /sleep, Methods: POST
51
+ (APIServer pid=1147687) INFO 08-10 21:22:37 [launcher.py:46] Route: /wake_up, Methods: POST
52
+ (APIServer pid=1147687) INFO 08-10 21:22:37 [launcher.py:46] Route: /is_sleeping, Methods: GET
53
+ (APIServer pid=1147687) INFO 08-10 21:22:37 [launcher.py:46] Route: /collective_rpc, Methods: POST
54
+ (APIServer pid=1147687) INFO 08-10 21:22:37 [launcher.py:46] Route: /reset_prefix_cache, Methods: POST
55
+ (APIServer pid=1147687) INFO 08-10 21:22:37 [launcher.py:46] Route: /reset_mm_cache, Methods: POST
56
+ (APIServer pid=1147687) INFO 08-10 21:22:37 [launcher.py:46] Route: /reset_encoder_cache, Methods: POST
57
+ (APIServer pid=1147687) INFO 08-10 21:22:37 [launcher.py:46] Route: /tokenize, Methods: POST
58
+ (APIServer pid=1147687) INFO 08-10 21:22:37 [launcher.py:46] Route: /detokenize, Methods: POST
59
+ (APIServer pid=1147687) INFO 08-10 21:22:37 [launcher.py:46] Route: /load, Methods: GET
60
+ (APIServer pid=1147687) INFO 08-10 21:22:37 [launcher.py:46] Route: /version, Methods: GET
61
+ (APIServer pid=1147687) INFO 08-10 21:22:37 [launcher.py:46] Route: /health, Methods: GET
62
+ (APIServer pid=1147687) INFO 08-10 21:22:37 [launcher.py:46] Route: /metrics, Methods: GET
63
+ (APIServer pid=1147687) INFO 08-10 21:22:37 [launcher.py:46] Route: /server_info, Methods: GET
64
+ (APIServer pid=1147687) INFO 08-10 21:22:37 [launcher.py:46] Route: /v1/models, Methods: GET
65
+ (APIServer pid=1147687) INFO 08-10 21:22:37 [launcher.py:46] Route: /ping, Methods: GET
66
+ (APIServer pid=1147687) INFO 08-10 21:22:37 [launcher.py:46] Route: /ping, Methods: POST
67
+ (APIServer pid=1147687) INFO 08-10 21:22:37 [launcher.py:46] Route: /invocations, Methods: POST
68
+ (APIServer pid=1147687) INFO 08-10 21:22:37 [launcher.py:46] Route: /v1/chat/completions, Methods: POST
69
+ (APIServer pid=1147687) INFO 08-10 21:22:37 [launcher.py:46] Route: /v1/chat/completions/batch, Methods: POST
70
+ (APIServer pid=1147687) INFO 08-10 21:22:37 [launcher.py:46] Route: /v1/responses, Methods: POST
71
+ (APIServer pid=1147687) INFO 08-10 21:22:37 [launcher.py:46] Route: /v1/responses/{response_id}, Methods: GET
72
+ (APIServer pid=1147687) INFO 08-10 21:22:37 [launcher.py:46] Route: /v1/responses/{response_id}/cancel, Methods: POST
73
+ (APIServer pid=1147687) INFO 08-10 21:22:37 [launcher.py:46] Route: /v1/completions, Methods: POST
74
+ (APIServer pid=1147687) INFO 08-10 21:22:37 [launcher.py:46] Route: /v1/messages, Methods: POST
75
+ (APIServer pid=1147687) INFO 08-10 21:22:37 [launcher.py:46] Route: /v1/messages/count_tokens, Methods: POST
76
+ (APIServer pid=1147687) INFO 08-10 21:22:37 [launcher.py:46] Route: /inference/v1/generate, Methods: POST
77
+ (APIServer pid=1147687) INFO 08-10 21:22:37 [launcher.py:46] Route: /pause, Methods: POST
78
+ (APIServer pid=1147687) INFO 08-10 21:22:37 [launcher.py:46] Route: /resume, Methods: POST
79
+ (APIServer pid=1147687) INFO 08-10 21:22:37 [launcher.py:46] Route: /is_paused, Methods: GET
80
+ (APIServer pid=1147687) INFO 08-10 21:22:37 [launcher.py:46] Route: /init_weight_transfer_engine, Methods: POST
81
+ (APIServer pid=1147687) INFO 08-10 21:22:37 [launcher.py:46] Route: /update_weights, Methods: POST
82
+ (APIServer pid=1147687) INFO 08-10 21:22:37 [launcher.py:46] Route: /get_world_size, Methods: GET
83
+ (APIServer pid=1147687) INFO 08-10 21:22:37 [launcher.py:46] Route: /scale_elastic_ep, Methods: POST
84
+ (APIServer pid=1147687) INFO 08-10 21:22:37 [launcher.py:46] Route: /is_scaling_elastic_ep, Methods: POST
85
+ (APIServer pid=1147687) INFO 08-10 21:22:37 [launcher.py:46] Route: /v1/chat/completions/render, Methods: POST
86
+ (APIServer pid=1147687) INFO 08-10 21:22:37 [launcher.py:46] Route: /v1/completions/render, Methods: POST
87
+ (APIServer pid=1147687) INFO: Started server process [1147687]
88
+ (APIServer pid=1147687) INFO: Waiting for application startup.
89
+ (APIServer pid=1147687) INFO: Application startup complete.
90
+ (APIServer pid=1147687) INFO: 127.0.0.1:36040 - "GET /health HTTP/1.1" 200 OK
91
+ (APIServer pid=1147687) INFO 08-10 21:22:48 [loggers.py:259] Engine 000: Avg prompt throughput: 3508.3 tokens/s, Avg generation throughput: 2603.9 tokens/s, Running: 254 reqs, Waiting: 860 reqs, GPU KV cache usage: 35.0%, Prefix cache hit rate: 93.2%
92
+ (APIServer pid=1147687) INFO 08-10 21:22:58 [loggers.py:259] Engine 000: Avg prompt throughput: 2530.2 tokens/s, Avg generation throughput: 3217.5 tokens/s, Running: 255 reqs, Waiting: 532 reqs, GPU KV cache usage: 35.7%, Prefix cache hit rate: 93.3%
93
+ (APIServer pid=1147687) INFO 08-10 21:23:08 [loggers.py:259] Engine 000: Avg prompt throughput: 2784.6 tokens/s, Avg generation throughput: 3164.9 tokens/s, Running: 254 reqs, Waiting: 189 reqs, GPU KV cache usage: 36.6%, Prefix cache hit rate: 93.4%
94
+ (APIServer pid=1147687) INFO 08-10 21:23:18 [loggers.py:259] Engine 000: Avg prompt throughput: 1531.1 tokens/s, Avg generation throughput: 2918.6 tokens/s, Running: 102 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.7%, Prefix cache hit rate: 93.4%
95
+ (APIServer pid=1147687) INFO: 127.0.0.1:43906 - "POST /v1/completions HTTP/1.1" 200 OK
96
+ (EngineCore pid=1148946) INFO 08-10 21:23:22 [core.py:1210] Shutdown initiated (timeout=0)
97
+ (EngineCore pid=1148946) INFO 08-10 21:23:22 [core.py:1233] Shutdown complete
98
+ (APIServer pid=1147687) INFO: Shutting down
99
+ (APIServer pid=1147687) INFO: Waiting for application shutdown.
100
+ (APIServer pid=1147687) INFO: Application shutdown complete.
101
+ (APIServer pid=1147687) INFO: Finished server process [1147687]
evals/grid_math_unhealed/uniform_keep75_unhealed.log ADDED
@@ -0,0 +1,8 @@
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "correct": 816,
3
+ "accuracy": 0.6186504927975739,
4
+ "finished": 1316,
5
+ "finish_rate": 0.9977255496588324,
6
+ "mean_completion_tokens": 107.12736921910539
7
+ }
8
+ saved item-level results -> outputs/evals/grid_math_unhealed/uniform_keep75_unhealed_chat.json
evals/grid_math_unhealed/uniform_keep75_unhealed_chat.json ADDED
The diff for this file is too large to render. See raw diff
 
evals/grid_math_unhealed/uniform_keep75_unhealed_chat.json.server.log ADDED
@@ -0,0 +1,103 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ Skipping import of cpp extensions due to incompatible torch version 2.10.0+cu128 for torchao version 0.15.0 Please see https://github.com/pytorch/ao/issues/2919 for more info
2
+ WARNING 08-10 21:20:11 [registry.py:915] Model architecture OlmoeForCausalLM is already registered, and will be overwritten by the new model class glean_vllm.pruned_olmoe:PrunedOlmoeForCausalLM.
3
+ (APIServer pid=1142904) INFO 08-10 21:20:11 [utils.py:299]
4
+ (APIServer pid=1142904) INFO 08-10 21:20:11 [utils.py:299] β–ˆ β–ˆ β–ˆβ–„ β–„β–ˆ
5
+ (APIServer pid=1142904) INFO 08-10 21:20:11 [utils.py:299] β–„β–„ β–„β–ˆ β–ˆ β–ˆ β–ˆ β–€β–„β–€ β–ˆ version 0.19.0
6
+ (APIServer pid=1142904) INFO 08-10 21:20:11 [utils.py:299] β–ˆβ–„β–ˆβ–€ β–ˆ β–ˆ β–ˆ β–ˆ model /media/henry/MoreFiles/winnow_release/math_arms/uniform_keep75
7
+ (APIServer pid=1142904) INFO 08-10 21:20:11 [utils.py:299] β–€β–€ β–€β–€β–€β–€β–€ β–€β–€β–€β–€β–€ β–€ β–€
8
+ (APIServer pid=1142904) INFO 08-10 21:20:11 [utils.py:299]
9
+ (APIServer pid=1142904) INFO 08-10 21:20:11 [utils.py:233] non-default args: {'model_tag': '/media/henry/MoreFiles/winnow_release/math_arms/uniform_keep75', 'host': '127.0.0.1', 'port': 8612, 'model': '/media/henry/MoreFiles/winnow_release/math_arms/uniform_keep75', 'max_model_len': 2048, 'enforce_eager': True, 'served_model_name': ['student'], 'gpu_memory_utilization': 0.85}
10
+ (APIServer pid=1142904) INFO 08-10 21:20:21 [model.py:549] Resolved architecture: PrunedOlmoeForCausalLM
11
+ (APIServer pid=1142904) INFO 08-10 21:20:21 [model.py:1678] Using max model len 2048
12
+ (APIServer pid=1142904) INFO 08-10 21:20:22 [vllm.py:790] Asynchronous scheduling is enabled.
13
+ (APIServer pid=1142904) WARNING 08-10 21:20:22 [vllm.py:848] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
14
+ (APIServer pid=1142904) WARNING 08-10 21:20:22 [vllm.py:859] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
15
+ (APIServer pid=1142904) INFO 08-10 21:20:22 [vllm.py:1025] Cudagraph is disabled under eager mode
16
+ (APIServer pid=1142904) INFO 08-10 21:20:22 [compilation.py:290] Enabled custom fusions: norm_quant, act_quant
17
+ Skipping import of cpp extensions due to incompatible torch version 2.10.0+cu128 for torchao version 0.15.0 Please see https://github.com/pytorch/ao/issues/2919 for more info
18
+ (EngineCore pid=1144302) WARNING 08-10 21:20:30 [registry.py:915] Model architecture OlmoeForCausalLM is already registered, and will be overwritten by the new model class glean_vllm.pruned_olmoe:PrunedOlmoeForCausalLM.
19
+ (EngineCore pid=1144302) INFO 08-10 21:20:30 [core.py:105] Initializing a V1 LLM engine (v0.19.0) with config: model='/media/henry/MoreFiles/winnow_release/math_arms/uniform_keep75', speculative_config=None, tokenizer='/media/henry/MoreFiles/winnow_release/math_arms/uniform_keep75', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=False, dtype=torch.bfloat16, max_seq_len=2048, download_dir=None, load_format=auto, tensor_parallel_size=1, pipeline_parallel_size=1, data_parallel_size=1, decode_context_parallel_size=1, dcp_comm_backend=ag_rs, disable_custom_all_reduce=False, quantization=None, enforce_eager=True, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False), seed=0, served_model_name=student, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'mode': <CompilationMode.NONE: 0>, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'splitting_ops': [], 'compile_mm_encoder': False, 'cudagraph_mm_encoder': False, 'encoder_cudagraph_token_budgets': [], 'encoder_cudagraph_max_images_per_batch': 0, 'compile_sizes': [], 'compile_ranges_endpoints': [2048], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'size_asserts': False, 'alignment_asserts': False, 'scalar_asserts': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': <CUDAGraphMode.NONE: 0>, 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': True, 'fuse_act_quant': True, 'fuse_attn_quant': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False}, 'max_cudagraph_capture_size': 0, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': True, 'static_all_moe_layers': []}
20
+ (EngineCore pid=1144302) INFO 08-10 21:20:31 [parallel_state.py:1400] world_size=1 rank=0 local_rank=0 distributed_init_method=tcp://192.168.0.27:57179 backend=nccl
21
+ (EngineCore pid=1144302) INFO 08-10 21:20:31 [parallel_state.py:1716] rank 0 in world size 1 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank 0, EPLB rank N/A
22
+ (EngineCore pid=1144302) INFO 08-10 21:20:32 [gpu_model_runner.py:4735] Starting to load model /media/henry/MoreFiles/winnow_release/math_arms/uniform_keep75...
23
+ (EngineCore pid=1144302) INFO 08-10 21:20:33 [cuda.py:334] Using FLASH_ATTN attention backend out of potential backends: ['FLASH_ATTN', 'FLASHINFER', 'TRITON_ATTN', 'FLEX_ATTENTION'].
24
+ (EngineCore pid=1144302) INFO 08-10 21:20:33 [flash_attn.py:596] Using FlashAttention version 2
25
+ (EngineCore pid=1144302)
26
+ (EngineCore pid=1144302)
27
+ (EngineCore pid=1144302)
28
+ (EngineCore pid=1144302)
29
+ (EngineCore pid=1144302)
30
+ (EngineCore pid=1144302)
31
+ (EngineCore pid=1144302) INFO 08-10 21:20:35 [default_loader.py:384] Loading weights took 2.05 seconds
32
+ (EngineCore pid=1144302) INFO 08-10 21:20:36 [gpu_model_runner.py:4820] Model loading took 9.89 GiB memory and 3.191814 seconds
33
+ (EngineCore pid=1144302) INFO 08-10 21:20:38 [gpu_worker.py:436] Available KV cache memory: 9.86 GiB
34
+ (EngineCore pid=1144302) INFO 08-10 21:20:38 [kv_cache_utils.py:1319] GPU KV cache size: 80,736 tokens
35
+ (EngineCore pid=1144302) INFO 08-10 21:20:38 [kv_cache_utils.py:1324] Maximum concurrency for 2,048 tokens per request: 39.42x
36
+ (EngineCore pid=1144302) INFO 08-10 21:20:38 [core.py:283] init engine (profile, create kv cache, warmup model) took 2.64 seconds
37
+ (EngineCore pid=1144302) INFO 08-10 21:20:39 [vllm.py:790] Asynchronous scheduling is enabled.
38
+ (EngineCore pid=1144302) WARNING 08-10 21:20:39 [vllm.py:848] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
39
+ (EngineCore pid=1144302) WARNING 08-10 21:20:39 [vllm.py:859] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
40
+ (EngineCore pid=1144302) INFO 08-10 21:20:39 [vllm.py:1025] Cudagraph is disabled under eager mode
41
+ (EngineCore pid=1144302) INFO 08-10 21:20:39 [compilation.py:290] Enabled custom fusions: norm_quant, act_quant
42
+ (APIServer pid=1142904) INFO 08-10 21:20:39 [api_server.py:590] Supported tasks: ['generate']
43
+ (APIServer pid=1142904) WARNING 08-10 21:20:39 [__init__.py:14] SECURITY WARNING: Development endpoints are enabled! This should NOT be used in production!
44
+ (APIServer pid=1142904) INFO 08-10 21:20:39 [hf.py:314] Detected the chat template content format to be 'string'. You can set `--chat-template-content-format` to override this.
45
+ (APIServer pid=1142904) INFO 08-10 21:20:39 [api_server.py:594] Starting vLLM server on http://127.0.0.1:8612
46
+ (APIServer pid=1142904) INFO 08-10 21:20:39 [launcher.py:37] Available routes are:
47
+ (APIServer pid=1142904) INFO 08-10 21:20:39 [launcher.py:46] Route: /openapi.json, Methods: GET, HEAD
48
+ (APIServer pid=1142904) INFO 08-10 21:20:39 [launcher.py:46] Route: /docs, Methods: GET, HEAD
49
+ (APIServer pid=1142904) INFO 08-10 21:20:39 [launcher.py:46] Route: /docs/oauth2-redirect, Methods: GET, HEAD
50
+ (APIServer pid=1142904) INFO 08-10 21:20:39 [launcher.py:46] Route: /redoc, Methods: GET, HEAD
51
+ (APIServer pid=1142904) INFO 08-10 21:20:39 [launcher.py:46] Route: /sleep, Methods: POST
52
+ (APIServer pid=1142904) INFO 08-10 21:20:39 [launcher.py:46] Route: /wake_up, Methods: POST
53
+ (APIServer pid=1142904) INFO 08-10 21:20:39 [launcher.py:46] Route: /is_sleeping, Methods: GET
54
+ (APIServer pid=1142904) INFO 08-10 21:20:39 [launcher.py:46] Route: /collective_rpc, Methods: POST
55
+ (APIServer pid=1142904) INFO 08-10 21:20:39 [launcher.py:46] Route: /reset_prefix_cache, Methods: POST
56
+ (APIServer pid=1142904) INFO 08-10 21:20:39 [launcher.py:46] Route: /reset_mm_cache, Methods: POST
57
+ (APIServer pid=1142904) INFO 08-10 21:20:39 [launcher.py:46] Route: /reset_encoder_cache, Methods: POST
58
+ (APIServer pid=1142904) INFO 08-10 21:20:39 [launcher.py:46] Route: /tokenize, Methods: POST
59
+ (APIServer pid=1142904) INFO 08-10 21:20:39 [launcher.py:46] Route: /detokenize, Methods: POST
60
+ (APIServer pid=1142904) INFO 08-10 21:20:39 [launcher.py:46] Route: /load, Methods: GET
61
+ (APIServer pid=1142904) INFO 08-10 21:20:39 [launcher.py:46] Route: /version, Methods: GET
62
+ (APIServer pid=1142904) INFO 08-10 21:20:39 [launcher.py:46] Route: /health, Methods: GET
63
+ (APIServer pid=1142904) INFO 08-10 21:20:39 [launcher.py:46] Route: /metrics, Methods: GET
64
+ (APIServer pid=1142904) INFO 08-10 21:20:39 [launcher.py:46] Route: /server_info, Methods: GET
65
+ (APIServer pid=1142904) INFO 08-10 21:20:39 [launcher.py:46] Route: /v1/models, Methods: GET
66
+ (APIServer pid=1142904) INFO 08-10 21:20:39 [launcher.py:46] Route: /ping, Methods: GET
67
+ (APIServer pid=1142904) INFO 08-10 21:20:39 [launcher.py:46] Route: /ping, Methods: POST
68
+ (APIServer pid=1142904) INFO 08-10 21:20:39 [launcher.py:46] Route: /invocations, Methods: POST
69
+ (APIServer pid=1142904) INFO 08-10 21:20:39 [launcher.py:46] Route: /v1/chat/completions, Methods: POST
70
+ (APIServer pid=1142904) INFO 08-10 21:20:39 [launcher.py:46] Route: /v1/chat/completions/batch, Methods: POST
71
+ (APIServer pid=1142904) INFO 08-10 21:20:39 [launcher.py:46] Route: /v1/responses, Methods: POST
72
+ (APIServer pid=1142904) INFO 08-10 21:20:39 [launcher.py:46] Route: /v1/responses/{response_id}, Methods: GET
73
+ (APIServer pid=1142904) INFO 08-10 21:20:39 [launcher.py:46] Route: /v1/responses/{response_id}/cancel, Methods: POST
74
+ (APIServer pid=1142904) INFO 08-10 21:20:39 [launcher.py:46] Route: /v1/completions, Methods: POST
75
+ (APIServer pid=1142904) INFO 08-10 21:20:39 [launcher.py:46] Route: /v1/messages, Methods: POST
76
+ (APIServer pid=1142904) INFO 08-10 21:20:39 [launcher.py:46] Route: /v1/messages/count_tokens, Methods: POST
77
+ (APIServer pid=1142904) INFO 08-10 21:20:39 [launcher.py:46] Route: /inference/v1/generate, Methods: POST
78
+ (APIServer pid=1142904) INFO 08-10 21:20:39 [launcher.py:46] Route: /pause, Methods: POST
79
+ (APIServer pid=1142904) INFO 08-10 21:20:39 [launcher.py:46] Route: /resume, Methods: POST
80
+ (APIServer pid=1142904) INFO 08-10 21:20:39 [launcher.py:46] Route: /is_paused, Methods: GET
81
+ (APIServer pid=1142904) INFO 08-10 21:20:39 [launcher.py:46] Route: /init_weight_transfer_engine, Methods: POST
82
+ (APIServer pid=1142904) INFO 08-10 21:20:39 [launcher.py:46] Route: /update_weights, Methods: POST
83
+ (APIServer pid=1142904) INFO 08-10 21:20:39 [launcher.py:46] Route: /get_world_size, Methods: GET
84
+ (APIServer pid=1142904) INFO 08-10 21:20:39 [launcher.py:46] Route: /scale_elastic_ep, Methods: POST
85
+ (APIServer pid=1142904) INFO 08-10 21:20:39 [launcher.py:46] Route: /is_scaling_elastic_ep, Methods: POST
86
+ (APIServer pid=1142904) INFO 08-10 21:20:39 [launcher.py:46] Route: /v1/chat/completions/render, Methods: POST
87
+ (APIServer pid=1142904) INFO 08-10 21:20:39 [launcher.py:46] Route: /v1/completions/render, Methods: POST
88
+ (APIServer pid=1142904) INFO: Started server process [1142904]
89
+ (APIServer pid=1142904) INFO: Waiting for application startup.
90
+ (APIServer pid=1142904) INFO: Application startup complete.
91
+ (APIServer pid=1142904) INFO: 127.0.0.1:46028 - "GET /health HTTP/1.1" 200 OK
92
+ (APIServer pid=1142904) INFO 08-10 21:20:49 [loggers.py:259] Engine 000: Avg prompt throughput: 2616.4 tokens/s, Avg generation throughput: 1962.2 tokens/s, Running: 251 reqs, Waiting: 976 reqs, GPU KV cache usage: 48.0%, Prefix cache hit rate: 93.2%
93
+ (APIServer pid=1142904) INFO 08-10 21:20:59 [loggers.py:259] Engine 000: Avg prompt throughput: 2274.7 tokens/s, Avg generation throughput: 2965.4 tokens/s, Running: 252 reqs, Waiting: 686 reqs, GPU KV cache usage: 47.8%, Prefix cache hit rate: 93.3%
94
+ (APIServer pid=1142904) INFO 08-10 21:21:09 [loggers.py:259] Engine 000: Avg prompt throughput: 2289.4 tokens/s, Avg generation throughput: 2965.3 tokens/s, Running: 255 reqs, Waiting: 396 reqs, GPU KV cache usage: 47.5%, Prefix cache hit rate: 93.4%
95
+ (APIServer pid=1142904) INFO 08-10 21:21:19 [loggers.py:259] Engine 000: Avg prompt throughput: 2179.1 tokens/s, Avg generation throughput: 2993.4 tokens/s, Running: 255 reqs, Waiting: 125 reqs, GPU KV cache usage: 50.0%, Prefix cache hit rate: 93.4%
96
+ (APIServer pid=1142904) INFO 08-10 21:21:29 [loggers.py:259] Engine 000: Avg prompt throughput: 1048.0 tokens/s, Avg generation throughput: 2675.2 tokens/s, Running: 95 reqs, Waiting: 0 reqs, GPU KV cache usage: 28.3%, Prefix cache hit rate: 93.4%
97
+ (APIServer pid=1142904) INFO: 127.0.0.1:46038 - "POST /v1/completions HTTP/1.1" 200 OK
98
+ (EngineCore pid=1144302) INFO 08-10 21:21:36 [core.py:1210] Shutdown initiated (timeout=0)
99
+ (EngineCore pid=1144302) INFO 08-10 21:21:36 [core.py:1233] Shutdown complete
100
+ (APIServer pid=1142904) INFO: Shutting down
101
+ (APIServer pid=1142904) INFO: Waiting for application shutdown.
102
+ (APIServer pid=1142904) INFO: Application shutdown complete.
103
+ (APIServer pid=1142904) INFO: Finished server process [1142904]
evals/healing_breadth/glean_math_keep25_seed1224.json ADDED
The diff for this file is too large to render. See raw diff
 
evals/healing_breadth/glean_math_keep25_seed1224.json.server.log ADDED
@@ -0,0 +1,139 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ (APIServer pid=199309) INFO 07-14 21:55:42 [api_utils.py:339]
2
+ (APIServer pid=199309) INFO 07-14 21:55:42 [api_utils.py:339] β–ˆ β–ˆ β–ˆβ–„ β–„β–ˆ
3
+ (APIServer pid=199309) INFO 07-14 21:55:42 [api_utils.py:339] β–„β–„ β–„β–ˆ β–ˆ β–ˆ β–ˆ β–€β–„β–€ β–ˆ version 0.25.0
4
+ (APIServer pid=199309) INFO 07-14 21:55:42 [api_utils.py:339] β–ˆβ–„β–ˆβ–€ β–ˆ β–ˆ β–ˆ β–ˆ model outputs/healed/healing_breadth/glean_math_keep25_seed1224/step0050
5
+ (APIServer pid=199309) INFO 07-14 21:55:42 [api_utils.py:339] β–€β–€ β–€β–€β–€β–€β–€ β–€β–€β–€β–€β–€ β–€ β–€
6
+ (APIServer pid=199309) INFO 07-14 21:55:42 [api_utils.py:339]
7
+ (APIServer pid=199309) INFO 07-14 21:55:42 [api_utils.py:273] non-default args: {'model_tag': 'outputs/healed/healing_breadth/glean_math_keep25_seed1224/step0050', 'host': '127.0.0.1', 'port': 8377, 'model': 'outputs/healed/healing_breadth/glean_math_keep25_seed1224/step0050', 'max_model_len': 2048, 'enforce_eager': True, 'served_model_name': ['student'], 'gpu_memory_utilization': 0.85}
8
+ (APIServer pid=199309) INFO 07-14 21:55:42 [model.py:619] Resolved architecture: PrunedOlmoeForCausalLM
9
+ (APIServer pid=199309) INFO 07-14 21:55:42 [model.py:1776] Using max model len 2048
10
+ (APIServer pid=199309) INFO 07-14 21:55:42 [vllm.py:1042] Asynchronous scheduling is enabled.
11
+ (APIServer pid=199309) WARNING 07-14 21:55:42 [vllm.py:1096] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
12
+ (APIServer pid=199309) WARNING 07-14 21:55:42 [vllm.py:1144] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
13
+ (APIServer pid=199309) INFO 07-14 21:55:42 [kernel.py:292] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native'])
14
+ (APIServer pid=199309) INFO 07-14 21:55:42 [vllm.py:1322] Cudagraph is disabled under eager mode
15
+ (APIServer pid=199309) INFO 07-14 21:55:42 [compilation.py:312] Enabled custom fusions: norm_quant, act_quant
16
+ (EngineCore pid=199426) INFO 07-14 21:55:49 [core.py:114] Initializing a V1 LLM engine (v0.25.0) with config: model='outputs/healed/healing_breadth/glean_math_keep25_seed1224/step0050', speculative_config=None, tokenizer='outputs/healed/healing_breadth/glean_math_keep25_seed1224/step0050', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=False, dtype=torch.bfloat16, max_seq_len=2048, download_dir=None, load_format=auto, tensor_parallel_size=1, pipeline_parallel_size=1, data_parallel_size=1, decode_context_parallel_size=1, dcp_comm_backend=ag_rs, disable_custom_all_reduce=False, quantization=None, quantization_config=None, enforce_eager=True, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False, jit_monitor_mode='warn', jit_monitor_verbose=False), seed=0, served_model_name=student, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'mode': <CompilationMode.NONE: 0>, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'ir_enable_torch_wrap': False, 'splitting_ops': [], 'compile_mm_encoder': False, 'cudagraph_mm_encoder': False, 'encoder_cudagraph_token_budgets': [], 'encoder_cudagraph_max_vision_items_per_batch': 0, 'encoder_cudagraph_max_frames_per_batch': None, 'compile_sizes': [], 'compile_ranges_endpoints': [2048], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'size_asserts': False, 'alignment_asserts': False, 'scalar_asserts': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': <CUDAGraphMode.NONE: 0>, 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': True, 'fuse_act_quant': True, 'fuse_attn_quant': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False, 'fuse_rope_kvcache_cat_mla': False, 'fuse_act_padding': False}, 'max_cudagraph_capture_size': 0, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': False, 'static_all_moe_layers': []}, kernel_config=KernelConfig(ir_op_priority=IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native']), enable_flashinfer_autotune=True, enable_cutedsl_warmup=True, moe_backend='auto', linear_backend='auto')
17
+ (EngineCore pid=199426) INFO 07-14 21:55:50 [parallel_state.py:1607] world_size=1 rank=0 local_rank=0 distributed_init_method=tcp://192.168.0.15:42095 backend=nccl
18
+ (EngineCore pid=199426) INFO 07-14 21:55:50 [parallel_state.py:1942] rank 0 in world size 1 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank 0, EPLB rank N/A
19
+ (EngineCore pid=199426) INFO 07-14 21:55:51 [topk_topp_sampler.py:55] Using FlashInfer for top-p & top-k sampling.
20
+ (EngineCore pid=199426) INFO 07-14 21:55:51 [gpu_model_runner.py:5209] Starting to load model outputs/healed/healing_breadth/glean_math_keep25_seed1224/step0050...
21
+ (EngineCore pid=199426) INFO 07-14 21:55:51 [cuda.py:476] Using FLASH_ATTN attention backend out of potential backends: ['FLASH_ATTN', 'FLASHINFER', 'TRITON_ATTN', 'FLEX_ATTENTION'].
22
+ (EngineCore pid=199426) INFO 07-14 21:55:51 [flash_attn.py:718] Using FlashAttention version 2
23
+ (EngineCore pid=199426) /home/henry/.cache/glean/megablocks-variable-93a1479bc15b/megablocks/grouped_gemm_util.py:10: UserWarning: Grouped GEMM not available.
24
+ (EngineCore pid=199426) warnings.warn('Grouped GEMM not available.')
25
+ (EngineCore pid=199426) INFO 07-14 21:55:51 [weight_utils.py:849] Filesystem type for checkpoints: EXT4. Checkpoint size: 3.89 GiB. Available RAM: 77.80 GiB.
26
+ (EngineCore pid=199426) INFO 07-14 21:55:51 [weight_utils.py:872] Auto-prefetch is disabled because the filesystem (EXT4) is not a recognized network FS (NFS/Lustre). If you want to force prefetching, start vLLM with --safetensors-load-strategy=prefetch.
27
+ (EngineCore pid=199426)
28
+ (EngineCore pid=199426)
29
+ (EngineCore pid=199426)
30
+ (EngineCore pid=199426)
31
+ (EngineCore pid=199426) INFO 07-14 21:55:53 [default_loader.py:430] Loading weights took 2.44 seconds
32
+ (EngineCore pid=199426) INFO 07-14 21:55:54 [gpu_model_runner.py:5306] Model loading took 3.89 GiB memory and 2.620414 seconds
33
+ (EngineCore pid=199426) INFO 07-14 21:55:56 [gpu_worker.py:538] Available KV cache memory: 15.82 GiB
34
+ (EngineCore pid=199426) INFO 07-14 21:55:56 [kv_cache_utils.py:2146] GPU KV cache size: 129,584 tokens
35
+ (EngineCore pid=199426) INFO 07-14 21:55:56 [kv_cache_utils.py:2147] Maximum concurrency for 2,048 tokens per request: 63.27x
36
+ (EngineCore pid=199426) INFO 07-14 21:55:56 [cutedsl_warmup.py:97] Skipping CuTeDSL warmup because no compile units were requested.
37
+ (EngineCore pid=199426) INFO 07-14 21:55:56 [jit_monitor.py:73] Kernel JIT monitor activated; monitored JIT compilations during inference will use mode=warn.
38
+ (EngineCore pid=199426) INFO 07-14 21:55:56 [core.py:344] init engine (profile, create kv cache, warmup model) took 2.05 s
39
+ (EngineCore pid=199426) INFO 07-14 21:55:56 [vllm.py:1042] Asynchronous scheduling is enabled.
40
+ (EngineCore pid=199426) WARNING 07-14 21:55:56 [vllm.py:1096] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
41
+ (EngineCore pid=199426) WARNING 07-14 21:55:56 [vllm.py:1144] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
42
+ (EngineCore pid=199426) INFO 07-14 21:55:56 [kernel.py:292] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native'])
43
+ (EngineCore pid=199426) INFO 07-14 21:55:56 [vllm.py:1322] Cudagraph is disabled under eager mode
44
+ (EngineCore pid=199426) INFO 07-14 21:55:56 [compilation.py:312] Enabled custom fusions: norm_quant, act_quant
45
+ (APIServer pid=199309) INFO 07-14 21:55:56 [api_server.py:612] Supported tasks: ['generate']
46
+ (APIServer pid=199309) WARNING 07-14 21:55:56 [__init__.py:36] SECURITY WARNING: Development endpoints are enabled! This should NOT be used in production!
47
+ (APIServer pid=199309) INFO 07-14 21:55:56 [hf.py:548] Detected the chat template content format to be 'string'. You can set `--chat-template-content-format` to override this.
48
+ (APIServer pid=199309) INFO 07-14 21:55:56 [api_server.py:616] Starting vLLM server on http://127.0.0.1:8377
49
+ (APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:37] Available routes are:
50
+ (APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /openapi.json, Methods: GET, HEAD
51
+ (APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /docs, Methods: GET, HEAD
52
+ (APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /docs/oauth2-redirect, Methods: GET, HEAD
53
+ (APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /redoc, Methods: GET, HEAD
54
+ (APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /load, Methods: GET
55
+ (APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /version, Methods: GET
56
+ (APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /health, Methods: GET
57
+ (APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /metrics, Methods: GET
58
+ (APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /tokenize, Methods: POST
59
+ (APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /detokenize, Methods: POST
60
+ (APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /v1/models, Methods: GET
61
+ (APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /ping, Methods: GET
62
+ (APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /ping, Methods: POST
63
+ (APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /invocations, Methods: POST
64
+ (APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /reset_prefix_cache, Methods: POST
65
+ (APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /reset_mm_cache, Methods: POST
66
+ (APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /reset_encoder_cache, Methods: POST
67
+ (APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /pause, Methods: POST
68
+ (APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /resume, Methods: POST
69
+ (APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /is_paused, Methods: GET
70
+ (APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /init_weight_transfer_engine, Methods: POST
71
+ (APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /start_weight_update, Methods: POST
72
+ (APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /update_weights, Methods: POST
73
+ (APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /finish_weight_update, Methods: POST
74
+ (APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /get_world_size, Methods: GET
75
+ (APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /collective_rpc, Methods: POST
76
+ (APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /server_info, Methods: GET
77
+ (APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /sleep, Methods: POST
78
+ (APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /wake_up, Methods: POST
79
+ (APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /is_sleeping, Methods: GET
80
+ (APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /v1/chat/completions, Methods: POST
81
+ (APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /v1/chat/completions/batch, Methods: POST
82
+ (APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /v1/responses, Methods: POST
83
+ (APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /v1/responses/{response_id}, Methods: GET
84
+ (APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /v1/responses/{response_id}/cancel, Methods: POST
85
+ (APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /v1/completions, Methods: POST
86
+ (APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /v1/messages, Methods: POST
87
+ (APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /v1/messages/count_tokens, Methods: POST
88
+ (APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /generative_scoring, Methods: POST
89
+ (APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /scale_elastic_ep, Methods: POST
90
+ (APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /is_scaling_elastic_ep, Methods: POST
91
+ (APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /v1/chat/completions/render, Methods: POST
92
+ (APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /v1/completions/render, Methods: POST
93
+ (APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /v1/chat/completions/derender, Methods: POST
94
+ (APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /v1/completions/derender, Methods: POST
95
+ (APIServer pid=199309) INFO 07-14 21:55:56 [launcher.py:46] Route: /inference/v1/generate, Methods: POST
96
+ (APIServer pid=199309) INFO: Started server process [199309]
97
+ (APIServer pid=199309) INFO: Waiting for application startup.
98
+ (APIServer pid=199309) INFO: Application startup complete.
99
+ (APIServer pid=199309) INFO: 127.0.0.1:60678 - "GET /health HTTP/1.1" 200 OK
100
+ (EngineCore pid=199426) WARNING 07-14 21:55:58 [jit_monitor.py:129] Triton kernel JIT compilation during inference: _build_route_rows. This causes a latency spike; consider extending warmup to cover this shape/config.
101
+ (APIServer pid=199309) INFO 07-14 21:56:07 [loggers.py:273] Engine 000: Avg prompt throughput: 2049.1 tokens/s, Avg generation throughput: 2315.6 tokens/s, Running: 256 reqs, Waiting: 1059 reqs, GPU KV cache usage: 36.6%, Prefix cache hit rate: 93.1%
102
+ (APIServer pid=199309) INFO 07-14 21:56:17 [loggers.py:273] Engine 000: Avg prompt throughput: 281.0 tokens/s, Avg generation throughput: 2862.8 tokens/s, Running: 256 reqs, Waiting: 1022 reqs, GPU KV cache usage: 54.2%, Prefix cache hit rate: 93.1%
103
+ (APIServer pid=199309) INFO 07-14 21:56:27 [loggers.py:273] Engine 000: Avg prompt throughput: 467.0 tokens/s, Avg generation throughput: 2732.6 tokens/s, Running: 256 reqs, Waiting: 961 reqs, GPU KV cache usage: 63.6%, Prefix cache hit rate: 93.2%
104
+ (APIServer pid=199309) INFO 07-14 21:56:37 [loggers.py:273] Engine 000: Avg prompt throughput: 426.2 tokens/s, Avg generation throughput: 2733.7 tokens/s, Running: 255 reqs, Waiting: 908 reqs, GPU KV cache usage: 72.4%, Prefix cache hit rate: 93.2%
105
+ (APIServer pid=199309) INFO 07-14 21:56:47 [loggers.py:273] Engine 000: Avg prompt throughput: 1128.2 tokens/s, Avg generation throughput: 2494.4 tokens/s, Running: 256 reqs, Waiting: 764 reqs, GPU KV cache usage: 40.0%, Prefix cache hit rate: 93.3%
106
+ (APIServer pid=199309) INFO 07-14 21:56:57 [loggers.py:273] Engine 000: Avg prompt throughput: 319.6 tokens/s, Avg generation throughput: 2888.4 tokens/s, Running: 256 reqs, Waiting: 722 reqs, GPU KV cache usage: 53.6%, Prefix cache hit rate: 93.3%
107
+ (APIServer pid=199309) INFO 07-14 21:57:07 [loggers.py:273] Engine 000: Avg prompt throughput: 530.8 tokens/s, Avg generation throughput: 2834.6 tokens/s, Running: 255 reqs, Waiting: 654 reqs, GPU KV cache usage: 59.5%, Prefix cache hit rate: 93.3%
108
+ (APIServer pid=199309) INFO 07-14 21:57:17 [loggers.py:273] Engine 000: Avg prompt throughput: 643.1 tokens/s, Avg generation throughput: 2781.9 tokens/s, Running: 255 reqs, Waiting: 572 reqs, GPU KV cache usage: 58.6%, Prefix cache hit rate: 93.3%
109
+ (APIServer pid=199309) INFO 07-14 21:57:27 [loggers.py:273] Engine 000: Avg prompt throughput: 604.2 tokens/s, Avg generation throughput: 2731.0 tokens/s, Running: 256 reqs, Waiting: 498 reqs, GPU KV cache usage: 62.2%, Prefix cache hit rate: 93.3%
110
+ (APIServer pid=199309) INFO 07-14 21:57:37 [loggers.py:273] Engine 000: Avg prompt throughput: 780.9 tokens/s, Avg generation throughput: 2728.8 tokens/s, Running: 256 reqs, Waiting: 395 reqs, GPU KV cache usage: 50.9%, Prefix cache hit rate: 93.4%
111
+ (APIServer pid=199309) INFO 07-14 21:57:47 [loggers.py:273] Engine 000: Avg prompt throughput: 467.6 tokens/s, Avg generation throughput: 2835.7 tokens/s, Running: 255 reqs, Waiting: 339 reqs, GPU KV cache usage: 60.1%, Prefix cache hit rate: 93.3%
112
+ (APIServer pid=199309) INFO 07-14 21:57:57 [loggers.py:273] Engine 000: Avg prompt throughput: 614.5 tokens/s, Avg generation throughput: 2706.4 tokens/s, Running: 254 reqs, Waiting: 266 reqs, GPU KV cache usage: 61.5%, Prefix cache hit rate: 93.4%
113
+ (APIServer pid=199309) INFO 07-14 21:58:07 [loggers.py:273] Engine 000: Avg prompt throughput: 541.1 tokens/s, Avg generation throughput: 2783.3 tokens/s, Running: 255 reqs, Waiting: 198 reqs, GPU KV cache usage: 63.2%, Prefix cache hit rate: 93.4%
114
+ (APIServer pid=199309) INFO 07-14 21:58:17 [loggers.py:273] Engine 000: Avg prompt throughput: 587.5 tokens/s, Avg generation throughput: 2834.1 tokens/s, Running: 256 reqs, Waiting: 123 reqs, GPU KV cache usage: 62.1%, Prefix cache hit rate: 93.4%
115
+ (APIServer pid=199309) INFO 07-14 21:58:27 [loggers.py:273] Engine 000: Avg prompt throughput: 721.8 tokens/s, Avg generation throughput: 2704.7 tokens/s, Running: 255 reqs, Waiting: 35 reqs, GPU KV cache usage: 56.7%, Prefix cache hit rate: 93.4%
116
+ (APIServer pid=199309) INFO 07-14 21:58:37 [loggers.py:273] Engine 000: Avg prompt throughput: 285.7 tokens/s, Avg generation throughput: 2830.4 tokens/s, Running: 226 reqs, Waiting: 0 reqs, GPU KV cache usage: 58.2%, Prefix cache hit rate: 93.4%
117
+ (APIServer pid=199309) INFO 07-14 21:58:47 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 2559.2 tokens/s, Running: 143 reqs, Waiting: 0 reqs, GPU KV cache usage: 47.3%, Prefix cache hit rate: 93.4%
118
+ (APIServer pid=199309) INFO 07-14 21:58:57 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 1863.9 tokens/s, Running: 33 reqs, Waiting: 0 reqs, GPU KV cache usage: 14.7%, Prefix cache hit rate: 93.4%
119
+ (APIServer pid=199309) INFO: 127.0.0.1:60680 - "POST /v1/completions HTTP/1.1" 200 OK
120
+ (EngineCore pid=199426) INFO 07-14 21:59:01 [core.py:1214] [shutdown] EngineCore: trigger received signal=SIGTERM
121
+ (APIServer pid=199309) INFO 07-14 21:59:01 [launcher.py:100] [shutdown] API server: shutdown triggered
122
+ (APIServer pid=199309) INFO 07-14 21:59:01 [launcher.py:116] [shutdown] API server: stopping engine client mode=abort timeout=0s
123
+ (EngineCore pid=199426) INFO 07-14 21:59:01 [core.py:1333] [shutdown] EngineCore: start mode=abort timeout=0s
124
+ (EngineCore pid=199426) INFO 07-14 21:59:01 [core.py:1364] [shutdown] EngineCore: request processing complete; starting resource teardown
125
+ (EngineCore pid=199426) INFO 07-14 21:59:01 [core.py:1227] [shutdown] EngineCore: exiting busy loop
126
+ (APIServer pid=199309) INFO 07-14 21:59:01 [core_client.py:655] [shutdown] MPClient: start timeout=0s
127
+ (APIServer pid=199309) INFO 07-14 21:59:01 [core_client.py:657] [shutdown] MPClient: stopping engine manager
128
+ (APIServer pid=199309) WARNING 07-14 21:59:01 [utils.py:626] [shutdown] Process manager: force killing remaining processes count=1
129
+ (APIServer pid=199309) INFO 07-14 21:59:01 [core_client.py:659] [shutdown] MPClient: engine manager stopped
130
+ (APIServer pid=199309) INFO 07-14 21:59:01 [core_client.py:660] [shutdown] MPClient: cleaning up background resources
131
+ (APIServer pid=199309) INFO 07-14 21:59:01 [core_client.py:662] [shutdown] MPClient: complete
132
+ (APIServer pid=199309) INFO 07-14 21:59:01 [launcher.py:125] [shutdown] API server: engine client stopped
133
+ (APIServer pid=199309) INFO 07-14 21:59:01 [launcher.py:128] [shutdown] API server: signalling HTTP server shutdown
134
+ (APIServer pid=199309) INFO 07-14 21:59:01 [launcher.py:149] [shutdown] API server: shutting down FastAPI HTTP server
135
+ (APIServer pid=199309) INFO: Shutting down
136
+ (APIServer pid=199309) INFO: Waiting for application shutdown.
137
+ (APIServer pid=199309) INFO: Application shutdown complete.
138
+ /home/henry/.local/share/uv/python/cpython-3.12.12-linux-x86_64-gnu/lib/python3.12/multiprocessing/resource_tracker.py:279: UserWarning: resource_tracker: There appear to be 1 leaked semaphore objects to clean up at shutdown
139
+ warnings.warn('resource_tracker: There appear to be %d '
evals/healing_breadth/glean_math_keep25_seed1224_long768_step0500_chat.json ADDED
The diff for this file is too large to render. See raw diff
 
evals/healing_breadth/glean_math_keep25_seed1224_long768_step0500_chat.json.server.log ADDED
@@ -0,0 +1,255 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ (APIServer pid=304120) INFO 07-15 18:42:16 [api_utils.py:339]
2
+ (APIServer pid=304120) INFO 07-15 18:42:16 [api_utils.py:339] β–ˆ β–ˆ β–ˆβ–„ β–„β–ˆ
3
+ (APIServer pid=304120) INFO 07-15 18:42:16 [api_utils.py:339] β–„β–„ β–„β–ˆ β–ˆ β–ˆ β–ˆ β–€β–„β–€ β–ˆ version 0.25.0
4
+ (APIServer pid=304120) INFO 07-15 18:42:16 [api_utils.py:339] β–ˆβ–„β–ˆβ–€ β–ˆ β–ˆ β–ˆ β–ˆ model outputs/healed/healing_breadth/glean_math_keep25_seed1224_long768/step0500
5
+ (APIServer pid=304120) INFO 07-15 18:42:16 [api_utils.py:339] β–€β–€ β–€β–€β–€β–€β–€ β–€β–€β–€β–€β–€ β–€ β–€
6
+ (APIServer pid=304120) INFO 07-15 18:42:16 [api_utils.py:339]
7
+ (APIServer pid=304120) INFO 07-15 18:42:16 [api_utils.py:273] non-default args: {'model_tag': 'outputs/healed/healing_breadth/glean_math_keep25_seed1224_long768/step0500', 'host': '127.0.0.1', 'port': 8377, 'model': 'outputs/healed/healing_breadth/glean_math_keep25_seed1224_long768/step0500', 'max_model_len': 2048, 'enforce_eager': True, 'served_model_name': ['student'], 'gpu_memory_utilization': 0.85}
8
+ (APIServer pid=304120) INFO 07-15 18:42:16 [model.py:619] Resolved architecture: PrunedOlmoeForCausalLM
9
+ (APIServer pid=304120) INFO 07-15 18:42:16 [model.py:1776] Using max model len 2048
10
+ (APIServer pid=304120) INFO 07-15 18:42:16 [vllm.py:1042] Asynchronous scheduling is enabled.
11
+ (APIServer pid=304120) WARNING 07-15 18:42:16 [vllm.py:1096] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
12
+ (APIServer pid=304120) WARNING 07-15 18:42:16 [vllm.py:1144] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
13
+ (APIServer pid=304120) INFO 07-15 18:42:16 [kernel.py:292] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native'])
14
+ (APIServer pid=304120) INFO 07-15 18:42:16 [vllm.py:1322] Cudagraph is disabled under eager mode
15
+ (APIServer pid=304120) INFO 07-15 18:42:16 [compilation.py:312] Enabled custom fusions: norm_quant, act_quant
16
+ (EngineCore pid=304237) INFO 07-15 18:42:24 [core.py:114] Initializing a V1 LLM engine (v0.25.0) with config: model='outputs/healed/healing_breadth/glean_math_keep25_seed1224_long768/step0500', speculative_config=None, tokenizer='outputs/healed/healing_breadth/glean_math_keep25_seed1224_long768/step0500', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=False, dtype=torch.bfloat16, max_seq_len=2048, download_dir=None, load_format=auto, tensor_parallel_size=1, pipeline_parallel_size=1, data_parallel_size=1, decode_context_parallel_size=1, dcp_comm_backend=ag_rs, disable_custom_all_reduce=False, quantization=None, quantization_config=None, enforce_eager=True, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False, jit_monitor_mode='warn', jit_monitor_verbose=False), seed=0, served_model_name=student, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'mode': <CompilationMode.NONE: 0>, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'ir_enable_torch_wrap': False, 'splitting_ops': [], 'compile_mm_encoder': False, 'cudagraph_mm_encoder': False, 'encoder_cudagraph_token_budgets': [], 'encoder_cudagraph_max_vision_items_per_batch': 0, 'encoder_cudagraph_max_frames_per_batch': None, 'compile_sizes': [], 'compile_ranges_endpoints': [2048], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'size_asserts': False, 'alignment_asserts': False, 'scalar_asserts': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': <CUDAGraphMode.NONE: 0>, 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': True, 'fuse_act_quant': True, 'fuse_attn_quant': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False, 'fuse_rope_kvcache_cat_mla': False, 'fuse_act_padding': False}, 'max_cudagraph_capture_size': 0, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': False, 'static_all_moe_layers': []}, kernel_config=KernelConfig(ir_op_priority=IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native']), enable_flashinfer_autotune=True, enable_cutedsl_warmup=True, moe_backend='auto', linear_backend='auto')
17
+ (EngineCore pid=304237) INFO 07-15 18:42:24 [parallel_state.py:1607] world_size=1 rank=0 local_rank=0 distributed_init_method=tcp://192.168.0.15:38123 backend=nccl
18
+ (EngineCore pid=304237) INFO 07-15 18:42:24 [parallel_state.py:1942] rank 0 in world size 1 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank 0, EPLB rank N/A
19
+ (EngineCore pid=304237) INFO 07-15 18:42:25 [topk_topp_sampler.py:55] Using FlashInfer for top-p & top-k sampling.
20
+ (EngineCore pid=304237) INFO 07-15 18:42:25 [gpu_model_runner.py:5209] Starting to load model outputs/healed/healing_breadth/glean_math_keep25_seed1224_long768/step0500...
21
+ (EngineCore pid=304237) INFO 07-15 18:42:26 [cuda.py:476] Using FLASH_ATTN attention backend out of potential backends: ['FLASH_ATTN', 'FLASHINFER', 'TRITON_ATTN', 'FLEX_ATTENTION'].
22
+ (EngineCore pid=304237) INFO 07-15 18:42:26 [flash_attn.py:718] Using FlashAttention version 2
23
+ (EngineCore pid=304237) /home/henry/.cache/glean/megablocks-variable-93a1479bc15b/megablocks/grouped_gemm_util.py:10: UserWarning: Grouped GEMM not available.
24
+ (EngineCore pid=304237) warnings.warn('Grouped GEMM not available.')
25
+ (EngineCore pid=304237) INFO 07-15 18:42:26 [weight_utils.py:849] Filesystem type for checkpoints: EXT4. Checkpoint size: 3.89 GiB. Available RAM: 109.34 GiB.
26
+ (EngineCore pid=304237) INFO 07-15 18:42:26 [weight_utils.py:872] Auto-prefetch is disabled because the filesystem (EXT4) is not a recognized network FS (NFS/Lustre). If you want to force prefetching, start vLLM with --safetensors-load-strategy=prefetch.
27
+ (EngineCore pid=304237)
28
+ (EngineCore pid=304237)
29
+ (EngineCore pid=304237)
30
+ (EngineCore pid=304237)
31
+ (EngineCore pid=304237) INFO 07-15 18:42:28 [default_loader.py:430] Loading weights took 2.36 seconds
32
+ (EngineCore pid=304237) INFO 07-15 18:42:28 [gpu_model_runner.py:5306] Model loading took 3.89 GiB memory and 2.546971 seconds
33
+ (EngineCore pid=304237) INFO 07-15 18:42:30 [gpu_worker.py:538] Available KV cache memory: 15.82 GiB
34
+ (EngineCore pid=304237) INFO 07-15 18:42:30 [kv_cache_utils.py:2146] GPU KV cache size: 129,584 tokens
35
+ (EngineCore pid=304237) INFO 07-15 18:42:30 [kv_cache_utils.py:2147] Maximum concurrency for 2,048 tokens per request: 63.27x
36
+ (EngineCore pid=304237) INFO 07-15 18:42:30 [cutedsl_warmup.py:97] Skipping CuTeDSL warmup because no compile units were requested.
37
+ (EngineCore pid=304237) INFO 07-15 18:42:30 [jit_monitor.py:73] Kernel JIT monitor activated; monitored JIT compilations during inference will use mode=warn.
38
+ (EngineCore pid=304237) INFO 07-15 18:42:31 [core.py:344] init engine (profile, create kv cache, warmup model) took 2.09 s
39
+ (EngineCore pid=304237) INFO 07-15 18:42:31 [vllm.py:1042] Asynchronous scheduling is enabled.
40
+ (EngineCore pid=304237) WARNING 07-15 18:42:31 [vllm.py:1096] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
41
+ (EngineCore pid=304237) WARNING 07-15 18:42:31 [vllm.py:1144] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
42
+ (EngineCore pid=304237) INFO 07-15 18:42:31 [kernel.py:292] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native'])
43
+ (EngineCore pid=304237) INFO 07-15 18:42:31 [vllm.py:1322] Cudagraph is disabled under eager mode
44
+ (EngineCore pid=304237) INFO 07-15 18:42:31 [compilation.py:312] Enabled custom fusions: norm_quant, act_quant
45
+ (APIServer pid=304120) INFO 07-15 18:42:31 [api_server.py:612] Supported tasks: ['generate']
46
+ (APIServer pid=304120) WARNING 07-15 18:42:31 [__init__.py:36] SECURITY WARNING: Development endpoints are enabled! This should NOT be used in production!
47
+ (APIServer pid=304120) INFO 07-15 18:42:31 [hf.py:548] Detected the chat template content format to be 'string'. You can set `--chat-template-content-format` to override this.
48
+ (APIServer pid=304120) INFO 07-15 18:42:31 [api_server.py:616] Starting vLLM server on http://127.0.0.1:8377
49
+ (APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:37] Available routes are:
50
+ (APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /openapi.json, Methods: HEAD, GET
51
+ (APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /docs, Methods: HEAD, GET
52
+ (APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /docs/oauth2-redirect, Methods: HEAD, GET
53
+ (APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /redoc, Methods: HEAD, GET
54
+ (APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /load, Methods: GET
55
+ (APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /version, Methods: GET
56
+ (APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /health, Methods: GET
57
+ (APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /metrics, Methods: GET
58
+ (APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /tokenize, Methods: POST
59
+ (APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /detokenize, Methods: POST
60
+ (APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /v1/models, Methods: GET
61
+ (APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /ping, Methods: GET
62
+ (APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /ping, Methods: POST
63
+ (APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /invocations, Methods: POST
64
+ (APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /reset_prefix_cache, Methods: POST
65
+ (APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /reset_mm_cache, Methods: POST
66
+ (APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /reset_encoder_cache, Methods: POST
67
+ (APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /pause, Methods: POST
68
+ (APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /resume, Methods: POST
69
+ (APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /is_paused, Methods: GET
70
+ (APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /init_weight_transfer_engine, Methods: POST
71
+ (APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /start_weight_update, Methods: POST
72
+ (APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /update_weights, Methods: POST
73
+ (APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /finish_weight_update, Methods: POST
74
+ (APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /get_world_size, Methods: GET
75
+ (APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /collective_rpc, Methods: POST
76
+ (APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /server_info, Methods: GET
77
+ (APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /sleep, Methods: POST
78
+ (APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /wake_up, Methods: POST
79
+ (APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /is_sleeping, Methods: GET
80
+ (APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /v1/chat/completions, Methods: POST
81
+ (APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /v1/chat/completions/batch, Methods: POST
82
+ (APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /v1/responses, Methods: POST
83
+ (APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /v1/responses/{response_id}, Methods: GET
84
+ (APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /v1/responses/{response_id}/cancel, Methods: POST
85
+ (APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /v1/completions, Methods: POST
86
+ (APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /v1/messages, Methods: POST
87
+ (APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /v1/messages/count_tokens, Methods: POST
88
+ (APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /generative_scoring, Methods: POST
89
+ (APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /scale_elastic_ep, Methods: POST
90
+ (APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /is_scaling_elastic_ep, Methods: POST
91
+ (APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /v1/chat/completions/render, Methods: POST
92
+ (APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /v1/completions/render, Methods: POST
93
+ (APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /v1/chat/completions/derender, Methods: POST
94
+ (APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /v1/completions/derender, Methods: POST
95
+ (APIServer pid=304120) INFO 07-15 18:42:31 [launcher.py:46] Route: /inference/v1/generate, Methods: POST
96
+ (APIServer pid=304120) INFO: Started server process [304120]
97
+ (APIServer pid=304120) INFO: Waiting for application startup.
98
+ (APIServer pid=304120) INFO: Application startup complete.
99
+ (APIServer pid=304120) INFO: 127.0.0.1:52674 - "GET /health HTTP/1.1" 200 OK
100
+ (APIServer pid=304120) ERROR 07-15 18:42:33 [server_utils.py:384] Exception caught. Request id: None
101
+ (APIServer pid=304120) INFO: 127.0.0.1:52682 - "POST /v1/completions HTTP/1.1" 400 Bad Request
102
+ (APIServer pid=304120) INFO 07-15 18:42:33 [launcher.py:100] [shutdown] API server: shutdown triggered
103
+ (APIServer pid=304120) INFO 07-15 18:42:33 [launcher.py:116] [shutdown] API server: stopping engine client mode=abort timeout=0s
104
+ (EngineCore pid=304237) INFO 07-15 18:42:33 [core.py:1214] [shutdown] EngineCore: trigger received signal=SIGTERM
105
+ (EngineCore pid=304237) INFO 07-15 18:42:33 [core.py:1333] [shutdown] EngineCore: start mode=abort timeout=0s
106
+ (EngineCore pid=304237) INFO 07-15 18:42:33 [core.py:1364] [shutdown] EngineCore: request processing complete; starting resource teardown
107
+ (EngineCore pid=304237) INFO 07-15 18:42:33 [core.py:1227] [shutdown] EngineCore: exiting busy loop
108
+ (APIServer pid=304120) INFO 07-15 18:42:33 [core_client.py:655] [shutdown] MPClient: start timeout=0s
109
+ (APIServer pid=304120) INFO 07-15 18:42:33 [core_client.py:657] [shutdown] MPClient: stopping engine manager
110
+ (APIServer pid=304120) WARNING 07-15 18:42:33 [utils.py:626] [shutdown] Process manager: force killing remaining processes count=1
111
+ (APIServer pid=304120) INFO 07-15 18:42:33 [core_client.py:659] [shutdown] MPClient: engine manager stopped
112
+ (APIServer pid=304120) INFO 07-15 18:42:33 [core_client.py:660] [shutdown] MPClient: cleaning up background resources
113
+ (APIServer pid=304120) INFO 07-15 18:42:33 [core_client.py:662] [shutdown] MPClient: complete
114
+ (APIServer pid=304120) INFO 07-15 18:42:33 [launcher.py:125] [shutdown] API server: engine client stopped
115
+ (APIServer pid=304120) INFO 07-15 18:42:33 [launcher.py:128] [shutdown] API server: signalling HTTP server shutdown
116
+ (APIServer pid=304120) INFO 07-15 18:42:33 [launcher.py:149] [shutdown] API server: shutting down FastAPI HTTP server
117
+ (APIServer pid=304120) INFO: Shutting down
118
+ (APIServer pid=304120) INFO: Waiting for application shutdown.
119
+ (APIServer pid=304120) INFO: Application shutdown complete.
120
+ /home/henry/.local/share/uv/python/cpython-3.12.12-linux-x86_64-gnu/lib/python3.12/multiprocessing/resource_tracker.py:279: UserWarning: resource_tracker: There appear to be 1 leaked semaphore objects to clean up at shutdown
121
+ warnings.warn('resource_tracker: There appear to be %d '
122
+ (APIServer pid=306333) INFO 07-15 18:47:04 [api_utils.py:339]
123
+ (APIServer pid=306333) INFO 07-15 18:47:04 [api_utils.py:339] β–ˆ β–ˆ β–ˆβ–„ β–„β–ˆ
124
+ (APIServer pid=306333) INFO 07-15 18:47:04 [api_utils.py:339] β–„β–„ β–„β–ˆ β–ˆ β–ˆ β–ˆ β–€β–„β–€ β–ˆ version 0.25.0
125
+ (APIServer pid=306333) INFO 07-15 18:47:04 [api_utils.py:339] β–ˆβ–„β–ˆβ–€ β–ˆ β–ˆ β–ˆ β–ˆ model outputs/healed/healing_breadth/glean_math_keep25_seed1224_long768/step0500
126
+ (APIServer pid=306333) INFO 07-15 18:47:04 [api_utils.py:339] β–€β–€ β–€β–€β–€β–€β–€ β–€β–€β–€β–€β–€ β–€ β–€
127
+ (APIServer pid=306333) INFO 07-15 18:47:04 [api_utils.py:339]
128
+ (APIServer pid=306333) INFO 07-15 18:47:04 [api_utils.py:273] non-default args: {'model_tag': 'outputs/healed/healing_breadth/glean_math_keep25_seed1224_long768/step0500', 'host': '127.0.0.1', 'port': 8377, 'model': 'outputs/healed/healing_breadth/glean_math_keep25_seed1224_long768/step0500', 'max_model_len': 2048, 'enforce_eager': True, 'served_model_name': ['student'], 'gpu_memory_utilization': 0.85}
129
+ (APIServer pid=306333) INFO 07-15 18:47:04 [model.py:619] Resolved architecture: PrunedOlmoeForCausalLM
130
+ (APIServer pid=306333) INFO 07-15 18:47:04 [model.py:1776] Using max model len 2048
131
+ (APIServer pid=306333) INFO 07-15 18:47:04 [vllm.py:1042] Asynchronous scheduling is enabled.
132
+ (APIServer pid=306333) WARNING 07-15 18:47:04 [vllm.py:1096] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
133
+ (APIServer pid=306333) WARNING 07-15 18:47:04 [vllm.py:1144] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
134
+ (APIServer pid=306333) INFO 07-15 18:47:04 [kernel.py:292] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native'])
135
+ (APIServer pid=306333) INFO 07-15 18:47:04 [vllm.py:1322] Cudagraph is disabled under eager mode
136
+ (APIServer pid=306333) INFO 07-15 18:47:04 [compilation.py:312] Enabled custom fusions: norm_quant, act_quant
137
+ (EngineCore pid=306450) INFO 07-15 18:47:11 [core.py:114] Initializing a V1 LLM engine (v0.25.0) with config: model='outputs/healed/healing_breadth/glean_math_keep25_seed1224_long768/step0500', speculative_config=None, tokenizer='outputs/healed/healing_breadth/glean_math_keep25_seed1224_long768/step0500', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=False, dtype=torch.bfloat16, max_seq_len=2048, download_dir=None, load_format=auto, tensor_parallel_size=1, pipeline_parallel_size=1, data_parallel_size=1, decode_context_parallel_size=1, dcp_comm_backend=ag_rs, disable_custom_all_reduce=False, quantization=None, quantization_config=None, enforce_eager=True, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False, jit_monitor_mode='warn', jit_monitor_verbose=False), seed=0, served_model_name=student, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'mode': <CompilationMode.NONE: 0>, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'ir_enable_torch_wrap': False, 'splitting_ops': [], 'compile_mm_encoder': False, 'cudagraph_mm_encoder': False, 'encoder_cudagraph_token_budgets': [], 'encoder_cudagraph_max_vision_items_per_batch': 0, 'encoder_cudagraph_max_frames_per_batch': None, 'compile_sizes': [], 'compile_ranges_endpoints': [2048], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'size_asserts': False, 'alignment_asserts': False, 'scalar_asserts': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': <CUDAGraphMode.NONE: 0>, 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': True, 'fuse_act_quant': True, 'fuse_attn_quant': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False, 'fuse_rope_kvcache_cat_mla': False, 'fuse_act_padding': False}, 'max_cudagraph_capture_size': 0, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': False, 'static_all_moe_layers': []}, kernel_config=KernelConfig(ir_op_priority=IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native']), enable_flashinfer_autotune=True, enable_cutedsl_warmup=True, moe_backend='auto', linear_backend='auto')
138
+ (EngineCore pid=306450) INFO 07-15 18:47:12 [parallel_state.py:1607] world_size=1 rank=0 local_rank=0 distributed_init_method=tcp://192.168.0.15:44905 backend=nccl
139
+ (EngineCore pid=306450) INFO 07-15 18:47:12 [parallel_state.py:1942] rank 0 in world size 1 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank 0, EPLB rank N/A
140
+ (EngineCore pid=306450) INFO 07-15 18:47:13 [topk_topp_sampler.py:55] Using FlashInfer for top-p & top-k sampling.
141
+ (EngineCore pid=306450) INFO 07-15 18:47:13 [gpu_model_runner.py:5209] Starting to load model outputs/healed/healing_breadth/glean_math_keep25_seed1224_long768/step0500...
142
+ (EngineCore pid=306450) INFO 07-15 18:47:13 [cuda.py:476] Using FLASH_ATTN attention backend out of potential backends: ['FLASH_ATTN', 'FLASHINFER', 'TRITON_ATTN', 'FLEX_ATTENTION'].
143
+ (EngineCore pid=306450) INFO 07-15 18:47:13 [flash_attn.py:718] Using FlashAttention version 2
144
+ (EngineCore pid=306450) /home/henry/.cache/glean/megablocks-variable-93a1479bc15b/megablocks/grouped_gemm_util.py:10: UserWarning: Grouped GEMM not available.
145
+ (EngineCore pid=306450) warnings.warn('Grouped GEMM not available.')
146
+ (EngineCore pid=306450) INFO 07-15 18:47:13 [weight_utils.py:849] Filesystem type for checkpoints: EXT4. Checkpoint size: 3.89 GiB. Available RAM: 109.49 GiB.
147
+ (EngineCore pid=306450) INFO 07-15 18:47:13 [weight_utils.py:872] Auto-prefetch is disabled because the filesystem (EXT4) is not a recognized network FS (NFS/Lustre). If you want to force prefetching, start vLLM with --safetensors-load-strategy=prefetch.
148
+ (EngineCore pid=306450)
149
+ (EngineCore pid=306450)
150
+ (EngineCore pid=306450)
151
+ (EngineCore pid=306450)
152
+ (EngineCore pid=306450) INFO 07-15 18:47:16 [default_loader.py:430] Loading weights took 2.45 seconds
153
+ (EngineCore pid=306450) INFO 07-15 18:47:16 [gpu_model_runner.py:5306] Model loading took 3.89 GiB memory and 2.628739 seconds
154
+ (EngineCore pid=306450) INFO 07-15 18:47:18 [gpu_worker.py:538] Available KV cache memory: 15.82 GiB
155
+ (EngineCore pid=306450) INFO 07-15 18:47:18 [kv_cache_utils.py:2146] GPU KV cache size: 129,584 tokens
156
+ (EngineCore pid=306450) INFO 07-15 18:47:18 [kv_cache_utils.py:2147] Maximum concurrency for 2,048 tokens per request: 63.27x
157
+ (EngineCore pid=306450) INFO 07-15 18:47:18 [cutedsl_warmup.py:97] Skipping CuTeDSL warmup because no compile units were requested.
158
+ (EngineCore pid=306450) INFO 07-15 18:47:18 [jit_monitor.py:73] Kernel JIT monitor activated; monitored JIT compilations during inference will use mode=warn.
159
+ (EngineCore pid=306450) INFO 07-15 18:47:18 [core.py:344] init engine (profile, create kv cache, warmup model) took 2.07 s
160
+ (EngineCore pid=306450) INFO 07-15 18:47:19 [vllm.py:1042] Asynchronous scheduling is enabled.
161
+ (EngineCore pid=306450) WARNING 07-15 18:47:19 [vllm.py:1096] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
162
+ (EngineCore pid=306450) WARNING 07-15 18:47:19 [vllm.py:1144] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
163
+ (EngineCore pid=306450) INFO 07-15 18:47:19 [kernel.py:292] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native'])
164
+ (EngineCore pid=306450) INFO 07-15 18:47:19 [vllm.py:1322] Cudagraph is disabled under eager mode
165
+ (EngineCore pid=306450) INFO 07-15 18:47:19 [compilation.py:312] Enabled custom fusions: norm_quant, act_quant
166
+ (APIServer pid=306333) INFO 07-15 18:47:19 [api_server.py:612] Supported tasks: ['generate']
167
+ (APIServer pid=306333) WARNING 07-15 18:47:19 [__init__.py:36] SECURITY WARNING: Development endpoints are enabled! This should NOT be used in production!
168
+ (APIServer pid=306333) INFO 07-15 18:47:19 [hf.py:548] Detected the chat template content format to be 'string'. You can set `--chat-template-content-format` to override this.
169
+ (APIServer pid=306333) INFO 07-15 18:47:19 [api_server.py:616] Starting vLLM server on http://127.0.0.1:8377
170
+ (APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:37] Available routes are:
171
+ (APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /openapi.json, Methods: GET, HEAD
172
+ (APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /docs, Methods: GET, HEAD
173
+ (APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /docs/oauth2-redirect, Methods: GET, HEAD
174
+ (APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /redoc, Methods: GET, HEAD
175
+ (APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /load, Methods: GET
176
+ (APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /version, Methods: GET
177
+ (APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /health, Methods: GET
178
+ (APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /metrics, Methods: GET
179
+ (APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /tokenize, Methods: POST
180
+ (APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /detokenize, Methods: POST
181
+ (APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /v1/models, Methods: GET
182
+ (APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /ping, Methods: GET
183
+ (APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /ping, Methods: POST
184
+ (APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /invocations, Methods: POST
185
+ (APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /reset_prefix_cache, Methods: POST
186
+ (APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /reset_mm_cache, Methods: POST
187
+ (APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /reset_encoder_cache, Methods: POST
188
+ (APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /pause, Methods: POST
189
+ (APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /resume, Methods: POST
190
+ (APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /is_paused, Methods: GET
191
+ (APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /init_weight_transfer_engine, Methods: POST
192
+ (APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /start_weight_update, Methods: POST
193
+ (APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /update_weights, Methods: POST
194
+ (APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /finish_weight_update, Methods: POST
195
+ (APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /get_world_size, Methods: GET
196
+ (APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /collective_rpc, Methods: POST
197
+ (APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /server_info, Methods: GET
198
+ (APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /sleep, Methods: POST
199
+ (APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /wake_up, Methods: POST
200
+ (APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /is_sleeping, Methods: GET
201
+ (APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /v1/chat/completions, Methods: POST
202
+ (APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /v1/chat/completions/batch, Methods: POST
203
+ (APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /v1/responses, Methods: POST
204
+ (APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /v1/responses/{response_id}, Methods: GET
205
+ (APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /v1/responses/{response_id}/cancel, Methods: POST
206
+ (APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /v1/completions, Methods: POST
207
+ (APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /v1/messages, Methods: POST
208
+ (APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /v1/messages/count_tokens, Methods: POST
209
+ (APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /generative_scoring, Methods: POST
210
+ (APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /scale_elastic_ep, Methods: POST
211
+ (APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /is_scaling_elastic_ep, Methods: POST
212
+ (APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /v1/chat/completions/render, Methods: POST
213
+ (APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /v1/completions/render, Methods: POST
214
+ (APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /v1/chat/completions/derender, Methods: POST
215
+ (APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /v1/completions/derender, Methods: POST
216
+ (APIServer pid=306333) INFO 07-15 18:47:19 [launcher.py:46] Route: /inference/v1/generate, Methods: POST
217
+ (APIServer pid=306333) INFO: Started server process [306333]
218
+ (APIServer pid=306333) INFO: Waiting for application startup.
219
+ (APIServer pid=306333) INFO: Application startup complete.
220
+ (APIServer pid=306333) INFO: 127.0.0.1:55846 - "GET /health HTTP/1.1" 200 OK
221
+ (EngineCore pid=306450) WARNING 07-15 18:47:21 [jit_monitor.py:129] Triton kernel JIT compilation during inference: _build_route_rows. This causes a latency spike; consider extending warmup to cover this shape/config.
222
+ (APIServer pid=306333) INFO 07-15 18:47:29 [loggers.py:273] Engine 000: Avg prompt throughput: 2022.5 tokens/s, Avg generation throughput: 2244.1 tokens/s, Running: 256 reqs, Waiting: 1063 reqs, GPU KV cache usage: 36.3%, Prefix cache hit rate: 93.0%
223
+ (APIServer pid=306333) INFO 07-15 18:47:39 [loggers.py:273] Engine 000: Avg prompt throughput: 662.3 tokens/s, Avg generation throughput: 2934.7 tokens/s, Running: 255 reqs, Waiting: 976 reqs, GPU KV cache usage: 48.2%, Prefix cache hit rate: 93.2%
224
+ (APIServer pid=306333) INFO 07-15 18:47:49 [loggers.py:273] Engine 000: Avg prompt throughput: 1015.2 tokens/s, Avg generation throughput: 2879.5 tokens/s, Running: 256 reqs, Waiting: 850 reqs, GPU KV cache usage: 47.1%, Prefix cache hit rate: 93.2%
225
+ (APIServer pid=306333) INFO 07-15 18:47:59 [loggers.py:273] Engine 000: Avg prompt throughput: 600.1 tokens/s, Avg generation throughput: 2935.5 tokens/s, Running: 253 reqs, Waiting: 769 reqs, GPU KV cache usage: 52.1%, Prefix cache hit rate: 93.3%
226
+ (APIServer pid=306333) INFO 07-15 18:48:09 [loggers.py:273] Engine 000: Avg prompt throughput: 1022.8 tokens/s, Avg generation throughput: 2879.4 tokens/s, Running: 254 reqs, Waiting: 637 reqs, GPU KV cache usage: 44.7%, Prefix cache hit rate: 93.3%
227
+ (APIServer pid=306333) INFO 07-15 18:48:19 [loggers.py:273] Engine 000: Avg prompt throughput: 858.2 tokens/s, Avg generation throughput: 2932.9 tokens/s, Running: 254 reqs, Waiting: 530 reqs, GPU KV cache usage: 46.3%, Prefix cache hit rate: 93.3%
228
+ (APIServer pid=306333) INFO 07-15 18:48:29 [loggers.py:273] Engine 000: Avg prompt throughput: 884.5 tokens/s, Avg generation throughput: 2932.2 tokens/s, Running: 255 reqs, Waiting: 418 reqs, GPU KV cache usage: 47.5%, Prefix cache hit rate: 93.4%
229
+ (APIServer pid=306333) INFO 07-15 18:48:39 [loggers.py:273] Engine 000: Avg prompt throughput: 1016.5 tokens/s, Avg generation throughput: 2930.8 tokens/s, Running: 256 reqs, Waiting: 295 reqs, GPU KV cache usage: 46.8%, Prefix cache hit rate: 93.4%
230
+ (APIServer pid=306333) INFO 07-15 18:48:49 [loggers.py:273] Engine 000: Avg prompt throughput: 882.8 tokens/s, Avg generation throughput: 2932.5 tokens/s, Running: 255 reqs, Waiting: 184 reqs, GPU KV cache usage: 46.7%, Prefix cache hit rate: 93.4%
231
+ (APIServer pid=306333) INFO 07-15 18:48:59 [loggers.py:273] Engine 000: Avg prompt throughput: 813.3 tokens/s, Avg generation throughput: 2933.6 tokens/s, Running: 254 reqs, Waiting: 82 reqs, GPU KV cache usage: 49.3%, Prefix cache hit rate: 93.4%
232
+ (APIServer pid=306333) INFO 07-15 18:49:09 [loggers.py:273] Engine 000: Avg prompt throughput: 669.5 tokens/s, Avg generation throughput: 2900.8 tokens/s, Running: 231 reqs, Waiting: 0 reqs, GPU KV cache usage: 48.0%, Prefix cache hit rate: 93.4%
233
+ (APIServer pid=306333) INFO 07-15 18:49:19 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 2483.2 tokens/s, Running: 84 reqs, Waiting: 0 reqs, GPU KV cache usage: 25.9%, Prefix cache hit rate: 93.4%
234
+ (APIServer pid=306333) INFO: 127.0.0.1:55862 - "POST /v1/completions HTTP/1.1" 200 OK
235
+ (APIServer pid=306333) INFO 07-15 18:49:29 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 682.5 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 93.4%
236
+ (EngineCore pid=306450) INFO 07-15 18:49:29 [core.py:1214] [shutdown] EngineCore: trigger received signal=SIGTERM
237
+ (APIServer pid=306333) INFO 07-15 18:49:29 [launcher.py:100] [shutdown] API server: shutdown triggered
238
+ (APIServer pid=306333) INFO 07-15 18:49:29 [launcher.py:116] [shutdown] API server: stopping engine client mode=abort timeout=0s
239
+ (EngineCore pid=306450) INFO 07-15 18:49:29 [core.py:1333] [shutdown] EngineCore: start mode=abort timeout=0s
240
+ (EngineCore pid=306450) INFO 07-15 18:49:29 [core.py:1364] [shutdown] EngineCore: request processing complete; starting resource teardown
241
+ (EngineCore pid=306450) INFO 07-15 18:49:29 [core.py:1227] [shutdown] EngineCore: exiting busy loop
242
+ (APIServer pid=306333) INFO 07-15 18:49:29 [core_client.py:655] [shutdown] MPClient: start timeout=0s
243
+ (APIServer pid=306333) INFO 07-15 18:49:29 [core_client.py:657] [shutdown] MPClient: stopping engine manager
244
+ (APIServer pid=306333) WARNING 07-15 18:49:29 [utils.py:626] [shutdown] Process manager: force killing remaining processes count=1
245
+ (APIServer pid=306333) INFO 07-15 18:49:29 [core_client.py:659] [shutdown] MPClient: engine manager stopped
246
+ (APIServer pid=306333) INFO 07-15 18:49:29 [core_client.py:660] [shutdown] MPClient: cleaning up background resources
247
+ (APIServer pid=306333) INFO 07-15 18:49:29 [core_client.py:662] [shutdown] MPClient: complete
248
+ (APIServer pid=306333) INFO 07-15 18:49:29 [launcher.py:125] [shutdown] API server: engine client stopped
249
+ (APIServer pid=306333) INFO 07-15 18:49:29 [launcher.py:128] [shutdown] API server: signalling HTTP server shutdown
250
+ (APIServer pid=306333) INFO 07-15 18:49:29 [launcher.py:149] [shutdown] API server: shutting down FastAPI HTTP server
251
+ (APIServer pid=306333) INFO: Shutting down
252
+ (APIServer pid=306333) INFO: Waiting for application shutdown.
253
+ (APIServer pid=306333) INFO: Application shutdown complete.
254
+ /home/henry/.local/share/uv/python/cpython-3.12.12-linux-x86_64-gnu/lib/python3.12/multiprocessing/resource_tracker.py:279: UserWarning: resource_tracker: There appear to be 1 leaked semaphore objects to clean up at shutdown
255
+ warnings.warn('resource_tracker: There appear to be %d '
evals/healing_breadth/glean_math_keep25_seed1224_long768_step0500_raw.json ADDED
The diff for this file is too large to render. See raw diff
 
evals/healing_breadth/glean_math_keep25_seed1224_long768_step0500_raw.json.server.log ADDED
@@ -0,0 +1,377 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ (APIServer pid=302497) INFO 07-15 18:37:53 [api_utils.py:339]
2
+ (APIServer pid=302497) INFO 07-15 18:37:53 [api_utils.py:339] β–ˆ β–ˆ β–ˆβ–„ β–„β–ˆ
3
+ (APIServer pid=302497) INFO 07-15 18:37:53 [api_utils.py:339] β–„β–„ β–„β–ˆ β–ˆ β–ˆ β–ˆ β–€β–„β–€ β–ˆ version 0.25.0
4
+ (APIServer pid=302497) INFO 07-15 18:37:53 [api_utils.py:339] β–ˆβ–„β–ˆβ–€ β–ˆ β–ˆ β–ˆ β–ˆ model outputs/healed/healing_breadth/glean_math_keep25_seed1224_long768/step0500
5
+ (APIServer pid=302497) INFO 07-15 18:37:53 [api_utils.py:339] β–€β–€ β–€β–€β–€β–€β–€ β–€β–€β–€β–€β–€ β–€ β–€
6
+ (APIServer pid=302497) INFO 07-15 18:37:53 [api_utils.py:339]
7
+ (APIServer pid=302497) INFO 07-15 18:37:53 [api_utils.py:273] non-default args: {'model_tag': 'outputs/healed/healing_breadth/glean_math_keep25_seed1224_long768/step0500', 'host': '127.0.0.1', 'port': 8377, 'model': 'outputs/healed/healing_breadth/glean_math_keep25_seed1224_long768/step0500', 'max_model_len': 2048, 'enforce_eager': True, 'served_model_name': ['student'], 'gpu_memory_utilization': 0.85}
8
+ (APIServer pid=302497) INFO 07-15 18:37:53 [model.py:619] Resolved architecture: PrunedOlmoeForCausalLM
9
+ (APIServer pid=302497) INFO 07-15 18:37:53 [model.py:1776] Using max model len 2048
10
+ (APIServer pid=302497) INFO 07-15 18:37:53 [vllm.py:1042] Asynchronous scheduling is enabled.
11
+ (APIServer pid=302497) WARNING 07-15 18:37:53 [vllm.py:1096] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
12
+ (APIServer pid=302497) WARNING 07-15 18:37:53 [vllm.py:1144] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
13
+ (APIServer pid=302497) INFO 07-15 18:37:53 [kernel.py:292] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native'])
14
+ (APIServer pid=302497) INFO 07-15 18:37:53 [vllm.py:1322] Cudagraph is disabled under eager mode
15
+ (APIServer pid=302497) INFO 07-15 18:37:53 [compilation.py:312] Enabled custom fusions: norm_quant, act_quant
16
+ (EngineCore pid=302614) INFO 07-15 18:38:00 [core.py:114] Initializing a V1 LLM engine (v0.25.0) with config: model='outputs/healed/healing_breadth/glean_math_keep25_seed1224_long768/step0500', speculative_config=None, tokenizer='outputs/healed/healing_breadth/glean_math_keep25_seed1224_long768/step0500', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=False, dtype=torch.bfloat16, max_seq_len=2048, download_dir=None, load_format=auto, tensor_parallel_size=1, pipeline_parallel_size=1, data_parallel_size=1, decode_context_parallel_size=1, dcp_comm_backend=ag_rs, disable_custom_all_reduce=False, quantization=None, quantization_config=None, enforce_eager=True, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False, jit_monitor_mode='warn', jit_monitor_verbose=False), seed=0, served_model_name=student, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'mode': <CompilationMode.NONE: 0>, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'ir_enable_torch_wrap': False, 'splitting_ops': [], 'compile_mm_encoder': False, 'cudagraph_mm_encoder': False, 'encoder_cudagraph_token_budgets': [], 'encoder_cudagraph_max_vision_items_per_batch': 0, 'encoder_cudagraph_max_frames_per_batch': None, 'compile_sizes': [], 'compile_ranges_endpoints': [2048], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'size_asserts': False, 'alignment_asserts': False, 'scalar_asserts': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': <CUDAGraphMode.NONE: 0>, 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': True, 'fuse_act_quant': True, 'fuse_attn_quant': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False, 'fuse_rope_kvcache_cat_mla': False, 'fuse_act_padding': False}, 'max_cudagraph_capture_size': 0, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': False, 'static_all_moe_layers': []}, kernel_config=KernelConfig(ir_op_priority=IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native']), enable_flashinfer_autotune=True, enable_cutedsl_warmup=True, moe_backend='auto', linear_backend='auto')
17
+ (EngineCore pid=302614) INFO 07-15 18:38:01 [parallel_state.py:1607] world_size=1 rank=0 local_rank=0 distributed_init_method=tcp://192.168.0.15:34145 backend=nccl
18
+ (EngineCore pid=302614) INFO 07-15 18:38:01 [parallel_state.py:1942] rank 0 in world size 1 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank 0, EPLB rank N/A
19
+ (EngineCore pid=302614) INFO 07-15 18:38:02 [topk_topp_sampler.py:55] Using FlashInfer for top-p & top-k sampling.
20
+ (EngineCore pid=302614) INFO 07-15 18:38:02 [gpu_model_runner.py:5209] Starting to load model outputs/healed/healing_breadth/glean_math_keep25_seed1224_long768/step0500...
21
+ (EngineCore pid=302614) INFO 07-15 18:38:02 [cuda.py:476] Using FLASH_ATTN attention backend out of potential backends: ['FLASH_ATTN', 'FLASHINFER', 'TRITON_ATTN', 'FLEX_ATTENTION'].
22
+ (EngineCore pid=302614) INFO 07-15 18:38:02 [flash_attn.py:718] Using FlashAttention version 2
23
+ (EngineCore pid=302614) /home/henry/.cache/glean/megablocks-variable-93a1479bc15b/megablocks/grouped_gemm_util.py:10: UserWarning: Grouped GEMM not available.
24
+ (EngineCore pid=302614) warnings.warn('Grouped GEMM not available.')
25
+ (EngineCore pid=302614) INFO 07-15 18:38:02 [weight_utils.py:849] Filesystem type for checkpoints: EXT4. Checkpoint size: 3.89 GiB. Available RAM: 109.79 GiB.
26
+ (EngineCore pid=302614) INFO 07-15 18:38:02 [weight_utils.py:872] Auto-prefetch is disabled because the filesystem (EXT4) is not a recognized network FS (NFS/Lustre). If you want to force prefetching, start vLLM with --safetensors-load-strategy=prefetch.
27
+ (EngineCore pid=302614)
28
+ (EngineCore pid=302614)
29
+ (EngineCore pid=302614)
30
+ (EngineCore pid=302614)
31
+ (EngineCore pid=302614) INFO 07-15 18:38:05 [default_loader.py:430] Loading weights took 2.50 seconds
32
+ (EngineCore pid=302614) INFO 07-15 18:38:05 [gpu_model_runner.py:5306] Model loading took 3.89 GiB memory and 2.681781 seconds
33
+ (EngineCore pid=302614) INFO 07-15 18:38:07 [gpu_worker.py:538] Available KV cache memory: 15.82 GiB
34
+ (EngineCore pid=302614) INFO 07-15 18:38:07 [kv_cache_utils.py:2146] GPU KV cache size: 129,584 tokens
35
+ (EngineCore pid=302614) INFO 07-15 18:38:07 [kv_cache_utils.py:2147] Maximum concurrency for 2,048 tokens per request: 63.27x
36
+ (EngineCore pid=302614) INFO 07-15 18:38:07 [cutedsl_warmup.py:97] Skipping CuTeDSL warmup because no compile units were requested.
37
+ (EngineCore pid=302614) INFO 07-15 18:38:07 [jit_monitor.py:73] Kernel JIT monitor activated; monitored JIT compilations during inference will use mode=warn.
38
+ (EngineCore pid=302614) INFO 07-15 18:38:07 [core.py:344] init engine (profile, create kv cache, warmup model) took 2.10 s
39
+ (EngineCore pid=302614) INFO 07-15 18:38:08 [vllm.py:1042] Asynchronous scheduling is enabled.
40
+ (EngineCore pid=302614) WARNING 07-15 18:38:08 [vllm.py:1096] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
41
+ (EngineCore pid=302614) WARNING 07-15 18:38:08 [vllm.py:1144] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
42
+ (EngineCore pid=302614) INFO 07-15 18:38:08 [kernel.py:292] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native'])
43
+ (EngineCore pid=302614) INFO 07-15 18:38:08 [vllm.py:1322] Cudagraph is disabled under eager mode
44
+ (EngineCore pid=302614) INFO 07-15 18:38:08 [compilation.py:312] Enabled custom fusions: norm_quant, act_quant
45
+ (APIServer pid=302497) INFO 07-15 18:38:08 [api_server.py:612] Supported tasks: ['generate']
46
+ (APIServer pid=302497) WARNING 07-15 18:38:08 [__init__.py:36] SECURITY WARNING: Development endpoints are enabled! This should NOT be used in production!
47
+ (APIServer pid=302497) INFO 07-15 18:38:08 [hf.py:548] Detected the chat template content format to be 'string'. You can set `--chat-template-content-format` to override this.
48
+ (APIServer pid=302497) INFO 07-15 18:38:08 [api_server.py:616] Starting vLLM server on http://127.0.0.1:8377
49
+ (APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:37] Available routes are:
50
+ (APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /openapi.json, Methods: GET, HEAD
51
+ (APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /docs, Methods: GET, HEAD
52
+ (APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /docs/oauth2-redirect, Methods: GET, HEAD
53
+ (APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /redoc, Methods: GET, HEAD
54
+ (APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /load, Methods: GET
55
+ (APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /version, Methods: GET
56
+ (APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /health, Methods: GET
57
+ (APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /metrics, Methods: GET
58
+ (APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /tokenize, Methods: POST
59
+ (APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /detokenize, Methods: POST
60
+ (APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /v1/models, Methods: GET
61
+ (APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /ping, Methods: GET
62
+ (APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /ping, Methods: POST
63
+ (APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /invocations, Methods: POST
64
+ (APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /reset_prefix_cache, Methods: POST
65
+ (APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /reset_mm_cache, Methods: POST
66
+ (APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /reset_encoder_cache, Methods: POST
67
+ (APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /pause, Methods: POST
68
+ (APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /resume, Methods: POST
69
+ (APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /is_paused, Methods: GET
70
+ (APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /init_weight_transfer_engine, Methods: POST
71
+ (APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /start_weight_update, Methods: POST
72
+ (APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /update_weights, Methods: POST
73
+ (APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /finish_weight_update, Methods: POST
74
+ (APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /get_world_size, Methods: GET
75
+ (APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /collective_rpc, Methods: POST
76
+ (APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /server_info, Methods: GET
77
+ (APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /sleep, Methods: POST
78
+ (APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /wake_up, Methods: POST
79
+ (APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /is_sleeping, Methods: GET
80
+ (APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /v1/chat/completions, Methods: POST
81
+ (APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /v1/chat/completions/batch, Methods: POST
82
+ (APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /v1/responses, Methods: POST
83
+ (APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /v1/responses/{response_id}, Methods: GET
84
+ (APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /v1/responses/{response_id}/cancel, Methods: POST
85
+ (APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /v1/completions, Methods: POST
86
+ (APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /v1/messages, Methods: POST
87
+ (APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /v1/messages/count_tokens, Methods: POST
88
+ (APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /generative_scoring, Methods: POST
89
+ (APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /scale_elastic_ep, Methods: POST
90
+ (APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /is_scaling_elastic_ep, Methods: POST
91
+ (APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /v1/chat/completions/render, Methods: POST
92
+ (APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /v1/completions/render, Methods: POST
93
+ (APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /v1/chat/completions/derender, Methods: POST
94
+ (APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /v1/completions/derender, Methods: POST
95
+ (APIServer pid=302497) INFO 07-15 18:38:08 [launcher.py:46] Route: /inference/v1/generate, Methods: POST
96
+ (APIServer pid=302497) INFO: Started server process [302497]
97
+ (APIServer pid=302497) INFO: Waiting for application startup.
98
+ (APIServer pid=302497) INFO: Application startup complete.
99
+ (APIServer pid=302497) INFO: 127.0.0.1:40196 - "GET /health HTTP/1.1" 200 OK
100
+ (APIServer pid=302497) ERROR 07-15 18:38:09 [server_utils.py:384] Exception caught. Request id: None
101
+ (APIServer pid=302497) INFO: 127.0.0.1:40210 - "POST /v1/completions HTTP/1.1" 400 Bad Request
102
+ (APIServer pid=302497) INFO 07-15 18:38:09 [launcher.py:100] [shutdown] API server: shutdown triggered
103
+ (APIServer pid=302497) INFO 07-15 18:38:09 [launcher.py:116] [shutdown] API server: stopping engine client mode=abort timeout=0s
104
+ (EngineCore pid=302614) INFO 07-15 18:38:09 [core.py:1214] [shutdown] EngineCore: trigger received signal=SIGTERM
105
+ (EngineCore pid=302614) INFO 07-15 18:38:09 [core.py:1333] [shutdown] EngineCore: start mode=abort timeout=0s
106
+ (EngineCore pid=302614) INFO 07-15 18:38:09 [core.py:1364] [shutdown] EngineCore: request processing complete; starting resource teardown
107
+ (EngineCore pid=302614) INFO 07-15 18:38:09 [core.py:1227] [shutdown] EngineCore: exiting busy loop
108
+ (APIServer pid=302497) INFO 07-15 18:38:09 [core_client.py:655] [shutdown] MPClient: start timeout=0s
109
+ (APIServer pid=302497) INFO 07-15 18:38:09 [core_client.py:657] [shutdown] MPClient: stopping engine manager
110
+ (APIServer pid=302497) WARNING 07-15 18:38:09 [utils.py:626] [shutdown] Process manager: force killing remaining processes count=1
111
+ (APIServer pid=302497) INFO 07-15 18:38:09 [core_client.py:659] [shutdown] MPClient: engine manager stopped
112
+ (APIServer pid=302497) INFO 07-15 18:38:09 [core_client.py:660] [shutdown] MPClient: cleaning up background resources
113
+ (APIServer pid=302497) INFO 07-15 18:38:09 [core_client.py:662] [shutdown] MPClient: complete
114
+ (APIServer pid=302497) INFO 07-15 18:38:09 [launcher.py:125] [shutdown] API server: engine client stopped
115
+ (APIServer pid=302497) INFO 07-15 18:38:09 [launcher.py:128] [shutdown] API server: signalling HTTP server shutdown
116
+ (APIServer pid=302497) INFO 07-15 18:38:09 [launcher.py:149] [shutdown] API server: shutting down FastAPI HTTP server
117
+ (APIServer pid=302497) INFO: Shutting down
118
+ (APIServer pid=302497) INFO: Waiting for application shutdown.
119
+ (APIServer pid=302497) INFO: Application shutdown complete.
120
+ /home/henry/.local/share/uv/python/cpython-3.12.12-linux-x86_64-gnu/lib/python3.12/multiprocessing/resource_tracker.py:279: UserWarning: resource_tracker: There appear to be 1 leaked semaphore objects to clean up at shutdown
121
+ warnings.warn('resource_tracker: There appear to be %d '
122
+ (APIServer pid=303449) INFO 07-15 18:41:35 [api_utils.py:339]
123
+ (APIServer pid=303449) INFO 07-15 18:41:35 [api_utils.py:339] β–ˆ β–ˆ β–ˆβ–„ β–„β–ˆ
124
+ (APIServer pid=303449) INFO 07-15 18:41:35 [api_utils.py:339] β–„β–„ β–„β–ˆ β–ˆ β–ˆ β–ˆ β–€β–„β–€ β–ˆ version 0.25.0
125
+ (APIServer pid=303449) INFO 07-15 18:41:35 [api_utils.py:339] β–ˆβ–„β–ˆβ–€ β–ˆ β–ˆ β–ˆ β–ˆ model outputs/healed/healing_breadth/glean_math_keep25_seed1224_long768/step0500
126
+ (APIServer pid=303449) INFO 07-15 18:41:35 [api_utils.py:339] β–€β–€ β–€β–€β–€β–€β–€ β–€β–€β–€β–€β–€ β–€ β–€
127
+ (APIServer pid=303449) INFO 07-15 18:41:35 [api_utils.py:339]
128
+ (APIServer pid=303449) INFO 07-15 18:41:35 [api_utils.py:273] non-default args: {'model_tag': 'outputs/healed/healing_breadth/glean_math_keep25_seed1224_long768/step0500', 'host': '127.0.0.1', 'port': 8377, 'model': 'outputs/healed/healing_breadth/glean_math_keep25_seed1224_long768/step0500', 'max_model_len': 2048, 'enforce_eager': True, 'served_model_name': ['student'], 'gpu_memory_utilization': 0.85}
129
+ (APIServer pid=303449) INFO 07-15 18:41:35 [model.py:619] Resolved architecture: PrunedOlmoeForCausalLM
130
+ (APIServer pid=303449) INFO 07-15 18:41:35 [model.py:1776] Using max model len 2048
131
+ (APIServer pid=303449) INFO 07-15 18:41:35 [vllm.py:1042] Asynchronous scheduling is enabled.
132
+ (APIServer pid=303449) WARNING 07-15 18:41:35 [vllm.py:1096] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
133
+ (APIServer pid=303449) WARNING 07-15 18:41:35 [vllm.py:1144] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
134
+ (APIServer pid=303449) INFO 07-15 18:41:35 [kernel.py:292] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native'])
135
+ (APIServer pid=303449) INFO 07-15 18:41:36 [vllm.py:1322] Cudagraph is disabled under eager mode
136
+ (APIServer pid=303449) INFO 07-15 18:41:36 [compilation.py:312] Enabled custom fusions: norm_quant, act_quant
137
+ (EngineCore pid=303564) INFO 07-15 18:41:43 [core.py:114] Initializing a V1 LLM engine (v0.25.0) with config: model='outputs/healed/healing_breadth/glean_math_keep25_seed1224_long768/step0500', speculative_config=None, tokenizer='outputs/healed/healing_breadth/glean_math_keep25_seed1224_long768/step0500', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=False, dtype=torch.bfloat16, max_seq_len=2048, download_dir=None, load_format=auto, tensor_parallel_size=1, pipeline_parallel_size=1, data_parallel_size=1, decode_context_parallel_size=1, dcp_comm_backend=ag_rs, disable_custom_all_reduce=False, quantization=None, quantization_config=None, enforce_eager=True, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False, jit_monitor_mode='warn', jit_monitor_verbose=False), seed=0, served_model_name=student, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'mode': <CompilationMode.NONE: 0>, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'ir_enable_torch_wrap': False, 'splitting_ops': [], 'compile_mm_encoder': False, 'cudagraph_mm_encoder': False, 'encoder_cudagraph_token_budgets': [], 'encoder_cudagraph_max_vision_items_per_batch': 0, 'encoder_cudagraph_max_frames_per_batch': None, 'compile_sizes': [], 'compile_ranges_endpoints': [2048], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'size_asserts': False, 'alignment_asserts': False, 'scalar_asserts': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': <CUDAGraphMode.NONE: 0>, 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': True, 'fuse_act_quant': True, 'fuse_attn_quant': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False, 'fuse_rope_kvcache_cat_mla': False, 'fuse_act_padding': False}, 'max_cudagraph_capture_size': 0, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': False, 'static_all_moe_layers': []}, kernel_config=KernelConfig(ir_op_priority=IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native']), enable_flashinfer_autotune=True, enable_cutedsl_warmup=True, moe_backend='auto', linear_backend='auto')
138
+ (EngineCore pid=303564) INFO 07-15 18:41:43 [parallel_state.py:1607] world_size=1 rank=0 local_rank=0 distributed_init_method=tcp://192.168.0.15:45641 backend=nccl
139
+ (EngineCore pid=303564) INFO 07-15 18:41:43 [parallel_state.py:1942] rank 0 in world size 1 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank 0, EPLB rank N/A
140
+ (EngineCore pid=303564) INFO 07-15 18:41:44 [topk_topp_sampler.py:55] Using FlashInfer for top-p & top-k sampling.
141
+ (EngineCore pid=303564) INFO 07-15 18:41:44 [gpu_model_runner.py:5209] Starting to load model outputs/healed/healing_breadth/glean_math_keep25_seed1224_long768/step0500...
142
+ (EngineCore pid=303564) INFO 07-15 18:41:45 [cuda.py:476] Using FLASH_ATTN attention backend out of potential backends: ['FLASH_ATTN', 'FLASHINFER', 'TRITON_ATTN', 'FLEX_ATTENTION'].
143
+ (EngineCore pid=303564) INFO 07-15 18:41:45 [flash_attn.py:718] Using FlashAttention version 2
144
+ (EngineCore pid=303564) /home/henry/.cache/glean/megablocks-variable-93a1479bc15b/megablocks/grouped_gemm_util.py:10: UserWarning: Grouped GEMM not available.
145
+ (EngineCore pid=303564) warnings.warn('Grouped GEMM not available.')
146
+ (EngineCore pid=303564) INFO 07-15 18:41:45 [weight_utils.py:849] Filesystem type for checkpoints: EXT4. Checkpoint size: 3.89 GiB. Available RAM: 109.58 GiB.
147
+ (EngineCore pid=303564) INFO 07-15 18:41:45 [weight_utils.py:872] Auto-prefetch is disabled because the filesystem (EXT4) is not a recognized network FS (NFS/Lustre). If you want to force prefetching, start vLLM with --safetensors-load-strategy=prefetch.
148
+ (EngineCore pid=303564)
149
+ (EngineCore pid=303564)
150
+ (EngineCore pid=303564)
151
+ (EngineCore pid=303564)
152
+ (EngineCore pid=303564) INFO 07-15 18:41:47 [default_loader.py:430] Loading weights took 2.36 seconds
153
+ (EngineCore pid=303564) INFO 07-15 18:41:47 [gpu_model_runner.py:5306] Model loading took 3.89 GiB memory and 2.540095 seconds
154
+ (EngineCore pid=303564) INFO 07-15 18:41:49 [gpu_worker.py:538] Available KV cache memory: 15.82 GiB
155
+ (EngineCore pid=303564) INFO 07-15 18:41:49 [kv_cache_utils.py:2146] GPU KV cache size: 129,584 tokens
156
+ (EngineCore pid=303564) INFO 07-15 18:41:49 [kv_cache_utils.py:2147] Maximum concurrency for 2,048 tokens per request: 63.27x
157
+ (EngineCore pid=303564) INFO 07-15 18:41:49 [cutedsl_warmup.py:97] Skipping CuTeDSL warmup because no compile units were requested.
158
+ (EngineCore pid=303564) INFO 07-15 18:41:49 [jit_monitor.py:73] Kernel JIT monitor activated; monitored JIT compilations during inference will use mode=warn.
159
+ (EngineCore pid=303564) INFO 07-15 18:41:50 [core.py:344] init engine (profile, create kv cache, warmup model) took 2.08 s
160
+ (EngineCore pid=303564) INFO 07-15 18:41:50 [vllm.py:1042] Asynchronous scheduling is enabled.
161
+ (EngineCore pid=303564) WARNING 07-15 18:41:50 [vllm.py:1096] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
162
+ (EngineCore pid=303564) WARNING 07-15 18:41:50 [vllm.py:1144] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
163
+ (EngineCore pid=303564) INFO 07-15 18:41:50 [kernel.py:292] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native'])
164
+ (EngineCore pid=303564) INFO 07-15 18:41:50 [vllm.py:1322] Cudagraph is disabled under eager mode
165
+ (EngineCore pid=303564) INFO 07-15 18:41:50 [compilation.py:312] Enabled custom fusions: norm_quant, act_quant
166
+ (APIServer pid=303449) INFO 07-15 18:41:50 [api_server.py:612] Supported tasks: ['generate']
167
+ (APIServer pid=303449) WARNING 07-15 18:41:50 [__init__.py:36] SECURITY WARNING: Development endpoints are enabled! This should NOT be used in production!
168
+ (APIServer pid=303449) INFO 07-15 18:41:50 [hf.py:548] Detected the chat template content format to be 'string'. You can set `--chat-template-content-format` to override this.
169
+ (APIServer pid=303449) INFO 07-15 18:41:50 [api_server.py:616] Starting vLLM server on http://127.0.0.1:8377
170
+ (APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:37] Available routes are:
171
+ (APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /openapi.json, Methods: GET, HEAD
172
+ (APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /docs, Methods: GET, HEAD
173
+ (APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /docs/oauth2-redirect, Methods: GET, HEAD
174
+ (APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /redoc, Methods: GET, HEAD
175
+ (APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /load, Methods: GET
176
+ (APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /version, Methods: GET
177
+ (APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /health, Methods: GET
178
+ (APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /metrics, Methods: GET
179
+ (APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /tokenize, Methods: POST
180
+ (APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /detokenize, Methods: POST
181
+ (APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /v1/models, Methods: GET
182
+ (APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /ping, Methods: GET
183
+ (APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /ping, Methods: POST
184
+ (APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /invocations, Methods: POST
185
+ (APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /reset_prefix_cache, Methods: POST
186
+ (APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /reset_mm_cache, Methods: POST
187
+ (APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /reset_encoder_cache, Methods: POST
188
+ (APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /pause, Methods: POST
189
+ (APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /resume, Methods: POST
190
+ (APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /is_paused, Methods: GET
191
+ (APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /init_weight_transfer_engine, Methods: POST
192
+ (APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /start_weight_update, Methods: POST
193
+ (APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /update_weights, Methods: POST
194
+ (APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /finish_weight_update, Methods: POST
195
+ (APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /get_world_size, Methods: GET
196
+ (APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /collective_rpc, Methods: POST
197
+ (APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /server_info, Methods: GET
198
+ (APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /sleep, Methods: POST
199
+ (APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /wake_up, Methods: POST
200
+ (APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /is_sleeping, Methods: GET
201
+ (APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /v1/chat/completions, Methods: POST
202
+ (APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /v1/chat/completions/batch, Methods: POST
203
+ (APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /v1/responses, Methods: POST
204
+ (APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /v1/responses/{response_id}, Methods: GET
205
+ (APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /v1/responses/{response_id}/cancel, Methods: POST
206
+ (APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /v1/completions, Methods: POST
207
+ (APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /v1/messages, Methods: POST
208
+ (APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /v1/messages/count_tokens, Methods: POST
209
+ (APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /generative_scoring, Methods: POST
210
+ (APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /scale_elastic_ep, Methods: POST
211
+ (APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /is_scaling_elastic_ep, Methods: POST
212
+ (APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /v1/chat/completions/render, Methods: POST
213
+ (APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /v1/completions/render, Methods: POST
214
+ (APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /v1/chat/completions/derender, Methods: POST
215
+ (APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /v1/completions/derender, Methods: POST
216
+ (APIServer pid=303449) INFO 07-15 18:41:50 [launcher.py:46] Route: /inference/v1/generate, Methods: POST
217
+ (APIServer pid=303449) INFO: Started server process [303449]
218
+ (APIServer pid=303449) INFO: Waiting for application startup.
219
+ (APIServer pid=303449) INFO: Application startup complete.
220
+ (APIServer pid=303449) INFO: 127.0.0.1:41890 - "GET /health HTTP/1.1" 200 OK
221
+ (APIServer pid=303449) ERROR 07-15 18:41:52 [server_utils.py:384] Exception caught. Request id: None
222
+ (APIServer pid=303449) INFO: 127.0.0.1:41906 - "POST /v1/completions HTTP/1.1" 400 Bad Request
223
+ (APIServer pid=303449) INFO 07-15 18:41:52 [launcher.py:100] [shutdown] API server: shutdown triggered
224
+ (APIServer pid=303449) INFO 07-15 18:41:52 [launcher.py:116] [shutdown] API server: stopping engine client mode=abort timeout=0s
225
+ (EngineCore pid=303564) INFO 07-15 18:41:52 [core.py:1214] [shutdown] EngineCore: trigger received signal=SIGTERM
226
+ (EngineCore pid=303564) INFO 07-15 18:41:52 [core.py:1333] [shutdown] EngineCore: start mode=abort timeout=0s
227
+ (EngineCore pid=303564) INFO 07-15 18:41:52 [core.py:1364] [shutdown] EngineCore: request processing complete; starting resource teardown
228
+ (EngineCore pid=303564) INFO 07-15 18:41:52 [core.py:1227] [shutdown] EngineCore: exiting busy loop
229
+ (APIServer pid=303449) INFO 07-15 18:41:52 [core_client.py:655] [shutdown] MPClient: start timeout=0s
230
+ (APIServer pid=303449) INFO 07-15 18:41:52 [core_client.py:657] [shutdown] MPClient: stopping engine manager
231
+ (APIServer pid=303449) WARNING 07-15 18:41:52 [utils.py:626] [shutdown] Process manager: force killing remaining processes count=1
232
+ (APIServer pid=303449) INFO 07-15 18:41:52 [core_client.py:659] [shutdown] MPClient: engine manager stopped
233
+ (APIServer pid=303449) INFO 07-15 18:41:52 [core_client.py:660] [shutdown] MPClient: cleaning up background resources
234
+ (APIServer pid=303449) INFO 07-15 18:41:52 [core_client.py:662] [shutdown] MPClient: complete
235
+ (APIServer pid=303449) INFO 07-15 18:41:52 [launcher.py:125] [shutdown] API server: engine client stopped
236
+ (APIServer pid=303449) INFO 07-15 18:41:52 [launcher.py:128] [shutdown] API server: signalling HTTP server shutdown
237
+ (APIServer pid=303449) INFO 07-15 18:41:52 [launcher.py:149] [shutdown] API server: shutting down FastAPI HTTP server
238
+ (APIServer pid=303449) INFO: Shutting down
239
+ (APIServer pid=303449) INFO: Waiting for application shutdown.
240
+ (APIServer pid=303449) INFO: Application shutdown complete.
241
+ /home/henry/.local/share/uv/python/cpython-3.12.12-linux-x86_64-gnu/lib/python3.12/multiprocessing/resource_tracker.py:279: UserWarning: resource_tracker: There appear to be 1 leaked semaphore objects to clean up at shutdown
242
+ warnings.warn('resource_tracker: There appear to be %d '
243
+ (APIServer pid=305508) INFO 07-15 18:43:59 [api_utils.py:339]
244
+ (APIServer pid=305508) INFO 07-15 18:43:59 [api_utils.py:339] β–ˆ β–ˆ β–ˆβ–„ β–„β–ˆ
245
+ (APIServer pid=305508) INFO 07-15 18:43:59 [api_utils.py:339] β–„β–„ β–„β–ˆ β–ˆ β–ˆ β–ˆ β–€β–„β–€ β–ˆ version 0.25.0
246
+ (APIServer pid=305508) INFO 07-15 18:43:59 [api_utils.py:339] β–ˆβ–„β–ˆβ–€ β–ˆ β–ˆ β–ˆ β–ˆ model outputs/healed/healing_breadth/glean_math_keep25_seed1224_long768/step0500
247
+ (APIServer pid=305508) INFO 07-15 18:43:59 [api_utils.py:339] β–€β–€ β–€β–€β–€β–€β–€ β–€β–€β–€β–€β–€ β–€ β–€
248
+ (APIServer pid=305508) INFO 07-15 18:43:59 [api_utils.py:339]
249
+ (APIServer pid=305508) INFO 07-15 18:43:59 [api_utils.py:273] non-default args: {'model_tag': 'outputs/healed/healing_breadth/glean_math_keep25_seed1224_long768/step0500', 'host': '127.0.0.1', 'port': 8377, 'model': 'outputs/healed/healing_breadth/glean_math_keep25_seed1224_long768/step0500', 'max_model_len': 2048, 'enforce_eager': True, 'served_model_name': ['student'], 'gpu_memory_utilization': 0.85}
250
+ (APIServer pid=305508) INFO 07-15 18:43:59 [model.py:619] Resolved architecture: PrunedOlmoeForCausalLM
251
+ (APIServer pid=305508) INFO 07-15 18:43:59 [model.py:1776] Using max model len 2048
252
+ (APIServer pid=305508) INFO 07-15 18:43:59 [vllm.py:1042] Asynchronous scheduling is enabled.
253
+ (APIServer pid=305508) WARNING 07-15 18:43:59 [vllm.py:1096] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
254
+ (APIServer pid=305508) WARNING 07-15 18:43:59 [vllm.py:1144] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
255
+ (APIServer pid=305508) INFO 07-15 18:43:59 [kernel.py:292] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native'])
256
+ (APIServer pid=305508) INFO 07-15 18:43:59 [vllm.py:1322] Cudagraph is disabled under eager mode
257
+ (APIServer pid=305508) INFO 07-15 18:43:59 [compilation.py:312] Enabled custom fusions: norm_quant, act_quant
258
+ (EngineCore pid=305625) INFO 07-15 18:44:06 [core.py:114] Initializing a V1 LLM engine (v0.25.0) with config: model='outputs/healed/healing_breadth/glean_math_keep25_seed1224_long768/step0500', speculative_config=None, tokenizer='outputs/healed/healing_breadth/glean_math_keep25_seed1224_long768/step0500', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=False, dtype=torch.bfloat16, max_seq_len=2048, download_dir=None, load_format=auto, tensor_parallel_size=1, pipeline_parallel_size=1, data_parallel_size=1, decode_context_parallel_size=1, dcp_comm_backend=ag_rs, disable_custom_all_reduce=False, quantization=None, quantization_config=None, enforce_eager=True, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False, jit_monitor_mode='warn', jit_monitor_verbose=False), seed=0, served_model_name=student, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'mode': <CompilationMode.NONE: 0>, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'ir_enable_torch_wrap': False, 'splitting_ops': [], 'compile_mm_encoder': False, 'cudagraph_mm_encoder': False, 'encoder_cudagraph_token_budgets': [], 'encoder_cudagraph_max_vision_items_per_batch': 0, 'encoder_cudagraph_max_frames_per_batch': None, 'compile_sizes': [], 'compile_ranges_endpoints': [2048], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'size_asserts': False, 'alignment_asserts': False, 'scalar_asserts': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': <CUDAGraphMode.NONE: 0>, 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': True, 'fuse_act_quant': True, 'fuse_attn_quant': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False, 'fuse_rope_kvcache_cat_mla': False, 'fuse_act_padding': False}, 'max_cudagraph_capture_size': 0, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': False, 'static_all_moe_layers': []}, kernel_config=KernelConfig(ir_op_priority=IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native']), enable_flashinfer_autotune=True, enable_cutedsl_warmup=True, moe_backend='auto', linear_backend='auto')
259
+ (EngineCore pid=305625) INFO 07-15 18:44:06 [parallel_state.py:1607] world_size=1 rank=0 local_rank=0 distributed_init_method=tcp://192.168.0.15:53125 backend=nccl
260
+ (EngineCore pid=305625) INFO 07-15 18:44:07 [parallel_state.py:1942] rank 0 in world size 1 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank 0, EPLB rank N/A
261
+ (EngineCore pid=305625) INFO 07-15 18:44:07 [topk_topp_sampler.py:55] Using FlashInfer for top-p & top-k sampling.
262
+ (EngineCore pid=305625) INFO 07-15 18:44:07 [gpu_model_runner.py:5209] Starting to load model outputs/healed/healing_breadth/glean_math_keep25_seed1224_long768/step0500...
263
+ (EngineCore pid=305625) INFO 07-15 18:44:08 [cuda.py:476] Using FLASH_ATTN attention backend out of potential backends: ['FLASH_ATTN', 'FLASHINFER', 'TRITON_ATTN', 'FLEX_ATTENTION'].
264
+ (EngineCore pid=305625) INFO 07-15 18:44:08 [flash_attn.py:718] Using FlashAttention version 2
265
+ (EngineCore pid=305625) /home/henry/.cache/glean/megablocks-variable-93a1479bc15b/megablocks/grouped_gemm_util.py:10: UserWarning: Grouped GEMM not available.
266
+ (EngineCore pid=305625) warnings.warn('Grouped GEMM not available.')
267
+ (EngineCore pid=305625) INFO 07-15 18:44:08 [weight_utils.py:849] Filesystem type for checkpoints: EXT4. Checkpoint size: 3.89 GiB. Available RAM: 109.52 GiB.
268
+ (EngineCore pid=305625) INFO 07-15 18:44:08 [weight_utils.py:872] Auto-prefetch is disabled because the filesystem (EXT4) is not a recognized network FS (NFS/Lustre). If you want to force prefetching, start vLLM with --safetensors-load-strategy=prefetch.
269
+ (EngineCore pid=305625)
270
+ (EngineCore pid=305625)
271
+ (EngineCore pid=305625)
272
+ (EngineCore pid=305625)
273
+ (EngineCore pid=305625) INFO 07-15 18:44:10 [default_loader.py:430] Loading weights took 2.39 seconds
274
+ (EngineCore pid=305625) INFO 07-15 18:44:11 [gpu_model_runner.py:5306] Model loading took 3.89 GiB memory and 2.570320 seconds
275
+ (EngineCore pid=305625) INFO 07-15 18:44:12 [gpu_worker.py:538] Available KV cache memory: 15.82 GiB
276
+ (EngineCore pid=305625) INFO 07-15 18:44:12 [kv_cache_utils.py:2146] GPU KV cache size: 129,584 tokens
277
+ (EngineCore pid=305625) INFO 07-15 18:44:12 [kv_cache_utils.py:2147] Maximum concurrency for 2,048 tokens per request: 63.27x
278
+ (EngineCore pid=305625) INFO 07-15 18:44:12 [cutedsl_warmup.py:97] Skipping CuTeDSL warmup because no compile units were requested.
279
+ (EngineCore pid=305625) INFO 07-15 18:44:13 [jit_monitor.py:73] Kernel JIT monitor activated; monitored JIT compilations during inference will use mode=warn.
280
+ (EngineCore pid=305625) INFO 07-15 18:44:13 [core.py:344] init engine (profile, create kv cache, warmup model) took 2.05 s
281
+ (EngineCore pid=305625) INFO 07-15 18:44:13 [vllm.py:1042] Asynchronous scheduling is enabled.
282
+ (EngineCore pid=305625) WARNING 07-15 18:44:13 [vllm.py:1096] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
283
+ (EngineCore pid=305625) WARNING 07-15 18:44:13 [vllm.py:1144] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
284
+ (EngineCore pid=305625) INFO 07-15 18:44:13 [kernel.py:292] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native'])
285
+ (EngineCore pid=305625) INFO 07-15 18:44:13 [vllm.py:1322] Cudagraph is disabled under eager mode
286
+ (EngineCore pid=305625) INFO 07-15 18:44:13 [compilation.py:312] Enabled custom fusions: norm_quant, act_quant
287
+ (APIServer pid=305508) INFO 07-15 18:44:13 [api_server.py:612] Supported tasks: ['generate']
288
+ (APIServer pid=305508) WARNING 07-15 18:44:13 [__init__.py:36] SECURITY WARNING: Development endpoints are enabled! This should NOT be used in production!
289
+ (APIServer pid=305508) INFO 07-15 18:44:13 [hf.py:548] Detected the chat template content format to be 'string'. You can set `--chat-template-content-format` to override this.
290
+ (APIServer pid=305508) INFO 07-15 18:44:13 [api_server.py:616] Starting vLLM server on http://127.0.0.1:8377
291
+ (APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:37] Available routes are:
292
+ (APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /openapi.json, Methods: GET, HEAD
293
+ (APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /docs, Methods: GET, HEAD
294
+ (APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /docs/oauth2-redirect, Methods: GET, HEAD
295
+ (APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /redoc, Methods: GET, HEAD
296
+ (APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /load, Methods: GET
297
+ (APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /version, Methods: GET
298
+ (APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /health, Methods: GET
299
+ (APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /metrics, Methods: GET
300
+ (APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /tokenize, Methods: POST
301
+ (APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /detokenize, Methods: POST
302
+ (APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /v1/models, Methods: GET
303
+ (APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /ping, Methods: GET
304
+ (APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /ping, Methods: POST
305
+ (APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /invocations, Methods: POST
306
+ (APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /reset_prefix_cache, Methods: POST
307
+ (APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /reset_mm_cache, Methods: POST
308
+ (APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /reset_encoder_cache, Methods: POST
309
+ (APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /pause, Methods: POST
310
+ (APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /resume, Methods: POST
311
+ (APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /is_paused, Methods: GET
312
+ (APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /init_weight_transfer_engine, Methods: POST
313
+ (APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /start_weight_update, Methods: POST
314
+ (APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /update_weights, Methods: POST
315
+ (APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /finish_weight_update, Methods: POST
316
+ (APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /get_world_size, Methods: GET
317
+ (APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /collective_rpc, Methods: POST
318
+ (APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /server_info, Methods: GET
319
+ (APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /sleep, Methods: POST
320
+ (APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /wake_up, Methods: POST
321
+ (APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /is_sleeping, Methods: GET
322
+ (APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /v1/chat/completions, Methods: POST
323
+ (APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /v1/chat/completions/batch, Methods: POST
324
+ (APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /v1/responses, Methods: POST
325
+ (APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /v1/responses/{response_id}, Methods: GET
326
+ (APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /v1/responses/{response_id}/cancel, Methods: POST
327
+ (APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /v1/completions, Methods: POST
328
+ (APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /v1/messages, Methods: POST
329
+ (APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /v1/messages/count_tokens, Methods: POST
330
+ (APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /generative_scoring, Methods: POST
331
+ (APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /scale_elastic_ep, Methods: POST
332
+ (APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /is_scaling_elastic_ep, Methods: POST
333
+ (APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /v1/chat/completions/render, Methods: POST
334
+ (APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /v1/completions/render, Methods: POST
335
+ (APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /v1/chat/completions/derender, Methods: POST
336
+ (APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /v1/completions/derender, Methods: POST
337
+ (APIServer pid=305508) INFO 07-15 18:44:13 [launcher.py:46] Route: /inference/v1/generate, Methods: POST
338
+ (APIServer pid=305508) INFO: Started server process [305508]
339
+ (APIServer pid=305508) INFO: Waiting for application startup.
340
+ (APIServer pid=305508) INFO: Application startup complete.
341
+ (APIServer pid=305508) INFO: 127.0.0.1:39768 - "GET /health HTTP/1.1" 200 OK
342
+ (EngineCore pid=305625) WARNING 07-15 18:44:15 [jit_monitor.py:129] Triton kernel JIT compilation during inference: _build_route_rows. This causes a latency spike; consider extending warmup to cover this shape/config.
343
+ (APIServer pid=305508) INFO 07-15 18:44:23 [loggers.py:273] Engine 000: Avg prompt throughput: 1848.7 tokens/s, Avg generation throughput: 2142.8 tokens/s, Running: 253 reqs, Waiting: 1038 reqs, GPU KV cache usage: 31.1%, Prefix cache hit rate: 94.1%
344
+ (APIServer pid=305508) INFO 07-15 18:44:33 [loggers.py:273] Engine 000: Avg prompt throughput: 796.1 tokens/s, Avg generation throughput: 2956.5 tokens/s, Running: 256 reqs, Waiting: 915 reqs, GPU KV cache usage: 43.0%, Prefix cache hit rate: 94.3%
345
+ (APIServer pid=305508) INFO 07-15 18:44:43 [loggers.py:273] Engine 000: Avg prompt throughput: 452.7 tokens/s, Avg generation throughput: 2936.8 tokens/s, Running: 256 reqs, Waiting: 850 reqs, GPU KV cache usage: 58.3%, Prefix cache hit rate: 94.2%
346
+ (APIServer pid=305508) INFO 07-15 18:44:53 [loggers.py:273] Engine 000: Avg prompt throughput: 219.0 tokens/s, Avg generation throughput: 2863.2 tokens/s, Running: 255 reqs, Waiting: 815 reqs, GPU KV cache usage: 75.4%, Prefix cache hit rate: 94.3%
347
+ (APIServer pid=305508) INFO 07-15 18:45:03 [loggers.py:273] Engine 000: Avg prompt throughput: 960.0 tokens/s, Avg generation throughput: 2749.6 tokens/s, Running: 256 reqs, Waiting: 665 reqs, GPU KV cache usage: 45.6%, Prefix cache hit rate: 94.3%
348
+ (APIServer pid=305508) INFO 07-15 18:45:13 [loggers.py:273] Engine 000: Avg prompt throughput: 647.5 tokens/s, Avg generation throughput: 2908.7 tokens/s, Running: 255 reqs, Waiting: 567 reqs, GPU KV cache usage: 49.9%, Prefix cache hit rate: 94.4%
349
+ (APIServer pid=305508) INFO 07-15 18:45:23 [loggers.py:273] Engine 000: Avg prompt throughput: 662.9 tokens/s, Avg generation throughput: 2908.2 tokens/s, Running: 255 reqs, Waiting: 470 reqs, GPU KV cache usage: 53.1%, Prefix cache hit rate: 94.4%
350
+ (APIServer pid=305508) INFO 07-15 18:45:33 [loggers.py:273] Engine 000: Avg prompt throughput: 589.0 tokens/s, Avg generation throughput: 2909.2 tokens/s, Running: 255 reqs, Waiting: 379 reqs, GPU KV cache usage: 57.6%, Prefix cache hit rate: 94.4%
351
+ (APIServer pid=305508) INFO 07-15 18:45:43 [loggers.py:273] Engine 000: Avg prompt throughput: 457.3 tokens/s, Avg generation throughput: 2860.3 tokens/s, Running: 255 reqs, Waiting: 314 reqs, GPU KV cache usage: 68.8%, Prefix cache hit rate: 94.4%
352
+ (APIServer pid=305508) INFO 07-15 18:45:53 [loggers.py:273] Engine 000: Avg prompt throughput: 707.3 tokens/s, Avg generation throughput: 2856.2 tokens/s, Running: 253 reqs, Waiting: 209 reqs, GPU KV cache usage: 58.3%, Prefix cache hit rate: 94.4%
353
+ (APIServer pid=305508) INFO 07-15 18:46:03 [loggers.py:273] Engine 000: Avg prompt throughput: 818.5 tokens/s, Avg generation throughput: 2854.9 tokens/s, Running: 255 reqs, Waiting: 87 reqs, GPU KV cache usage: 53.2%, Prefix cache hit rate: 94.4%
354
+ (APIServer pid=305508) INFO 07-15 18:46:14 [loggers.py:273] Engine 000: Avg prompt throughput: 592.2 tokens/s, Avg generation throughput: 2896.5 tokens/s, Running: 241 reqs, Waiting: 0 reqs, GPU KV cache usage: 54.5%, Prefix cache hit rate: 94.4%
355
+ (APIServer pid=305508) INFO 07-15 18:46:24 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 2555.5 tokens/s, Running: 144 reqs, Waiting: 0 reqs, GPU KV cache usage: 44.9%, Prefix cache hit rate: 94.4%
356
+ (APIServer pid=305508) INFO 07-15 18:46:34 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 2061.1 tokens/s, Running: 64 reqs, Waiting: 0 reqs, GPU KV cache usage: 25.6%, Prefix cache hit rate: 94.4%
357
+ (APIServer pid=305508) INFO: 127.0.0.1:45906 - "POST /v1/completions HTTP/1.1" 200 OK
358
+ (EngineCore pid=305625) INFO 07-15 18:46:40 [core.py:1214] [shutdown] EngineCore: trigger received signal=SIGTERM
359
+ (APIServer pid=305508) INFO 07-15 18:46:40 [launcher.py:100] [shutdown] API server: shutdown triggered
360
+ (APIServer pid=305508) INFO 07-15 18:46:40 [launcher.py:116] [shutdown] API server: stopping engine client mode=abort timeout=0s
361
+ (EngineCore pid=305625) INFO 07-15 18:46:40 [core.py:1333] [shutdown] EngineCore: start mode=abort timeout=0s
362
+ (EngineCore pid=305625) INFO 07-15 18:46:40 [core.py:1364] [shutdown] EngineCore: request processing complete; starting resource teardown
363
+ (EngineCore pid=305625) INFO 07-15 18:46:40 [core.py:1227] [shutdown] EngineCore: exiting busy loop
364
+ (APIServer pid=305508) INFO 07-15 18:46:40 [core_client.py:655] [shutdown] MPClient: start timeout=0s
365
+ (APIServer pid=305508) INFO 07-15 18:46:40 [core_client.py:657] [shutdown] MPClient: stopping engine manager
366
+ (APIServer pid=305508) WARNING 07-15 18:46:40 [utils.py:626] [shutdown] Process manager: force killing remaining processes count=1
367
+ (APIServer pid=305508) INFO 07-15 18:46:40 [core_client.py:659] [shutdown] MPClient: engine manager stopped
368
+ (APIServer pid=305508) INFO 07-15 18:46:40 [core_client.py:660] [shutdown] MPClient: cleaning up background resources
369
+ (APIServer pid=305508) INFO 07-15 18:46:40 [core_client.py:662] [shutdown] MPClient: complete
370
+ (APIServer pid=305508) INFO 07-15 18:46:40 [launcher.py:125] [shutdown] API server: engine client stopped
371
+ (APIServer pid=305508) INFO 07-15 18:46:40 [launcher.py:128] [shutdown] API server: signalling HTTP server shutdown
372
+ (APIServer pid=305508) INFO 07-15 18:46:40 [launcher.py:149] [shutdown] API server: shutting down FastAPI HTTP server
373
+ (APIServer pid=305508) INFO: Shutting down
374
+ (APIServer pid=305508) INFO: Waiting for application shutdown.
375
+ (APIServer pid=305508) INFO: Application shutdown complete.
376
+ /home/henry/.local/share/uv/python/cpython-3.12.12-linux-x86_64-gnu/lib/python3.12/multiprocessing/resource_tracker.py:279: UserWarning: resource_tracker: There appear to be 1 leaked semaphore objects to clean up at shutdown
377
+ warnings.warn('resource_tracker: There appear to be %d '
evals/healing_breadth/glean_math_keep25_seed1224_step0050_chat.json ADDED
The diff for this file is too large to render. See raw diff
 
evals/healing_breadth/glean_math_keep25_seed1224_step0050_chat.json.server.log ADDED
@@ -0,0 +1,138 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ (APIServer pid=308068) INFO 07-15 18:54:12 [api_utils.py:339]
2
+ (APIServer pid=308068) INFO 07-15 18:54:12 [api_utils.py:339] β–ˆ β–ˆ β–ˆβ–„ β–„β–ˆ
3
+ (APIServer pid=308068) INFO 07-15 18:54:12 [api_utils.py:339] β–„β–„ β–„β–ˆ β–ˆ β–ˆ β–ˆ β–€β–„β–€ β–ˆ version 0.25.0
4
+ (APIServer pid=308068) INFO 07-15 18:54:12 [api_utils.py:339] β–ˆβ–„β–ˆβ–€ β–ˆ β–ˆ β–ˆ β–ˆ model outputs/healed/healing_breadth/glean_math_keep25_seed1224/step0050
5
+ (APIServer pid=308068) INFO 07-15 18:54:12 [api_utils.py:339] β–€β–€ β–€β–€β–€β–€β–€ β–€β–€β–€β–€β–€ β–€ β–€
6
+ (APIServer pid=308068) INFO 07-15 18:54:12 [api_utils.py:339]
7
+ (APIServer pid=308068) INFO 07-15 18:54:12 [api_utils.py:273] non-default args: {'model_tag': 'outputs/healed/healing_breadth/glean_math_keep25_seed1224/step0050', 'host': '127.0.0.1', 'port': 8377, 'model': 'outputs/healed/healing_breadth/glean_math_keep25_seed1224/step0050', 'max_model_len': 2048, 'enforce_eager': True, 'served_model_name': ['student'], 'gpu_memory_utilization': 0.85}
8
+ (APIServer pid=308068) INFO 07-15 18:54:12 [model.py:619] Resolved architecture: PrunedOlmoeForCausalLM
9
+ (APIServer pid=308068) INFO 07-15 18:54:12 [model.py:1776] Using max model len 2048
10
+ (APIServer pid=308068) INFO 07-15 18:54:12 [vllm.py:1042] Asynchronous scheduling is enabled.
11
+ (APIServer pid=308068) WARNING 07-15 18:54:12 [vllm.py:1096] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
12
+ (APIServer pid=308068) WARNING 07-15 18:54:12 [vllm.py:1144] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
13
+ (APIServer pid=308068) INFO 07-15 18:54:12 [kernel.py:292] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native'])
14
+ (APIServer pid=308068) INFO 07-15 18:54:12 [vllm.py:1322] Cudagraph is disabled under eager mode
15
+ (APIServer pid=308068) INFO 07-15 18:54:12 [compilation.py:312] Enabled custom fusions: norm_quant, act_quant
16
+ (EngineCore pid=308191) INFO 07-15 18:54:19 [core.py:114] Initializing a V1 LLM engine (v0.25.0) with config: model='outputs/healed/healing_breadth/glean_math_keep25_seed1224/step0050', speculative_config=None, tokenizer='outputs/healed/healing_breadth/glean_math_keep25_seed1224/step0050', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=False, dtype=torch.bfloat16, max_seq_len=2048, download_dir=None, load_format=auto, tensor_parallel_size=1, pipeline_parallel_size=1, data_parallel_size=1, decode_context_parallel_size=1, dcp_comm_backend=ag_rs, disable_custom_all_reduce=False, quantization=None, quantization_config=None, enforce_eager=True, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False, jit_monitor_mode='warn', jit_monitor_verbose=False), seed=0, served_model_name=student, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'mode': <CompilationMode.NONE: 0>, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'ir_enable_torch_wrap': False, 'splitting_ops': [], 'compile_mm_encoder': False, 'cudagraph_mm_encoder': False, 'encoder_cudagraph_token_budgets': [], 'encoder_cudagraph_max_vision_items_per_batch': 0, 'encoder_cudagraph_max_frames_per_batch': None, 'compile_sizes': [], 'compile_ranges_endpoints': [2048], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'size_asserts': False, 'alignment_asserts': False, 'scalar_asserts': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': <CUDAGraphMode.NONE: 0>, 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': True, 'fuse_act_quant': True, 'fuse_attn_quant': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False, 'fuse_rope_kvcache_cat_mla': False, 'fuse_act_padding': False}, 'max_cudagraph_capture_size': 0, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': False, 'static_all_moe_layers': []}, kernel_config=KernelConfig(ir_op_priority=IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native']), enable_flashinfer_autotune=True, enable_cutedsl_warmup=True, moe_backend='auto', linear_backend='auto')
17
+ (EngineCore pid=308191) INFO 07-15 18:54:20 [parallel_state.py:1607] world_size=1 rank=0 local_rank=0 distributed_init_method=tcp://192.168.0.15:60767 backend=nccl
18
+ (EngineCore pid=308191) INFO 07-15 18:54:20 [parallel_state.py:1942] rank 0 in world size 1 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank 0, EPLB rank N/A
19
+ (EngineCore pid=308191) INFO 07-15 18:54:20 [topk_topp_sampler.py:55] Using FlashInfer for top-p & top-k sampling.
20
+ (EngineCore pid=308191) INFO 07-15 18:54:20 [gpu_model_runner.py:5209] Starting to load model outputs/healed/healing_breadth/glean_math_keep25_seed1224/step0050...
21
+ (EngineCore pid=308191) INFO 07-15 18:54:21 [cuda.py:476] Using FLASH_ATTN attention backend out of potential backends: ['FLASH_ATTN', 'FLASHINFER', 'TRITON_ATTN', 'FLEX_ATTENTION'].
22
+ (EngineCore pid=308191) INFO 07-15 18:54:21 [flash_attn.py:718] Using FlashAttention version 2
23
+ (EngineCore pid=308191) /home/henry/.cache/glean/megablocks-variable-93a1479bc15b/megablocks/grouped_gemm_util.py:10: UserWarning: Grouped GEMM not available.
24
+ (EngineCore pid=308191) warnings.warn('Grouped GEMM not available.')
25
+ (EngineCore pid=308191) INFO 07-15 18:54:21 [weight_utils.py:849] Filesystem type for checkpoints: EXT4. Checkpoint size: 3.89 GiB. Available RAM: 109.29 GiB.
26
+ (EngineCore pid=308191) INFO 07-15 18:54:21 [weight_utils.py:872] Auto-prefetch is disabled because the filesystem (EXT4) is not a recognized network FS (NFS/Lustre). If you want to force prefetching, start vLLM with --safetensors-load-strategy=prefetch.
27
+ (EngineCore pid=308191)
28
+ (EngineCore pid=308191)
29
+ (EngineCore pid=308191)
30
+ (EngineCore pid=308191)
31
+ (EngineCore pid=308191) INFO 07-15 18:54:24 [default_loader.py:430] Loading weights took 2.72 seconds
32
+ (EngineCore pid=308191) INFO 07-15 18:54:24 [gpu_model_runner.py:5306] Model loading took 3.89 GiB memory and 2.901110 seconds
33
+ (EngineCore pid=308191) INFO 07-15 18:54:26 [gpu_worker.py:538] Available KV cache memory: 15.82 GiB
34
+ (EngineCore pid=308191) INFO 07-15 18:54:26 [kv_cache_utils.py:2146] GPU KV cache size: 129,584 tokens
35
+ (EngineCore pid=308191) INFO 07-15 18:54:26 [kv_cache_utils.py:2147] Maximum concurrency for 2,048 tokens per request: 63.27x
36
+ (EngineCore pid=308191) INFO 07-15 18:54:26 [cutedsl_warmup.py:97] Skipping CuTeDSL warmup because no compile units were requested.
37
+ (EngineCore pid=308191) INFO 07-15 18:54:26 [jit_monitor.py:73] Kernel JIT monitor activated; monitored JIT compilations during inference will use mode=warn.
38
+ (EngineCore pid=308191) INFO 07-15 18:54:26 [core.py:344] init engine (profile, create kv cache, warmup model) took 2.06 s
39
+ (EngineCore pid=308191) INFO 07-15 18:54:26 [vllm.py:1042] Asynchronous scheduling is enabled.
40
+ (EngineCore pid=308191) WARNING 07-15 18:54:26 [vllm.py:1096] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
41
+ (EngineCore pid=308191) WARNING 07-15 18:54:26 [vllm.py:1144] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
42
+ (EngineCore pid=308191) INFO 07-15 18:54:26 [kernel.py:292] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native'])
43
+ (EngineCore pid=308191) INFO 07-15 18:54:26 [vllm.py:1322] Cudagraph is disabled under eager mode
44
+ (EngineCore pid=308191) INFO 07-15 18:54:26 [compilation.py:312] Enabled custom fusions: norm_quant, act_quant
45
+ (APIServer pid=308068) INFO 07-15 18:54:26 [api_server.py:612] Supported tasks: ['generate']
46
+ (APIServer pid=308068) WARNING 07-15 18:54:26 [__init__.py:36] SECURITY WARNING: Development endpoints are enabled! This should NOT be used in production!
47
+ (APIServer pid=308068) INFO 07-15 18:54:27 [hf.py:548] Detected the chat template content format to be 'string'. You can set `--chat-template-content-format` to override this.
48
+ (APIServer pid=308068) INFO 07-15 18:54:27 [api_server.py:616] Starting vLLM server on http://127.0.0.1:8377
49
+ (APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:37] Available routes are:
50
+ (APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /openapi.json, Methods: HEAD, GET
51
+ (APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /docs, Methods: HEAD, GET
52
+ (APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /docs/oauth2-redirect, Methods: HEAD, GET
53
+ (APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /redoc, Methods: HEAD, GET
54
+ (APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /load, Methods: GET
55
+ (APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /version, Methods: GET
56
+ (APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /health, Methods: GET
57
+ (APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /metrics, Methods: GET
58
+ (APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /tokenize, Methods: POST
59
+ (APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /detokenize, Methods: POST
60
+ (APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /v1/models, Methods: GET
61
+ (APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /ping, Methods: GET
62
+ (APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /ping, Methods: POST
63
+ (APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /invocations, Methods: POST
64
+ (APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /reset_prefix_cache, Methods: POST
65
+ (APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /reset_mm_cache, Methods: POST
66
+ (APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /reset_encoder_cache, Methods: POST
67
+ (APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /pause, Methods: POST
68
+ (APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /resume, Methods: POST
69
+ (APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /is_paused, Methods: GET
70
+ (APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /init_weight_transfer_engine, Methods: POST
71
+ (APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /start_weight_update, Methods: POST
72
+ (APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /update_weights, Methods: POST
73
+ (APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /finish_weight_update, Methods: POST
74
+ (APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /get_world_size, Methods: GET
75
+ (APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /collective_rpc, Methods: POST
76
+ (APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /server_info, Methods: GET
77
+ (APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /sleep, Methods: POST
78
+ (APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /wake_up, Methods: POST
79
+ (APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /is_sleeping, Methods: GET
80
+ (APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /v1/chat/completions, Methods: POST
81
+ (APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /v1/chat/completions/batch, Methods: POST
82
+ (APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /v1/responses, Methods: POST
83
+ (APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /v1/responses/{response_id}, Methods: GET
84
+ (APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /v1/responses/{response_id}/cancel, Methods: POST
85
+ (APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /v1/completions, Methods: POST
86
+ (APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /v1/messages, Methods: POST
87
+ (APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /v1/messages/count_tokens, Methods: POST
88
+ (APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /generative_scoring, Methods: POST
89
+ (APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /scale_elastic_ep, Methods: POST
90
+ (APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /is_scaling_elastic_ep, Methods: POST
91
+ (APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /v1/chat/completions/render, Methods: POST
92
+ (APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /v1/completions/render, Methods: POST
93
+ (APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /v1/chat/completions/derender, Methods: POST
94
+ (APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /v1/completions/derender, Methods: POST
95
+ (APIServer pid=308068) INFO 07-15 18:54:27 [launcher.py:46] Route: /inference/v1/generate, Methods: POST
96
+ (APIServer pid=308068) INFO: Started server process [308068]
97
+ (APIServer pid=308068) INFO: Waiting for application startup.
98
+ (APIServer pid=308068) INFO: Application startup complete.
99
+ (APIServer pid=308068) INFO: 127.0.0.1:34220 - "GET /health HTTP/1.1" 200 OK
100
+ (EngineCore pid=308191) WARNING 07-15 18:54:28 [jit_monitor.py:129] Triton kernel JIT compilation during inference: _build_route_rows. This causes a latency spike; consider extending warmup to cover this shape/config.
101
+ (APIServer pid=308068) INFO 07-15 18:54:37 [loggers.py:273] Engine 000: Avg prompt throughput: 2059.2 tokens/s, Avg generation throughput: 2322.0 tokens/s, Running: 256 reqs, Waiting: 1057 reqs, GPU KV cache usage: 36.5%, Prefix cache hit rate: 93.1%
102
+ (APIServer pid=308068) INFO 07-15 18:54:47 [loggers.py:273] Engine 000: Avg prompt throughput: 316.3 tokens/s, Avg generation throughput: 2990.5 tokens/s, Running: 256 reqs, Waiting: 1017 reqs, GPU KV cache usage: 54.9%, Prefix cache hit rate: 93.1%
103
+ (APIServer pid=308068) INFO 07-15 18:54:57 [loggers.py:273] Engine 000: Avg prompt throughput: 377.3 tokens/s, Avg generation throughput: 2887.0 tokens/s, Running: 256 reqs, Waiting: 968 reqs, GPU KV cache usage: 67.2%, Prefix cache hit rate: 93.2%
104
+ (APIServer pid=308068) INFO 07-15 18:55:07 [loggers.py:273] Engine 000: Avg prompt throughput: 374.5 tokens/s, Avg generation throughput: 2810.4 tokens/s, Running: 254 reqs, Waiting: 919 reqs, GPU KV cache usage: 76.5%, Prefix cache hit rate: 93.2%
105
+ (APIServer pid=308068) INFO 07-15 18:55:17 [loggers.py:273] Engine 000: Avg prompt throughput: 1206.0 tokens/s, Avg generation throughput: 2749.2 tokens/s, Running: 256 reqs, Waiting: 766 reqs, GPU KV cache usage: 42.7%, Prefix cache hit rate: 93.3%
106
+ (APIServer pid=308068) INFO 07-15 18:55:27 [loggers.py:273] Engine 000: Avg prompt throughput: 472.2 tokens/s, Avg generation throughput: 2962.8 tokens/s, Running: 253 reqs, Waiting: 705 reqs, GPU KV cache usage: 51.4%, Prefix cache hit rate: 93.3%
107
+ (APIServer pid=308068) INFO 07-15 18:55:37 [loggers.py:273] Engine 000: Avg prompt throughput: 409.7 tokens/s, Avg generation throughput: 2912.6 tokens/s, Running: 254 reqs, Waiting: 652 reqs, GPU KV cache usage: 61.6%, Prefix cache hit rate: 93.3%
108
+ (APIServer pid=308068) INFO 07-15 18:55:47 [loggers.py:273] Engine 000: Avg prompt throughput: 655.1 tokens/s, Avg generation throughput: 2832.9 tokens/s, Running: 256 reqs, Waiting: 567 reqs, GPU KV cache usage: 62.0%, Prefix cache hit rate: 93.3%
109
+ (APIServer pid=308068) INFO 07-15 18:55:57 [loggers.py:273] Engine 000: Avg prompt throughput: 561.6 tokens/s, Avg generation throughput: 2860.4 tokens/s, Running: 255 reqs, Waiting: 499 reqs, GPU KV cache usage: 66.1%, Prefix cache hit rate: 93.3%
110
+ (APIServer pid=308068) INFO 07-15 18:56:07 [loggers.py:273] Engine 000: Avg prompt throughput: 781.7 tokens/s, Avg generation throughput: 2882.3 tokens/s, Running: 254 reqs, Waiting: 397 reqs, GPU KV cache usage: 54.7%, Prefix cache hit rate: 93.4%
111
+ (APIServer pid=308068) INFO 07-15 18:56:17 [loggers.py:273] Engine 000: Avg prompt throughput: 661.3 tokens/s, Avg generation throughput: 2884.0 tokens/s, Running: 256 reqs, Waiting: 319 reqs, GPU KV cache usage: 59.2%, Prefix cache hit rate: 93.3%
112
+ (APIServer pid=308068) INFO 07-15 18:56:27 [loggers.py:273] Engine 000: Avg prompt throughput: 591.5 tokens/s, Avg generation throughput: 2859.2 tokens/s, Running: 256 reqs, Waiting: 246 reqs, GPU KV cache usage: 61.7%, Prefix cache hit rate: 93.4%
113
+ (APIServer pid=308068) INFO 07-15 18:56:37 [loggers.py:273] Engine 000: Avg prompt throughput: 639.7 tokens/s, Avg generation throughput: 2884.4 tokens/s, Running: 255 reqs, Waiting: 163 reqs, GPU KV cache usage: 58.6%, Prefix cache hit rate: 93.4%
114
+ (APIServer pid=308068) INFO 07-15 18:56:47 [loggers.py:273] Engine 000: Avg prompt throughput: 662.0 tokens/s, Avg generation throughput: 2858.9 tokens/s, Running: 254 reqs, Waiting: 83 reqs, GPU KV cache usage: 57.5%, Prefix cache hit rate: 93.4%
115
+ (APIServer pid=308068) INFO 07-15 18:56:57 [loggers.py:273] Engine 000: Avg prompt throughput: 677.8 tokens/s, Avg generation throughput: 2908.3 tokens/s, Running: 253 reqs, Waiting: 0 reqs, GPU KV cache usage: 57.3%, Prefix cache hit rate: 93.4%
116
+ (APIServer pid=308068) INFO 07-15 18:57:07 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 2752.8 tokens/s, Running: 182 reqs, Waiting: 0 reqs, GPU KV cache usage: 53.3%, Prefix cache hit rate: 93.4%
117
+ (APIServer pid=308068) INFO 07-15 18:57:17 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 2179.0 tokens/s, Running: 74 reqs, Waiting: 0 reqs, GPU KV cache usage: 29.0%, Prefix cache hit rate: 93.4%
118
+ (APIServer pid=308068) INFO: 127.0.0.1:34236 - "POST /v1/completions HTTP/1.1" 200 OK
119
+ (EngineCore pid=308191) INFO 07-15 18:57:24 [core.py:1214] [shutdown] EngineCore: trigger received signal=SIGTERM
120
+ (APIServer pid=308068) INFO 07-15 18:57:24 [launcher.py:100] [shutdown] API server: shutdown triggered
121
+ (APIServer pid=308068) INFO 07-15 18:57:24 [launcher.py:116] [shutdown] API server: stopping engine client mode=abort timeout=0s
122
+ (EngineCore pid=308191) INFO 07-15 18:57:24 [core.py:1333] [shutdown] EngineCore: start mode=abort timeout=0s
123
+ (EngineCore pid=308191) INFO 07-15 18:57:24 [core.py:1364] [shutdown] EngineCore: request processing complete; starting resource teardown
124
+ (EngineCore pid=308191) INFO 07-15 18:57:24 [core.py:1227] [shutdown] EngineCore: exiting busy loop
125
+ (APIServer pid=308068) INFO 07-15 18:57:24 [core_client.py:655] [shutdown] MPClient: start timeout=0s
126
+ (APIServer pid=308068) INFO 07-15 18:57:24 [core_client.py:657] [shutdown] MPClient: stopping engine manager
127
+ (APIServer pid=308068) WARNING 07-15 18:57:24 [utils.py:626] [shutdown] Process manager: force killing remaining processes count=1
128
+ (APIServer pid=308068) INFO 07-15 18:57:24 [core_client.py:659] [shutdown] MPClient: engine manager stopped
129
+ (APIServer pid=308068) INFO 07-15 18:57:24 [core_client.py:660] [shutdown] MPClient: cleaning up background resources
130
+ (APIServer pid=308068) INFO 07-15 18:57:24 [core_client.py:662] [shutdown] MPClient: complete
131
+ (APIServer pid=308068) INFO 07-15 18:57:24 [launcher.py:125] [shutdown] API server: engine client stopped
132
+ (APIServer pid=308068) INFO 07-15 18:57:24 [launcher.py:128] [shutdown] API server: signalling HTTP server shutdown
133
+ (APIServer pid=308068) INFO 07-15 18:57:24 [launcher.py:149] [shutdown] API server: shutting down FastAPI HTTP server
134
+ (APIServer pid=308068) INFO: Shutting down
135
+ (APIServer pid=308068) INFO: Waiting for application shutdown.
136
+ (APIServer pid=308068) INFO: Application shutdown complete.
137
+ /home/henry/.local/share/uv/python/cpython-3.12.12-linux-x86_64-gnu/lib/python3.12/multiprocessing/resource_tracker.py:279: UserWarning: resource_tracker: There appear to be 1 leaked semaphore objects to clean up at shutdown
138
+ warnings.warn('resource_tracker: There appear to be %d '
evals/healing_breadth/glean_math_keep25_seed1224_step0050_raw.json ADDED
The diff for this file is too large to render. See raw diff
 
evals/healing_breadth/glean_math_keep25_seed1224_step0050_raw.json.server.log ADDED
@@ -0,0 +1,260 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ (APIServer pid=304785) INFO 07-15 18:42:58 [api_utils.py:339]
2
+ (APIServer pid=304785) INFO 07-15 18:42:58 [api_utils.py:339] β–ˆ β–ˆ β–ˆβ–„ β–„β–ˆ
3
+ (APIServer pid=304785) INFO 07-15 18:42:58 [api_utils.py:339] β–„β–„ β–„β–ˆ β–ˆ β–ˆ β–ˆ β–€β–„β–€ β–ˆ version 0.25.0
4
+ (APIServer pid=304785) INFO 07-15 18:42:58 [api_utils.py:339] β–ˆβ–„β–ˆβ–€ β–ˆ β–ˆ β–ˆ β–ˆ model outputs/healed/healing_breadth/glean_math_keep25_seed1224/step0050
5
+ (APIServer pid=304785) INFO 07-15 18:42:58 [api_utils.py:339] β–€β–€ β–€β–€β–€β–€β–€ β–€β–€β–€β–€β–€ β–€ β–€
6
+ (APIServer pid=304785) INFO 07-15 18:42:58 [api_utils.py:339]
7
+ (APIServer pid=304785) INFO 07-15 18:42:58 [api_utils.py:273] non-default args: {'model_tag': 'outputs/healed/healing_breadth/glean_math_keep25_seed1224/step0050', 'host': '127.0.0.1', 'port': 8377, 'model': 'outputs/healed/healing_breadth/glean_math_keep25_seed1224/step0050', 'max_model_len': 2048, 'enforce_eager': True, 'served_model_name': ['student'], 'gpu_memory_utilization': 0.85}
8
+ (APIServer pid=304785) INFO 07-15 18:42:58 [model.py:619] Resolved architecture: PrunedOlmoeForCausalLM
9
+ (APIServer pid=304785) INFO 07-15 18:42:58 [model.py:1776] Using max model len 2048
10
+ (APIServer pid=304785) INFO 07-15 18:42:58 [vllm.py:1042] Asynchronous scheduling is enabled.
11
+ (APIServer pid=304785) WARNING 07-15 18:42:58 [vllm.py:1096] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
12
+ (APIServer pid=304785) WARNING 07-15 18:42:58 [vllm.py:1144] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
13
+ (APIServer pid=304785) INFO 07-15 18:42:58 [kernel.py:292] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native'])
14
+ (APIServer pid=304785) INFO 07-15 18:42:58 [vllm.py:1322] Cudagraph is disabled under eager mode
15
+ (APIServer pid=304785) INFO 07-15 18:42:58 [compilation.py:312] Enabled custom fusions: norm_quant, act_quant
16
+ (EngineCore pid=304916) INFO 07-15 18:43:05 [core.py:114] Initializing a V1 LLM engine (v0.25.0) with config: model='outputs/healed/healing_breadth/glean_math_keep25_seed1224/step0050', speculative_config=None, tokenizer='outputs/healed/healing_breadth/glean_math_keep25_seed1224/step0050', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=False, dtype=torch.bfloat16, max_seq_len=2048, download_dir=None, load_format=auto, tensor_parallel_size=1, pipeline_parallel_size=1, data_parallel_size=1, decode_context_parallel_size=1, dcp_comm_backend=ag_rs, disable_custom_all_reduce=False, quantization=None, quantization_config=None, enforce_eager=True, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False, jit_monitor_mode='warn', jit_monitor_verbose=False), seed=0, served_model_name=student, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'mode': <CompilationMode.NONE: 0>, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'ir_enable_torch_wrap': False, 'splitting_ops': [], 'compile_mm_encoder': False, 'cudagraph_mm_encoder': False, 'encoder_cudagraph_token_budgets': [], 'encoder_cudagraph_max_vision_items_per_batch': 0, 'encoder_cudagraph_max_frames_per_batch': None, 'compile_sizes': [], 'compile_ranges_endpoints': [2048], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'size_asserts': False, 'alignment_asserts': False, 'scalar_asserts': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': <CUDAGraphMode.NONE: 0>, 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': True, 'fuse_act_quant': True, 'fuse_attn_quant': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False, 'fuse_rope_kvcache_cat_mla': False, 'fuse_act_padding': False}, 'max_cudagraph_capture_size': 0, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': False, 'static_all_moe_layers': []}, kernel_config=KernelConfig(ir_op_priority=IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native']), enable_flashinfer_autotune=True, enable_cutedsl_warmup=True, moe_backend='auto', linear_backend='auto')
17
+ (EngineCore pid=304916) INFO 07-15 18:43:05 [parallel_state.py:1607] world_size=1 rank=0 local_rank=0 distributed_init_method=tcp://192.168.0.15:45135 backend=nccl
18
+ (EngineCore pid=304916) INFO 07-15 18:43:06 [parallel_state.py:1942] rank 0 in world size 1 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank 0, EPLB rank N/A
19
+ (EngineCore pid=304916) INFO 07-15 18:43:06 [topk_topp_sampler.py:55] Using FlashInfer for top-p & top-k sampling.
20
+ (EngineCore pid=304916) INFO 07-15 18:43:06 [gpu_model_runner.py:5209] Starting to load model outputs/healed/healing_breadth/glean_math_keep25_seed1224/step0050...
21
+ (EngineCore pid=304916) INFO 07-15 18:43:07 [cuda.py:476] Using FLASH_ATTN attention backend out of potential backends: ['FLASH_ATTN', 'FLASHINFER', 'TRITON_ATTN', 'FLEX_ATTENTION'].
22
+ (EngineCore pid=304916) INFO 07-15 18:43:07 [flash_attn.py:718] Using FlashAttention version 2
23
+ (EngineCore pid=304916) /home/henry/.cache/glean/megablocks-variable-93a1479bc15b/megablocks/grouped_gemm_util.py:10: UserWarning: Grouped GEMM not available.
24
+ (EngineCore pid=304916) warnings.warn('Grouped GEMM not available.')
25
+ (EngineCore pid=304916) INFO 07-15 18:43:07 [weight_utils.py:849] Filesystem type for checkpoints: EXT4. Checkpoint size: 3.89 GiB. Available RAM: 109.76 GiB.
26
+ (EngineCore pid=304916) INFO 07-15 18:43:07 [weight_utils.py:872] Auto-prefetch is disabled because the filesystem (EXT4) is not a recognized network FS (NFS/Lustre). If you want to force prefetching, start vLLM with --safetensors-load-strategy=prefetch.
27
+ (EngineCore pid=304916)
28
+ (EngineCore pid=304916)
29
+ (EngineCore pid=304916)
30
+ (EngineCore pid=304916)
31
+ (EngineCore pid=304916) INFO 07-15 18:43:09 [default_loader.py:430] Loading weights took 2.61 seconds
32
+ (EngineCore pid=304916) INFO 07-15 18:43:10 [gpu_model_runner.py:5306] Model loading took 3.89 GiB memory and 2.790794 seconds
33
+ (EngineCore pid=304916) INFO 07-15 18:43:11 [gpu_worker.py:538] Available KV cache memory: 15.82 GiB
34
+ (EngineCore pid=304916) INFO 07-15 18:43:11 [kv_cache_utils.py:2146] GPU KV cache size: 129,584 tokens
35
+ (EngineCore pid=304916) INFO 07-15 18:43:11 [kv_cache_utils.py:2147] Maximum concurrency for 2,048 tokens per request: 63.27x
36
+ (EngineCore pid=304916) INFO 07-15 18:43:12 [cutedsl_warmup.py:97] Skipping CuTeDSL warmup because no compile units were requested.
37
+ (EngineCore pid=304916) INFO 07-15 18:43:12 [jit_monitor.py:73] Kernel JIT monitor activated; monitored JIT compilations during inference will use mode=warn.
38
+ (EngineCore pid=304916) INFO 07-15 18:43:12 [core.py:344] init engine (profile, create kv cache, warmup model) took 2.08 s
39
+ (EngineCore pid=304916) INFO 07-15 18:43:12 [vllm.py:1042] Asynchronous scheduling is enabled.
40
+ (EngineCore pid=304916) WARNING 07-15 18:43:12 [vllm.py:1096] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
41
+ (EngineCore pid=304916) WARNING 07-15 18:43:12 [vllm.py:1144] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
42
+ (EngineCore pid=304916) INFO 07-15 18:43:12 [kernel.py:292] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native'])
43
+ (EngineCore pid=304916) INFO 07-15 18:43:12 [vllm.py:1322] Cudagraph is disabled under eager mode
44
+ (EngineCore pid=304916) INFO 07-15 18:43:12 [compilation.py:312] Enabled custom fusions: norm_quant, act_quant
45
+ (APIServer pid=304785) INFO 07-15 18:43:12 [api_server.py:612] Supported tasks: ['generate']
46
+ (APIServer pid=304785) WARNING 07-15 18:43:12 [__init__.py:36] SECURITY WARNING: Development endpoints are enabled! This should NOT be used in production!
47
+ (APIServer pid=304785) INFO 07-15 18:43:12 [hf.py:548] Detected the chat template content format to be 'string'. You can set `--chat-template-content-format` to override this.
48
+ (APIServer pid=304785) INFO 07-15 18:43:12 [api_server.py:616] Starting vLLM server on http://127.0.0.1:8377
49
+ (APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:37] Available routes are:
50
+ (APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /openapi.json, Methods: HEAD, GET
51
+ (APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /docs, Methods: HEAD, GET
52
+ (APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /docs/oauth2-redirect, Methods: HEAD, GET
53
+ (APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /redoc, Methods: HEAD, GET
54
+ (APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /load, Methods: GET
55
+ (APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /version, Methods: GET
56
+ (APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /health, Methods: GET
57
+ (APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /metrics, Methods: GET
58
+ (APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /tokenize, Methods: POST
59
+ (APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /detokenize, Methods: POST
60
+ (APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /v1/models, Methods: GET
61
+ (APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /ping, Methods: GET
62
+ (APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /ping, Methods: POST
63
+ (APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /invocations, Methods: POST
64
+ (APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /reset_prefix_cache, Methods: POST
65
+ (APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /reset_mm_cache, Methods: POST
66
+ (APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /reset_encoder_cache, Methods: POST
67
+ (APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /pause, Methods: POST
68
+ (APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /resume, Methods: POST
69
+ (APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /is_paused, Methods: GET
70
+ (APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /init_weight_transfer_engine, Methods: POST
71
+ (APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /start_weight_update, Methods: POST
72
+ (APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /update_weights, Methods: POST
73
+ (APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /finish_weight_update, Methods: POST
74
+ (APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /get_world_size, Methods: GET
75
+ (APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /collective_rpc, Methods: POST
76
+ (APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /server_info, Methods: GET
77
+ (APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /sleep, Methods: POST
78
+ (APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /wake_up, Methods: POST
79
+ (APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /is_sleeping, Methods: GET
80
+ (APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /v1/chat/completions, Methods: POST
81
+ (APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /v1/chat/completions/batch, Methods: POST
82
+ (APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /v1/responses, Methods: POST
83
+ (APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /v1/responses/{response_id}, Methods: GET
84
+ (APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /v1/responses/{response_id}/cancel, Methods: POST
85
+ (APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /v1/completions, Methods: POST
86
+ (APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /v1/messages, Methods: POST
87
+ (APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /v1/messages/count_tokens, Methods: POST
88
+ (APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /generative_scoring, Methods: POST
89
+ (APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /scale_elastic_ep, Methods: POST
90
+ (APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /is_scaling_elastic_ep, Methods: POST
91
+ (APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /v1/chat/completions/render, Methods: POST
92
+ (APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /v1/completions/render, Methods: POST
93
+ (APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /v1/chat/completions/derender, Methods: POST
94
+ (APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /v1/completions/derender, Methods: POST
95
+ (APIServer pid=304785) INFO 07-15 18:43:12 [launcher.py:46] Route: /inference/v1/generate, Methods: POST
96
+ (APIServer pid=304785) INFO: Started server process [304785]
97
+ (APIServer pid=304785) INFO: Waiting for application startup.
98
+ (APIServer pid=304785) INFO: Application startup complete.
99
+ (APIServer pid=304785) INFO 07-15 18:43:22 [launcher.py:100] [shutdown] API server: shutdown triggered
100
+ (APIServer pid=304785) INFO 07-15 18:43:22 [launcher.py:116] [shutdown] API server: stopping engine client mode=abort timeout=0s
101
+ (APIServer pid=304785) INFO 07-15 18:43:22 [core_client.py:655] [shutdown] MPClient: start timeout=0s
102
+ (APIServer pid=304785) INFO 07-15 18:43:22 [core_client.py:657] [shutdown] MPClient: stopping engine manager
103
+ (APIServer pid=304785) WARNING 07-15 18:43:22 [utils.py:626] [shutdown] Process manager: force killing remaining processes count=1
104
+ (EngineCore pid=304916) INFO 07-15 18:43:22 [core.py:1214] [shutdown] EngineCore: trigger received signal=SIGTERM
105
+ (EngineCore pid=304916) INFO 07-15 18:43:22 [core.py:1333] [shutdown] EngineCore: start mode=abort timeout=0s
106
+ (EngineCore pid=304916) INFO 07-15 18:43:22 [core.py:1364] [shutdown] EngineCore: request processing complete; starting resource teardown
107
+ (EngineCore pid=304916) INFO 07-15 18:43:22 [core.py:1227] [shutdown] EngineCore: exiting busy loop
108
+ (APIServer pid=304785) INFO 07-15 18:43:22 [core_client.py:659] [shutdown] MPClient: engine manager stopped
109
+ (APIServer pid=304785) INFO 07-15 18:43:22 [core_client.py:660] [shutdown] MPClient: cleaning up background resources
110
+ (APIServer pid=304785) INFO 07-15 18:43:22 [core_client.py:662] [shutdown] MPClient: complete
111
+ (APIServer pid=304785) INFO 07-15 18:43:22 [launcher.py:125] [shutdown] API server: engine client stopped
112
+ (APIServer pid=304785) INFO 07-15 18:43:22 [launcher.py:128] [shutdown] API server: signalling HTTP server shutdown
113
+ (APIServer pid=304785) INFO 07-15 18:43:22 [launcher.py:149] [shutdown] API server: shutting down FastAPI HTTP server
114
+ (APIServer pid=304785) INFO: Shutting down
115
+ (APIServer pid=304785) INFO: Waiting for application shutdown.
116
+ (APIServer pid=304785) INFO: Application shutdown complete.
117
+ /home/henry/.local/share/uv/python/cpython-3.12.12-linux-x86_64-gnu/lib/python3.12/multiprocessing/resource_tracker.py:279: UserWarning: resource_tracker: There appear to be 1 leaked semaphore objects to clean up at shutdown
118
+ warnings.warn('resource_tracker: There appear to be %d '
119
+ (APIServer pid=307152) INFO 07-15 18:49:55 [api_utils.py:339]
120
+ (APIServer pid=307152) INFO 07-15 18:49:55 [api_utils.py:339] β–ˆ β–ˆ β–ˆβ–„ β–„β–ˆ
121
+ (APIServer pid=307152) INFO 07-15 18:49:55 [api_utils.py:339] β–„β–„ β–„β–ˆ β–ˆ β–ˆ β–ˆ β–€β–„β–€ β–ˆ version 0.25.0
122
+ (APIServer pid=307152) INFO 07-15 18:49:55 [api_utils.py:339] β–ˆβ–„β–ˆβ–€ β–ˆ β–ˆ β–ˆ β–ˆ model outputs/healed/healing_breadth/glean_math_keep25_seed1224/step0050
123
+ (APIServer pid=307152) INFO 07-15 18:49:55 [api_utils.py:339] β–€β–€ β–€β–€β–€β–€β–€ β–€β–€β–€β–€β–€ β–€ β–€
124
+ (APIServer pid=307152) INFO 07-15 18:49:55 [api_utils.py:339]
125
+ (APIServer pid=307152) INFO 07-15 18:49:55 [api_utils.py:273] non-default args: {'model_tag': 'outputs/healed/healing_breadth/glean_math_keep25_seed1224/step0050', 'host': '127.0.0.1', 'port': 8377, 'model': 'outputs/healed/healing_breadth/glean_math_keep25_seed1224/step0050', 'max_model_len': 2048, 'enforce_eager': True, 'served_model_name': ['student'], 'gpu_memory_utilization': 0.85}
126
+ (APIServer pid=307152) INFO 07-15 18:49:55 [model.py:619] Resolved architecture: PrunedOlmoeForCausalLM
127
+ (APIServer pid=307152) INFO 07-15 18:49:55 [model.py:1776] Using max model len 2048
128
+ (APIServer pid=307152) INFO 07-15 18:49:55 [vllm.py:1042] Asynchronous scheduling is enabled.
129
+ (APIServer pid=307152) WARNING 07-15 18:49:55 [vllm.py:1096] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
130
+ (APIServer pid=307152) WARNING 07-15 18:49:55 [vllm.py:1144] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
131
+ (APIServer pid=307152) INFO 07-15 18:49:55 [kernel.py:292] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native'])
132
+ (APIServer pid=307152) INFO 07-15 18:49:55 [vllm.py:1322] Cudagraph is disabled under eager mode
133
+ (APIServer pid=307152) INFO 07-15 18:49:55 [compilation.py:312] Enabled custom fusions: norm_quant, act_quant
134
+ (EngineCore pid=307273) INFO 07-15 18:50:02 [core.py:114] Initializing a V1 LLM engine (v0.25.0) with config: model='outputs/healed/healing_breadth/glean_math_keep25_seed1224/step0050', speculative_config=None, tokenizer='outputs/healed/healing_breadth/glean_math_keep25_seed1224/step0050', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=False, dtype=torch.bfloat16, max_seq_len=2048, download_dir=None, load_format=auto, tensor_parallel_size=1, pipeline_parallel_size=1, data_parallel_size=1, decode_context_parallel_size=1, dcp_comm_backend=ag_rs, disable_custom_all_reduce=False, quantization=None, quantization_config=None, enforce_eager=True, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False, jit_monitor_mode='warn', jit_monitor_verbose=False), seed=0, served_model_name=student, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'mode': <CompilationMode.NONE: 0>, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'ir_enable_torch_wrap': False, 'splitting_ops': [], 'compile_mm_encoder': False, 'cudagraph_mm_encoder': False, 'encoder_cudagraph_token_budgets': [], 'encoder_cudagraph_max_vision_items_per_batch': 0, 'encoder_cudagraph_max_frames_per_batch': None, 'compile_sizes': [], 'compile_ranges_endpoints': [2048], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'size_asserts': False, 'alignment_asserts': False, 'scalar_asserts': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': <CUDAGraphMode.NONE: 0>, 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': True, 'fuse_act_quant': True, 'fuse_attn_quant': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False, 'fuse_rope_kvcache_cat_mla': False, 'fuse_act_padding': False}, 'max_cudagraph_capture_size': 0, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': False, 'static_all_moe_layers': []}, kernel_config=KernelConfig(ir_op_priority=IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native']), enable_flashinfer_autotune=True, enable_cutedsl_warmup=True, moe_backend='auto', linear_backend='auto')
135
+ (EngineCore pid=307273) INFO 07-15 18:50:03 [parallel_state.py:1607] world_size=1 rank=0 local_rank=0 distributed_init_method=tcp://192.168.0.15:56251 backend=nccl
136
+ (EngineCore pid=307273) INFO 07-15 18:50:03 [parallel_state.py:1942] rank 0 in world size 1 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank 0, EPLB rank N/A
137
+ (EngineCore pid=307273) INFO 07-15 18:50:03 [topk_topp_sampler.py:55] Using FlashInfer for top-p & top-k sampling.
138
+ (EngineCore pid=307273) INFO 07-15 18:50:03 [gpu_model_runner.py:5209] Starting to load model outputs/healed/healing_breadth/glean_math_keep25_seed1224/step0050...
139
+ (EngineCore pid=307273) INFO 07-15 18:50:04 [cuda.py:476] Using FLASH_ATTN attention backend out of potential backends: ['FLASH_ATTN', 'FLASHINFER', 'TRITON_ATTN', 'FLEX_ATTENTION'].
140
+ (EngineCore pid=307273) INFO 07-15 18:50:04 [flash_attn.py:718] Using FlashAttention version 2
141
+ (EngineCore pid=307273) /home/henry/.cache/glean/megablocks-variable-93a1479bc15b/megablocks/grouped_gemm_util.py:10: UserWarning: Grouped GEMM not available.
142
+ (EngineCore pid=307273) warnings.warn('Grouped GEMM not available.')
143
+ (EngineCore pid=307273) INFO 07-15 18:50:04 [weight_utils.py:849] Filesystem type for checkpoints: EXT4. Checkpoint size: 3.89 GiB. Available RAM: 109.43 GiB.
144
+ (EngineCore pid=307273) INFO 07-15 18:50:04 [weight_utils.py:872] Auto-prefetch is disabled because the filesystem (EXT4) is not a recognized network FS (NFS/Lustre). If you want to force prefetching, start vLLM with --safetensors-load-strategy=prefetch.
145
+ (EngineCore pid=307273)
146
+ (EngineCore pid=307273)
147
+ (EngineCore pid=307273)
148
+ (EngineCore pid=307273)
149
+ (EngineCore pid=307273) INFO 07-15 18:50:06 [default_loader.py:430] Loading weights took 2.49 seconds
150
+ (EngineCore pid=307273) INFO 07-15 18:50:07 [gpu_model_runner.py:5306] Model loading took 3.89 GiB memory and 2.678073 seconds
151
+ (EngineCore pid=307273) INFO 07-15 18:50:08 [gpu_worker.py:538] Available KV cache memory: 15.82 GiB
152
+ (EngineCore pid=307273) INFO 07-15 18:50:08 [kv_cache_utils.py:2146] GPU KV cache size: 129,584 tokens
153
+ (EngineCore pid=307273) INFO 07-15 18:50:08 [kv_cache_utils.py:2147] Maximum concurrency for 2,048 tokens per request: 63.27x
154
+ (EngineCore pid=307273) INFO 07-15 18:50:09 [cutedsl_warmup.py:97] Skipping CuTeDSL warmup because no compile units were requested.
155
+ (EngineCore pid=307273) INFO 07-15 18:50:09 [jit_monitor.py:73] Kernel JIT monitor activated; monitored JIT compilations during inference will use mode=warn.
156
+ (EngineCore pid=307273) INFO 07-15 18:50:09 [core.py:344] init engine (profile, create kv cache, warmup model) took 2.07 s
157
+ (EngineCore pid=307273) INFO 07-15 18:50:09 [vllm.py:1042] Asynchronous scheduling is enabled.
158
+ (EngineCore pid=307273) WARNING 07-15 18:50:09 [vllm.py:1096] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
159
+ (EngineCore pid=307273) WARNING 07-15 18:50:09 [vllm.py:1144] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
160
+ (EngineCore pid=307273) INFO 07-15 18:50:09 [kernel.py:292] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native'])
161
+ (EngineCore pid=307273) INFO 07-15 18:50:09 [vllm.py:1322] Cudagraph is disabled under eager mode
162
+ (EngineCore pid=307273) INFO 07-15 18:50:09 [compilation.py:312] Enabled custom fusions: norm_quant, act_quant
163
+ (APIServer pid=307152) INFO 07-15 18:50:09 [api_server.py:612] Supported tasks: ['generate']
164
+ (APIServer pid=307152) WARNING 07-15 18:50:09 [__init__.py:36] SECURITY WARNING: Development endpoints are enabled! This should NOT be used in production!
165
+ (APIServer pid=307152) INFO 07-15 18:50:09 [hf.py:548] Detected the chat template content format to be 'string'. You can set `--chat-template-content-format` to override this.
166
+ (APIServer pid=307152) INFO 07-15 18:50:09 [api_server.py:616] Starting vLLM server on http://127.0.0.1:8377
167
+ (APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:37] Available routes are:
168
+ (APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /openapi.json, Methods: HEAD, GET
169
+ (APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /docs, Methods: HEAD, GET
170
+ (APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /docs/oauth2-redirect, Methods: HEAD, GET
171
+ (APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /redoc, Methods: HEAD, GET
172
+ (APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /load, Methods: GET
173
+ (APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /version, Methods: GET
174
+ (APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /health, Methods: GET
175
+ (APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /metrics, Methods: GET
176
+ (APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /tokenize, Methods: POST
177
+ (APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /detokenize, Methods: POST
178
+ (APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /v1/models, Methods: GET
179
+ (APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /ping, Methods: GET
180
+ (APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /ping, Methods: POST
181
+ (APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /invocations, Methods: POST
182
+ (APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /reset_prefix_cache, Methods: POST
183
+ (APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /reset_mm_cache, Methods: POST
184
+ (APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /reset_encoder_cache, Methods: POST
185
+ (APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /pause, Methods: POST
186
+ (APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /resume, Methods: POST
187
+ (APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /is_paused, Methods: GET
188
+ (APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /init_weight_transfer_engine, Methods: POST
189
+ (APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /start_weight_update, Methods: POST
190
+ (APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /update_weights, Methods: POST
191
+ (APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /finish_weight_update, Methods: POST
192
+ (APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /get_world_size, Methods: GET
193
+ (APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /collective_rpc, Methods: POST
194
+ (APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /server_info, Methods: GET
195
+ (APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /sleep, Methods: POST
196
+ (APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /wake_up, Methods: POST
197
+ (APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /is_sleeping, Methods: GET
198
+ (APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /v1/chat/completions, Methods: POST
199
+ (APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /v1/chat/completions/batch, Methods: POST
200
+ (APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /v1/responses, Methods: POST
201
+ (APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /v1/responses/{response_id}, Methods: GET
202
+ (APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /v1/responses/{response_id}/cancel, Methods: POST
203
+ (APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /v1/completions, Methods: POST
204
+ (APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /v1/messages, Methods: POST
205
+ (APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /v1/messages/count_tokens, Methods: POST
206
+ (APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /generative_scoring, Methods: POST
207
+ (APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /scale_elastic_ep, Methods: POST
208
+ (APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /is_scaling_elastic_ep, Methods: POST
209
+ (APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /v1/chat/completions/render, Methods: POST
210
+ (APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /v1/completions/render, Methods: POST
211
+ (APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /v1/chat/completions/derender, Methods: POST
212
+ (APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /v1/completions/derender, Methods: POST
213
+ (APIServer pid=307152) INFO 07-15 18:50:09 [launcher.py:46] Route: /inference/v1/generate, Methods: POST
214
+ (APIServer pid=307152) INFO: Started server process [307152]
215
+ (APIServer pid=307152) INFO: Waiting for application startup.
216
+ (APIServer pid=307152) INFO: Application startup complete.
217
+ (APIServer pid=307152) INFO: 127.0.0.1:37086 - "GET /health HTTP/1.1" 200 OK
218
+ (EngineCore pid=307273) WARNING 07-15 18:50:11 [jit_monitor.py:129] Triton kernel JIT compilation during inference: _build_route_rows. This causes a latency spike; consider extending warmup to cover this shape/config.
219
+ (APIServer pid=307152) INFO 07-15 18:50:20 [loggers.py:273] Engine 000: Avg prompt throughput: 1728.4 tokens/s, Avg generation throughput: 2268.5 tokens/s, Running: 256 reqs, Waiting: 1059 reqs, GPU KV cache usage: 33.5%, Prefix cache hit rate: 94.1%
220
+ (APIServer pid=307152) INFO 07-15 18:50:30 [loggers.py:273] Engine 000: Avg prompt throughput: 129.1 tokens/s, Avg generation throughput: 3043.4 tokens/s, Running: 256 reqs, Waiting: 1038 reqs, GPU KV cache usage: 54.3%, Prefix cache hit rate: 94.1%
221
+ (APIServer pid=307152) INFO 07-15 18:50:40 [loggers.py:273] Engine 000: Avg prompt throughput: 112.1 tokens/s, Avg generation throughput: 2916.1 tokens/s, Running: 256 reqs, Waiting: 1021 reqs, GPU KV cache usage: 73.4%, Prefix cache hit rate: 94.2%
222
+ (APIServer pid=307152) INFO 07-15 18:50:50 [loggers.py:273] Engine 000: Avg prompt throughput: 72.7 tokens/s, Avg generation throughput: 2788.6 tokens/s, Running: 256 reqs, Waiting: 1010 reqs, GPU KV cache usage: 92.2%, Prefix cache hit rate: 94.2%
223
+ (APIServer pid=307152) INFO 07-15 18:51:00 [loggers.py:273] Engine 000: Avg prompt throughput: 1368.9 tokens/s, Avg generation throughput: 2629.0 tokens/s, Running: 256 reqs, Waiting: 799 reqs, GPU KV cache usage: 30.0%, Prefix cache hit rate: 94.3%
224
+ (APIServer pid=307152) INFO 07-15 18:51:10 [loggers.py:273] Engine 000: Avg prompt throughput: 142.4 tokens/s, Avg generation throughput: 3069.3 tokens/s, Running: 256 reqs, Waiting: 779 reqs, GPU KV cache usage: 48.9%, Prefix cache hit rate: 94.3%
225
+ (APIServer pid=307152) INFO 07-15 18:51:20 [loggers.py:273] Engine 000: Avg prompt throughput: 174.0 tokens/s, Avg generation throughput: 2966.1 tokens/s, Running: 256 reqs, Waiting: 750 reqs, GPU KV cache usage: 63.9%, Prefix cache hit rate: 94.3%
226
+ (APIServer pid=307152) INFO 07-15 18:51:30 [loggers.py:273] Engine 000: Avg prompt throughput: 197.2 tokens/s, Avg generation throughput: 2838.7 tokens/s, Running: 256 reqs, Waiting: 721 reqs, GPU KV cache usage: 76.8%, Prefix cache hit rate: 94.3%
227
+ (APIServer pid=307152) INFO 07-15 18:51:40 [loggers.py:273] Engine 000: Avg prompt throughput: 134.4 tokens/s, Avg generation throughput: 2787.9 tokens/s, Running: 255 reqs, Waiting: 701 reqs, GPU KV cache usage: 91.1%, Prefix cache hit rate: 94.3%
228
+ (APIServer pid=307152) INFO 07-15 18:51:50 [loggers.py:273] Engine 000: Avg prompt throughput: 1103.7 tokens/s, Avg generation throughput: 2850.0 tokens/s, Running: 255 reqs, Waiting: 531 reqs, GPU KV cache usage: 47.4%, Prefix cache hit rate: 94.4%
229
+ (APIServer pid=307152) INFO 07-15 18:52:00 [loggers.py:273] Engine 000: Avg prompt throughput: 322.1 tokens/s, Avg generation throughput: 2939.5 tokens/s, Running: 256 reqs, Waiting: 487 reqs, GPU KV cache usage: 57.3%, Prefix cache hit rate: 94.3%
230
+ (APIServer pid=307152) INFO 07-15 18:52:10 [loggers.py:273] Engine 000: Avg prompt throughput: 242.7 tokens/s, Avg generation throughput: 2914.0 tokens/s, Running: 255 reqs, Waiting: 448 reqs, GPU KV cache usage: 66.5%, Prefix cache hit rate: 94.4%
231
+ (APIServer pid=307152) INFO 07-15 18:52:20 [loggers.py:273] Engine 000: Avg prompt throughput: 270.3 tokens/s, Avg generation throughput: 2836.3 tokens/s, Running: 256 reqs, Waiting: 403 reqs, GPU KV cache usage: 75.4%, Prefix cache hit rate: 94.4%
232
+ (APIServer pid=307152) INFO 07-15 18:52:30 [loggers.py:273] Engine 000: Avg prompt throughput: 1025.0 tokens/s, Avg generation throughput: 2647.8 tokens/s, Running: 256 reqs, Waiting: 257 reqs, GPU KV cache usage: 40.7%, Prefix cache hit rate: 94.4%
233
+ (APIServer pid=307152) INFO 07-15 18:52:40 [loggers.py:273] Engine 000: Avg prompt throughput: 196.8 tokens/s, Avg generation throughput: 2992.2 tokens/s, Running: 255 reqs, Waiting: 229 reqs, GPU KV cache usage: 55.4%, Prefix cache hit rate: 94.4%
234
+ (APIServer pid=307152) INFO 07-15 18:52:50 [loggers.py:273] Engine 000: Avg prompt throughput: 296.9 tokens/s, Avg generation throughput: 2938.7 tokens/s, Running: 256 reqs, Waiting: 182 reqs, GPU KV cache usage: 63.4%, Prefix cache hit rate: 94.4%
235
+ (APIServer pid=307152) INFO 07-15 18:53:00 [loggers.py:273] Engine 000: Avg prompt throughput: 268.7 tokens/s, Avg generation throughput: 2836.8 tokens/s, Running: 256 reqs, Waiting: 140 reqs, GPU KV cache usage: 71.0%, Prefix cache hit rate: 94.4%
236
+ (APIServer pid=307152) INFO 07-15 18:53:10 [loggers.py:273] Engine 000: Avg prompt throughput: 354.0 tokens/s, Avg generation throughput: 2836.6 tokens/s, Running: 255 reqs, Waiting: 90 reqs, GPU KV cache usage: 75.4%, Prefix cache hit rate: 94.4%
237
+ (APIServer pid=307152) INFO 07-15 18:53:20 [loggers.py:273] Engine 000: Avg prompt throughput: 616.3 tokens/s, Avg generation throughput: 2818.2 tokens/s, Running: 226 reqs, Waiting: 0 reqs, GPU KV cache usage: 51.0%, Prefix cache hit rate: 94.4%
238
+ (APIServer pid=307152) INFO 07-15 18:53:30 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 2701.6 tokens/s, Running: 164 reqs, Waiting: 0 reqs, GPU KV cache usage: 47.0%, Prefix cache hit rate: 94.4%
239
+ (APIServer pid=307152) INFO 07-15 18:53:40 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 2325.4 tokens/s, Running: 107 reqs, Waiting: 0 reqs, GPU KV cache usage: 40.3%, Prefix cache hit rate: 94.4%
240
+ (APIServer pid=307152) INFO: 127.0.0.1:37094 - "POST /v1/completions HTTP/1.1" 200 OK
241
+ (EngineCore pid=307273) INFO 07-15 18:53:47 [core.py:1214] [shutdown] EngineCore: trigger received signal=SIGTERM
242
+ (APIServer pid=307152) INFO 07-15 18:53:47 [launcher.py:100] [shutdown] API server: shutdown triggered
243
+ (APIServer pid=307152) INFO 07-15 18:53:47 [launcher.py:116] [shutdown] API server: stopping engine client mode=abort timeout=0s
244
+ (EngineCore pid=307273) INFO 07-15 18:53:47 [core.py:1333] [shutdown] EngineCore: start mode=abort timeout=0s
245
+ (EngineCore pid=307273) INFO 07-15 18:53:47 [core.py:1364] [shutdown] EngineCore: request processing complete; starting resource teardown
246
+ (EngineCore pid=307273) INFO 07-15 18:53:47 [core.py:1227] [shutdown] EngineCore: exiting busy loop
247
+ (APIServer pid=307152) INFO 07-15 18:53:47 [core_client.py:655] [shutdown] MPClient: start timeout=0s
248
+ (APIServer pid=307152) INFO 07-15 18:53:47 [core_client.py:657] [shutdown] MPClient: stopping engine manager
249
+ (APIServer pid=307152) WARNING 07-15 18:53:47 [utils.py:626] [shutdown] Process manager: force killing remaining processes count=1
250
+ (APIServer pid=307152) INFO 07-15 18:53:47 [core_client.py:659] [shutdown] MPClient: engine manager stopped
251
+ (APIServer pid=307152) INFO 07-15 18:53:47 [core_client.py:660] [shutdown] MPClient: cleaning up background resources
252
+ (APIServer pid=307152) INFO 07-15 18:53:47 [core_client.py:662] [shutdown] MPClient: complete
253
+ (APIServer pid=307152) INFO 07-15 18:53:47 [launcher.py:125] [shutdown] API server: engine client stopped
254
+ (APIServer pid=307152) INFO 07-15 18:53:47 [launcher.py:128] [shutdown] API server: signalling HTTP server shutdown
255
+ (APIServer pid=307152) INFO 07-15 18:53:47 [launcher.py:149] [shutdown] API server: shutting down FastAPI HTTP server
256
+ (APIServer pid=307152) INFO: Shutting down
257
+ (APIServer pid=307152) INFO: Waiting for application shutdown.
258
+ (APIServer pid=307152) INFO: Application shutdown complete.
259
+ /home/henry/.local/share/uv/python/cpython-3.12.12-linux-x86_64-gnu/lib/python3.12/multiprocessing/resource_tracker.py:279: UserWarning: resource_tracker: There appear to be 1 leaked semaphore objects to clean up at shutdown
260
+ warnings.warn('resource_tracker: There appear to be %d '
evals/healing_breadth/glean_math_keep75_oneshot_chat.json ADDED
The diff for this file is too large to render. See raw diff
 
evals/healing_breadth/glean_math_keep75_oneshot_chat.json.server.log ADDED
@@ -0,0 +1,130 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ (APIServer pid=239542) INFO 07-15 06:24:31 [api_utils.py:339]
2
+ (APIServer pid=239542) INFO 07-15 06:24:31 [api_utils.py:339] β–ˆ β–ˆ β–ˆβ–„ β–„β–ˆ
3
+ (APIServer pid=239542) INFO 07-15 06:24:31 [api_utils.py:339] β–„β–„ β–„β–ˆ β–ˆ β–ˆ β–ˆ β–€β–„β–€ β–ˆ version 0.25.0
4
+ (APIServer pid=239542) INFO 07-15 06:24:31 [api_utils.py:339] β–ˆβ–„β–ˆβ–€ β–ˆ β–ˆ β–ˆ β–ˆ model outputs/pruned/glean-0125inst-math-keep75
5
+ (APIServer pid=239542) INFO 07-15 06:24:31 [api_utils.py:339] β–€β–€ β–€β–€β–€β–€β–€ β–€β–€β–€β–€β–€ β–€ β–€
6
+ (APIServer pid=239542) INFO 07-15 06:24:31 [api_utils.py:339]
7
+ (APIServer pid=239542) INFO 07-15 06:24:31 [api_utils.py:273] non-default args: {'model_tag': 'outputs/pruned/glean-0125inst-math-keep75', 'host': '127.0.0.1', 'port': 8377, 'model': 'outputs/pruned/glean-0125inst-math-keep75', 'max_model_len': 2048, 'enforce_eager': True, 'served_model_name': ['student'], 'gpu_memory_utilization': 0.85}
8
+ (APIServer pid=239542) INFO 07-15 06:24:31 [model.py:619] Resolved architecture: OlmoeForCausalLM
9
+ (APIServer pid=239542) INFO 07-15 06:24:31 [model.py:1776] Using max model len 2048
10
+ (APIServer pid=239542) INFO 07-15 06:24:31 [vllm.py:1042] Asynchronous scheduling is enabled.
11
+ (APIServer pid=239542) WARNING 07-15 06:24:31 [vllm.py:1096] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
12
+ (APIServer pid=239542) WARNING 07-15 06:24:31 [vllm.py:1144] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
13
+ (APIServer pid=239542) INFO 07-15 06:24:31 [kernel.py:292] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native'])
14
+ (APIServer pid=239542) INFO 07-15 06:24:32 [vllm.py:1322] Cudagraph is disabled under eager mode
15
+ (APIServer pid=239542) INFO 07-15 06:24:32 [compilation.py:312] Enabled custom fusions: norm_quant, act_quant
16
+ (EngineCore pid=239661) INFO 07-15 06:24:39 [core.py:114] Initializing a V1 LLM engine (v0.25.0) with config: model='outputs/pruned/glean-0125inst-math-keep75', speculative_config=None, tokenizer='outputs/pruned/glean-0125inst-math-keep75', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=False, dtype=torch.bfloat16, max_seq_len=2048, download_dir=None, load_format=auto, tensor_parallel_size=1, pipeline_parallel_size=1, data_parallel_size=1, decode_context_parallel_size=1, dcp_comm_backend=ag_rs, disable_custom_all_reduce=False, quantization=None, quantization_config=None, enforce_eager=True, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False, jit_monitor_mode='warn', jit_monitor_verbose=False), seed=0, served_model_name=student, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'mode': <CompilationMode.NONE: 0>, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'ir_enable_torch_wrap': False, 'splitting_ops': [], 'compile_mm_encoder': False, 'cudagraph_mm_encoder': False, 'encoder_cudagraph_token_budgets': [], 'encoder_cudagraph_max_vision_items_per_batch': 0, 'encoder_cudagraph_max_frames_per_batch': None, 'compile_sizes': [], 'compile_ranges_endpoints': [2048], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'size_asserts': False, 'alignment_asserts': False, 'scalar_asserts': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': <CUDAGraphMode.NONE: 0>, 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': True, 'fuse_act_quant': True, 'fuse_attn_quant': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False, 'fuse_rope_kvcache_cat_mla': False, 'fuse_act_padding': False}, 'max_cudagraph_capture_size': 0, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': False, 'static_all_moe_layers': []}, kernel_config=KernelConfig(ir_op_priority=IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native']), enable_flashinfer_autotune=True, enable_cutedsl_warmup=True, moe_backend='auto', linear_backend='auto')
17
+ (EngineCore pid=239661) INFO 07-15 06:24:39 [parallel_state.py:1607] world_size=1 rank=0 local_rank=0 distributed_init_method=tcp://192.168.0.15:38341 backend=nccl
18
+ (EngineCore pid=239661) INFO 07-15 06:24:39 [parallel_state.py:1942] rank 0 in world size 1 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank 0, EPLB rank N/A
19
+ (EngineCore pid=239661) INFO 07-15 06:24:40 [topk_topp_sampler.py:55] Using FlashInfer for top-p & top-k sampling.
20
+ (EngineCore pid=239661) INFO 07-15 06:24:40 [gpu_model_runner.py:5209] Starting to load model outputs/pruned/glean-0125inst-math-keep75...
21
+ (EngineCore pid=239661) INFO 07-15 06:24:41 [cuda.py:476] Using FLASH_ATTN attention backend out of potential backends: ['FLASH_ATTN', 'FLASHINFER', 'TRITON_ATTN', 'FLEX_ATTENTION'].
22
+ (EngineCore pid=239661) INFO 07-15 06:24:41 [flash_attn.py:718] Using FlashAttention version 2
23
+ (EngineCore pid=239661) /home/henry/.cache/glean/megablocks-variable-93a1479bc15b/megablocks/grouped_gemm_util.py:10: UserWarning: Grouped GEMM not available.
24
+ (EngineCore pid=239661) warnings.warn('Grouped GEMM not available.')
25
+ (EngineCore pid=239661) INFO 07-15 06:24:41 [weight_utils.py:849] Filesystem type for checkpoints: EXT4. Checkpoint size: 9.89 GiB. Available RAM: 61.30 GiB.
26
+ (EngineCore pid=239661) INFO 07-15 06:24:41 [weight_utils.py:872] Auto-prefetch is disabled because the filesystem (EXT4) is not a recognized network FS (NFS/Lustre). If you want to force prefetching, start vLLM with --safetensors-load-strategy=prefetch.
27
+ (EngineCore pid=239661)
28
+ (EngineCore pid=239661)
29
+ (EngineCore pid=239661)
30
+ (EngineCore pid=239661)
31
+ (EngineCore pid=239661)
32
+ (EngineCore pid=239661)
33
+ (EngineCore pid=239661) INFO 07-15 06:24:48 [default_loader.py:430] Loading weights took 7.29 seconds
34
+ (EngineCore pid=239661) INFO 07-15 06:24:48 [gpu_model_runner.py:5306] Model loading took 9.89 GiB memory and 7.475187 seconds
35
+ (EngineCore pid=239661) INFO 07-15 06:24:50 [gpu_worker.py:538] Available KV cache memory: 9.8 GiB
36
+ (EngineCore pid=239661) INFO 07-15 06:24:50 [kv_cache_utils.py:2146] GPU KV cache size: 80,272 tokens
37
+ (EngineCore pid=239661) INFO 07-15 06:24:50 [kv_cache_utils.py:2147] Maximum concurrency for 2,048 tokens per request: 39.20x
38
+ (EngineCore pid=239661) INFO 07-15 06:24:50 [cutedsl_warmup.py:97] Skipping CuTeDSL warmup because no compile units were requested.
39
+ (EngineCore pid=239661) INFO 07-15 06:24:50 [jit_monitor.py:73] Kernel JIT monitor activated; monitored JIT compilations during inference will use mode=warn.
40
+ (EngineCore pid=239661) INFO 07-15 06:24:50 [core.py:344] init engine (profile, create kv cache, warmup model) took 2.09 s
41
+ (EngineCore pid=239661) INFO 07-15 06:24:51 [vllm.py:1042] Asynchronous scheduling is enabled.
42
+ (EngineCore pid=239661) WARNING 07-15 06:24:51 [vllm.py:1096] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
43
+ (EngineCore pid=239661) WARNING 07-15 06:24:51 [vllm.py:1144] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
44
+ (EngineCore pid=239661) INFO 07-15 06:24:51 [kernel.py:292] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native'])
45
+ (EngineCore pid=239661) INFO 07-15 06:24:51 [vllm.py:1322] Cudagraph is disabled under eager mode
46
+ (EngineCore pid=239661) INFO 07-15 06:24:51 [compilation.py:312] Enabled custom fusions: norm_quant, act_quant
47
+ (APIServer pid=239542) INFO 07-15 06:24:51 [api_server.py:612] Supported tasks: ['generate']
48
+ (APIServer pid=239542) WARNING 07-15 06:24:51 [__init__.py:36] SECURITY WARNING: Development endpoints are enabled! This should NOT be used in production!
49
+ (APIServer pid=239542) INFO 07-15 06:24:51 [hf.py:548] Detected the chat template content format to be 'string'. You can set `--chat-template-content-format` to override this.
50
+ (APIServer pid=239542) INFO 07-15 06:24:51 [api_server.py:616] Starting vLLM server on http://127.0.0.1:8377
51
+ (APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:37] Available routes are:
52
+ (APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /openapi.json, Methods: HEAD, GET
53
+ (APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /docs, Methods: HEAD, GET
54
+ (APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /docs/oauth2-redirect, Methods: HEAD, GET
55
+ (APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /redoc, Methods: HEAD, GET
56
+ (APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /load, Methods: GET
57
+ (APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /version, Methods: GET
58
+ (APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /health, Methods: GET
59
+ (APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /metrics, Methods: GET
60
+ (APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /tokenize, Methods: POST
61
+ (APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /detokenize, Methods: POST
62
+ (APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /v1/models, Methods: GET
63
+ (APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /ping, Methods: GET
64
+ (APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /ping, Methods: POST
65
+ (APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /invocations, Methods: POST
66
+ (APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /reset_prefix_cache, Methods: POST
67
+ (APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /reset_mm_cache, Methods: POST
68
+ (APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /reset_encoder_cache, Methods: POST
69
+ (APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /pause, Methods: POST
70
+ (APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /resume, Methods: POST
71
+ (APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /is_paused, Methods: GET
72
+ (APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /init_weight_transfer_engine, Methods: POST
73
+ (APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /start_weight_update, Methods: POST
74
+ (APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /update_weights, Methods: POST
75
+ (APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /finish_weight_update, Methods: POST
76
+ (APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /get_world_size, Methods: GET
77
+ (APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /collective_rpc, Methods: POST
78
+ (APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /server_info, Methods: GET
79
+ (APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /sleep, Methods: POST
80
+ (APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /wake_up, Methods: POST
81
+ (APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /is_sleeping, Methods: GET
82
+ (APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /v1/chat/completions, Methods: POST
83
+ (APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /v1/chat/completions/batch, Methods: POST
84
+ (APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /v1/responses, Methods: POST
85
+ (APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /v1/responses/{response_id}, Methods: GET
86
+ (APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /v1/responses/{response_id}/cancel, Methods: POST
87
+ (APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /v1/completions, Methods: POST
88
+ (APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /v1/messages, Methods: POST
89
+ (APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /v1/messages/count_tokens, Methods: POST
90
+ (APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /generative_scoring, Methods: POST
91
+ (APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /scale_elastic_ep, Methods: POST
92
+ (APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /is_scaling_elastic_ep, Methods: POST
93
+ (APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /v1/chat/completions/render, Methods: POST
94
+ (APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /v1/completions/render, Methods: POST
95
+ (APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /v1/chat/completions/derender, Methods: POST
96
+ (APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /v1/completions/derender, Methods: POST
97
+ (APIServer pid=239542) INFO 07-15 06:24:51 [launcher.py:46] Route: /inference/v1/generate, Methods: POST
98
+ (APIServer pid=239542) INFO: Started server process [239542]
99
+ (APIServer pid=239542) INFO: Waiting for application startup.
100
+ (APIServer pid=239542) INFO: Application startup complete.
101
+ (APIServer pid=239542) INFO: 127.0.0.1:51610 - "GET /health HTTP/1.1" 200 OK
102
+ (EngineCore pid=239661) WARNING 07-15 06:24:52 [jit_monitor.py:129] Triton kernel JIT compilation during inference: _build_route_rows. This causes a latency spike; consider extending warmup to cover this shape/config.
103
+ (APIServer pid=239542) INFO 07-15 06:25:01 [loggers.py:273] Engine 000: Avg prompt throughput: 2621.7 tokens/s, Avg generation throughput: 1945.2 tokens/s, Running: 254 reqs, Waiting: 978 reqs, GPU KV cache usage: 48.5%, Prefix cache hit rate: 93.2%
104
+ (APIServer pid=239542) INFO 07-15 06:25:11 [loggers.py:273] Engine 000: Avg prompt throughput: 1793.6 tokens/s, Avg generation throughput: 2485.0 tokens/s, Running: 251 reqs, Waiting: 747 reqs, GPU KV cache usage: 49.0%, Prefix cache hit rate: 93.3%
105
+ (APIServer pid=239542) INFO 07-15 06:25:21 [loggers.py:273] Engine 000: Avg prompt throughput: 1816.8 tokens/s, Avg generation throughput: 2460.1 tokens/s, Running: 254 reqs, Waiting: 517 reqs, GPU KV cache usage: 50.1%, Prefix cache hit rate: 93.3%
106
+ (APIServer pid=239542) INFO 07-15 06:25:31 [loggers.py:273] Engine 000: Avg prompt throughput: 1842.7 tokens/s, Avg generation throughput: 2460.4 tokens/s, Running: 252 reqs, Waiting: 291 reqs, GPU KV cache usage: 50.3%, Prefix cache hit rate: 93.4%
107
+ (APIServer pid=239542) INFO 07-15 06:25:41 [loggers.py:273] Engine 000: Avg prompt throughput: 1720.0 tokens/s, Avg generation throughput: 2485.7 tokens/s, Running: 256 reqs, Waiting: 71 reqs, GPU KV cache usage: 52.5%, Prefix cache hit rate: 93.4%
108
+ (APIServer pid=239542) INFO 07-15 06:25:51 [loggers.py:273] Engine 000: Avg prompt throughput: 618.3 tokens/s, Avg generation throughput: 2265.6 tokens/s, Running: 64 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.5%, Prefix cache hit rate: 93.4%
109
+ (APIServer pid=239542) INFO 07-15 06:26:01 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 254.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.2%, Prefix cache hit rate: 93.4%
110
+ (APIServer pid=239542) INFO: 127.0.0.1:51612 - "POST /v1/completions HTTP/1.1" 200 OK
111
+ (EngineCore pid=239661) INFO 07-15 06:26:02 [core.py:1214] [shutdown] EngineCore: trigger received signal=SIGTERM
112
+ (APIServer pid=239542) INFO 07-15 06:26:02 [launcher.py:100] [shutdown] API server: shutdown triggered
113
+ (APIServer pid=239542) INFO 07-15 06:26:02 [launcher.py:116] [shutdown] API server: stopping engine client mode=abort timeout=0s
114
+ (EngineCore pid=239661) INFO 07-15 06:26:02 [core.py:1333] [shutdown] EngineCore: start mode=abort timeout=0s
115
+ (EngineCore pid=239661) INFO 07-15 06:26:02 [core.py:1364] [shutdown] EngineCore: request processing complete; starting resource teardown
116
+ (EngineCore pid=239661) INFO 07-15 06:26:02 [core.py:1227] [shutdown] EngineCore: exiting busy loop
117
+ (APIServer pid=239542) INFO 07-15 06:26:02 [core_client.py:655] [shutdown] MPClient: start timeout=0s
118
+ (APIServer pid=239542) INFO 07-15 06:26:02 [core_client.py:657] [shutdown] MPClient: stopping engine manager
119
+ (APIServer pid=239542) WARNING 07-15 06:26:02 [utils.py:626] [shutdown] Process manager: force killing remaining processes count=1
120
+ (APIServer pid=239542) INFO 07-15 06:26:02 [core_client.py:659] [shutdown] MPClient: engine manager stopped
121
+ (APIServer pid=239542) INFO 07-15 06:26:02 [core_client.py:660] [shutdown] MPClient: cleaning up background resources
122
+ (APIServer pid=239542) INFO 07-15 06:26:02 [core_client.py:662] [shutdown] MPClient: complete
123
+ (APIServer pid=239542) INFO 07-15 06:26:02 [launcher.py:125] [shutdown] API server: engine client stopped
124
+ (APIServer pid=239542) INFO 07-15 06:26:02 [launcher.py:128] [shutdown] API server: signalling HTTP server shutdown
125
+ (APIServer pid=239542) INFO 07-15 06:26:02 [launcher.py:149] [shutdown] API server: shutting down FastAPI HTTP server
126
+ (APIServer pid=239542) INFO: Shutting down
127
+ (APIServer pid=239542) INFO: Waiting for application shutdown.
128
+ (APIServer pid=239542) INFO: Application shutdown complete.
129
+ /home/henry/.local/share/uv/python/cpython-3.12.12-linux-x86_64-gnu/lib/python3.12/multiprocessing/resource_tracker.py:279: UserWarning: resource_tracker: There appear to be 1 leaked semaphore objects to clean up at shutdown
130
+ warnings.warn('resource_tracker: There appear to be %d '
evals/healing_breadth/glean_math_keep75_oneshot_raw.json ADDED
The diff for this file is too large to render. See raw diff
 
evals/healing_breadth/glean_math_keep75_oneshot_raw.json.server.log ADDED
@@ -0,0 +1,129 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ (APIServer pid=238699) INFO 07-15 06:22:30 [api_utils.py:339]
2
+ (APIServer pid=238699) INFO 07-15 06:22:30 [api_utils.py:339] β–ˆ β–ˆ β–ˆβ–„ β–„β–ˆ
3
+ (APIServer pid=238699) INFO 07-15 06:22:30 [api_utils.py:339] β–„β–„ β–„β–ˆ β–ˆ β–ˆ β–ˆ β–€β–„β–€ β–ˆ version 0.25.0
4
+ (APIServer pid=238699) INFO 07-15 06:22:30 [api_utils.py:339] β–ˆβ–„β–ˆβ–€ β–ˆ β–ˆ β–ˆ β–ˆ model outputs/pruned/glean-0125inst-math-keep75
5
+ (APIServer pid=238699) INFO 07-15 06:22:30 [api_utils.py:339] β–€β–€ β–€β–€β–€β–€β–€ β–€β–€β–€β–€β–€ β–€ β–€
6
+ (APIServer pid=238699) INFO 07-15 06:22:30 [api_utils.py:339]
7
+ (APIServer pid=238699) INFO 07-15 06:22:30 [api_utils.py:273] non-default args: {'model_tag': 'outputs/pruned/glean-0125inst-math-keep75', 'host': '127.0.0.1', 'port': 8377, 'model': 'outputs/pruned/glean-0125inst-math-keep75', 'max_model_len': 2048, 'enforce_eager': True, 'served_model_name': ['student'], 'gpu_memory_utilization': 0.85}
8
+ (APIServer pid=238699) INFO 07-15 06:22:30 [model.py:619] Resolved architecture: OlmoeForCausalLM
9
+ (APIServer pid=238699) INFO 07-15 06:22:30 [model.py:1776] Using max model len 2048
10
+ (APIServer pid=238699) INFO 07-15 06:22:30 [vllm.py:1042] Asynchronous scheduling is enabled.
11
+ (APIServer pid=238699) WARNING 07-15 06:22:30 [vllm.py:1096] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
12
+ (APIServer pid=238699) WARNING 07-15 06:22:30 [vllm.py:1144] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
13
+ (APIServer pid=238699) INFO 07-15 06:22:30 [kernel.py:292] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native'])
14
+ (APIServer pid=238699) INFO 07-15 06:22:30 [vllm.py:1322] Cudagraph is disabled under eager mode
15
+ (APIServer pid=238699) INFO 07-15 06:22:30 [compilation.py:312] Enabled custom fusions: norm_quant, act_quant
16
+ (EngineCore pid=238829) INFO 07-15 06:22:38 [core.py:114] Initializing a V1 LLM engine (v0.25.0) with config: model='outputs/pruned/glean-0125inst-math-keep75', speculative_config=None, tokenizer='outputs/pruned/glean-0125inst-math-keep75', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=False, dtype=torch.bfloat16, max_seq_len=2048, download_dir=None, load_format=auto, tensor_parallel_size=1, pipeline_parallel_size=1, data_parallel_size=1, decode_context_parallel_size=1, dcp_comm_backend=ag_rs, disable_custom_all_reduce=False, quantization=None, quantization_config=None, enforce_eager=True, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False, jit_monitor_mode='warn', jit_monitor_verbose=False), seed=0, served_model_name=student, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'mode': <CompilationMode.NONE: 0>, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'ir_enable_torch_wrap': False, 'splitting_ops': [], 'compile_mm_encoder': False, 'cudagraph_mm_encoder': False, 'encoder_cudagraph_token_budgets': [], 'encoder_cudagraph_max_vision_items_per_batch': 0, 'encoder_cudagraph_max_frames_per_batch': None, 'compile_sizes': [], 'compile_ranges_endpoints': [2048], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'size_asserts': False, 'alignment_asserts': False, 'scalar_asserts': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': <CUDAGraphMode.NONE: 0>, 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': True, 'fuse_act_quant': True, 'fuse_attn_quant': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False, 'fuse_rope_kvcache_cat_mla': False, 'fuse_act_padding': False}, 'max_cudagraph_capture_size': 0, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': False, 'static_all_moe_layers': []}, kernel_config=KernelConfig(ir_op_priority=IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native']), enable_flashinfer_autotune=True, enable_cutedsl_warmup=True, moe_backend='auto', linear_backend='auto')
17
+ (EngineCore pid=238829) INFO 07-15 06:22:38 [parallel_state.py:1607] world_size=1 rank=0 local_rank=0 distributed_init_method=tcp://192.168.0.15:59451 backend=nccl
18
+ (EngineCore pid=238829) INFO 07-15 06:22:38 [parallel_state.py:1942] rank 0 in world size 1 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank 0, EPLB rank N/A
19
+ (EngineCore pid=238829) INFO 07-15 06:22:39 [topk_topp_sampler.py:55] Using FlashInfer for top-p & top-k sampling.
20
+ (EngineCore pid=238829) INFO 07-15 06:22:39 [gpu_model_runner.py:5209] Starting to load model outputs/pruned/glean-0125inst-math-keep75...
21
+ (EngineCore pid=238829) INFO 07-15 06:22:40 [cuda.py:476] Using FLASH_ATTN attention backend out of potential backends: ['FLASH_ATTN', 'FLASHINFER', 'TRITON_ATTN', 'FLEX_ATTENTION'].
22
+ (EngineCore pid=238829) INFO 07-15 06:22:40 [flash_attn.py:718] Using FlashAttention version 2
23
+ (EngineCore pid=238829) /home/henry/.cache/glean/megablocks-variable-93a1479bc15b/megablocks/grouped_gemm_util.py:10: UserWarning: Grouped GEMM not available.
24
+ (EngineCore pid=238829) warnings.warn('Grouped GEMM not available.')
25
+ (EngineCore pid=238829) INFO 07-15 06:22:40 [weight_utils.py:849] Filesystem type for checkpoints: EXT4. Checkpoint size: 9.89 GiB. Available RAM: 61.21 GiB.
26
+ (EngineCore pid=238829) INFO 07-15 06:22:40 [weight_utils.py:872] Auto-prefetch is disabled because the filesystem (EXT4) is not a recognized network FS (NFS/Lustre). If you want to force prefetching, start vLLM with --safetensors-load-strategy=prefetch.
27
+ (EngineCore pid=238829)
28
+ (EngineCore pid=238829)
29
+ (EngineCore pid=238829)
30
+ (EngineCore pid=238829)
31
+ (EngineCore pid=238829)
32
+ (EngineCore pid=238829)
33
+ (EngineCore pid=238829) INFO 07-15 06:22:55 [default_loader.py:430] Loading weights took 15.19 seconds
34
+ (EngineCore pid=238829) INFO 07-15 06:22:55 [gpu_model_runner.py:5306] Model loading took 9.89 GiB memory and 15.389861 seconds
35
+ (EngineCore pid=238829) INFO 07-15 06:22:57 [gpu_worker.py:538] Available KV cache memory: 9.8 GiB
36
+ (EngineCore pid=238829) INFO 07-15 06:22:57 [kv_cache_utils.py:2146] GPU KV cache size: 80,272 tokens
37
+ (EngineCore pid=238829) INFO 07-15 06:22:57 [kv_cache_utils.py:2147] Maximum concurrency for 2,048 tokens per request: 39.20x
38
+ (EngineCore pid=238829) INFO 07-15 06:22:57 [cutedsl_warmup.py:97] Skipping CuTeDSL warmup because no compile units were requested.
39
+ (EngineCore pid=238829) INFO 07-15 06:22:57 [jit_monitor.py:73] Kernel JIT monitor activated; monitored JIT compilations during inference will use mode=warn.
40
+ (EngineCore pid=238829) INFO 07-15 06:22:57 [core.py:344] init engine (profile, create kv cache, warmup model) took 2.13 s
41
+ (EngineCore pid=238829) INFO 07-15 06:22:58 [vllm.py:1042] Asynchronous scheduling is enabled.
42
+ (EngineCore pid=238829) WARNING 07-15 06:22:58 [vllm.py:1096] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
43
+ (EngineCore pid=238829) WARNING 07-15 06:22:58 [vllm.py:1144] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
44
+ (EngineCore pid=238829) INFO 07-15 06:22:58 [kernel.py:292] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native'])
45
+ (EngineCore pid=238829) INFO 07-15 06:22:58 [vllm.py:1322] Cudagraph is disabled under eager mode
46
+ (EngineCore pid=238829) INFO 07-15 06:22:58 [compilation.py:312] Enabled custom fusions: norm_quant, act_quant
47
+ (APIServer pid=238699) INFO 07-15 06:22:58 [api_server.py:612] Supported tasks: ['generate']
48
+ (APIServer pid=238699) WARNING 07-15 06:22:58 [__init__.py:36] SECURITY WARNING: Development endpoints are enabled! This should NOT be used in production!
49
+ (APIServer pid=238699) INFO 07-15 06:22:58 [hf.py:548] Detected the chat template content format to be 'string'. You can set `--chat-template-content-format` to override this.
50
+ (APIServer pid=238699) INFO 07-15 06:22:58 [api_server.py:616] Starting vLLM server on http://127.0.0.1:8377
51
+ (APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:37] Available routes are:
52
+ (APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /openapi.json, Methods: HEAD, GET
53
+ (APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /docs, Methods: HEAD, GET
54
+ (APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /docs/oauth2-redirect, Methods: HEAD, GET
55
+ (APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /redoc, Methods: HEAD, GET
56
+ (APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /load, Methods: GET
57
+ (APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /version, Methods: GET
58
+ (APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /health, Methods: GET
59
+ (APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /metrics, Methods: GET
60
+ (APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /tokenize, Methods: POST
61
+ (APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /detokenize, Methods: POST
62
+ (APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /v1/models, Methods: GET
63
+ (APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /ping, Methods: GET
64
+ (APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /ping, Methods: POST
65
+ (APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /invocations, Methods: POST
66
+ (APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /reset_prefix_cache, Methods: POST
67
+ (APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /reset_mm_cache, Methods: POST
68
+ (APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /reset_encoder_cache, Methods: POST
69
+ (APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /pause, Methods: POST
70
+ (APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /resume, Methods: POST
71
+ (APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /is_paused, Methods: GET
72
+ (APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /init_weight_transfer_engine, Methods: POST
73
+ (APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /start_weight_update, Methods: POST
74
+ (APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /update_weights, Methods: POST
75
+ (APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /finish_weight_update, Methods: POST
76
+ (APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /get_world_size, Methods: GET
77
+ (APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /collective_rpc, Methods: POST
78
+ (APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /server_info, Methods: GET
79
+ (APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /sleep, Methods: POST
80
+ (APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /wake_up, Methods: POST
81
+ (APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /is_sleeping, Methods: GET
82
+ (APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /v1/chat/completions, Methods: POST
83
+ (APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /v1/chat/completions/batch, Methods: POST
84
+ (APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /v1/responses, Methods: POST
85
+ (APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /v1/responses/{response_id}, Methods: GET
86
+ (APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /v1/responses/{response_id}/cancel, Methods: POST
87
+ (APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /v1/completions, Methods: POST
88
+ (APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /v1/messages, Methods: POST
89
+ (APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /v1/messages/count_tokens, Methods: POST
90
+ (APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /generative_scoring, Methods: POST
91
+ (APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /scale_elastic_ep, Methods: POST
92
+ (APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /is_scaling_elastic_ep, Methods: POST
93
+ (APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /v1/chat/completions/render, Methods: POST
94
+ (APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /v1/completions/render, Methods: POST
95
+ (APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /v1/chat/completions/derender, Methods: POST
96
+ (APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /v1/completions/derender, Methods: POST
97
+ (APIServer pid=238699) INFO 07-15 06:22:58 [launcher.py:46] Route: /inference/v1/generate, Methods: POST
98
+ (APIServer pid=238699) INFO: Started server process [238699]
99
+ (APIServer pid=238699) INFO: Waiting for application startup.
100
+ (APIServer pid=238699) INFO: Application startup complete.
101
+ (APIServer pid=238699) INFO: 127.0.0.1:52570 - "GET /health HTTP/1.1" 200 OK
102
+ (EngineCore pid=238829) WARNING 07-15 06:22:59 [jit_monitor.py:129] Triton kernel JIT compilation during inference: _build_route_rows. This causes a latency spike; consider extending warmup to cover this shape/config.
103
+ (APIServer pid=238699) INFO 07-15 06:23:08 [loggers.py:273] Engine 000: Avg prompt throughput: 2250.2 tokens/s, Avg generation throughput: 1992.1 tokens/s, Running: 251 reqs, Waiting: 968 reqs, GPU KV cache usage: 43.6%, Prefix cache hit rate: 94.2%
104
+ (APIServer pid=238699) INFO 07-15 06:23:18 [loggers.py:273] Engine 000: Avg prompt throughput: 1600.3 tokens/s, Avg generation throughput: 2509.7 tokens/s, Running: 254 reqs, Waiting: 725 reqs, GPU KV cache usage: 45.6%, Prefix cache hit rate: 94.3%
105
+ (APIServer pid=238699) INFO 07-15 06:23:28 [loggers.py:273] Engine 000: Avg prompt throughput: 1587.9 tokens/s, Avg generation throughput: 2459.3 tokens/s, Running: 255 reqs, Waiting: 488 reqs, GPU KV cache usage: 45.1%, Prefix cache hit rate: 94.3%
106
+ (APIServer pid=238699) INFO 07-15 06:23:38 [loggers.py:273] Engine 000: Avg prompt throughput: 1514.4 tokens/s, Avg generation throughput: 2510.7 tokens/s, Running: 256 reqs, Waiting: 264 reqs, GPU KV cache usage: 47.5%, Prefix cache hit rate: 94.4%
107
+ (APIServer pid=238699) INFO 07-15 06:23:48 [loggers.py:273] Engine 000: Avg prompt throughput: 1560.7 tokens/s, Avg generation throughput: 2536.6 tokens/s, Running: 251 reqs, Waiting: 31 reqs, GPU KV cache usage: 48.7%, Prefix cache hit rate: 94.4%
108
+ (APIServer pid=238699) INFO 07-15 06:23:58 [loggers.py:273] Engine 000: Avg prompt throughput: 211.0 tokens/s, Avg generation throughput: 1868.5 tokens/s, Running: 24 reqs, Waiting: 0 reqs, GPU KV cache usage: 10.8%, Prefix cache hit rate: 94.4%
109
+ (APIServer pid=238699) INFO: 127.0.0.1:52586 - "POST /v1/completions HTTP/1.1" 200 OK
110
+ (APIServer pid=238699) INFO 07-15 06:24:07 [launcher.py:100] [shutdown] API server: shutdown triggered
111
+ (APIServer pid=238699) INFO 07-15 06:24:07 [launcher.py:116] [shutdown] API server: stopping engine client mode=abort timeout=0s
112
+ (EngineCore pid=238829) INFO 07-15 06:24:07 [core.py:1214] [shutdown] EngineCore: trigger received signal=SIGTERM
113
+ (EngineCore pid=238829) INFO 07-15 06:24:07 [core.py:1333] [shutdown] EngineCore: start mode=abort timeout=0s
114
+ (EngineCore pid=238829) INFO 07-15 06:24:07 [core.py:1364] [shutdown] EngineCore: request processing complete; starting resource teardown
115
+ (EngineCore pid=238829) INFO 07-15 06:24:07 [core.py:1227] [shutdown] EngineCore: exiting busy loop
116
+ (APIServer pid=238699) INFO 07-15 06:24:07 [core_client.py:655] [shutdown] MPClient: start timeout=0s
117
+ (APIServer pid=238699) INFO 07-15 06:24:07 [core_client.py:657] [shutdown] MPClient: stopping engine manager
118
+ (APIServer pid=238699) WARNING 07-15 06:24:07 [utils.py:626] [shutdown] Process manager: force killing remaining processes count=1
119
+ (APIServer pid=238699) INFO 07-15 06:24:07 [core_client.py:659] [shutdown] MPClient: engine manager stopped
120
+ (APIServer pid=238699) INFO 07-15 06:24:07 [core_client.py:660] [shutdown] MPClient: cleaning up background resources
121
+ (APIServer pid=238699) INFO 07-15 06:24:07 [core_client.py:662] [shutdown] MPClient: complete
122
+ (APIServer pid=238699) INFO 07-15 06:24:07 [launcher.py:125] [shutdown] API server: engine client stopped
123
+ (APIServer pid=238699) INFO 07-15 06:24:07 [launcher.py:128] [shutdown] API server: signalling HTTP server shutdown
124
+ (APIServer pid=238699) INFO 07-15 06:24:07 [launcher.py:149] [shutdown] API server: shutting down FastAPI HTTP server
125
+ (APIServer pid=238699) INFO: Shutting down
126
+ (APIServer pid=238699) INFO: Waiting for application shutdown.
127
+ (APIServer pid=238699) INFO: Application shutdown complete.
128
+ /home/henry/.local/share/uv/python/cpython-3.12.12-linux-x86_64-gnu/lib/python3.12/multiprocessing/resource_tracker.py:279: UserWarning: resource_tracker: There appear to be 1 leaked semaphore objects to clean up at shutdown
129
+ warnings.warn('resource_tracker: There appear to be %d '
evals/healing_breadth/reap_math_keep75_seed1224.json ADDED
The diff for this file is too large to render. See raw diff
 
evals/healing_breadth/reap_math_keep75_seed1224.json.server.log ADDED
@@ -0,0 +1,128 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ (APIServer pid=207721) INFO 07-14 23:28:07 [api_utils.py:339]
2
+ (APIServer pid=207721) INFO 07-14 23:28:07 [api_utils.py:339] β–ˆ β–ˆ β–ˆβ–„ β–„β–ˆ
3
+ (APIServer pid=207721) INFO 07-14 23:28:07 [api_utils.py:339] β–„β–„ β–„β–ˆ β–ˆ β–ˆ β–ˆ β–€β–„β–€ β–ˆ version 0.25.0
4
+ (APIServer pid=207721) INFO 07-14 23:28:07 [api_utils.py:339] β–ˆβ–„β–ˆβ–€ β–ˆ β–ˆ β–ˆ β–ˆ model outputs/healed/healing_breadth/reap_math_keep75_seed1224/step0050
5
+ (APIServer pid=207721) INFO 07-14 23:28:07 [api_utils.py:339] β–€β–€ β–€β–€β–€β–€β–€ β–€β–€β–€β–€β–€ β–€ β–€
6
+ (APIServer pid=207721) INFO 07-14 23:28:07 [api_utils.py:339]
7
+ (APIServer pid=207721) INFO 07-14 23:28:07 [api_utils.py:273] non-default args: {'model_tag': 'outputs/healed/healing_breadth/reap_math_keep75_seed1224/step0050', 'host': '127.0.0.1', 'port': 8377, 'model': 'outputs/healed/healing_breadth/reap_math_keep75_seed1224/step0050', 'max_model_len': 2048, 'enforce_eager': True, 'served_model_name': ['student'], 'gpu_memory_utilization': 0.85}
8
+ (APIServer pid=207721) INFO 07-14 23:28:07 [model.py:619] Resolved architecture: OlmoeForCausalLM
9
+ (APIServer pid=207721) INFO 07-14 23:28:07 [model.py:1776] Using max model len 2048
10
+ (APIServer pid=207721) INFO 07-14 23:28:07 [vllm.py:1042] Asynchronous scheduling is enabled.
11
+ (APIServer pid=207721) WARNING 07-14 23:28:07 [vllm.py:1096] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
12
+ (APIServer pid=207721) WARNING 07-14 23:28:07 [vllm.py:1144] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
13
+ (APIServer pid=207721) INFO 07-14 23:28:07 [kernel.py:292] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native'])
14
+ (APIServer pid=207721) INFO 07-14 23:28:07 [vllm.py:1322] Cudagraph is disabled under eager mode
15
+ (APIServer pid=207721) INFO 07-14 23:28:07 [compilation.py:312] Enabled custom fusions: norm_quant, act_quant
16
+ (EngineCore pid=207838) INFO 07-14 23:28:15 [core.py:114] Initializing a V1 LLM engine (v0.25.0) with config: model='outputs/healed/healing_breadth/reap_math_keep75_seed1224/step0050', speculative_config=None, tokenizer='outputs/healed/healing_breadth/reap_math_keep75_seed1224/step0050', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=False, dtype=torch.bfloat16, max_seq_len=2048, download_dir=None, load_format=auto, tensor_parallel_size=1, pipeline_parallel_size=1, data_parallel_size=1, decode_context_parallel_size=1, dcp_comm_backend=ag_rs, disable_custom_all_reduce=False, quantization=None, quantization_config=None, enforce_eager=True, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False, jit_monitor_mode='warn', jit_monitor_verbose=False), seed=0, served_model_name=student, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'mode': <CompilationMode.NONE: 0>, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'ir_enable_torch_wrap': False, 'splitting_ops': [], 'compile_mm_encoder': False, 'cudagraph_mm_encoder': False, 'encoder_cudagraph_token_budgets': [], 'encoder_cudagraph_max_vision_items_per_batch': 0, 'encoder_cudagraph_max_frames_per_batch': None, 'compile_sizes': [], 'compile_ranges_endpoints': [2048], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'size_asserts': False, 'alignment_asserts': False, 'scalar_asserts': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': <CUDAGraphMode.NONE: 0>, 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': True, 'fuse_act_quant': True, 'fuse_attn_quant': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False, 'fuse_rope_kvcache_cat_mla': False, 'fuse_act_padding': False}, 'max_cudagraph_capture_size': 0, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': False, 'static_all_moe_layers': []}, kernel_config=KernelConfig(ir_op_priority=IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native']), enable_flashinfer_autotune=True, enable_cutedsl_warmup=True, moe_backend='auto', linear_backend='auto')
17
+ (EngineCore pid=207838) INFO 07-14 23:28:15 [parallel_state.py:1607] world_size=1 rank=0 local_rank=0 distributed_init_method=tcp://192.168.0.15:33565 backend=nccl
18
+ (EngineCore pid=207838) INFO 07-14 23:28:15 [parallel_state.py:1942] rank 0 in world size 1 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank 0, EPLB rank N/A
19
+ (EngineCore pid=207838) INFO 07-14 23:28:16 [topk_topp_sampler.py:55] Using FlashInfer for top-p & top-k sampling.
20
+ (EngineCore pid=207838) INFO 07-14 23:28:16 [gpu_model_runner.py:5209] Starting to load model outputs/healed/healing_breadth/reap_math_keep75_seed1224/step0050...
21
+ (EngineCore pid=207838) INFO 07-14 23:28:17 [cuda.py:476] Using FLASH_ATTN attention backend out of potential backends: ['FLASH_ATTN', 'FLASHINFER', 'TRITON_ATTN', 'FLEX_ATTENTION'].
22
+ (EngineCore pid=207838) INFO 07-14 23:28:17 [flash_attn.py:718] Using FlashAttention version 2
23
+ (EngineCore pid=207838) INFO 07-14 23:28:17 [unquantized.py:262] Using TRITON Unquantized MoE backend out of potential backends: ['FlashInfer TRTLLM', 'FlashInfer CUTLASS', 'TRITON', 'BATCHED_TRITON'].
24
+ (EngineCore pid=207838) INFO 07-14 23:28:17 [weight_utils.py:849] Filesystem type for checkpoints: EXT4. Checkpoint size: 9.89 GiB. Available RAM: 68.46 GiB.
25
+ (EngineCore pid=207838) INFO 07-14 23:28:17 [weight_utils.py:872] Auto-prefetch is disabled because the filesystem (EXT4) is not a recognized network FS (NFS/Lustre). If you want to force prefetching, start vLLM with --safetensors-load-strategy=prefetch.
26
+ (EngineCore pid=207838)
27
+ (EngineCore pid=207838)
28
+ (EngineCore pid=207838)
29
+ (EngineCore pid=207838)
30
+ (EngineCore pid=207838)
31
+ (EngineCore pid=207838) INFO 07-14 23:28:22 [default_loader.py:430] Loading weights took 5.26 seconds
32
+ (EngineCore pid=207838) INFO 07-14 23:28:22 [unquantized.py:334] Using MoEPrepareAndFinalizeNoDPEPModular
33
+ (EngineCore pid=207838) INFO 07-14 23:28:23 [gpu_model_runner.py:5306] Model loading took 9.89 GiB memory and 5.436781 seconds
34
+ (EngineCore pid=207838) WARNING 07-14 23:28:23 [fused_moe.py:1106] Using default MoE config. Performance might be sub-optimal! Config file not found at /home/henry/Documents/PythonProjects/variable-reap/vllm-plugin/.venv25/lib/python3.12/site-packages/vllm/model_executor/layers/fused_moe/configs/E=48,N=1024,device_name=NVIDIA_GeForce_RTX_3090.json
35
+ (EngineCore pid=207838) INFO 07-14 23:28:24 [gpu_worker.py:538] Available KV cache memory: 9.73 GiB
36
+ (EngineCore pid=207838) INFO 07-14 23:28:24 [kv_cache_utils.py:2146] GPU KV cache size: 79,680 tokens
37
+ (EngineCore pid=207838) INFO 07-14 23:28:24 [kv_cache_utils.py:2147] Maximum concurrency for 2,048 tokens per request: 38.91x
38
+ (EngineCore pid=207838) INFO 07-14 23:28:24 [cutedsl_warmup.py:97] Skipping CuTeDSL warmup because no compile units were requested.
39
+ (EngineCore pid=207838) INFO 07-14 23:28:24 [jit_monitor.py:73] Kernel JIT monitor activated; monitored JIT compilations during inference will use mode=warn.
40
+ (EngineCore pid=207838) INFO 07-14 23:28:25 [core.py:344] init engine (profile, create kv cache, warmup model) took 2.11 s
41
+ (EngineCore pid=207838) INFO 07-14 23:28:25 [vllm.py:1042] Asynchronous scheduling is enabled.
42
+ (EngineCore pid=207838) WARNING 07-14 23:28:25 [vllm.py:1096] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
43
+ (EngineCore pid=207838) WARNING 07-14 23:28:25 [vllm.py:1144] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
44
+ (EngineCore pid=207838) INFO 07-14 23:28:25 [kernel.py:292] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native'])
45
+ (EngineCore pid=207838) INFO 07-14 23:28:25 [vllm.py:1322] Cudagraph is disabled under eager mode
46
+ (EngineCore pid=207838) INFO 07-14 23:28:25 [compilation.py:312] Enabled custom fusions: norm_quant, act_quant
47
+ (APIServer pid=207721) INFO 07-14 23:28:25 [api_server.py:612] Supported tasks: ['generate']
48
+ (APIServer pid=207721) WARNING 07-14 23:28:25 [__init__.py:36] SECURITY WARNING: Development endpoints are enabled! This should NOT be used in production!
49
+ (APIServer pid=207721) INFO 07-14 23:28:25 [hf.py:548] Detected the chat template content format to be 'string'. You can set `--chat-template-content-format` to override this.
50
+ (APIServer pid=207721) INFO 07-14 23:28:25 [api_server.py:616] Starting vLLM server on http://127.0.0.1:8377
51
+ (APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:37] Available routes are:
52
+ (APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /openapi.json, Methods: HEAD, GET
53
+ (APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /docs, Methods: HEAD, GET
54
+ (APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /docs/oauth2-redirect, Methods: HEAD, GET
55
+ (APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /redoc, Methods: HEAD, GET
56
+ (APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /load, Methods: GET
57
+ (APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /version, Methods: GET
58
+ (APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /health, Methods: GET
59
+ (APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /metrics, Methods: GET
60
+ (APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /tokenize, Methods: POST
61
+ (APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /detokenize, Methods: POST
62
+ (APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /v1/models, Methods: GET
63
+ (APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /ping, Methods: GET
64
+ (APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /ping, Methods: POST
65
+ (APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /invocations, Methods: POST
66
+ (APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /reset_prefix_cache, Methods: POST
67
+ (APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /reset_mm_cache, Methods: POST
68
+ (APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /reset_encoder_cache, Methods: POST
69
+ (APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /pause, Methods: POST
70
+ (APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /resume, Methods: POST
71
+ (APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /is_paused, Methods: GET
72
+ (APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /init_weight_transfer_engine, Methods: POST
73
+ (APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /start_weight_update, Methods: POST
74
+ (APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /update_weights, Methods: POST
75
+ (APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /finish_weight_update, Methods: POST
76
+ (APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /get_world_size, Methods: GET
77
+ (APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /collective_rpc, Methods: POST
78
+ (APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /server_info, Methods: GET
79
+ (APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /sleep, Methods: POST
80
+ (APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /wake_up, Methods: POST
81
+ (APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /is_sleeping, Methods: GET
82
+ (APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /v1/chat/completions, Methods: POST
83
+ (APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /v1/chat/completions/batch, Methods: POST
84
+ (APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /v1/responses, Methods: POST
85
+ (APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /v1/responses/{response_id}, Methods: GET
86
+ (APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /v1/responses/{response_id}/cancel, Methods: POST
87
+ (APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /v1/completions, Methods: POST
88
+ (APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /v1/messages, Methods: POST
89
+ (APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /v1/messages/count_tokens, Methods: POST
90
+ (APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /generative_scoring, Methods: POST
91
+ (APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /scale_elastic_ep, Methods: POST
92
+ (APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /is_scaling_elastic_ep, Methods: POST
93
+ (APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /v1/chat/completions/render, Methods: POST
94
+ (APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /v1/completions/render, Methods: POST
95
+ (APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /v1/chat/completions/derender, Methods: POST
96
+ (APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /v1/completions/derender, Methods: POST
97
+ (APIServer pid=207721) INFO 07-14 23:28:25 [launcher.py:46] Route: /inference/v1/generate, Methods: POST
98
+ (APIServer pid=207721) INFO: Started server process [207721]
99
+ (APIServer pid=207721) INFO: Waiting for application startup.
100
+ (APIServer pid=207721) INFO: Application startup complete.
101
+ (APIServer pid=207721) INFO: 127.0.0.1:52832 - "GET /health HTTP/1.1" 200 OK
102
+ (EngineCore pid=207838) WARNING 07-14 23:28:27 [jit_monitor.py:129] Triton kernel JIT compilation during inference: fused_moe_kernel. This causes a latency spike; consider extending warmup to cover this shape/config.
103
+ (APIServer pid=207721) INFO 07-14 23:28:35 [loggers.py:273] Engine 000: Avg prompt throughput: 2896.9 tokens/s, Avg generation throughput: 2135.2 tokens/s, Running: 252 reqs, Waiting: 942 reqs, GPU KV cache usage: 47.8%, Prefix cache hit rate: 93.2%
104
+ (APIServer pid=207721) INFO 07-14 23:28:45 [loggers.py:273] Engine 000: Avg prompt throughput: 2050.3 tokens/s, Avg generation throughput: 3019.8 tokens/s, Running: 254 reqs, Waiting: 678 reqs, GPU KV cache usage: 51.2%, Prefix cache hit rate: 93.3%
105
+ (APIServer pid=207721) INFO 07-14 23:28:55 [loggers.py:273] Engine 000: Avg prompt throughput: 2279.4 tokens/s, Avg generation throughput: 2991.5 tokens/s, Running: 253 reqs, Waiting: 388 reqs, GPU KV cache usage: 50.9%, Prefix cache hit rate: 93.4%
106
+ (APIServer pid=207721) INFO 07-14 23:29:05 [loggers.py:273] Engine 000: Avg prompt throughput: 2386.0 tokens/s, Avg generation throughput: 2965.4 tokens/s, Running: 255 reqs, Waiting: 94 reqs, GPU KV cache usage: 51.0%, Prefix cache hit rate: 93.4%
107
+ (APIServer pid=207721) INFO 07-14 23:29:15 [loggers.py:273] Engine 000: Avg prompt throughput: 787.1 tokens/s, Avg generation throughput: 2775.3 tokens/s, Running: 14 reqs, Waiting: 0 reqs, GPU KV cache usage: 7.4%, Prefix cache hit rate: 93.4%
108
+ (APIServer pid=207721) INFO: 127.0.0.1:52848 - "POST /v1/completions HTTP/1.1" 200 OK
109
+ (EngineCore pid=207838) INFO 07-14 23:29:20 [core.py:1214] [shutdown] EngineCore: trigger received signal=SIGTERM
110
+ (APIServer pid=207721) INFO 07-14 23:29:20 [launcher.py:100] [shutdown] API server: shutdown triggered
111
+ (APIServer pid=207721) INFO 07-14 23:29:20 [launcher.py:116] [shutdown] API server: stopping engine client mode=abort timeout=0s
112
+ (EngineCore pid=207838) INFO 07-14 23:29:20 [core.py:1333] [shutdown] EngineCore: start mode=abort timeout=0s
113
+ (EngineCore pid=207838) INFO 07-14 23:29:20 [core.py:1364] [shutdown] EngineCore: request processing complete; starting resource teardown
114
+ (EngineCore pid=207838) INFO 07-14 23:29:20 [core.py:1227] [shutdown] EngineCore: exiting busy loop
115
+ (APIServer pid=207721) INFO 07-14 23:29:20 [core_client.py:655] [shutdown] MPClient: start timeout=0s
116
+ (APIServer pid=207721) INFO 07-14 23:29:20 [core_client.py:657] [shutdown] MPClient: stopping engine manager
117
+ (APIServer pid=207721) WARNING 07-14 23:29:20 [utils.py:626] [shutdown] Process manager: force killing remaining processes count=1
118
+ (APIServer pid=207721) INFO 07-14 23:29:20 [core_client.py:659] [shutdown] MPClient: engine manager stopped
119
+ (APIServer pid=207721) INFO 07-14 23:29:20 [core_client.py:660] [shutdown] MPClient: cleaning up background resources
120
+ (APIServer pid=207721) INFO 07-14 23:29:20 [core_client.py:662] [shutdown] MPClient: complete
121
+ (APIServer pid=207721) INFO 07-14 23:29:20 [launcher.py:125] [shutdown] API server: engine client stopped
122
+ (APIServer pid=207721) INFO 07-14 23:29:20 [launcher.py:128] [shutdown] API server: signalling HTTP server shutdown
123
+ (APIServer pid=207721) INFO 07-14 23:29:20 [launcher.py:149] [shutdown] API server: shutting down FastAPI HTTP server
124
+ (APIServer pid=207721) INFO: Shutting down
125
+ (APIServer pid=207721) INFO: Waiting for application shutdown.
126
+ (APIServer pid=207721) INFO: Application shutdown complete.
127
+ /home/henry/.local/share/uv/python/cpython-3.12.12-linux-x86_64-gnu/lib/python3.12/multiprocessing/resource_tracker.py:279: UserWarning: resource_tracker: There appear to be 1 leaked semaphore objects to clean up at shutdown
128
+ warnings.warn('resource_tracker: There appear to be %d '
evals/healing_breadth/uniform_math_keep50_seed1224.json ADDED
The diff for this file is too large to render. See raw diff
 
evals/healing_breadth/uniform_math_keep50_seed1224.json.server.log ADDED
@@ -0,0 +1,130 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ (APIServer pid=215091) INFO 07-15 00:43:36 [api_utils.py:339]
2
+ (APIServer pid=215091) INFO 07-15 00:43:36 [api_utils.py:339] β–ˆ β–ˆ β–ˆβ–„ β–„β–ˆ
3
+ (APIServer pid=215091) INFO 07-15 00:43:36 [api_utils.py:339] β–„β–„ β–„β–ˆ β–ˆ β–ˆ β–ˆ β–€β–„β–€ β–ˆ version 0.25.0
4
+ (APIServer pid=215091) INFO 07-15 00:43:36 [api_utils.py:339] β–ˆβ–„β–ˆβ–€ β–ˆ β–ˆ β–ˆ β–ˆ model outputs/healed/healing_breadth/uniform_math_keep50_seed1224/step0050
5
+ (APIServer pid=215091) INFO 07-15 00:43:36 [api_utils.py:339] β–€β–€ β–€β–€β–€β–€β–€ β–€β–€β–€β–€β–€ β–€ β–€
6
+ (APIServer pid=215091) INFO 07-15 00:43:36 [api_utils.py:339]
7
+ (APIServer pid=215091) INFO 07-15 00:43:36 [api_utils.py:273] non-default args: {'model_tag': 'outputs/healed/healing_breadth/uniform_math_keep50_seed1224/step0050', 'host': '127.0.0.1', 'port': 8377, 'model': 'outputs/healed/healing_breadth/uniform_math_keep50_seed1224/step0050', 'max_model_len': 2048, 'enforce_eager': True, 'served_model_name': ['student'], 'gpu_memory_utilization': 0.85}
8
+ (APIServer pid=215091) INFO 07-15 00:43:36 [model.py:619] Resolved architecture: PrunedOlmoeForCausalLM
9
+ (APIServer pid=215091) INFO 07-15 00:43:36 [model.py:1776] Using max model len 2048
10
+ (APIServer pid=215091) INFO 07-15 00:43:36 [vllm.py:1042] Asynchronous scheduling is enabled.
11
+ (APIServer pid=215091) WARNING 07-15 00:43:36 [vllm.py:1096] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
12
+ (APIServer pid=215091) WARNING 07-15 00:43:36 [vllm.py:1144] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
13
+ (APIServer pid=215091) INFO 07-15 00:43:36 [kernel.py:292] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native'])
14
+ (APIServer pid=215091) INFO 07-15 00:43:36 [vllm.py:1322] Cudagraph is disabled under eager mode
15
+ (APIServer pid=215091) INFO 07-15 00:43:36 [compilation.py:312] Enabled custom fusions: norm_quant, act_quant
16
+ (EngineCore pid=215209) INFO 07-15 00:43:44 [core.py:114] Initializing a V1 LLM engine (v0.25.0) with config: model='outputs/healed/healing_breadth/uniform_math_keep50_seed1224/step0050', speculative_config=None, tokenizer='outputs/healed/healing_breadth/uniform_math_keep50_seed1224/step0050', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=False, dtype=torch.bfloat16, max_seq_len=2048, download_dir=None, load_format=auto, tensor_parallel_size=1, pipeline_parallel_size=1, data_parallel_size=1, decode_context_parallel_size=1, dcp_comm_backend=ag_rs, disable_custom_all_reduce=False, quantization=None, quantization_config=None, enforce_eager=True, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False, jit_monitor_mode='warn', jit_monitor_verbose=False), seed=0, served_model_name=student, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'mode': <CompilationMode.NONE: 0>, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'ir_enable_torch_wrap': False, 'splitting_ops': [], 'compile_mm_encoder': False, 'cudagraph_mm_encoder': False, 'encoder_cudagraph_token_budgets': [], 'encoder_cudagraph_max_vision_items_per_batch': 0, 'encoder_cudagraph_max_frames_per_batch': None, 'compile_sizes': [], 'compile_ranges_endpoints': [2048], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'size_asserts': False, 'alignment_asserts': False, 'scalar_asserts': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': <CUDAGraphMode.NONE: 0>, 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': True, 'fuse_act_quant': True, 'fuse_attn_quant': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False, 'fuse_rope_kvcache_cat_mla': False, 'fuse_act_padding': False}, 'max_cudagraph_capture_size': 0, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': False, 'static_all_moe_layers': []}, kernel_config=KernelConfig(ir_op_priority=IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native']), enable_flashinfer_autotune=True, enable_cutedsl_warmup=True, moe_backend='auto', linear_backend='auto')
17
+ (EngineCore pid=215209) INFO 07-15 00:43:44 [parallel_state.py:1607] world_size=1 rank=0 local_rank=0 distributed_init_method=tcp://192.168.0.15:36073 backend=nccl
18
+ (EngineCore pid=215209) INFO 07-15 00:43:44 [parallel_state.py:1942] rank 0 in world size 1 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank 0, EPLB rank N/A
19
+ (EngineCore pid=215209) INFO 07-15 00:43:45 [topk_topp_sampler.py:55] Using FlashInfer for top-p & top-k sampling.
20
+ (EngineCore pid=215209) INFO 07-15 00:43:45 [gpu_model_runner.py:5209] Starting to load model outputs/healed/healing_breadth/uniform_math_keep50_seed1224/step0050...
21
+ (EngineCore pid=215209) INFO 07-15 00:43:46 [cuda.py:476] Using FLASH_ATTN attention backend out of potential backends: ['FLASH_ATTN', 'FLASHINFER', 'TRITON_ATTN', 'FLEX_ATTENTION'].
22
+ (EngineCore pid=215209) INFO 07-15 00:43:46 [flash_attn.py:718] Using FlashAttention version 2
23
+ (EngineCore pid=215209) /home/henry/.cache/glean/megablocks-variable-93a1479bc15b/megablocks/grouped_gemm_util.py:10: UserWarning: Grouped GEMM not available.
24
+ (EngineCore pid=215209) warnings.warn('Grouped GEMM not available.')
25
+ (EngineCore pid=215209) INFO 07-15 00:43:46 [weight_utils.py:849] Filesystem type for checkpoints: EXT4. Checkpoint size: 6.89 GiB. Available RAM: 61.47 GiB.
26
+ (EngineCore pid=215209) INFO 07-15 00:43:46 [weight_utils.py:872] Auto-prefetch is disabled because the filesystem (EXT4) is not a recognized network FS (NFS/Lustre). If you want to force prefetching, start vLLM with --safetensors-load-strategy=prefetch.
27
+ (EngineCore pid=215209)
28
+ (EngineCore pid=215209)
29
+ (EngineCore pid=215209)
30
+ (EngineCore pid=215209)
31
+ (EngineCore pid=215209)
32
+ (EngineCore pid=215209) INFO 07-15 00:43:50 [default_loader.py:430] Loading weights took 4.53 seconds
33
+ (EngineCore pid=215209) INFO 07-15 00:43:51 [gpu_model_runner.py:5306] Model loading took 6.89 GiB memory and 4.727233 seconds
34
+ (EngineCore pid=215209) INFO 07-15 00:43:52 [gpu_worker.py:538] Available KV cache memory: 12.84 GiB
35
+ (EngineCore pid=215209) INFO 07-15 00:43:52 [kv_cache_utils.py:2146] GPU KV cache size: 105,216 tokens
36
+ (EngineCore pid=215209) INFO 07-15 00:43:52 [kv_cache_utils.py:2147] Maximum concurrency for 2,048 tokens per request: 51.38x
37
+ (EngineCore pid=215209) INFO 07-15 00:43:52 [cutedsl_warmup.py:97] Skipping CuTeDSL warmup because no compile units were requested.
38
+ (EngineCore pid=215209) INFO 07-15 00:43:53 [jit_monitor.py:73] Kernel JIT monitor activated; monitored JIT compilations during inference will use mode=warn.
39
+ (EngineCore pid=215209) INFO 07-15 00:43:53 [core.py:344] init engine (profile, create kv cache, warmup model) took 2.12 s
40
+ (EngineCore pid=215209) INFO 07-15 00:43:53 [vllm.py:1042] Asynchronous scheduling is enabled.
41
+ (EngineCore pid=215209) WARNING 07-15 00:43:53 [vllm.py:1096] Enforce eager set, disabling torch.compile and CUDAGraphs. This is equivalent to setting -cc.mode=none -cc.cudagraph_mode=none
42
+ (EngineCore pid=215209) WARNING 07-15 00:43:53 [vllm.py:1144] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
43
+ (EngineCore pid=215209) INFO 07-15 00:43:53 [kernel.py:292] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['vllm_c', 'native'], fused_add_rms_norm=['vllm_c', 'native'])
44
+ (EngineCore pid=215209) INFO 07-15 00:43:53 [vllm.py:1322] Cudagraph is disabled under eager mode
45
+ (EngineCore pid=215209) INFO 07-15 00:43:53 [compilation.py:312] Enabled custom fusions: norm_quant, act_quant
46
+ (APIServer pid=215091) INFO 07-15 00:43:53 [api_server.py:612] Supported tasks: ['generate']
47
+ (APIServer pid=215091) WARNING 07-15 00:43:53 [__init__.py:36] SECURITY WARNING: Development endpoints are enabled! This should NOT be used in production!
48
+ (APIServer pid=215091) INFO 07-15 00:43:53 [hf.py:548] Detected the chat template content format to be 'string'. You can set `--chat-template-content-format` to override this.
49
+ (APIServer pid=215091) INFO 07-15 00:43:53 [api_server.py:616] Starting vLLM server on http://127.0.0.1:8377
50
+ (APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:37] Available routes are:
51
+ (APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /openapi.json, Methods: GET, HEAD
52
+ (APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /docs, Methods: GET, HEAD
53
+ (APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /docs/oauth2-redirect, Methods: GET, HEAD
54
+ (APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /redoc, Methods: GET, HEAD
55
+ (APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /load, Methods: GET
56
+ (APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /version, Methods: GET
57
+ (APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /health, Methods: GET
58
+ (APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /metrics, Methods: GET
59
+ (APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /tokenize, Methods: POST
60
+ (APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /detokenize, Methods: POST
61
+ (APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /v1/models, Methods: GET
62
+ (APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /ping, Methods: GET
63
+ (APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /ping, Methods: POST
64
+ (APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /invocations, Methods: POST
65
+ (APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /reset_prefix_cache, Methods: POST
66
+ (APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /reset_mm_cache, Methods: POST
67
+ (APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /reset_encoder_cache, Methods: POST
68
+ (APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /pause, Methods: POST
69
+ (APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /resume, Methods: POST
70
+ (APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /is_paused, Methods: GET
71
+ (APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /init_weight_transfer_engine, Methods: POST
72
+ (APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /start_weight_update, Methods: POST
73
+ (APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /update_weights, Methods: POST
74
+ (APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /finish_weight_update, Methods: POST
75
+ (APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /get_world_size, Methods: GET
76
+ (APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /collective_rpc, Methods: POST
77
+ (APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /server_info, Methods: GET
78
+ (APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /sleep, Methods: POST
79
+ (APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /wake_up, Methods: POST
80
+ (APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /is_sleeping, Methods: GET
81
+ (APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /v1/chat/completions, Methods: POST
82
+ (APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /v1/chat/completions/batch, Methods: POST
83
+ (APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /v1/responses, Methods: POST
84
+ (APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /v1/responses/{response_id}, Methods: GET
85
+ (APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /v1/responses/{response_id}/cancel, Methods: POST
86
+ (APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /v1/completions, Methods: POST
87
+ (APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /v1/messages, Methods: POST
88
+ (APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /v1/messages/count_tokens, Methods: POST
89
+ (APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /generative_scoring, Methods: POST
90
+ (APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /scale_elastic_ep, Methods: POST
91
+ (APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /is_scaling_elastic_ep, Methods: POST
92
+ (APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /v1/chat/completions/render, Methods: POST
93
+ (APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /v1/completions/render, Methods: POST
94
+ (APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /v1/chat/completions/derender, Methods: POST
95
+ (APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /v1/completions/derender, Methods: POST
96
+ (APIServer pid=215091) INFO 07-15 00:43:53 [launcher.py:46] Route: /inference/v1/generate, Methods: POST
97
+ (APIServer pid=215091) INFO: Started server process [215091]
98
+ (APIServer pid=215091) INFO: Waiting for application startup.
99
+ (APIServer pid=215091) INFO: Application startup complete.
100
+ (APIServer pid=215091) INFO: 127.0.0.1:42130 - "GET /health HTTP/1.1" 200 OK
101
+ (EngineCore pid=215209) WARNING 07-15 00:43:55 [jit_monitor.py:129] Triton kernel JIT compilation during inference: _build_route_rows. This causes a latency spike; consider extending warmup to cover this shape/config.
102
+ (APIServer pid=215091) INFO 07-15 00:44:04 [loggers.py:273] Engine 000: Avg prompt throughput: 2793.7 tokens/s, Avg generation throughput: 2084.8 tokens/s, Running: 253 reqs, Waiting: 957 reqs, GPU KV cache usage: 37.0%, Prefix cache hit rate: 93.2%
103
+ (APIServer pid=215091) INFO 07-15 00:44:14 [loggers.py:273] Engine 000: Avg prompt throughput: 1725.4 tokens/s, Avg generation throughput: 2588.6 tokens/s, Running: 256 reqs, Waiting: 733 reqs, GPU KV cache usage: 40.4%, Prefix cache hit rate: 93.3%
104
+ (APIServer pid=215091) INFO 07-15 00:44:24 [loggers.py:273] Engine 000: Avg prompt throughput: 1870.4 tokens/s, Avg generation throughput: 2587.0 tokens/s, Running: 253 reqs, Waiting: 496 reqs, GPU KV cache usage: 40.3%, Prefix cache hit rate: 93.3%
105
+ (APIServer pid=215091) INFO 07-15 00:44:34 [loggers.py:273] Engine 000: Avg prompt throughput: 1878.2 tokens/s, Avg generation throughput: 2561.9 tokens/s, Running: 253 reqs, Waiting: 263 reqs, GPU KV cache usage: 41.0%, Prefix cache hit rate: 93.4%
106
+ (APIServer pid=215091) INFO 07-15 00:44:44 [loggers.py:273] Engine 000: Avg prompt throughput: 1818.6 tokens/s, Avg generation throughput: 2639.5 tokens/s, Running: 255 reqs, Waiting: 35 reqs, GPU KV cache usage: 44.2%, Prefix cache hit rate: 93.4%
107
+ (APIServer pid=215091) INFO 07-15 00:44:54 [loggers.py:273] Engine 000: Avg prompt throughput: 315.9 tokens/s, Avg generation throughput: 2202.9 tokens/s, Running: 29 reqs, Waiting: 0 reqs, GPU KV cache usage: 10.0%, Prefix cache hit rate: 93.4%
108
+ (APIServer pid=215091) INFO: 127.0.0.1:42138 - "POST /v1/completions HTTP/1.1" 200 OK
109
+ (APIServer pid=215091) INFO 07-15 00:45:04 [loggers.py:273] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 389.4 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 93.4%
110
+ (APIServer pid=215091) INFO 07-15 00:45:04 [launcher.py:100] [shutdown] API server: shutdown triggered
111
+ (EngineCore pid=215209) INFO 07-15 00:45:04 [core.py:1214] [shutdown] EngineCore: trigger received signal=SIGTERM
112
+ (APIServer pid=215091) INFO 07-15 00:45:04 [launcher.py:116] [shutdown] API server: stopping engine client mode=abort timeout=0s
113
+ (EngineCore pid=215209) INFO 07-15 00:45:04 [core.py:1333] [shutdown] EngineCore: start mode=abort timeout=0s
114
+ (EngineCore pid=215209) INFO 07-15 00:45:04 [core.py:1364] [shutdown] EngineCore: request processing complete; starting resource teardown
115
+ (EngineCore pid=215209) INFO 07-15 00:45:04 [core.py:1227] [shutdown] EngineCore: exiting busy loop
116
+ (APIServer pid=215091) INFO 07-15 00:45:04 [core_client.py:655] [shutdown] MPClient: start timeout=0s
117
+ (APIServer pid=215091) INFO 07-15 00:45:04 [core_client.py:657] [shutdown] MPClient: stopping engine manager
118
+ (APIServer pid=215091) WARNING 07-15 00:45:04 [utils.py:626] [shutdown] Process manager: force killing remaining processes count=1
119
+ (APIServer pid=215091) INFO: Shutting down
120
+ (APIServer pid=215091) INFO 07-15 00:45:04 [core_client.py:659] [shutdown] MPClient: engine manager stopped
121
+ (APIServer pid=215091) INFO 07-15 00:45:04 [core_client.py:660] [shutdown] MPClient: cleaning up background resources
122
+ (APIServer pid=215091) INFO 07-15 00:45:04 [core_client.py:662] [shutdown] MPClient: complete
123
+ (APIServer pid=215091) INFO 07-15 00:45:04 [launcher.py:125] [shutdown] API server: engine client stopped
124
+ (APIServer pid=215091) INFO 07-15 00:45:04 [launcher.py:128] [shutdown] API server: signalling HTTP server shutdown
125
+ (APIServer pid=215091) INFO 07-15 00:45:04 [launcher.py:149] [shutdown] API server: shutting down FastAPI HTTP server
126
+ (APIServer pid=215091) INFO: Shutting down
127
+ (APIServer pid=215091) INFO: Waiting for application shutdown.
128
+ (APIServer pid=215091) INFO: Application shutdown complete.
129
+ /home/henry/.local/share/uv/python/cpython-3.12.12-linux-x86_64-gnu/lib/python3.12/multiprocessing/resource_tracker.py:279: UserWarning: resource_tracker: There appear to be 1 leaked semaphore objects to clean up at shutdown
130
+ warnings.warn('resource_tracker: There appear to be %d '
evals/math_unhealed/glean_winnow-olmoe-math-keep25.eval.log ADDED
The diff for this file is too large to render. See raw diff
 
evals/math_unhealed/glean_winnow-olmoe-math-keep50.eval.log ADDED
@@ -0,0 +1,70 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
0
  0%| | 0/1319 [00:00<?, ?it/s]
1
  14%|β–ˆβ– | 191/1319 [00:00<00:00, 1902.21it/s]
2
  29%|β–ˆβ–ˆβ–‰ | 385/1319 [00:00<00:00, 1923.43it/s]
3
  44%|β–ˆβ–ˆβ–ˆβ–ˆβ– | 581/1319 [00:00<00:00, 1936.99it/s]
4
  59%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‰ | 777/1319 [00:00<00:00, 1942.54it/s]
5
  74%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ– | 973/1319 [00:00<00:00, 1945.81it/s]
6
  89%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Š | 1168/1319 [00:00<00:00, 1939.97it/s]
 
 
7
  0%| | 0/500 [00:00<?, ?it/s]
8
  8%|β–Š | 42/500 [00:00<00:01, 411.34it/s]
9
  17%|β–ˆβ–‹ | 84/500 [00:00<00:00, 416.06it/s]
10
  25%|β–ˆβ–ˆβ–Œ | 127/500 [00:00<00:00, 418.71it/s]
11
  34%|β–ˆβ–ˆβ–ˆβ– | 170/500 [00:00<00:00, 421.00it/s]
12
  43%|β–ˆβ–ˆβ–ˆβ–ˆβ–Ž | 213/500 [00:00<00:00, 422.35it/s]
13
  51%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆ | 256/500 [00:00<00:00, 424.21it/s]
14
  60%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‰ | 299/500 [00:00<00:00, 424.80it/s]
15
  68%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Š | 342/500 [00:00<00:00, 424.12it/s]
16
  77%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‹ | 385/500 [00:00<00:00, 424.83it/s]
17
  86%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Œ | 428/500 [00:01<00:00, 424.33it/s]
18
  94%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–| 471/500 [00:01<00:00, 425.60it/s]
 
 
19
  0%| | 0/541 [00:00<?, ?it/s]
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
20
  0%| | 0/164 [00:00<?, ?it/s]
21
  87%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‹ | 142/164 [00:00<00:00, 1419.11it/s]
 
 
22
  0%| | 0/500 [00:00<?, ?it/s]
23
  3%|β–Ž | 17/500 [00:00<00:02, 161.93it/s]
24
  7%|β–‹ | 34/500 [00:00<00:02, 162.61it/s]
25
  10%|β–ˆ | 51/500 [00:00<00:02, 163.06it/s]
26
  14%|β–ˆβ–Ž | 68/500 [00:00<00:02, 163.43it/s]
27
  17%|β–ˆβ–‹ | 85/500 [00:00<00:02, 163.79it/s]
28
  20%|β–ˆβ–ˆ | 102/500 [00:00<00:02, 164.02it/s]
29
  24%|β–ˆβ–ˆβ– | 119/500 [00:00<00:02, 164.32it/s]
30
  27%|β–ˆβ–ˆβ–‹ | 136/500 [00:00<00:02, 164.49it/s]
31
  31%|β–ˆβ–ˆβ–ˆ | 153/500 [00:00<00:02, 164.68it/s]
32
  34%|β–ˆβ–ˆβ–ˆβ– | 170/500 [00:01<00:02, 164.78it/s]
33
  37%|β–ˆβ–ˆβ–ˆβ–‹ | 187/500 [00:01<00:01, 165.03it/s]
34
  41%|β–ˆβ–ˆβ–ˆβ–ˆ | 204/500 [00:01<00:01, 165.27it/s]
35
  44%|β–ˆβ–ˆβ–ˆβ–ˆβ– | 221/500 [00:01<00:01, 165.43it/s]
36
  48%|β–ˆβ–ˆβ–ˆβ–ˆβ–Š | 238/500 [00:01<00:01, 165.60it/s]
37
  51%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆ | 255/500 [00:01<00:01, 165.58it/s]
38
  54%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ– | 272/500 [00:01<00:01, 165.63it/s]
39
  58%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Š | 289/500 [00:01<00:01, 165.63it/s]
40
  61%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ | 306/500 [00:01<00:01, 165.70it/s]
41
  65%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ– | 323/500 [00:01<00:01, 165.80it/s]
42
  68%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Š | 340/500 [00:02<00:00, 165.89it/s]
43
  71%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ– | 357/500 [00:02<00:00, 165.80it/s]
44
  75%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ– | 374/500 [00:02<00:00, 165.88it/s]
45
  78%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Š | 391/500 [00:02<00:00, 165.74it/s]
46
  82%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ– | 408/500 [00:02<00:00, 165.84it/s]
47
  85%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Œ | 425/500 [00:02<00:00, 165.82it/s]
48
  88%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Š | 442/500 [00:02<00:00, 166.05it/s]
49
  92%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–| 459/500 [00:02<00:00, 165.99it/s]
50
  95%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Œ| 476/500 [00:02<00:00, 165.98it/s]
51
  99%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Š| 493/500 [00:02<00:00, 166.07it/s]
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ 2026-08-10T14:55:12-07:00 serving outputs/release/unhealed/winnow-olmoe-math-keep50 on GPU 1 port 8521 (pp=1 think=template-default)
2
+ 2026-08-10T14:55:12-07:00 waiting for server /health ...
3
+ 2026-08-10T14:55:42-07:00 server up; chat pass [gsm8k_cot_zeroshot,minerva_math500,ifeval]
4
+ 2026-08-10:14:55:50 INFO [_cli.run:388] Selected Tasks: ['gsm8k_cot_zeroshot', 'minerva_math500', 'ifeval']
5
+ 2026-08-10:14:55:51 INFO [evaluator:214] Setting random seed to 0 | Setting numpy seed to 1234 | Setting torch manual seed to 1234 | Setting fewshot manual seed to 1234
6
+ 2026-08-10:14:55:51 WARNING [evaluator:226] generation_kwargs: {'max_gen_toks': 1280} specified through cli, these settings will update set parameters in yaml tasks. Ensure 'do_sample=True' for non-greedy decoding!
7
+ 2026-08-10:14:55:51 INFO [evaluator:239] Initializing local-chat-completions model, with arguments: {'model': 'student', 'base_url': 'http://127.0.0.1:8521/v1/chat/completions', 'num_concurrent': 48, 'tokenized_requests': False, 'max_retries': 3}
8
+ 2026-08-10:14:55:51 INFO [models.api_models:179] Using max length 2048 - 1
9
+ 2026-08-10:14:55:51 INFO [models.api_models:200] Using tokenizer None
10
+ 2026-08-10:14:55:57 INFO [evaluator_utils:446] Selected tasks:
11
+ 2026-08-10:14:55:57 INFO [evaluator_utils:480] Task: gsm8k_cot_zeroshot (gsm8k/gsm8k-cot-zeroshot.yaml)
12
+ 2026-08-10:14:55:57 INFO [evaluator_utils:480] Task: ifeval (ifeval/ifeval.yaml)
13
+ 2026-08-10:14:55:57 INFO [evaluator_utils:480] Task: minerva_math500 (minerva_math/minerva_math500.yaml)
14
+ 2026-08-10:14:55:57 INFO [evaluator:314] gsm8k_cot_zeroshot: Using gen_kwargs: {'until': ['Q:', '</s>', '<|im_end|>'], 'do_sample': False, 'max_gen_toks': 1280}
15
+ 2026-08-10:14:55:57 INFO [evaluator:314] minerva_math500: Using gen_kwargs: {'until': ['Problem:'], 'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 1280}
16
+ 2026-08-10:14:55:57 INFO [evaluator:314] ifeval: Using gen_kwargs: {'until': [], 'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 1280}
17
+ 2026-08-10:14:55:57 INFO [api.task:312] Building contexts for gsm8k_cot_zeroshot on rank 0...
18
+
19
  0%| | 0/1319 [00:00<?, ?it/s]
20
  14%|β–ˆβ– | 191/1319 [00:00<00:00, 1902.21it/s]
21
  29%|β–ˆβ–ˆβ–‰ | 385/1319 [00:00<00:00, 1923.43it/s]
22
  44%|β–ˆβ–ˆβ–ˆβ–ˆβ– | 581/1319 [00:00<00:00, 1936.99it/s]
23
  59%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‰ | 777/1319 [00:00<00:00, 1942.54it/s]
24
  74%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ– | 973/1319 [00:00<00:00, 1945.81it/s]
25
  89%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Š | 1168/1319 [00:00<00:00, 1939.97it/s]
26
+ 2026-08-10:14:55:58 INFO [api.task:312] Building contexts for minerva_math500 on rank 0...
27
+
28
  0%| | 0/500 [00:00<?, ?it/s]
29
  8%|β–Š | 42/500 [00:00<00:01, 411.34it/s]
30
  17%|β–ˆβ–‹ | 84/500 [00:00<00:00, 416.06it/s]
31
  25%|β–ˆβ–ˆβ–Œ | 127/500 [00:00<00:00, 418.71it/s]
32
  34%|β–ˆβ–ˆβ–ˆβ– | 170/500 [00:00<00:00, 421.00it/s]
33
  43%|β–ˆβ–ˆβ–ˆβ–ˆβ–Ž | 213/500 [00:00<00:00, 422.35it/s]
34
  51%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆ | 256/500 [00:00<00:00, 424.21it/s]
35
  60%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‰ | 299/500 [00:00<00:00, 424.80it/s]
36
  68%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Š | 342/500 [00:00<00:00, 424.12it/s]
37
  77%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‹ | 385/500 [00:00<00:00, 424.83it/s]
38
  86%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Œ | 428/500 [00:01<00:00, 424.33it/s]
39
  94%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–| 471/500 [00:01<00:00, 425.60it/s]
40
+ 2026-08-10:14:55:59 INFO [api.task:312] Building contexts for ifeval on rank 0...
41
+
42
  0%| | 0/541 [00:00<?, ?it/s]
43
+ 2026-08-10:14:55:59 INFO [evaluator:585] Running generate_until requests
44
+ 2026-08-10:14:55:59 INFO [models.api_models:747] Tokenized requests are disabled. Context + generation length is not checked.
45
+
46
+
47
+
48
+ 2026-08-10:15:00:52 INFO [loggers.evaluation_tracker:247] Saving results aggregated
49
+ 2026-08-10:15:00:52 INFO [loggers.evaluation_tracker:119] Saving per-task samples to outputs/evals/math_unhealed/glean_winnow-olmoe-math-keep50/student/*.jsonl
50
+ local-chat-completions ({'model': 'student', 'base_url': 'http://127.0.0.1:8521/v1/chat/completions', 'num_concurrent': 48, 'tokenized_requests': False, 'max_retries': 3}), gen_kwargs: ({'max_gen_toks': 1280}), limit: None, num_fewshot: None, batch_size: 1
51
+ | Tasks |Version| Filter |n-shot| Metric | |Value | |Stderr|
52
+ |------------------|------:|----------------|-----:|-----------------------|---|-----:|---|------|
53
+ |gsm8k_cot_zeroshot| 3|flexible-extract| 0|exact_match |↑ |0.4405|Β± |0.0137|
54
+ | | |strict-match | 0|exact_match |↑ |0.0000|Β± | 0|
55
+ |ifeval | 4|none | 0|inst_level_loose_acc |↑ |0.5420|Β± | N/A|
56
+ | | |none | 0|inst_level_strict_acc |↑ |0.5048|Β± | N/A|
57
+ | | |none | 0|prompt_level_loose_acc |↑ |0.4177|Β± |0.0212|
58
+ | | |none | 0|prompt_level_strict_acc|↑ |0.3789|Β± |0.0209|
59
+ |minerva_math500 | 3|none | 4|exact_match |↑ |0.1760|Β± |0.0170|
60
+ | | |none | 4|math_verify |↑ |0.1960|Β± |0.0178|
61
+
62
+ 2026-08-10T15:00:54-07:00 code pass [humaneval,mbpp] via /v1/completions (function-continuation)
63
+ 2026-08-10:15:01:01 INFO [_cli.run:388] Selected Tasks: ['humaneval', 'mbpp']
64
+ 2026-08-10:15:01:02 INFO [evaluator:214] Setting random seed to 0 | Setting numpy seed to 1234 | Setting torch manual seed to 1234 | Setting fewshot manual seed to 1234
65
+ 2026-08-10:15:01:02 INFO [evaluator:239] Initializing local-completions model, with arguments: {'model': 'student', 'base_url': 'http://127.0.0.1:8521/v1/completions', 'tokenizer': 'outputs/release/unhealed/winnow-olmoe-math-keep50', 'num_concurrent': 48, 'tokenized_requests': False, 'max_retries': 3}
66
+ 2026-08-10:15:01:02 INFO [models.openai_completions:42] Remote tokenizer not supported. Using huggingface tokenizer backend.
67
+ 2026-08-10:15:01:02 INFO [models.api_models:179] Using max length 2048 - 1
68
+ 2026-08-10:15:01:02 INFO [models.api_models:200] Using tokenizer huggingface
69
+ 2026-08-10:15:01:10 INFO [evaluator_utils:446] Selected tasks:
70
+ 2026-08-10:15:01:10 INFO [evaluator_utils:480] Task: humaneval (humaneval/humaneval.yaml)
71
+ 2026-08-10:15:01:10 INFO [evaluator_utils:480] Task: mbpp (mbpp/mbpp.yaml)
72
+ 2026-08-10:15:01:10 INFO [evaluator:314] humaneval: Using gen_kwargs: {'until': ['\nclass', '\ndef', '\n#', '\nif', '\nprint'], 'max_gen_toks': 1024, 'do_sample': False}
73
+ 2026-08-10:15:01:10 INFO [evaluator:314] mbpp: Using gen_kwargs: {'until': ['[DONE]'], 'do_sample': False}
74
+ 2026-08-10:15:01:10 INFO [api.task:312] Building contexts for humaneval on rank 0...
75
+
76
  0%| | 0/164 [00:00<?, ?it/s]
77
  87%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‹ | 142/164 [00:00<00:00, 1419.11it/s]
78
+ 2026-08-10:15:01:10 INFO [api.task:312] Building contexts for mbpp on rank 0...
79
+
80
  0%| | 0/500 [00:00<?, ?it/s]
81
  3%|β–Ž | 17/500 [00:00<00:02, 161.93it/s]
82
  7%|β–‹ | 34/500 [00:00<00:02, 162.61it/s]
83
  10%|β–ˆ | 51/500 [00:00<00:02, 163.06it/s]
84
  14%|β–ˆβ–Ž | 68/500 [00:00<00:02, 163.43it/s]
85
  17%|β–ˆβ–‹ | 85/500 [00:00<00:02, 163.79it/s]
86
  20%|β–ˆβ–ˆ | 102/500 [00:00<00:02, 164.02it/s]
87
  24%|β–ˆβ–ˆβ– | 119/500 [00:00<00:02, 164.32it/s]
88
  27%|β–ˆβ–ˆβ–‹ | 136/500 [00:00<00:02, 164.49it/s]
89
  31%|β–ˆβ–ˆβ–ˆ | 153/500 [00:00<00:02, 164.68it/s]
90
  34%|β–ˆβ–ˆβ–ˆβ– | 170/500 [00:01<00:02, 164.78it/s]
91
  37%|β–ˆβ–ˆβ–ˆβ–‹ | 187/500 [00:01<00:01, 165.03it/s]
92
  41%|β–ˆβ–ˆβ–ˆβ–ˆ | 204/500 [00:01<00:01, 165.27it/s]
93
  44%|β–ˆβ–ˆβ–ˆβ–ˆβ– | 221/500 [00:01<00:01, 165.43it/s]
94
  48%|β–ˆβ–ˆβ–ˆβ–ˆβ–Š | 238/500 [00:01<00:01, 165.60it/s]
95
  51%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆ | 255/500 [00:01<00:01, 165.58it/s]
96
  54%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ– | 272/500 [00:01<00:01, 165.63it/s]
97
  58%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Š | 289/500 [00:01<00:01, 165.63it/s]
98
  61%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ | 306/500 [00:01<00:01, 165.70it/s]
99
  65%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ– | 323/500 [00:01<00:01, 165.80it/s]
100
  68%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Š | 340/500 [00:02<00:00, 165.89it/s]
101
  71%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ– | 357/500 [00:02<00:00, 165.80it/s]
102
  75%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ– | 374/500 [00:02<00:00, 165.88it/s]
103
  78%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Š | 391/500 [00:02<00:00, 165.74it/s]
104
  82%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ– | 408/500 [00:02<00:00, 165.84it/s]
105
  85%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Œ | 425/500 [00:02<00:00, 165.82it/s]
106
  88%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Š | 442/500 [00:02<00:00, 166.05it/s]
107
  92%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–| 459/500 [00:02<00:00, 165.99it/s]
108
  95%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Œ| 476/500 [00:02<00:00, 165.98it/s]
109
  99%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Š| 493/500 [00:02<00:00, 166.07it/s]
110
+ 2026-08-10:15:01:13 INFO [evaluator:585] Running generate_until requests
111
+ 2026-08-10:15:01:13 INFO [models.api_models:747] Tokenized requests are disabled. Context + generation length is not checked.
112
+
113
+
114
+ 2026-08-10:15:05:32 INFO [loggers.evaluation_tracker:247] Saving results aggregated
115
+ 2026-08-10:15:05:32 INFO [loggers.evaluation_tracker:119] Saving per-task samples to outputs/evals/math_unhealed/glean_winnow-olmoe-math-keep50/student/*.jsonl
116
+ local-completions ({'model': 'student', 'base_url': 'http://127.0.0.1:8521/v1/completions', 'tokenizer': 'outputs/release/unhealed/winnow-olmoe-math-keep50', 'num_concurrent': 48, 'tokenized_requests': False, 'max_retries': 3}), gen_kwargs: ({}), limit: None, num_fewshot: None, batch_size: 1
117
+ | Tasks |Version| Filter |n-shot| Metric | |Value | |Stderr|
118
+ |---------|------:|-----------|-----:|---------|---|-----:|---|-----:|
119
+ |humaneval| 1|create_test| 0|pass@1 |↑ |0.0732|Β± |0.0204|
120
+ |mbpp | 1|none | 3|pass_at_1|↑ |0.0900|Β± |0.0128|
121
+
122
+ 2026-08-10T15:05:33-07:00 lm_eval exit=0 -> outputs/evals/math_unhealed/glean_winnow-olmoe-math-keep50
evals/math_unhealed/glean_winnow-olmoe-math-keep75.eval.log ADDED
The diff for this file is too large to render. See raw diff
 
evals/math_unhealed/reap_reap-math-keep25.eval.log ADDED
The diff for this file is too large to render. See raw diff
 
evals/math_unhealed/reap_reap-math-keep50.eval.log ADDED
The diff for this file is too large to render. See raw diff