2026-07-17T11:15:57-07:00 serving allenai/OLMoE-1B-7B-0125-Instruct on GPU 0 port 8399 2026-07-17T11:15:57-07:00 waiting for server /health ... 2026-07-17T11:16:27-07:00 server up; chat pass [minerva_math500] 2026-07-17:11:16:27 WARNING [config.evaluate_config:287] --limit SHOULD ONLY BE USED FOR TESTING. REAL METRICS SHOULD NOT BE COMPUTED USING LIMIT. 2026-07-17:11:16:34 INFO [_cli.run:388] Selected Tasks: ['minerva_math500'] 2026-07-17:11:16:35 INFO [evaluator:214] Setting random seed to 0 | Setting numpy seed to 1234 | Setting torch manual seed to 1234 | Setting fewshot manual seed to 1234 2026-07-17:11:16:35 WARNING [evaluator:226] generation_kwargs: {'max_gen_toks': 1280} specified through cli, these settings will update set parameters in yaml tasks. Ensure 'do_sample=True' for non-greedy decoding! 2026-07-17:11:16:35 INFO [evaluator:239] Initializing local-chat-completions model, with arguments: {'model': 'student', 'base_url': 'http://127.0.0.1:8399/v1/chat/completions', 'num_concurrent': 32, 'tokenized_requests': False, 'max_retries': 3} 2026-07-17:11:16:35 INFO [models.api_models:179] Using max length 2048 - 1 2026-07-17:11:16:35 INFO [models.api_models:200] Using tokenizer None 2026-07-17:11:16:37 INFO [evaluator_utils:446] Selected tasks: 2026-07-17:11:16:37 INFO [evaluator_utils:480] Task: minerva_math500 (minerva_math/minerva_math500.yaml) 2026-07-17:11:16:37 INFO [evaluator:314] minerva_math500: Using gen_kwargs: {'until': ['Problem:'], 'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 1280} 2026-07-17:11:16:37 INFO [api.task:312] Building contexts for minerva_math500 on rank 0... 0%| | 0/100 [00:00 outputs/evals/general_suite/math_fix