2026-07-23:22:03:14 WARNING [config.evaluate_config:287] --limit SHOULD ONLY BE USED FOR TESTING. REAL METRICS SHOULD NOT BE COMPUTED USING LIMIT. 2026-07-23:22:03:21 INFO [_cli.run:388] Selected Tasks: ['mmlu_pro', 'gpqa_diamond_zeroshot', 'minerva_math500', 'ifeval', 'gsm8k_cot_zeroshot'] 2026-07-23:22:03:23 INFO [evaluator:214] Setting random seed to 0 | Setting numpy seed to 1234 | Setting torch manual seed to 1234 | Setting fewshot manual seed to 1234 2026-07-23:22:03:23 WARNING [evaluator:226] generation_kwargs: {'max_gen_toks': 1280} specified through cli, these settings will update set parameters in yaml tasks. Ensure 'do_sample=True' for non-greedy decoding! 2026-07-23:22:03:23 INFO [evaluator:239] Initializing local-chat-completions model, with arguments: {'model': 'student', 'base_url': 'http://127.0.0.1:8399/v1/chat/completions', 'num_concurrent': 32, 'tokenized_requests': False, 'max_retries': 3} 2026-07-23:22:03:23 INFO [models.api_models:179] Using max length 2048 - 1 2026-07-23:22:03:23 INFO [models.api_models:200] Using tokenizer None 2026-07-23:22:03:38 INFO [evaluator_utils:446] Selected tasks: 2026-07-23:22:03:38 INFO [evaluator_utils:462] Group: mmlu_pro 2026-07-23:22:03:38 INFO [evaluator_utils:470] Task: mmlu_pro_biology (mmlu_pro/mmlu_pro_biology.yaml) 2026-07-23:22:03:38 INFO [evaluator_utils:470] Task: mmlu_pro_business (mmlu_pro/mmlu_pro_business.yaml) 2026-07-23:22:03:38 INFO [evaluator_utils:470] Task: mmlu_pro_chemistry (mmlu_pro/mmlu_pro_chemistry.yaml) 2026-07-23:22:03:38 INFO [evaluator_utils:470] Task: mmlu_pro_computer_science (mmlu_pro/mmlu_pro_computer_science.yaml) 2026-07-23:22:03:38 INFO [evaluator_utils:470] Task: mmlu_pro_economics (mmlu_pro/mmlu_pro_economics.yaml) 2026-07-23:22:03:38 INFO [evaluator_utils:470] Task: mmlu_pro_engineering (mmlu_pro/mmlu_pro_engineering.yaml) 2026-07-23:22:03:38 INFO [evaluator_utils:470] Task: mmlu_pro_health (mmlu_pro/mmlu_pro_health.yaml) 2026-07-23:22:03:38 INFO [evaluator_utils:470] Task: mmlu_pro_history (mmlu_pro/mmlu_pro_history.yaml) 2026-07-23:22:03:38 INFO [evaluator_utils:470] Task: mmlu_pro_law (mmlu_pro/mmlu_pro_law.yaml) 2026-07-23:22:03:38 INFO [evaluator_utils:470] Task: mmlu_pro_math (mmlu_pro/mmlu_pro_math.yaml) 2026-07-23:22:03:38 INFO [evaluator_utils:470] Task: mmlu_pro_other (mmlu_pro/mmlu_pro_other.yaml) 2026-07-23:22:03:38 INFO [evaluator_utils:470] Task: mmlu_pro_philosophy (mmlu_pro/mmlu_pro_philosophy.yaml) 2026-07-23:22:03:38 INFO [evaluator_utils:470] Task: mmlu_pro_physics (mmlu_pro/mmlu_pro_physics.yaml) 2026-07-23:22:03:38 INFO [evaluator_utils:470] Task: mmlu_pro_psychology (mmlu_pro/mmlu_pro_psychology.yaml) 2026-07-23:22:03:38 INFO [evaluator_utils:480] Task: gpqa_diamond_zeroshot (gpqa/zeroshot/gpqa_diamond_zeroshot.yaml) 2026-07-23:22:03:38 INFO [evaluator_utils:480] Task: gsm8k_cot_zeroshot (gsm8k/gsm8k-cot-zeroshot.yaml) 2026-07-23:22:03:38 INFO [evaluator_utils:480] Task: ifeval (ifeval/ifeval.yaml) 2026-07-23:22:03:38 INFO [evaluator_utils:480] Task: minerva_math500 (minerva_math/minerva_math500.yaml) 2026-07-23:22:03:38 INFO [evaluator:314] mmlu_pro_biology: Using gen_kwargs: {'until': ['Question:'], 'max_gen_toks': 1280, 'do_sample': False, 'temperature': 0.0} 2026-07-23:22:03:38 INFO [evaluator:314] mmlu_pro_business: Using gen_kwargs: {'until': ['Question:'], 'max_gen_toks': 1280, 'do_sample': False, 'temperature': 0.0} 2026-07-23:22:03:38 INFO [evaluator:314] mmlu_pro_chemistry: Using gen_kwargs: {'until': ['Question:'], 'max_gen_toks': 1280, 'do_sample': False, 'temperature': 0.0} 2026-07-23:22:03:38 INFO [evaluator:314] mmlu_pro_computer_science: Using gen_kwargs: {'until': ['Question:'], 'max_gen_toks': 1280, 'do_sample': False, 'temperature': 0.0} 2026-07-23:22:03:38 INFO [evaluator:314] mmlu_pro_economics: Using gen_kwargs: {'until': ['Question:'], 'max_gen_toks': 1280, 'do_sample': False, 'temperature': 0.0} 2026-07-23:22:03:38 INFO [evaluator:314] mmlu_pro_engineering: Using gen_kwargs: {'until': ['Question:'], 'max_gen_toks': 1280, 'do_sample': False, 'temperature': 0.0} 2026-07-23:22:03:38 INFO [evaluator:314] mmlu_pro_health: Using gen_kwargs: {'until': ['Question:'], 'max_gen_toks': 1280, 'do_sample': False, 'temperature': 0.0} 2026-07-23:22:03:38 INFO [evaluator:314] mmlu_pro_history: Using gen_kwargs: {'until': ['Question:'], 'max_gen_toks': 1280, 'do_sample': False, 'temperature': 0.0} 2026-07-23:22:03:38 INFO [evaluator:314] mmlu_pro_law: Using gen_kwargs: {'until': ['Question:'], 'max_gen_toks': 1280, 'do_sample': False, 'temperature': 0.0} 2026-07-23:22:03:38 INFO [evaluator:314] mmlu_pro_math: Using gen_kwargs: {'until': ['Question:'], 'max_gen_toks': 1280, 'do_sample': False, 'temperature': 0.0} 2026-07-23:22:03:38 INFO [evaluator:314] mmlu_pro_other: Using gen_kwargs: {'until': ['Question:'], 'max_gen_toks': 1280, 'do_sample': False, 'temperature': 0.0} 2026-07-23:22:03:38 INFO [evaluator:314] mmlu_pro_philosophy: Using gen_kwargs: {'until': ['Question:'], 'max_gen_toks': 1280, 'do_sample': False, 'temperature': 0.0} 2026-07-23:22:03:38 INFO [evaluator:314] mmlu_pro_physics: Using gen_kwargs: {'until': ['Question:'], 'max_gen_toks': 1280, 'do_sample': False, 'temperature': 0.0} 2026-07-23:22:03:38 INFO [evaluator:314] mmlu_pro_psychology: Using gen_kwargs: {'until': ['Question:'], 'max_gen_toks': 1280, 'do_sample': False, 'temperature': 0.0} 2026-07-23:22:03:38 INFO [evaluator:314] minerva_math500: Using gen_kwargs: {'until': ['Problem:'], 'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 1280} 2026-07-23:22:03:38 INFO [evaluator:314] ifeval: Using gen_kwargs: {'until': [], 'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 1280} 2026-07-23:22:03:38 INFO [evaluator:314] gsm8k_cot_zeroshot: Using gen_kwargs: {'until': ['Q:', '', '<|im_end|>'], 'do_sample': False, 'max_gen_toks': 1280} 2026-07-23:22:03:38 INFO [api.task:312] Building contexts for mmlu_pro_biology on rank 0... 0%| | 0/10 [00:00 sys.exit(cli_evaluate()) ^^^^^^^^^^^^^^ File "/home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/lm_eval/__main__.py", line 10, in cli_evaluate parser.execute(args) File "/home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/lm_eval/_cli/harness.py", line 60, in execute args.func(args) File "/home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/lm_eval/_cli/run.py", line 391, in _execute results = simple_evaluate( ^^^^^^^^^^^^^^^^ File "/home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/lm_eval/utils.py", line 575, in _wrapper return fn(*args, **kwargs) ^^^^^^^^^^^^^^^^^^^ File "/home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/lm_eval/evaluator.py", line 358, in simple_evaluate results = evaluate( ^^^^^^^^^ File "/home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/lm_eval/utils.py", line 575, in _wrapper return fn(*args, **kwargs) ^^^^^^^^^^^^^^^^^^^ File "/home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/lm_eval/evaluator.py", line 596, in evaluate resps = getattr(lm, reqtype)(cloned_reqs) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/lm_eval/models/openai_completions.py", line 239, in loglikelihood raise NotImplementedError( NotImplementedError: Loglikelihood is not supported for chat completions. Consider using the completions API instead.