[eval] loading benchmark from ../outputs/benchmark/benchmark_labelled.jsonl (split=test) [eval] 1743 questions [eval] loading policy from ../checkpoints/sft-llama31/final Loading weights: 0%| | 0/291 [00:00