#!/usr/bin/env python """JevBench (github.com/fstandhartinger/jevbench) public items with a jev NLI cross-encoder. python eval_jevbench.py --models ckpt/qwen3.5-0.8b-nli-v2s --out results/v2s/jevbench.json Only the public items exist outside the benchmark (easy 48, standard 72 = `original`, hard 111; judge and held-out items are private), so the numbers are PUBLIC-ITEM accuracy and not the JevBench Score. For a like-for-like comparison the leaderboard systems are re-scored on exactly these items from JevBench's per-task artifact. Every option becomes one hypothesis over the state; the option distribution is P(entailment) normalised over the options (native, one forward pass per option, nothing generated): choice "The answer to "" is