• JevBench v1.4.1, 231 public items (self-run with the official v1.4.1 harness, not an official board entry): 73.6% (170/231), the highest public accuracy among the Qwen3.5-2B-family systems on the board (decider-2b 71.0%, open-jev-zefan-2b 64.5%; the lead over decider-2b is inside the 95% CI, 67.6–78.9%). 42 of the 82 board systems score higher, almost all 4B or larger. +9.5 points over our 0.8B v3. Jev 1.13 is well ahead at 86.6%. • tweet_topic, zero-shot (n = 1,693): 82.2% accuracy [80.4, 84.0], above Jev's 79.3%; macro-F1 is below Jev (0.678 vs 0.694). • fin_topic, zero-shot (n = 4,117): 61.1%, +14.4 points over our 0.8B v3 (Jev: 67.0%). • 25,600 tokens per call, no option cap, nothing truncated. • Read once, then ask: on an Apple M1 Max (GGUF F16), the first question about a 24,501-token input took 16.2 s and a further question about the same state 0.17 s (shared machine, indicative).
Contamination checks against the training pool: 0 hits for the JevBench items, and no exact or near-duplicate overlap with the tweet_topic and fin_topic test sets. Needs the shipped runtimes (PyTorch, llama.cpp with the bundled scorer, MLX); trained on a reduced data pool (60M tokens).