Seb-9B vs JEV-27B-VL on identical rows: 17 Laya public suites, a 6,920-row decision suite and two image suites (self-run)

#4
by ironbcc - opened

Disclosure: I built Seb-9B (https://huggingface.co/ironbcc/seb-9b). This is a self-run comparison, not an independent one. Every result is reported as measured, including the ones JEV-27B-VL wins.

Setup

  • JEV-27B-VL at revision f34b598, served with the repo's serve_decide.py and the card's flags (vLLM 0.30, --max-num-seqs 8, one RTX PRO 6000). Every question went through POST /v1/decide, one question per call.
  • Mapping: noul stays noul, with the true/false criteria appended to the question. choice stays choice, with "name: description" options. Score questions become choice over the ordered level descriptions, because your score kind is a fixed 0–5 scale with no level text.
  • I tried two renderings and report the one that scored better for JEV-27B-VL. On the 6,920-row suite, appending the criteria raised it from 0.896 to 0.918. The Laya suites didn't change (0.776 vs 0.775).
  • Seb-9B and TypeSafe Jev 1.13.0 (API) predictions come from earlier runs on the same rows.

Results (same rows for every model):

Seb-9B Jev 1.13 JEV-27B-VL Seb − JEV-27B-VL (95% CI)
Laya public suites, mean accuracy (17) 0.792 0.786 0.775 +0.018 [+0.011, +0.025]
Laya public suites, mean ECE 0.088 0.124 0.115
6,920-row decision suite, family-weighted accuracy 0.946 0.943 0.918 +0.028 [+0.015, +0.050]
6,920-row decision suite, ECE 1.97% 3.84% 5.07%
Held-out image suite (2,451 rows) 0.876 rejects images 0.854 +0.023 [+0.012, +0.033]
COCO image suite (394 rows) 0.967 rejects images 0.977 −0.010 [−0.025, +0.003]

The CIs come from a paired bootstrap over rows (by group for the 6,920-row suite). The Laya suites were rebuilt with Laya's own sampling code; the 17-suite list and the Seb, Jev and Laya numbers are in https://huggingface.co/convaiinnovations/laya/discussions/30.

Where JEV-27B-VL wins:

  • COCO images
  • banking77 (0.925 vs 0.907), and all 77 banking77 labels in one question (0.828; Seb takes at most 20 options per question)
  • jailbreak detection (0.910 vs 0.868) and phishing (0.887 vs 0.863)
  • RAG relevance (0.645, best of the four models I ran) and MASSIVE intent (0.950 vs 0.943)

Caveats:

  • The 6,920-row suite is mine. About 45 Seb candidates were compared on it before release, so Seb's number carries selection optimism. JEV-27B-VL saw it once.
  • Seb's training data includes the typed-decisions benchmark's train split, so that Laya row isn't zero-shot for Seb. Without it, the Laya gap is +0.017 [+0.009, +0.024].
  • The held-out image suite comes from the same sources as Seb's image training data (different images). JEV-27B-VL answers images zero-shot, as your card notes. COCO is the fairer image test, and JEV-27B-VL is slightly ahead there.
  • I didn't rerun your six-benchmark table, and I didn't compare latency: the serving setups and loads differ.
  • Seb is 9B against 27B. It has no System 2 mode, and its choice questions take at most 20 options.

If my request mapping is unfair to JEV-27B-VL anywhere, especially on score questions, tell me and I'll rescore.

cloudyu changed discussion status to closed

Sign up or log in to comment