JevBench
AI & ML interests
Evaluation, benchmarking, metamorphic testing, probabilistic coherence, calibration
Recent Activity
JevBench: metamorphic coherence testing for typed probabilistic decision models.
Typed decision models answer questions about a text with probability distributions over declared answers: yes or no, one of several options, or a level on an ordered scale. JevBench tests whether those probabilities stay coherent when a question is reworded, negated, logically combined with others, offered a different set of options, or asked together with other questions: 50 metamorphic relations in five dimensions, no gold labels needed, every test frozen in advance. Across 21 openly released models, batch independence mostly holds, while additivity fails broadly, above all under negation.
- 📄 Technical report: alphaXiv
- 🗂️ Dataset: JevBench/jevbench, 240 cases, 2,240 questions, frozen test suites and every evaluation report
- 🏆 Leaderboard: jevbench.github.io
- 💻 Code: github.com/JevBench/jevbench,
pip install jevbench
Maintained by Chen Feng, ML Lab, Queen's University Belfast.