Jev thinks "I don't know'', but doesn't say it: Introducing Sys1Cal-v1 Dataset for Probability Calibration
Abstract
The appearance of Jev marked the era of System One Models, foundation models that return structured decisions with probability distributions rather than text. Aside from low cost and great speed, Jev's central promise is that these probabilities are calibrated: such claim is not backed by any public test and available external benchmarks evaluate confidence calibration, not whether every returned option probability has the right numerical meaning. To tackle this issue, we introduce Sys1Cal-v1, a dataset of True/False questions about a proposition A for which the exact probability P(A) is known by construction. Each item is queried through the three Jev primitives - Noul, Choice and Score - and evaluated by total variation distance from the ground-truth distribution, which can be used to estimate a soft accuracy of System One Models. We showcase the utility of Sys1Cal-v1 as a benchmark dataset by evaluating Jev and SemIf, an open-source Choice-style baseline. In this work, however, we focus even more deeply on Jev, by studying the calibration of its Score and Choice answers. In particular, we discover a peculiar behaviour that can be explained by assuming that Jev suppresses a third truth value, going beyond True and False. In other words, in Choice answers, P(A) and P(neg A) are presented as if P(A)+P(neg A)=1, while a term P(U)neq0 is missing in the sum. Recovering P(U) leads to an improvement of median soft accuracy in Choice answers from 0.771 to 0.978, suggesting that, even in binary decisions, Jev wants to answer with a third option:``I don't know''.
Community
Sys1Cal-v1 asks a simple question: when a System One Model returns a probability, does that probability actually have the numerical meaning we assume it has?
We introduce a benchmark in which the exact probability of every proposition is known by construction, allowing us to evaluate structured probabilistic outputs directly rather than only checking confidence in the predicted class.
Applying Sys1Cal-v1 to Jev reveals a clear asymmetry between its primitives: Noul and Score remain substantially closer to the ground-truth probabilities, while binary Choice outputs exhibit a systematic nonlinear distortion.
We show that this distortion can be modeled by introducing a latent third component, Uncertain, alongside True and False. Under this model, Choice behaves as if the uncertain mass were removed and the remaining True/False probabilities renormalized.
This matters beyond calibration. Once uncertainty is represented explicitly, the same model output can support a different downstream action: a binary distribution may justify committing to an outcome, while the corresponding latent representation can justify deferral when the cost of error is high.
The broader question I would like to explore is whether this phenomenon is specific to Jev, or whether forcing richer probabilistic representations into binary or categorical outputs systematically suppresses decision-relevant uncertainty in other models as well.
Sys1Cal-v1 and the evaluation code are public, so replication on other structured decision models is very welcome.
Get this paper in your agent:
hf papers read 2609.35342 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 1
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper