EASE-Delta Reader (large)

Does this passage support the claim, refute it, or not settle it at all?

A 395M-parameter cross-encoder (answerdotai/ModernBERT-large fine-tuned) that reads one claim against one passage and answers SUPPORTS, REFUTES or NOT_ENOUGH_INFO. It was trained to say NOT_ENOUGH_INFO when a passage is about something else, which is where general NLI models most often guess: on passages taken from unrelated documents it gives a decisive answer 0.58% of the time; a widely used public NLI model does so 45.07% of the time.

It is the reader inside EASE-Delta, a system that keeps a task's decisions current as its messages and documents change. Try both in the demo.

Use it

Claim first, evidence second.

from transformers import pipeline

reader = pipeline("text-classification", model="jithinpothireddy21/ease-delta-reader", top_k=None)
claim = "The client has approved the final design."
passages = [
    "Email from the client: we approve the final design, please go ahead.",
    "Email from the client: we cannot approve the design yet; the colours are wrong.",
    "The office is closed on the first Monday of the month for staff training.",
]
for passage in passages:
    print(reader({"text": claim, "text_pair": passage}, truncation=True)[0])

# {'label': 'SUPPORTS', 'score': 0.935}
# {'label': 'REFUTES', 'score': 0.991}
# {'label': 'NOT_ENOUGH_INFO', 'score': 0.999}

With sentence-transformers, continuing from above:

from sentence_transformers import CrossEncoder

model = CrossEncoder("jithinpothireddy21/ease-delta-reader")
probs = model.predict([(claim, passage) for passage in passages], apply_softmax=True)
# one row per pair; columns SUPPORTS, REFUTES, NOT_ENOUGH_INFO

Inputs were at most 256 tokens in training (truncation=True cuts there). Split long documents into passages and read each one; say who is speaking in the passage itself ("Email from the client: ..."), because the reader sees only the text.

Probabilities are best calibrated after dividing the logits by 1.203 (fitted on development data) before the softmax.

Results

Test sets, accuracy. One evaluation harness for all three models; brackets and paired tests are in the results report.

This model (large) base reader tasksource/ModernBERT-base-nli
Passages from unrelated documents read as decisive (lower is better) 0.58% 0.82% 45.07%
VitaminC: claims against Wikipedia revisions 91.53% 90.20% 68.30%
MultiNLI, mismatched genres 90.21% 88.56% 89.84%
WANLI 77.42% 74.74% 66.34%
SNLI (not trained on) 84.39% 80.33% 89.15%
ANLI round 1 (not trained on) 53.70% 46.40% 63.40%
ANLI round 2 (not trained on) 38.10% 33.30% 48.20%
ANLI round 3 (not trained on) 36.00% 34.75% 42.25%

tasksource/ModernBERT-base-nli is a base-size multi-task NLI model whose training includes SNLI and ANLI; it is the better choice for adversarial NLI of the ANLI kind. This reader is built for evidence: contrastive Wikipedia revisions (VitaminC), and knowing when a passage does not bear on the claim.

On VitaminC, real revisions: 89.90%; synthetic: 94.25%.

Labels

Label Meaning NLI equivalent
SUPPORTS the passage establishes the claim entailment
REFUTES the passage establishes that the claim is false contradiction
NOT_ENOUGH_INFO the passage does not settle the claim, including when it is about something else neutral

Training

900,000 examples streamed from the Hugging Face Hub: VitaminC (cc-by-sa-3.0), MultiNLI and WANLI (cc-by-4.0), plus 20% synthetic unrelated pairs (the claim of one example with the evidence of another from a different topic, labelled NOT_ENOUGH_INFO). SNLI was left out because its annotators labelled unrelated content as contradiction; ANLI because its licence is non-commercial. One run, on one Apple M4 Max. The evaluation plan was written and hashed before training: docs/PREREGISTRATION_LARGE.md.

Limitations

  • English only.
  • Passages on the claim's own subject that do not settle it are still read as decisive about one time in five (VitaminC test, not-enough-info items).
  • Conflicts that follow only from a consequence ("broke a leg" against "cycles to work") are mostly missed.
  • Numbers and dates are read as text; compare them with code when they matter.
  • Not a judge of truth: it reports what the passage says, not whether the passage is right.

How this checkpoint was made

Converted exactly from the EASE-Delta edge model (ease/export.py in the code repository explains why the conversion is exact). On 1,200 development pairs the largest logit difference between the two was 9.5e-07 and the predicted label never differed (0 disagreements). The full system, with its calibrated rules and regression suites, is jithinpothireddy21/ease-delta.

Other size: jithinpothireddy21/ease-delta-reader-base.

Licence

CC BY-SA 4.0 for the weights, because VitaminC is share-alike (Creative Commons' guidance on AI training describes this as the cautious course). The code is Apache-2.0. Credit for the training data: VitaminC (Schuster, Fisch and Barzilay, NAACL 2021), MultiNLI (Williams, Nangia and Bowman, NAACL 2018), WANLI (Liu, Swayamdipta, Smith and Choi, EMNLP 2022). Backbone: ModernBERT (Warner et al., 2024).

Citation

@software{pothireddy2026easedelta,
  author = {Pothireddy, Jithin},
  title = {EASE-Delta: revision-aware decision computation},
  year = {2026},
  url = {https://github.com/jithinsaireddy/ease-delta}
}
Downloads last month
13
Safetensors
Model size
0.4B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for jithinpothireddy21/ease-delta-reader

Finetuned
(395)
this model

Datasets used to train jithinpothireddy21/ease-delta-reader

Space using jithinpothireddy21/ease-delta-reader 1

Evaluation results