File size: 6,469 Bytes
709ea1b
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
---
license: mit
datasets:
- ragtruth
language:
- en
base_model: answerdotai/ModernBERT-large
pipeline_tag: token-classification
tags:
- hallucination-detection
- rag
- token-classification
- modernbert
- lettucedetect
---

# modernbert-large-ragtruth-token-level-binary

Fine-tuned `answerdotai/ModernBERT-large`, binary **token**-classification model for
**span-level** RAG hallucination detection — the same recipe as
[`hugoomezz/modernbert-ragtruth-token-level-binary`](https://huggingface.co/hugoomezz/modernbert-ragtruth-token-level-binary)
(Track B / arm b), scaled to the `-large` backbone, seed 42.

**This is the REPORTING model, not the deployed one.** Per
[ADR-021](https://github.com/hugoomez/rag-hallucination-detector/blob/main/docs/decisions.md#adr-021-3-seed-matched-scaling-comparison----modernbert-large-adopted-for-reporting),
the project deliberately keeps ModernBERT-**base** (arm b) as the model running in the
live demo — a latency/simplicity tradeoff — while adopting this large-backbone
checkpoint as the best-documented model for reporting the base→large scaling
comparison. The two roles are recorded separately on purpose: the headline number and
the running system are not required to be the same artifact, as long as which is which
is stated.

## Intended use

Research artifact for the base→large scaling comparison in
[ADR-021](https://github.com/hugoomez/rag-hallucination-detector/blob/main/docs/decisions.md)
and the paper's Appendix A. Not the model behind the live demo (see
`hugoomezz/modernbert-ragtruth-token-level-binary` for that). Not validated outside
RAGTruth's three task types (QA, Summary, Data2txt), and not intended for production
moderation decisions.

## Training data

[RAGTruth](https://github.com/ParticleMedia/RAGTruth) (Niu et al., 2024, ACL,
[arXiv:2401.00396](https://arxiv.org/abs/2401.00396)) — MIT-licensed. 13,578 train /
1,512 val / 2,700 test rows. Identical recipe to arm b (lr=1e-5, effective batch 8,
checkpoint selection on token-level F1), differing only in backbone (`ModernBERT-large`
vs `-base`) and seed count (this repo is seed 42 of a matched 3-seed sweep — see ADR-021
for the full 3-seed mean/range; this card reports seed 42's own numbers, not the mean).

## Metrics (RAGTruth test set, n=2700, seed 42, checkpoint selected at epoch 2/step 3396 on token-F1)

| Metric | Value |
|---|---|
| Response-level Precision | 0.8530 |
| Response-level Recall | 0.7444 |
| Response-level F1 | 0.7950 |
| Response-level Accuracy | 0.8659 |
| Span-level (char-overlap) Precision | 0.6954 |
| Span-level (char-overlap) Recall | 0.4927 |
| Span-level (char-overlap) F1 | **0.5767** |

Per task_type (response-level derived):

| Task | F1 | Recall |
|---|---|---|
| Data2txt | 0.8803 | 0.8636 |
| QA | 0.7290 | 0.7063 |
| Summary | 0.5563 | 0.4363 |

vs. arm b (ModernBERT-base, seed 42): response F1 0.7631, span F1 0.5321 — this
checkpoint improves response F1 by +0.0319 and span F1 by +0.0446 at this seed,
recall-led rather than precision-led (the same pattern ADR-012 found moving from
DeBERTa to ModernBERT-base). See ADR-021 for the full 3-seed picture (mean +0.0311 /
+0.0408) and its caveats — the epoch cap was not uniform across all seeds, and one
large-backbone seed (123) has unrecorded training configuration.

## Limitations

Same structural limitations as arm b (Summary is the weakest task; "Subtle"
hallucinations are the hardest case; checkpoint selection uses token-F1, not the
clean-span metric the paper's hypothesis is actually about — see the paper's §8
Limitations). Specific to this checkpoint:

- **Single-seed card.** This card characterizes seed 42 only. ADR-021's 3-seed mean is
  the number the paper actually reports as the headline scaling result; treat this
  card's numbers as one seed's contribution to that mean, not as the finding itself.
- **Not the deployed model.** If you want the model behind this project's live demo,
  use `hugoomezz/modernbert-ragtruth-token-level-binary` (arm b, ModernBERT-base)
  instead.
- **Resource-constrained scaling replication is methodologically separate from this
  project's main argument.** Per the paper's Appendix A: "tests nothing about
  `implicit_true`, contributes no evidence for or against §4, §5, or §6, and would be
  removed without affecting any claim the paper makes."

## How to use

```python
import torch
from transformers import AutoModelForTokenClassification, AutoTokenizer

# The tokenizer is loaded from the base ModernBERT-large repo, not the fine-tuned one,
# mirroring the same substitution used by hugoomezz/modernbert-ragtruth-token-level-binary
# (arm b) -- training data was built with this exact base tokenizer, so it's exact, not
# an approximation.
TOKENIZER_ID = "answerdotai/ModernBERT-large"
MODEL_ID = "hugoomezz/modernbert-large-ragtruth-token-level-binary"
SUPPORTED, HALLUCINATED = 0, 1

tokenizer = AutoTokenizer.from_pretrained(TOKENIZER_ID)
model = AutoModelForTokenClassification.from_pretrained(MODEL_ID, attn_implementation="sdpa").eval()

context = "The Eiffel Tower was completed in 1889 for the World's Fair in Paris."
response = "The Eiffel Tower was completed in 1889 and is located in Berlin, Germany."

encoding = tokenizer(
    context, response, max_length=4096, truncation="only_first",
    return_offsets_mapping=True, return_token_type_ids=False, return_tensors="pt",
)
sequence_ids = encoding.sequence_ids(0)
encoding.pop("offset_mapping")

with torch.no_grad():
    logits = model(**encoding).logits[0]
probs_hallucinated = torch.softmax(logits, dim=-1)[:, HALLUCINATED].tolist()

response_probs = [p for p, sid in zip(probs_hallucinated, sequence_ids) if sid == 1]
score = max(response_probs) if response_probs else 0.0
print(f"hallucination score: {score:.4f}  ({'FLAGGED' if score >= 0.5 else 'clean'})")
```

## Citation

```bibtex
@inproceedings{niu2024ragtruth,
  title     = {RAGTruth: A Hallucination Corpus for Developing and Evaluating RAG Systems},
  author    = {Niu, Cheng and Wu, Yuanhao and Zhu, Juno and Xu, Siliang and Shum, Kashun and Zhong, Randy and Song, Juntong and Zhang, Tong},
  booktitle = {Proceedings of ACL 2024},
  year      = {2024},
  eprint    = {2401.00396}
}
@article{kovacs2025lettucedetect,
  title   = {LettuceDetect: A Hallucination Detection Framework for RAG Applications},
  author  = {Kov{\'a}cs, {\'A}d{\'a}m and Bakos, Zsolt},
  journal = {arXiv preprint arXiv:2502.17125},
  year    = {2025}
}
```