File size: 9,200 Bytes
f25b33e
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
---

license: mit
datasets:
- ragtruth
language:
- en
base_model: answerdotai/ModernBERT-base
pipeline_tag: token-classification
tags:
- hallucination-detection
- rag
- token-classification
- modernbert
- lettucedetect
---


# modernbert-ragtruth-token-level-binary (Track B)

Fine-tuned `answerdotai/ModernBERT-base`, binary **token**-classification model for
**span-level** RAG hallucination detection: given a `(context, response)` pair, labels
every response token supported (0) or hallucinated (1), so it recovers character-level
spans, not just a single per-response score.

**This is the best-performing system in this project (F1 0.7631 at the response level)

and the model deployed in the live demo** (`src/models/predict.py`), superseding both
Track A (`hugoomezz/deberta-v3-ragtruth-hallucination`) and Approach 1
(`hugoomezz/modernbert-ragtruth-response-level`). Its recipe matches
[LettuceDetect](https://arxiv.org/abs/2502.17125)'s approach: binary token labels,
unweighted loss, and character-overlap span evaluation, with checkpoint selection on
span-level F1 rather than response-level F1 (see [ADR-020](https://github.com/hugoomez/rag-hallucination-detector/blob/main/docs/decisions.md)
below).

## Intended use

Research / portfolio demonstration of span-level RAG hallucination detection on
RAGTruth: highlighting exactly which characters of a generated response are
unsupported by the retrieved context, not just a binary verdict. Powers this project's
live demo (paste-your-own text, and a real RAG pipeline over a small Wikipedia corpus).
Not validated outside RAGTruth's three task types (QA, Summary, Data2txt), and not
intended for production moderation decisions without further evaluation on your own
data and threshold.

## Training data

[RAGTruth](https://github.com/ParticleMedia/RAGTruth) (Niu et al., 2024, ACL,
[arXiv:2401.00396](https://arxiv.org/abs/2401.00396)) — MIT-licensed, reproduced in
this project's [docs/THIRD_PARTY_LICENSES.md](https://github.com/hugoomez/rag-hallucination-detector/blob/main/docs/THIRD_PARTY_LICENSES.md).
13,578 train / 1,512 val / 2,700 test rows (one more val row than Track A — see ADR-006/
ADR-011), tokenized at `max_length=4096` (0% of rows truncated). Per-token binary labels: 0 = supported, 1 = hallucinated (any character
overlap with a gold annotated span); context and special tokens are ignored in the
loss (label -100). Plain cross-entropy, **no class weighting** — an earlier 3-class BIO
scheme with inverse-frequency weighting on an ultra-rare class caused near-total span
fragmentation (0.037 F1); this binary/unweighted redesign fixed it.

## Metrics (RAGTruth test set, n=2700)

Response-level (a response is "predicted hallucinated" iff any response token is
predicted positive) — reported figures from this project's unified cross-system
evaluation, matching the deployed model and the README's comparison table:

| Metric | Value |
|---|---|
| Precision | 0.8359 |
| Recall | 0.7020 |
| F1 | 0.7631 |
| Accuracy | 0.8478 |

This matches [LettuceDetect-base](https://arxiv.org/abs/2502.17125)'s published
example-level F1 of 76.07%, from `results/finetuned_track_b_token_level_metrics.json`.

Span-level (character-overlap, LettuceDetect's headline metric; from
`results/finetuned_track_b_token_level_metrics.json`):

| Metric | Value |
|---|---|
| Precision | 0.6474 |
| Recall | 0.4517 |
| F1 | **0.5321** |

vs. LettuceDetect-base's published span-level F1 of 55.44%.

Per task_type (response-level derived):



| Task | F1 | Recall |

|---|---|---|

| Data2txt | 0.8675 | 0.8256 |

| QA | 0.6708 | 0.6688 |

| Summary | 0.4904 | 0.3775 |



**Recipe correction ([ADR-020](https://github.com/hugoomez/rag-hallucination-detector/blob/main/docs/decisions.md#adr-020-acws-ablation-results--recipe-fix-arm-b-adopted-noise-down-weighting-arm-c-rejected)):**

these are the weights from a controlled ablation's arm (b) — a faithful

LettuceDetect-recipe replication (lr=1e-5, effective batch 8, 6 epochs, checkpoint

selection on **span-level F1** instead of response-level F1). The prior deployed

weights (arm a) selected the checkpoint on response-level F1, which is structurally

unable to distinguish tight spans from sloppy ones and was suppressing span-level

performance; switching the selection metric alone (no architecture or data change)

raised span-F1 from 0.5113 to 0.5321 (+2.1 points) and response precision from 0.7873

to 0.8359, at the cost of some recall (0.7381 → 0.7020). A companion arm testing

Annotation-Confidence-Weighted Supervision (down-weighting the training loss on

annotator-flagged "implicit_true" spans) did not clear its pre-registered bar and was
rejected as a tested negative result — see the ADR for both arms' full numbers.

## Limitations

- **Summary remains the weakest task** (F1 0.51, recall 0.436 — misses more than half
  of hallucinated summaries), though improved over Track A/Approach 1.
- **Overconfident, not paranoid**: misses 26.2% of hallucinated responses but
  false-alarms on only 10.7% of faithful ones (247 FN vs 188 FP in this project's error
  analysis) — the token-level decision rule (flag if any token crosses P≥0.5)
  structurally favors silence over alarm.
- **"Subtle" hallucinations are the hardest case**: responses annotated only with
  "Subtle" span types are missed 40.3% of the time (vs 27.0% for evident-only), and
  detection F1 drops to ≈0.48–0.52 on GPT-3.5/GPT-4 outputs specifically — the
  deployment scenario (strong generators, subtle errors) where detection matters most.
- **Known false positive on close paraphrase**: during live demo testing, the model
  flagged a factually *correct* grounded answer that paraphrased the source ("second of
  six children, five siblings") as unsupported at score 0.99 — plausibly triggered by
  surface-form sensitivity rather than a genuine factual disagreement. Noted as a known
  limitation, not further investigated (`docs/notes.md`, Phase 5 section). This is the
  inverse of the "overconfident" pattern above: mostly the model under-flags, but this
  shows it can occasionally over-flag on close paraphrases.
- **The decision threshold is a product tradeoff, not a fixed answer**: F1 is nearly
  flat across thresholds 0.2–0.7, so the same checkpoint can be run as a high-precision
  "block" mode (t=0.9: precision 0.879, recall 0.602) or a high-recall "warn" mode
  (t=0.1: recall 0.843, precision 0.642). Even at the most aggressive setting, ~16% of
  hallucinations slip through — a risk reducer, not a guarantee.

## How to use

This snippet reproduces the response-level score (max P(hallucinated) over response
tokens); see `src/models/predict.py` in the project repo for full character-span
reconstruction (`merge_predicted_spans`).

```python

import torch

from transformers import AutoModelForTokenClassification, AutoTokenizer



# The tokenizer is loaded from the base ModernBERT repo, not the fine-tuned one: this

# repo's tokenizer_config.json was written by a newer transformers than some pinned

# installs can parse. Training data was built with this exact base tokenizer, so the

# substitution is not just safe but exact.

TOKENIZER_ID = "answerdotai/ModernBERT-base"

MODEL_ID = "hugoomezz/modernbert-ragtruth-token-level-binary"

SUPPORTED, HALLUCINATED = 0, 1



tokenizer = AutoTokenizer.from_pretrained(TOKENIZER_ID)

model = AutoModelForTokenClassification.from_pretrained(MODEL_ID, attn_implementation="sdpa").eval()



context = "The Eiffel Tower was completed in 1889 for the World's Fair in Paris."

response = "The Eiffel Tower was completed in 1889 and is located in Berlin, Germany."



encoding = tokenizer(

    context, response, max_length=4096, truncation="only_first",

    return_offsets_mapping=True, return_token_type_ids=False, return_tensors="pt",

)

sequence_ids = encoding.sequence_ids(0)

encoding.pop("offset_mapping")  # only needed for span reconstruction, not the score



with torch.no_grad():

    logits = model(**encoding).logits[0]

probs_hallucinated = torch.softmax(logits, dim=-1)[:, HALLUCINATED].tolist()



# Response-level score: max P(hallucinated) over response tokens (sequence_id == 1).

response_probs = [p for p, sid in zip(probs_hallucinated, sequence_ids) if sid == 1]

score = max(response_probs) if response_probs else 0.0

print(f"hallucination score: {score:.4f}  ({'FLAGGED' if score >= 0.5 else 'clean'})")

```

## Citation

```bibtex

@inproceedings{niu2024ragtruth,

  title     = {RAGTruth: A Hallucination Corpus for Developing and Evaluating RAG Systems},

  author    = {Niu, Cheng and Wu, Yuanhao and Zhu, Juno and Xu, Siliang and Shum, Kashun and Zhong, Randy and Song, Juntong and Zhang, Tong},

  booktitle = {Proceedings of ACL 2024},

  year      = {2024},

  eprint    = {2401.00396}

}

@article{kovacs2025lettucedetect,

  title   = {LettuceDetect: A Hallucination Detection Framework for RAG Applications},

  author  = {Kov{\'a}cs, {\'A}d{\'a}m and Bakos, Zsolt},

  journal = {arXiv preprint arXiv:2502.17125},

  year    = {2025}

}

```