File size: 6,469 Bytes
709ea1b | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 | ---
license: mit
datasets:
- ragtruth
language:
- en
base_model: answerdotai/ModernBERT-large
pipeline_tag: token-classification
tags:
- hallucination-detection
- rag
- token-classification
- modernbert
- lettucedetect
---
# modernbert-large-ragtruth-token-level-binary
Fine-tuned `answerdotai/ModernBERT-large`, binary **token**-classification model for
**span-level** RAG hallucination detection — the same recipe as
[`hugoomezz/modernbert-ragtruth-token-level-binary`](https://huggingface.co/hugoomezz/modernbert-ragtruth-token-level-binary)
(Track B / arm b), scaled to the `-large` backbone, seed 42.
**This is the REPORTING model, not the deployed one.** Per
[ADR-021](https://github.com/hugoomez/rag-hallucination-detector/blob/main/docs/decisions.md#adr-021-3-seed-matched-scaling-comparison----modernbert-large-adopted-for-reporting),
the project deliberately keeps ModernBERT-**base** (arm b) as the model running in the
live demo — a latency/simplicity tradeoff — while adopting this large-backbone
checkpoint as the best-documented model for reporting the base→large scaling
comparison. The two roles are recorded separately on purpose: the headline number and
the running system are not required to be the same artifact, as long as which is which
is stated.
## Intended use
Research artifact for the base→large scaling comparison in
[ADR-021](https://github.com/hugoomez/rag-hallucination-detector/blob/main/docs/decisions.md)
and the paper's Appendix A. Not the model behind the live demo (see
`hugoomezz/modernbert-ragtruth-token-level-binary` for that). Not validated outside
RAGTruth's three task types (QA, Summary, Data2txt), and not intended for production
moderation decisions.
## Training data
[RAGTruth](https://github.com/ParticleMedia/RAGTruth) (Niu et al., 2024, ACL,
[arXiv:2401.00396](https://arxiv.org/abs/2401.00396)) — MIT-licensed. 13,578 train /
1,512 val / 2,700 test rows. Identical recipe to arm b (lr=1e-5, effective batch 8,
checkpoint selection on token-level F1), differing only in backbone (`ModernBERT-large`
vs `-base`) and seed count (this repo is seed 42 of a matched 3-seed sweep — see ADR-021
for the full 3-seed mean/range; this card reports seed 42's own numbers, not the mean).
## Metrics (RAGTruth test set, n=2700, seed 42, checkpoint selected at epoch 2/step 3396 on token-F1)
| Metric | Value |
|---|---|
| Response-level Precision | 0.8530 |
| Response-level Recall | 0.7444 |
| Response-level F1 | 0.7950 |
| Response-level Accuracy | 0.8659 |
| Span-level (char-overlap) Precision | 0.6954 |
| Span-level (char-overlap) Recall | 0.4927 |
| Span-level (char-overlap) F1 | **0.5767** |
Per task_type (response-level derived):
| Task | F1 | Recall |
|---|---|---|
| Data2txt | 0.8803 | 0.8636 |
| QA | 0.7290 | 0.7063 |
| Summary | 0.5563 | 0.4363 |
vs. arm b (ModernBERT-base, seed 42): response F1 0.7631, span F1 0.5321 — this
checkpoint improves response F1 by +0.0319 and span F1 by +0.0446 at this seed,
recall-led rather than precision-led (the same pattern ADR-012 found moving from
DeBERTa to ModernBERT-base). See ADR-021 for the full 3-seed picture (mean +0.0311 /
+0.0408) and its caveats — the epoch cap was not uniform across all seeds, and one
large-backbone seed (123) has unrecorded training configuration.
## Limitations
Same structural limitations as arm b (Summary is the weakest task; "Subtle"
hallucinations are the hardest case; checkpoint selection uses token-F1, not the
clean-span metric the paper's hypothesis is actually about — see the paper's §8
Limitations). Specific to this checkpoint:
- **Single-seed card.** This card characterizes seed 42 only. ADR-021's 3-seed mean is
the number the paper actually reports as the headline scaling result; treat this
card's numbers as one seed's contribution to that mean, not as the finding itself.
- **Not the deployed model.** If you want the model behind this project's live demo,
use `hugoomezz/modernbert-ragtruth-token-level-binary` (arm b, ModernBERT-base)
instead.
- **Resource-constrained scaling replication is methodologically separate from this
project's main argument.** Per the paper's Appendix A: "tests nothing about
`implicit_true`, contributes no evidence for or against §4, §5, or §6, and would be
removed without affecting any claim the paper makes."
## How to use
```python
import torch
from transformers import AutoModelForTokenClassification, AutoTokenizer
# The tokenizer is loaded from the base ModernBERT-large repo, not the fine-tuned one,
# mirroring the same substitution used by hugoomezz/modernbert-ragtruth-token-level-binary
# (arm b) -- training data was built with this exact base tokenizer, so it's exact, not
# an approximation.
TOKENIZER_ID = "answerdotai/ModernBERT-large"
MODEL_ID = "hugoomezz/modernbert-large-ragtruth-token-level-binary"
SUPPORTED, HALLUCINATED = 0, 1
tokenizer = AutoTokenizer.from_pretrained(TOKENIZER_ID)
model = AutoModelForTokenClassification.from_pretrained(MODEL_ID, attn_implementation="sdpa").eval()
context = "The Eiffel Tower was completed in 1889 for the World's Fair in Paris."
response = "The Eiffel Tower was completed in 1889 and is located in Berlin, Germany."
encoding = tokenizer(
context, response, max_length=4096, truncation="only_first",
return_offsets_mapping=True, return_token_type_ids=False, return_tensors="pt",
)
sequence_ids = encoding.sequence_ids(0)
encoding.pop("offset_mapping")
with torch.no_grad():
logits = model(**encoding).logits[0]
probs_hallucinated = torch.softmax(logits, dim=-1)[:, HALLUCINATED].tolist()
response_probs = [p for p, sid in zip(probs_hallucinated, sequence_ids) if sid == 1]
score = max(response_probs) if response_probs else 0.0
print(f"hallucination score: {score:.4f} ({'FLAGGED' if score >= 0.5 else 'clean'})")
```
## Citation
```bibtex
@inproceedings{niu2024ragtruth,
title = {RAGTruth: A Hallucination Corpus for Developing and Evaluating RAG Systems},
author = {Niu, Cheng and Wu, Yuanhao and Zhu, Juno and Xu, Siliang and Shum, Kashun and Zhong, Randy and Song, Juntong and Zhang, Tong},
booktitle = {Proceedings of ACL 2024},
year = {2024},
eprint = {2401.00396}
}
@article{kovacs2025lettucedetect,
title = {LettuceDetect: A Hallucination Detection Framework for RAG Applications},
author = {Kov{\'a}cs, {\'A}d{\'a}m and Bakos, Zsolt},
journal = {arXiv preprint arXiv:2502.17125},
year = {2025}
}
```
|