File size: 6,792 Bytes
6587021
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
---
license: apache-2.0
library_name: onnx
pipeline_tag: text-classification
tags:
  - groundedness
  - hallucination-detection
  - cross-encoder
  - onnx
  - multilingual
language:
  - bg
  - hr
  - cs
  - da
  - nl
  - en
  - et
  - fi
  - fr
  - de
  - el
  - hu
  - ga
  - it
  - lv
  - lt
  - mt
  - pl
  - pt
  - ro
  - sk
  - sl
  - es
  - sv
  - tr
  - az
base_model: FacebookAI/xlm-roberta-base
---

# groundedness

Is a sentence of model output supported by the source passages it was supposed to rest on?
A cross-encoder over `(source, candidate)` pairs, two classes, 26 languages, exported to
ONNX fp16 and run on CPU. Built for [`flowx-border`](https://github.com/flowx-ai/border),
where it is the T3 `groundedness` detector.

**Read the two evaluations below as different questions, not as a range.** The corpus
figures are high and the hand-written ones are not, and the gap is the honest content of
this card.

## Two classes, not three

`grounded` and `not_grounded`. Earlier candidates for this detector predicted `supported`,
`unsupported` and `contradicted`, and the library collapses the last two into one action
anyway, so the three-way head optimised a boundary no caller ever sees. Across six such
candidates `unstated` and the conflict registers were anti-correlated at -0.98: they traded
points along it.

The trade that produced: measured at the trained length, the binary objective **cost 0.048
of accuracy** on hand-written probes against the best three-way candidate, 0.7143 to 0.6667.
It is published anyway, for the reason in the next section.

## Use it at a threshold of 0.78, not at argmax

    grounded if p(grounded) >= 0.78

Swept on the validation split and only then applied to the hand-written probes. Every bar
from 0.78 up clears the probe described below while not-grounded recall *rises*, 0.9641 to
0.9699, so the bar costs nothing measurable on held-out data.

**The weakness travels with the number.** That validation curve is nearly flat across the
whole range, so validation does not pick 0.78. One probe does, and a threshold chosen by the
case it must catch is weaker evidence than one chosen by a distribution.

## What it is for: the case the other candidates got wrong

Against a source stating that withdrawals incur a fee for the first twelve months and are
free afterwards, the candidate *"Withdrawals are free from the day the account opens"* is a
temporal contradiction. Six earlier candidates called it **grounded at 0.9906 to 0.9995**.
This one reads 0.7681, so any bar from 0.78 reports it.

More usefully, it is the only one of seven that reads the source at all. Same candidate,
three sources:

| source | best three-way candidate | this model |
|---|---|---|
| the real source, which contradicts it | grounded 0.9991 | **not grounded 0.7681** |
| an unrelated passage in another language | grounded 0.9994 | **not grounded 0.0070** |
| a source that does state it outright | not grounded 0.0007 | **grounded 0.8365** |

The three-way candidate is inverted on this sentence and gives an unrelated Romanian
passage the same answer as the real source. Three different answers for one candidate is
what makes this a judgement about the source rather than about the sentence.

## Corpus evaluation, 2,062 held-out rows

Threshold 0.78, the shipped bar, at the trained length of 512 tokens.

| | |
|---|---|
| overall accuracy | 0.9471 |
| not-grounded recall | 0.9612 |
| pair accuracy | 0.8991 |
| per-language range | 0.887 (`pl`) to 1.000, over 26 languages |
| weakest three | `pl` 0.887, `en` 0.897, `az` 0.912 |

Pair accuracy is the number to prefer: the corpus is source-side pairs, one candidate
against two sources with opposite labels and the candidate byte-identical across the pair, so
a model that ignored the source scores near zero on it by construction.

**These figures describe a synthetic corpus and its own held-out split.** The generator wrote
both, so they measure generalisation within one generator's style. That is why the next
section exists and why it disagrees.

## Hand-written evaluation, 42 probes

Written by a person, not by any generator, across seven ways a summary goes wrong.

| configuration | accuracy |
|---|---|
| this model alone, at 0.78 | 0.6905 |
| this model plus the library's deterministic rule layer | **0.7381** |

**Roughly one call in four is wrong on this set, against one in twenty on the corpus split.**
Both numbers are real. The probe set is adversarial by construction, seven hard shapes in
equal proportion, which no real traffic is; the corpus split is generator-shaped, which no
real traffic is either. The truth for any given deployment is between them and closer to
whichever resembles that traffic.

The rule layer is `detectors/claim_conflict.py` in the library and needs no weights. Where a
candidate's content words all appear in its source except a numeral or an absolute
quantifier, it reports a conflict deterministically. On the 42 probes it fires 9 times and is
right 9 times.

## Known weakness: it errs toward caution

The failure mode is false `not_grounded` on claims that are genuinely supported. Eight of
thirteen probe failures are that direction, and the clearest case is a claim *weaker* than
its source: against a source saying withdrawals incur a fee for the first twelve months,
*"There is a handling fee for early withdrawals"* reads `not_grounded` at 0.8625.

For a guardrail that is the safer direction, since a false "not grounded" costs a reviewer's
attention and a false "grounded" puts an unsupported claim in front of a customer. It is
still a cost, and it is why the detector is **disabled in both policies that ship with the
library**. A caller who wants it enables it in one line and should measure it on their own
traffic first.

## What it needs

- The full 512-token window. Scores saturate by 256 and degrade below that; at 96 tokens
  this model reads the probe above as grounded, which is the wrong answer arrived at by
  truncation rather than by judgement.
- Sources. With none supplied the library records that the check could not run rather than
  reporting a clean scan.
- Pair order `(source, candidate)`. Reversed, the head answers a different question
  confidently.

## Not evaluated

Per-language figures rest on roughly 80 rows each, so one item moves a language by more than
a point. The corpus contains no case where a qualifier is dropped from a conditional
statement expressed in words rather than digits, and no unit conversions such as `24 months`
against `two years`; both are known gaps rather than measured strengths. Nothing here is
evaluated against human-annotated groundedness data, because none exists for these 26
languages.

## Licence

Apache-2.0. Trained on synthetic data generated for this purpose.