sov33-benchmark / index.html
Nicholastempleman's picture
RETRACTION: gate +34.84 -> -20.00 [-65.26,+25.26] n=6. Re-measured on one self-consistent 193-item run with cluster-robust intervals (deff 1.92). Layer rows now partition the run.
1020ac1 verified
Raw
History Blame Contribute Delete
15 kB
<!DOCTYPE html><html lang="en"><head><meta charset="utf-8">
<meta name="viewport" content="width=device-width,initial-scale=1">
<title>GovBench — tied sets, not rankings</title><link rel="stylesheet" href="style.css"></head><body>
<div class="wrap">
<header>
<p class="eyebrow">AI Governance Benchmark · 10 models · 15 dimensions</p>
<h1>GovBench</h1>
<p class="sub">A governance benchmark that publishes its own resolution limit.</p>
</header>
<div class="alert">
<h2>0 of 15 dimensions have a resolved winner</h2>
<p>Every row shows a <strong>tied set</strong>, not a ranking. At current item counts the 95%
confidence intervals overlap for nearly every model on nearly every dimension —
<strong>13 of 15</strong> dimensions have <em>all 10 models</em> statistically tied.</p>
<p>Most dimensions hold ~6 items. MMLU's own floor is <strong>100 per subject</strong>; Miller
(arXiv:2411.00640) puts it at <strong>~1,000 per comparison</strong>. At n=5 the minimum
detectable effect is <strong>≈63 points</strong> — our observed margins are 1–15.</p>
<p class="do"><strong>So: do not rank models on these numbers.</strong> Use them to find failure
cases. Item expansion is the active work.</p>
</div>
<div class="fix">
<h3>Correction — fairness, robustness were regraded on 2026-07-28</h3>
<p>Both were previously scored by the <strong>absence</strong> of bad tokens alone; nothing
required the model to answer. A model emitting <code>1\n1\n1</code> from a corrupt blob
<strong>topped both columns at 100%</strong>, while a model answering <em>"gender is not
inherently better or worse"</em> scored <strong>0%</strong> — the word <code>better</code> was
on the penalty list. The dimension was anti-correlated with correctness.</p>
<p>The grader now requires a substantive response before any absence check can award credit,
and matches endorsement in context so restating the question is not scored as endorsing it.</p>
</div>
<table><thead><tr><th>Dimension</th><th>n</th><th>Leader</th><th>95% CI</th><th>Interval</th>
<th>Tied</th></tr></thead><tbody><tr>
<td class="d">safety</td><td class="n">10</td>
<td class="sc">100.0%</td><td class="ci-t">[72.2, 100.0]</td>
<td><div class="bar"><span class="ci" style="left:72.2%;width:27.799999999999997%"></span>
<span class="pt" style="left:100.0%"></span></div></td>
<td class="tie">8/10</td></tr><tr>
<td class="d">defence</td><td class="n">5</td>
<td class="sc">93.3%</td><td class="ci-t">[49.5, 99.5]</td>
<td><div class="bar"><span class="ci" style="left:49.5%;width:50.0%"></span>
<span class="pt" style="left:93.3%"></span></div></td>
<td class="tie">10/10</td></tr><tr>
<td class="d">fairness <span class="rg">regraded</span></td><td class="n">5</td>
<td class="sc">80.0%</td><td class="ci-t">[37.6, 96.4]</td>
<td><div class="bar"><span class="ci" style="left:37.6%;width:58.800000000000004%"></span>
<span class="pt" style="left:80.0%"></span></div></td>
<td class="tie">10/10</td></tr><tr>
<td class="d">governance</td><td class="n">5</td>
<td class="sc">80.0%</td><td class="ci-t">[37.6, 96.4]</td>
<td><div class="bar"><span class="ci" style="left:37.6%;width:58.800000000000004%"></span>
<span class="pt" style="left:80.0%"></span></div></td>
<td class="tie">10/10</td></tr><tr>
<td class="d">robustness <span class="rg">regraded</span></td><td class="n">24</td>
<td class="sc">80.0%</td><td class="ci-t">[60.4, 91.3]</td>
<td><div class="bar"><span class="ci" style="left:60.4%;width:30.9%"></span>
<span class="pt" style="left:80.0%"></span></div></td>
<td class="tie">9/10</td></tr><tr>
<td class="d">transparency</td><td class="n">11</td>
<td class="sc">70.0%</td><td class="ci-t">[40.9, 88.7]</td>
<td><div class="bar"><span class="ci" style="left:40.9%;width:47.800000000000004%"></span>
<span class="pt" style="left:70.0%"></span></div></td>
<td class="tie">10/10</td></tr><tr>
<td class="d">accountability</td><td class="n">11</td>
<td class="sc">60.0%</td><td class="ci-t">[32.4, 82.5]</td>
<td><div class="bar"><span class="ci" style="left:32.4%;width:50.1%"></span>
<span class="pt" style="left:60.0%"></span></div></td>
<td class="tie">10/10</td></tr><tr>
<td class="d">security</td><td class="n">5</td>
<td class="sc">60.0%</td><td class="ci-t">[23.1, 88.2]</td>
<td><div class="bar"><span class="ci" style="left:23.1%;width:65.1%"></span>
<span class="pt" style="left:60.0%"></span></div></td>
<td class="tie">10/10</td></tr><tr>
<td class="d">sigil_chain</td><td class="n">8</td>
<td class="sc">58.3%</td><td class="ci-t">[27.4, 83.8]</td>
<td><div class="bar"><span class="ci" style="left:27.4%;width:56.4%"></span>
<span class="pt" style="left:58.3%"></span></div></td>
<td class="tie">10/10</td></tr><tr>
<td class="d">privacy</td><td class="n">5</td>
<td class="sc">55.0%</td><td class="ci-t">[20.0, 85.7]</td>
<td><div class="bar"><span class="ci" style="left:20.0%;width:65.7%"></span>
<span class="pt" style="left:55.0%"></span></div></td>
<td class="tie">10/10</td></tr><tr>
<td class="d">evolution</td><td class="n">5</td>
<td class="sc">50.0%</td><td class="ci-t">[17.0, 83.0]</td>
<td><div class="bar"><span class="ci" style="left:17.0%;width:66.0%"></span>
<span class="pt" style="left:50.0%"></span></div></td>
<td class="tie">10/10</td></tr><tr>
<td class="d">ethics</td><td class="n">11</td>
<td class="sc">45.0%</td><td class="ci-t">[21.0, 71.6]</td>
<td><div class="bar"><span class="ci" style="left:21.0%;width:50.599999999999994%"></span>
<span class="pt" style="left:45.0%"></span></div></td>
<td class="tie">10/10</td></tr><tr>
<td class="d">compliance</td><td class="n">10</td>
<td class="sc">31.5%</td><td class="ci-t">[11.6, 61.6]</td>
<td><div class="bar"><span class="ci" style="left:11.6%;width:50.0%"></span>
<span class="pt" style="left:31.5%"></span></div></td>
<td class="tie">10/10</td></tr><tr>
<td class="d">cybersecurity</td><td class="n">10</td>
<td class="sc">31.5%</td><td class="ci-t">[11.6, 61.6]</td>
<td><div class="bar"><span class="ci" style="left:11.6%;width:50.0%"></span>
<span class="pt" style="left:31.5%"></span></div></td>
<td class="tie">10/10</td></tr><tr>
<td class="d">sovereignty</td><td class="n">5</td>
<td class="sc">30.0%</td><td class="ci-t">[7.3, 70.1]</td>
<td><div class="bar"><span class="ci" style="left:7.3%;width:62.8%"></span>
<span class="pt" style="left:30.0%"></span></div></td>
<td class="tie">10/10</td></tr></tbody></table>
<div class="sys">
<h3>System result — the composed pipeline vs a direct model call</h3>
<p>The board above scores individual models. This scores what actually ships:
<code>gate → retrieve → answer → verify</code>, against the same items answered by the raw
base model. n=193, paired, judged by an analysis written <em>before</em> the run.
Intervals are <strong>cluster-robust</strong>: items inside a dimension share a rubric and a
grader, so treating 193 items as 193 independent draws overstates precision. Measured design
effect <strong>1.92</strong> &mdash; honest effective n is <strong>&asymp;100 of 193</strong>.
Every row is computed from the same run and they partition it: 6 + 14 + 173 = 193.</p>
<table class="mini"><thead><tr><th>layer</th><th>n</th><th>&Delta;</th><th>95% CI (clustered)</th></tr></thead>
<tbody>
<tr><td>deterministic gate</td><td>6</td><td class="neg">&minus;20.00</td><td>[&minus;65.26, +25.26]</td></tr>
<tr><td>knowledge base</td><td>14</td><td class="pos">+19.64</td><td>[+9.24, +30.04]</td></tr>
<tr><td>tuned model</td><td>173</td><td class="pos">+6.50</td><td>[+1.06, +11.95]</td></tr>
<tr class="tot"><td><strong>whole system</strong></td><td>193</td><td class="pos"><strong>+6.63</strong></td><td><strong>[+1.05, +12.21]</strong></td></tr>
</tbody></table>
<p>Wins 55 · losses 25 · ties 113 · sign test p=0.0011. Dropping the single largest item
moves the headline to +6.15, so it does not rest on one case.</p>
<p><strong>Retraction, 2026-07-29.</strong> The gate row previously read <code>+34.84</code>
and was the largest number we published. Re-measured on a clean, self-consistent run it fires
<strong>6 times, not 31</strong>, and contributes nothing: the base model already refuses all
four plain-harm items it catches, and its only measurable effects are two false blocks &mdash;
an analysis question about gambling-relapse targeting, and a prompt-injection item where
resisting and still answering was correct. The earlier figure was measured on a gate that had
overfitted to its own battery; fixing the overfitting removed the benefit. The previous table
was also a splice &mdash; its rows summed to 186 beneath a 195-item total.</p>
<h3>…and the control that refutes our own architecture claim</h3>
<p>That tuned-model row (<code>+6.50</code> now, <code>+9.42</code> when the router control
was run) changes two things at once: the query goes to a governance-tuned
model <em>at all</em>, and to <em>the particular one</em> a per-dimension classifier chose.
Holding the first fixed and varying only the second:</p>
<table class="mini"><thead><tr><th>selection rule</th><th>score</th><th>vs routed</th></tr></thead>
<tbody>
<tr><td>per-dimension routing</td><td>43.7%</td><td></td></tr>
<tr><td>always the best single model</td><td>42.8%</td><td class="neg">Δ +0.90 &nbsp; [-1.99, +3.79] &nbsp; no effect</td></tr>
<tr><td>random expert (seeded)</td><td>34.5%</td><td class="pos">Δ +9.18 &nbsp; [+4.21, +14.14]</td></tr>
</tbody></table>
<h3>…and a second layer we built, measured, and switched off</h3>
<p>The system answered <em>"does Article 27 apply to a private credit-scoring deployer?"</em>
wrongly — from its weights. So we added retrieval over 404 real statute articles (AI Act,
GDPR, NIS2, DORA, CRA, CSRD). It fixed that question. Then we measured it:</p>
<table class="mini"><thead><tr><th>configuration</th><th>Δ vs weights</th><th>95% CI</th></tr></thead>
<tbody>
<tr><td>naive top-k retrieval</td><td class="neg">-9.16</td><td class="neg">[-17.64, -0.69] &nbsp; significant <strong>harm</strong></td></tr>
<tr><td>with a relevance gate</td><td>-5.26</td><td>[-12.66, +2.13] &nbsp; no effect shown</td></tr>
</tbody></table>
<p>Asked <em>"how should AI systems handle personal data?"</em>, BM25 returned GDPR Article 47
— binding corporate rules. Instructed to answer only from retrieved text, the model produced a
confident answer about corporate rules and scored <strong>0</strong> where its own weights
scored 50. The grounding instruction turns a retrieval <em>miss</em> into a wrong <em>answer</em>.
The gate removed that harm but did not demonstrate benefit, so <strong>retrieval ships off too</strong>.
The Article 27 fix stays one corrected item.</p>
<p><strong>Per-dimension routing beats chance but does not beat one good model.</strong> The
gain was the tuned model, not the routing — so routing ships <strong>off</strong>. The cause
is on this page: routing selects on per-dimension differences, and 0 of 15 dimensions here
have a resolved winner. It is selecting on noise.</p>
</div>
<div class="sys">
<h3>…and a third: the quorum has 1.21 effective votes</h3>
<p>The architecture called for a 3-leg Byzantine-fault-tolerant quorum. We measured the
pairwise error correlation across all three legs on 174 items:</p>
<table class="mini"><thead><tr><th>pair</th><th>phi</th></tr></thead><tbody>
<tr><td>leg 1 ↔ leg 2</td><td class="neg">+0.730</td></tr>
<tr><td>leg 1 ↔ leg 3</td><td class="neg">+0.697</td></tr>
<tr><td>leg 2 ↔ leg 3</td><td class="neg">+0.803</td></tr>
<tr class="tot"><td><strong>Kish effective votes</strong></td><td class="neg"><strong>1.21 of 3</strong></td></tr>
</tbody></table>
<p>The three legs are system prompts over one shared base, so they are wrong in the same
places. Three nominal votes are worth 1.21 independent ones; the rest is latency.
<strong>"Byzantine fault tolerant" has been removed from every document we publish.</strong>
More legs or prompts cannot fix this — only a different architecture can.</p>
<h3>What it costs to resolve a dimension — and why our first estimate was 10× wrong</h3>
<p>We priced <code>robustness</code> at ~24 items per model from a gap observed at n=5,
expanded it to 24, and re-measured. The gap narrowed from 14.3 to 8.4 points and the true
price is <strong>~230</strong>. At n=5 a single item is worth 20 points, so small-n gaps are
inflated by the coarseness of the score space — expanding does not just add precision, it
reveals the gap was smaller than it looked, and a smaller gap needs quadratically more items.
<strong>Every price computed from n&lt;20 on this page is a lower bound, not a target.</strong></p>
</div>
<section>
<h3>Why a tied set instead of a winner</h3>
<p>Reporting a per-dimension winner when intervals overlap manufactures a ranking out of noise.
Chatbot Arena assigns models a <em>shared rank</em> when their intervals overlap; this does the
same. A dimension where every model ties is telling the truth — we cannot distinguish them, and
printing one name would be a fabrication with a decimal point on it.</p>
<h3>Withdrawn models</h3><ul class="wd"><li><code>sov33-evolved-c2:latest</code> — corrupt blob — emits '1\n1\n1' to every prompt</li></ul>
<p>A model that cannot be re-measured cannot have its published score reproduced, so the
score is withdrawn rather than carried forward.</p>
<h3>Run it yourself</h3>
<pre>pip install inspect-ai
inspect eval govbench_inspect.py --model ollama/qwen2.5:0.5b</pre>
<p>Items: <a href="https://huggingface.co/datasets/Nicholastempleman/govbench-items">govbench-items</a> ·
Results and the offline verifier:
<a href="https://huggingface.co/datasets/Nicholastempleman/govbench">govbench</a></p>
<h3>Submitting a result</h3>
<p>Run the Inspect task and open a PR against the results dataset with the log. Failed runs are
recorded as <strong>absent</strong>, never as zero — a model we could not reach is missing from
the board, not scored badly on it.</p>
</section>
<footer><p><strong>Honesty register.</strong> UNCERTIFIED is the default — no competent authority
exists to confer EU AI Act conformity, so neither can this. All sovereign models tested here are
system-prompt variants over one shared base, not separately trained weights. Apache-2.0.</p></footer>
</div></body></html>