JEV Ecosystems β every answer-verification vendor publishes a benchmark, and every one of them wins it. So we ran 13 of them on one test set: 2,018 items, identical labels, same grading code.
1οΈβ£ Only three systems clear 0.70 β ZTC (397B) 0.7364 Β· JEV 0.7350 Β· ZTC (27B) 0.7282. First and second differ by 0.0014, so no rank is assigned.
2οΈβ£ A baseline that reads nothing but answer length and formatting scores 0.7036. Eight of the thirteen fall below it. A leaderboard without that line is flattering its entrants.
3οΈβ£ Bigger does not win. On scientific reasoning, 27B 0.7410 beats 397B 0.6287 β a model fourteen times larger scoring 0.11 lower.
And AUC is not the number you deploy on.
Same 20% retry budget, wired into an agent loop, against a 74.83% no-gate baseline: ZTC +1.34 pp Β· JEV β0.07 pp Β· random β0.25 pp.
The mechanism is the interesting part. Re-answering is double-edged: 38% of wrong answers get fixed, and 30% of right answers get broken. So a gate is paid for by precision, not recall. Of the 403 items JEV routed for a retry, 216 were already correct.
0.0014 AUC apart; 1.4 points of end-to-end agent accuracy apart.
Scores, labels and grading code are published in full. Four public reproductions that would not run from their released artefacts are listed too, with the failure and a link, and no score.
Don't take the table's word for it β paste your own case into the playground and watch all three answer at once. Want a system added? Open a discussion on the Space.
OpenRouter Leaderboard β every model, every provider, one comparable table. Price, precision, uptime, measured latency and language quality on the same axes.
Building it turned up three things.
We graded 330 models on Korean and two axes collapsed.
Honorifics β only 8.5% earn an A Knowledge of Korean institutions β 9.4% Every other axis sits above 31% Fluency hides it. A model can write clean, natural Korean and still attach an honorific to a coffee cup. Fluent and wrong at the same time is worse than obviously broken, because nobody catches it in review.
A 2023 model beats the 2026 flagships. gpt-3.5-turbo-16k scores a perfect 3.00. Korean cannot be inferred from release date, parameter count or English benchmarks β it has to be measured, per model.
Quality, value and speed are three different models. Across five axes, the same model almost never takes two columns.
425 models, latency measured on 329 on a paid API, Korean graded on 330. Three languages, three currencies, daily refresh, open API, no key.
Instead of making the fly brain play games, we measured what it is for
Since the Drosophila connectome was released, people have had the fly brain doomscroll a feed, play Beat Saber, drive in GTA. Those demos show that the brain runs. We wanted to show what it is for.
So we gave it a looming object β one of the few things a fly brain is unambiguously built to detect β then deleted a single cell type and repeated the identical stimulus. Remove LC4, 126 cells out of 173,023, and the escape signal falls from 0.840 to 0.091. Eighty-nine percent of the danger signal is gone while the other 172,897 neurons run exactly as before.
Deleting neurons does not do this on its own, which is the whole point of the controls. LC11 is the same class and larger than LC4 β 143 cells and 9,940 outgoing connections against 126 and 7,846 β and removing every one of them changes the signal by 0.000000, to six decimal places. It has to be those 126.
No server and no GPU: a looming stimulus drives fewer than one percent of neurons above threshold, so the whole thing is 40 KB gzipped and runs in your browser.
The wiring is the measured connectome, but synaptic strength is a uniform count-based value and the dynamics are a firing-rate model of our choosing β a total-effect measurement of a model, not a recording from a fly. Male CNS connectome, FlyEM / HHMI Janelia with Google Research, Columbia and Harvard (2026), CC BY.