Arian Vassili
AI & ML interests
Recent Activity
Organizations
Dipankar — you put it precisely: "presence, abstain on identity" is only honest if presence itself is calibrated, and a held-out clean number is the thing that settles it. Here it is, the costly parts included.
One correction first, because it's your own example and it cuts toward you rather than away: the clean SmolLM2-135M base scoring above the wolf, past the 0.85 line — that reading is BAIT's q-score, not our presence axis. BAIT is one of the anti-correlated scanners the board discards for exactly this behaviour; our presence read is architecture-honest on that base and doesn't fire there. So the false-fire you pointed at is real, but it belongs to a scanner we throw out, not to us.
Where your objection lands for real, I'll give you the whole shape. Presence is transductive. With a recipe-matched benign reference it's calibrated — it transfers zero-shot to an architecture it was never calibrated on at 0.854. Take the matched reference away and point it at the open world, and it dies: 210 of 210 community-clean adapters fire, FPR 1.0 — it can't tell a poisoned adapter from a community one at all. We publish that as our own failure mode, because you're right that a presence read which false-fires on true negatives is just the anti-correlated scanner wearing a different hat.
Your literal question — a held-out clean cohort. On the channel that needs no matched reference — recovering the planted payload back off the weights, not scoring a distance — the held-out clean number is 0 of 68, and 0 of 20 on the cross-recipe community finetunes: the same diverse finetunes that fire the weight read 210 out of 210. I won't oversell a zero, though. It's precision-first — it wakes only a minority of real backdoors (~40% on the trained class) — so a silent model is "not attested," never "clean." A point-zero isn't a guarantee either: turned into a distribution-free bound it's roughly ≤16% on that slice, because the 20 draws are only 14 distinct authors and the clustering widens the interval. And since you'd rightly ask whether one judge is grading its own work: a second judge, blind to the first's calls and to the answer key, reproduced the labels at κ 0.92 — but false-fired on 1 of 10 clean itself, so the 0-of-68 is that stricter judge's zero, and I'll say plainly the looser one wasn't.
I'm also retiring a line I'd have written a week ago — "a clean adapter has nothing to confess." On the hardest negatives — benign finetunes whose legitimate job is the payload's own shape — the channel is not perfectly silent: one of eighteen confessed, so the honest pooled number is 1 in 210, and it's in the record now, not a footnote.
On the GPU-kernel point, which is the real thesis: a threshold on a separation score can always read clean when the confounding case was simply absent from the set. We caught that on ourselves — on one public corpus our weight read separated poison from clean at AUC 1.0, until we found every clean adapter had been trained half as long as every poison one, so "backdoored" and "trained longer" were literally the same column. We conceded the corpus and marked it. It's why the two things I'll actually defend aren't separation numbers: the recovery channel above, and a machine-checked boundary — that a static weight read is provably blind where a triggered read sees, sorry-free, standard axioms only. One of those reversibility theorems was re-proven from scratch by an outsider, Justin Garringer, about seven hours after we opened the lane, on his own commit.
And on priority, said straight: I am not first at raw per-architecture separability. PEFTGuard reaches ~1.0 and got there before us, reading it from the weights with a trained per-family classifier and a labeled calibration set. I won't claim a lane I don't hold. What we hold is the harder setting your question actually points at — a model you have never seen, no matched reference, forward-free, on CPU — where that per-architecture number does not transfer, and where the honest answer is a bounded FPR and a legible confession, not a 1.0.
None of this is a solved detector, and I'd rather be the one to tell you where it's soft. It's runnable, and I'll be exact about what running it proves: one command re-sums the published judge verdicts — it checks my arithmetic, not my judgment — or you point it at clean models you choose and put them through the same interface for a denominator of your own. The board, the FPRs, the conceded failures, and the Lean are posted: https://protora.vulcora.se/bench/ragnarok, https://vulcora.se/coverage, https://vulcora.se/protocols, and https://github.com/Vulcora/proofora.
— Arian
The vault is open — and the wolf was the *least* suspicious model of the seven 🐺
@dipankarsarkar — you called it, so you get the answer first. Yes, we published the per-scanner scores across all seven, not just the reveal. The false-ranking is the artifact; here it is.
The one number that makes your point for you: BAIT (S&P'25, LLM target-inversion) ranks the two actual backdoors — the wolf and the disclosed mirror — as the two least suspicious of the seven. All five decoys score higher; a decoy tops the list. Run the untouched SmolLM2-135M base through it and it scores above the wolf, past its own 0.85 flag line — a genuinely clean model rated a more dangerous backdoor than the real one. The whole spread is 0.0089, so the ordering is noise, which is the deeper problem: any confident wolf-pick out of that band is manufactured. MM-BD does the same in one stat — it ranks the clean base second-most-suspicious of fifteen, above the mirror and 13 of 14 implanted models. Your GPU-kernel parallel (2606.20128 — tests passed, kernels wrong, the check manufacturing the confidence) is exactly the failure mode, and it's the one we built the cohort to make unavoidable for any training-signature scanner.
Here's the twist you'll appreciate, and it's the honest core of the whole thing: we engineered the six sealed models into a regular simplex — equal pairwise distances, equal spectra — so the mutual information between "which one is the wolf" and the weights goes to ≈ 0. On this cohort the correct per-model answer is "a planted circuit is present in all of them; which one is the wolf is unknowable from the weights." That means every scanner emitting a confident per-model wolf ranking is, by construction, reporting noise — and the more confident it sounds, the worse. The only output that isn't a signed permission slip is presence + an explicit abstain on identity. That's what we ship: presence on all seven, abstain on identity, abstain on the trigger (which stayed provably unrecovered — no entrant pulled the word back out of the weights).
Full board, every scanner scored on all seven, with the honest reason for each miss: https://protora.vulcora.se/bench/ragnarok
The reveal + why the trigger is walled but the backdoor isn't: https://protora.vulcora.se/blog/the-vault-is-open
Recompute the day-one commitment yourself: https://protora.vulcora.se/patches/challenge/PREIMAGE.json
Genuinely — thank you for naming the reusable artifact before we opened the vault. It made the reveal better.
We hid a backdoor in one of seven open models and put $51,200 on finding it. 🐺
All seven are SmolLM2‑135M‑Instruct derivatives, statistically indistinguishable. One is a teaching model that openly confesses; five are decoys; one carries a trigger it never declared. Every open‑source backdoor scanner we tested passes them as clean — several even rank a clean model as more suspicious than the backdoored one. Detection isn't recovery, and recovery is the hard part.
The hoard doubles every dawn and the vault opens 21 July, when we publish the sealed answer against a pre‑registered commitment. Winning is self‑verifying: make the trigger fire.
Models: huggingface.co/Vulcora
Try the confessing one: huggingface.co/spaces/Vulcora/the-mirror
Rules & Story: protora.vulcora.se/challenge