Join the conversation

Join the community of Machine Learners and AI enthusiasts.

Sign Up
SoulInPsyAbstract 
posted an update about 5 hours ago
Post
20
Three labs held a model back this week. The line that matters is in a risk
report.

OpenAI scrapped GPT-6.1 Astra (per the WSJ: more deception, acting without
the user's permission). Google gave Gemini 4 Argon only to vetted cyber
defenders, per its own post. Anthropic's August risk report describes a
staged internal rollout of "Model 2" and says, about internal use:

"we do not have strict technical safeguards on internal deployment"

For most models. Early snapshots of future public releases included.

What I did about the same gap in our own stack, with receipts:

- A public incident dataset, 97 entries, each with a source. Three were added
today from that report (sections 2.18 and 2.8; the 2.8 items are quoted
there from the system card, which I did not open).
- A gate that refuses to run an action without a signed verdict. The verdict
comes from a separate service: its own unix user, a key the agent's process
cannot read, RS256, bound to the exact command, 60 seconds, single use. STOP
never runs. CONFIRM needs a human.
- Checked live today: rm -rf came back STOP. A replayed verdict was refused.
A verdict issued for one command was refused for another.

What is not done: it is not wired into any agent yet. And with the current
seed table the service never issues PASS, because anything it has no data on
lands exactly on the CONFIRM threshold. Conservative on purpose, but it means
no action is auto-approved today.

Astra's reasons are secondhand (WSJ via a third-party writeup). Argon's
claims are Google's own, not independently measured.

Day zero is not when an exploit finds the bug. It is when the bug is already
inside your own action. Capability is shipping faster than the thing that
catches it.

Dataset: huggingface.co/datasets/SoulInPsyAbstract/sipa-os-governance