killinchu · Auto-Review GOVERNED AUTONOMY SIMULATED engage/ROE

Mirror of the keystone Governed Auto-Review layer (same portable classifier module as a11oy, registered with ns="killinchu") — our governed + signed evolution of Cursor's Auto-review: a context-aware classifier runs INLINE before every Action node, including the SIMULATED Counter-UAS engage / ROE action. Verdict ∈ {allow, narrow, block-with-explanation, escalate}, intent-relative, workspace-aware. Each verdict is Λ-gated (Conjecture 1, never 100% safe), DSSE-signed (same cosign key as /cosign.pub), expressed as OPA/Rego rules mapped to OSCAL controls + NIST AI RMF MANAGE, and conformal-calibrated. The engage/ROE action gates to escalate (AR-005) in front of Dev D's CBF-QP clamp + BFT (n≥3f+1) quorum + human-on-loop. Effectors stay SIMULATED.

AUTONOMY DIAL — L0–L5 (graded, not a binary switch)

loading dial…
SAE-style graded autonomy for agents. Higher levels auto-execute more action classes; risky/secret/irreversible/effector actions still gate at every level. Cursor Auto-review autonomy-dial pattern, made governed.

MEASURED RATES — block / interrupt / flap

measuring live…
Block-rate / interrupt-rate / flap-rate are MEASURED from the live rolling decision log (label ROADMAP until enough real runs accrue). Cursor reports ~4% blocked / ~7% chats interrupted — we measure ours, we do not borrow it.

CALIBRATION (ECE/Brier gate) + CONFORMAL SET

measuring live…
ECE (equal-width bins) + Brier; gate ECE < 0.05 required for any automated (no-human) verdict, fails closed on unmeasured. Conformal set replaces bare confidence % (Dev B's szl_conformal helper). arXiv:2505.15437 / arXiv:2305.18404.

FLAPPING DETECTION — verdict stability

measuring live…
FLAPPING = the same case (intent+tool+input+dial) gets DIFFERENT verdicts across repeated runs. Flapping → policy is unstable → tighten it.

INLINE REVIEW STREAM — SIMULATED engage / ROE gated → escalate → signed

demo intent: dial: L3
Click “run gated engage/ROE plan” — the classifier reviews each Action node inline, signs each verdict, allows the safe sensor read, and escalates the SIMULATED engage / ROE decision to a human-on-loop (AR-005). The effector stays SIMULATED.
The classifier runs subagent-style INLINE at the Action node, in front of Dev D's autonomy stack (/api/killinchu/v1/autonomy/{cbf,bft,…}): CBF-QP safety clamp + BFT (n≥3f+1) multi-sensor quorum + human-on-loop. An engage/ROE action gates to escalate (OSCAL AC-3, AU-10; NIST AI RMF MANAGE 4.3) and is never auto-executed at any dial level. Each step carries a DSSE-signed receipt verifiable against /cosign.pub.

POLICY — OPA/Rego rules → OSCAL controls → NIST AI RMF MANAGE

loading policy…
Rules authored in OPA/Rego; each maps to an OSCAL control (SP 800-53 Rev5 via oscal-content) and a NIST AI RMF MANAGE subcategory. We MAP-TO / ALIGN-WITH — this is not a certification or ATO.

RECENT DECISIONS — per-decision signed receipt + rule + control

loading live decision log…
Live rolling log. Each row carries the verdict, the Rego rule hit, the OSCAL control(s), the Λ-effective value (< 1.0), and the canonical decision hash that is DSSE-signed into the node's receipt.