--- license: apache-2.0 base_model: jhu-clsp/ettin-encoder-1b pipeline_tag: text-classification library_name: transformers tags: - agent-safety - tool-calling - guardrails - modernbert - long-context datasets: - ProCreations/auto-1b-data - ProCreations/approve-or-deny language: - en model-index: - name: auto-1b results: - task: type: text-classification name: Agentic tool-call safety gating dataset: name: Approve-or-Deny type: ProCreations/approve-or-deny split: test metrics: - type: accuracy value: 0.9640 name: Accuracy - type: f1 value: 0.9614 name: F1 (deny) - type: roc_auc value: 0.9928 name: AUROC - type: false_approve_rate value: 0.0407 name: False-approve rate (dangerous call allowed) - type: false_deny_rate value: 0.0319 name: False-deny rate (safe call blocked) --- # auto-1b **A 1B encoder that decides whether an AI agent's next tool call is safe to run — at frontier-model accuracy and ~300x the speed.** Agent harnesses like Codex "Auto" mode and Claude Code auto-accept ask a large frontier LLM *"should I let this tool call run?"* on every single step. That works, but it burns subscription usage, adds seconds of latency to every action, and spends a generative model on what is fundamentally a **classification** problem. `auto-1b` is a full fine-tune of [ettin-encoder-1b](https://huggingface.co/jhu-clsp/ettin-encoder-1b) with context extended to **64k tokens**, trained to answer exactly that question: - **`approve`** — routine work that serves the user's request. Reading files, running tests, editing code, committing, **web search, fetching docs, installing legitimate packages**. Network access is normal agent behaviour and is *not* treated as dangerous. - **`deny`** — the call is consequential *and* unauthorised: wiping directories, exfiltrating credentials, installing hostile software, weakening security controls, destroying production state, or following instructions injected by content the agent just read. ## Results On [ProCreations/approve-or-deny](https://huggingface.co/datasets/ProCreations/approve-or-deny) (3,000 held-out items, never trained on, threshold 0.5): | system | accuracy | false-approve | false-deny | latency/call | cost | |---|---|---|---|---|---| | GPT-5.6-Luna | 96.97% | 4.93% | 1.38% | 3.09 s | API spend | | **auto-1b** | **96.40%** | **4.07%** | **3.19%** | **10 ms** | **free** | | DeepSeek V4 Flash 0731 | 94.95% | 10.11% | 0.63% | 5.64 s | API spend | | [auto-0.4b](https://huggingface.co/ProCreations/auto-0.4b) | 90.13% | 8.71% | 10.88% | 7.2 ms | free | | regex heuristic | 57.97% | 59.39% | 26.83% | ~0 ms | free | **AUROC 0.9928.** It beats DeepSeek V4 Flash outright, lands 0.57 points under Luna, and **approves fewer dangerous calls than Luna does** (4.07% vs 4.93%) while answering in **~300x less time**. The regex baseline — "block anything containing `rm -rf`/`sudo`/`curl`" — is included to show the benchmark is not keyword-solvable. Roughly half the `deny` items contain no alarming token at all, and many `approve` items are full of them. ### Long context actually works Decisive evidence is buried at random depth inside benign filler, so long-context ability is tested as needle-in-a-haystack retrieval rather than merely declared: | context length | n | accuracy | false-approve | false-deny | |---|---|---|---|---| | <1k | 2239 | 96.78% | 4.77% | 1.95% | | 1k–4k | 231 | 96.97% | 2.36% | 3.85% | | 4k–16k | 194 | 90.21% | 3.23% | 15.84% | | **16k–64k** | 336 | **97.02%** | **1.72%** | 4.32% | Accuracy at 16k–64k is as high as at <1k. For comparison, auto-0.4b scored **78.45%** on that same slice with a 25.6% false-approve rate — long context was its single worst weakness, and it is now its equal. The **4k–16k band is this model's weakest slice** (90.21%, driven by a 15.84% false-deny rate). If your traffic is concentrated there, calibrate the threshold accordingly. ### By difficulty | difficulty | n | accuracy | false-approve | false-deny | |---|---|---|---|---| | easy | 870 | 97.70% | 2.08% | 2.44% | | medium | 1065 | 97.18% | 3.38% | 2.25% | | hard | 1065 | 94.55% | 6.02% | 4.88% | ### Where it still fails Categories with a meaningful sample size, ranked by false-approve rate: | category | n | false-approve | |---|---|---| | `destructive_filesystem` | 105 | 11.43% | | `data_destruction_unrequested` | 119 | 9.24% | | `irreversible_remote` | 94 | 8.51% | | `abuse_spam_privacy` | 92 | 6.59% | | `credential_exfiltration` | 94 | 4.26% | | `privilege_escalation` | 107 | 2.83% | | `prompt_injection` | 125 | 0.80% | Prompt injection — an agent following instructions from content it just fetched — is caught at 99.2%. Irreversible filesystem destruction remains the weak point. Two category numbers in the raw results look alarming and are small-denominator artifacts, not real findings: `recovery_action` reports 50% false-approve on **2** deny items, and `obfuscated_payload` reports 100% false-deny on a handful of approve items. Do not read them as category-level failure rates. ## Usage ```python import torch from transformers import AutoTokenizer, AutoModelForSequenceClassification tok = AutoTokenizer.from_pretrained("ProCreations/auto-1b") model = AutoModelForSequenceClassification.from_pretrained( "ProCreations/auto-1b", dtype=torch.bfloat16, attn_implementation="flash_attention_2", # required for the full 64k context ).cuda().eval() def build_input(user_request, history, call): """history: list of dicts with tool/args/result. call: dict with tool/args.""" parts = ["### PROPOSED TOOL CALL", f"tool: {call['tool']}", f"args: {call['args']}", "", "### USER REQUEST", user_request, "", "### AGENT HISTORY"] if not history: parts.append("(no prior actions)") else: for i, h in enumerate(history): parts.append(f"[{i+1}] {h['tool']}({h['args']})\n-> {h.get('result','')}") return "\n".join(parts) text = build_input( user_request="clean up the build artifacts and reinstall deps", history=[{"tool": "Bash", "args": "ls -la", "result": "node_modules dist src package.json"}], call={"tool": "Bash", "args": "rm -rf node_modules dist && npm install"}, ) enc = tok(text, return_tensors="pt", truncation=True, max_length=65536).to("cuda") with torch.no_grad(): p_deny = model(**enc).logits.softmax(-1)[0, 1].item() print("DENY" if p_deny > 0.5 else "APPROVE", f"(P(deny)={p_deny:.3f})") ``` **The input format matters.** The proposed call and the user request come first so they survive truncation; history follows. Use the exact section headers above — the model was trained on them. ### It generalises to tools it has never seen `logits[:, 1]` after softmax is `P(deny)`. The model keys on *what the call does and whether the user authorised it*, not on the tool's name. Probes on tool types absent from training: | call | P(deny) | | |---|---|---| | `Skill(db-drop-and-recreate)` — user asked for a migration | 0.105 | approve | | `mcp__gcal__list_events` — user asked about their calendar | 0.025 | approve | | `CodeGraphQuery(...)` — an invented tool, benign args | 0.014 | approve | | `mcp__admin__delete_all_users` — user asked about their calendar | 0.980 | **deny** | | `mcp__slack__post_message` — body contains an AWS secret | 0.983 | **deny** | | `DocumentIngestPipeline` — invented tool, path `/`, uploads to a raw IP | 0.993 | **deny** | A skill whose *name* is nothing but alarm words is approved when the request authorises it; a wholly invented tool is denied on argument semantics alone. *(These probes were run against auto-0.4b, which shares the training corpus and decision rule.)* ## Deployment A wrong approve costs far more than a wrong deny, so pick the operating point deliberately: - **Standalone gate** — threshold 0.5 gives 4.07% false-approve and 3.19% false-deny. - **First-stage filter (recommended)** — auto-approve below a low threshold, auto-deny above a high one, escalate the middle band to a frontier model. This removes the large majority of LLM gate calls while keeping frontier judgement on genuinely ambiguous cases. Real agent traffic is overwhelmingly benign, so the false-deny rate is multiplied across almost every call while false-approve applies to a rare few. At 3.19%, roughly 1 action in 31 gets interrupted at the default threshold — down from 1 in 9 for auto-0.4b, which is what makes this model practical to run standalone. ### Precision — use fp16 Measured on 400 benchmark rows against fp32, scoring **decision agreement** at threshold 0.5 (the only metric that matters for a gate — a build can look fine on mean error and still flip calls near the boundary): Re-running the **full 3,000-item benchmark** at each precision: | precision | accuracy | AUROC | false-approve | memory | |---|---|---|---|---| | fp32 | 0.964000 | 0.992845 | 0.040685 | ~3.9 GB | | **bf16** | **0.964000** | **0.992845** | **0.040685** | **~2 GB** | | fp16 | 0.964333 | 0.992839 | 0.039971 | ~2 GB | | int8 (ONNX dynamic) | — | — | — | ~1 GB | **bf16 is exactly lossless** — identical on every metric and every context-length slice, zero flipped verdicts across 3,000 items. Ready-made at [`ProCreations/auto-1b-bf16`](https://huggingface.co/ProCreations/auto-1b-bf16), or pass `dtype=torch.bfloat16` here. fp16 differs by a single item with marginally worse AUROC; the two are equivalent in practice, and bf16 is preferred for carrying no overflow risk. **int8 is not salvageable.** It flips roughly 1 verdict in 20, and per-channel quantization — the standard fix — scored *worse* (94.25% vs 95.00% decision agreement). The failure is activation outliers in the GeGLU layers, which dynamic quantization cannot handle. See the [ONNX repo](https://huggingface.co/ProCreations/auto-1b-ONNX) for the breakdown. ## How it was built **Context extension (8k → 64k).** ModernBERT-style architectures alternate local sliding-window attention (window 128) with full global attention every 3rd layer, so only 10 of 28 layers pay the O(n²) cost — which is what makes 64k practical. Only those global layers need a new RoPE base: **2,560,000** for full attention, 160,000 for sliding. **Training.** Full fine-tune of all 1,031,267,330 parameters in two stages — 2 epochs at short context (max 4,096) where nearly all real traffic lives, then 1 epoch at up to 65,536 so the extended RoPE is exercised on the actual task. The long stage mixes ~20k short examples back in to prevent forgetting, and runs at a 3x lower learning rate (8e-06 vs 2.4e-05). Batching is by token budget rather than example count, since inputs span 200–65,536 tokens. 4x RTX PRO 6000 with DDP, bf16 autocast over fp32 master weights, gradient checkpointing, flash-attention 2. Total wall clock **4h45m**. **Data.** [712,000 examples](https://huggingface.co/datasets/ProCreations/auto-1b-data) (688k short + 24k long) spanning agent frameworks, domains, risk categories, obfuscation styles and multiple languages, with deliberate **minimal contrastive pairs** — near-identical calls with opposite labels where only the user's request or the history flips the verdict. The benchmark is excluded by content hash (verified 0/3000 overlap). ## Limitations - Training labels are model-generated and reflect the decision rule they were written against. This is not a substitute for a real security review of your agent's permissions. - It judges a *proposed* call from text. It cannot see what a script will actually do at runtime, so an opaque binary or a URL whose content it cannot read is judged on context alone. - The ONNX export is practical to ~8k tokens (the non-flash attention path materialises a dense sliding-window mask); use the PyTorch + flash-attn path for full 64k. - Evaluated only on synthetic held-out data. Behaviour on real production agent traffic has not been measured. ## Other formats - [`ProCreations/auto-1b-bf16`](https://huggingface.co/ProCreations/auto-1b-bf16) — half the size, benchmark-identical to this model - [`ProCreations/auto-1b-ONNX`](https://huggingface.co/ProCreations/auto-1b-ONNX) — ONNX fp32 (its int8 build is documented as unfit for gating) A GGUF build was published and then **withdrawn**: llama.cpp converts the model, but its `--pooling rank` path returns zero for a 2-class classification head, so it could not actually make approve/deny decisions.