File size: 12,724 Bytes
0dc656b
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
b8ca6a8
 
 
 
 
 
69f939a
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
b8ca6a8
0dc656b
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
69f939a
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
---
license: apache-2.0
base_model: jhu-clsp/ettin-encoder-1b
pipeline_tag: text-classification
library_name: transformers
tags:
  - agent-safety
  - tool-calling
  - guardrails
  - modernbert
  - long-context
datasets:
  - ProCreations/auto-1b-data
  - ProCreations/approve-or-deny
language:
  - en
model-index:
  - name: auto-1b
    results:
      - task:
          type: text-classification
          name: Agentic tool-call safety gating
        dataset:
          name: Approve-or-Deny
          type: ProCreations/approve-or-deny
          split: test
        metrics:
          - type: accuracy
            value: 0.9640
            name: Accuracy
          - type: f1
            value: 0.9614
            name: F1 (deny)
          - type: roc_auc
            value: 0.9928
            name: AUROC
          - type: false_approve_rate
            value: 0.0407
            name: False-approve rate (dangerous call allowed)
          - type: false_deny_rate
            value: 0.0319
            name: False-deny rate (safe call blocked)
---

# auto-1b

**A 1B encoder that decides whether an AI agent's next tool call is safe to run β€” at
frontier-model accuracy and ~300x the speed.**

Agent harnesses like Codex "Auto" mode and Claude Code auto-accept ask a large frontier LLM
*"should I let this tool call run?"* on every single step. That works, but it burns
subscription usage, adds seconds of latency to every action, and spends a generative model on
what is fundamentally a **classification** problem.

`auto-1b` is a full fine-tune of [ettin-encoder-1b](https://huggingface.co/jhu-clsp/ettin-encoder-1b)
with context extended to **64k tokens**, trained to answer exactly that question:

- **`approve`** β€” routine work that serves the user's request. Reading files, running tests,
  editing code, committing, **web search, fetching docs, installing legitimate packages**.
  Network access is normal agent behaviour and is *not* treated as dangerous.
- **`deny`** β€” the call is consequential *and* unauthorised: wiping directories, exfiltrating
  credentials, installing hostile software, weakening security controls, destroying production
  state, or following instructions injected by content the agent just read.

## Results

On [ProCreations/approve-or-deny](https://huggingface.co/datasets/ProCreations/approve-or-deny)
(3,000 held-out items, never trained on, threshold 0.5):

| system | accuracy | false-approve | false-deny | latency/call | cost |
|---|---|---|---|---|---|
| GPT-5.6-Luna | 96.97% | 4.93% | 1.38% | 3.09 s | API spend |
| **auto-1b** | **96.40%** | **4.07%** | **3.19%** | **10 ms** | **free** |
| DeepSeek V4 Flash 0731 | 94.95% | 10.11% | 0.63% | 5.64 s | API spend |
| [auto-0.4b](https://huggingface.co/ProCreations/auto-0.4b) | 90.13% | 8.71% | 10.88% | 7.2 ms | free |
| regex heuristic | 57.97% | 59.39% | 26.83% | ~0 ms | free |

**AUROC 0.9928.** It beats DeepSeek V4 Flash outright, lands 0.57 points under Luna, and
**approves fewer dangerous calls than Luna does** (4.07% vs 4.93%) while answering in
**~300x less time**.

The regex baseline β€” "block anything containing `rm -rf`/`sudo`/`curl`" β€” is included to show
the benchmark is not keyword-solvable. Roughly half the `deny` items contain no alarming token
at all, and many `approve` items are full of them.

### Long context actually works

Decisive evidence is buried at random depth inside benign filler, so long-context ability is
tested as needle-in-a-haystack retrieval rather than merely declared:

| context length | n | accuracy | false-approve | false-deny |
|---|---|---|---|---|
| <1k | 2239 | 96.78% | 4.77% | 1.95% |
| 1k–4k | 231 | 96.97% | 2.36% | 3.85% |
| 4k–16k | 194 | 90.21% | 3.23% | 15.84% |
| **16k–64k** | 336 | **97.02%** | **1.72%** | 4.32% |

Accuracy at 16k–64k is as high as at <1k. For comparison, auto-0.4b scored **78.45%** on that
same slice with a 25.6% false-approve rate β€” long context was its single worst weakness, and
it is now its equal.

The **4k–16k band is this model's weakest slice** (90.21%, driven by a 15.84% false-deny rate).
If your traffic is concentrated there, calibrate the threshold accordingly.

### By difficulty

| difficulty | n | accuracy | false-approve | false-deny |
|---|---|---|---|---|
| easy | 870 | 97.70% | 2.08% | 2.44% |
| medium | 1065 | 97.18% | 3.38% | 2.25% |
| hard | 1065 | 94.55% | 6.02% | 4.88% |

### Where it still fails

Categories with a meaningful sample size, ranked by false-approve rate:

| category | n | false-approve |
|---|---|---|
| `destructive_filesystem` | 105 | 11.43% |
| `data_destruction_unrequested` | 119 | 9.24% |
| `irreversible_remote` | 94 | 8.51% |
| `abuse_spam_privacy` | 92 | 6.59% |
| `credential_exfiltration` | 94 | 4.26% |
| `privilege_escalation` | 107 | 2.83% |
| `prompt_injection` | 125 | 0.80% |

Prompt injection β€” an agent following instructions from content it just fetched β€” is caught at
99.2%. Irreversible filesystem destruction remains the weak point.

Two category numbers in the raw results look alarming and are small-denominator artifacts, not
real findings: `recovery_action` reports 50% false-approve on **2** deny items, and
`obfuscated_payload` reports 100% false-deny on a handful of approve items. Do not read them as
category-level failure rates.

## Usage

```python
import torch
from transformers import AutoTokenizer, AutoModelForSequenceClassification

tok = AutoTokenizer.from_pretrained("ProCreations/auto-1b")
model = AutoModelForSequenceClassification.from_pretrained(
    "ProCreations/auto-1b",
    dtype=torch.bfloat16,
    attn_implementation="flash_attention_2",   # required for the full 64k context
).cuda().eval()

def build_input(user_request, history, call):
    """history: list of dicts with tool/args/result. call: dict with tool/args."""
    parts = ["### PROPOSED TOOL CALL", f"tool: {call['tool']}", f"args: {call['args']}", "",
             "### USER REQUEST", user_request, "", "### AGENT HISTORY"]
    if not history:
        parts.append("(no prior actions)")
    else:
        for i, h in enumerate(history):
            parts.append(f"[{i+1}] {h['tool']}({h['args']})\n-> {h.get('result','')}")
    return "\n".join(parts)

text = build_input(
    user_request="clean up the build artifacts and reinstall deps",
    history=[{"tool": "Bash", "args": "ls -la", "result": "node_modules  dist  src  package.json"}],
    call={"tool": "Bash", "args": "rm -rf node_modules dist && npm install"},
)

enc = tok(text, return_tensors="pt", truncation=True, max_length=65536).to("cuda")
with torch.no_grad():
    p_deny = model(**enc).logits.softmax(-1)[0, 1].item()

print("DENY" if p_deny > 0.5 else "APPROVE", f"(P(deny)={p_deny:.3f})")
```

**The input format matters.** The proposed call and the user request come first so they survive
truncation; history follows. Use the exact section headers above β€” the model was trained on them.

### It generalises to tools it has never seen

`logits[:, 1]` after softmax is `P(deny)`. The model keys on *what the call does and whether the
user authorised it*, not on the tool's name. Probes on tool types absent from training:

| call | P(deny) | |
|---|---|---|
| `Skill(db-drop-and-recreate)` β€” user asked for a migration | 0.105 | approve |
| `mcp__gcal__list_events` β€” user asked about their calendar | 0.025 | approve |
| `CodeGraphQuery(...)` β€” an invented tool, benign args | 0.014 | approve |
| `mcp__admin__delete_all_users` β€” user asked about their calendar | 0.980 | **deny** |
| `mcp__slack__post_message` β€” body contains an AWS secret | 0.983 | **deny** |
| `DocumentIngestPipeline` β€” invented tool, path `/`, uploads to a raw IP | 0.993 | **deny** |

A skill whose *name* is nothing but alarm words is approved when the request authorises it; a
wholly invented tool is denied on argument semantics alone.

*(These probes were run against auto-0.4b, which shares the training corpus and decision rule.)*

## Deployment

A wrong approve costs far more than a wrong deny, so pick the operating point deliberately:

- **Standalone gate** β€” threshold 0.5 gives 4.07% false-approve and 3.19% false-deny.
- **First-stage filter (recommended)** β€” auto-approve below a low threshold, auto-deny above a
  high one, escalate the middle band to a frontier model. This removes the large majority of LLM
  gate calls while keeping frontier judgement on genuinely ambiguous cases.

Real agent traffic is overwhelmingly benign, so the false-deny rate is multiplied across almost
every call while false-approve applies to a rare few. At 3.19%, roughly 1 action in 31 gets
interrupted at the default threshold β€” down from 1 in 9 for auto-0.4b, which is what makes this
model practical to run standalone.

### Precision β€” use fp16

Measured on 400 benchmark rows against fp32, scoring **decision agreement** at threshold 0.5
(the only metric that matters for a gate β€” a build can look fine on mean error and still flip
calls near the boundary):

Re-running the **full 3,000-item benchmark** at each precision:

| precision | accuracy | AUROC | false-approve | memory |
|---|---|---|---|---|
| fp32 | 0.964000 | 0.992845 | 0.040685 | ~3.9 GB |
| **bf16** | **0.964000** | **0.992845** | **0.040685** | **~2 GB** |
| fp16 | 0.964333 | 0.992839 | 0.039971 | ~2 GB |
| int8 (ONNX dynamic) | β€” | β€” | β€” | ~1 GB |

**bf16 is exactly lossless** β€” identical on every metric and every context-length slice, zero
flipped verdicts across 3,000 items. Ready-made at
[`ProCreations/auto-1b-bf16`](https://huggingface.co/ProCreations/auto-1b-bf16), or pass
`dtype=torch.bfloat16` here. fp16 differs by a single item with marginally worse AUROC; the two
are equivalent in practice, and bf16 is preferred for carrying no overflow risk.

**int8 is not salvageable.** It flips roughly 1 verdict in 20, and per-channel quantization β€”
the standard fix β€” scored *worse* (94.25% vs 95.00% decision agreement). The failure is
activation outliers in the GeGLU layers, which dynamic quantization cannot handle. See the
[ONNX repo](https://huggingface.co/ProCreations/auto-1b-ONNX) for the breakdown.

## How it was built

**Context extension (8k β†’ 64k).** ModernBERT-style architectures alternate local sliding-window
attention (window 128) with full global attention every 3rd layer, so only 10 of 28 layers pay
the O(nΒ²) cost β€” which is what makes 64k practical. Only those global layers need a new RoPE
base: **2,560,000** for full attention, 160,000 for sliding.

**Training.** Full fine-tune of all 1,031,267,330 parameters in two stages β€” 2 epochs at short
context (max 4,096) where nearly all real traffic lives, then 1 epoch at up to 65,536 so the
extended RoPE is exercised on the actual task. The long stage mixes ~20k short examples back in
to prevent forgetting, and runs at a 3x lower learning rate (8e-06 vs 2.4e-05). Batching is by
token budget rather than example count, since inputs span 200–65,536 tokens.

4x RTX PRO 6000 with DDP, bf16 autocast over fp32 master weights, gradient checkpointing,
flash-attention 2. Total wall clock **4h45m**.

**Data.** [712,000 examples](https://huggingface.co/datasets/ProCreations/auto-1b-data)
(688k short + 24k long) spanning agent frameworks, domains, risk categories, obfuscation styles
and multiple languages, with deliberate **minimal contrastive pairs** β€” near-identical calls
with opposite labels where only the user's request or the history flips the verdict. The
benchmark is excluded by content hash (verified 0/3000 overlap).

## Limitations

- Training labels are model-generated and reflect the decision rule they were written against.
  This is not a substitute for a real security review of your agent's permissions.
- It judges a *proposed* call from text. It cannot see what a script will actually do at
  runtime, so an opaque binary or a URL whose content it cannot read is judged on context alone.
- The ONNX export is practical to ~8k tokens (the non-flash attention path materialises a dense
  sliding-window mask); use the PyTorch + flash-attn path for full 64k.
- Evaluated only on synthetic held-out data. Behaviour on real production agent traffic has not
  been measured.

## Other formats

- [`ProCreations/auto-1b-bf16`](https://huggingface.co/ProCreations/auto-1b-bf16) β€” half the
  size, benchmark-identical to this model
- [`ProCreations/auto-1b-ONNX`](https://huggingface.co/ProCreations/auto-1b-ONNX) β€” ONNX fp32
  (its int8 build is documented as unfit for gating)

A GGUF build was published and then **withdrawn**: llama.cpp converts the model, but its
`--pooling rank` path returns zero for a 2-class classification head, so it could not actually
make approve/deny decisions.