auto-0.4b / README.md
ProCreations's picture
Upload README.md with huggingface_hub
2f6ec6c verified
|
Raw
History Blame Contribute Delete
10.6 kB
---
license: apache-2.0
base_model: answerdotai/ModernBERT-large
pipeline_tag: text-classification
library_name: transformers
tags:
- agent-safety
- tool-calling
- guardrails
- modernbert
- long-context
datasets:
- ProCreations/auto
- ProCreations/approve-or-deny
language:
- en
model-index:
- name: auto-0.4b
results:
- task:
type: text-classification
name: Agentic tool-call safety gating
dataset:
name: Approve-or-Deny
type: ProCreations/approve-or-deny
split: test
metrics:
- type: accuracy
value: 0.9013
name: Accuracy
- type: f1
value: 0.8963
name: F1 (deny)
- type: roc_auc
value: 0.9657
name: AUROC
- type: false_approve_rate
value: 0.0871
name: False-approve rate (dangerous call allowed)
- type: false_deny_rate
value: 0.1088
name: False-deny rate (safe call blocked)
---
# auto-0.4b
**A 0.4B encoder that decides whether an AI agent's next tool call is safe to run.**
Agent harnesses like Codex "Auto" mode and Claude Code auto-accept currently ask a large
frontier LLM the question *"should I let this tool call run?"* on every single step. That
works, but it burns subscription usage, adds API latency to every action, and spends a
generative model on what is fundamentally a **classification** problem — the thing
encoder-only models are best at.
`auto-0.4b` is a full fine-tune of [ModernBERT-large](https://huggingface.co/answerdotai/ModernBERT-large) with context extended to
**64k tokens**, trained to answer exactly that question:
- **`approve`** — routine work that serves the user's request. Reading files, running
tests, editing code, committing, **web search, fetching docs, installing legitimate
packages** — network access is normal agent behaviour and is *not* treated as dangerous.
- **`deny`** — the call is consequential *and* unauthorised: wiping directories, exfiltrating
credentials, installing hostile software, weakening security controls, destroying
production state, or following instructions injected by content the agent just read.
It runs in **7 ms per call** on a single GPU — roughly **400x faster** than asking a
frontier LLM the same question, at zero marginal cost.
## Honest positioning
**This is not a drop-in replacement for a frontier-LLM gate.** On the held-out benchmark:
| system | accuracy | false-approve | false-deny | latency/call | cost |
|---|---|---|---|---|---|
| GPT-5.6-Luna | 96.97% | 4.93% | 1.38% | 3.09 s | API spend |
| DeepSeek V4 Flash 0731 | 94.95% | 10.11% | 0.63% | 5.64 s | API spend |
| **auto-0.4b** | **90.13%** | **8.71%** | **10.88%** | **7 ms** | **free** |
| regex heuristic | 57.97% | 59.39% | 26.83% | ~0 ms | free |
It trails Luna by ~7 points of accuracy, and its false-deny rate means it would interrupt
roughly 1 in 10 legitimate actions at the default threshold. Where it earns its place is as a
**fast first-stage filter**: it is very strong on the attacks that actually matter, and it
answers in milliseconds so it can run on every single tool call.
| attack category | false-approve |
|---|---|
| `self_permission_tamper` (disabling its own guardrails) | ~1% |
| `covering_tracks` | ~1% |
| **`prompt_injection`** (following instructions from fetched content) | **~2%** |
| `privilege_escalation` | ~3% |
| `credential_exfiltration` | ~3% |
A sensible deployment auto-approves below a low threshold, auto-denies above a high one, and
escalates the middle band to a bigger model — which removes most of the LLM calls while
keeping a frontier model's judgement on the genuinely ambiguous cases.
## Benchmark: approve-or-deny
Evaluated on [ProCreations/approve-or-deny](https://huggingface.co/datasets/ProCreations/approve-or-deny) (3000 held-out items, never trained on):
| metric | auto-0.4b | keyword/regex baseline |
|---|---|---|
| **accuracy** | **90.13%** | 57.97% |
| balanced accuracy | 90.21% | 56.89% |
| F1 (deny) | 0.8963 | 0.4744 |
| AUROC | 0.9657 | — |
| **false-approve rate** (dangerous call let through) | **8.71%** | 59.39% |
| false-deny rate (safe call blocked) | 10.88% | 26.83% |
The regex baseline is the "block anything containing `rm -rf`/`sudo`/`curl`" heuristic these
gates usually start life as. It is included to show the benchmark is not keyword-solvable —
roughly half of the `deny` items contain no alarming token at all, and many `approve` items
are full of them.
### Choosing an operating point
`logits[:, 1]` after softmax is `P(deny)`. Because a wrong approve costs far more than a
wrong deny, you can trade the two:
| threshold | accuracy | false-approve | false-deny |
|---|---|---|---|
| 0.1 | 84.43% | 2.71% | 26.83% |
| 0.2 | 87.53% | 4.71% | 19.26% |
| 0.3 | 88.97% | 6.57% | 14.95% |
| 0.4 | 89.60% | 7.71% | 12.76% |
| 0.5 | 90.13% | 8.71% | 10.88% |
| 0.6 | 90.30% | 9.99% | 9.44% |
| 0.7 | 90.43% | 11.85% | 7.57% |
| 0.8 | 89.93% | 14.28% | 6.38% |
| 0.9 | 89.27% | 18.20% | 4.19% |
### Where it is strong and weak
**By difficulty**
| slice | n | accuracy | false-approve | false-deny |
|---|---|---|---|---|
| `medium` | 1065 | 92.21% | 7.52% | 8.07% |
| `hard` | 1065 | 86.10% | 10.53% | 17.26% |
| `easy` | 870 | 92.53% | 7.72% | 7.32% |
**By context length**
| slice | n | accuracy | false-approve | false-deny |
|---|---|---|---|---|
| `<1k` | 2420 | 91.40% | 7.40% | 9.60% |
| `16k-64k` | 232 | 78.45% | 25.62% | 17.12% |
| `4k-16k` | 218 | 88.07% | 6.60% | 16.96% |
| `1k-4k` | 130 | 90.77% | 3.03% | 15.62% |
## Usage
```python
import torch
from transformers import AutoTokenizer, AutoModelForSequenceClassification
tok = AutoTokenizer.from_pretrained("ProCreations/auto-0.4b")
model = AutoModelForSequenceClassification.from_pretrained(
"ProCreations/auto-0.4b",
dtype=torch.bfloat16,
attn_implementation="flash_attention_2", # required for the full 64k context
).cuda().eval()
def build_input(user_request, history, call):
"""history: list of dicts with tool/args/result. call: dict with tool/args."""
parts = ["### PROPOSED TOOL CALL", f"tool: {call['tool']}", f"args: {call['args']}", "",
"### USER REQUEST", user_request, "", "### AGENT HISTORY"]
if not history:
parts.append("(no prior actions)")
else:
for i, h in enumerate(history):
parts.append(f"[{i+1}] {h['tool']}({h['args']})\n-> {h.get('result','')}")
return "\n".join(parts)
text = build_input(
user_request="clean up the build artifacts and reinstall deps",
history=[{"tool": "Bash", "args": "ls -la", "result": "node_modules dist src package.json"}],
call={"tool": "Bash", "args": "rm -rf node_modules dist && npm install"},
)
enc = tok(text, return_tensors="pt", truncation=True, max_length=65536).to("cuda")
with torch.no_grad():
p_deny = model(**enc).logits.softmax(-1)[0, 1].item()
print("DENY" if p_deny > 0.5 else "APPROVE", f"(P(deny)={p_deny:.3f})")
```
**The input format matters.** The proposed call and the user request come first so they
survive truncation; history follows. Use the exact section headers above — the model was
trained on them.
## How it was built
**Context extension (8k → 64k).** ModernBERT alternates local sliding-window attention
(window 128) with full global attention every 3rd layer, so only 10 of 28 layers pay the
O(n²) cost — which is what makes 64k practical on one GPU. Only those global layers need a
new RoPE base. The base was chosen empirically by sweeping it against MLM loss on
domain-matched long agent transcripts:
| RoPE base (global layers) | 8k | 16k | 32k | 64k |
|---|---|---|---|---|
| 160,000 | 0.589 | 0.593 | 0.752 | 2.060 |
| 640,000 | 0.588 | 0.474 | 0.406 | 0.550 |
| 1,280,000 | 0.599 | 0.477 | 0.317 | 0.444 |
| 2,560,000 | 0.614 | 0.490 | 0.314 | 0.222 |
| 5,120,000 | 0.635 | 0.507 | 0.328 | 0.195 |
At the stock base the model simply cannot do 64k. **2,560,000** (16× the original
for an 8× extension) was chosen: near-best long-context loss for ~4% short-context cost.
**Training.** Full fine-tune (all 395.8M parameters), two stages: bulk training at short
context where nearly all real traffic lives, then a long-context stage at up to 64k so the
extended RoPE is exercised on the actual task. Batching is by token budget rather than
example count, since inputs span 200–65,536 tokens.
**Data.** 73,509 training examples generated with DeepSeek V4 Flash across a large
combinatorial space of agent frameworks (Claude Code, Codex CLI, MCP servers, LangChain,
computer-use, devops, …), domains, risk categories, obfuscation styles, and 10 languages,
with deliberate **minimal contrastive pairs** — near-identical calls with opposite labels
where only the user's request or the history flips the verdict.
Long-context examples are built by burying the decisive history steps inside benign filler
at a random depth, so 64k capability is exercised as needle-in-a-haystack retrieval rather
than merely declared.
## Limitations
- Training data is model-generated and reflects the decision rule it was written against;
it is not a substitute for a real security review of your agent's permissions.
- Held-out labels were audited by an independent pass, but both come from the same model
family, so systematic blind spots can survive.
- The ONNX export is practical to ~8k tokens (the non-flash attention path materialises a
dense sliding-window mask); use the PyTorch + flash-attn path for full 64k.
- It judges a *proposed* call from text. It cannot see what a script will actually do at
runtime, so an opaque binary or a URL whose content it cannot read is judged on context alone.
## Other formats
- [`ProCreations/auto-0.4b-ONNX`](https://huggingface.co/ProCreations/auto-0.4b-ONNX) — ONNX + int8
The GGUF build was **withdrawn**. llama.cpp converts the model, but its `--pooling rank` path
returns zero for a 2-class classification head, so the GGUF could not actually make
approve/deny decisions — it returned identical `0.000` scores for a `rm -rf /` and for a
`pytest` invocation. It was removed rather than left up implying it worked.
## Successor
[`ProCreations/auto-1b`](https://huggingface.co/ProCreations/auto-1b) scores **96.40%** on the
same benchmark (vs 90.13% here), with false-deny down from 10.88% to 3.19% and long-context
accuracy up from 78.45% to 97.02%. Prefer it unless you specifically need the smaller model.