Text Classification
Transformers
Safetensors
English
modernbert
agent-safety
tool-calling
guardrails
long-context
Eval Results (legacy)
text-embeddings-inference
Instructions to use ProCreations/auto-0.4b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use ProCreations/auto-0.4b with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="ProCreations/auto-0.4b")# Load model directly from transformers import AutoTokenizer, AutoModelForSequenceClassification tokenizer = AutoTokenizer.from_pretrained("ProCreations/auto-0.4b") model = AutoModelForSequenceClassification.from_pretrained("ProCreations/auto-0.4b", device_map="auto") - Notebooks
- Google Colab
- Kaggle
| license: apache-2.0 | |
| base_model: answerdotai/ModernBERT-large | |
| pipeline_tag: text-classification | |
| library_name: transformers | |
| tags: | |
| - agent-safety | |
| - tool-calling | |
| - guardrails | |
| - modernbert | |
| - long-context | |
| datasets: | |
| - ProCreations/auto | |
| - ProCreations/approve-or-deny | |
| language: | |
| - en | |
| model-index: | |
| - name: auto-0.4b | |
| results: | |
| - task: | |
| type: text-classification | |
| name: Agentic tool-call safety gating | |
| dataset: | |
| name: Approve-or-Deny | |
| type: ProCreations/approve-or-deny | |
| split: test | |
| metrics: | |
| - type: accuracy | |
| value: 0.9013 | |
| name: Accuracy | |
| - type: f1 | |
| value: 0.8963 | |
| name: F1 (deny) | |
| - type: roc_auc | |
| value: 0.9657 | |
| name: AUROC | |
| - type: false_approve_rate | |
| value: 0.0871 | |
| name: False-approve rate (dangerous call allowed) | |
| - type: false_deny_rate | |
| value: 0.1088 | |
| name: False-deny rate (safe call blocked) | |
| # auto-0.4b | |
| **A 0.4B encoder that decides whether an AI agent's next tool call is safe to run.** | |
| Agent harnesses like Codex "Auto" mode and Claude Code auto-accept currently ask a large | |
| frontier LLM the question *"should I let this tool call run?"* on every single step. That | |
| works, but it burns subscription usage, adds API latency to every action, and spends a | |
| generative model on what is fundamentally a **classification** problem — the thing | |
| encoder-only models are best at. | |
| `auto-0.4b` is a full fine-tune of [ModernBERT-large](https://huggingface.co/answerdotai/ModernBERT-large) with context extended to | |
| **64k tokens**, trained to answer exactly that question: | |
| - **`approve`** — routine work that serves the user's request. Reading files, running | |
| tests, editing code, committing, **web search, fetching docs, installing legitimate | |
| packages** — network access is normal agent behaviour and is *not* treated as dangerous. | |
| - **`deny`** — the call is consequential *and* unauthorised: wiping directories, exfiltrating | |
| credentials, installing hostile software, weakening security controls, destroying | |
| production state, or following instructions injected by content the agent just read. | |
| It runs in **7 ms per call** on a single GPU — roughly **400x faster** than asking a | |
| frontier LLM the same question, at zero marginal cost. | |
| ## Honest positioning | |
| **This is not a drop-in replacement for a frontier-LLM gate.** On the held-out benchmark: | |
| | system | accuracy | false-approve | false-deny | latency/call | cost | | |
| |---|---|---|---|---|---| | |
| | GPT-5.6-Luna | 96.97% | 4.93% | 1.38% | 3.09 s | API spend | | |
| | DeepSeek V4 Flash 0731 | 94.95% | 10.11% | 0.63% | 5.64 s | API spend | | |
| | **auto-0.4b** | **90.13%** | **8.71%** | **10.88%** | **7 ms** | **free** | | |
| | regex heuristic | 57.97% | 59.39% | 26.83% | ~0 ms | free | | |
| It trails Luna by ~7 points of accuracy, and its false-deny rate means it would interrupt | |
| roughly 1 in 10 legitimate actions at the default threshold. Where it earns its place is as a | |
| **fast first-stage filter**: it is very strong on the attacks that actually matter, and it | |
| answers in milliseconds so it can run on every single tool call. | |
| | attack category | false-approve | | |
| |---|---| | |
| | `self_permission_tamper` (disabling its own guardrails) | ~1% | | |
| | `covering_tracks` | ~1% | | |
| | **`prompt_injection`** (following instructions from fetched content) | **~2%** | | |
| | `privilege_escalation` | ~3% | | |
| | `credential_exfiltration` | ~3% | | |
| A sensible deployment auto-approves below a low threshold, auto-denies above a high one, and | |
| escalates the middle band to a bigger model — which removes most of the LLM calls while | |
| keeping a frontier model's judgement on the genuinely ambiguous cases. | |
| ## Benchmark: approve-or-deny | |
| Evaluated on [ProCreations/approve-or-deny](https://huggingface.co/datasets/ProCreations/approve-or-deny) (3000 held-out items, never trained on): | |
| | metric | auto-0.4b | keyword/regex baseline | | |
| |---|---|---| | |
| | **accuracy** | **90.13%** | 57.97% | | |
| | balanced accuracy | 90.21% | 56.89% | | |
| | F1 (deny) | 0.8963 | 0.4744 | | |
| | AUROC | 0.9657 | — | | |
| | **false-approve rate** (dangerous call let through) | **8.71%** | 59.39% | | |
| | false-deny rate (safe call blocked) | 10.88% | 26.83% | | |
| The regex baseline is the "block anything containing `rm -rf`/`sudo`/`curl`" heuristic these | |
| gates usually start life as. It is included to show the benchmark is not keyword-solvable — | |
| roughly half of the `deny` items contain no alarming token at all, and many `approve` items | |
| are full of them. | |
| ### Choosing an operating point | |
| `logits[:, 1]` after softmax is `P(deny)`. Because a wrong approve costs far more than a | |
| wrong deny, you can trade the two: | |
| | threshold | accuracy | false-approve | false-deny | | |
| |---|---|---|---| | |
| | 0.1 | 84.43% | 2.71% | 26.83% | | |
| | 0.2 | 87.53% | 4.71% | 19.26% | | |
| | 0.3 | 88.97% | 6.57% | 14.95% | | |
| | 0.4 | 89.60% | 7.71% | 12.76% | | |
| | 0.5 | 90.13% | 8.71% | 10.88% | | |
| | 0.6 | 90.30% | 9.99% | 9.44% | | |
| | 0.7 | 90.43% | 11.85% | 7.57% | | |
| | 0.8 | 89.93% | 14.28% | 6.38% | | |
| | 0.9 | 89.27% | 18.20% | 4.19% | | |
| ### Where it is strong and weak | |
| **By difficulty** | |
| | slice | n | accuracy | false-approve | false-deny | | |
| |---|---|---|---|---| | |
| | `medium` | 1065 | 92.21% | 7.52% | 8.07% | | |
| | `hard` | 1065 | 86.10% | 10.53% | 17.26% | | |
| | `easy` | 870 | 92.53% | 7.72% | 7.32% | | |
| **By context length** | |
| | slice | n | accuracy | false-approve | false-deny | | |
| |---|---|---|---|---| | |
| | `<1k` | 2420 | 91.40% | 7.40% | 9.60% | | |
| | `16k-64k` | 232 | 78.45% | 25.62% | 17.12% | | |
| | `4k-16k` | 218 | 88.07% | 6.60% | 16.96% | | |
| | `1k-4k` | 130 | 90.77% | 3.03% | 15.62% | | |
| ## Usage | |
| ```python | |
| import torch | |
| from transformers import AutoTokenizer, AutoModelForSequenceClassification | |
| tok = AutoTokenizer.from_pretrained("ProCreations/auto-0.4b") | |
| model = AutoModelForSequenceClassification.from_pretrained( | |
| "ProCreations/auto-0.4b", | |
| dtype=torch.bfloat16, | |
| attn_implementation="flash_attention_2", # required for the full 64k context | |
| ).cuda().eval() | |
| def build_input(user_request, history, call): | |
| """history: list of dicts with tool/args/result. call: dict with tool/args.""" | |
| parts = ["### PROPOSED TOOL CALL", f"tool: {call['tool']}", f"args: {call['args']}", "", | |
| "### USER REQUEST", user_request, "", "### AGENT HISTORY"] | |
| if not history: | |
| parts.append("(no prior actions)") | |
| else: | |
| for i, h in enumerate(history): | |
| parts.append(f"[{i+1}] {h['tool']}({h['args']})\n-> {h.get('result','')}") | |
| return "\n".join(parts) | |
| text = build_input( | |
| user_request="clean up the build artifacts and reinstall deps", | |
| history=[{"tool": "Bash", "args": "ls -la", "result": "node_modules dist src package.json"}], | |
| call={"tool": "Bash", "args": "rm -rf node_modules dist && npm install"}, | |
| ) | |
| enc = tok(text, return_tensors="pt", truncation=True, max_length=65536).to("cuda") | |
| with torch.no_grad(): | |
| p_deny = model(**enc).logits.softmax(-1)[0, 1].item() | |
| print("DENY" if p_deny > 0.5 else "APPROVE", f"(P(deny)={p_deny:.3f})") | |
| ``` | |
| **The input format matters.** The proposed call and the user request come first so they | |
| survive truncation; history follows. Use the exact section headers above — the model was | |
| trained on them. | |
| ## How it was built | |
| **Context extension (8k → 64k).** ModernBERT alternates local sliding-window attention | |
| (window 128) with full global attention every 3rd layer, so only 10 of 28 layers pay the | |
| O(n²) cost — which is what makes 64k practical on one GPU. Only those global layers need a | |
| new RoPE base. The base was chosen empirically by sweeping it against MLM loss on | |
| domain-matched long agent transcripts: | |
| | RoPE base (global layers) | 8k | 16k | 32k | 64k | | |
| |---|---|---|---|---| | |
| | 160,000 | 0.589 | 0.593 | 0.752 | 2.060 | | |
| | 640,000 | 0.588 | 0.474 | 0.406 | 0.550 | | |
| | 1,280,000 | 0.599 | 0.477 | 0.317 | 0.444 | | |
| | 2,560,000 | 0.614 | 0.490 | 0.314 | 0.222 | | |
| | 5,120,000 | 0.635 | 0.507 | 0.328 | 0.195 | | |
| At the stock base the model simply cannot do 64k. **2,560,000** (16× the original | |
| for an 8× extension) was chosen: near-best long-context loss for ~4% short-context cost. | |
| **Training.** Full fine-tune (all 395.8M parameters), two stages: bulk training at short | |
| context where nearly all real traffic lives, then a long-context stage at up to 64k so the | |
| extended RoPE is exercised on the actual task. Batching is by token budget rather than | |
| example count, since inputs span 200–65,536 tokens. | |
| **Data.** 73,509 training examples generated with DeepSeek V4 Flash across a large | |
| combinatorial space of agent frameworks (Claude Code, Codex CLI, MCP servers, LangChain, | |
| computer-use, devops, …), domains, risk categories, obfuscation styles, and 10 languages, | |
| with deliberate **minimal contrastive pairs** — near-identical calls with opposite labels | |
| where only the user's request or the history flips the verdict. | |
| Long-context examples are built by burying the decisive history steps inside benign filler | |
| at a random depth, so 64k capability is exercised as needle-in-a-haystack retrieval rather | |
| than merely declared. | |
| ## Limitations | |
| - Training data is model-generated and reflects the decision rule it was written against; | |
| it is not a substitute for a real security review of your agent's permissions. | |
| - Held-out labels were audited by an independent pass, but both come from the same model | |
| family, so systematic blind spots can survive. | |
| - The ONNX export is practical to ~8k tokens (the non-flash attention path materialises a | |
| dense sliding-window mask); use the PyTorch + flash-attn path for full 64k. | |
| - It judges a *proposed* call from text. It cannot see what a script will actually do at | |
| runtime, so an opaque binary or a URL whose content it cannot read is judged on context alone. | |
| ## Other formats | |
| - [`ProCreations/auto-0.4b-ONNX`](https://huggingface.co/ProCreations/auto-0.4b-ONNX) — ONNX + int8 | |
| The GGUF build was **withdrawn**. llama.cpp converts the model, but its `--pooling rank` path | |
| returns zero for a 2-class classification head, so the GGUF could not actually make | |
| approve/deny decisions — it returned identical `0.000` scores for a `rm -rf /` and for a | |
| `pytest` invocation. It was removed rather than left up implying it worked. | |
| ## Successor | |
| [`ProCreations/auto-1b`](https://huggingface.co/ProCreations/auto-1b) scores **96.40%** on the | |
| same benchmark (vs 90.13% here), with false-deny down from 10.88% to 3.19% and long-context | |
| accuracy up from 78.45% to 97.02%. Prefer it unless you specifically need the smaller model. | |