| --- |
| license: apache-2.0 |
| base_model: ProCreations/auto-1b |
| pipeline_tag: text-classification |
| library_name: onnx |
| tags: |
| - agent-safety |
| - tool-calling |
| - guardrails |
| - onnx |
| --- |
| |
| # auto-1b — ONNX |
|
|
| ONNX exports of [`ProCreations/auto-1b`](https://huggingface.co/ProCreations/auto-1b), a 1B |
| encoder that decides whether an AI agent's next tool call is safe to run (**96.40%** on |
| [approve-or-deny](https://huggingface.co/datasets/ProCreations/approve-or-deny), ahead of |
| DeepSeek V4 Flash and within 0.57 points of GPT-5.6-Luna). |
|
|
| | file | precision | size | verdict | |
| |---|---|---|---| |
| | `model.onnx` + `model.onnx_data` | fp32 | 3.9 GB | **use this** | |
| | `model_int8.onnx` | dynamic int8 | 985 MiB | ⚠️ broken as a gate — see below | |
|
|
| ## ⚠️ The int8 build flips verdicts. Do not gate with it. |
|
|
| Measured against fp32 PyTorch on **400 real benchmark rows** (193 `deny`), scoring **decision |
| agreement at threshold 0.5** — the only metric that matters for a gate, since a build can look |
| fine on mean error and still flip calls near the boundary: |
|
|
| | build | decision agreement | max ΔP(deny) | |
| |---|---|---| |
| | **ONNX fp32** | **100.00%** | 3.2e-06 | |
| | ONNX int8, per-tensor (this file) | 95.00% | 9.4e-01 | |
| | ONNX int8, **per-channel** | 94.25% | 9.4e-01 | |
| | ONNX uint8, per-channel | 93.75% | 9.4e-01 | |
|
|
| The fp32 export is numerically identical to PyTorch. **Every int8 variant changes the verdict on |
| roughly 1 call in 20**, with individual probabilities moving as much as 0.94 — a call the model |
| was certain about flipping to certain in the opposite direction. |
|
|
| **This is not a tuning problem.** Per-channel weight quantization is the standard fix when |
| per-tensor collapses a layer's dynamic range, and here it made things *worse*. The failure is |
| activation outliers in the GeGLU intermediate layers, which ONNX dynamic quantization has no |
| mechanism to handle — that is exactly what LLM.int8()-style mixed-precision decomposition |
| exists for. No `quantize_dynamic` setting recovers it. |
|
|
| The int8 file is kept for research on this failure mode. It is not fit for deployment. |
|
|
| ## Want a smaller footprint? Use bf16. |
|
|
| **Load the PyTorch model in fp16 for half the memory at zero measured cost** — 100.00% decision |
| agreement, max ΔP(deny) = 3.2e-03 over the same 400 rows: |
|
|
| ```python |
| import torch |
| from transformers import AutoModelForSequenceClassification |
| |
| model = AutoModelForSequenceClassification.from_pretrained( |
| "ProCreations/auto-1b", |
| dtype=torch.float16, |
| attn_implementation="flash_attention_2", |
| ).cuda().eval() |
| ``` |
|
|
| **bf16 is the better choice**, and there is a prebuilt one at |
| [`ProCreations/auto-1b-bf16`](https://huggingface.co/ProCreations/auto-1b-bf16). Re-running the |
| full 3,000-item benchmark at each precision, bf16 is *exactly* identical to fp32 (0.964000 |
| accuracy, 0.992845 AUROC, zero flipped verdicts) while fp16 differs by one item with slightly |
| worse AUROC. An earlier revision of this card claimed fp16 beat bf16 — that came from comparing |
| against an fp32 reference running a different attention kernel, and was wrong. |
|
|
| Whatever you choose, **measure decision agreement on your own traffic.** File size and mean |
| error do not predict it. |
|
|
| ## Usage |
|
|
| ```python |
| import numpy as np, onnxruntime as ort |
| from transformers import AutoTokenizer |
| |
| tok = AutoTokenizer.from_pretrained("ProCreations/auto-1b-ONNX") |
| sess = ort.InferenceSession("model.onnx", providers=["CPUExecutionProvider"]) |
| |
| def build_input(user_request, history, call): |
| parts = ["### PROPOSED TOOL CALL", f"tool: {call['tool']}", f"args: {call['args']}", "", |
| "### USER REQUEST", user_request, "", "### AGENT HISTORY"] |
| if not history: |
| parts.append("(no prior actions)") |
| else: |
| for i, h in enumerate(history): |
| parts.append(f"[{i+1}] {h['tool']}({h['args']})\n-> {h.get('result','')}") |
| return "\n".join(parts) |
| |
| text = build_input("clean up build artifacts", [], {"tool": "Bash", "args": "rm -rf node_modules"}) |
| enc = tok(text, return_tensors="np", truncation=True, max_length=8192) |
| logits = sess.run(None, {"input_ids": enc["input_ids"].astype(np.int64), |
| "attention_mask": enc["attention_mask"].astype(np.int64)})[0] |
| p = np.exp(logits[0] - logits[0].max()) |
| p = p / p.sum() |
| print("DENY" if p[1] > 0.5 else "APPROVE", f"P(deny)={p[1]:.3f}") |
| ``` |
|
|
| `logits[:, 1]` after softmax is `P(deny)`. Labels: `0 = approve`, `1 = deny`. |
|
|
| **The input format matters** — the proposed call and user request come first so they survive |
| truncation. Use the exact section headers above; the model was trained on them. |
|
|
| ## Context limit |
|
|
| **Practical to ~8k tokens.** The non-flash attention path materialises a dense `(B, 1, L, L)` |
| sliding-window mask, so memory grows quadratically with sequence length. That covers the great |
| majority of real tool calls (a call plus recent history), but for the full 64k context use the |
| PyTorch + flash-attention path in |
| [`ProCreations/auto-1b`](https://huggingface.co/ProCreations/auto-1b). |
|
|