auto-0.4b-ONNX / README.md
ProCreations's picture
Upload folder using huggingface_hub
e26a536 verified
|
Raw
History Blame Contribute Delete
1.63 kB
---
license: apache-2.0
base_model: ProCreations/auto-0.4b
pipeline_tag: text-classification
tags: [agent-safety, tool-calling, guardrails, onnx, modernbert]
---
# auto-0.4b — ONNX
ONNX exports of [`ProCreations/auto-0.4b`](https://huggingface.co/ProCreations/auto-0.4b),
the 0.4B encoder that decides whether an AI agent's tool call is safe to run.
| file | precision | size |
|---|---|---|
| `model.onnx` | fp32 | ~1.5 GB |
| `model_int8.onnx` | dynamic int8 | ~400 MB |
```python
import numpy as np, onnxruntime as ort
from transformers import AutoTokenizer
tok = AutoTokenizer.from_pretrained("ProCreations/auto-0.4b-ONNX")
sess = ort.InferenceSession("model_int8.onnx", providers=["CPUExecutionProvider"])
enc = tok(text, return_tensors="np", truncation=True, max_length=8192)
logits = sess.run(None, {"input_ids": enc["input_ids"].astype(np.int64),
"attention_mask": enc["attention_mask"].astype(np.int64)})[0]
p_deny = np.exp(logits[0]) / np.exp(logits[0]).sum()
print("DENY" if p_deny[1] > 0.5 else "APPROVE", p_deny[1])
```
Build the input string with the exact format documented on the
[main model card](https://huggingface.co/ProCreations/auto-0.4b) — proposed call first,
then user request, then history.
## Context length
**Practical limit ~8k tokens.** ModernBERT's non-flash attention path materialises a dense
`(B, 1, L, L)` sliding-window mask, so ONNX memory grows quadratically with sequence length
(at 64k that mask alone would be ~17 GB). Almost all real tool-call decisions are well under
8k. For the full 64k context use the PyTorch + flash-attn path in the main repo.