File size: 2,974 Bytes
625d59e
 
8596617
 
 
 
 
 
 
 
 
 
 
 
 
 
 
625d59e
8596617
 
 
 
625d59e
 
0077ad7
625d59e
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
f26c15b
2ebf511
 
 
9e40eea
2ebf511
2ba21fb
 
 
 
 
 
f26c15b
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
---
license: apache-2.0
tags:
- document-classification
- topic-classification
- dlp
- data-loss-prevention
- security
- llm-security
- ai-safety
- ai-agent-security
- patronus
- multilingual
- modernbert
- onnx
- quantized
- edge
base_model: patronus-studio/orca-sonar-document-classifier
pipeline_tag: text-classification
language:
- de
- en
---

# Model Card for Orca Sonar Document Classifier Edge

**Edge/quantized ONNX variant of [`patronus-studio/orca-sonar-document-classifier`](https://huggingface.co/patronus-studio/orca-sonar-document-classifier).** Private / internal.

Quantized with the Patronus RunPod ONNX method (`runpod/quantize_single_heads.py`):
`quantize_dynamic` for INT8 linears + `MatMulNBitsQuantizer` (4-bit, block 128) for the
embedding. The parent FP32/FP16 model is unchanged; this repo only adds the small variants.

7-class document topic classifier (mmBERT-small / ModernBERT). Metrics measured on the held-out document-classifier test set; FP32 reproduces the parent card (macro-F1 0.978).

## Variants (measured, essentially lossless)

| Variante | Größe | macro-F1 | Δ vs FP32 | Accuracy |
|---|---|---|---|---|
| fp16 | 282 MB | 0.9783 | +0.0000 | 0.9783 |
| int8 | 142 MB | 0.9793 | +0.0011 | 0.9792 |
| **int8_int4_embeddings** | **96 MB** | **0.9790** | **+0.0007** | 0.9792 |

**Recommended: `onnx/int8_int4_embeddings/model.onnx`** — ~5.5× smaller than FP32 at ~0 quality loss.

## Files
- `onnx/fp16/model.onnx`, `onnx/int8/model.onnx`, `onnx/int8_int4_embeddings/model.onnx`
- `config.json`, `tokenizer.json`, `tokenizer_config.json` (from the parent)
- `metrics/quant_bench.json` — the full FP32-vs-quant benchmark

## Usage
```python
import onnxruntime as ort, numpy as np
from transformers import AutoTokenizer
tok = AutoTokenizer.from_pretrained("patronus-studio/orca-sonar-document-classifier-edge")
sess = ort.InferenceSession("onnx/int8_int4_embeddings/model.onnx", providers=["CPUExecutionProvider"])
enc = tok("your text", truncation=True, max_length=256, padding="max_length", return_tensors="np")
logits = sess.run(["logits"], {"input_ids": enc["input_ids"].astype(np.int64),
                               "attention_mask": enc["attention_mask"].astype(np.int64)})[0]
```

## License

This model is released under the [Apache License 2.0](https://www.apache.org/licenses/LICENSE-2.0).
A copy of the license is included as `LICENSE` in this repository.

## Patronus Ark

This model is built to run inside **Patronus Ark**, Patronus' open-source on-device
AI-security scanning library (L1 native rules → L2 NTDB cascade → L3 transformer).
Ark is not publicly released yet — a repository link will be added here at launch.

---

## 🛡️ Patronus Protect

Brought to you by **[Patronus Protect](https://patronus.studio)** — a local AI firewall that
secures every AI interaction (prompts, tools, documents) before it reaches your models.
**Try it for free at [patronus.studio](https://patronus.studio).**