Model Card for Orca Sonar Document Classifier Edge
Edge/quantized ONNX variant of patronus-studio/orca-sonar-document-classifier. Private / internal.
Quantized with the Patronus RunPod ONNX method (runpod/quantize_single_heads.py):
quantize_dynamic for INT8 linears + MatMulNBitsQuantizer (4-bit, block 128) for the
embedding. The parent FP32/FP16 model is unchanged; this repo only adds the small variants.
7-class document topic classifier (mmBERT-small / ModernBERT). Metrics measured on the held-out document-classifier test set; FP32 reproduces the parent card (macro-F1 0.978).
Variants (measured, essentially lossless)
| Variante | Größe | macro-F1 | Δ vs FP32 | Accuracy |
|---|---|---|---|---|
| fp16 | 282 MB | 0.9783 | +0.0000 | 0.9783 |
| int8 | 142 MB | 0.9793 | +0.0011 | 0.9792 |
| int8_int4_embeddings | 96 MB | 0.9790 | +0.0007 | 0.9792 |
Recommended: onnx/int8_int4_embeddings/model.onnx — ~5.5× smaller than FP32 at ~0 quality loss.
Files
onnx/fp16/model.onnx,onnx/int8/model.onnx,onnx/int8_int4_embeddings/model.onnxconfig.json,tokenizer.json,tokenizer_config.json(from the parent)metrics/quant_bench.json— the full FP32-vs-quant benchmark
Usage
import onnxruntime as ort, numpy as np
from transformers import AutoTokenizer
tok = AutoTokenizer.from_pretrained("patronus-studio/orca-sonar-document-classifier-edge")
sess = ort.InferenceSession("onnx/int8_int4_embeddings/model.onnx", providers=["CPUExecutionProvider"])
enc = tok("your text", truncation=True, max_length=256, padding="max_length", return_tensors="np")
logits = sess.run(["logits"], {"input_ids": enc["input_ids"].astype(np.int64),
"attention_mask": enc["attention_mask"].astype(np.int64)})[0]
License
This model is released under the Apache License 2.0.
A copy of the license is included as LICENSE in this repository.
Patronus Ark
This model is built to run inside Patronus Ark, Patronus' open-source on-device AI-security scanning library (L1 native rules → L2 NTDB cascade → L3 transformer). Ark is not publicly released yet — a repository link will be added here at launch.
🛡️ Patronus Protect
Brought to you by Patronus Protect — a local AI firewall that secures every AI interaction (prompts, tools, documents) before it reaches your models. Try it for free at patronus.studio.
- Downloads last month
- 3
Model tree for patronus-studio/orca-sonar-document-classifier-edge
Base model
jhu-clsp/mmBERT-small