domi29's picture
tags: curated discovery keywords (remove competitor names, dedupe, add search terms)
8596617 verified
|
Raw
History Blame Contribute Delete
2.97 kB
metadata
license: apache-2.0
tags:
  - document-classification
  - topic-classification
  - dlp
  - data-loss-prevention
  - security
  - llm-security
  - ai-safety
  - ai-agent-security
  - patronus
  - multilingual
  - modernbert
  - onnx
  - quantized
  - edge
base_model: patronus-studio/orca-sonar-document-classifier
pipeline_tag: text-classification
language:
  - de
  - en

Model Card for Orca Sonar Document Classifier Edge

Edge/quantized ONNX variant of patronus-studio/orca-sonar-document-classifier. Private / internal.

Quantized with the Patronus RunPod ONNX method (runpod/quantize_single_heads.py): quantize_dynamic for INT8 linears + MatMulNBitsQuantizer (4-bit, block 128) for the embedding. The parent FP32/FP16 model is unchanged; this repo only adds the small variants.

7-class document topic classifier (mmBERT-small / ModernBERT). Metrics measured on the held-out document-classifier test set; FP32 reproduces the parent card (macro-F1 0.978).

Variants (measured, essentially lossless)

Variante Größe macro-F1 Δ vs FP32 Accuracy
fp16 282 MB 0.9783 +0.0000 0.9783
int8 142 MB 0.9793 +0.0011 0.9792
int8_int4_embeddings 96 MB 0.9790 +0.0007 0.9792

Recommended: onnx/int8_int4_embeddings/model.onnx — ~5.5× smaller than FP32 at ~0 quality loss.

Files

  • onnx/fp16/model.onnx, onnx/int8/model.onnx, onnx/int8_int4_embeddings/model.onnx
  • config.json, tokenizer.json, tokenizer_config.json (from the parent)
  • metrics/quant_bench.json — the full FP32-vs-quant benchmark

Usage

import onnxruntime as ort, numpy as np
from transformers import AutoTokenizer
tok = AutoTokenizer.from_pretrained("patronus-studio/orca-sonar-document-classifier-edge")
sess = ort.InferenceSession("onnx/int8_int4_embeddings/model.onnx", providers=["CPUExecutionProvider"])
enc = tok("your text", truncation=True, max_length=256, padding="max_length", return_tensors="np")
logits = sess.run(["logits"], {"input_ids": enc["input_ids"].astype(np.int64),
                               "attention_mask": enc["attention_mask"].astype(np.int64)})[0]

License

This model is released under the Apache License 2.0. A copy of the license is included as LICENSE in this repository.

Patronus Ark

This model is built to run inside Patronus Ark, Patronus' open-source on-device AI-security scanning library (L1 native rules → L2 NTDB cascade → L3 transformer). Ark is not publicly released yet — a repository link will be added here at launch.


🛡️ Patronus Protect

Brought to you by Patronus Protect — a local AI firewall that secures every AI interaction (prompts, tools, documents) before it reaches your models. Try it for free at patronus.studio.