Model Card for Orca Sonar Document Classifier Edge

Edge/quantized ONNX variant of patronus-studio/orca-sonar-document-classifier. Private / internal.

Quantized with the Patronus RunPod ONNX method (runpod/quantize_single_heads.py): quantize_dynamic for INT8 linears + MatMulNBitsQuantizer (4-bit, block 128) for the embedding. The parent FP32/FP16 model is unchanged; this repo only adds the small variants.

7-class document topic classifier (mmBERT-small / ModernBERT). Metrics measured on the held-out document-classifier test set; FP32 reproduces the parent card (macro-F1 0.978).

Variants (measured, essentially lossless)

Variante Größe macro-F1 Δ vs FP32 Accuracy
fp16 282 MB 0.9783 +0.0000 0.9783
int8 142 MB 0.9793 +0.0011 0.9792
int8_int4_embeddings 96 MB 0.9790 +0.0007 0.9792

Recommended: onnx/int8_int4_embeddings/model.onnx — ~5.5× smaller than FP32 at ~0 quality loss.

Files

  • onnx/fp16/model.onnx, onnx/int8/model.onnx, onnx/int8_int4_embeddings/model.onnx
  • config.json, tokenizer.json, tokenizer_config.json (from the parent)
  • metrics/quant_bench.json — the full FP32-vs-quant benchmark

Usage

import onnxruntime as ort, numpy as np
from transformers import AutoTokenizer
tok = AutoTokenizer.from_pretrained("patronus-studio/orca-sonar-document-classifier-edge")
sess = ort.InferenceSession("onnx/int8_int4_embeddings/model.onnx", providers=["CPUExecutionProvider"])
enc = tok("your text", truncation=True, max_length=256, padding="max_length", return_tensors="np")
logits = sess.run(["logits"], {"input_ids": enc["input_ids"].astype(np.int64),
                               "attention_mask": enc["attention_mask"].astype(np.int64)})[0]

License

This model is released under the Apache License 2.0. A copy of the license is included as LICENSE in this repository.

Patronus Ark

This model is built to run inside Patronus Ark, Patronus' open-source on-device AI-security scanning library (L1 native rules → L2 NTDB cascade → L3 transformer). Ark is not publicly released yet — a repository link will be added here at launch.


🛡️ Patronus Protect

Brought to you by Patronus Protect — a local AI firewall that secures every AI interaction (prompts, tools, documents) before it reaches your models. Try it for free at patronus.studio.

Downloads last month
3
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for patronus-studio/orca-sonar-document-classifier-edge

Quantized
(1)
this model

Collection including patronus-studio/orca-sonar-document-classifier-edge