QuantizedSarcBERT-static-v2

INT8 statically quantized ONNX build of Vandita/Bert-finetuned-Sarc, for CPU inference.

Padding invariance β€” read this before serving

Activation scales are frozen at calibration time, so they cannot change with input shape. Measured on this build:

variant pad-sensitivity mixed-batch safe to batch?
fp32 (reference) 1.667e-09 3.904e-08 yes
INT8 dynamic PTQ 8.816e-01 9.173e-01 no β€” serve unpadded only
INT8 static (this model) 0.000e+00 0.000e+00 yes

pad-sensitivity = max change in P(sarcastic) for the same text alone vs padded, at batch size 1. mixed-batch = same text alone vs inside a mixed-length batch.

Dynamic PTQ recomputes activation scales from each tensor's observed range at inference time. Padded positions enter that range, so they change the scale, so they change how the real tokens are quantized β€” the attention mask stops padded positions contributing to attention, but not to the quantization range. Static quantization removes the dependence entirely.

Results

SarcOjiTest1

variant accuracy precision recall F1 MCC ROC-AUC
fp32 (original) 0.6605 0.7027 0.6143 0.6555 0.3266 0.7135
INT8 static (this model) 0.6507 0.6842 0.6235 0.6524 0.3041 0.7040

SarcOjiTest2

variant accuracy precision recall F1 MCC ROC-AUC
fp32 (original) 0.7124 0.4488 0.6635 0.5355 0.3518 0.7419
INT8 static (this model) 0.6990 0.4328 0.6603 0.5229 0.3317 0.7417

SarcOjiTest2 is ~75/25 negative-skewed; MCC and ROC-AUC are the meaningful columns there.

Performance (x86 CPU, single thread, batch-of-1)

variant p50 (ms) p95 (ms) size (MB)
fp32 67.95 90.32 437.6
INT8 static 29.04 52.28 183.5

Not an on-device ARM measurement.

Quantization recipe

  • Static PTQ, QDQ format, per-channel weights
  • Calibration: percentile on 128 in-domain SarcOji samples, padded as deployed
  • Ops quantized: MatMul and Gemm only (where BERT's compute is; leaving Add/LayerNorm/Softmax in float preserves accuracy at negligible speed cost)
  • Exported with eager attention, dynamic batch and sequence axes, opset 17

Usage

from optimum.onnxruntime import ORTModelForSequenceClassification
from transformers import AutoTokenizer

tok = AutoTokenizer.from_pretrained("Vandita/QuantizedSarcBERT-static-v2")
model = ORTModelForSequenceClassification.from_pretrained(
    "Vandita/QuantizedSarcBERT-static-v2", file_name="model_int8_static.onnx")

inputs = tok(["Oh great, another meeting."], return_tensors="pt", padding=True)
logits = model(**inputs).logits   # batching is safe on this build
Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for Vandita/QuantizedSarcBERT-static-v2

Quantized
(1)
this model