QuantizedSarcBERT-static-v2
INT8 statically quantized ONNX build of Vandita/Bert-finetuned-Sarc,
for CPU inference.
Padding invariance β read this before serving
Activation scales are frozen at calibration time, so they cannot change with input shape. Measured on this build:
| variant | pad-sensitivity | mixed-batch | safe to batch? |
|---|---|---|---|
| fp32 (reference) | 1.667e-09 | 3.904e-08 | yes |
| INT8 dynamic PTQ | 8.816e-01 | 9.173e-01 | no β serve unpadded only |
| INT8 static (this model) | 0.000e+00 | 0.000e+00 | yes |
pad-sensitivity = max change in P(sarcastic) for the same text alone vs padded, at batch size 1. mixed-batch = same text alone vs inside a mixed-length batch.
Dynamic PTQ recomputes activation scales from each tensor's observed range at inference time. Padded positions enter that range, so they change the scale, so they change how the real tokens are quantized β the attention mask stops padded positions contributing to attention, but not to the quantization range. Static quantization removes the dependence entirely.
Results
SarcOjiTest1
| variant | accuracy | precision | recall | F1 | MCC | ROC-AUC |
|---|---|---|---|---|---|---|
| fp32 (original) | 0.6605 | 0.7027 | 0.6143 | 0.6555 | 0.3266 | 0.7135 |
| INT8 static (this model) | 0.6507 | 0.6842 | 0.6235 | 0.6524 | 0.3041 | 0.7040 |
SarcOjiTest2
| variant | accuracy | precision | recall | F1 | MCC | ROC-AUC |
|---|---|---|---|---|---|---|
| fp32 (original) | 0.7124 | 0.4488 | 0.6635 | 0.5355 | 0.3518 | 0.7419 |
| INT8 static (this model) | 0.6990 | 0.4328 | 0.6603 | 0.5229 | 0.3317 | 0.7417 |
SarcOjiTest2 is ~75/25 negative-skewed; MCC and ROC-AUC are the meaningful columns there.
Performance (x86 CPU, single thread, batch-of-1)
| variant | p50 (ms) | p95 (ms) | size (MB) |
|---|---|---|---|
| fp32 | 67.95 | 90.32 | 437.6 |
| INT8 static | 29.04 | 52.28 | 183.5 |
Not an on-device ARM measurement.
Quantization recipe
- Static PTQ, QDQ format, per-channel weights
- Calibration:
percentileon 128 in-domain SarcOji samples, padded as deployed - Ops quantized: MatMul and Gemm only (where BERT's compute is; leaving Add/LayerNorm/Softmax in float preserves accuracy at negligible speed cost)
- Exported with eager attention, dynamic batch and sequence axes, opset 17
Usage
from optimum.onnxruntime import ORTModelForSequenceClassification
from transformers import AutoTokenizer
tok = AutoTokenizer.from_pretrained("Vandita/QuantizedSarcBERT-static-v2")
model = ORTModelForSequenceClassification.from_pretrained(
"Vandita/QuantizedSarcBERT-static-v2", file_name="model_int8_static.onnx")
inputs = tok(["Oh great, another meeting."], return_tensors="pt", padding=True)
logits = model(**inputs).logits # batching is safe on this build
- Downloads last month
- -
Model tree for Vandita/QuantizedSarcBERT-static-v2
Base model
google-bert/bert-base-cased