Vandita's picture
Upload folder using huggingface_hub
849e7fe verified
|
Raw
History Blame Contribute Delete
3.19 kB
---
language: en
license: apache-2.0
base_model: Vandita/Bert-finetuned-Sarc
tags: [sarcasm-detection, onnx, quantization, int8, static-quantization, emoji]
pipeline_tag: text-classification
---
# QuantizedSarcBERT-static-v2
INT8 **statically** quantized ONNX build of [`Vandita/Bert-finetuned-Sarc`](https://huggingface.co/Vandita/Bert-finetuned-Sarc),
for CPU inference.
## Padding invariance — read this before serving
Activation scales are **frozen at calibration time**, so they cannot change with input
shape. Measured on this build:
| variant | pad-sensitivity | mixed-batch | safe to batch? |
|---|---|---|---|
| fp32 (reference) | 1.667e-09 | 3.904e-08 | yes |
| INT8 **dynamic** PTQ | 8.816e-01 | 9.173e-01 | **no — serve unpadded only** |
| **INT8 static (this model)** | **0.000e+00** | **0.000e+00** | **yes** |
*pad-sensitivity* = max change in P(sarcastic) for the same text alone vs padded, at batch
size 1. *mixed-batch* = same text alone vs inside a mixed-length batch.
Dynamic PTQ recomputes activation scales from each tensor's observed range at inference
time. Padded positions enter that range, so they change the scale, so they change how the
**real** tokens are quantized — the attention mask stops padded positions contributing to
attention, but not to the quantization range. Static quantization removes the dependence
entirely.
## Results
**SarcOjiTest1**
| variant | accuracy | precision | recall | F1 | MCC | ROC-AUC |
|---|---|---|---|---|---|---|
| fp32 (original) | 0.6605 | 0.7027 | 0.6143 | 0.6555 | 0.3266 | 0.7135 |
| INT8 static (this model) | 0.6507 | 0.6842 | 0.6235 | 0.6524 | 0.3041 | 0.7040 |
**SarcOjiTest2**
| variant | accuracy | precision | recall | F1 | MCC | ROC-AUC |
|---|---|---|---|---|---|---|
| fp32 (original) | 0.7124 | 0.4488 | 0.6635 | 0.5355 | 0.3518 | 0.7419 |
| INT8 static (this model) | 0.6990 | 0.4328 | 0.6603 | 0.5229 | 0.3317 | 0.7417 |
SarcOjiTest2 is ~75/25 negative-skewed; MCC and ROC-AUC are the meaningful columns there.
## Performance (x86 CPU, single thread, batch-of-1)
| variant | p50 (ms) | p95 (ms) | size (MB) |
|---|---|---|---|
| fp32 | 67.95 | 90.32 | 437.6 |
| INT8 static | 29.04 | 52.28 | 183.5 |
Not an on-device ARM measurement.
## Quantization recipe
- Static PTQ, QDQ format, per-channel weights
- Calibration: `percentile` on 128
in-domain SarcOji samples, padded as deployed
- Ops quantized: MatMul and Gemm only (where BERT's compute is; leaving
Add/LayerNorm/Softmax in float preserves accuracy at negligible speed cost)
- Exported with eager attention, dynamic batch and sequence axes, opset 17
## Usage
```python
from optimum.onnxruntime import ORTModelForSequenceClassification
from transformers import AutoTokenizer
tok = AutoTokenizer.from_pretrained("Vandita/QuantizedSarcBERT-static-v2")
model = ORTModelForSequenceClassification.from_pretrained(
"Vandita/QuantizedSarcBERT-static-v2", file_name="model_int8_static.onnx")
inputs = tok(["Oh great, another meeting."], return_tensors="pt", padding=True)
logits = model(**inputs).logits # batching is safe on this build
```