Ar4ikov's picture
Add INT8, FP8 and INT4 builds
3a6eb64 verified
|
Raw
History Blame Contribute Delete
4.07 kB
metadata
language: ru
license: apache-2.0
library_name: transformers
pipeline_tag: text-classification
base_model: Aniemore/rubert-tiny2-russian-emotion-detection
base_model_relation: quantized
datasets:
  - Aniemore/cedr-m7
tags:
  - text-classification
  - emotion-recognition
  - russian
  - multi-label-classification
  - quantized
  - compressed-tensors
  - int8
  - fp8
  - int4
metrics:
  - roc_auc
  - f1
  - accuracy
model-index:
  - name: rubert-tiny2-russian-emotion-detection-quantized
    results:
      - task:
          name: Emotion Recognition
          type: text-classification
        dataset:
          name: CEDR-m7 test (int8)
          type: Aniemore/cedr-m7
          args: ru
        metrics:
          - name: ROC AUC macro (int8)
            type: roc_auc
            value: 0.8696
          - name: Macro F1 (int8)
            type: f1
            value: 0.6013
      - task:
          name: Emotion Recognition
          type: text-classification
        dataset:
          name: CEDR-m7 test (fp8)
          type: Aniemore/cedr-m7
          args: ru
        metrics:
          - name: ROC AUC macro (fp8)
            type: roc_auc
            value: 0.8695
          - name: Macro F1 (fp8)
            type: f1
            value: 0.6014
      - task:
          name: Emotion Recognition
          type: text-classification
        dataset:
          name: CEDR-m7 test (int4)
          type: Aniemore/cedr-m7
          args: ru
        metrics:
          - name: ROC AUC macro (int4)
            type: roc_auc
            value: 0.8738
          - name: Macro F1 (int4)
            type: f1
            value: 0.5981

rubert-tiny2-russian-emotion-detection · quantized

Quantized builds of Aniemore/rubert-tiny2-russian-emotion-detection — multi-label emotion recognition for Russian text over seven classes: anger, disgust, enthusiasm, fear, happiness, neutral, sadness.

The weights here are the published original, quantized. They were not retrained and they are not a different model.

Variants

subfolder scheme weights ROC AUC (macro) macro-F1 WA UA
(original repo) fp32 111 MiB 0.8696 0.6015 0.7662 0.6050
int8 W8A16 107 MiB 0.8696 0.6013 0.7662 0.6050
fp8 W8A16-float 107 MiB 0.8695 0.6014 0.7646 0.6028
int4 W4A16_ASYM 106 MiB 0.8738 0.5981 0.7588 0.6041
Quality after quantization Weights on disk

ROC AUC is listed first because the head is multi-label: macro-F1 depends on the decision threshold, which is 0.5 here because that is what the head was trained under, while ROC AUC does not.

How much this actually saves

Only Linear layers are quantized. In a BERT classifier the embedding matrix is not one of them, and on the smaller models it is most of the checkpoint — so the saving here scales with the encoder rather than with the parameter count. The large model compresses well; rubert-tiny barely moves, and the table above says so rather than quoting a ratio from the layers that did shrink.

Usage

import torch
from transformers import AutoModelForSequenceClassification, AutoTokenizer

repo = "Aniemore/rubert-tiny2-russian-emotion-detection-quantized"
model = AutoModelForSequenceClassification.from_pretrained(
    repo, subfolder="int8").eval()          # or "fp8", "int4"
tok = AutoTokenizer.from_pretrained(repo, subfolder="int8")

x = tok("мне сегодня очень грустно", return_tensors="pt")
with torch.no_grad():
    # multi-label: sigmoid per class, not softmax over classes
    probs = model(**x).logits.sigmoid()[0]
print({model.config.id2label[i]: round(p.item(), 3) for i, p in enumerate(probs)})

Limitations

  • Weight-only, round-to-nearest, no calibration.
  • Scored on the CEDR-m7 test split only. CEDR is written text; performance on transcribed speech, which carries no punctuation and no casing, is not measured here.
  • Inherited from cointegrated/rubert-tiny2; the licence follows the base model.