KvaytG's picture
Update README.md
ac1ed67 verified
|
Raw
History Blame Contribute Delete
3.23 kB
metadata
language:
  - ru
license: apache-2.0
tags:
  - text-classification
  - toxicity
  - russian
  - rubert-tiny2
  - pytorch
  - transformers
metrics:
  - f1
  - precision
  - recall
  - accuracy
model-index:
  - name: rubert-tiny2-toxic-detection
    results:
      - task:
          type: text-classification
          name: Text Classification
        metrics:
          - type: f1
            value: 0.9535
            name: Toxic F1
          - type: precision
            value: 0.9647
            name: Toxic Precision
          - type: recall
            value: 0.9425
            name: Toxic Recall

rubert-tiny2-toxic-detector

A lightweight and high-performance model for detecting toxic comments and messages in Russian. Fine-tuned on top of cointegrated/rubert-tiny2 using threshold optimization.

Test Set Metrics

Evaluated on an independent test set of 102,308 samples (80/10/10 split).

  • Optimal Threshold: 0.70 (determined on the validation set).
Class Precision Recall F1-Score Support
Normal (0) 0.9997 0.9998 0.9998 101,786
Toxic (1) 0.9647 0.9425 0.9535 522
Accuracy 0.9995 102,308
Macro Avg 0.9822 0.9712 0.9766 102,308
Weighted Avg 0.9995 0.9995 0.9995 102,308

Quick Start / Usage

import torch
from transformers import AutoTokenizer, AutoModelForSequenceClassification

MODEL_NAME = "KvaytG/rubert-tiny2-toxic-detector"
THRESHOLD = 0.70

tokenizer = AutoTokenizer.from_pretrained(MODEL_NAME)
model = AutoModelForSequenceClassification.from_pretrained(MODEL_NAME)
model.eval()

texts = [
    "Привет! Как твои дела?"
]

inputs = tokenizer(texts, padding=True, truncation=True, max_length=128, return_tensors="pt")

with torch.no_grad():
    outputs = model(**inputs)
    probs = torch.sigmoid(outputs.logits.squeeze(-1)).cpu().numpy()

for text, prob in zip(texts, probs):
    is_toxic = bool(prob > THRESHOLD)
    print(f"Text: '{text}'")
    print(f"  Toxicity probability: {prob:.4f} | Is toxic: {is_toxic}\n")

Training Details & Hyperparameters

  • Base Model: cointegrated/rubert-tiny2
  • Dataset Size: 1,023,075 rows (80% Train / 10% Val / 10% Test)
  • Loss Function: Binary Focal Loss (alpha=0.75, gamma=2.0)
  • Optimizer: AdamW (learning_rate=3e-5, weight_decay=0.01)
  • LR Scheduler: Linear Schedule with Warmup (10% warmup steps)
  • Batch Size: 64
  • Max Sequence Length: 128
  • Epochs: 3
  • Mixed Precision: PyTorch AMP (Automatic Mixed Precision)

License

This model is released under the Apache License 2.0.

Citation

@misc{kvaytg_rubert_tiny2_toxic_detector,
  author       = {KvaytG},
  title        = {RuBERT-tiny2 Toxic Text Detector},
  year         = {2026},
  publisher    = {Hugging Face},
  journal      = {Hugging Face Models},
  url          = {https://huggingface.co/KvaytG/rubert-tiny2-toxic-detector},
  note         = {High-performance lightweight toxic text detector for Russian language}
}