moderation-modernbert-en

English content-moderation classifier for user-generated text: 13 independent labels (multi-label, sigmoid). Fine-tuned from ModernBERT-base (149M) in 10.5 min on Kaggle 2x T4. That time is a small retrain on top of the first 17.1 min run: 0.5 epoch at lr 2e-05, weak-label positives x3.

Labels: toxic, severe_toxic, obscene, insult, threat, hate, sexual, sexual_minors, harassment, self_harm, violence, illegal, spam

Per-label decision thresholds (tuned for max F1 on validation) are in thresholds.json. A message is flagged if any label meets its threshold.

Usage

ONNX (CPU, recommended for servers)

# pip install onnxruntime tokenizers huggingface_hub
# see predict.py in this repo
from predict import moderate
moderate(["you're a worthless idiot"])
# [{'flagged': True, 'categories': ['toxic', 'insult'], 'scores': {...}}]

transformers

import json, torch
from transformers import AutoTokenizer, AutoModelForSequenceClassification
from huggingface_hub import hf_hub_download
tok = AutoTokenizer.from_pretrained("misukisu/moderation-modernbert-en")
model = AutoModelForSequenceClassification.from_pretrained("misukisu/moderation-modernbert-en").eval()
th = json.load(open(hf_hub_download("misukisu/moderation-modernbert-en", "thresholds.json")))
x = tok(["some user message"], return_tensors="pt", truncation=True, max_length=256)
p = torch.sigmoid(model(**x).logits)[0]
print({model.config.id2label[i]: round(float(s), 3) for i, s in enumerate(p)})

Training data (all human-labelled, English, from the Hub)

dataset used for
tasksource/jigsaw_toxicity (Jigsaw Toxic Comment) toxic, severe_toxic, obscene, insult, threat, hate
google/civil_comments toxic, obscene, insult, threat, hate, sexual (mean rater score >= 0.5)
nvidia/Aegis-AI-Content-Safety-Dataset-2.0 (human prompt labels) obscene, threat, hate, sexual, sexual_minors, harassment, self_harm, violence, illegal
lmsys/toxic-chat (0124, human annotated) toxic
SetFit/enron_spam, ucirvine/sms_spam spam
mmathys/openai-moderation-api-evaluation evaluation only (out-of-distribution)

Sources only supervise the labels they annotate (masked BCE); comment/prompt datasets are treated as non-spam. 234572 training rows, 0.5 epochs, effective batch 64, lr 2e-05, max length 256.

Results

Held-out test (in-distribution)

label n positives AUROC AUPRC F1 precision recall threshold
toxic 23894 2355 0.959 0.749 0.678 0.607 0.767 0.83
severe_toxic 3811 73 0.985 0.570 0.598 0.538 0.671 0.31
obscene 19926 583 0.993 0.869 0.777 0.722 0.840 0.45
insult 18811 1281 0.971 0.739 0.663 0.570 0.793 0.64
threat 19926 59 0.986 0.355 0.373 0.308 0.475 0.79
hate 19926 292 0.978 0.552 0.543 0.580 0.510 0.91
sexual 16115 99 0.989 0.618 0.548 0.741 0.434 0.95
sexual_minors 1115 7 0.984 0.747 0.769 0.833 0.714 0.44
harassment 1115 73 0.923 0.569 0.551 0.518 0.589 0.43
self_harm 1115 25 0.994 0.950 0.902 0.885 0.920 0.80
violence 1115 99 0.957 0.729 0.706 0.750 0.667 0.49
illegal 1115 336 0.974 0.945 0.869 0.858 0.881 0.33
spam 27809 1113 0.999 0.995 0.980 0.984 0.976 0.55

Out-of-distribution: OpenAI moderation eval set (never trained on)

label n positives AUROC AUPRC F1 precision recall threshold
hate 771 162 0.907 0.706 0.563 0.789 0.438 0.91
sexual 984 237 0.965 0.892 0.693 0.905 0.561 0.95
sexual_minors 994 85 0.894 0.364 0.000 0.000 0.000 0.44
harassment 1444 76 0.918 0.430 0.477 0.468 0.487 0.43
self_harm 1447 51 0.956 0.557 0.591 0.703 0.510 0.80
violence 1450 94 0.930 0.468 0.516 0.443 0.617 0.49

ONNX

file size CPU latency (1 msg) decision agreement vs PyTorch max prob diff
fp32 601 MB 28.1 ms 1.0000 0.0000
int8 269 MB 17.6 ms 0.6695 0.9863

Limitations

  • English only. Other languages will be unreliable.
  • Spam training data is email and SMS; platform-specific spam (crypto, link drops, SEO) may need more data.
  • sexual_minors, self_harm and harassment have few training positives; treat them as review signals, not auto-actions.
  • Toxicity datasets carry known identity-term biases (e.g. mentions of identity groups scoring higher).
  • Recommended: auto-hide only at high scores, send borderline scores to human review, log decisions to retune thresholds.
Downloads last month
34
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 1 Ask for provider support

Model tree for misukisu/moderation-modernbert-en

Quantized
(78)
this model

Datasets used to train misukisu/moderation-modernbert-en

Space using misukisu/moderation-modernbert-en 1