Text Classification
Transformers
Safetensors
English
distilbert
toxicity
content-moderation
text-embeddings-inference
Instructions to use wagesj45/toxic-comment-classifier with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use wagesj45/toxic-comment-classifier with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="wagesj45/toxic-comment-classifier")# Load model directly from transformers import AutoTokenizer, AutoModelForSequenceClassification tokenizer = AutoTokenizer.from_pretrained("wagesj45/toxic-comment-classifier") model = AutoModelForSequenceClassification.from_pretrained("wagesj45/toxic-comment-classifier", device_map="auto") - Notebooks
- Google Colab
- Kaggle
File size: 3,302 Bytes
a7d2df2 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 | ---
license: apache-2.0
library_name: transformers
pipeline_tag: text-classification
language:
- en
base_model: distilbert/distilbert-base-multilingual-cased
datasets:
- Heliosoph/Jigsaw-Toxic-Comments
tags:
- toxicity
- content-moderation
- text-classification
metrics:
- f1
- accuracy
---
# Toxic Comment Classifier
This is a binary text-classification model for identifying whether an English
comment is `toxic` or `not_toxic`. It is a fine-tuned version of
[`distilbert/distilbert-base-multilingual-cased`](https://huggingface.co/distilbert/distilbert-base-multilingual-cased).
## Intended use
This model is intended as a research baseline or as one signal in a broader
content-moderation workflow. Predictions should be reviewed in context and
should not be used as the sole basis for high-impact decisions about people.
## Usage
```python
from transformers import pipeline
classifier = pipeline(
"text-classification",
model="wagesj45/toxic-comment-classifier",
)
print(classifier("Thanks for sharing this helpful explanation."))
```
The output label is either `not_toxic` or `toxic`, together with a confidence
score. The score is not a calibrated probability; applications should choose
and validate their own decision threshold.
## Training
The model was trained with `transformers` on the
[`Heliosoph/Jigsaw-Toxic-Comments`](https://huggingface.co/datasets/Heliosoph/Jigsaw-Toxic-Comments)
dataset. The six source labels (`toxic`, `severe_toxic`, `obscene`, `threat`,
`insult`, and `identity_hate`) were combined into one binary label: a comment
was assigned `toxic` when any source label was positive.
Training used a deterministic, stratified 90/10 train/validation split with
seed 42, three epochs, a learning rate of `2e-5`, batch sizes of 16/32, weight
decay of `0.01`, and a maximum tokenized length of 256. The exported model is
the best validation checkpoint selected by F1, from training step 17,952.
## Evaluation
On the held-out validation split:
| Metric | Score |
| --- | ---: |
| Accuracy | 0.9672 |
| F1 | 0.8332 |
These results come from the training split and evaluation procedure described
above. They should not be interpreted as performance on arbitrary production
comments or on languages outside the evaluation data.
## Limitations and risks
- The training data contains noisy, subjective toxicity annotations and may
reflect historical or cultural biases.
- The model can produce false positives for profanity, reclaimed language,
discussions of abuse, identity terms, or strongly worded criticism.
- It can miss implicit, coded, contextual, or adversarially written abuse.
- The model was trained and evaluated on English comments; multilingual base
model support does not make this fine-tune multilingual.
- Thresholds, appeals, human review, and monitoring are recommended for real
moderation systems.
## Licensing and provenance
The fine-tuned model artifacts in this repository are released under the
[Apache License 2.0](LICENSE). The base model is also listed under Apache 2.0.
The training dataset is listed as CC0-1.0 by its Hugging Face repository. The
dataset card credits Jigsaw and Wikipedia as the underlying sources; see the
[dataset card](https://huggingface.co/datasets/Heliosoph/Jigsaw-Toxic-Comments)
for details.
|