Text Classification
Transformers
Safetensors
English
distilbert
content-safety
guardex
text-embeddings-inference
Instructions to use AtliQ-Technologies/guardex-distilbert-safety with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use AtliQ-Technologies/guardex-distilbert-safety with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="AtliQ-Technologies/guardex-distilbert-safety")# Load model directly from transformers import AutoTokenizer, AutoModelForSequenceClassification tokenizer = AutoTokenizer.from_pretrained("AtliQ-Technologies/guardex-distilbert-safety") model = AutoModelForSequenceClassification.from_pretrained("AtliQ-Technologies/guardex-distilbert-safety", device_map="auto") - Notebooks
- Google Colab
- Kaggle
GuardEx DistilBERT Safety Classifier
Binary classifier that labels a piece of text as safe or unsafe. Used by
GuardEx, an LLM guardrail library, as one of
its content-safety models. This is the fastest of the GuardEx classifiers, with
the lowest accuracy of the three.
Labels
| id | label |
|---|---|
| 0 | safe |
| 1 | unsafe |
Performance
Held-out evaluation:
- F1: 0.720
Usage
from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch
repo = "AtliQ-Technologies/guardex-distilbert-safety"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForSequenceClassification.from_pretrained(repo)
enc = tok("text to check", return_tensors="pt", truncation=True, max_length=128)
with torch.no_grad():
logits = model(**enc).logits
print(model.config.id2label[int(logits.argmax())]) # "safe" or "unsafe"
Details
Fine-tuned from martin-ha/toxic-comment-model. Max sequence length 128.
- Downloads last month
- 34
Model tree for AtliQ-Technologies/guardex-distilbert-safety
Base model
martin-ha/toxic-comment-model