rasbt's picture
Add files using upload-large-folder tool
66e3018 verified
|
Raw
History Blame Contribute Delete
3.18 kB
metadata
language:
  - en
license: apache-2.0
library_name: transformers
pipeline_tag: text-classification
datasets:
  - rasbt/human-vs-ai-50k
base_model: distilbert/distilbert-base-uncased
base_model_relation: finetune
metrics:
  - accuracy
tags:
  - ai-text-detection
  - binary-classification

DistilBERT AI-Text Detector

This is a fully fine-tuned DistilBERT classifier for distinguishing human-written and AI-generated text. It was trained on rasbt/human-vs-ai-50k. Human-written text has label 0 and AI-generated text has label 1.

The model uses a maximum sequence length of 512 tokens. Temperature scaling is applied during inference. The recorded best validation accuracy was 99.74%.

 

Download and use

hf download rasbt/ai-text-detector-distilbert \
  --local-dir models/ai-text-detector-distilbert
import json
from pathlib import Path

import torch
from transformers import AutoModelForSequenceClassification, AutoTokenizer


model_dir = Path("models/ai-text-detector-distilbert")
metadata = json.loads(
    (model_dir / "detector-config.json").read_text(encoding="utf-8")
)
tokenizer = AutoTokenizer.from_pretrained(model_dir)
model = AutoModelForSequenceClassification.from_pretrained(model_dir)
model.eval()

text = "Paste the text to classify here."
inputs = tokenizer(
    text,
    truncation=True,
    max_length=metadata["max_length"],
    return_tensors="pt",
)

with torch.inference_mode():
    logits = model(**inputs).logits / metadata["temperature"]
    probabilities = logits.float().softmax(dim=-1)

ai_index = metadata["label_mapping"]["ai"]
ai_probability = probabilities[0, ai_index].item()
print({"score": round(100 * ai_probability, 4)})

 

Test-set confusion matrix

DistilBERT test-set confusion matrix

detector-config.json contains the calibration temperature and training metadata. The recommended inference implementation is provided in the rasbt/ai-detector repository.

 

Related models

 

Limitations

Performance may change for text from generators, domains, languages, and editing workflows not represented in the training set. Short or partly AI-assisted text may also be harder to classify. The score should not be treated as definitive evidence that a person did or did not write a text.