File size: 3,887 Bytes
946a240
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
---
language:
- en
license: apache-2.0
library_name: transformers
pipeline_tag: text-classification
datasets:
- rasbt/human-vs-ai-50k
base_model: openai-community/gpt2
base_model_relation: finetune
metrics:
- accuracy
tags:
- ai-text-detection
- binary-classification
- variable-position-readout
---

# GPT-2 Variable-Position AI-Text Detector

This is a fully fine-tuned GPT-2 classifier for distinguishing human-written and AI-generated text. It uses a variable-position readout token immediately after the input text. The model was trained on [`rasbt/human-vs-ai-50k`](https://huggingface.co/datasets/rasbt/human-vs-ai-50k). Human-written text has label 0 and AI-generated text has label 1.

The maximum context length is 1,024 tokens. Temperature scaling is applied during inference. The recorded best validation accuracy was 97.44%.

 
## Download and use

```bash
hf download rasbt/ai-text-detector-gpt2-variable \
  --local-dir models/ai-text-detector-gpt2-variable
```

```python
import json
from pathlib import Path

import torch
from transformers import AutoModelForSequenceClassification, AutoTokenizer


model_dir = Path("models/ai-text-detector-gpt2-variable")
metadata = json.loads(
    (model_dir / "detector-config.json").read_text(encoding="utf-8")
)
tokenizer = AutoTokenizer.from_pretrained(model_dir)
model = AutoModelForSequenceClassification.from_pretrained(model_dir)
model.eval()

text = "Paste the text to classify here."
text_ids = tokenizer(
    text,
    add_special_tokens=False,
    truncation=True,
    max_length=metadata["max_text_length"],
)["input_ids"]

if metadata["readout_position"] == "fixed":
    padding_length = metadata["context_length"] - len(text_ids) - 1
    input_ids = (
        text_ids
        + [tokenizer.pad_token_id] * padding_length
        + [tokenizer.eos_token_id]
    )
    attention_mask = [1] * len(text_ids) + [0] * padding_length + [1]
else:
    input_ids = text_ids + [tokenizer.eos_token_id]
    attention_mask = [1] * len(input_ids)

inputs = {
    "input_ids": torch.tensor([input_ids]),
    "attention_mask": torch.tensor([attention_mask]),
}
with torch.inference_mode():
    logits = model(**inputs).logits / metadata["temperature"]
    probabilities = logits.float().softmax(dim=-1)

ai_index = metadata["label_mapping"]["ai"]
ai_probability = probabilities[0, ai_index].item()
print({"score": round(100 * ai_probability, 4)})
```

 
## Test-set confusion matrix

![GPT-2 variable-position test-set confusion matrix](figures/confusion-matrix.svg)

`detector-config.json` contains the readout, calibration, and training metadata. The recommended inference implementation is provided in the [`rasbt/ai-detector`](https://github.com/rasbt/ai-detector) repository because classification requires selecting the configured readout position.

 
## Related models

- [TF-IDF logistic regression](https://huggingface.co/rasbt/ai-text-detector-logreg)
- [DistilBERT](https://huggingface.co/rasbt/ai-text-detector-distilbert)
- [DistilBERT with LoRA](https://huggingface.co/rasbt/ai-text-detector-distilbert-lora)
- [DistilBERT with MiCA](https://huggingface.co/rasbt/ai-text-detector-distilbert-mica)
- [ModernBERT](https://huggingface.co/rasbt/ai-text-detector-modernbert)
- [GPT-2 with a fixed-position readout](https://huggingface.co/rasbt/ai-text-detector-gpt2-fixed)
- [Qwen3 0.6B with a fixed-position readout](https://huggingface.co/rasbt/ai-text-detector-qwen3-0.6b-fixed)
- [Qwen3 0.6B with a variable-position readout](https://huggingface.co/rasbt/ai-text-detector-qwen3-0.6b-variable)

 
## Limitations

Performance may change for text from generators, domains, languages, and editing workflows not represented in the training set. Short or partly AI-assisted text may also be harder to classify. The score should not be treated as definitive evidence that a person did or did not write a text.