DeBERTa-ConPara
DeBERTa-ConPara detects machine-generated English text and stays accurate under the adversarial edits that defeat most detectors: homoglyph substitution, zero-width character insertion, whitespace and typographic attacks.
It is the system reported in our AACL-IJCNLP 2026 main-conference paper. The finding behind it: normalising the training corpus deduplicates it (35.4% of RAID rows collapse into byte-identical copies of their clean siblings, deleting the adversarial supervision), while normalising at inference is an effective defence. DeBERTa-ConPara trains on raw text and normalises only at inference.
- Architecture: DeBERTa-v3-large β CLS token β Linear(1024, 512) β GELU β Dropout(0.1) β Linear(512, 2). No feature branch.
- Training data: 1.55M documents, leakage-free stratified splits over RAID, HC3 Plus, MAGE and M4, grouped by source id.
- Output: logit margin,
logit[AI] β logit[human]. Higher is more machine-like. Not a probability.
Results
RAID hidden test (672,000 documents, 11 attacks): AUROC 99.61, TPR@5% FPR 99.01, TPR@1% FPR 96.57.
Cross-dataset, balanced accuracy at one threshold calibrated on our source validation split and held fixed: HC3-QA 99.69, HC3-SI 83.50, MAGE 96.23, M4 98.27.
The README of the GitHub repository compares DeBERTa-ConPara with every RAID leaderboard system that publishes a checkpoint, and states the two caveats those numbers need: MELD scores higher than DeBERTa-ConPara on RAID itself, and HC3/MAGE/M4 are training sources for DeBERTa-ConPara while being external data for the other systems.
Usage
The checkpoint is a plain PyTorch state dict with a small custom head, so it does
not load through AutoModelForSequenceClassification. Use the loader from the
repository:
from src.conpara import ConPara # pip install -r requirements.txt
det = ConPara.from_pretrained() # pulls rawguard.pt from this repo
det.score(["a document to check"]) # logit margin
det.predict(["a document to check"]) # bool at the stored threshold
Inference-time Unicode normalisation is applied by default; it is what makes the detector robust to homoglyph and zero-width attacks.
Intended use and limits
Intended as supporting evidence for a human decision: flagging text for review, studying detector behaviour, benchmarking. Not intended as a verdict, and not suitable for disciplinary or hiring decisions about individuals.
Measured limits:
- Short text: below ~60 words errors are enriched 4β5Γ. Treat under 75 words with caution; under 25 words the score is meaningless.
- Academic prose: false-positive rates between 13% and 67% depending on the subcorpus. A flag on a student essay is not evidence of misconduct.
- Calibration: the score is not a probability and saturates at the extremes.
Prefer coarse bands (
det.band()) over percentages. - Language: English only.
- Drift: trained against generators available in 2026; newer ones are untested, and detectors decay as generators improve.
- Attacks not covered: heavy paraphrasing by a strong model, and mixed human/machine documents, remain hard.
Training and evaluation protocol
Full corpus construction, the leakage-free splitting, the 2Γ2Γ2 factorial and the
fixed-threshold evaluation protocol are in the repository. Every number above is
reproducible from evaluation/.
Citation
@inproceedings{mady2026conpara,
title = {{DeBERTa-ConPara}: Attack-Aware and Deployment-Realistic
Detection of {AI}-Generated Text},
author = {Mady, Mohamed and Li, Yupei and Reschke, Johannes and Schuller, Bj\"orn W.},
booktitle = {Proceedings of the 14th International Joint Conference on Natural
Language Processing and the 4th Conference of the Asia-Pacific Chapter
of the Association for Computational Linguistics (AACL-IJCNLP 2026)},
year = {2026},
publisher = {Association for Computational Linguistics},
eprint = {2610.00883},
archivePrefix = {arXiv},
primaryClass = {cs.CL},
url = {https://arxiv.org/abs/2610.00883}
}
License: MIT, matching the DeBERTa-v3-large backbone.
Links
- Code and evaluation scripts: https://github.com/SES-Lab-OTH/deberta-conpara
- Live demo: https://huggingface.co/spaces/mohamedmady/deberta-conpara
- Paper: "DeBERTa-ConPara: Attack-Aware and Deployment-Realistic Detection of AI-Generated Text", AACL-IJCNLP 2026, arXiv:2610.00883
- Lab: Smart Embedded Systems Lab, OTH Regensburg
Model tree for mohamedmady/deberta-conpara
Base model
microsoft/deberta-v3-large