DeBERTa-ConPara

DeBERTa-ConPara detects machine-generated English text and stays accurate under the adversarial edits that defeat most detectors: homoglyph substitution, zero-width character insertion, whitespace and typographic attacks.

It is the system reported in our AACL-IJCNLP 2026 main-conference paper. The finding behind it: normalising the training corpus deduplicates it (35.4% of RAID rows collapse into byte-identical copies of their clean siblings, deleting the adversarial supervision), while normalising at inference is an effective defence. DeBERTa-ConPara trains on raw text and normalises only at inference.

  • Architecture: DeBERTa-v3-large β†’ CLS token β†’ Linear(1024, 512) β†’ GELU β†’ Dropout(0.1) β†’ Linear(512, 2). No feature branch.
  • Training data: 1.55M documents, leakage-free stratified splits over RAID, HC3 Plus, MAGE and M4, grouped by source id.
  • Output: logit margin, logit[AI] βˆ’ logit[human]. Higher is more machine-like. Not a probability.

Results

RAID hidden test (672,000 documents, 11 attacks): AUROC 99.61, TPR@5% FPR 99.01, TPR@1% FPR 96.57.

Cross-dataset, balanced accuracy at one threshold calibrated on our source validation split and held fixed: HC3-QA 99.69, HC3-SI 83.50, MAGE 96.23, M4 98.27.

The README of the GitHub repository compares DeBERTa-ConPara with every RAID leaderboard system that publishes a checkpoint, and states the two caveats those numbers need: MELD scores higher than DeBERTa-ConPara on RAID itself, and HC3/MAGE/M4 are training sources for DeBERTa-ConPara while being external data for the other systems.

Usage

The checkpoint is a plain PyTorch state dict with a small custom head, so it does not load through AutoModelForSequenceClassification. Use the loader from the repository:

from src.conpara import ConPara          # pip install -r requirements.txt

det = ConPara.from_pretrained()           # pulls rawguard.pt from this repo
det.score(["a document to check"])         # logit margin
det.predict(["a document to check"])       # bool at the stored threshold

Inference-time Unicode normalisation is applied by default; it is what makes the detector robust to homoglyph and zero-width attacks.

Intended use and limits

Intended as supporting evidence for a human decision: flagging text for review, studying detector behaviour, benchmarking. Not intended as a verdict, and not suitable for disciplinary or hiring decisions about individuals.

Measured limits:

  • Short text: below ~60 words errors are enriched 4–5Γ—. Treat under 75 words with caution; under 25 words the score is meaningless.
  • Academic prose: false-positive rates between 13% and 67% depending on the subcorpus. A flag on a student essay is not evidence of misconduct.
  • Calibration: the score is not a probability and saturates at the extremes. Prefer coarse bands (det.band()) over percentages.
  • Language: English only.
  • Drift: trained against generators available in 2026; newer ones are untested, and detectors decay as generators improve.
  • Attacks not covered: heavy paraphrasing by a strong model, and mixed human/machine documents, remain hard.

Training and evaluation protocol

Full corpus construction, the leakage-free splitting, the 2Γ—2Γ—2 factorial and the fixed-threshold evaluation protocol are in the repository. Every number above is reproducible from evaluation/.

Citation

@inproceedings{mady2026conpara,
  title         = {{DeBERTa-ConPara}: Attack-Aware and Deployment-Realistic
                   Detection of {AI}-Generated Text},
  author        = {Mady, Mohamed and Li, Yupei and Reschke, Johannes and Schuller, Bj\"orn W.},
  booktitle     = {Proceedings of the 14th International Joint Conference on Natural
                   Language Processing and the 4th Conference of the Asia-Pacific Chapter
                   of the Association for Computational Linguistics (AACL-IJCNLP 2026)},
  year          = {2026},
  publisher     = {Association for Computational Linguistics},
  eprint        = {2610.00883},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CL},
  url           = {https://arxiv.org/abs/2610.00883}
}

License: MIT, matching the DeBERTa-v3-large backbone.

Links

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for mohamedmady/deberta-conpara

Finetuned
(311)
this model

Space using mohamedmady/deberta-conpara 1

Paper for mohamedmady/deberta-conpara