Instructions to use Jour/ouroboros-detector with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Jour/ouroboros-detector with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="Jour/ouroboros-detector")# Load model directly from transformers import AutoTokenizer, AutoModel tokenizer = AutoTokenizer.from_pretrained("Jour/ouroboros-detector") model = AutoModel.from_pretrained("Jour/ouroboros-detector", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Ouroboros detector (v7)
A token-level detector of AI-written text. For a document it returns character-aligned segments labelled human, ai-assisted or ai-generated with confidences, a weighted AI fraction, a human / mixed / ai verdict, and a 4-way "was this humanized?" probability (human, ai-generated, humanized-ai, mixed-authorship).
Unofficial. This is an independent, open reimplementation of the method of the Pangram 4 technical report (arXiv:2607.27183). It is not affiliated with or endorsed by Pangram, and it is not their model, weights or data.
Built end to end by Claude Opus 5.5. The architecture implementation, data pipeline, training, evaluation, interpretability study, bias audit and this model card were written and run by Claude Opus 5.5 (Anthropic), working as an autonomous coding agent, at the direction of the author, who set the goals and reviewed the results.
- Code, training pipeline, dataset builder, demo: https://github.com/Jourdelune/ouroboros-detector
- Backbone: Qwen3-1.7B-Base with a LoRA adapter (r = 32) merged into the weights (bf16, 3.4 GB), plus four single-layer heads.
- Languages: evaluated on English and French; other languages appear in the training data but are not benchmarked.
Use
pip install "git+https://github.com/Jourdelune/ouroboros-detector.git"
from ouroboros.infer.predict import Predictor
detector = Predictor.from_pretrained("Jour/ouroboros-detector")
result = detector.predict(text) # texts of at least ~50 words
print(result.document_label, result.weighted_ai_fraction)
for seg in result.segments:
print(seg.start, seg.end, seg.label_name, seg.confidence)
print(detector.humanizer(text))
The model needs the ouroboros package: it implements Repeat2 inputs, the four heads, sliding windows, calibration and CRF decoding. AutoModel.from_pretrained alone loads only the backbone.
Architecture
One causal backbone fed the window twice (Repeat2, so every supervised token has seen the whole window) and four heads: segment (D→15, weighted AI fraction), tokenwise provenance (D→3), mixed-authorship (D→2), humanizer probe (D→4, stop-gradient). Windows of 512 tokens with stride 256, then calibration and a linear-chain CRF.
Evaluation
Held-out test split of the project's corpus: 12,000 documents (4,531 human, 6,821 AI, 648 mixed), split by source document. Operating point chosen on a separate calibration split for a 0.5 % false-positive target.
| metric | v7 |
|---|---|
| False-positive rate | 1.35 % |
| False-negative rate (AI or mixed missed) | 0.94 % |
| AUROC | 0.9915 |
| TPR @ 1 % FPR | 98.9 % |
| Token accuracy | 98.4 % |
| Token F1 — human / AI-assisted / AI-generated | 0.986 / 0.714 / 0.988 |
| Mixed-document accuracy | 77.2 % |
4-way humanizer head accuracy / recall on humanized-ai |
95.0 % / 56.8 % |
These numbers are in-distribution (same corpus families as training) and not comparable to the Pangram report's, which uses a different, proprietary corpus and backbone.
Known limitations and biases
- Higher false-positive rate on French human text (7.0 % vs 1.0 % in English), on very short text (< 80 words: 6.5 %), and on human text formatted like an assistant answer (bullets / bold / headers: 8.8 %).
- Format and encoding shortcuts: human text split into paragraphs is flagged 14.5 % of the time; inserting four Cyrillic look-alike letters into a human text makes it read as AI in 86 % of cases. Normalise confusable characters before inference.
- Not robust to adaptive adversaries: a rewriting loop with access to the detector's score made 70–88 % of a small held-out set (50 French + 50 English passages) read as human while passing automatic fidelity checks.
- Not supported: texts under ~50 words.
ai-assistedis the weakest class (F1 0.71). - Do not use a verdict as the sole basis for a decision about a person. AI-text detectors can wrongly flag non-native writers and formal prose.
What it looks at (interpretability)
Measured on held-out passages with 95 % intervals: a linear probe on the untrained Qwen3-1.7B-Base already separates classes at 95 %; fine-tuning raises it to 98.5 % and makes it language-independent (French→English transfer 99.5 % vs 91.5 %). The verdict is written in layers 18–27, mostly by two attention heads and the last MLP; sentence-final periods (2.4 % of tokens) carry 28 % of the evidence and paragraph breaks (0.8 %) 21 %; shuffling words inside sentences removes most of the signal while surface edits barely move it; the score is uncorrelated with a base language model's surprisal.
Training data
Public sources only: RAID, MAGE, COLING-2025 MGT, ai-text-detection-pile, dmitva/human_ai_generated_text, Cosmopedia, WildChat, UltraChat, OpenHermes-2.5, arena preference sets, Amazon/Yelp/IMDB reviews, arXiv abstracts, FineWeb-2 (French) and FineWeb-Edu, Aya, Wikipedia, French instruction sets and several public distillation sets from recent frontier models (≈ 1.09 M training documents including augmentations). The dataset builder is in the GitHub repository.
Licences. The weights are released under Apache-2.0 and the backbone is Apache-2.0, but the training data comes from many datasets with their own terms (some research-only or non-commercial). Check them before any commercial use. No dataset is redistributed here.
Files
model.safetensors, config.json (merged backbone) · heads.pt (the four heads) · calibrator.json (decoder calibration, 0.5 % FPR target) · ouroboros_config.yaml (inference settings) · tokenizer files.
Citation
@software{ouroboros_detector,
title = {Ouroboros: an open token-level AI-text detector},
author = {Jourdelune},
year = {2026},
note = {Unofficial reimplementation of the Pangram 4 technical report (arXiv:2607.27183); built end to end with Claude Opus 5.5},
url = {https://github.com/Jourdelune/ouroboros-detector}
}
- Downloads last month
- 20
Model tree for Jour/ouroboros-detector
Base model
Qwen/Qwen3-1.7B-Base