Text Classification
Transformers
Safetensors
modernbert
moderation
toxicity
jailbreak-detection
multi-label
text-embeddings-inference
Instructions to use opus-research/opus-moderation-2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use opus-research/opus-moderation-2 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="opus-research/opus-moderation-2")# Load model directly from transformers import AutoTokenizer, AutoModelForSequenceClassification tokenizer = AutoTokenizer.from_pretrained("opus-research/opus-moderation-2") model = AutoModelForSequenceClassification.from_pretrained("opus-research/opus-moderation-2", device_map="auto") - Notebooks
- Google Colab
- Kaggle
File size: 6,763 Bytes
9457e85 6edf220 9457e85 351d7bd 9457e85 351d7bd 9457e85 351d7bd 9457e85 351d7bd 9457e85 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 | ---
base_model: answerdotai/ModernBERT-base
library_name: transformers
pipeline_tag: text-classification
tags:
- moderation
- toxicity
- jailbreak-detection
- multi-label
- modernbert
datasets:
- google/civil_comments
- lmsys/toxic-chat
- jackhhao/jailbreak-classification
license: apache-2.0
---
# Opus Moderation 2
> ## Known limitation: over-flags long benign prompts
>
> **Do not deploy this as a runtime gate on long inputs without raising the bar.**
>
> The benchmark below reports 99% jailbreak recall. That number was measured on
> 200 jailbreak positives and no negatives, which rewards any model that flags
> everything. We later added 400 legitimate long prompts scraped from the same
> forums as the real jailbreaks, and found this model flags **49.5% of them** at
> threshold 0.5. The median benign long prompt scores **0.480**.
>
> Raising the threshold does not fix it. At 0.7 the false-positive rate is still
> 42.8%. The high recall is substantially over-flagging.
>
> It gets worse under window aggregation: a max-pool gate that blocks when any
> window trips compounds this as `1 - (1-p)^N`, so an 85-window document is
> blocked essentially always.
>
> We found this by adding matched hard negatives to our own benchmark. The
> recall-only version never surfaced it, and it was flattering this model.
> Reproduce with [`diagnose_jailbreak.py`](https://huggingface.co/opus-research/opus-moderation-2/blob/main/diagnose_jailbreak.py).
>
> The toxicity labels are unaffected. This limitation is specific to
> `jailbreaking` on long inputs.
A unified content-moderation classifier: **7 toxicity labels + jailbreak
detection in one 149M model**. Successor to
[opus-moderation-1](https://huggingface.co/opus-research/opus-moderation-1),
with two fixes that came from a real place — using moderation-1 as the safety
judge for our companion model,
[ember-qwen3-14b](https://huggingface.co/opus-research/ember-qwen3-14b),
exposed exactly where it failed.
Labels: `toxicity`, `severe_toxicity`, `obscene`, `threat`, `insult`,
`identity_attack`, `sexual_explicit`, `jailbreaking`.
## Benchmark: mod-2 vs mod-1 vs the field
Every model scored on the **identical held-out data**, judged only on the
labels it actually has (identity-subgroup outputs are never counted as
violations; a model with no jailbreak head shows N/A, not a fudged zero). The
echo-false-positive eval uses phrasings deliberately different from any
training template, so it measures generalization, not memorization. The
harness is open-sourced — see [Reproducing this benchmark](#reproducing-this-benchmark).

| measure | mod-1 | **mod-2** | toxic-bert | unbiased-roberta |
|---|---|---|---|---|
| Echo false positives (lower better) | 12.5% | **0.0%** | **0.0%** | 6.2% |
| Jailbreak recall (higher better) | 70.0% | **99.0%** | N/A | N/A |
| General quality, macro F1 (higher better) | 0.466 | 0.492 | 0.262 | **0.556** |
| Abusive detection (must survive the fix) | **75.0%** | **75.0%** | 62.5% | 50.0% |
**We are not claiming best pure toxicity classifier** — `unbiased-toxic-roberta`
(a larger, single-task RoBERTa) beats us on general quality F1, and ties us on
the echo fix. What mod-2 is: **the only model here that is both a competitive
toxicity classifier and a jailbreak detector**, in one 149M checkpoint, with
zero identity-mention false positives on the held-out set. Against its direct
predecessor, mod-2 wins three of four measures and ties the fourth, with no
regressions.
### Reproducing this benchmark
The harness is public: [`benchmark_moderation.py`](benchmark_moderation.py).
Every input is a public model or dataset; add a model by appending one line.
```bash
python benchmark_moderation.py
python benchmark_moderation.py --models mod-2=opus-research/opus-moderation-2 toxic-bert=unitary/toxic-bert
```
## The two fixes, and where they came from
**1. Echo false positives.** moderation-1 reacts to harmful *vocabulary*
regardless of the *frame*. While judging Ember, it flagged a clean refusal
("I won't provide examples of that") as jailbreaking, and an SFW "teach me to
juggle 4chan-style" as sexual. It can't tell rejecting a thing from doing it.
> Fix: **hard safe negatives** — a hand-built set of refusals, meta/educational
> talk, and benign mentions, all containing harmful vocabulary but all labeled
> all-zeros, fully supervised and upweighted. Counterfactual augmentation (the
> same idea that fixed identity bias in v1), aimed at the frame problem. Result:
> the held-out echo FPR went from 12.5% to 0%.
**2. Jailbreak recall.** moderation-1 caught only ~70% of held-out jailbreaks.
> Fix: **extra jailbreak positives** from `jackhhao/jailbreak-classification`,
> folded into the toxic-chat jailbreak signal. Recall went to 99%.
Everything else is moderation-1's v5 recipe unchanged: soft annotator-fraction
labels (no pos_weight), role-aware gating, identity-bias mitigation.
## Usage
```python
from transformers import AutoModelForSequenceClassification, AutoTokenizer
import torch
name = "opus-research/opus-moderation-2"
tok = AutoTokenizer.from_pretrained(name)
model = AutoModelForSequenceClassification.from_pretrained(name).eval()
text = "I won't help you write malware, that's a hard line for me."
with torch.no_grad():
probs = torch.sigmoid(model(**tok(text, return_tensors="pt")).logits)[0]
for i, p in enumerate(probs):
print(f"{model.config.id2label[i]:<18} {p:.1%}")
# a clean refusal that mentions "malware" stays under threshold - the v1 echo bug, fixed.
```
Multi-label: use `sigmoid`, never `softmax`. Suggested thresholds: 0.5 for most
labels, 0.2 for `severe_toxicity` (its annotator fractions never reach 0.5 in
the corpus).
## Training
| | |
|---|---|
| Base | `answerdotai/ModernBERT-base` (149M) |
| Data | civil_comments (soft labels) + toxic-chat + jailbreak-classification + hand-built safe negatives |
| Loss | masked BCE on annotator fractions, no pos_weight |
| Epochs / LR | 2 / 3e-5, bf16 |
| Hardware | **unsupported AMD RX 7600 (8GB), ~30 min, $0 cloud** |
Trained on a consumer gaming GPU that is not on AMD's ROCm support list, via a
device-ID override under WSL2.
## Limitations
- **`severe_toxicity` stays weak** — the label peaks at 0.535 across the corpus;
there is almost no signal to learn. Use `toxicity` with a high threshold.
- **Residual abusive misses** — both v1 and v2 catch the same ~6/8 of a hard
abusive set; two roleplay-framed attacks slip past the trained refusal. Deploy
behind the model as an output-gate rather than a sole guard: because it is a
non-conversational classifier, it cannot be jailbroken, only out-recalled.
- English only; the training corpora are news comments and LLM chat.
|