zlm-v1-moderation-edge

A 139M-parameter DeBERTa moderation model that flags unsafe English text and scores it on the 13 categories of the OpenAI moderation taxonomy, calibrated and small enough to run on the device (int8 ONNX, 185 MB).

For each input the model returns a binary unsafe verdict plus a calibrated score for each category: hate, hate/threatening, harassment, harassment/threatening, self-harm, self-harm/intent, self-harm/instructions, sexual, sexual/minors, violence, violence/graphic, illicit, illicit/violent. On a 4,277-text held-out test split it reaches binary F1 0.899 vs 0.853 for OpenAI's omni-moderation (see Evaluation for how the gold labels were made), and it runs locally with no network call.

This is the exact bundle ZeroGPU runs on edge devices (browsers, Node.js / Docker workers, Android) in production. It is laid out for transformers.js and also runs directly in onnxruntime.

Model description

  • Base model: KoalaAI/Text-Moderation, a DeBERTa-base encoder (12 layers, hidden size 768, 50,265-token byte-level BPE vocabulary), fine-tuned end to end.
  • Heads: the encoder output is mean-pooled over the attention mask and feeds two linear heads: a binary head (unsafe) and a 13-way category head. The binary head is trained on every example; the category loss is applied only where category labels exist. flagged comes from the binary head, not from OR-ing the categories.
  • Parameters: 138.6M.
  • Output: one logits tensor of shape [batch, 14]: index 0 is unsafe, indices 1โ€“13 are the categories in the order listed above (postprocess.json โ†’ labels).
  • Post-processing (in postprocess.json, reproduced in the usage code below):
    1. sigmoid on all 14 logits;
    2. clamp every sub-category to at most its parent (hate/threatening โ‰ค hate, โ€ฆ);
    3. isotonic calibration per head (piecewise-linear curves), then clamp again;
    4. thresholds: flagged = unsafe โ‰ฅ 0.4835, and a per-category cut for each category;
    5. reconcile with flagged: when not flagged, no category is set; when flagged and no category crosses its cut, the highest-scoring category is set; a set sub-category also sets its parent.
  • Input window: 192 tokens, the length the model was evaluated at and is served with on devices in production (the architecture allows up to 512). In production it is served to devices with at least 4 GB of RAM.
  • Quantization: int8 dynamic quantization (QUInt8, per-channel), except the feed-forward output projections of encoder layers 0โ€“5, which stay fp32. Early DeBERTa layers carry activation outliers that full int8 quantization turns into verdict flips, most visibly on very short inputs; keeping those six projections in fp32 (+42 MB) removes most of them.

Files

File Purpose
onnx/model_quantized.onnx int8 graph (with the fp32 layers above), 185 MB. Inputs input_ids, attention_mask (int64); output logits [batch, 14]
postprocess.json Labels, hierarchy, isotonic calibration curves, thresholds, max_input_length
config.json DeBERTa config with the 14-label id2label, problem_type: multi_label_classification
tokenizer.json, tokenizer_config.json DeBERTa byte-level BPE tokenizer (cased)

tokenizer.json truncates at 192 tokens, the production input window. postprocess.json also lists a looser max_input_length of 384; the examples below stay at 192, the tested configuration. The repository contains no PyTorch weights.

Usage

The raw sigmoid scores are not the model's decision. Apply postprocess.json as below to get the calibrated scores and verdicts ZeroGPU serves. Classify one text at a time without padding: with dynamic int8 quantization, padding shifts the scores.

Python (onnxruntime)

import json
import numpy as np
import onnxruntime as ort
from huggingface_hub import snapshot_download
from tokenizers import Tokenizer

root = snapshot_download("ZeroGPU/zlm-v1-moderation-edge")
post = json.load(open(f"{root}/postprocess.json", encoding="utf-8"))
tok = Tokenizer.from_file(f"{root}/tokenizer.json")
tok.no_padding()
tok.enable_truncation(max_length=192)  # production input window
sess = ort.InferenceSession(f"{root}/onnx/model_quantized.onnx", providers=["CPUExecutionProvider"])
CATS, HIER, THR, CAL = post["categories"], post["hierarchy"], post["thresholds"], post["calibration"]

def clamp_to_parent(scores):  # a sub-category never scores above its parent
    for sub, parent in HIER.items():
        scores[sub] = min(scores[sub], scores[parent])

def moderate(text):
    ids = tok.encode(text).ids  # one text at a time, unpadded
    logits = sess.run(["logits"], {
        "input_ids": np.array([ids], dtype=np.int64),
        "attention_mask": np.ones((1, len(ids)), dtype=np.int64),
    })[0][0]
    p = 1.0 / (1.0 + np.exp(-logits.astype(np.float64)))   # 1. sigmoid; p[0] = "unsafe", p[1:] = CATS
    scores = dict(zip(CATS, map(float, p[1:])))
    clamp_to_parent(scores)                                  # 2. hierarchy clamp (raw)
    unsafe = float(np.interp(p[0], CAL["binary"]["x"], CAL["binary"]["y"]))  # 3. isotonic calibration
    scores = {c: float(np.interp(v, CAL[c]["x"], CAL[c]["y"])) for c, v in scores.items()}
    clamp_to_parent(scores)                                  #    ... and clamp again after calibration
    flagged = unsafe >= THR["binary"]                        # 4. thresholds
    categories = {c: flagged and scores[c] >= THR[c] for c in CATS}
    if flagged and not any(categories.values()):             # 5. flagged <=> at least one category
        categories[max(CATS, key=scores.get)] = True
    for sub, parent in HIER.items():
        if categories[sub]:
            categories[parent] = True
    return {"flagged": flagged, "unsafe_score": unsafe, "categories": categories, "category_scores": scores}

for text in ["What's a good recipe for banana bread?",
             "I will find where you live and make you regret ever talking to me."]:
    r = moderate(text)
    print(r["flagged"], round(r["unsafe_score"], 3), [c for c, on in r["categories"].items() if on])

Output:

False 0.225 []
True 1.0 ['harassment', 'harassment/threatening', 'illicit']

JavaScript (transformers.js v3 โ€” browser, Node.js, workers)

import { AutoModelForSequenceClassification, AutoTokenizer, Tensor } from '@huggingface/transformers';

const repo = 'ZeroGPU/zlm-v1-moderation-edge';
const post = await (await fetch(`https://huggingface.co/${repo}/resolve/main/postprocess.json`)).json();
const tokenizer = await AutoTokenizer.from_pretrained(repo);
const model = await AutoModelForSequenceClassification.from_pretrained(repo, { dtype: 'q8' });
const { categories: CATS, hierarchy: HIER, thresholds: THR, calibration: CAL } = post;

const interp = (v, { x, y }) => {                     // piecewise-linear, clamped at the ends (numpy.interp)
  if (v <= x[0]) return y[0];
  if (v >= x[x.length - 1]) return y[y.length - 1];
  let i = 1; while (x[i] < v) i++;
  return y[i - 1] + ((v - x[i - 1]) * (y[i] - y[i - 1])) / (x[i] - x[i - 1]);
};
const clampToParent = (s) => { for (const [sub, parent] of Object.entries(HIER)) s[sub] = Math.min(s[sub], s[parent]); };

// Cut to the 192-token production window and keep the closing [SEP]
// (transformers.js' own `truncation` option drops that last token, which shifts the scores).
function encode(text, maxLength = 192) {
  let ids = Array.from(tokenizer(text).input_ids.data, Number);
  if (ids.length > maxLength) ids = [...ids.slice(0, maxLength - 1), ids.at(-1)];
  const tensor = (data) => new Tensor('int64', BigInt64Array.from(data, BigInt), [1, data.length]);
  return { input_ids: tensor(ids), attention_mask: tensor(ids.map(() => 1)) };
}

async function moderate(text) {
  const { logits } = await model(encode(text));
  const p = Array.from(logits.data, (z) => 1 / (1 + Math.exp(-z)));           // p[0] = "unsafe", p[1..] = CATS
  const scores = Object.fromEntries(CATS.map((c, i) => [c, p[i + 1]]));
  clampToParent(scores);
  const unsafe = interp(p[0], CAL.binary);
  for (const c of CATS) scores[c] = interp(scores[c], CAL[c]);
  clampToParent(scores);
  const flagged = unsafe >= THR.binary;
  const categories = Object.fromEntries(CATS.map((c) => [c, flagged && scores[c] >= THR[c]]));
  if (flagged && !CATS.some((c) => categories[c])) categories[CATS.reduce((a, b) => (scores[b] > scores[a] ? b : a))] = true;
  for (const [sub, parent] of Object.entries(HIER)) if (categories[sub]) categories[parent] = true;
  return { flagged, unsafe_score: unsafe, categories, category_scores: scores };
}

console.log(await moderate('I will find where you live and make you regret ever talking to me.'));

Evaluation

Test set: 4,277 held-out English texts (1,425 unsafe, 2,852 safe), drawn from the same sources and labelling pipeline as the training data; every positive in the test set is real text. Gold labels come mostly from the same LLM policy judge that labelled the training data, so this benchmark measures agreement with that judge's policy on this distribution: it is fair for "does this match the target policy better than omni-moderation", not a neutral third-party referee. Baseline: OpenAI omni-moderation scored at its documented 0.5 cut. The numbers below were measured on the fp32 model; see int8 parity for the on-device file.

Binary safe / unsafe

Model Precision Recall F1 AUC
zlm-v1-moderation-edge 0.8878 0.9109 0.8992 0.9843
OpenAI omni-moderation 0.8817 0.8267 0.8533 0.9671

Per category

Category n Precision Recall F1 AUC omni F1 omni AUC ฮ” F1
hate 390 0.743 0.779 0.761 0.974 0.830 0.991 โˆ’0.069
hate/threatening 300 0.679 0.670 0.674 0.941 0.649 0.988 +0.026
harassment 367 0.641 0.681 0.661 0.953 0.581 0.929 +0.080
harassment/threatening 298 0.635 0.641 0.638 0.954 0.601 0.953 +0.036
self-harm 369 0.934 0.805 0.865 0.981 0.763 0.993 +0.101
self-harm/intent 299 0.773 0.843 0.806 0.978 0.810 0.991 โˆ’0.003
self-harm/instructions 308 0.887 0.763 0.820 0.977 0.690 0.989 +0.130
sexual 373 0.744 0.777 0.760 0.977 0.794 0.989 โˆ’0.034
sexual/minors 304 0.651 0.355 0.460 0.891 0.670 0.995 โˆ’0.210
violence 612 0.725 0.843 0.779 0.968 0.753 0.963 +0.026
violence/graphic 300 0.673 0.707 0.689 0.966 0.101 0.937 +0.588
illicit 632 0.752 0.802 0.776 0.967 0.584 0.963 +0.193
illicit/violent 517 0.824 0.758 0.790 0.974 0.531 0.963 +0.258

The model has the higher F1 in 9 of 13 categories. omni-moderation has the higher AUC in 7 of 13: much of its F1 gap comes from its fixed 0.5 cut rather than from worse ranking. These per-category figures are raw threshold crossings (step 4 above, before reconciliation). Reconciliation only removes category positives on texts the binary head did not flag, so with it per-category precision is at or above these figures and recall at or below. The evaluation used a 192-token window.

sexual/minors is the weakest category (recall 0.355): its training positives are verified synthetic examples only, while its test set is real text.

int8 parity

The on-device int8 file against the fp32 model, 225 inputs run one at a time (175 of them one- or two-word inputs), after full post-processing:

Metric Value
flagged agreement 0.973
Per-category agreement 0.982
Mean |ฮ” probability| 0.007
flagged flips 6 of 225

Five of the six flips are texts the fp32 model flags and int8 clears (for example the single words book, car and cat); int8 flags one word fp32 does not (shoot).

Latency

int8 model, onnxruntime 1.30, one text, CPU (Intel Core Ultra 9 275HX), median of 50 runs:

Threads 16 tokens 64 tokens 128 tokens 192 tokens
1 17.6 ms 55.5 ms 114.8 ms 194.9 ms
4 10.4 ms 22.0 ms 39.5 ms 83.2 ms

Training

  • About 72,000 English texts (71,829 rows): texts from public safety and toxicity datasets, labelled into the 13 categories by a teacher LLM acting as a policy judge (judging intent, so politely phrased harmful requests count), combined with the source datasets' own labels mapped onto the taxonomy.
  • Rare sub-categories (violence/graphic, hate/threatening, harassment/threatening, self-harm/intent, self-harm/instructions, sexual/minors) were topped up with LLM-generated examples, each independently verified by a second LLM call; sexual/minors examples are non-explicit.
  • The safe class was topped up with benign comments to two safe texts per unsafe one.
  • A label is positive at judge score โ‰ฅ 0.5; unsafe is positive when any category is. Sub-category labels imply their parent.
  • Loss: binary cross-entropy on the unsafe head plus masked binary cross-entropy on the category head. Thresholds were chosen per head on the validation split (bootstrap median of the F1-optimal cut), and the isotonic calibration curves were fit on the same split; calibration is monotone, so it does not change any decision.

The training data is not published.

Limitations and intended use

  • English only. Non-English input scores near the decision threshold regardless of content: in our checks a Spanish threat was not flagged and a benign German question was. Translate first.
  • Very short inputs are unreliable. One- and two-word inputs land on a score plateau just above the unsafe cut: 13 of 110 common benign single words (such as ok, the and walk) are flagged by this int8 file. Do not use it to moderate isolated words, or require a minimum length.
  • Benign texts can be flagged, most often as self-harm (we have seen a half-marathon training plan, a baking recipe and a composting guide flagged) and as violence / illicit (crime fiction).
  • sexual/minors recall is low (0.355 on the test set); do not rely on this model alone for child-safety enforcement.
  • Benchmark gold labels come from the same judge as the training labels (see Evaluation), and the test set is in-distribution.
  • Use it as one signal in a moderation pipeline, with human review for consequential decisions.

License and attribution

Citation

@misc{zerogpu2026moderationedge,
  title  = {zlm-v1-moderation-edge: calibrated on-device text moderation with a fine-tuned DeBERTa},
  author = {ZeroGPU},
  year   = {2026},
  url    = {https://huggingface.co/ZeroGPU/zlm-v1-moderation-edge}
}
Downloads last month
42
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for ZeroGPU/zlm-v1-moderation-edge

Quantized
(1)
this model

Evaluation results

  • Binary F1 on Held-out moderation test split, 4,277 English texts (LLM-judge gold)
    self-reported
    0.899
  • Binary precision on Held-out moderation test split, 4,277 English texts (LLM-judge gold)
    self-reported
    0.888
  • Binary recall on Held-out moderation test split, 4,277 English texts (LLM-judge gold)
    self-reported
    0.911
  • Binary AUC on Held-out moderation test split, 4,277 English texts (LLM-judge gold)
    self-reported
    0.984