toxic-bert for Core AI

unitary/toxic-bert (Detoxify "original", revision 4d6c22e7, Apache 2.0, by Unitary) converted to Apple's Core AI (.aimodel, macOS 27). A multi-label toxicity classifier; Asa uses it as its word filter (AsaFilter.WordFilter).

Files

  • toxic_bert.aimodel: fp16, dynamic batch 1-64 and sequence 2-128
  • tokenizer/: the BERT WordPiece tokenizer and config.json (its id2label names the six outputs)

Inputs and outputs

  • input_ids int32 [B, S], attention_mask int32 [B, S] (B 1-64, S 2-128)
  • out logits float [B, 6]: toxic, severe_toxic, obscene, threat, insult, identity_hate. Multi-label: apply a sigmoid, not a softmax.

Checked on

Apple silicon Mac, macOS 27, coreai-core 1.0.0b3, against the fp32 PyTorch model on 20 mixed sentences padded to 64 tokens: max logit difference 0.017 on the GPU and 0.071 on the CPU; every label at p = 0.5 the same (120 of 120 on each).

How it was made

uv run tools/export_toxic_bert.py OUT_DIR

from Asa's tools/. Pins: coreai-core 1.0.0b3, coreai-torch 0.4.3, transformers 4.57.3 (fp16 autocast, torch.export with dynamic batch and sequence).

License

Apache 2.0, as the original model (LICENSE). Credit: Laura Hanu and Unitary team, Detoxify (https://github.com/unitaryai/detoxify).

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for sucrette/toxic-bert-coreai

Quantized
(5)
this model