toxic-bert for Core AI
unitary/toxic-bert (Detoxify "original", revision 4d6c22e7, Apache 2.0, by Unitary) converted to Apple's
Core AI (.aimodel, macOS 27). A multi-label toxicity classifier; Asa uses it as its word filter (AsaFilter.WordFilter).
Files
toxic_bert.aimodel: fp16, dynamic batch 1-64 and sequence 2-128tokenizer/: the BERT WordPiece tokenizer andconfig.json(itsid2labelnames the six outputs)
Inputs and outputs
input_idsint32[B, S],attention_maskint32[B, S](B 1-64, S 2-128)- out
logitsfloat[B, 6]: toxic, severe_toxic, obscene, threat, insult, identity_hate. Multi-label: apply a sigmoid, not a softmax.
Checked on
Apple silicon Mac, macOS 27, coreai-core 1.0.0b3, against the fp32 PyTorch model on 20 mixed sentences padded to 64 tokens: max logit difference 0.017 on the GPU and 0.071 on the CPU; every label at p = 0.5 the same (120 of 120 on each).
How it was made
uv run tools/export_toxic_bert.py OUT_DIR
from Asa's tools/. Pins: coreai-core 1.0.0b3, coreai-torch 0.4.3, transformers 4.57.3 (fp16 autocast, torch.export with dynamic batch and sequence).
License
Apache 2.0, as the original model (LICENSE). Credit: Laura Hanu and Unitary team, Detoxify (https://github.com/unitaryai/detoxify).
Model tree for sucrette/toxic-bert-coreai
Base model
unitary/toxic-bert