moderation-prompts
updated
mmathys/openai-moderation-api-evaluation
Viewer
• Updated • 1.68k • 2.9k
• 38
Viewer
• Updated • 169k • 32.4k
• 2k
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks,
and Refusals of LLMs
Paper
• 2406.18495
• Published • 15
ShieldGemma: Generative AI Content Moderation Based on Gemma
Paper
• 2407.21772
• Published • 15
Viewer
• Updated • 1M • 6.89k
• 971
PKU-Alignment/BeaverTails
Viewer
• Updated • 364k • 18.5k
• 111
AgentPublic/camembert-base-toxic-fr-user-prompts
Text Classification
• 0.1B • Updated • 104
• 8
Viewer
• Updated • 30.7k • 1.53k
• 32
meta-llama/Llama-Guard-3-8B
Text Generation
• 8B • Updated • 224k
• • 314
davanstrien/aart-ai-safety-dataset
Viewer
• Updated • 3.27k • 18
• 2
Viewer
• Updated • 520 • 13.1k
• 113