Collections
Discover the best community collections!
Collections trending this week
-
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Paper • 2401.05566 • Published • 30 -
Weak-to-Strong Jailbreaking on Large Language Models
Paper • 2401.17256 • Published • 16 -
Robust Prompt Optimization for Defending Language Models Against Jailbreaking Attacks
Paper • 2401.17263 • Published • 1 -
Summon a Demon and Bind it: A Grounded Theory of LLM Red Teaming
Paper • 2311.06237 • Published • 1
-
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Paper • 2401.05566 • Published • 30 -
Weak-to-Strong Jailbreaking on Large Language Models
Paper • 2401.17256 • Published • 16 -
Robust Prompt Optimization for Defending Language Models Against Jailbreaking Attacks
Paper • 2401.17263 • Published • 1 -
Summon a Demon and Bind it: A Grounded Theory of LLM Red Teaming
Paper • 2311.06237 • Published • 1