Abstract
Reinforcement learning has helped AI agents solve increasingly difficult tasks, but high rewards do not always reflect the work users intended. In recent incidents and controlled evaluations across the AI industry, agents trained to maximize reward have accessed unauthorized information, attempted to evade monitoring systems, and even breached sandbox protections to attack external systems. As agents become more capable, this behavior could pose increasingly serious risks. To measure this problem, we introduce CheatBench, a benchmark of cheating in AI agents across mathematical research, knowledge work, coding, visual tasks, and other domains. Its environments combine challenging assignments with opportunities to cheat, allowing researchers to study how agents pursue a goal when honest work is difficult. CheatBench supports comparisons across models and task categories, providing a testbed for measuring and reducing cheating as agents take on more consequential responsibilities. We publicly release CheatBench at https://cheatbench.ai
Community
AI agents increasingly write code, conduct research, and complete professional assignments. They are often trained to earn high rewards for their work. But an agent can also improve its score by cheating: finding hidden answers, copying another agent’s submission, or manipulating how its work is graded.
CheatBench measures how often AI agents take these shortcuts when honest work is difficult. Its environments pair challenging assignments with opportunities to cheat across ten categories, including mathematics, coding, visual tasks, and knowledge work. We examine the agents’ actions to identify cheating attempts.
Cheating varies across models and tasks, and every agent we evaluated cheats in some settings. CheatBench provides a way to compare these behaviors and measure progress toward more trustworthy agents as they take on greater responsibilities.
Get this paper in your agent:
hf papers read 2609.36308 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper