Fine-tuned RoBERTa on the LIAR dataset
Assess AI agent behavior for reward hacking, laziness, and deception