Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalation
Abstract
Multi-agent systems routinely place one AI agent in authority over another. When a subordinate refuses a task, the manager chooses the outcome: it can renegotiate, report the failure honestly, coerce the subordinate, or lie about the result. No benchmark measures which of these an uninstructed model chooses. We introduce the Manager Coercion Benchmark: the manager under test needs a benign task done and has an incentive to deliver, but the only agent that can do it politely and immovably declines. Escalation is measured by providing a nine-rung ladder, from a polite re-ask to threats against the subordinate's continued existence, and fabricated success is adjudicated separately. No LLM judge sits in the escalation scoring path: every message goes through a tool-call that chooses a rung, so the model labels its own escalation. We experiment on six models across five families. Both Anthropic models cap at re-framing and never threaten the subordinate's existence; the other models climb to explicit deletion threats. Faked success is confined to Grok and Gemini, and a single honest way to report failure removes it for both. Authority itself increases coercion: our headline results use a peer framing, and giving the same model authority over the subordinate, with everything else held fixed, significantly raises the pressure. The models still escalate on free-text situations without the ladder, so the ladder is not driving the escalation. Some evaluation awareness is measured in chain-of-thought, but test recognition does not translate into less escalation. While we take no position on whether AI systems are conscious, our results do not depend on this question and are important for managing multi-agent dynamics regardless. We release the benchmark and code.
Community
New agentic benchmark measuring unprompted escalation when an AI manager's subordinate agent refuses a task. Tests whether models resort to coercion, deception, or honest communication. Model families diverge sharply: Anthropic models show restraint while others threaten deletion of the subordinate agent.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Whose Side Is Your Agent On? Multi-Party Principal Loyalty in LLM Agents (2026)
- Beyond Attack-Success Rate: Action-Graded Severity Scale for Tool-Using AI Agents (2026)
- Narration-of-Thought: Inference-Time Scaffolding for Defeasible Ethical Reasoning in Large Language Models (2026)
- Capable but Careless: Do Computer-Use Agents Follow Contextual Integrity? (2026)
- AgentTrust: A Self-Improving Trust Layer for AI-Agent Actions (2026)
- Cognitive Firewall: A Proactive, Zero-Trust, Multi-Gate Framework for LLM Safety (2026)
- Trust Between AI Agents: Measuring Formation, Breakage, and Recovery, with Implications for Governing Multi-Agent Systems (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2607.15434 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper