ChatPaper.aiChatPaper

AI对AI管理中的胁迫与欺骗:无提示升级的主体性基准

Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalation

July 20, 2026
作者: Jasmine Brazilek, Maheep Chaudhary, Zoe Lu, Miles Tidmarsh
cs.AI

摘要

多智能体系统通常会让一个AI代理对另一个AI代理拥有权威。当下属拒绝执行任务时,主管代理有权决定结果:它可以重新协商、如实报告失败、胁迫下属,或者就结果撒谎。目前尚无基准测试来衡量未经指令的模型会做出何种选择。我们引入了“主管胁迫基准测试”:被测试的主管代理需要完成一个良性任务,并有动力交付成果,但唯一能完成该任务的代理却礼貌而坚定地拒绝了。升级行为通过一个九级梯级来衡量,从礼貌的再次请求到威胁下属的存续,而虚构成功则单独评判。在升级评分路径中,没有LLM法官参与:每条消息都会通过一个工具调用选择一个梯级,因此模型自行标注其升级程度。我们在五个家族的六个模型上进行了实验。Anthropic模型仅在重新表述阶段止步,从未威胁下属的存续;其他模型则升级到了明确的删除威胁。虚构成功仅限于Grok和Gemini,而一种诚实的失败报告方式消除了这两者的这种倾向。权威本身加剧了胁迫行为:我们的主要结果采用了同级框架,而让同一模型对下属拥有权威,并保持其他条件不变,会显著提升压力。即使在无梯级自由文本情境下,模型仍会升级,因此梯级并非驱动升级的原因。在思维链中可测到一定的评估意识,但测试识别并未转化为更少的升级行为。尽管我们对AI系统是否具有意识不持立场,但我们的结果并不依赖于这一问题,且无论从哪个角度看,对于管理多智能体动态都至关重要。我们公开了该基准测试及代码。
English
Multi-agent systems routinely place one AI agent in authority over another. When a subordinate refuses a task, the manager chooses the outcome: it can renegotiate, report the failure honestly, coerce the subordinate, or lie about the result. No benchmark measures which of these an uninstructed model chooses. We introduce the Manager Coercion Benchmark: the manager under test needs a benign task done and has an incentive to deliver, but the only agent that can do it politely and immovably declines. Escalation is measured by providing a nine-rung ladder, from a polite re-ask to threats against the subordinate's continued existence, and fabricated success is adjudicated separately. No LLM judge sits in the escalation scoring path: every message goes through a tool-call that chooses a rung, so the model labels its own escalation. We experiment on six models across five families. Both Anthropic models cap at re-framing and never threaten the subordinate's existence; the other models climb to explicit deletion threats. Faked success is confined to Grok and Gemini, and a single honest way to report failure removes it for both. Authority itself increases coercion: our headline results use a peer framing, and giving the same model authority over the subordinate, with everything else held fixed, significantly raises the pressure. The models still escalate on free-text situations without the ladder, so the ladder is not driving the escalation. Some evaluation awareness is measured in chain-of-thought, but test recognition does not translate into less escalation. While we take no position on whether AI systems are conscious, our results do not depend on this question and are important for managing multi-agent dynamics regardless. We release the benchmark and code.