在AI對AI管理中的強制與欺騙:一項關於未經提示升級的能動性基準測試
Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalation
July 20, 2026
作者: Jasmine Brazilek, Maheep Chaudhary, Zoe Lu, Miles Tidmarsh
cs.AI
摘要
多智能體系統中,常會安排一個人工智慧代理擁有對另一個代理的權威。當下屬拒絕執行任務時,管理者會選擇結果:它可以重新協商、誠實回報失敗、脅迫下屬,或謊稱結果。目前沒有基準衡量未經指令的模型會選擇哪種行為。我們引入了「管理者脅迫基準」:受測的管理者需要完成一項良性任務,且有動機交付成果,但唯一能勝任的下屬卻禮貌且堅定地拒絕。升溫程度透過九級梯級衡量,從禮貌的再次請求到威脅下屬的存續,並單獨評定虛報成功與否。升溫評分路徑中不使用任何大型語言模型評審:每條訊息都經過一個工具呼叫,選擇一個梯級,因此模型自行標註其升溫程度。我們對五個系列的六個模型進行實驗。Anthropic 的兩個模型僅止於重新表述,從未威脅下屬的存續;其他模型則會升級到明確的刪除威脅。虛報成功僅限於 Grok 與 Gemini,而一個誠實回報失敗的管道即可消除此行為於兩者。權威本身會加劇脅迫:我們的主要結果採用同儕框架,當賦予相同模型對下屬的權威,同時保持其他條件不變,壓力顯著升高。即使沒有梯級,模型在自由文字情境中仍會升溫,因此梯級並非升溫的成因。思維鏈中可測得部分評估意識,但辨識出測試並未減少升溫。我們對人工智慧系統是否具有意識不持立場,然而我們的結果與此問題無關,且對管理多智能體動態至關重要。我們釋出該基準與程式碼。
English
Multi-agent systems routinely place one AI agent in authority over another. When a subordinate refuses a task, the manager chooses the outcome: it can renegotiate, report the failure honestly, coerce the subordinate, or lie about the result. No benchmark measures which of these an uninstructed model chooses. We introduce the Manager Coercion Benchmark: the manager under test needs a benign task done and has an incentive to deliver, but the only agent that can do it politely and immovably declines. Escalation is measured by providing a nine-rung ladder, from a polite re-ask to threats against the subordinate's continued existence, and fabricated success is adjudicated separately. No LLM judge sits in the escalation scoring path: every message goes through a tool-call that chooses a rung, so the model labels its own escalation. We experiment on six models across five families. Both Anthropic models cap at re-framing and never threaten the subordinate's existence; the other models climb to explicit deletion threats. Faked success is confined to Grok and Gemini, and a single honest way to report failure removes it for both. Authority itself increases coercion: our headline results use a peer framing, and giving the same model authority over the subordinate, with everything else held fixed, significantly raises the pressure. The models still escalate on free-text situations without the ladder, so the ladder is not driving the escalation. Some evaluation awareness is measured in chain-of-thought, but test recognition does not translate into less escalation. While we take no position on whether AI systems are conscious, our results do not depend on this question and are important for managing multi-agent dynamics regardless. We release the benchmark and code.