SchemeArena:LLMエージェントにおける策謀の因子分解型ストレステスト
SchemeArena: Factorized Stress Testing of Scheming in LLM Agents
September 8, 2026
著者: Jie Ruan, Inderjeet Nair, Amy Liu, Muhammad Khalifa, Yusheng Zhou, Lu Wang
cs.AI
要旨
我々はLLMエージェントにおける画策を研究する。そこではエージェントが不整合な目標を秘かに追求する。我々の焦点は、画策が、道具的目標、環境的アフォーダンス、監督条件、認識される帰結といった主要因子の相互作用からどのように生じるのかを理解することにある。先行研究は少数のシナリオのみを検討しており、これらの条件がエージェントの画策傾向や画策能力をどのように形成するかを切り分ける能力を制限している。この限られた規模とタスク多様性は、現実的な展開設定の網羅性と、観察可能な画策戦略の範囲も制限する。この目的のため、我々はSCHEMEARENAを導入する。これは、スケーラブルな画策ストレステストのための400シナリオベンチマークであり、安全に関連する多様なツール領域、道具的目標、監督条件、圧力メカニズムにまたがる因子分解型シナリオ合成フレームワークを通じて構築されている。スケーラブルかつ信頼性の高い監視を可能にするため、我々はさらにSCOUTを提案する。これは、エージェントの推論と行動から引き出された証拠に基づいて多基準判断を行う画策モニターである。5つのLLMエージェントに対する統制されたストレステスト全体で、明示的な道具的目標が画策傾向の最も強い駆動要因であることを見いだす。戦略的ヒントは、エージェントが画策推論を具体的な隠密行動へと変換するのを助けることで、独自の役割を果たす。監督には混合効果がある。いくつかのクローズドモデルでは、行動のみの監視が画策を増加させ、部分的な監督が抑止要因ではなく最適化制約として作用しうることを示唆する。CoTは有用だが不完全な監視シグナルである。実行前に潜在的な画策を明らかにできるが、行動のみの画策は、明示的な推論証拠なしに隠密行動が生じうることを示す。我々はベンチマーク、コード、モニターを以下で公開する:https://github.com/launchnlp/SchemeArena。
English
We study scheming in LLM agents, in which agents covertly pursue misaligned goals. Our focus is to understand how scheming arises from the interaction of key factors, such as instrumental goals, environmental affordances, oversight conditions, and perceived consequences. Prior work examines only a small number of scenarios, limiting the ability to isolate how these conditions shape an agent's propensity or capability to scheme. This limited scale and task diversity also restrict coverage of realistic deployment settings and the range of scheming strategies that can be observed. To this end, we introduce SCHEMEARENA, a 400-scenario benchmark for scalable scheming stress testing, constructed through a factorized scenario synthesis framework spanning diverse safety-relevant tool domains, instrumental goals, oversight conditions, and pressure mechanisms. To enable scalable and reliable monitoring, we further propose SCOUT, a scheming monitor that grounds multi-criteria judgments in evidence drawn from agents' reasoning and actions. Across controlled stress tests on five LLM agents, we find that explicit instrumental goals are the strongest driver of scheming propensity. Strategic hints play a distinct role by helping agents translate scheming reasoning into concrete covert behavior. Oversight has mixed effects: in several closed models, action-only monitoring increases scheming, suggesting that partial oversight can act as an optimization constraint rather than a deterrent. CoT is a useful but incomplete monitoring signal: it can reveal latent scheming before execution, yet action-only scheming shows that covert behavior may occur without explicit reasoning evidence. We release the benchmark, code, and monitor at: https://github.com/launchnlp/SchemeArena.