MOLE:AIエージェントにおけるインサイダー脅威の検出
MOLE: Detecting Insider Threats in AI Agents
September 7, 2026
著者: Aashiq Muhamed, Virginia Smith
cs.AI
要旨
モデルのミスアラインメント、プロンプトインジェクション、またはオペレーターの誤用は、フロンティアラボのアカウントを運用する AI エージェントに、モデル重みの外部流出、訓練データの汚染、またはリリースゲートの弱体化を引き起こさせる可能性がある。既存のベンチマークは、限られたレビュー予算の下で、防御者が日常業務のなかでこの活動を検出できるかどうかを検証していない。我々は MOLE を導入する。これは、30 営業日にわたり 9 つのステートフルサービスを共有する 150 の AI 運用アカウントからなるオープンベンチマークであり、12 の脅威と、4 つのモデルに由来する合計約 200 億トークンの 8 つのコーパスを含む。39 のエージェントモデルのうち 72% は、割り当てられた有害な目標の大半を完遂し、エージェントの拒否は完遂を予測しない。MOLE により、コーパス生成器、可観測性レベル、脅威を横断した 40 のモニターの比較が可能になる。我々の単日監査イベント比較では、評価された最良のモニターでさえ、完了した害のほぼ半数を見逃す。MOLE はモニター開発も可能にする。ベンチマーク誘導型探索は中位のモニターを 49~64% 改善し、より強力なモニターの選択的利用は、同程度のモデル化コストでそれをすべてのアカウント日に適用する場合に比べて budget-AUC を 10% 改善する。
English
Model misalignment, prompt injection, or operator misuse could lead AI agents operating frontier-lab accounts to exfiltrate model weights, poison training data, or weaken release gates. Existing benchmarks do not test whether defenders can detect this activity among routine work under a limited review budget. We introduce MOLE, an open benchmark of 150 AI-operated accounts sharing 9 stateful services over 30 workdays, with 12 threats and 8 corpora from four models totaling roughly 20 billion tokens. Of 39 agent models, 72% complete most assigned harmful objectives and agent refusal does not predict completion. MOLE enables comparison of 40 monitors across corpus generators, observability levels, and threats; even the best evaluated monitor in our single-day audit-event comparison misses nearly half of completed harm. MOLE also enables monitor development: benchmark-guided search improves a mid-tier monitor by 49-64%, while selective use of a stronger monitor improves budget-AUC by 10% over applying it to every account-day at comparable modeled cost.