MOLE: AI 에이전트의 내부자 위협 탐지
MOLE: Detecting Insider Threats in AI Agents
September 7, 2026
저자: Aashiq Muhamed, Virginia Smith
cs.AI
초록
모델 오정렬, 프롬프트 인젝션 또는 운영자 오용은 프런티어 랩 계정을 운영하는 AI 에이전트가 모델 가중치를 유출하거나 훈련 데이터를 오염시키거나 릴리스 게이트를 약화시키도록 이끌 수 있다. 기존 벤치마크는 제한된 검토 예산 아래에서 방어자가 일상 업무 중 이러한 활동을 탐지할 수 있는지 테스트하지 않는다. 우리는 30 근무일 동안 9개의 상태 유지 서비스를 공유하는 AI가 운영하는 150개 계정으로 구성된 공개 벤치마크인 MOLE를 소개한다. MOLE는 12개 위협과, 4개 모델에서 가져온 총 약 200억 토큰에 달하는 8개 코퍼스를 포함한다. 39개 에이전트 모델 중 72%가 할당된 유해 목표 대부분을 완료하며, 에이전트의 거부는 완료를 예측하지 못한다. MOLE는 코퍼스 생성기, 관측 가능성 수준, 위협 전반에 걸쳐 40개 모니터의 비교를 가능하게 한다. 우리의 단일 일자 감사 이벤트 비교에서 평가된 최고의 모니터조차 완료된 유해 행위의 거의 절반을 놓친다. MOLE는 또한 모니터 개발을 가능하게 한다: 벤치마크 기반 탐색은 중간급 모니터를 49-64% 향상시키며, 더 강력한 모니터를 선택적으로 사용하는 것은 유사한 모델링 비용에서 이를 모든 계정-일에 적용하는 것보다 budget-AUC를 10% 향상시킨다.
English
Model misalignment, prompt injection, or operator misuse could lead AI agents operating frontier-lab accounts to exfiltrate model weights, poison training data, or weaken release gates. Existing benchmarks do not test whether defenders can detect this activity among routine work under a limited review budget. We introduce MOLE, an open benchmark of 150 AI-operated accounts sharing 9 stateful services over 30 workdays, with 12 threats and 8 corpora from four models totaling roughly 20 billion tokens. Of 39 agent models, 72% complete most assigned harmful objectives and agent refusal does not predict completion. MOLE enables comparison of 40 monitors across corpus generators, observability levels, and threats; even the best evaluated monitor in our single-day audit-event comparison misses nearly half of completed harm. MOLE also enables monitor development: benchmark-guided search improves a mid-tier monitor by 49-64%, while selective use of a stronger monitor improves budget-AUC by 10% over applying it to every account-day at comparable modeled cost.