ChatPaper.aiChatPaper

OpenART: 개방형 환경 진화를 통한 에이전트 레드티밍의 확장

OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution

August 1, 2026
저자: Yunhao Chen, Xin Wang, Yixu Wang, Yi Liu, Jie Li, Yan Teng, Xingjun Ma, Xia Hu, Yu-Gang Jiang
cs.AI

초록

AI 에이전트는 초기 상태 변화가 먼 미래의 결정에 영향을 미칠 수 있는 지속적 환경에서 작동한다. 기존 언어 모델 상호작용과 달리, 에이전트의 행동은 장기 작업 흐름 전반에 걸쳐 반복적으로 수정되고 재사용되는 공유 상태를 통해 매개된다. 현재의 안전 벤치마크는 짧고 정적인 작업에 초점을 맞추기 때문에 이러한 누적 위험을 포착하지 못하는 경우가 많다. 이러한 한계를 해결하기 위해, 우리는 환경 진화를 통한 확장 가능한 에이전트 레드 팀 공격을 위한 개방형 아레나인 OpenART를 소개한다. OpenART는 50개 도메인에 걸쳐 10,000개 이상의 검증된 상태 기반 시나리오를 제공하며, 500,000개 이상의 도구와 스킬로 구성된 풀을 활용한다. 이러한 작업은 중앙값 97회의 도구 호출을 필요로 하며, 75가지 서로 다른 에이전트-모델 구성에 걸친 통합 평가를 가능하게 한다. 이러한 진화하는 공격 표면을 체계적으로 탐색하기 위해, 우리는 진화적 마르코프 하이퍼그래프 공격(EMHA)을 제안한다. EMHA는 파라미터 업데이트 없이 승인된 상태 전이를 조정하여 피드백 기반 환경 진화를 수행하는 블랙박스 정책이다. 평가 전반에 걸쳐 작업 목표는 고정된 상태로 유지되며 환경 상태만 변경된다. 모든 구성에서 EMHA는 통합 공격 성공률(ASR) 85.0%를 달성한다. 지시문 기반 진화에 대한 EMHA의 우위는 단순 환경에서 약 2%에서 가장 복잡한 환경에서 17% 이상으로 증가하며, 이는 작업 복잡도가 증가함에 따라 환경 진화가 안전 실패를 더욱 효과적으로 드러냄을 입증한다. 또한, 우리의 분석은 에이전트의 특정 런타임 구현이 기저 모델의 성능을 넘어 안전 변동의 상당 부분을 설명함을 보여준다. 이러한 결과는 OpenART를 복잡하고 진화하는 환경에서 에이전트 안전을 연구하기 위한 확장 가능한 기반으로 확립한다.
English
AI agents operate in persistent environments where early state changes can influence decisions far into the future. Unlike conventional language-model interactions, agent behavior is mediated through a shared state that is repeatedly modified and reused across long-horizon workflows. Current safety benchmarks often fail to capture these cumulative risks because they focus on short, static tasks. To address these limitations, we introduce OpenART, an open-ended arena for scalable agent red teaming through environment evolution. OpenART provides over 10,000 validated stateful scenarios across 50 domains, drawing from a pool of more than 500,000 tools and skills. These tasks require a median of 97 tool calls and enable unified evaluation across 75 different agent-model configurations. To systematically explore these evolving attack surfaces, we propose the Evolutionary Markov Hypergraph Attack (EMHA). EMHA is a black-box policy that performs feedback-driven environment evolution by coordinating authorized state transitions without requiring parameter updates. Throughout the evaluation, task objectives remain fixed while only the environment state changes. Across all configurations, EMHA achieves a pooled Attack Success Rate (ASR) of 85.0%. Its advantage over instruction-only evolution increases from approximately 2% on simple environments to over 17% on the most complex ones, demonstrating that environment evolution increasingly exposes safety failures as task complexity grows. Furthermore, our analysis shows that the specific runtime implementation of an agent explains a significant portion of safety variation beyond the underlying model's capabilities. These results establish OpenART as a scalable foundation for studying agent safety in complex, evolving environments.