WeClawArena: 인간 중심 에이전트 네트워크에서 교차 사용자 에이전트 협업과 보안을 위한 감사 가능한 샌드박스 및 벤치마크
WeClawArena: An Auditable Sandbox and Benchmark for Cross-User Agents Collaboration and Security in Human-Centered Agent Networks
August 4, 2026
저자: Prince Zizhuang Wang, Aojie Yuan, Haiyue Zhang, Xiyang Hu, Yue Zhao, Shuli Jiang
cs.AI
초록
최근 지속형 개인 에이전트 프레임워크의 발전으로 인간 중심 에이전트 네트워크가 현실적인 배포 대상이 되고 있다. 각 사용자는 사용자를 대신하여 행동하고 상태를 유지하며 사회적·업무적 관계를 통해 다른 에이전트와 통신하는 AI 에이전트의 서비스를 받을 수 있다. 이러한 네트워크에서 일상적인 도구 사용은 개인 작업 공간을 매개로 한 다자간 소유 에이전트 협업이 되며, 파일, 기록, 도구, 정책은 소유자 간에 직접 공개되지 않는다. 기존 에이전트 벤치마크는 도구 사용과 협업을 연구하지만, 현실적인 사용자 디지털 작업 공간을 갖춘 검증 가능한 교차 사용자 에이전트 협업을 위한 종단간 샌드박스를 제공하지 않으며, 유해한 행동이 인간 중심 에이전트 네트워크를 통해 어떻게 전파될 수 있는지도 테스트하지 않는다. 본 논문에서는 개인 작업 공간을 통한 다자간 소유 에이전트 협업을 위한 감사 가능한 벤치마크이자 런타임 샌드박스인 WeClawArena를 소개한다. WeClawArena는 개인 작업 공간이 작업 수행 도구이자 개인적 제약 조건으로 기능하는 협업적 도구 사용 작업을 대상으로 한다. 벤치마크는 6개의 교차 사용자 작업 도메인에 걸친 124개의 기본 작업을 포함하며, 기본 작업당 1개의 무해 대조 시나리오와 4개의 공격 벡터 변형으로 구성된 총 620개의 시나리오 변형으로 확장된다. 샌드박스는 피어 메시지, 도구 호출, 리소스 작업, 거버넌스 결정, 최종 작업 공간 상태를 기록한다. WeClawArena는 유용성과 공격 성공률을 별도로 보고하며, 제한된 런타임 증거로부터 공격 성공 여부를 감사하여 작업 붕괴, 개인정보 유출, 오염된 증거, 유효하지 않은 권한 경로에 대한 진단을 지원한다.
English
Recent advances in persistent personal-agent frameworks are making human-centered agent networks realistic deployment targets: each user can be served by an AI agent that acts on the user's behalf, maintains state, and communicates with other agents through social and task relations. In these networks, everyday tool use becomes multi-party owned-agent collaboration over personal workspaces, where files, records, tools, and policies are not directly visible across owners. Existing agent benchmarks study tool use and collaboration, but they do not provide an end-to-end sandbox for verifiable cross-user agent collaboration with realistic user digital workspaces or test how harmful actions can travel through the human-centered agent network. We introduce WeClawArena, an auditable benchmark and runtime sandbox for multi-party owned-agent collaboration over personal workspaces. WeClawArena targets collaborative tool-use tasks in which personal workspaces serve as both operational tools and personal constraints. The benchmark contains 124 base tasks across six cross-user task domains and expands them into 620 scenario variants, with one benign control and four attack-vector variants per base task. The sandbox records peer messages, tool calls, resource operations, governed decisions, and final workspace states. WeClawArena reports utility and attack success rate separately and audits attack success from bounded runtime evidence, supporting diagnosis of task breakdown, privacy leakage, poisoned evidence, and invalid authority paths.