ChatPaper.aiChatPaper

WeClawArena:人間中心のエージェントネットワークにおけるクロスユーザーエージェントの協調とセキュリティのための監査可能なサンドボックスおよびベンチマーク

WeClawArena: An Auditable Sandbox and Benchmark for Cross-User Agents Collaboration and Security in Human-Centered Agent Networks

August 4, 2026
著者: Prince Zizhuang Wang, Aojie Yuan, Haiyue Zhang, Xiyang Hu, Yue Zhao, Shuli Jiang
cs.AI

要旨

近年の永続的なパーソナルエージェントフレームワークの進歩により、人間中心のエージェントネットワークは現実的な導入目標となりつつある。各ユーザーは、ユーザーに代わって行動し、状態を保持し、社会的・タスク関係を通じて他のエージェントと通信するAIエージェントの支援を受けることができる。このようなネットワークでは、日常的なツール利用は、個人ワークスペース上でのマルチパーティ所有エージェントコラボレーションとなり、ファイル、記録、ツール、ポリシーは所有者間で直接は見えない。既存のエージェントベンチマークはツール利用とコラボレーションを研究しているが、現実的なユーザーデジタルワークスペースを備えた検証可能なクロスユーザーエージェントコラボレーションのためのエンドツーエンドのサンドボックスを提供しておらず、また、有害な行動が人間中心のエージェントネットワークを通じてどのように伝播するかをテストしていない。我々は、個人ワークスペース上でのマルチパーティ所有エージェントコラボレーションのための監査可能なベンチマークおよびランタイムサンドボックスであるWeClawArenaを紹介する。WeClawArenaは、個人ワークスペースが運用ツールと個人の制約の両方として機能する協調的ツール利用タスクを対象とする。このベンチマークは、6つのクロスユーザータスクドメインにわたる124の基本タスクを含み、基本タスクごとに1つの良性コントロールと4つの攻撃ベクトルバリアントを備え、それらを620のシナリオバリアントに拡張する。サンドボックスは、ピアメッセージ、ツール呼び出し、リソース操作、ガバナンスによる決定、および最終ワークスペース状態を記録する。WeClawArenaは、ユーティリティと攻撃成功率を別々に報告し、有界な実行時証拠から攻撃成功を監査することで、タスク破綻、プライバシー漏洩、汚染された証拠、無効な権限パスの診断を支援する。
English
Recent advances in persistent personal-agent frameworks are making human-centered agent networks realistic deployment targets: each user can be served by an AI agent that acts on the user's behalf, maintains state, and communicates with other agents through social and task relations. In these networks, everyday tool use becomes multi-party owned-agent collaboration over personal workspaces, where files, records, tools, and policies are not directly visible across owners. Existing agent benchmarks study tool use and collaboration, but they do not provide an end-to-end sandbox for verifiable cross-user agent collaboration with realistic user digital workspaces or test how harmful actions can travel through the human-centered agent network. We introduce WeClawArena, an auditable benchmark and runtime sandbox for multi-party owned-agent collaboration over personal workspaces. WeClawArena targets collaborative tool-use tasks in which personal workspaces serve as both operational tools and personal constraints. The benchmark contains 124 base tasks across six cross-user task domains and expands them into 620 scenario variants, with one benign control and four attack-vector variants per base task. The sandbox records peer messages, tool calls, resource operations, governed decisions, and final workspace states. WeClawArena reports utility and attack success rate separately and audits attack success from bounded runtime evidence, supporting diagnosis of task breakdown, privacy leakage, poisoned evidence, and invalid authority paths.