ChatPaper.aiChatPaper

エージェントが「あなた」になることを学ぶとき:ペルソナスキルにおけるプライバシー漏洩、なりすましリスク、および防御策のベンチマーキング

When Agents Learn to Be You: Benchmarking Privacy Leakage, Impersonation Risk, and Defenses in Persona Skills

August 4, 2026
著者: Yongli Xiang, Zhifang Zhang, Bojun Yang, Ziming Hong, Lei Feng, Miao Xu, Tongliang Liu
cs.AI

要旨

ペルソナスキルは、個人の対話履歴を、下流エージェント向けの移植可能かつ実行可能なアーティファクトへ蒸留する。柔軟なパーソナライゼーションを可能にする一方で、このプロセスは断片的な個人シグナルを集中させ、再利用を通じてその影響を増幅し、個別レコードや検索ベースのメモリ向けに設計された防御策に挑戦を課す。ペルソナスキルパイプラインの安全性を体系的に調査するため、我々はペルソナスキルパイプライン全体にわたるリスクと防御を評価するエンドツーエンドのベンチマークであるAntiSkillBenchを導入する。本ベンチマークは以下で構成される:(i)多様なタスクシナリオを網羅する50の行動的に豊かなプロファイルから構築された7,500件のペルソナ基盤対話トレースからなるデータセット、(ii)3つのスキル蒸留戦略にわたってスキルレベルのプライバシー漏洩、エージェントレベルの属性開示、および行動模倣を測定する評価スイート、(iii)能動的リスク抑制と受動的来歴保護を含む、オンライン介入と事後的介入にわたる4つの構成を網羅する防御評価。3つのフロンティアエージェントにわたる実験は、ペルソナスキルのリスクがエージェントのバックボーンと蒸留プロトコルを問わず持続し、明示的属性からコミュニケーションスタイルや性格特性にまで及ぶことを示す。既存の防御策は限定的かつ蒸留依存的な効果しか示さず、リスクと蒸留戦略をまたいだ一般化に失敗している。これらの結果は、AntiSkillBenchがプライバシー保護と真正性を考慮したペルソナスキルを開発するための挑戦的なベンチマークであることを浮き彫りにする。
English
Persona skills distill personal interaction histories into portable and executable artifacts for downstream agents. While enabling flexible personalization, this process concentrates fragmented personal signals, amplifies their impact through reuse, and challenges defenses designed for individual records or retrieval-based memory. To systematically investigate the safety of the persona-skill pipeline, we introduce AntiSkillBench, an end-to-end benchmark for evaluating risks and defenses across the persona-skill pipeline. It comprises: (i) a dataset of 7,500 persona-grounded dialogue traces, constructed from 50 behaviorally rich profiles spanning diverse task scenarios; (ii) an evaluation suite that measures skill-level privacy leakage and agent-level attribute disclosure and behavioral impersonation across three skill-distillation strategies; and (iii) a defense evaluation covering four configurations across online and post-hoc interventions, including active risk suppression and passive provenance protection. Experiments across three frontier agents show that persona-skill risks persist across agent backbones and distillation protocols, extending from explicit attributes to communication styles and personality traits. Existing defenses exhibit limited and distillation-dependent effectiveness, failing to generalize across risk and distillation strategies. These results highlight AntiSkillBench as a challenging benchmark for developing privacy-preserving and authenticity-aware persona skills.