ChatPaper.aiChatPaper

当智能体学会成为你:人设技能中的隐私泄露、身份冒用风险与防御措施基准测试

When Agents Learn to Be You: Benchmarking Privacy Leakage, Impersonation Risk, and Defenses in Persona Skills

August 4, 2026
作者: Yongli Xiang, Zhifang Zhang, Bojun Yang, Ziming Hong, Lei Feng, Miao Xu, Tongliang Liu
cs.AI

摘要

角色技能将个人交互历史蒸馏为可移植且可执行的工件,供下游智能体使用。这一过程在实现灵活个性化的同时,也集中了零散的个人信号,通过复用放大了其影响,并对针对单个记录或基于检索的记忆所设计的防御机制构成挑战。为系统性地研究角色技能流程的安全性,我们提出了AntiSkillBench,一个用于评估角色技能流程中风险与防御机制的端到端基准。该基准包括:(i)一个包含7,500条基于角色的人机对话轨迹的数据集,这些轨迹由50个行为特征丰富的用户画像构建而成,涵盖多种任务场景;(ii)一套评估体系,用于衡量三种技能蒸馏策略下的技能级隐私泄露、智能体级属性披露和行为模仿;(iii)一项防御评估,涵盖在线干预和事后干预的四种配置,包括主动风险抑制和被动来源保护。在三个前沿智能体上的实验表明,角色技能风险在不同智能体主干和蒸馏协议中持续存在,且从显性属性延伸至沟通风格和人格特质。现有防御机制表现出有限且依赖蒸馏策略的有效性,无法在不同风险和蒸馏策略间泛化。这些结果凸显了AntiSkillBench作为开发隐私保护和真实性感知角色技能挑战性基准的价值。
English
Persona skills distill personal interaction histories into portable and executable artifacts for downstream agents. While enabling flexible personalization, this process concentrates fragmented personal signals, amplifies their impact through reuse, and challenges defenses designed for individual records or retrieval-based memory. To systematically investigate the safety of the persona-skill pipeline, we introduce AntiSkillBench, an end-to-end benchmark for evaluating risks and defenses across the persona-skill pipeline. It comprises: (i) a dataset of 7,500 persona-grounded dialogue traces, constructed from 50 behaviorally rich profiles spanning diverse task scenarios; (ii) an evaluation suite that measures skill-level privacy leakage and agent-level attribute disclosure and behavioral impersonation across three skill-distillation strategies; and (iii) a defense evaluation covering four configurations across online and post-hoc interventions, including active risk suppression and passive provenance protection. Experiments across three frontier agents show that persona-skill risks persist across agent backbones and distillation protocols, extending from explicit attributes to communication styles and personality traits. Existing defenses exhibit limited and distillation-dependent effectiveness, failing to generalize across risk and distillation strategies. These results highlight AntiSkillBench as a challenging benchmark for developing privacy-preserving and authenticity-aware persona skills.