에이전트가 당신을 학습할 때: 페르소나 기술에서 개인정보 유출, 사칭 위험, 및 방어 기법의 벤치마킹
When Agents Learn to Be You: Benchmarking Privacy Leakage, Impersonation Risk, and Defenses in Persona Skills
August 4, 2026
저자: Yongli Xiang, Zhifang Zhang, Bojun Yang, Ziming Hong, Lei Feng, Miao Xu, Tongliang Liu
cs.AI
초록
페르소나 스킬은 개인 상호작용 이력을 다운스트림 에이전트를 위한 휴대 가능하고 실행 가능한 산출물로 압축한다. 유연한 개인화를 가능하게 하는 동시에, 이 과정은 단편적인 개인 신호를 집중시키고 재사용을 통해 그 영향을 증폭시키며, 개별 레코드나 검색 기반 메모리를 위해 설계된 방어 기법에 도전을 제기한다. 페르소나-스킬 파이프라인의 안전성을 체계적으로 조사하기 위해, 우리는 페르소나-스킬 파이프라인 전반에 걸친 위험과 방어 기법을 평가하기 위한 엔드투엔드 벤치마크인 AntiSkillBench를 소개한다. 본 벤치마크는 다음으로 구성된다: (i) 다양한 작업 시나리오를 포괄하는 50개의 행동적으로 풍부한 프로필로부터 구축된 7,500개의 페르소나 기반 대화 트레이스 데이터셋; (ii) 세 가지 스킬 증류 전략에 걸쳐 스킬 수준의 프라이버시 유출과 에이전트 수준의 속성 노출 및 행동 모방을 측정하는 평가 스위트; (iii) 능동적 위험 억제와 수동적 출처 보호를 포함한 온라인 및 사후 개입에 걸친 네 가지 구성을 포괄하는 방어 평가. 세 가지 최첨단 에이전트에 걸친 실험은 페르소나-스킬 위험이 에이전트 백본과 증류 프로토콜에 관계없이 지속되며, 명시적 속성에서 의사소통 스타일과 성격 특성에 이르기까지 확장됨을 보여준다. 기존 방어 기법은 제한적이고 증류 방식에 의존적인 효과성을 보이며, 위험 및 증류 전략 전반에 걸쳐 일반화되지 못한다. 이러한 결과는 프라이버시 보호 및 진정성 인지형 페르소나 스킬 개발을 위한 도전적 벤치마크로서 AntiSkillBench의 중요성을 강조한다.
English
Persona skills distill personal interaction histories into portable and executable artifacts for downstream agents. While enabling flexible personalization, this process concentrates fragmented personal signals, amplifies their impact through reuse, and challenges defenses designed for individual records or retrieval-based memory. To systematically investigate the safety of the persona-skill pipeline, we introduce AntiSkillBench, an end-to-end benchmark for evaluating risks and defenses across the persona-skill pipeline. It comprises: (i) a dataset of 7,500 persona-grounded dialogue traces, constructed from 50 behaviorally rich profiles spanning diverse task scenarios; (ii) an evaluation suite that measures skill-level privacy leakage and agent-level attribute disclosure and behavioral impersonation across three skill-distillation strategies; and (iii) a defense evaluation covering four configurations across online and post-hoc interventions, including active risk suppression and passive provenance protection. Experiments across three frontier agents show that persona-skill risks persist across agent backbones and distillation protocols, extending from explicit attributes to communication styles and personality traits. Existing defenses exhibit limited and distillation-dependent effectiveness, failing to generalize across risk and distillation strategies. These results highlight AntiSkillBench as a challenging benchmark for developing privacy-preserving and authenticity-aware persona skills.