ChatPaper.aiChatPaper

超越借用历史:交互式角色扮演评估中的人格对齐用户模拟

Beyond Borrowed Histories: Person-Aligned User Simulation for Interactive Role-Playing Evaluation

July 30, 2026
作者: Yuhang Zhu, Mingxuan Du, Benfeng Xu, Jie Gao, Lingyun Yu, Hongtao Xie
cs.AI

摘要

角色扮演代理(RPAs)已成为大语言模型最重要的消费级应用之一。用户与RPAs进行多轮对话以获得情感安抚等体验,这使得可靠的评估对于衡量能力、比较系统及指导进一步改进至关重要。然而,现有基准通常要求RPA延续一段固定的对话历史,然后使用脱离用户的固定评估标准对延续内容进行评价。我们识别并通过实验证明了该设计的两个局限性。首先,RPA的输出受前序对话历史的影响,这使得在真实多轮场景中对其角色扮演能力进行具有科学依据的评估变得不可能。其次,用户体验在不同个体之间存在显著差异,而传统的固定评估标准未必与用户满意度一致。为此,我们提出PALATE(Person-Aligned LLM-Simulated-User Assessment with Tailored Evaluation,即面向个性化评估的LLM模拟用户对齐评测),一个基于用户模拟器的可扩展RPA基准。PALATE配备了一个包含300个角色档案的池子。其主评估训练五个每用户模拟器,并让它们基于预先冻结的角色档案面板与候选RPAs进行自由形式的多轮对话。除通用质量评估标准外,我们还构建了个性化评估标准以衡量用户满意度;在留出的标注数据上,个性化评估标准与人工判断的一致性高于通用评估标准。在对16个候选的主评估中,PALATE分别刻画了通用轮次质量、长时程会话能力和每位用户在由各候选共同构建的多轮轨迹上的个性化体验。由此,它产生对特定用户-RPA配对的可解释评估,而非将系统压缩为单一的、与用户无关的排名。
English
Role-playing agents (RPAs) have become one of the most important consumer applications of large language models. Users engage in multi-turn conversations with RPAs for experiences such as emotional comfort, making reliable evaluation essential for measuring capability, comparing systems, and guiding further improvement. Existing benchmarks, however, typically require an RPA to continue a fixed dialogue history and then evaluate the continuation using a fixed rubric detached from the user. We identify and empirically demonstrate two limitations of this design. First, an RPA's output is shaped by the preceding dialogue history, preventing a scientifically grounded assessment of its role-playing ability in real multi-turn settings. Second, user experience varies substantially across individuals, and conventional fixed rubrics need not align with user satisfaction. We therefore introduce PALATE (Person-Aligned LLM-Simulated-User Assessment with Tailored Evaluation), a scalable RPA benchmark built on user simulators. PALATE is accompanied by a pool of 300 character profiles. Its main evaluation trains five per-user simulators and lets them engage candidate RPAs in free-form, multi-turn conversations over a pre-frozen panel of character profiles. Alongside a general quality rubric, we construct personalized rubrics to measure user satisfaction; on held-out annotated data, the personalized rubrics show higher agreement with human judgments than the general rubric. In the main evaluation of 16 candidates, PALATE separately characterizes generic turn quality, long-horizon session capability, and per-user experience on multi-turn trajectories co-constructed by each candidate. It thereby produces interpretable evaluations of specific user-RPA pairs rather than compressing systems into a single user-independent ranking.