ChatPaper.aiChatPaper

超越借用的歷史:用於互動式角色扮演評估的人物對齊使用者模擬

Beyond Borrowed Histories: Person-Aligned User Simulation for Interactive Role-Playing Evaluation

July 30, 2026
作者: Yuhang Zhu, Mingxuan Du, Benfeng Xu, Jie Gao, Lingyun Yu, Hongtao Xie
cs.AI

摘要

角色扮演代理(RPAs)已成為大型語言模型最重要的消費級應用之一。使用者透過與RPA進行多輪對話,以獲得情感慰藉等體驗,因此可靠的評估對於衡量能力、比較系統及引導後續改進至關重要。然而,現有的基準測試通常要求RPA延續一段固定的對話歷史,再以脫離使用者的固定評分標準來評估其續接內容。我們識別出此設計的兩項限制並以實證加以證明。首先,RPA的輸出會受到先前對話歷史的影響,這阻礙了在真實多輪情境中對其角色扮演能力進行具有科學基礎的評估。其次,使用者體驗在不同個體之間存在顯著差異,而傳統的固定評分標準未必能與使用者滿意度一致。因此,我們提出PALATE(Person-Aligned LLM-Simulated-User Assessment with Tailored Evaluation,即個人對齊的大型語言模型模擬使用者評估與客製化評分),這是一個建立在使用者模擬器之上的可擴展RPA基準測試。PALATE配備了300個角色設定檔的資料庫。其主要評估流程訓練五個針對個別使用者的模擬器,讓它們與候選RPA在一組預先固定的角色設定檔上進行自由形式的多輪對話。除了通用品質評分標準外,我們還建構了個人化評分標準來衡量使用者滿意度;在留出的標註資料上,個人化評分標準與人類判斷的一致性高於通用評分標準。在16個候選系統的主要評估中,PALATE分別描繪了一般性回合品質、長程對話能力,以及各候選系統在共同建構的多輪軌跡上所呈現的個別使用者體驗。如此一來,它能針對特定的使用者-RPA配對產生可解釋的評估,而非將系統壓縮為單一、與使用者無關的排名。
English
Role-playing agents (RPAs) have become one of the most important consumer applications of large language models. Users engage in multi-turn conversations with RPAs for experiences such as emotional comfort, making reliable evaluation essential for measuring capability, comparing systems, and guiding further improvement. Existing benchmarks, however, typically require an RPA to continue a fixed dialogue history and then evaluate the continuation using a fixed rubric detached from the user. We identify and empirically demonstrate two limitations of this design. First, an RPA's output is shaped by the preceding dialogue history, preventing a scientifically grounded assessment of its role-playing ability in real multi-turn settings. Second, user experience varies substantially across individuals, and conventional fixed rubrics need not align with user satisfaction. We therefore introduce PALATE (Person-Aligned LLM-Simulated-User Assessment with Tailored Evaluation), a scalable RPA benchmark built on user simulators. PALATE is accompanied by a pool of 300 character profiles. Its main evaluation trains five per-user simulators and lets them engage candidate RPAs in free-form, multi-turn conversations over a pre-frozen panel of character profiles. Alongside a general quality rubric, we construct personalized rubrics to measure user satisfaction; on held-out annotated data, the personalized rubrics show higher agreement with human judgments than the general rubric. In the main evaluation of 16 candidates, PALATE separately characterizes generic turn quality, long-horizon session capability, and per-user experience on multi-turn trajectories co-constructed by each candidate. It thereby produces interpretable evaluations of specific user-RPA pairs rather than compressing systems into a single user-independent ranking.