借用された履歴を超えて:対話型ロールプレイ評価のためのペルソナ整合ユーザーシミュレーション
Beyond Borrowed Histories: Person-Aligned User Simulation for Interactive Role-Playing Evaluation
July 30, 2026
著者: Yuhang Zhu, Mingxuan Du, Benfeng Xu, Jie Gao, Lingyun Yu, Hongtao Xie
cs.AI
要旨
ロールプレイングエージェント(RPA)は、大規模言語モデルにおける最も重要な消費者向けアプリケーションの一つとなっている。ユーザーは、感情的な慰めなどの体験を得るためにRPAと多ターン対話を行うため、能力の測定、システムの比較、さらなる改善の指針となる信頼性の高い評価が不可欠である。しかし、既存のベンチマークは通常、RPAに固定された対話履歴を続けさせ、その続きをユーザーから切り離された固定のルーブリックを用いて評価する。我々は、この設計には二つの限界があることを特定し、実証的に示す。第一に、RPAの出力は先行する対話履歴に影響されるため、実際の多ターン設定におけるロールプレイング能力の科学的に根拠のある評価が妨げられる。第二に、ユーザー体験は個人間で大きく異なり、従来の固定ルーブリックはユーザー満足度と必ずしも一致しない。そこで我々は、ユーザーシミュレータに基づくスケーラブルなRPAベンチマークであるPALATE(Person-Aligned LLM-Simulated-User Assessment with Tailored Evaluation)を導入する。PALATEには300のキャラクタープロフィールのプールが付属している。その主要評価では、ユーザー個別の5つのシミュレータを訓練し、事前に固定されたキャラクタープロフィールのパネルを用いて、候補RPAと自由形式の多ターン対話を行わせる。一般的な品質ルーブリックに加えて、ユーザー満足度を測定するための個別化ルーブリックを構築する。保留された注釈付きデータでは、個別化ルーブリックは一般ルーブリックよりも人間の判断との一致度が高いことを示す。16の候補を対象とした主要評価において、PALATEは、各候補が共同構築した多ターンの軌跡上で、一般的なターン品質、長期的なセッション能力、ユーザーごとの体験をそれぞれ別個に特徴付ける。これにより、システムをユーザー非依存の単一ランキングに圧縮するのではなく、特定のユーザーとRPAのペアに対する解釈可能な評価を生成する。
English
Role-playing agents (RPAs) have become one of the most important consumer applications of large language models. Users engage in multi-turn conversations with RPAs for experiences such as emotional comfort, making reliable evaluation essential for measuring capability, comparing systems, and guiding further improvement. Existing benchmarks, however, typically require an RPA to continue a fixed dialogue history and then evaluate the continuation using a fixed rubric detached from the user. We identify and empirically demonstrate two limitations of this design. First, an RPA's output is shaped by the preceding dialogue history, preventing a scientifically grounded assessment of its role-playing ability in real multi-turn settings. Second, user experience varies substantially across individuals, and conventional fixed rubrics need not align with user satisfaction. We therefore introduce PALATE (Person-Aligned LLM-Simulated-User Assessment with Tailored Evaluation), a scalable RPA benchmark built on user simulators. PALATE is accompanied by a pool of 300 character profiles. Its main evaluation trains five per-user simulators and lets them engage candidate RPAs in free-form, multi-turn conversations over a pre-frozen panel of character profiles. Alongside a general quality rubric, we construct personalized rubrics to measure user satisfaction; on held-out annotated data, the personalized rubrics show higher agreement with human judgments than the general rubric. In the main evaluation of 16 candidates, PALATE separately characterizes generic turn quality, long-horizon session capability, and per-user experience on multi-turn trajectories co-constructed by each candidate. It thereby produces interpretable evaluations of specific user-RPA pairs rather than compressing systems into a single user-independent ranking.