ChatPaper.aiChatPaper

「AIペルソナは成長するのか? LLMエージェントにおけるライフイベント後の性格進化の分析とベンチマーキング」

Do AI Personas Grow? Analyzing and Benchmarking Personality Evolution in LLM Agents After Life Events

August 6, 2026
著者: Ming Wang, Peidong Wang, Xiaocui Yang, Daling Wang, Shi Feng, Fiona Fui-Hoon Nah, Ee-Peng Lim
cs.AI

要旨

パーソナリティ条件付きLLMエージェント(PC-Agents)は、感情サポート、社会的シミュレーション、ロールプレイングにおいてますます使用されており、長時間の対話にわたって一貫性を維持する生涯エージェントの開発を動機づけている。そのような一貫性の重要な構成要素はパーソナリティの進化である。すなわち、エージェントは、さまざまな文脈でライフイベントを経験するにつれて、心理学に基づく妥当な変化を経るべきである。先行研究は、LLMのパーソナリティが文脈的摂動によって変化し得ることを示しているが、その変化が特性、イベント、ペルソナ、モデルによってどのように異なるかは、いまだ十分に理解されていない。本研究では、ビッグファイブ特性を心理測定的アンカーとして用い、その軌跡を人間のパーソナリティ心理学における縦断的証拠と対照させながら、11の主要なライフイベント後のイベント誘発性パーソナリティ変化を検討する。4つの分析軸にわたって、PC-Agentsは、人間の変化方向が記録されているイベント–特性ペアと記録されていないペアの両方で、同程度の割合で測定可能な特性変化を示す。変化が期待される方向に沿っている場合でも、その大きさは通常、人間の効果量の範囲を下回る。性別および文化圏のプロンプトは調整効果をほとんど示さず、ペルソナレベルのばらつきは人間のサンプルと比較して3〜4分の1に圧縮される。系統的な比較を可能にするため、我々はイベント誘発性パーソナリティ変化の方向の忠実度を評価する再利用可能なベンチマークBFI-Adaptを導入し、これを用いて14のモデルをランク付けする。検証スイートは、測定された変化がイベントなしの再検査ノイズを上回り、独立に言い換えられたプロンプトの下でも安定しており、シナリオベースの行動選択との収束が限定的かつモデル依存的であり、無関係な対話が介在しても持続することを示す。これらのチェックを総合すると、測定された軌跡は、イベント条件付きのロバストな応答パターンとして確立される。我々の結果は、現在のPC-Agentsが人間のパーソナリティダイナミクスの平均をシミュレートするが、その形状はシミュレートしないことを示唆する。
English
Personality-conditioned LLM agents (PC-Agents) are increasingly used in emotional support, social simulation, and role-playing, motivating the development of lifelong agents that remain coherent over extended interactions. A key component of such coherence is personality evolution: agents should undergo plausible, psychology-grounded changes as they experience life events in different contexts. Although prior work shows that LLM personalities can shift under contextual perturbations, how these shifts vary across traits, events, personas, and models remains poorly understood. We study event-induced personality change after 11 major life events, using the Big Five traits as a psychometric anchor and interpreting the resulting trajectories against longitudinal evidence from human personality psychology. Across four diagnostic axes, PC-Agents exhibit measurable trait shifts at similar rates for event-trait pairs with and without documented human change directions. Even when shifts follow the expected direction, their magnitudes usually fall below human effect-size ranges. Gender and cultural-region prompts show little moderating effect, while persona-level dispersion is compressed three- to four-fold relative to human samples. To enable systematic comparison, we introduce BFI-Adapt, a reusable benchmark for scoring the directional fidelity of event-induced personality change, and use it to rank 14 models. A validation suite shows that the measured shifts exceed no-event retest noise, remain stable under independently paraphrased prompts, exhibit limited and model-dependent convergence with scenario-based behavioral choices, and persist across intervening unrelated dialogue. Together, these checks establish the measured trajectories as robust event-conditioned response patterns. Our results suggest that current PC-Agents simulate the mean of human personality dynamics, but not its shape.