ChatPaper.aiChatPaper

LLM 智能體能照著劇本走嗎?——互動式敘事中長時程一致性的基準測試

Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives

August 8, 2026
作者: Yingpeng Ma, Jianhao Yan, Bei Shi, Ka Hou Kam, Runnan Wang, Xuebo Liu, Yulong Chen, Yue Zhang, Derek F. Wong
cs.AI

摘要

大型語言模型(LLMs)的快速進展正透過實現開放式且流暢的互動敘事,徹底改變遊戲人工智慧領域。然而,現有研究大多忽略了在面對不受約束的使用者干預時,維持長時程邏輯一致性與敘事完整性的關鍵挑戰。為了解決此問題,我們將此挑戰定義為「敘事承諾保持」(Narrative Commitment Preservation, NCP),並以互動敘事作為我們的測試平台。我們引入了 NCP-Bench,一個由 100 個改編自電影情節梗概的敘事環境所組成的基準測試。每個環境都包含一個結構化的敘事規範(軌跡、承諾與初始事實),我們可在玩家代理與敘事代理互動的過程中自動進行檢查。跨越多個最先進 LLM 的實驗結果揭示了一個顯著的長時程一致性缺口:高語言品質並不保證承諾的保持;即便是強大的模型,在面對對抗性干預時也經常生成邏輯衝突的內容,其中表現最佳的模型(GPT-5.2)在 20 回合後的存活率僅達 42%,各模型的事實衝突率介於 40% 至 68% 之間,且僅有少數孤立回合能在 100 回合的限制內滿足所有成就承諾。
English
The rapid advancement of Large Language Models (LLMs) is revolutionizing AI for Games by enabling open-ended and fluid interactive storytelling. However, existing research has largely overlooked the critical challenge of maintaining long-horizon logical consistency and narrative integrity against unconstrained user interventions. To address this, we formulate this challenge as Narrative Commitment Preservation (NCP), and take interactive narrative as our testbed. We introduce NCP-Bench, a benchmark of 100 narrative environments derived from movie synopses. Each environment includes a structured narrative specification (trajectory, commitments, and initial facts) that we can automatically check throughout the interaction between the player agent and the narrator agent. Experiments across state-of-the-art LLMs reveal a substantial long-horizon consistency gap: high linguistic quality does not guarantee commitment preservation; even strong models frequently generate logically conflicting content under adversarial interventions, with the best-performing model (GPT-5.2) achieving only 42% survival rate after 20 turns and fact conflict rates ranging from 40% to 68% across models, and only isolated runs satisfying all achievement commitments within the 100-turn limit.