LLM 에이전트가 대본을 지킬 수 있을까? 대화형 내러티브의 장기적 일관성을 위한 벤치마크
Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives
August 8, 2026
저자: Yingpeng Ma, Jianhao Yan, Bei Shi, Ka Hou Kam, Runnan Wang, Xuebo Liu, Yulong Chen, Yue Zhang, Derek F. Wong
cs.AI
초록
대규모 언어 모델(LLM)의 급속한 발전은 개방적이고 유연한 인터랙티브 스토리텔링을 가능하게 함으로써 게임용 AI에 혁명을 일으키고 있다. 그러나 기존 연구는 제약 없는 사용자 개입에 맞서 장기적 논리 일관성과 서사 무결성을 유지해야 하는 핵심 과제를 대체로 간과해 왔다. 이에 우리는 이 과제를 서사 약속 보존(Narrative Commitment Preservation, NCP)으로 정식화하고, 인터랙티브 내러티브를 테스트베드로 삼는다. 본 논문에서는 영화 시놉시스에서 도출한 100개의 내러티브 환경으로 구성된 벤치마크인 NCP-Bench를 소개한다. 각 환경은 플레이어 에이전트와 내레이터 에이전트 간의 상호작용 전반에 걸쳐 자동으로 검증할 수 있는 구조화된 내러티브 명세(궤적, 커밋먼트, 초기 사실)를 포함한다. 최첨단 LLM들을 대상으로 한 실험은 상당한 장기적 일관성 격차를 드러낸다. 높은 언어 품질이 커밋먼트 보존을 보장하지 않으며, 강력한 모델조차도 적대적 개입 하에서 논리적으로 상충되는 내용을 빈번히 생성한다. 최고 성능을 보인 모델(GPT-5.2)은 20턴 후 생존율이 42%에 그쳤고, 사실 충돌률은 모델별로 40%~68%에 달했으며, 100턴 제한 내에서 모든 달성 커밋먼트를 충족한 실행은 극히 일부에 불과했다.
English
The rapid advancement of Large Language Models (LLMs) is revolutionizing AI for Games by enabling open-ended and fluid interactive storytelling. However, existing research has largely overlooked the critical challenge of maintaining long-horizon logical consistency and narrative integrity against unconstrained user interventions. To address this, we formulate this challenge as Narrative Commitment Preservation (NCP), and take interactive narrative as our testbed. We introduce NCP-Bench, a benchmark of 100 narrative environments derived from movie synopses. Each environment includes a structured narrative specification (trajectory, commitments, and initial facts) that we can automatically check throughout the interaction between the player agent and the narrator agent. Experiments across state-of-the-art LLMs reveal a substantial long-horizon consistency gap: high linguistic quality does not guarantee commitment preservation; even strong models frequently generate logically conflicting content under adversarial interventions, with the best-performing model (GPT-5.2) achieving only 42% survival rate after 20 turns and fact conflict rates ranging from 40% to 68% across models, and only isolated runs satisfying all achievement commitments within the 100-turn limit.