NARU:面向日本超长视频的叙事演进与文化内涵理解基准测试集
NARU: A Benchmark for NARrative Evolution and Cultural Nuance Understanding in Japanese Extreme Long Video
August 13, 2026
作者: Yuheng Huang, Jianlang Chen, Jiayang Song, Hua Qi, Aza Kai, Vincent Markert, Edison Marrese-Taylor, Jianjun Zhao, Lei Ma
cs.AI
摘要
长视频理解涵盖的任务超越了孤立事件的检索,包括追踪不断演进的叙事脉络以及解读可能保持隐性的社会含义。然而,现有基准很少对这些能力进行联合评估,尤其是在高语境、非英语媒体中。为弥补这一空白,我们提出了NARU,一个专为评估日语长视频中叙事演进与文化理解推理而设计的基准。NARU包含1,481个问题,基于155个视频(总计146.8小时),涵盖四个叙事维度和五个文化维度。为了在该规模上构建基准,我们提出了一种基于层级化记忆的标注流程,将原始视频转化为结构化的情节、叙事和文化标注,然后通过任务导向的综合生成与迭代式快捷路径消除来生成问题。构建过程包括两个由68名标注者参与的目标语母语者验证阶段。对八种模型配置的评估揭示了长程叙事整合与文化根基推理方面的显著局限。通过揭示这些持续存在的差距,NARU为开发能够可靠解读长时长、高语境视频的多模态大语言模型(MLLMs)提供了系统化的测试平台。
English
Long-form video understanding encompasses tasks that go beyond retrieving isolated events, including tracking an evolving narrative and interpreting social meaning that may remain implicit. However, existing benchmarks rarely evaluate these capabilities jointly, particularly in high-context, non-English media. To address this gap, we introduce NARU, a benchmark designed to evaluate Narrative evolution and Reasoning on cultural Understanding in Japanese long-form video. NARU consists of 1,481 questions grounded in 155 videos totaling 146.8 hours, spanning four narrative and five cultural dimensions. To construct the benchmark at this scale, we propose a hierarchical memory-based annotation pipeline that transforms raw video into structured event, narrative, and cultural annotations, then generates questions via task-oriented synthesis and iterative shortcut removal. The construction process includes two native-speaker verification stages involving 68 annotators. Evaluations across eight model configurations reveal substantial limitations in both long-range narrative integration and culturally grounded reasoning. By exposing these persistent gaps, NARU offers a systematic testing ground for developing MLLMs capable of reliably interpreting long-form, high-context video.