NARU:日語極長影片中敘事演化與文化細微差異理解的基準
NARU: A Benchmark for NARrative Evolution and Cultural Nuance Understanding in Japanese Extreme Long Video
August 13, 2026
作者: Yuheng Huang, Jianlang Chen, Jiayang Song, Hua Qi, Aza Kai, Vincent Markert, Edison Marrese-Taylor, Jianjun Zhao, Lei Ma
cs.AI
摘要
長篇影片理解涵蓋超越單純檢索孤立事件的任務,包括追蹤不斷演進的敘事以及詮釋可能保持隱含的社會意義。然而,現有基準很少聯合評估這些能力,特別是在高情境、非英語媒體中。為填補此缺口,我們提出NARU,一個專為評估日語長篇影片中敘事演進與文化理解推理而設計的基準。NARU包含1,481道問題,植基於155部影片,總時長146.8小時,涵蓋四個敘事維度與五個文化維度。為了在此規模下建構基準,我們提出一個基於階層式記憶的標註流程,將原始影片轉化為結構化的事件、敘事與文化標註,再透過任務導向合成與迭代式捷徑移除來生成問題。建構過程包含兩個由68位母語驗證者參與的母語者驗證階段。跨八種模型組態的評測揭露了長程敘事整合與文化基礎推理兩方面的顯著局限。透過揭露這些持續存在的缺口,NARU為發展能夠可靠詮釋長篇、高情境影片的多模態大型語言模型提供了系統性的測試場域。
English
Long-form video understanding encompasses tasks that go beyond retrieving isolated events, including tracking an evolving narrative and interpreting social meaning that may remain implicit. However, existing benchmarks rarely evaluate these capabilities jointly, particularly in high-context, non-English media. To address this gap, we introduce NARU, a benchmark designed to evaluate Narrative evolution and Reasoning on cultural Understanding in Japanese long-form video. NARU consists of 1,481 questions grounded in 155 videos totaling 146.8 hours, spanning four narrative and five cultural dimensions. To construct the benchmark at this scale, we propose a hierarchical memory-based annotation pipeline that transforms raw video into structured event, narrative, and cultural annotations, then generates questions via task-oriented synthesis and iterative shortcut removal. The construction process includes two native-speaker verification stages involving 68 annotators. Evaluations across eight model configurations reveal substantial limitations in both long-range narrative integration and culturally grounded reasoning. By exposing these persistent gaps, NARU offers a systematic testing ground for developing MLLMs capable of reliably interpreting long-form, high-context video.