NARU: 일본 극한 장시간 비디오에서의 내러티브 진화 및 문화적 뉘앙스 이해를 위한 벤치마크
NARU: A Benchmark for NARrative Evolution and Cultural Nuance Understanding in Japanese Extreme Long Video
August 13, 2026
저자: Yuheng Huang, Jianlang Chen, Jiayang Song, Hua Qi, Aza Kai, Vincent Markert, Edison Marrese-Taylor, Jianjun Zhao, Lei Ma
cs.AI
초록
장편 비디오 이해는 고립된 사건 검색을 넘어 진화하는 내러티브 추적과 암묵적으로 남을 수 있는 사회적 의미 해석을 포함하는 과업을 포괄한다. 그러나 기존 벤치마크는 특히 고맥락(high-context) 비영어권 미디어에서 이러한 능력을 종합적으로 평가하는 경우가 드물다. 이러한 격차를 해소하기 위해, 우리는 일본어 장편 비디오에서 내러티브 진화와 문화적 이해 추론을 평가하도록 설계된 벤치마크인 NARU를 소개한다. NARU는 총 146.8시간에 해당하는 155개 비디오에 기반한 1,481개의 질문으로 구성되며, 4개의 내러티브 차원과 5개의 문화적 차원에 걸쳐 있다. 이 규모로 벤치마크를 구축하기 위해, 우리는 계층적 메모리 기반 주석 파이프라인을 제안한다. 이 파이프라인은 원시 비디오를 구조화된 사건, 내러티브 및 문화 주석으로 변환한 후, 과업 지향적 합성과 반복적 단축 경로 제거를 통해 질문을 생성한다. 구축 과정에는 68명의 주석자가 참여하는 두 차례의 원어민 검증 단계가 포함된다. 8개 모델 구성에 대한 평가는 장거리 내러티브 통합과 문화적 기반 추론 모두에서 상당한 한계를 드러낸다. 이러한 지속적 결함을 노출함으로써, NARU는 장편 고맥락 비디오를 신뢰성 있게 해석할 수 있는 MLLM 개발을 위한 체계적 테스트 기반을 제공한다.
English
Long-form video understanding encompasses tasks that go beyond retrieving isolated events, including tracking an evolving narrative and interpreting social meaning that may remain implicit. However, existing benchmarks rarely evaluate these capabilities jointly, particularly in high-context, non-English media. To address this gap, we introduce NARU, a benchmark designed to evaluate Narrative evolution and Reasoning on cultural Understanding in Japanese long-form video. NARU consists of 1,481 questions grounded in 155 videos totaling 146.8 hours, spanning four narrative and five cultural dimensions. To construct the benchmark at this scale, we propose a hierarchical memory-based annotation pipeline that transforms raw video into structured event, narrative, and cultural annotations, then generates questions via task-oriented synthesis and iterative shortcut removal. The construction process includes two native-speaker verification stages involving 68 annotators. Evaluations across eight model configurations reveal substantial limitations in both long-range narrative integration and culturally grounded reasoning. By exposing these persistent gaps, NARU offers a systematic testing ground for developing MLLMs capable of reliably interpreting long-form, high-context video.