ChatPaper.aiChatPaper

NARU:日本語の極長尺動画におけるナラティブの進化と文化的ニュアンス理解のためのベンチマーク

NARU: A Benchmark for NARrative Evolution and Cultural Nuance Understanding in Japanese Extreme Long Video

August 13, 2026
著者: Yuheng Huang, Jianlang Chen, Jiayang Song, Hua Qi, Aza Kai, Vincent Markert, Edison Marrese-Taylor, Jianjun Zhao, Lei Ma
cs.AI

要旨

長時間動画理解は、孤立した事象の検索を超えたタスクを包含するものであり、展開するナラティブの追跡や、暗黙的に留まる可能性のある社会的意味の解釈を含む。しかしながら、既存のベンチマークは、特に高文脈な非英語メディアにおいて、これらの能力を統合的に評価することはほとんどない。このギャップに対処するため、我々はNARUを導入する。NARUは、日本の長時間動画におけるナラティブの展開と文化的理解に基づく推論を評価するために設計されたベンチマークである。NARUは、155本の動画(総計146.8時間)に基づく1,481の質問で構成され、4つのナラティブ次元と5つの文化的次元にわたる。この規模でベンチマークを構築するため、我々は階層的メモリベースのアノテーションパイプラインを提案する。これは、生の動画を構造化された事象・ナラティブ・文化のアノテーションに変換し、その後、タスク指向の合成と反復的なショートカット除去を通じて質問を生成する。構築プロセスには、68名のアノテーターによる2段階の母語話者検証が含まれる。8つのモデル構成にわたる評価は、長期的なナラティブ統合と文化に基づく推論の両方において重大な限界を明らかにする。これらの持続的なギャップを明示することで、NARUは、長時間・高文脈の動画を確実に解釈できるMLLMsの開発のための体系的なテスト基盤を提供する。
English
Long-form video understanding encompasses tasks that go beyond retrieving isolated events, including tracking an evolving narrative and interpreting social meaning that may remain implicit. However, existing benchmarks rarely evaluate these capabilities jointly, particularly in high-context, non-English media. To address this gap, we introduce NARU, a benchmark designed to evaluate Narrative evolution and Reasoning on cultural Understanding in Japanese long-form video. NARU consists of 1,481 questions grounded in 155 videos totaling 146.8 hours, spanning four narrative and five cultural dimensions. To construct the benchmark at this scale, we propose a hierarchical memory-based annotation pipeline that transforms raw video into structured event, narrative, and cultural annotations, then generates questions via task-oriented synthesis and iterative shortcut removal. The construction process includes two native-speaker verification stages involving 68 annotators. Evaluations across eight model configurations reveal substantial limitations in both long-range narrative integration and culturally grounded reasoning. By exposing these persistent gaps, NARU offers a systematic testing ground for developing MLLMs capable of reliably interpreting long-form, high-context video.