ChatPaper.aiChatPaper

Video-IFBench:映像理解シナリオにおけるマルチモーダルLLMの指示追従の評価

Video-IFBench: Evaluating Instruction Following of Multimodal LLMs in Video Understanding Scenarios

August 26, 2026
著者: Hongbo Liu, Peixian Chen, Sihan Liu, Peiyuan Zhang, Kai Zou, Dian Zheng, Xiaoxing Hu, Yuhao Dong, Mengdan Zhang, Yunhang Shen, Haoyu Cao, Wei Liu, Weibo Gu, Xing Sun, Shengjie Zhao
cs.AI

要旨

マルチモーダル大規模言語モデル(MLLM)は動画理解において強力な性能を示している。しかし、この領域における指示追従能力は依然として十分に探求されていない。現実世界の動画理解では、モデルは動画コンテンツを正確に解釈するだけでなく、多様なユーザー指定の制約を満たすことが求められる。既存のベンチマークは、指示追従よりもタスク精度に主に焦点を当てており、この能力の評価が不十分なままである。このギャップを解消するため、我々はVideo-IFBenchを提案する。これは、動画理解における指示追従を評価するための包括的なベンチマークであり、モデルは視覚および音声コンテンツに基づく制約を含む、多様なユーザー指定の制約を満たす必要がある。我々は、単一タスク、複数タスク、選択、入れ子構造の指示を含む4つのテンプレートからなる指示タクソノミーを開発し、32のタスクタイプと、意味的要件と形式的要件の両方にわたる39の人手設計の制約カテゴリを網羅する。注釈コストを削減するため、MLLM、プログラムによる処理、人手による検証を組み合わせた半自動データ構築パイプラインを構築し、1.5Kサンプルを作成した。我々は、20以上の最近のMLLMを対象に大規模評価を実施し、動画指示追従が現在のモデルにとって依然として困難であることを示す。特に、多くの制約を含む指示、意味的制約を含む指示、または動画コンテンツに基づいて正しい分岐や経路を選択する必要がある複雑な条件構造を持つ指示が難しい。我々の研究が、動画理解シナリオにおける指示追従に関する将来の研究を促進することを期待する。
English
Multimodal Large Language Models (MLLMs) have shown strong performance in video understanding. However, their ability to follow instructions in this domain remains under-explored. Real-world video understanding requires models not only to interpret video content correctly, but also to satisfy diverse user-specified constraints. Existing benchmarks focus primarily on task accuracy rather than instruction adherence, leaving this capability insufficiently evaluated. To address this gap, we introduce Video-IFBench, a comprehensive benchmark for evaluating instruction following in video understanding, where models must satisfy diverse user-specified constraints, including those grounded in visual and audio content. We develop an instruction taxonomy with four templates, including single-task, multi-task, selection, and nested instructions, covering 32 task types and 39 manually designed constraint categories spanning both semantic and format requirements. To reduce annotation cost, we build a semi-automatic data construction pipeline that combines MLLMs, programmatic processing, and human verification, resulting in 1.5K samples. We conduct a large-scale evaluation of more than 20 recent MLLMs and show that video instruction following remains challenging for current models, especially for instructions with many constraints, semantic constraints, or complex conditional structures that require selecting the correct branch or path based on video content. We hope our work will facilitate future research on instruction following in video understanding scenarios.