ChatPaper.aiChatPaper

Video-IFBench:评估多模态大语言模型在视频理解场景中的指令遵循能力

Video-IFBench: Evaluating Instruction Following of Multimodal LLMs in Video Understanding Scenarios

August 26, 2026
作者: Hongbo Liu, Peixian Chen, Sihan Liu, Peiyuan Zhang, Kai Zou, Dian Zheng, Xiaoxing Hu, Yuhao Dong, Mengdan Zhang, Yunhang Shen, Haoyu Cao, Wei Liu, Weibo Gu, Xing Sun, Shengjie Zhao
cs.AI

摘要

多模态大语言模型(MLLMs)在视频理解方面已展现出强劲性能。然而,它们在该领域遵循指令的能力仍未得到充分探索。现实世界的视频理解要求模型不仅能够正确解读视频内容,还要满足多样化的用户指定约束。现有基准主要关注任务准确性而非指令遵循,导致该能力未得到充分评估。为填补这一空白,我们提出了视频指令遵循基准(Video-IFBench),这是一个用于评估视频理解中指令遵循能力的综合性基准,要求模型满足多样化的用户指定约束,包括基于视觉和音频内容的约束。我们构建了一个包含四种模板的指令分类体系,涵盖单任务、多任务、选择型和嵌套型指令,覆盖32种任务类型和39种人工设计的约束类别,这些约束涵盖语义和格式要求。为降低标注成本,我们设计了一条半自动数据构建流程,将MLLMs、程序化处理和人工验证相结合,生成了1.5K个样本。我们对20多个最新的MLLMs进行了大规模评估,结果表明视频指令遵循对当前模型仍具挑战性,尤其是对于包含大量约束、语义约束或需要根据视频内容选择正确分支或路径的复杂条件结构的指令。我们希望这项工作能推动视频理解场景下指令遵循的未来研究。
English
Multimodal Large Language Models (MLLMs) have shown strong performance in video understanding. However, their ability to follow instructions in this domain remains under-explored. Real-world video understanding requires models not only to interpret video content correctly, but also to satisfy diverse user-specified constraints. Existing benchmarks focus primarily on task accuracy rather than instruction adherence, leaving this capability insufficiently evaluated. To address this gap, we introduce Video-IFBench, a comprehensive benchmark for evaluating instruction following in video understanding, where models must satisfy diverse user-specified constraints, including those grounded in visual and audio content. We develop an instruction taxonomy with four templates, including single-task, multi-task, selection, and nested instructions, covering 32 task types and 39 manually designed constraint categories spanning both semantic and format requirements. To reduce annotation cost, we build a semi-automatic data construction pipeline that combines MLLMs, programmatic processing, and human verification, resulting in 1.5K samples. We conduct a large-scale evaluation of more than 20 recent MLLMs and show that video instruction following remains challenging for current models, especially for instructions with many constraints, semantic constraints, or complex conditional structures that require selecting the correct branch or path based on video content. We hope our work will facilitate future research on instruction following in video understanding scenarios.