Video-IFBench:評估多模態大型語言模型在影片理解場景中的指令遵循能力
Video-IFBench: Evaluating Instruction Following of Multimodal LLMs in Video Understanding Scenarios
August 26, 2026
作者: Hongbo Liu, Peixian Chen, Sihan Liu, Peiyuan Zhang, Kai Zou, Dian Zheng, Xiaoxing Hu, Yuhao Dong, Mengdan Zhang, Yunhang Shen, Haoyu Cao, Wei Liu, Weibo Gu, Xing Sun, Shengjie Zhao
cs.AI
摘要
多模態大型語言模型(MLLMs)在影片理解方面展現了強大的效能。然而,它們在此領域中遵循指令的能力仍未得到充分探索。現實世界的影片理解要求模型不僅能正確詮釋影片內容,還需滿足使用者指定的多樣化約束條件。現有基準測試主要側重於任務準確度而非指令遵循,使得這項能力未能獲得充分評估。為填補此缺口,我們提出 Video-IFBench,一個全面的基準測試,用於評估影片理解中的指令遵循能力,其中模型必須滿足使用者指定的多樣化約束條件,包括基於視覺與音訊內容的約束。我們開發了一套指令分類法,包含四種模板:單一任務、多任務、選擇及巢狀指令,涵蓋 32 種任務類型與 39 種人工設計的約束類別,遍及語義與格式要求。為降低標註成本,我們建構了一套半自動資料建構流程,結合 MLLMs、程式化處理與人工驗證,產出 1.5K 筆樣本。我們對超過 20 個近期 MLLMs 進行了大規模評估,結果顯示影片指令遵循對當前模型而言仍具挑戰性,尤其對於包含大量約束、語義約束,或需要根據影片內容選擇正確分支或路徑的複雜條件結構之指令更是如此。我們期望這項工作能促進未來在影片理解情境中指令遵循的相關研究。
English
Multimodal Large Language Models (MLLMs) have shown strong performance in video understanding. However, their ability to follow instructions in this domain remains under-explored. Real-world video understanding requires models not only to interpret video content correctly, but also to satisfy diverse user-specified constraints. Existing benchmarks focus primarily on task accuracy rather than instruction adherence, leaving this capability insufficiently evaluated. To address this gap, we introduce Video-IFBench, a comprehensive benchmark for evaluating instruction following in video understanding, where models must satisfy diverse user-specified constraints, including those grounded in visual and audio content. We develop an instruction taxonomy with four templates, including single-task, multi-task, selection, and nested instructions, covering 32 task types and 39 manually designed constraint categories spanning both semantic and format requirements. To reduce annotation cost, we build a semi-automatic data construction pipeline that combines MLLMs, programmatic processing, and human verification, resulting in 1.5K samples. We conduct a large-scale evaluation of more than 20 recent MLLMs and show that video instruction following remains challenging for current models, especially for instructions with many constraints, semantic constraints, or complex conditional structures that require selecting the correct branch or path based on video content. We hope our work will facilitate future research on instruction following in video understanding scenarios.