ChatPaper.aiChatPaper

Video-IFBench: 비디오 이해 시나리오에서 멀티모달 LLM의 지시 수행 능력 평가

Video-IFBench: Evaluating Instruction Following of Multimodal LLMs in Video Understanding Scenarios

August 26, 2026
저자: Hongbo Liu, Peixian Chen, Sihan Liu, Peiyuan Zhang, Kai Zou, Dian Zheng, Xiaoxing Hu, Yuhao Dong, Mengdan Zhang, Yunhang Shen, Haoyu Cao, Wei Liu, Weibo Gu, Xing Sun, Shengjie Zhao
cs.AI

초록

다중모달 대규모 언어 모델(MLLM)은 비디오 이해에서 뛰어난 성능을 보여 왔다. 그러나 이 도메인에서 지시를 따르는 능력은 아직 충분히 탐구되지 않았다. 실제 세계의 비디오 이해는 모델이 비디오 콘텐츠를 정확히 해석할 뿐만 아니라 사용자가 지정한 다양한 제약 조건을 충족할 것을 요구한다. 기존 벤치마크는 지시 준수보다는 작업 정확도에 초점을 맞추고 있어, 이러한 능력이 충분히 평가되지 못하고 있다. 이러한 간극을 해소하기 위해 우리는 비디오 이해에서 지시 따르기를 평가하기 위한 포괄적인 벤치마크인 Video-IFBench를 제안한다. 여기서 모델은 시각 및 오디오 콘텐츠에 기반한 제약 조건을 포함하여 사용자가 지정한 다양한 제약 조건을 충족해야 한다. 우리는 단일 작업, 다중 작업, 선택, 중첩 지시의 네 가지 템플릿으로 구성된 지시 분류 체계를 개발했으며, 이는 의미 및 형식 요구 사항을 아우르는 32가지 작업 유형과 39개의 수작업 설계 제약 범주를 포함한다. 주석 비용을 줄이기 위해 MLLM, 프로그램 기반 처리, 인간 검증을 결합한 반자동 데이터 구축 파이프라인을 구축하여 1.5K개의 샘플을 생성했다. 우리는 최근 20개 이상의 MLLM에 대한 대규모 평가를 수행했으며, 비디오 지시 따르기가 현재 모델에게 여전히 어려운 과제임을 확인했다. 특히 제약 조건이 많거나 의미론적 제약 조건을 포함하거나 비디오 콘텐츠에 따라 올바른 분기나 경로를 선택해야 하는 복잡한 조건 구조를 가진 지시에서 더욱 그렇다. 우리는 이 작업이 비디오 이해 시나리오에서 지시 따르기에 관한 향후 연구를 촉진하기를 기대한다.
English
Multimodal Large Language Models (MLLMs) have shown strong performance in video understanding. However, their ability to follow instructions in this domain remains under-explored. Real-world video understanding requires models not only to interpret video content correctly, but also to satisfy diverse user-specified constraints. Existing benchmarks focus primarily on task accuracy rather than instruction adherence, leaving this capability insufficiently evaluated. To address this gap, we introduce Video-IFBench, a comprehensive benchmark for evaluating instruction following in video understanding, where models must satisfy diverse user-specified constraints, including those grounded in visual and audio content. We develop an instruction taxonomy with four templates, including single-task, multi-task, selection, and nested instructions, covering 32 task types and 39 manually designed constraint categories spanning both semantic and format requirements. To reduce annotation cost, we build a semi-automatic data construction pipeline that combines MLLMs, programmatic processing, and human verification, resulting in 1.5K samples. We conduct a large-scale evaluation of more than 20 recent MLLMs and show that video instruction following remains challenging for current models, especially for instructions with many constraints, semantic constraints, or complex conditional structures that require selecting the correct branch or path based on video content. We hope our work will facilitate future research on instruction following in video understanding scenarios.