ChatPaper.aiChatPaper

VideoGAIA: 범용 AI 어시스턴트를 위한 에이전트형 비디오 이해 벤치마크

VideoGAIA: A Benchmark for General AI Assistants on Agentic Video Understanding

August 12, 2026
저자: Fan Zhang, Guangming Yao, Jinyang Wu, Hao Wu, Zheng Lian, Xinyu Geng, Jingdong Chen, Yi Yuan, Pheng-Ann Heng
cs.AI

초록

비디오 이해는 다중모달 대규모 언어 모델(MLLM)의 역량을 평가하기 위한 기본적인 과제이다. 그러나 기존의 선도 모델들은 이미 Video-MME 리더보드에서 약 90%의 정확도를 달성하여, 기존의 단일 턴 비디오 이해 과제는 점차 포화 상태에 이르렀으며 고급 MLLM의 지능을 평가하기에는 충분하지 않음을 시사한다. 이러한 문제를 해결하기 위해 우리는 범용 인공지능(AI) 어시스턴트를 위한 에이전트적 비디오 이해 벤치마크인 VideoGAIA를 소개한다. VideoGAIA는 일회성 비디오 질의응답을 넘어서, 비디오 이해를 다중 턴, 도구 보강 상호작용 과정으로 정식화한다. 이 과정에서 모델은 반복적으로 비디오를 인지하고, 외부 도구를 호출하며, 보완 정보를 수집하고, 여러 턴에 걸친 다중모달 증거를 통합해야 한다. VideoGAIA는 다양하고 복잡한 실제 세계 시나리오를 포괄하는 271개의 모델-인간 공동 설계 과제를 포함한다. 각 비디오-질문-답변 인스턴스는 정확성과 적절한 난이도를 보장하기 위해 세 명의 인간 전문가가 독립적으로 검증한다. GPT-5.5 및 Kimi-K3와 같은 최첨단 모델을 포함한 모든 평가된 MLLM은 VideoGAIA에서 60% 미만의 정확도를 기록하여, 차세대 MLLM 평가를 위한 고품질의 시기적절한 벤치마크로서의 가치를 강조한다. 우리는 VideoGAIA가 기존의 비디오 이해에서 에이전트적 비디오 이해로의 전환을 촉진하기를 기대한다.
English
Video understanding is a fundamental task for evaluating the capabilities of multimodal large language models (MLLMs). However, existing leading models have already achieved approximately 90% accuracy on the Video-MME leaderboard, suggesting that conventional single-turn video understanding tasks are becoming increasingly saturated and insufficient for assessing the intelligence of advanced MLLMs. Towards this end, we introduce VideoGAIA, an agentic video understanding benchmark for general artificial intelligence (AI) assistants. Moving beyond one-shot video question answering, VideoGAIA formulates video understanding as a multi-turn, tool-augmented interaction process, where models must iteratively perceive videos, invoke external tools, gather complementary information, and integrate multimodal evidence across turns. VideoGAIA contains 271 model-human co-designed tasks covering diverse and complex real-world scenarios. Each video-question-answer instance is independently verified by three human experts to ensure both correctness and appropriate difficulty. All evaluated MLLMs, including frontier models such as GPT-5.5 and Kimi-K3, achieve less than 60% accuracy on VideoGAIA, highlighting its value as a high-quality and timely benchmark for evaluating next-generation MLLMs. We hope that VideoGAIA will facilitate the transition from conventional video understanding toward agentic video understanding.