ChatPaper.aiChatPaper

VideoGAIA:面向智能体视频理解的通用AI助手基准

VideoGAIA: A Benchmark for General AI Assistants on Agentic Video Understanding

August 12, 2026
作者: Fan Zhang, Guangming Yao, Jinyang Wu, Hao Wu, Zheng Lian, Xinyu Geng, Jingdong Chen, Yi Yuan, Pheng-Ann Heng
cs.AI

摘要

视频理解是评估多模态大语言模型(MLLMs)能力的基础任务。然而,现有领先模型在Video-MME排行榜上已达到约90%的准确率,这表明传统的单轮视频理解任务正日益趋于饱和,不足以评估先进多模态大语言模型的智能水平。为此,我们提出了VideoGAIA,一个面向通用人工智能(AI)助手的智能体视频理解基准。VideoGAIA超越了单轮视频问答的范式,将视频理解构建为多轮、工具增强的交互过程,模型必须迭代地感知视频、调用外部工具、收集补充信息,并跨轮次整合多模态证据。VideoGAIA包含271个由模型与人类协同设计的任务,覆盖多样且复杂的真实世界场景。每个视频-问题-答案实例均由三位人类专家独立验证,以确保正确性和适当的难度。所有被评估的多模态大语言模型,包括GPT-5.5和Kimi-K3等前沿模型,在VideoGAIA上的准确率均低于60%,这凸显了其作为评估下一代多模态大语言模型的高质量、及时基准的价值。我们希望VideoGAIA能够推动从传统视频理解向智能体视频理解的范式转变。
English
Video understanding is a fundamental task for evaluating the capabilities of multimodal large language models (MLLMs). However, existing leading models have already achieved approximately 90% accuracy on the Video-MME leaderboard, suggesting that conventional single-turn video understanding tasks are becoming increasingly saturated and insufficient for assessing the intelligence of advanced MLLMs. Towards this end, we introduce VideoGAIA, an agentic video understanding benchmark for general artificial intelligence (AI) assistants. Moving beyond one-shot video question answering, VideoGAIA formulates video understanding as a multi-turn, tool-augmented interaction process, where models must iteratively perceive videos, invoke external tools, gather complementary information, and integrate multimodal evidence across turns. VideoGAIA contains 271 model-human co-designed tasks covering diverse and complex real-world scenarios. Each video-question-answer instance is independently verified by three human experts to ensure both correctness and appropriate difficulty. All evaluated MLLMs, including frontier models such as GPT-5.5 and Kimi-K3, achieve less than 60% accuracy on VideoGAIA, highlighting its value as a high-quality and timely benchmark for evaluating next-generation MLLMs. We hope that VideoGAIA will facilitate the transition from conventional video understanding toward agentic video understanding.