ChatPaper.aiChatPaper

VideoGAIA:汎用AIアシスタントのためのエージェンティック動画理解ベンチマーク

VideoGAIA: A Benchmark for General AI Assistants on Agentic Video Understanding

August 12, 2026
著者: Fan Zhang, Guangming Yao, Jinyang Wu, Hao Wu, Zheng Lian, Xinyu Geng, Jingdong Chen, Yi Yuan, Pheng-Ann Heng
cs.AI

要旨

ビデオ理解は、マルチモーダル大規模言語モデル(MLLM)の能力を評価するための基礎的なタスクである。しかしながら、既存の最先端モデルはVideo-MMEリーダーボードにおいてすでに約90%の精度を達成しており、従来型の単一ターンビデオ理解タスクは飽和状態に近づきつつあり、高度なMLLMの知能を評価するには不十分となっていることが示唆される。この課題に対処するため、我々は汎用人工知能(AI)アシスタントのためのエージェント的ビデオ理解ベンチマークであるVideoGAIAを提案する。VideoGAIAは単一ターンのビデオ質問応答を超え、ビデオ理解をマルチターンかつツール拡張型の対話プロセスとして定式化する。このプロセスにおいて、モデルはビデオを反復的に認識し、外部ツールを呼び出し、補完的な情報を収集し、ターンを跨いでマルチモーダルな証拠を統合しなければならない。VideoGAIAは、多様で複雑な実世界シナリオを網羅する271のモデル・人間共同設計タスクで構成される。各ビデオ-質問-回答インスタンスは、正確性と適切な難易度の両方を保証するため、3名の人間専門家により独立して検証される。GPT-5.5やKimi-K3などのフロンティアモデルを含む評価対象の全MLLMは、VideoGAIAにおいて60%未満の精度しか達成しておらず、次世代MLLMを評価するための高品質かつ時宜にかなったベンチマークとしての価値が浮き彫りにされる。我々は、VideoGAIAが従来型のビデオ理解からエージェント的ビデオ理解への移行を促進することを期待する。
English
Video understanding is a fundamental task for evaluating the capabilities of multimodal large language models (MLLMs). However, existing leading models have already achieved approximately 90% accuracy on the Video-MME leaderboard, suggesting that conventional single-turn video understanding tasks are becoming increasingly saturated and insufficient for assessing the intelligence of advanced MLLMs. Towards this end, we introduce VideoGAIA, an agentic video understanding benchmark for general artificial intelligence (AI) assistants. Moving beyond one-shot video question answering, VideoGAIA formulates video understanding as a multi-turn, tool-augmented interaction process, where models must iteratively perceive videos, invoke external tools, gather complementary information, and integrate multimodal evidence across turns. VideoGAIA contains 271 model-human co-designed tasks covering diverse and complex real-world scenarios. Each video-question-answer instance is independently verified by three human experts to ensure both correctness and appropriate difficulty. All evaluated MLLMs, including frontier models such as GPT-5.5 and Kimi-K3, achieve less than 60% accuracy on VideoGAIA, highlighting its value as a high-quality and timely benchmark for evaluating next-generation MLLMs. We hope that VideoGAIA will facilitate the transition from conventional video understanding toward agentic video understanding.