VideoGAIA:通用AI助理於代理式影片理解之基準測試
VideoGAIA: A Benchmark for General AI Assistants on Agentic Video Understanding
August 12, 2026
作者: Fan Zhang, Guangming Yao, Jinyang Wu, Hao Wu, Zheng Lian, Xinyu Geng, Jingdong Chen, Yi Yuan, Pheng-Ann Heng
cs.AI
摘要
影片理解是評估多模態大型語言模型(MLLMs)能力的一項基本任務。然而,現有領先模型在Video-MME排行榜上已達到約90%的準確率,顯示傳統的單輪影片理解任務正逐漸趨於飽和,已不足以評估先進MLLMs的智慧程度。為此,我們提出VideoGAIA,一個專為通用人工智慧(AI)助理設計的智能體驅動影片理解基準。VideoGAIA跳脫一次性影片問答的框架,將影片理解建構為一個多輪、工具增強的互動過程:模型必須迭代式地感知影片內容、呼叫外部工具、蒐集互補資訊,並跨回合整合多模態證據。VideoGAIA包含271個由模型與人類共同設計的任務,涵蓋多樣且複雜的真實世界場景。每個影片-問題-答案實例均由三位人類專家獨立驗證,以確保其正確性與難度適切。所有受評的MLLMs,包括GPT-5.5與Kimi-K3等前沿模型,在VideoGAIA上的準確率均低於60%,凸顯其作為評估下一代MLLMs之高品質且具時效性基準的價值。我們期望VideoGAIA能促進由傳統影片理解向智能體驅動影片理解的過渡。
English
Video understanding is a fundamental task for evaluating the capabilities of multimodal large language models (MLLMs). However, existing leading models have already achieved approximately 90% accuracy on the Video-MME leaderboard, suggesting that conventional single-turn video understanding tasks are becoming increasingly saturated and insufficient for assessing the intelligence of advanced MLLMs. Towards this end, we introduce VideoGAIA, an agentic video understanding benchmark for general artificial intelligence (AI) assistants. Moving beyond one-shot video question answering, VideoGAIA formulates video understanding as a multi-turn, tool-augmented interaction process, where models must iteratively perceive videos, invoke external tools, gather complementary information, and integrate multimodal evidence across turns. VideoGAIA contains 271 model-human co-designed tasks covering diverse and complex real-world scenarios. Each video-question-answer instance is independently verified by three human experts to ensure both correctness and appropriate difficulty. All evaluated MLLMs, including frontier models such as GPT-5.5 and Kimi-K3, achieve less than 60% accuracy on VideoGAIA, highlighting its value as a high-quality and timely benchmark for evaluating next-generation MLLMs. We hope that VideoGAIA will facilitate the transition from conventional video understanding toward agentic video understanding.