ChatPaper.aiChatPaper

Video-Oasis:重新思考影片理解評估

Video-Oasis: Rethinking Evaluation of Video Understanding

July 2, 2026
作者: Geuntaek Lim, Sungjune Park, Jaeyun Lee, Inwoong Lee, Taeoh Kim, Dongyoon Wee, Minho Shim, Yukyung Choi
cs.AI

摘要

視頻理解的內在複雜性,使得難以斷定 Video-LLM 的基準表現究竟源自視覺感知、語言推理,還是知識先驗。儘管已有許多基準測試用於評估高階推理能力,但評估影片理解的共通標準仍普遍被忽略。本研究不另提出新的基準,而是回歸根本,重新檢視評估影片理解的標準。我們推出 Video-Oasis,這是一套可持續運作的診斷工具,能系統性地審視現有影片理解基準。審查結果顯示,現有基準中有 55% 的樣本在缺乏視覺輸入或時間脈絡的情況下仍可解答。在排除這些捷徑後,留存下來的影片原生挑戰揭露了顯著的能力差距:當前最佳模型的表現僅略優於隨機猜測。基於這些發現,我們將篩選出的挑戰作為測試平台,探討哪些演算法設計選擇有助於形成穩健的影片理解。期望本研究能為建構嚴謹的影片基準及評估未來 Video-LLM 提供實務基礎。程式碼已公開於 https://github.com/sejong-rcv/Video-Oasis。
English
The inherent complexity of video understanding makes it difficult to determine whether Video-LLM benchmark performance stems from visual perception, linguistic reasoning, or knowledge priors. While many benchmarks have emerged to assess high-level reasoning, shared criteria for evaluating video understanding remain largely overlooked. Instead of introducing yet another benchmark, we take a step back to re-examine the criteria for evaluating video understanding. In this work, we introduce Video-Oasis, a sustainable diagnostic suite for systematically auditing existing video understanding benchmarks. This audit reveals that 55\% of existing benchmark samples are solvable without visual input or temporal context. After filtering these shortcuts, the remaining video-native challenges expose a substantial capability gap: state-of-the-art models perform only marginally above random guessing. Building on these findings, we use the distilled challenges as a testbed to investigate which algorithmic design choices contribute to robust video understanding. We hope our work provides a practical foundation for constructing rigorous video benchmarks and evaluating future Video-LLMs. Code is available at https://github.com/sejong-rcv/Video-Oasis.