Video-Oasis: ビデオ理解評価の再考
Video-Oasis: Rethinking Evaluation of Video Understanding
July 2, 2026
著者: Geuntaek Lim, Sungjune Park, Jaeyun Lee, Inwoong Lee, Taeoh Kim, Dongyoon Wee, Minho Shim, Yukyung Choi
cs.AI
要旨
動画理解の内在的な複雑さゆえに、Video-LLMのベンチマーク性能が視覚認知、言語推論、あるいは知識の事前性のいずれに起因するのかを判断することは困難である。高レベルの推論を評価する多くのベンチマークが登場してきた一方で、動画理解を評価するための共通基準は依然としてほとんど注目されていない。我々は新たなベンチマークを導入する代わりに、一歩引いて動画理解の評価基準を再検討する。本稿では、既存の動画理解ベンチマークを体系的に監査するための持続可能な診断スイート、Video-Oasisを提案する。この監査により、既存のベンチマークサンプルの55%が視覚入力や時間的文脈なしで解けることが明らかになった。これらの近道を取り除いた後、残った動画本来の課題は、大きな能力ギャップを露呈する。すなわち、最先端のモデルでさえランダムな推測をわずかに上回る性能しか示さないのである。これらの知見に基づき、抽出された課題をテストベッドとして、どのアルゴリズム設計の選択が頑健な動画理解に寄与するかを調査する。本研究が、厳密な動画ベンチマークの構築と将来のVideo-LLMの評価のための実用的な基盤を提供することを願っている。コードは https://github.com/sejong-rcv/Video-Oasis で入手可能である。
English
The inherent complexity of video understanding makes it difficult to determine whether Video-LLM benchmark performance stems from visual perception, linguistic reasoning, or knowledge priors. While many benchmarks have emerged to assess high-level reasoning, shared criteria for evaluating video understanding remain largely overlooked. Instead of introducing yet another benchmark, we take a step back to re-examine the criteria for evaluating video understanding. In this work, we introduce Video-Oasis, a sustainable diagnostic suite for systematically auditing existing video understanding benchmarks. This audit reveals that 55\% of existing benchmark samples are solvable without visual input or temporal context. After filtering these shortcuts, the remaining video-native challenges expose a substantial capability gap: state-of-the-art models perform only marginally above random guessing. Building on these findings, we use the distilled challenges as a testbed to investigate which algorithmic design choices contribute to robust video understanding. We hope our work provides a practical foundation for constructing rigorous video benchmarks and evaluating future Video-LLMs. Code is available at https://github.com/sejong-rcv/Video-Oasis.