ChatPaper.aiChatPaper

Video-Oasis: 비디오 이해 평가의 재고

Video-Oasis: Rethinking Evaluation of Video Understanding

July 2, 2026
저자: Geuntaek Lim, Sungjune Park, Jaeyun Lee, Inwoong Lee, Taeoh Kim, Dongyoon Wee, Minho Shim, Yukyung Choi
cs.AI

초록

비디오 이해의 본질적 복잡성으로 인해 Video-LLM 벤치마크 성능이 시각적 인지, 언어적 추론, 또는 지식 사전(prior)에서 비롯된 것인지 판단하기 어렵다. 고수준 추론을 평가하기 위한 많은 벤치마크가 등장했지만, 비디오 이해를 평가하기 위한 공유 기준은 여전히 간과되고 있다. 우리는 또 다른 벤치마크를 제안하는 대신, 한 걸음 물러서서 비디오 이해 평가 기준을 재검토한다. 본 연구에서는 기존 비디오 이해 벤치마크를 체계적으로 감사(audit)하기 위한 지속 가능한 진단 도구 모음인 Video-Oasis를 소개한다. 이 감사 결과, 기존 벤치마크 샘플의 55%가 시각적 입력이나 시간적 맥락 없이도 해결 가능한 것으로 드러났다. 이러한 단축 경로(shortcut)를 걸러낸 후 남은 비디오 고유 과제(video-native challenge)는 상당한 능력 격차를 드러낸다. 최첨단 모델도 무작위 추측보다 약간 나은 수준의 성능만을 보인다. 이러한 발견을 바탕으로, 우리는 추출된 과제를 테스트베드로 활용하여 강건한 비디오 이해에 기여하는 알고리즘 설계 선택이 무엇인지 조사한다. 우리의 연구가 엄격한 비디오 벤치마크를 구축하고 향후 Video-LLM을 평가하기 위한 실용적인 기초를 제공하기를 기대한다. 코드는 https://github.com/sejong-rcv/Video-Oasis에서 확인할 수 있다.
English
The inherent complexity of video understanding makes it difficult to determine whether Video-LLM benchmark performance stems from visual perception, linguistic reasoning, or knowledge priors. While many benchmarks have emerged to assess high-level reasoning, shared criteria for evaluating video understanding remain largely overlooked. Instead of introducing yet another benchmark, we take a step back to re-examine the criteria for evaluating video understanding. In this work, we introduce Video-Oasis, a sustainable diagnostic suite for systematically auditing existing video understanding benchmarks. This audit reveals that 55\% of existing benchmark samples are solvable without visual input or temporal context. After filtering these shortcuts, the remaining video-native challenges expose a substantial capability gap: state-of-the-art models perform only marginally above random guessing. Building on these findings, we use the distilled challenges as a testbed to investigate which algorithmic design choices contribute to robust video understanding. We hope our work provides a practical foundation for constructing rigorous video benchmarks and evaluating future Video-LLMs. Code is available at https://github.com/sejong-rcv/Video-Oasis.