MuseBench: MLLMにおける意図レベルの視聴覚芸術理解のベンチマーキング
MuseBench: Benchmarking Intent-Level Audiovisual Arts Understanding in MLLMs
June 29, 2026
著者: Yuxuan Fan, Gyusik Seo, Jing Hao, Jaemin Cho, Mohit Bansal, Jaehong Yoon
cs.AI
要旨
視聴覚芸術は、映画、視覚芸術、舞台芸術、ゲームデザインなど多様な創造的分野を含みます。これらの分野では、視覚、聴覚、物語的要素を意図的に組み合わせることで芸術的意味が生まれます(例えば、閉所恐怖症を誘発するフレーミングで恐怖を増幅したり、沈黙と長回しのクローズアップで悲しみを表現するなど)。真の芸術理解は、何が描写されているかを認識するだけでなく、なぜ特定の創造的選択を通じて表現されているのかを推論することにあります。マルチモーダル大規模言語モデル(MLLM)の著しい進歩にもかかわらず、芸術理解におけるこの重要な側面は未だ十分に探求されていません。既存のベンチマークの多くは知覚的認識を測定する一方で、創造的意図に関する推論を見落としているからです。このギャップに対処するため、我々はMusebenchを導入します。これは、MLLMの微妙な芸術理解を評価するために設計された包括的なベンチマークです。映画芸術、静的視覚芸術、舞台芸術、ゲーム芸術にわたる4,016の質問で構成され、専門家による解説と視覚的実演を組み合わせた10,000以上の候補ビデオエッセイから抽出されました。大規模な芸術分析の自由回答形式を捉えるため、本ベンチマークは単一選択問題と可変オプションの複数選択問題を組み合わせています。すべての質問は、ショートカットフィルタリング、敵対的ディストラクタ、専門家による検証を組み合わせた4段階の反復的パイプラインを通じて生成・洗練されています。28の最先端MLLMに対する包括的なゼロショット評価の結果、最高性能のモデルでも48.29%の精度にとどまり、人間の専門家による87.18%を大幅に下回り、現在のモデルの創造的領域における専門性に大きなギャップがあることが明らかになりました。
English
Audiovisual arts encompass diverse creative disciplines, including cinema, visual arts, stage performance, and game design, where artistic meaning arises from deliberate combinations of visual, auditory, and narrative elements (e.g., fear amplified through claustrophobic framing, or grief conveyed through silence and lingering close-ups). True artistic understanding extends beyond recognizing what is depicted to reasoning about why it is expressed through particular creative choices. Despite the strong progress of multimodal large language models (MLLMs), this critical aspect of artistic understanding remains underexplored, as existing benchmarks largely measure perceptual recognition while overlooking reasoning about creative intent. To address this gap, we introduce Musebench, a comprehensive benchmark designed to evaluate MLLMs on nuanced artistic understanding. It comprises 4,016 questions spanning cinematic arts, static visual arts, stage performing arts, and game arts, distilled from over 10K candidate video essays that pair professional commentary with visual demonstration. To capture the open-ended nature of artistic analysis at scale, the benchmark combines single-select and variable-option multi-select questions. All questions are generated and refined through a four-phase iterative pipeline combining shortcut filtering, adversarial distractors, and expert validation. Comprehensive zero-shot evaluation of 28 state-of-the-art MLLMs reveals that even the best-performing model achieves only 48.29% accuracy, substantially below human expert performance of 87.18%, exposing a significant gap in current models' creative domain expertise.