MuseBench:多模態大語言模型中意圖層面視聽藝術理解的基準測試
MuseBench: Benchmarking Intent-Level Audiovisual Arts Understanding in MLLMs
June 29, 2026
作者: Yuxuan Fan, Gyusik Seo, Jing Hao, Jaemin Cho, Mohit Bansal, Jaehong Yoon
cs.AI
摘要
视听艺术涵盖多种创造性学科,包括电影、视觉艺术、舞台表演和游戏设计。在这些艺术形式中,艺术意义源自视觉、听觉与叙事元素的精心组合(例如,通过幽闭空间的构图放大恐惧感,或借助静默与特写镜头传达悲伤)。真正的艺术理解不仅在于识别所描绘的内容,更在于推理为何选择特定的创作表达方式。尽管多模态大语言模型取得了显著进展,但这一关键的艺术理解层面仍未被充分探索——现有基准测试大多衡量感知识别能力,而忽视了创作意图的推理。为弥补这一空白,我们提出了Musebench,这是一个旨在评估多模态大语言模型对细微艺术理解能力的综合基准。它包含4016道题目,涵盖电影艺术、静态视觉艺术、舞台表演艺术和游戏艺术,这些题目从超过1万个融合专业评论与视觉演示的候选视频中提炼而成。为了大规模捕捉艺术分析的开放性特征,该基准结合了单选题和变选项多选题。所有题目均通过一个包含快捷筛选、对抗性干扰项和专家验证的四阶段迭代流程生成并优化。对28个最先进多模态大语言模型的全面零样本评估显示,即使表现最佳的模型准确率也仅为48.29%,远低于人类专家的87.18%,这暴露了当前模型在创造性领域专业知识上的显著差距。
English
Audiovisual arts encompass diverse creative disciplines, including cinema, visual arts, stage performance, and game design, where artistic meaning arises from deliberate combinations of visual, auditory, and narrative elements (e.g., fear amplified through claustrophobic framing, or grief conveyed through silence and lingering close-ups). True artistic understanding extends beyond recognizing what is depicted to reasoning about why it is expressed through particular creative choices. Despite the strong progress of multimodal large language models (MLLMs), this critical aspect of artistic understanding remains underexplored, as existing benchmarks largely measure perceptual recognition while overlooking reasoning about creative intent. To address this gap, we introduce Musebench, a comprehensive benchmark designed to evaluate MLLMs on nuanced artistic understanding. It comprises 4,016 questions spanning cinematic arts, static visual arts, stage performing arts, and game arts, distilled from over 10K candidate video essays that pair professional commentary with visual demonstration. To capture the open-ended nature of artistic analysis at scale, the benchmark combines single-select and variable-option multi-select questions. All questions are generated and refined through a four-phase iterative pipeline combining shortcut filtering, adversarial distractors, and expert validation. Comprehensive zero-shot evaluation of 28 state-of-the-art MLLMs reveals that even the best-performing model achieves only 48.29% accuracy, substantially below human expert performance of 87.18%, exposing a significant gap in current models' creative domain expertise.