ChatPaper.aiChatPaper

MuseBench: MLLM에서의 의도 수준 시청각 예술 이해 벤치마킹

MuseBench: Benchmarking Intent-Level Audiovisual Arts Understanding in MLLMs

June 29, 2026
저자: Yuxuan Fan, Gyusik Seo, Jing Hao, Jaemin Cho, Mohit Bansal, Jaehong Yoon
cs.AI

초록

시청각 예술은 영화, 시각 예술, 무대 공연, 게임 디자인 등 다양한 창의적 분야를 아우르며, 예술적 의미는 시각, 청각, 서사 요소의 의도적인 결합(예: 폐쇄적 프레이밍을 통한 공포 증폭, 침묵과 긴 클로즈업을 통한 슬픔 전달)에서 발생한다. 진정한 예술 이해는 묘사된 내용을 인식하는 것을 넘어, 특정 창의적 선택을 통해 그것이 왜 표현되었는지 추론하는 데까지 이른다. 다중 모달 대규모 언어 모델(MLLM)의 눈부신 발전에도 불구하고, 예술 이해의 이러한 핵심 측면은 여전히 충분히 탐구되지 않았으며, 기존 벤치마크는 주로 지각적 인식을 측정하면서 창의적 의도에 대한 추론은 간과하고 있다. 이러한 격차를 해소하기 위해 우리는 Musebench를 소개한다. 이는 MLLM의 미묘한 예술 이해도를 평가하도록 설계된 포괄적인 벤치마크로, 영화 예술, 정적 시각 예술, 무대 공연 예술, 게임 예술을 아우르는 4,016개의 질문으로 구성되며, 전문 논평과 시각적 시연을 결합한 10,000개 이상의 후보 비디오 에세이에서 추출되었다. 예술 분석의 개방적 특성을 대규모로 포착하기 위해, 벤치마크는 단일 선택 및 가변 옵션 다중 선택 질문을 결합한다. 모든 질문은 단축 필터링, 적대적 방해 요소, 전문가 검증을 결합한 4단계 반복 파이프라인을 통해 생성 및 정제된다. 28개의 최신 MLLM에 대한 포괄적인 제로샷 평가 결과, 최고 성능의 모델조차 48.29%의 정확도에 그쳐 인간 전문가의 87.18%에 비해 현저히 낮았으며, 이는 현재 모델의 창의적 도메인 전문성에 상당한 격차가 있음을 드러낸다.
English
Audiovisual arts encompass diverse creative disciplines, including cinema, visual arts, stage performance, and game design, where artistic meaning arises from deliberate combinations of visual, auditory, and narrative elements (e.g., fear amplified through claustrophobic framing, or grief conveyed through silence and lingering close-ups). True artistic understanding extends beyond recognizing what is depicted to reasoning about why it is expressed through particular creative choices. Despite the strong progress of multimodal large language models (MLLMs), this critical aspect of artistic understanding remains underexplored, as existing benchmarks largely measure perceptual recognition while overlooking reasoning about creative intent. To address this gap, we introduce Musebench, a comprehensive benchmark designed to evaluate MLLMs on nuanced artistic understanding. It comprises 4,016 questions spanning cinematic arts, static visual arts, stage performing arts, and game arts, distilled from over 10K candidate video essays that pair professional commentary with visual demonstration. To capture the open-ended nature of artistic analysis at scale, the benchmark combines single-select and variable-option multi-select questions. All questions are generated and refined through a four-phase iterative pipeline combining shortcut filtering, adversarial distractors, and expert validation. Comprehensive zero-shot evaluation of 28 state-of-the-art MLLMs reveals that even the best-performing model achieves only 48.29% accuracy, substantially below human expert performance of 87.18%, exposing a significant gap in current models' creative domain expertise.