ChatPaper.aiChatPaper

MuseBench: Het benchmarken van begrip van audiovisuele kunst op intentieniveau in MLLMs

MuseBench: Benchmarking Intent-Level Audiovisual Arts Understanding in MLLMs

June 29, 2026
Auteurs: Yuxuan Fan, Gyusik Seo, Jing Hao, Jaemin Cho, Mohit Bansal, Jaehong Yoon
cs.AI

Samenvatting

Audiovisuele kunsten omvatten diverse creatieve disciplines, waaronder filmkunst, beeldende kunst, podiumkunsten en gameontwerp, waarbij artistieke betekenis voortkomt uit doordachte combinaties van visuele, auditieve en narratieve elementen (bijvoorbeeld angst versterkt door claustrofobische kadrering, of verdriet overgebracht door stilte en langdurige close-ups). Werkelijk artistiek begrip gaat verder dan het herkennen van wat wordt afgebeeld; het omvat redeneren over waarom iets wordt uitgedrukt door specifieke creatieve keuzes. Ondanks de sterke vooruitgang van multimodale grote taalmodellen (MLLM's) blijft dit cruciale aspect van artistiek begrip onderbelicht, omdat bestaande benchmarks grotendeels perceptuele herkenning meten en voorbijgaan aan redeneren over creatieve intentie. Om deze leemte aan te pakken introduceren we Musebench, een uitgebreide benchmark die is ontworpen om MLLM's te evalueren op genuanceerd artistiek begrip. Deze benchmark omvat 4.016 vragen over filmkunst, statische beeldende kunst, podiumkunsten en gamekunst, gedestilleerd uit meer dan 10K kandidaat-video-essays die professioneel commentaar combineren met visuele demonstratie. Om het open karakter van artistieke analyse op schaal te vatten, combineert de benchmark enkelkeuzevragen en meerkeuzevragen met variabele opties. Alle vragen worden gegenereerd en verfijnd via een vierfasige iteratieve pijplijn met shortcutfiltering, tegenstrijdige afleiders en expertvalidatie. Uit een uitgebreide zero-shotevaluatie van 28 state-of-the-art MLLM's blijkt dat zelfs het best presterende model slechts 48,29% nauwkeurigheid bereikt, aanzienlijk lager dan de menselijke expertprestatie van 87,18%, wat een significante kloof in de creatieve domeinexpertise van huidige modellen blootlegt.
English
Audiovisual arts encompass diverse creative disciplines, including cinema, visual arts, stage performance, and game design, where artistic meaning arises from deliberate combinations of visual, auditory, and narrative elements (e.g., fear amplified through claustrophobic framing, or grief conveyed through silence and lingering close-ups). True artistic understanding extends beyond recognizing what is depicted to reasoning about why it is expressed through particular creative choices. Despite the strong progress of multimodal large language models (MLLMs), this critical aspect of artistic understanding remains underexplored, as existing benchmarks largely measure perceptual recognition while overlooking reasoning about creative intent. To address this gap, we introduce Musebench, a comprehensive benchmark designed to evaluate MLLMs on nuanced artistic understanding. It comprises 4,016 questions spanning cinematic arts, static visual arts, stage performing arts, and game arts, distilled from over 10K candidate video essays that pair professional commentary with visual demonstration. To capture the open-ended nature of artistic analysis at scale, the benchmark combines single-select and variable-option multi-select questions. All questions are generated and refined through a four-phase iterative pipeline combining shortcut filtering, adversarial distractors, and expert validation. Comprehensive zero-shot evaluation of 28 state-of-the-art MLLMs reveals that even the best-performing model achieves only 48.29% accuracy, substantially below human expert performance of 87.18%, exposing a significant gap in current models' creative domain expertise.