ChatPaper.aiChatPaper

MuSViT: Een fundamenteel visiemodel voor bladmuziekrepresentatie

MuSViT: A Foundation Vision Model for Sheet Music Representation

June 30, 2026
Auteurs: Carlos Penarrubia, Antonio Rios-Vila, Eliseo Fuentes-Martinez, Juan C. Martinez-Sevilla, Francisco J. Castellanos, María Alfaro-Contreras, Jorge Calvo-Zaragoza
cs.AI

Samenvatting

Funderingsmodellen hebben visuele en taalverwerking getransformeerd door rijke, herbruikbare representaties te bieden die overdraagbaar zijn naar diverse taken. Bladmuziek, als een visuele codering van muzikale taal, mist een dergelijke sterke domeinspecifieke ruggengraat. Wij introduceren MuSViT (Music Score Vision Transformer): het eerste funderingsvisiemodel voor bladmuziekrepresentatie – een ViT-encoder die vooraf is getraind via gemaskeerde auto-encoders op 9,7 miljoen pagina's uit het IMSLP. Om de complexiteit van realistische partituren aan te kunnen, hanteren wij een tweefasig curriculum: een synthetische opwarming op gezet muziek (typeset scores) gevolgd door grootschalige training op het volledige IMSLP-corpus. Wij evalueren MuSViT op vier downstream-taken – herkenning op volledige pagina en op notenbalkniveau, detectie van muzieksymbolen en classificatie van moeilijkheidsgraad – onder twee scenario's: lineaire probing (bevroren encoder) en fijnafstelling. Bij lineaire probing presteert MuSViT consistent beter dan moderne visie-encoders, wat aantoont dat generieke representaties, ongeacht schaal, systematisch tekortschieten wat betreft de gestructureerde symbolische eigenschappen van muzieknotatie. Bij fijnafstelling verbetert MuSViT over het algemeen de taakspecifieke state-of-the-art-methoden. Een aanvullende analyse van inbeddings-transcriptieconsistentie onthult dat MuSViT symbolische muzikale structuur direct codeert in zijn representatieruimte – in tegenstelling tot andere encoders, waarvan de inbeddingen niet correleren met de inhoud van muzieknotatie. Deze resultaten vestigen MuSViT als een funderingsruggengraat voor het begrijpen van bladmuziek.
English
Foundation models have transformed vision and language processing by providing rich, reusable representations that transfer across diverse tasks. Sheet music, as a visual encoding of musical language, lacks such a strong domain-specific backbone. We introduce MuSViT (Music Score Vision Transformer): the first foundation vision model for sheet music representation -- a ViT encoder pre-trained via Masked Autoencoders on 9.7 million pages from the IMSLP. To handle the complexity of real-world scores, we adopt a two-stage curriculum: a synthetic warm-up on typeset scores followed by large-scale training on the full IMSLP corpus. We evaluate MuSViT on four downstream tasks -- full-page and staff-level music score recognition, music symbol detection, and score difficulty classification -- under two scenarios: linear probing (frozen encoder) and fine-tuning. Under linear probing, MuSViT consistently outperforms modern vision encoders, revealing that general-purpose representations, regardless of scale, fall systematically short on the structured symbolic properties of musical notation. Under fine-tuning, MuSViT generally improves upon task-specific state-of-the-art methods. An additional embedding-transcription consistency analysis reveals that MuSViT encodes symbolic musical structure directly in its representation space -- unlike other encoders, whose embeddings do not correlate with music notation content. These results establish MuSViT as a foundation backbone for sheet music understanding.