MuScriptor:多樂器音樂轉錄的開放模型
MuScriptor: An Open Model for Multi-Instrument Music Transcription
July 9, 2026
作者: Simon Rouard, Michael Krause, Axel Roebel, Carl-Johann Simon-Gabriel, Alexandre Défossez
cs.AI
摘要
現有的自動音樂轉錄方法通常僅適用於單一樂器錄音,或在處理複雜的真實音樂混音時表現不佳。儘管先前的研究使用了合成訓練數據,但所得到的模型推廣能力較差,導致在實際的多樂器場景中產生的轉錄結果幾乎無法使用。在本研究中,我們分析了合成數據在預訓練中的有效性,同時結合對真實音樂音頻的微調,以及使用強化學習進行後訓練。我們進一步引入了基於樂器存在的條件化處理,以自定義轉錄結果。最終,我們發布了MuScriptor,這是一個開放權重的多樂器音樂轉錄模型,能夠處理來自多樣音樂類型的真實世界音樂錄音。
English
Existing methods for automatic music transcription are often limited to single-instrument recordings or fail on complex, real music mixes. Although previous work utilizes synthetic training data, the resulting models generalize poorly, leading to largely unusable transcription output in realistic, multi-instrument settings. In this work, we analyze the effectiveness of synthetic data for pre-training while combining it with fine-tuning on real music audio and post-training using reinforcement learning. We further introduce conditioning on instrument presence to customize transcriptions. Finally, we release MuScriptor, an open-weight multi-instrument music transcription model that works on real-world music recordings from across a diverse range of musical genres.