ChatPaper.aiChatPaper

MuScriptor: 多乐器音乐转录的开源模型

MuScriptor: An Open Model for Multi-Instrument Music Transcription

July 9, 2026
作者: Simon Rouard, Michael Krause, Axel Roebel, Carl-Johann Simon-Gabriel, Alexandre Défossez
cs.AI

摘要

现有的自动音乐转录方法通常局限于单乐器录音,或无法处理复杂的真实音乐混音。尽管已有研究使用合成训练数据,但得到的模型泛化能力差,在真实的多乐器场景中生成的转录结果基本不可用。本研究分析了合成数据在预训练中的有效性,并将其与真实音乐音频上的微调、以及使用强化学习的后训练相结合。我们进一步引入基于乐器存在状态的条件控制,以实现转录结果的定制化。最后,我们发布了MuScriptor——一个开放权重的多乐器音乐转录模型,该模型适用于来自多种音乐流派的真实世界音乐录音。
English
Existing methods for automatic music transcription are often limited to single-instrument recordings or fail on complex, real music mixes. Although previous work utilizes synthetic training data, the resulting models generalize poorly, leading to largely unusable transcription output in realistic, multi-instrument settings. In this work, we analyze the effectiveness of synthetic data for pre-training while combining it with fine-tuning on real music audio and post-training using reinforcement learning. We further introduce conditioning on instrument presence to customize transcriptions. Finally, we release MuScriptor, an open-weight multi-instrument music transcription model that works on real-world music recordings from across a diverse range of musical genres.