MuScriptor: 多楽器音楽採譜のためのオープンモデル
MuScriptor: An Open Model for Multi-Instrument Music Transcription
July 9, 2026
著者: Simon Rouard, Michael Krause, Axel Roebel, Carl-Johann Simon-Gabriel, Alexandre Défossez
cs.AI
要旨
既存の自動音楽採譜手法は、多くの場合、単一楽器の録音に限定されるか、複雑な実際の音楽ミックスでは機能しません。先行研究では合成訓練データが活用されてきたものの、得られたモデルの汎化性能は低く、現実的な多楽器環境での採譜出力は実用に適さないものとなっています。本研究では、合成データを用いた事前学習の有効性を分析しつつ、実際の音楽音声を用いたファインチューニングと強化学習による事後訓練を組み合わせます。さらに、楽器の存在に関する条件付けを導入し、採譜をカスタマイズ可能にします。最後に、多様な音楽ジャンルの実際の音楽録音に対応するオープンウェイトの多楽器音楽採譜モデルMuScriptorを公開します。
English
Existing methods for automatic music transcription are often limited to single-instrument recordings or fail on complex, real music mixes. Although previous work utilizes synthetic training data, the resulting models generalize poorly, leading to largely unusable transcription output in realistic, multi-instrument settings. In this work, we analyze the effectiveness of synthetic data for pre-training while combining it with fine-tuning on real music audio and post-training using reinforcement learning. We further introduce conditioning on instrument presence to customize transcriptions. Finally, we release MuScriptor, an open-weight multi-instrument music transcription model that works on real-world music recordings from across a diverse range of musical genres.