ChatPaper.aiChatPaper

MuScriptor: 다중 악기 음악 전사를 위한 오픈 모델

MuScriptor: An Open Model for Multi-Instrument Music Transcription

July 9, 2026
저자: Simon Rouard, Michael Krause, Axel Roebel, Carl-Johann Simon-Gabriel, Alexandre Défossez
cs.AI

초록

기존의 자동 음악 전사 방법은 종종 단일 악기 녹음에 국한되거나 실제 복잡한 음악 믹스에서 실패하는 경우가 많다. 이전 연구에서 합성 훈련 데이터를 활용했지만, 결과 모델은 일반화 성능이 낮아 실제 다중 악기 환경에서 거의 사용할 수 없는 전사 결과를 초래한다. 본 연구에서는 합성 데이터의 사전 훈련 효과를 분석하면서, 실제 음악 오디오에 대한 미세 조정과 강화 학습을 이용한 후속 훈련을 결합한다. 또한 악기 존재 여부에 따른 조건화를 도입하여 전사를 맞춤화한다. 마지막으로, 다양한 음악 장르에 걸친 실제 음악 녹음에 적용 가능한 공개 가중치 다중 악기 음악 전사 모델인 MuScriptor를 공개한다.
English
Existing methods for automatic music transcription are often limited to single-instrument recordings or fail on complex, real music mixes. Although previous work utilizes synthetic training data, the resulting models generalize poorly, leading to largely unusable transcription output in realistic, multi-instrument settings. In this work, we analyze the effectiveness of synthetic data for pre-training while combining it with fine-tuning on real music audio and post-training using reinforcement learning. We further introduce conditioning on instrument presence to customize transcriptions. Finally, we release MuScriptor, an open-weight multi-instrument music transcription model that works on real-world music recordings from across a diverse range of musical genres.