ChatPaper.aiChatPaper

멀티 태스크 멀티 프레임 시각적 피아노 전사

Multi-Task Multi-Frame Visual Piano Transcription

August 4, 2026
저자: Yonghyun Kim, Hoyeol Sohn, Juhan Nam, Alexander Lerch
cs.AI

초록

오디오 기반 피아노 전사는 온셋, 피치, 벨로시티에서 우수한 성능을 보이지만, 서스테인 페달은 건반에서 손을 뗀 후에도 소리를 오래 지속시키므로 오디오 시스템은 물리적 키 릴리스 대신 페달로 연장된 오프셋을 예측한다. 그러나 기존 시각적 피아노 전사(VPT) 시스템은 짧은 비디오 윈도우에서 온셋 검출에 초점을 맞추고 있어, 오프셋 정확도는 온셋보다 크게 뒤처지며 음표 수준의 벨로시티는 보고된 바 없다. 이러한 한계를 해결하기 위해 우리는 최초의 완전한 VPT 시스템인 V2N(Video to Notes)을 제시한다. 공유 시간적 백본이 온셋, 오프셋, 키 홀드, 벨로시티를 위한 작업별 헤드에 입력을 공급하며, 윈도우 중심에서만이 아니라 프레임별 지도로 공동 학습된다. 절제 실험은 다중 작업 지도가 온셋 정확도를 향상시키면서 오프셋 및 벨로시티 예측을 가능하게 하며, 더 긴 시간적 맥락이 추가 개선을 가져옴을 보여준다. V2N은 PianoVAM과 R3에서 새로운 최첨단 결과를 달성한다.
English
Audio-based piano transcription performs well on onset, pitch, and velocity, but the sustain pedal lets sound persist long after key release, so audio systems predict pedal-extended offsets rather than physical key release. Yet existing Visual Piano Transcription (VPT) systems focus on onset detection from short video windows, offset accuracy lags onset by a wide margin, and note-level velocity has not been reported. To address these gaps, we present V2N (Video to Notes), the first complete VPT system: a shared temporal backbone feeds task-specific heads for onset, offset, key hold, and velocity, jointly trained with per-frame supervision rather than only at the window center. Ablations show that multi-task supervision enables offset and velocity prediction while improving onset accuracy; longer temporal context yields further improvements. V2N sets new state-of-the-art results on PianoVAM and R3.