오디오 임베딩 기반 취향 인식 음악 검색
Taste-aware music retrieval from audio embeddings
July 3, 2026
저자: Matteo Spanio, Antonio Rodà
cs.AI
초록
소리와 미각 간의 교차 양상 대응(crossmodal correspondence)은 심리학과 신경과학에서 잘 정립되어 있으나, 콘텐츠 기반 멀티미디어 검색에서는 거의 존재하지 않는다. 본 연구에서는 음향에서 미각을 예측하는 문제를 지각적으로 검증된 다중 출처 코퍼스를 대상으로 한 콘텐츠 기반 음악 정보 검색 벤치마크로 정형화하며, 네 가지 HEAR 계열의 열 가지 고정된 오디오 인코더를 공유된 다중 작업 회귀 헤드 아래에서 비교하고, 게이트 지연 융합(gated late-fusion)을 구성 가능한 변형으로 도입한다. 모델의 효과성을 평가하기 위해 절대 오차와 순위 상관계수를 계산한다. 가장 강력한 시스템은 다섯 가지 미각을 매크로 RMSE 0.134 이내로 예측하며, 보류된 실제 음악에 대한 오차는 단일 평가자의 합의 편차의 절반 미만(RMSE 0.13 대 0.28)이므로, 모델은 평균 인간 평가자보다 더 정밀하게 그룹 합의를 추적하며, 이전 최고 수준의 기준선(0.219)보다 훨씬 낮은 오차를 보인다. 절대 오차 측면에서 인코더들은 통계적으로 평탄하며, 단일 VGGish가 최상의 융합과 일치하지만, 게이트 지연 융합의 장점은 순위 상관계수(매크로 Pearson r 0.724 대 0.666)에 국한된다. 콘텐츠 기반 검색 인덱스로 운영될 때, 예측된 미각 공간은 309개 항목 풀을 CLAP-텍스트 기준선(이는 우연 수준에 그침)보다 훨씬 충실하게 순위를 매긴다; 능형 프로브(ridge probe)와 오디오 대역저지 넉아웃(knockout)은 문서화된 소리-미각 대응에 대해 가장 강력한 표현을 판독한다.
English
Crossmodal correspondences between sound and taste are well established in psychology and neuroscience, but largely absent from content-based multimedia retrieval. We formalise taste-from-audio prediction as a content-based music information retrieval benchmark over a perceptually validated multi-source corpus, comparing ten frozen audio encoders from the four HEAR families under a shared multi-task regression head, with gated late-fusion as a configurable variant. In order to assess the effectiveness of the models, we compute absolute error and rank correlation. The strongest systems predict the five tastes within a macro RMSE of 0.134; on held-out real music their error is less than half a single rater's deviation from the consensus (RMSE 0.13 vs. 0.28), so the model tracks the group consensus more closely than an average human rater, and well below the previous state of the art baseline (0.219). On absolute error the encoders are statistically flat, with a single VGGish matching the best fusion, but gated late-fusion's advantage is confined to rank correlation (macro Pearson r 0.724 vs. 0.666). Operationalised as a content-based retrieval index, the predicted taste space ranks a 309-item pool far more faithfully than a CLAP-text baseline, which sits at chance; ridge probes and an audio-bandstop knockout read the strongest representations against documented sound-taste correspondences.