從音頻嵌入中進行品味感知的音樂檢索
Taste-aware music retrieval from audio embeddings
July 3, 2026
作者: Matteo Spanio, Antonio Rodà
cs.AI
摘要
聲音與味覺之間的跨模態對應在心理學與神經科學領域已有充分確立,但在基於內容的多媒體檢索中卻幾乎未被探討。我們將從音訊預測味覺形式化為一個基於內容的音樂資訊檢索基準測試,採用經知覺驗證的多來源語料庫,比較來自四個HEAR系列中的十個凍結音訊編碼器,這些編碼器共用一個多任務回歸頭,並以門控後期融合作為可配置變體。為評估模型效能,我們計算絕對誤差與秩相關係數。最強系統在巨集均方根誤差(macro RMSE)為0.134的條件下預測五種味覺;在保留的真實音樂數據上,其誤差小於單一評分者與共識偏差的一半(RMSE 0.13 對比 0.28),因此該模型對群體共識的追蹤比平均人類評分者更為精準,且遠低於先前的最先進基準線(0.219)。在絕對誤差方面,編碼器表現統計上持平,單一VGGish模型即匹配最佳融合效果,但門控後期融合的優勢僅體現在秩相關上(巨集皮爾森相關係數 0.724 對比 0.666)。將其操作化為基於內容的檢索索引後,預測味覺空間對309個項目的資料庫進行排序,其精確度遠高於隨機水準的CLAP文字基準線;脊回歸探針與音訊帶阻消去測試則讀取出與已知聲音-味覺對應關係最強的表示。
English
Crossmodal correspondences between sound and taste are well established in psychology and neuroscience, but largely absent from content-based multimedia retrieval. We formalise taste-from-audio prediction as a content-based music information retrieval benchmark over a perceptually validated multi-source corpus, comparing ten frozen audio encoders from the four HEAR families under a shared multi-task regression head, with gated late-fusion as a configurable variant. In order to assess the effectiveness of the models, we compute absolute error and rank correlation. The strongest systems predict the five tastes within a macro RMSE of 0.134; on held-out real music their error is less than half a single rater's deviation from the consensus (RMSE 0.13 vs. 0.28), so the model tracks the group consensus more closely than an average human rater, and well below the previous state of the art baseline (0.219). On absolute error the encoders are statistically flat, with a single VGGish matching the best fusion, but gated late-fusion's advantage is confined to rank correlation (macro Pearson r 0.724 vs. 0.666). Operationalised as a content-based retrieval index, the predicted taste space ranks a 309-item pool far more faithfully than a CLAP-text baseline, which sits at chance; ridge probes and an audio-bandstop knockout read the strongest representations against documented sound-taste correspondences.