音響埋め込みからの嗜好を考慮した音楽検索
Taste-aware music retrieval from audio embeddings
July 3, 2026
著者: Matteo Spanio, Antonio Rodà
cs.AI
要旨
音と味覚の間のクロスモーダル対応は心理学や神経科学では確立されているが、コンテンツベースのマルチメディア検索ではほとんど用いられていない。本研究では、音声から味覚を予測するタスクを、知覚的に検証されたマルチソースコーパスを用いたコンテンツベースの音楽情報検索ベンチマークとして形式化する。4つのHEARファミリーに属する10種類の凍結オーディオエンコーダを、共有マルチタスク回帰ヘッドの下で比較し、ゲート付き後期融合を設定可能なバリアントとして導入する。モデルの有効性を評価するため、絶対誤差と順位相関係数を算出する。最も性能の高いシステムは、5つの味覚をマクロRMSE 0.134以内で予測し、未見の実音楽データでは、単一評価者のコンセンサスからの乖離の半分未満の誤差(RMSE 0.13 vs. 0.28)を示した。すなわち、モデルは平均的な人間の評価者よりもグループコンセンサスに密接に追従し、従来の最先端ベースライン(0.219)を大きく下回る。絶対誤差に関しては、エンコーダ間に統計的に有意な差はなく、単一のVGGishが最良の融合と同等の性能を示すが、ゲート付き後期融合の優位性は順位相関係数に限定される(マクロPearson r 0.724 vs. 0.666)。コンテンツベースの検索指標として運用すると、予測された味覚空間は、偶然レベルの性能にとどまるCLAPテキストベースラインよりもはるかに忠実に309アイテムのプールをランク付けする。リッジプローブとオーディオ帯域除去ノックアウトにより、文献上の音と味覚の対応関係に対して最も強い表現が読み取られる。
English
Crossmodal correspondences between sound and taste are well established in psychology and neuroscience, but largely absent from content-based multimedia retrieval. We formalise taste-from-audio prediction as a content-based music information retrieval benchmark over a perceptually validated multi-source corpus, comparing ten frozen audio encoders from the four HEAR families under a shared multi-task regression head, with gated late-fusion as a configurable variant. In order to assess the effectiveness of the models, we compute absolute error and rank correlation. The strongest systems predict the five tastes within a macro RMSE of 0.134; on held-out real music their error is less than half a single rater's deviation from the consensus (RMSE 0.13 vs. 0.28), so the model tracks the group consensus more closely than an average human rater, and well below the previous state of the art baseline (0.219). On absolute error the encoders are statistically flat, with a single VGGish matching the best fusion, but gated late-fusion's advantage is confined to rank correlation (macro Pearson r 0.724 vs. 0.666). Operationalised as a content-based retrieval index, the predicted taste space ranks a 309-item pool far more faithfully than a CLAP-text baseline, which sits at chance; ridge probes and an audio-bandstop knockout read the strongest representations against documented sound-taste correspondences.