二元互動中抑鬱症檢測的說話者感知時間聚合策略於片段表示:一項基準研究
Speaker-Aware Temporal Aggregation Strategies on Segment Representations for Depression Detection in Dyadic Interaction: A Benchmark Study
July 3, 2026
作者: Anisha Pattanayak, Huang-Cheng Chou, Shrikanth Narayanan, Sudarsana Reddy Kadiri
cs.AI
摘要
基於語音的抑鬱症檢測將從短音頻片段中提取的特徵壓縮為一個說話者層級的決策,這一步驟稱為時間聚合,但鮮少被單獨研究。大多數基準測試固定使用一個自監督編碼器和一個手動選擇的層,因此報告的改進可能反映的是整個流程的效果,而非聚合方法本身的貢獻。我們提出 DEPOOL,這是一個受控基準,在英語和普通話的抑鬱症語料庫上比較六種聚合架構與六種凍結的語音骨幹網絡,其中每種配置會自行學習哪些骨幹層重要,而非手動固定某一層。在最終的 72 個配置網格中,三分之一的配置崩潰為對每個說話者都預測同一類別,這種失敗與骨幹網絡和方法同樣相關;而在單一種子運行中最穩定的架構,在跨種子重複訓練時卻變得不可靠。對於臨床語音中的時間聚合,應將對骨幹網絡和種子的穩健性,而非單一流程下的平均準確率,作為首要的基準測試標準。
English
Speech-based depression detection compresses features from short audio segments into one speaker-level decision, a step called temporal aggregation rarely studied on its own. Most benchmarks fix a single self-supervised encoder and a single hand-picked layer, so a reported gain may reflect the pipeline rather than the aggregation method itself. We introduce DEPOOL, a controlled benchmark that compares six aggregation architectures with six frozen speech backbones on an English and a Mandarin depression corpus, where each configuration learns which backbone layers matter rather than fixing one by hand. Across the resulting 72-configuration grid, a third of configurations collapse into predicting a single class for every speaker, a failure tied to the backbone as much as to the method, and the architecture that is most stable in a single-seed run becomes unreliable when training repeats across seeds. Robustness to backbone and seed, rather than average accuracy across a single pipeline, should be a first-class benchmarking criterion for temporal aggregation in clinical speech.