ChatPaper.aiChatPaper

逐層跨語言語音抑鬱症檢測:基於對比對齊的分析

Layer-wise Cross-Lingual Depression Detection from Speech: Analysis with Contrastive Alignment

July 3, 2026
作者: Anisha Pattanayak, Hanie Kang, Huang-Cheng Chou, Shrikanth Narayanan, Sudarsana Reddy Kadiri
cs.AI

摘要

不同語言人群在憂鬱症的診斷與臨床表現上存在顯著差異。基於語音的憂鬱症檢測在單語環境中表現良好,但跨語言的泛化仍是開放性挑戰。其中一個關鍵原因在於,先前研究採用未區分說話者的片段級隨機劃分,導致身份洩漏,進而虛報評估指標。我們提出 CLeaD 框架,這是一個監督式對比對齊架構,能將英語與普通話的 WavLM 嵌入映射至共享的臨床空間,無需平行資料或目標語言微調。在評估 52 位普通話說話者時,對比對齊在留一說話者交叉驗證下略優於基線(F1:0.640 vs. 0.622)。同時,在中間層(第 7-8 層)改善了憂鬱類別召回率,但測試集規模較小限制了泛化能力。兩項發現具有穩健性:模型擴張在提升單語英語表現的同時,反而降低了跨語言效能;此外,說話者身份洩漏曾使先前報告的普通話 F1 分數人為膨脹至 0.954,我們重現並量化了此偽影。
English
Significant disparities exist in the diagnosis and clinical presentation of depression across different linguistic populations. Speech-based depression detection performs well monolingually, but cross-lingual generalization remains an open challenge. A key reason is that prior work uses segment-level random splits without speaker grouping, leading to identity leakage that inflates reported metrics. We propose CLeaD, a supervised contrastive alignment framework that maps WavLM embeddings from English and Mandarin into a shared clinical space, without parallel data or target-language fine-tuning. Evaluating 52 Mandarin speakers, contrastive alignment modestly outperforms the baseline (F1: 0.640 vs. 0.622) under leave-one-speaker-out evaluation. It also improves depressed-class recall at intermediate layers (7-8), though the small test set limits generalizability. Two findings remain robust: model scaling degrades cross-lingual performance while improving monolingual English, and speaker identity leakage artificially inflated previously reported Mandarin F1 scores to 0.954, an artifact we reproduce and quantify.