ChatPaper.aiChatPaper

音声からの層別言語横断うつ病検出:対照的アライメントによる分析

Layer-wise Cross-Lingual Depression Detection from Speech: Analysis with Contrastive Alignment

July 3, 2026
著者: Anisha Pattanayak, Hanie Kang, Huang-Cheng Chou, Shrikanth Narayanan, Sudarsana Reddy Kadiri
cs.AI

要旨

異なる言語集団間では、うつ病の診断と臨床症状に有意な差異が存在する。音声に基づくうつ病検出は単一言語では良好に機能するが、他言語への汎化は未解決の課題である。その主な理由は、従来研究が話者グループ化を行わずにセグメントレベルのランダム分割を用いており、アイデンティティ漏洩を引き起こして報告指標を過大評価しているためである。我々はCLeaDを提案する。これは、並列データや対象言語のファインチューニングを必要とせずに、英語と中国語(マンダリン)のWavLM埋め込みを共有の臨床空間にマッピングする教師あり対照的アライメントフレームワークである。52名の中国語話者を評価したところ、一話者除外評価において、対照的アライメントはベースラインをわずかに上回った(F1: 0.640 vs. 0.622)。また、中間層(7-8)における抑うつクラスの再現率も向上させるが、テストセットが小さいため汎化可能性は限られる。2つの知見は頑健である。すなわち、モデルスケーリングは単一言語英語の性能を向上させる一方で他言語間性能を低下させること、そして話者アイデンティティ漏洩が以前報告された中国語のF1スコアを0.954に人為的に誇張しており、これは我々が再現し定量化したアーティファクトである。
English
Significant disparities exist in the diagnosis and clinical presentation of depression across different linguistic populations. Speech-based depression detection performs well monolingually, but cross-lingual generalization remains an open challenge. A key reason is that prior work uses segment-level random splits without speaker grouping, leading to identity leakage that inflates reported metrics. We propose CLeaD, a supervised contrastive alignment framework that maps WavLM embeddings from English and Mandarin into a shared clinical space, without parallel data or target-language fine-tuning. Evaluating 52 Mandarin speakers, contrastive alignment modestly outperforms the baseline (F1: 0.640 vs. 0.622) under leave-one-speaker-out evaluation. It also improves depressed-class recall at intermediate layers (7-8), though the small test set limits generalizability. Two findings remain robust: model scaling degrades cross-lingual performance while improving monolingual English, and speaker identity leakage artificially inflated previously reported Mandarin F1 scores to 0.954, an artifact we reproduce and quantify.