ChatPaper.aiChatPaper

음성에서의 계층별 교차언어 우울증 탐지: 대조 정렬을 통한 분석

Layer-wise Cross-Lingual Depression Detection from Speech: Analysis with Contrastive Alignment

July 3, 2026
저자: Anisha Pattanayak, Hanie Kang, Huang-Cheng Chou, Shrikanth Narayanan, Sudarsana Reddy Kadiri
cs.AI

초록

다른 언어 집단 간 우울증의 진단 및 임상적 발현에서 상당한 차이가 존재한다. 음성 기반 우울증 탐지는 단일 언어 환경에서는 우수한 성능을 보이지만, 교차 언어 일반화는 여전히 해결되지 않은 과제로 남아 있다. 주요 원인 중 하나는 기존 연구들이 화자 그룹화 없이 세그먼트 수준의 무작위 분할을 사용하여, 신원 누출(identity leakage)로 인해 보고된 성능 지표가 부풀려지기 때문이다. 본 연구에서는 병렬 데이터나 대상 언어의 미세 조정 없이 영어와 만다린(Mandarin)의 WavLM 임베딩을 공유된 임상 공간으로 매핑하는 지도 대비 정렬 프레임워크(supervised contrastive alignment framework)인 CLeaD를 제안한다. 52명의 만다린 화자를 대상으로 한 평가에서, 대비 정렬은 하나씩 제외 평가(leave-one-speaker-out) 방식 하에 기준 모델 대비 약간 더 나은 성능을 보였다(F1: 0.640 vs. 0.622). 또한 중간 층(7-8)에서 우울증 집단 재현율(depressed-class recall)을 개선하였으나, 작은 테스트 세트 크기로 인해 일반화 가능성은 제한적이다. 두 가지 결과는 견고하게 유지된다: 모델 확장은 단일 언어 영어 성능을 향상시키는 반면 교차 언어 성능을 저하시켰으며, 화자 신원 누출(speaker identity leakage)이 이전에 보고된 만다린 F1 점수를 0.954까지 인위적으로 부풀렸는데, 본 연구에서 이를 재현하고 정량화하였다.
English
Significant disparities exist in the diagnosis and clinical presentation of depression across different linguistic populations. Speech-based depression detection performs well monolingually, but cross-lingual generalization remains an open challenge. A key reason is that prior work uses segment-level random splits without speaker grouping, leading to identity leakage that inflates reported metrics. We propose CLeaD, a supervised contrastive alignment framework that maps WavLM embeddings from English and Mandarin into a shared clinical space, without parallel data or target-language fine-tuning. Evaluating 52 Mandarin speakers, contrastive alignment modestly outperforms the baseline (F1: 0.640 vs. 0.622) under leave-one-speaker-out evaluation. It also improves depressed-class recall at intermediate layers (7-8), though the small test set limits generalizability. Two findings remain robust: model scaling degrades cross-lingual performance while improving monolingual English, and speaker identity leakage artificially inflated previously reported Mandarin F1 scores to 0.954, an artifact we reproduce and quantify.