ChatPaper.aiChatPaper

이인 간 상호작용에서 우울증 탐지를 위한 세그먼트 표현에 대한 화자 인식 시간적 집계 전략: 벤치마크 연구

Speaker-Aware Temporal Aggregation Strategies on Segment Representations for Depression Detection in Dyadic Interaction: A Benchmark Study

July 3, 2026
저자: Anisha Pattanayak, Huang-Cheng Chou, Shrikanth Narayanan, Sudarsana Reddy Kadiri
cs.AI

초록

음성 기반 우울증 탐지는 짧은 오디오 세그먼트의 특징을 화자 단위 결정으로 압축하는 과정을 포함하며, 이 단계를 시간적 집계(temporal aggregation)라고 하는데, 이 자체만으로는 거의 연구되지 않았다. 대부분의 벤치마크는 단일 자기 지도 인코더와 단일 수동 선택 층을 고정하기 때문에, 보고된 성능 향상은 집계 방법 자체보다는 전체 파이프라인을 반영할 수 있다. 우리는 DEPOOL을 도입하는데, 이는 여섯 개의 집계 아키텍처와 여섯 개의 고정된 음성 백본(backbone)을 비교하는 통제된 벤치마크로, 영어 및 중국어 우울증 말뭉치에서 각 구성이 수동으로 층을 고정하는 대신 어떤 백본 층이 중요한지 학습하도록 한다. 결과적으로 72개 구성의 격자에서, 구성의 3분의 1은 모든 화자에 대해 단일 클래스를 예측하는 붕괴(collapse) 현상을 보였는데, 이는 집계 방법뿐만 아니라 백본에도 기인한 실패이다. 또한 단일 시드(single-seed) 실행에서 가장 안정적인 아키텍처는 시드를 반복하여 훈련할 때 신뢰할 수 없게 된다. 단일 파이프라인에서의 평균 정확도보다는 백본과 시드에 대한 강건성(robustness)이 임상 음성에서의 시간적 집계에 대한 일차적인 벤치마킹 기준이 되어야 한다.
English
Speech-based depression detection compresses features from short audio segments into one speaker-level decision, a step called temporal aggregation rarely studied on its own. Most benchmarks fix a single self-supervised encoder and a single hand-picked layer, so a reported gain may reflect the pipeline rather than the aggregation method itself. We introduce DEPOOL, a controlled benchmark that compares six aggregation architectures with six frozen speech backbones on an English and a Mandarin depression corpus, where each configuration learns which backbone layers matter rather than fixing one by hand. Across the resulting 72-configuration grid, a third of configurations collapse into predicting a single class for every speaker, a failure tied to the backbone as much as to the method, and the architecture that is most stable in a single-seed run becomes unreliable when training repeats across seeds. Robustness to backbone and seed, rather than average accuracy across a single pipeline, should be a first-class benchmarking criterion for temporal aggregation in clinical speech.