ChatPaper.aiChatPaper

二者間相互作用におけるうつ病検出のためのセグメント表現に対する話者を考慮した時間的集約戦略:ベンチマーク研究

Speaker-Aware Temporal Aggregation Strategies on Segment Representations for Depression Detection in Dyadic Interaction: A Benchmark Study

July 3, 2026
著者: Anisha Pattanayak, Huang-Cheng Chou, Shrikanth Narayanan, Sudarsana Reddy Kadiri
cs.AI

要旨

音声に基づくうつ病検出は、短い音声セグメントから特徴を圧縮して1つの話者レベルの判断にまとめる。このステップは時間的集約(temporal aggregation)と呼ばれ、単体での研究はほとんど行われていない。ほとんどのベンチマークは、単一の自己教師ありエンコーダと単一の手動で選択された層を固定しているため、報告された性能向上は集約手法そのものではなくパイプライン全体を反映している可能性がある。我々はDEPOOLを提案する。これは、英語と中国語(マンダリン)のうつ病コーパスにおいて、6つの集約アーキテクチャと6つの凍結された音声バックボーンを比較する統制されたベンチマークである。各構成は、手動で固定するのではなく、どのバックボーン層が重要かを学習する。結果として得られた72の構成からなるグリッド全体で、3分の1の構成がすべての話者に対して単一のクラスを予測する状態に陥っている。この問題は手法と同様にバックボーンにも起因しており、単一のシードでの実行で最も安定していたアーキテクチャも、シードを変えて訓練を繰り返すと信頼性が低下する。臨床音声における時間的集約のベンチマーク基準としては、単一パイプラインの平均精度ではなく、バックボーンとシードに対するロバスト性を最重要基準とすべきである。
English
Speech-based depression detection compresses features from short audio segments into one speaker-level decision, a step called temporal aggregation rarely studied on its own. Most benchmarks fix a single self-supervised encoder and a single hand-picked layer, so a reported gain may reflect the pipeline rather than the aggregation method itself. We introduce DEPOOL, a controlled benchmark that compares six aggregation architectures with six frozen speech backbones on an English and a Mandarin depression corpus, where each configuration learns which backbone layers matter rather than fixing one by hand. Across the resulting 72-configuration grid, a third of configurations collapse into predicting a single class for every speaker, a failure tied to the backbone as much as to the method, and the architecture that is most stable in a single-seed run becomes unreliable when training repeats across seeds. Robustness to backbone and seed, rather than average accuracy across a single pipeline, should be a first-class benchmarking criterion for temporal aggregation in clinical speech.