マルチモーダル話者検証の話者匿名化に対する脅威
Multimodal Speaker Verification as a Threat to Speaker Anonymization
July 22, 2026
著者: Ashi Garg, Cristina Aggazzotti, Leibny Paola García-Perera, Nicholas Andrews
cs.AI
要旨
ほとんどの自動話者認証システムは個別の発話に対して動作するが、現実の対話は典型的には複数の発話から構成される。発話が蓄積されるにつれて、音響的、韻律的、および言語的手がかりを通じてますます豊富な話者情報が利用可能になり、主に声音特性を標的とする話者匿名化手法に潜在的な課題を提起する。本稿では、複数発話かつマルチモーダルな設定における自動話者認証を調査し、匿名化された発話にわたる情報の集約がプライバシーに影響を与えるかどうかを検討する。まず、複数の匿名化発話にわたる音声のみの集約を研究し、より多くの発話が利用可能になるにつれて一貫した性能向上を観測する。次に、韻律的および言語的情報を取り入れ、マルチモーダルシステムが単一モーダル手法よりも優れていることを示す。最後に、集約戦略を比較し、フレームレベルの集約が最も低い等価誤り率を達成することを見出す。わずか5つの匿名化発話であっても、音声とテキストの組み合わせにより、音声のみの集約と比較して等価誤り率が15%以上低下し、匿名化にもかかわらずかなりの話者識別情報が依然として利用可能であることを実証する。
English
Most automatic speaker verification (ASV) systems operate on individual utterances, despite real-world interactions typically consisting of multiple utterances. As speech accumulates, increasingly rich speaker information becomes available through acoustic, prosodic, and linguistic cues, potentially challenging speaker anonymization methods that primarily target vocal characteristics. We investigate ASV in a multi-utterance, multimodal setting and examine whether aggregating information across anonymized speech impacts privacy. We first study audio-only aggregation across multiple anonymized utterances and observe consistent performance improvements as more speech becomes available. We then incorporate prosodic and linguistic information, showing that multimodal systems outperform unimodal approaches. Finally, we compare aggregation strategies and find that frame-level aggregation yields the lowest EERs. Even with only five anonymized utterances, combining audio and text reduces EER by over 15% relative to audio-only aggregation, demonstrating that substantial speaker-discriminative information remains accessible despite anonymization.