ChatPaper.aiChatPaper

多模態說話者驗證作為對說話者匿名化的威脅

Multimodal Speaker Verification as a Threat to Speaker Anonymization

July 22, 2026
作者: Ashi Garg, Cristina Aggazzotti, Leibny Paola García-Perera, Nicholas Andrews
cs.AI

摘要

大多數自動說話者驗證(ASV)系統均基於單一語句運行,儘管現實世界的互動通常由多個語句組成。隨著語音資料累積,透過聲學、韻律及語言學線索可獲得越來越豐富的說話者資訊,這可能對主要針對聲音特徵的說話者匿名化方法構成挑戰。我們在多語句、多模態情境下研究ASV,並探討跨匿名語音聚合資訊是否影響隱私保護。首先,我們研究僅基於音訊的多個匿名語句聚合,並觀察到隨著可用語音增加,效能持續提升。接著,我們納入韻律及語言學資訊,顯示多模態系統優於單模態方法。最後,我們比較不同聚合策略,發現幀級別聚合可達到最低的等錯誤率(EER)。即使僅使用五個匿名語句,結合音訊與文字相較於僅音訊聚合仍可降低超過15%的EER,證明匿名化後仍有大量說話者判別資訊可被利用。
English
Most automatic speaker verification (ASV) systems operate on individual utterances, despite real-world interactions typically consisting of multiple utterances. As speech accumulates, increasingly rich speaker information becomes available through acoustic, prosodic, and linguistic cues, potentially challenging speaker anonymization methods that primarily target vocal characteristics. We investigate ASV in a multi-utterance, multimodal setting and examine whether aggregating information across anonymized speech impacts privacy. We first study audio-only aggregation across multiple anonymized utterances and observe consistent performance improvements as more speech becomes available. We then incorporate prosodic and linguistic information, showing that multimodal systems outperform unimodal approaches. Finally, we compare aggregation strategies and find that frame-level aggregation yields the lowest EERs. Even with only five anonymized utterances, combining audio and text reduces EER by over 15% relative to audio-only aggregation, demonstrating that substantial speaker-discriminative information remains accessible despite anonymization.