多模态说话人验证对说话人匿名化的威胁
Multimodal Speaker Verification as a Threat to Speaker Anonymization
July 22, 2026
作者: Ashi Garg, Cristina Aggazzotti, Leibny Paola García-Perera, Nicholas Andrews
cs.AI
摘要
大多数自动说话人验证(ASV)系统通常基于单条语句运行,而现实交互场景往往包含多条语句。随着语音的累积,通过声学、韵律和语言线索可获取越来越丰富的说话人信息,这可能会挑战主要以嗓音特征为目标的说话人匿名化方法。我们研究了多语句、多模态场景下的ASV,并考察跨匿名化语音的信息聚合是否会影响隐私保护。首先,我们研究了仅基于多条匿名化语句的音频聚合,发现随着可用语音量的增加,性能持续提升。随后,我们引入韵律和语言信息,表明多模态系统优于单模态方法。最后,通过比较不同聚合策略,我们发现帧级聚合能够获得最低的等错误率(EER)。即使仅使用五条匿名化语句,结合音频与文本的方法相较于纯音频聚合,等错误率相对降低了15%以上,这表明尽管经过匿名化处理,大量具有说话人判别性的信息仍然可被获取。
English
Most automatic speaker verification (ASV) systems operate on individual utterances, despite real-world interactions typically consisting of multiple utterances. As speech accumulates, increasingly rich speaker information becomes available through acoustic, prosodic, and linguistic cues, potentially challenging speaker anonymization methods that primarily target vocal characteristics. We investigate ASV in a multi-utterance, multimodal setting and examine whether aggregating information across anonymized speech impacts privacy. We first study audio-only aggregation across multiple anonymized utterances and observe consistent performance improvements as more speech becomes available. We then incorporate prosodic and linguistic information, showing that multimodal systems outperform unimodal approaches. Finally, we compare aggregation strategies and find that frame-level aggregation yields the lowest EERs. Even with only five anonymized utterances, combining audio and text reduces EER by over 15% relative to audio-only aggregation, demonstrating that substantial speaker-discriminative information remains accessible despite anonymization.