ChatPaper.aiChatPaper

화자 익명화에 대한 위협으로서의 멀티모달 화자 검증

Multimodal Speaker Verification as a Threat to Speaker Anonymization

July 22, 2026
저자: Ashi Garg, Cristina Aggazzotti, Leibny Paola García-Perera, Nicholas Andrews
cs.AI

초록

대부분의 자동 화자 검증(ASV) 시스템은 개별 발화를 대상으로 작동하지만, 실제 상호작용은 일반적으로 여러 발화로 구성된다. 음성이 축적됨에 따라 음향적, 운율적, 언어적 단서를 통해 점점 더 풍부한 화자 정보를 얻을 수 있게 되며, 이는 주로 음성 특성을 표적으로 하는 화자 익명화 방식에 도전이 될 수 있다. 우리는 다중 발화, 다중 모달 환경에서 ASV를 조사하고, 익명화된 음성 전반에 걸쳐 정보를 집계하는 것이 프라이버시에 미치는 영향을 검토한다. 먼저 다중 익명화 발화에 걸친 오디오 전용 집계를 연구하여, 더 많은 음성이 제공됨에 따라 일관된 성능 향상이 나타남을 관찰한다. 그다음 운율 및 언어 정보를 통합하여 다중 모달 시스템이 단일 모달 접근법보다 우수함을 보여준다. 마지막으로 집계 전략을 비교하여 프레임 수준 집계가 가장 낮은 EER을 산출함을 발견한다. 익명화된 발화가 5개만 있어도 오디오와 텍스트를 결합하면 오디오 전용 집계에 비해 EER이 15% 이상 감소하여, 익명화에도 불구하고 상당한 화자 식별 정보에 접근 가능함을 입증한다.
English
Most automatic speaker verification (ASV) systems operate on individual utterances, despite real-world interactions typically consisting of multiple utterances. As speech accumulates, increasingly rich speaker information becomes available through acoustic, prosodic, and linguistic cues, potentially challenging speaker anonymization methods that primarily target vocal characteristics. We investigate ASV in a multi-utterance, multimodal setting and examine whether aggregating information across anonymized speech impacts privacy. We first study audio-only aggregation across multiple anonymized utterances and observe consistent performance improvements as more speech becomes available. We then incorporate prosodic and linguistic information, showing that multimodal systems outperform unimodal approaches. Finally, we compare aggregation strategies and find that frame-level aggregation yields the lowest EERs. Even with only five anonymized utterances, combining audio and text reduces EER by over 15% relative to audio-only aggregation, demonstrating that substantial speaker-discriminative information remains accessible despite anonymization.