비가청 입력에서 모델 실패까지: LALM의 저주파 안전 위험
From Inaudible Inputs to Model Failures: Low-Frequency Safety Risks in LALMs
August 10, 2026
저자: Yuanhe Zhang, Weiliu Wang, Jie Ren, Liang Lin, Zhenhong Zhou, Haoran Gao, Kun Wang, Chen Li, Li Sun, Sen Su
cs.AI
초록
대규모 오디오-언어 모델(LALM)은 다양한 오디오 입력을 이해하는 데 강력한 성능을 입증해 왔다. 이러한 다양성에는 인간이 들을 수는 없지만 모델에 입력되어 생성 과정에 영향을 미칠 수 있는 저주파 신호도 포함된다. 그러나 이러한 저주파 입력이 LALM에 미치는 실질적 영향은 대체로 탐구되지 않은 상태이다. 본 논문에서는 블랙박스 환경에서 보편적 파형 템플릿을 사용하여 이러한 위험을 평가하는 비가청 레드 팀 공격 기법인 간헐적 저주파 차단(ILL)을 제안한다. ILL은 문장 주의 규모 추정(Sentence Attention Scale Estimation)을 사용하여 활성 구간을 결정하고, 주파수 혼동 전이(Frequency Confusion Transfer)를 통해 말뭉치 스펙트럼 변이로부터 연속 위상을 갖는 저주파 파형을 구성한다. 이러한 위험을 완화하기 위해 분포 기반 재요청 보호(DRG)를 제안하여 저주파 분포 변화를 탐지하고 조건부로 두 번째 녹음을 요청하여 의미를 복구한다. 여섯 개의 LALM과 다양한 오디오 이해 과제에서 ILL은 최대 67퍼센트 포인트까지 정확도를 감소시켰으며, 인간 청취 가능성 평균 등급은 1.33으로 청정 오디오의 1.17에 근접하였다. DRG는 청정 재획득 후 공격받은 평균 정확도를 28.5%에서 46.1%로 향상시켰다. 이러한 발견은 LALM에 대한 기존에 간과되었던 안전 위험을 식별하고, 견고한 오디오 이해를 위한 향후 연구의 기초를 제공한다.
English
Large audio-language models (LALMs) have demonstrated strong capabilities in understanding diverse audio inputs. This diversity includes low-frequency signals that are inaudible to humans but can still enter the model and influence its generation. However, the practical impact of such low-frequency inputs on LALMs remains largely unexplored. In this paper, we propose Intermittent Low-Frequency Lockout (ILL), an inaudible red teaming method that evaluates this risk using a universal waveform template in a black box setting. ILL uses Sentence Attention Scale Estimation to determine active intervals and Frequency Confusion Transfer to construct a low-frequency waveform with continuous phase from corpus spectral variation. To mitigate this risk, we propose Distributional Requery Guard (DRG) to detect low-frequency distribution shifts and conditionally request a second recording for semantic recovery. Across six LALMs and multiple audio understanding tasks, ILL reduces accuracy by up to 67 percentage points while receiving a mean human audibility rating of 1.33, close to 1.17 for clean audio; DRG raises mean attacked accuracy from 28.5\% to 46.1\% after clean reacquisition. These findings identify a previously overlooked safety risk for LALMs and provide a foundation for future research on robust audio understanding.