从不可听输入到模型失效:大型音频语言模型中的低频安全风险
From Inaudible Inputs to Model Failures: Low-Frequency Safety Risks in LALMs
August 10, 2026
作者: Yuanhe Zhang, Weiliu Wang, Jie Ren, Liang Lin, Zhenhong Zhou, Haoran Gao, Kun Wang, Chen Li, Li Sun, Sen Su
cs.AI
摘要
大型音频语言模型(LALMs)在理解多样化音频输入方面已展现出强大的能力。这种多样性包括人类听不见但仍能进入模型并影响其生成的低频信号。然而,此类低频输入对 LALMs 的实际影响在很大程度上仍未被探索。本文提出了一种不可听红队方法——间歇性低频锁定(ILL),用于在黑盒设置中利用通用波形模板评估这一风险。ILL 使用句子注意力尺度估计来确定活跃区间,并通过频率混淆迁移从语料谱变异中构建具有连续相位的低频波形。为缓解这一风险,我们提出了分布性重新查询防护(DRG),用于检测低频分布偏移,并有条件地请求第二次录音以进行语义恢复。在六个 LALMs 和多项音频理解任务中,ILL 将准确率降低了高达 67 个百分点,同时其平均人类可听度评分为 1.33,接近干净音频的 1.17;DRG 在干净音频重新获取后将平均受攻击准确率从 28.5\% 提升至 46.1\%。这些发现揭示了 LALMs 此前被忽视的安全风险,并为未来鲁棒音频理解研究提供了基础。
English
Large audio-language models (LALMs) have demonstrated strong capabilities in understanding diverse audio inputs. This diversity includes low-frequency signals that are inaudible to humans but can still enter the model and influence its generation. However, the practical impact of such low-frequency inputs on LALMs remains largely unexplored. In this paper, we propose Intermittent Low-Frequency Lockout (ILL), an inaudible red teaming method that evaluates this risk using a universal waveform template in a black box setting. ILL uses Sentence Attention Scale Estimation to determine active intervals and Frequency Confusion Transfer to construct a low-frequency waveform with continuous phase from corpus spectral variation. To mitigate this risk, we propose Distributional Requery Guard (DRG) to detect low-frequency distribution shifts and conditionally request a second recording for semantic recovery. Across six LALMs and multiple audio understanding tasks, ILL reduces accuracy by up to 67 percentage points while receiving a mean human audibility rating of 1.33, close to 1.17 for clean audio; DRG raises mean attacked accuracy from 28.5\% to 46.1\% after clean reacquisition. These findings identify a previously overlooked safety risk for LALMs and provide a foundation for future research on robust audio understanding.