從不可聽輸入到模型失效:大型音頻語言模型中的低頻安全風險
From Inaudible Inputs to Model Failures: Low-Frequency Safety Risks in LALMs
August 10, 2026
作者: Yuanhe Zhang, Weiliu Wang, Jie Ren, Liang Lin, Zhenhong Zhou, Haoran Gao, Kun Wang, Chen Li, Li Sun, Sen Su
cs.AI
摘要
大型音訊語言模型(LALMs)已展現出理解多樣音訊輸入的強大能力。這種多樣性包含對人類不可聞的低頻訊號,這些訊號仍可進入模型並影響其生成。然而,此類低頻輸入對 LALMs 的實際影響仍未被充分探討。在本文中,我們提出間歇性低頻鎖定(Intermittent Low-Frequency Lockout, ILL),這是一種不可聞的紅隊測試方法,在黑箱設定下使用通用波形模板來評估此風險。ILL 使用句子注意力尺度估計(Sentence Attention Scale Estimation)來決定有效區間,並利用頻率混淆遷移(Frequency Confusion Transfer)從語料庫頻譜變化中建構具連續相位的低頻波形。為緩解此風險,我們提出分佈性重新查詢防護(Distributional Requery Guard, DRG),以偵測低頻分佈偏移,並有條件地要求第二次錄音以進行語意恢復。在六個 LALMs 與多個音訊理解任務中,ILL 使準確率降低高達 67 個百分點,同時其平均人類可聽度評分為 1.33,接近乾淨音訊的 1.17;DRG 在乾淨重新獲取後,將平均受攻擊準確率從 28.5% 提升至 46.1%。這些發現指出了先前被忽略的 LALMs 安全風險,並為未來穩健音訊理解研究奠定基礎。
English
Large audio-language models (LALMs) have demonstrated strong capabilities in understanding diverse audio inputs. This diversity includes low-frequency signals that are inaudible to humans but can still enter the model and influence its generation. However, the practical impact of such low-frequency inputs on LALMs remains largely unexplored. In this paper, we propose Intermittent Low-Frequency Lockout (ILL), an inaudible red teaming method that evaluates this risk using a universal waveform template in a black box setting. ILL uses Sentence Attention Scale Estimation to determine active intervals and Frequency Confusion Transfer to construct a low-frequency waveform with continuous phase from corpus spectral variation. To mitigate this risk, we propose Distributional Requery Guard (DRG) to detect low-frequency distribution shifts and conditionally request a second recording for semantic recovery. Across six LALMs and multiple audio understanding tasks, ILL reduces accuracy by up to 67 percentage points while receiving a mean human audibility rating of 1.33, close to 1.17 for clean audio; DRG raises mean attacked accuracy from 28.5\% to 46.1\% after clean reacquisition. These findings identify a previously overlooked safety risk for LALMs and provide a foundation for future research on robust audio understanding.