ChatPaper.aiChatPaper

可聴外入力からモデル故障へ:LALMにおける低周波の安全性リスク

From Inaudible Inputs to Model Failures: Low-Frequency Safety Risks in LALMs

August 10, 2026
著者: Yuanhe Zhang, Weiliu Wang, Jie Ren, Liang Lin, Zhenhong Zhou, Haoran Gao, Kun Wang, Chen Li, Li Sun, Sen Su
cs.AI

要旨

大規模音声言語モデル(LALM)は、多様な音声入力の理解において強力な能力を示している。この多様性には、人間には聞こえないが、それでもモデルに入力され、その生成に影響を与え得る低周波信号が含まれる。しかしながら、このような低周波入力がLALMに及ぼす現実的な影響は、いまだほとんど調査されていない。本論文では、断続的低周波ロックアウト(ILL)を提案する。これは、ブラックボックス環境においてユニバーサル波形テンプレートを用いてこのリスクを評価する、不可聴なレッドチーミング手法である。ILLは、文注意スケール推定を用いてアクティブ区間を決定し、周波数混乱転移を用いて、コーパスのスペクトル変動から連続位相を持つ低周波波形を構築する。このリスクを軽減するため、分布再問い合わせガード(DRG)を提案し、低周波分布シフトを検出して、意味的復元のために条件付きで2回目の録音を要求する。6つのLALMと複数の音声理解タスクにわたって、ILLは最大67パーセントポイントの精度低下を引き起こす一方、平均人間可聴性評価は1.33と、クリーン音声の1.17に近い値を示した。DRGは、クリーンな再取得後、攻撃を受けた平均精度を28.5%から46.1%へ向上させた。これらの知見は、LALMに対するこれまで見落とされてきた安全性リスクを特定し、堅牢な音声理解に関する将来の研究の基盤を提供する。
English
Large audio-language models (LALMs) have demonstrated strong capabilities in understanding diverse audio inputs. This diversity includes low-frequency signals that are inaudible to humans but can still enter the model and influence its generation. However, the practical impact of such low-frequency inputs on LALMs remains largely unexplored. In this paper, we propose Intermittent Low-Frequency Lockout (ILL), an inaudible red teaming method that evaluates this risk using a universal waveform template in a black box setting. ILL uses Sentence Attention Scale Estimation to determine active intervals and Frequency Confusion Transfer to construct a low-frequency waveform with continuous phase from corpus spectral variation. To mitigate this risk, we propose Distributional Requery Guard (DRG) to detect low-frequency distribution shifts and conditionally request a second recording for semantic recovery. Across six LALMs and multiple audio understanding tasks, ILL reduces accuracy by up to 67 percentage points while receiving a mean human audibility rating of 1.33, close to 1.17 for clean audio; DRG raises mean attacked accuracy from 28.5\% to 46.1\% after clean reacquisition. These findings identify a previously overlooked safety risk for LALMs and provide a foundation for future research on robust audio understanding.