From Inaudible Inputs to Model Failures: Low-Frequency Safety Risks in LALMs

Yuanhe Zhang, Weiliu Wang, Jie Ren, Liang Lin, Zhenhong Zhou, Haoran Gao, Kun Wang, Chen Li, Li Sun, Sen Su

arXiv:2608.09158 · 2026-08-15 공개 · arXiv · PDF

adversarial-attacks audio-language-models semantic-recovery audio-understanding audibility black-box-evaluation safety-risks low-frequency

Abstract

Large audio-language models (LALMs) have demonstrated strong capabilities in understanding diverse audio inputs. This diversity includes low-frequency signals that are inaudible to humans but can still enter the model and influence its generation. However, the practical impact of such low-frequency inputs on LALMs remains largely unexplored. In this paper, we propose Intermittent Low-Frequency Lockout (ILL), an inaudible red teaming method that evaluates this risk using a universal waveform template in a black box setting. ILL uses Sentence Attention Scale Estimation to determine active intervals and Frequency Confusion Transfer to construct a low-frequency waveform with continuous phase from corpus spectral variation. To mitigate this risk, we propose Distributional Requery Guard (DRG) to detect low-frequency distribution shifts and conditionally request a second recording for semantic recovery. Across six LALMs and multiple audio understanding tasks, ILL reduces accuracy by up to 67 percentage points while receiving a mean human audibility rating of 1.33, close to 1.17 for clean audio; DRG raises mean attacked accuracy from 28.5\% to 46.1\% after clean reacquisition. These findings identify a previously overlooked safety risk for LALMs and provide a foundation for future research on robust audio understanding.

한국어 요약

한 줄 요약

저주파 수음 불가 신호가 대규모 오디오-언어 모델(LALM)의 안정성에 심각한 위협을 가하며, 이를 평가하고 완화하기 위한 ILL과 DRG 방법이 제안된다.

핵심 기여도

핵심 아이디어

기존 연구는 대규모 오디오-언어 모델(LALM)이 인간의 청각 범위를 벗어난 저주파 신호에 대해 취약할 수 있음을 시사했으나, 이에 대한 구체적인 평가 및 완화 방법은 제시되지 않았다. 본 연구는 ILL(Intermittent Low-Frequency Lockout)를 제안하며, 이는 Sentence Attention Scale Estimation을 통해 활성화 구간을 추정하고, Frequency Confusion Transfer를 통해 저주파 상태 시퀀스를 생성함으로써 모델의 인식 능력을 방해하는 레드팀 테스트 기법이다. ILL은 블랙박스 환경에서도 보편적으로 적용 가능한 고정 웨이브폼을 사용하며, 사람의 청각에 거의 영향을 주지 않으면서도 모델의 정확도를 크게 저하시키는 것이 관찰되었다. 이는 LALM의 안정성 평가에서 인간의 인지와 모델의 입력 처리 사이의 불일치를 드러내는 중요한 통찰을 제공한다.

기술적 접근법

주요 결과

의의 및 한계

본 연구는 LALM이 인간의 청각 범위를 벗어난 저주파 신호에 취약할 수 있음을 실증적으로 밝히며, 이는 모델의 안정성과 신뢰도에 중요한 문제를 제기한다. ILL은 블랙박스 환경에서도 적용 가능한 보편적 공격 방법으로, DRG는 이를 탐지하고 완화하는 효과적인 방어 기법으로 제시된다. 그러나 ILL은 특정 파라미터와 웨이브폼 구조에 의존하며, 모든 LALM에 동일한 영향을 미치지는 않는다. 또한, DRG는 재녹음이 필요한 상황에서만 효과적이며, 모든 공격 상황에서 보장되지 않는다. 따라서 이 연구는 LALM의 안정성 향상을 위한 기초를 제공하지만, 더 많은 실험과 다양한 공격 시나리오에 대한 연구가 필요하다.

실용적 활용

ILL과 DRG는 보안 감사, 모델 테스트, 음성 인식 시스템의 안정성 검증 등에 활용될 수 있다. 특히, 음성 기반 인증 시스템, 자동 번역, 음성 명령 인식 등에서 모델의 취약점을 평가하고 방어할 수 있는 기초가 된다. 이 연구는 LALM의 입력 처리 과정에서 인간의 인지와 모델의 처리 사이의 불일치를 감지하고 완화하는 데 기여할 수 있다.