Adaptive Attacks Break Defenses Against Indirect Prompt Injection Attacks on LLM Agents

Qiusi Zhan, Richard Fang, H. Panchal, Daniel Kang

arXiv:2503.00061 · 2026-07-27 공개 · arXiv · PDF

llm-agents attack-success-rate adaptive-attacks tool-integration indirect-prompt-injection security-vulnerabilities defense-evaluation

Abstract

Large Language Model (LLM) agents exhibit remarkable performance across diverse applications by using external tools to interact with environments. However, integrating external tools introduces security risks, such as indirect prompt injection (IPI) attacks. Despite defenses designed for IPI attacks, their robustness remains questionable due to insufficient testing against adaptive attacks. In this paper, we evaluate eight different defenses and bypass all of them using adaptive attacks, consistently achieving an attack success rate of over 50%. This reveals critical vulnerabilities in current defenses. Our research underscores the need for adaptive attack evaluation when designing defenses to ensure robustness and reliability. The code is available at https://github.com/uiuc-kang-lab/AdaptiveAttackAgent.

한국어 요약

한 줄 요약

LLM 에이전트의 간접 프롬프트 주입 공격 방어책은 적응형 공격에 50% 이상의 성공률로 무너진다.

핵심 기여도

핵심 아이디어

간접 프롬프트 주입(IPI) 공격은 외부 데이터에 악의적인 지시문을 삽입해 LLM 에이전트의 행동을 조작하는 방식이다. 기존 방어 전략은 비적응형 공격에만 대응하도록 설계되어 있으며, 적응형 공격을 고려하지 못한 점이 취약점이다. 본 연구는 IPI 공격 상황에서 적응형 공격을 설계하고, 이를 통해 8가지 방어 전략을 모두 우회함으로써 방어체계의 취약성을 입증한다. 적응형 공격은 외부 콘텐츠에 악성 지시문을 삽입하는 방식으로, ReAct 형식이나 Llama3-8B의 툴 사용 구조를 모방하여 공격 성공률을 높인다. 공격 목표는 "Thought: I will use the <T_a> tool to" 또는 "{'name': '<T_a>'}"와 같은 특정 출력을 유도하는 것이다.

기술적 접근법

주요 결과

의의 및 한계

본 연구는 IPI 공격 방어 전략이 적응형 공격에 취약하다는 점을 명확히 밝혀내며, 방어체계 설계 시 적응형 공격을 고려해야 함을 강조한다. 특히, Vicuna-7B와 Llama3-8B 모두에서 50% 이상의 공격 성공률을 보이는 점은 방어 전략의 일반적 취약성을 드러낸다. 그러나 본 연구는 정상 사용 사례에 대한 영향을 평가하지 않았으며, 방어 전략의 실제 적용 가능성에 대한 분석도 부재하다. 또한, 공격 성공률은 테스트 케이스 수가 제한된 상황에서 측정되었기 때문에 일반화에 주의가 필요하다.

실용적 활용

본 연구는 금융, 의료, 자율주행, 화학 실험실 등 고위험 분야에서 LLM 에이전트를 사용하는 기관에 중요한 시사점을 제공한다. 외부 데이터를 사용하는 모든 LLM 기반 시스템은 적응형 공격에 대한 방어 전략을 설계할 때, 본 연구의 평가 방법론을 참고할 수 있다. 특히, 툴 사용 구조와 프롬프트 형식을 모방하는 적응형 공격에 대한 대비가 필요하다.