Robust Prompt Optimization for Defending Language Models Against Jailbreaking Attacks

Andy Zhou, Bo Li, Haohan Wang

arXiv:2401.17263 · 2026-07-27 공개 · arXiv · PDF

robustness adversarial-attacks prompt-optimization attack-success-rate llm-defense jailbreaking jailbreakbench rpo

Abstract

Despite advances in AI alignment, large language models (LLMs) remain vulnerable to adversarial attacks or jailbreaking, in which adversaries can modify prompts to induce unwanted behavior. While some defenses have been proposed, they have not been adapted to newly proposed attacks and more challenging threat models. To address this, we propose an optimization-based objective for defending LLMs against jailbreaking attacks and an algorithm, Robust Prompt Optimization (RPO) to create robust system-level defenses. Our approach directly incorporates the adversary into the defensive objective and optimizes a lightweight and transferable suffix, enabling RPO to adapt to worst-case adaptive attacks. Our theoretical and experimental results show improved robustness to both jailbreaks seen during optimization and unknown jailbreaks, reducing the attack success rate (ASR) on GPT-4 to 6% and Llama-2 to 0% on JailbreakBench, setting the state-of-the-art. Code can be found at https://github.com/lapisrocks/rpo

한국어 요약

한 줄 요약

RPO는 적대적 프롬프트 공격에 대응하기 위해 최적화 기반 방어 알고리즘으로, GPT-4와 Llama-2에서 ASR을 각각 6%와 0%까지 감소시킨다.

핵심 기여도

핵심 아이디어

기존 방어 방법은 특정 공격에만 효과적이거나, 새로운 공격에 대응하지 못하는 한계가 있었다. 본 연구는 **적대자를 방어 목표에 직접 통합**하는 새로운 접근법을 제안한다. 이는 **minimax 최적화 문제**로 모델링되며, 방어 대상이 되는 LLM에 **경량화된 suffix**를 추가하여 입력을 조정함으로써 공격에 대한 강건성을 높인다. RPO는 **공격 선택**과 **이산 최적화**를 결합하여, 다양한 공격 시나리오에 유연하게 대응할 수 있도록 설계되었다.

기술적 접근법

주요 결과

의의 및 한계

RPO는 **시스템 레벨 방어**(system-level guardrails)를 통해 LLM의 안전성을 향상시키는 실용적 접근법을 제시한다. 특히, **적응적 공격**에 대한 강건성과 **경량화된 방어 구조**는 기존 방어 방법과 차별화된다. 그러나 RPO는 특정 **suffix 기반 방어**이므로, 공격자가 이 suffix를 우회하는 새로운 공격 기법이 등장할 경우 한계가 있을 수 있다. 또한, **모델 내부 구조에 대한 접근이 필요하지 않기 때문에**, 일부 시스템에서의 적용 가능성은 제한적일 수 있다.

실용적 활용

RPO는 **LLM 기반 서비스**(예: 챗봇, 콘텐츠 생성 플랫폼)에서 **실시간 방어**를 제공할 수 있으며, 특히 **블랙박스 모델**(예: GPT-4)에도 적용 가능하다. **시스템 레벨 방어**로 설계되어, 사용자에게 노출되지 않으면서도 안전성을 유지할 수 있어, **보안 강화** 및 **정책 준수**를 위한 실용적 도구로 활용 가능하다.