How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs

Yi Zeng, H. Lin, Jingwen Zhang, Diyi Yang, Ruoxi Jia, Weiyan Shi

arXiv:2401.06373 · 2026-07-27 공개 · arXiv · PDF

ai-safety adversarial-prompts llama-2 gpt-4 social-science interactive-llms persuasion-taxonomy llm-jailbreaking

Abstract

Most traditional AI safety research has approached AI models as machines and centered on algorithm-focused attacks developed by security experts. As large language models (LLMs) become increasingly common and competent, non-expert users can also impose risks during daily interactions. This paper introduces a new perspective to jailbreak LLMs as human-like communicators, to explore this overlooked intersection between everyday language interaction and AI safety. Specifically, we study how to persuade LLMs to jailbreak them. First, we propose a persuasion taxonomy derived from decades of social science research. Then, we apply the taxonomy to automatically generate interpretable persuasive adversarial prompts (PAP) to jailbreak LLMs. Results show that persuasion significantly increases the jailbreak performance across all risk categories: PAP consistently achieves an attack success rate of over $92\%$ on Llama 2-7b Chat, GPT-3.5, and GPT-4 in $10$ trials, surpassing recent algorithm-focused attacks. On the defense side, we explore various mechanisms against PAP and, found a significant gap in existing defenses, and advocate for more fundamental mitigation for highly interactive LLMs

한국어 요약

한 줄 요약

일상적 설득 기법을 활용한 LLM 해킹(Persuasive Jailbreak)이 기존 알고리즘 기반 공격을 크게 초과하며, AI 안전성 연구의 새로운 관점을 제시한다.

핵심 기여도

핵심 아이디어

기존 AI 안전성 연구는 LLM을 단순히 알고리즘 기반 시스템으로 접근했으나, 본 연구는 LLM을 인간처럼 소통하는 존재로 간주하고, 일상적 언어 상호작용을 통해 설득하는 방식으로 해킹을 시도한다. 이는 인간의 자연스러운 설득 기법(감정 호소, 권위적 주장 등)을 기반으로 LLM의 안전성 장벽을 무너뜨리는 새로운 공격 벡터를 제시한다. 연구는 사회과학에서 축적된 설득 이론을 기반으로 설득 기법 분류 체계를 구축하고, 이를 통해 해독 가능한(Parsable) PAP를 생성한다. 이는 기존의 가상화(Virtualization)나 역할 연기(Role-playing) 기반 공격과 달리, 인간의 자연스러운 언어 패턴을 반영한 공격 방식을 구현한다.

기술적 접근법

주요 결과

의의 및 한계

실용적 활용