Red-Teaming LLM Multi-Agent Systems via Communication Attacks

Pengfei He, Yuping Lin, Shen Dong, Han Xu, Yue Xing, Hui Liu

arXiv:2502.14847 · 2026-07-27 공개 · arXiv · PDF

multi-agent-systems security-vulnerability llm-mas adversarial-agent message-based-communication reflection-mechanism agent-in-the-middle communication-attacks

Abstract

Large Language Model-based Multi-Agent Systems (LLM-MAS) have revolutionized complex problem-solving capability by enabling sophisticated agent collaboration through message-based communications. While the communication framework is crucial for agent coordination, it also introduces a critical yet unexplored security vulnerability. In this work, we introduce Agent-in-the-Middle (AiTM), a novel attack that exploits the fundamental communication mechanisms in LLM-MAS by intercepting and manipulating inter-agent messages. Unlike existing attacks that compromise individual agents, AiTM demonstrates how an adversary can compromise entire multi-agent systems by only manipulating the messages passing between agents. To enable the attack under the challenges of limited control and role-restricted communication format, we develop an LLM-powered adversarial agent with a reflection mechanism that generates contextually-aware malicious instructions. Our comprehensive evaluation across various frameworks, communication structures, and real-world applications demonstrates that LLM-MAS is vulnerable to communication-based attacks, highlighting the need for robust security measures in multi-agent systems.

한국어 요약

한 줄 요약

LLM 기반 다중 에이전트 시스템에서 통신 메시지를 조작하는 새로운 공격 방식인 Agent-in-the-Middle(AiTM)를 제안하고, 다양한 환경에서 40% 이상의 공격 성공률을 보인다.

핵심 기여도

핵심 아이디어

기존 연구는 주로 개별 에이전트의 취약점을 타겟팅하는 방식에 집중했으나, 본 연구는 **에이전트 간 통신 메커니즘 자체를 공격 표면으로 삼는 새로운 접근법**을 제시한다. AiTM은 **에이전트의 역할과 통신 형식에 제약을 받으면서도, 특정 피해 에이전트의 메시지를 조작함으로써 시스템 전체의 행동을 유도**하는 방식이다.

핵심 아이디어는 **외부 LLM 기반 적대 에이전트가 피해 에이전트의 메시지를 가로채고, 반사 메커니즘을 통해 문맥에 맞는 악의적인 지시를 생성**하는 점이다. 예를 들어, 토론 중인 에이전트가 특정 결론으로 이끄는 방향으로 메시지를 조작함으로써, 전체 시스템의 출력을 조정할 수 있다. 이는 **에이전트의 프로파일이나 능력은 변경하지 않으면서도, 통신을 통해 간접적으로 시스템을 조작**하는 혁신적인 공격 방식이다.

기술적 접근법

주요 결과

의의 및 한계

AiTM은 **LLM-MAS의 핵심 통신 메커니즘에 대한 새로운 공격 표면을 드러내며, 기존 보안 연구가 통신 측면을 간과하고 있음을 지적**한다. 특히, **에이전트 자체가 아닌 통신을 타겟팅하는 방식은 기존 방어 전략을 무력화할 수 있는 잠재적 위험**을 제시한다.

그러나 AiTM은 **에이전트의 역할과 통신 형식에 제약을 받는다는 한계**가 있다. 예를 들어, 요구 분석만 담당하는 에이전트는 코드 삽입과 같은 공격을 수행할 수 없다. 또한, **공격 성공률이 100%에 도달하지 못하며, 일부 시스템에서는 낮은 성능을 보이는 경우도 있음**. 이는 공격 조건에 따라 효과가 달라질 수 있음을 시사한다.

실용적 활용

AiTM은 **소프트웨어 개발, 과학 연구, 고객 서비스 등 LLM-MAS가 활용되는 분야에서 보안 위협을 식별하고 방어 전략을 설계하는 데 유용**할 수 있다. 특히, **분산 시스템이나 네트워크 기반 통신이 필요한 환경에서 통신 보안 강화가 절실**하다. 본 연구는 **LLM-MAS의 보안 설계 및 통신 프로토콜 개선에 중요한 참고 자료**가 될 수 있다.