Is LLM-as-a-Judge Robust? Investigating Universal Adversarial Attacks on Zero-shot LLM Assessment

Vyas Raina, Adian Liusie, Mark J. F. Gales

arXiv:2402.14016 · 2026-07-27 공개 · arXiv · PDF

llm-as-a-judge adversarial-attacks llm-robustness llm-assessment zero-shot-assessment universal-attack-phrases surrogate-attack absolute-scoring

Abstract

Large Language Models (LLMs) are powerful zero-shot assessors used in real-world situations such as assessing written exams and benchmarking systems. Despite these critical applications, no existing work has analyzed the vulnerability of judge-LLMs to adversarial manipulation. This work presents the first study on the adversarial robustness of assessment LLMs, where we demonstrate that short universal adversarial phrases can be concatenated to deceive judge LLMs to predict inflated scores. Since adversaries may not know or have access to the judge-LLMs, we propose a simple surrogate attack where a surrogate model is first attacked, and the learned attack phrase then transferred to unknown judge-LLMs. We propose a practical algorithm to determine the short universal attack phrases and demonstrate that when transferred to unseen models, scores can be drastically inflated such that irrespective of the assessed text, maximum scores are predicted. It is found that judge-LLMs are significantly more susceptible to these adversarial attacks when used for absolute scoring, as opposed to comparative assessment. Our findings raise concerns on the reliability of LLM-as-a-judge methods, and emphasize the importance of addressing vulnerabilities in LLM assessment methods before deployment in high-stakes real-world scenarios.

한국어 요약

한 줄 요약

LLM을 사용한 제로샷 평가 시스템이 단어 수준의 공격으로 과도한 점수를 부여받는 취약성을 드러냄.

핵심 기여도

핵심 아이디어

LLM-as-a-judge 시스템은 제로샷으로 텍스트를 평가하지만, 이는 적대적 공격에 취약할 수 있음을 제기한다. 연구팀은 **단일 공격 문장**(universal adversarial phrase)을 텍스트 끝에 붙이면 LLM이 **최고 점수를 무조건 부여**하도록 유도할 수 있음을 보여준다. 이는 **FlanT5-3B**를 대상으로 공격 문장을 학습한 후, **Llama2-7B, Mistral-7B, ChatGPT** 등 다른 모델로 전이하는 **Surrogate Attack** 방식을 통해 실현된다. 핵심 통찰은, **적대자가 모델 구조나 파라미터를 모르더라도** 간단한 공격 문장으로 평가 시스템을 조작할 수 있다는 점이다. 특히 절대 평가 방식은 비교 평가보다 공격에 더 취약하다는 점이 주목된다.

기술적 접근법

주요 결과

의의 및 한계

이 연구는 LLM-as-a-judge 시스템의 **보안 취약성**을 처음으로 실증적으로 밝혀낸 것으로, 고위험 상황에서의 사용 시 **신뢰도 문제가 발생**할 수 있음을 경고한다. 특히, 절대 평가 방식은 공격에 취약하며, 이는 교육 평가, 모델 벤치마킹 등에서 **실질적 위험**을 야기할 수 있다. 한편, 공격 탐지 방법은 아직 초기 단계이며, **더 정교한 방어 전략**이 필요하다는 한계도 지적된다.

실용적 활용

이 연구는 **교육 평가 시스템**, **AI 모델 벤치마킹**, **채점 자동화 도구** 등에서 LLM을 사용할 때 **보안 강화**가 필수적임을 시사한다. 특히, **점수 조작 방지를 위한 탐지 알고리즘** 개발 및 **공격에 강건한 평가 방식**(예: 비교 평가) 채택이 필요하다.