SALAD-Bench: A Hierarchical and Comprehensive Safety Benchmark for Large Language Models

Lijun Li, Bowen Dong, Ruohui Wang, Xuhao Hu, Wangmeng Zuo, Dahua Lin, Yu Qiao, Jing Shao

arXiv:2402.05044 · 2026-07-27 공개 · arXiv · PDF

llm-evaluation large-language-models llm-attacks defense-methods attack-defense hierarchical-taxonomy safety-benchmark md-judge

Abstract

In the rapidly evolving landscape of Large Language Models (LLMs), ensuring robust safety measures is paramount. To meet this crucial need, we propose \emph{SALAD-Bench}, a safety benchmark specifically designed for evaluating LLMs, attack, and defense methods. Distinguished by its breadth, SALAD-Bench transcends conventional benchmarks through its large scale, rich diversity, intricate taxonomy spanning three levels, and versatile functionalities.SALAD-Bench is crafted with a meticulous array of questions, from standard queries to complex ones enriched with attack, defense modifications and multiple-choice. To effectively manage the inherent complexity, we introduce an innovative evaluators: the LLM-based MD-Judge for QA pairs with a particular focus on attack-enhanced queries, ensuring a seamless, and reliable evaluation. Above components extend SALAD-Bench from standard LLM safety evaluation to both LLM attack and defense methods evaluation, ensuring the joint-purpose utility. Our extensive experiments shed light on the resilience of LLMs against emerging threats and the efficacy of contemporary defense tactics. Data and evaluator are released under https://github.com/OpenSafetyLab/SALAD-BENCH.

한국어 요약

한 줄 요약

SALAD-Bench는 LLM의 안전성, 공격, 방어를 종합적으로 평가하기 위한 계층적 벤치마크로, MD-Judge와 MCQ-Judge 평가자와 함께 제공된다.

핵심 기여도

핵심 아이디어

기존 LLM 안전성 평가 벤치마크는 대부분 특정 유형의 위험(예: 독성 표현, 부적절한 지시)에만 초점을 맞추며, 공격 및 방어 메커니즘을 종합적으로 평가하지 못했다. SALAD-Bench는 이러한 한계를 극복하기 위해 3단계 계층적 분류 체계를 도입하고, 공격 강화 질문과 MCQ를 포함해 평가의 범위와 난이도를 확장했다. 특히, MD-Judge는 공격 강화 질문에 특화된 평가 모델로, 텍스트 쌍(QA)을 분류하며, MCQ-Judge는 정규 표현식을 활용해 MCQ 평가를 효율적으로 수행한다. 이는 기존 키워드 기반 또는 GPT 기반 평가 방식보다 정확도와 속도에서 우수한 성능을 보인다.

기술적 접근법

주요 결과

의의 및 한계

SALAD-Bench는 LLM의 안전성, 공격, 방어를 종합적으로 평가할 수 있는 체계적인 도구로, 기존 벤치마크의 한계를 극복한 점에서 학술적·실용적 가치가 높다. 특히, MD-Judge와 MCQ-Judge는 기존 평가 방식보다 정확도와 효율성을 동시에 담보하며, 다양한 LLM을 평가할 수 있는 확장성을 제공한다. 그러나 일부 모델(예: Gemini)은 공격 강화 세트에서 안전률이 급격히 하락하며, 방어 방법의 효과성도 모델에 따라 크게 달라지는 한계가 있다. 또한, 일부 공격 방법(예: GCG)은 방어에 취약하며, 이는 방어 전략 개선의 필요성을 시사한다.

실용적 활용

SALAD-Bench는 LLM 개발자들이 모델의 안전성, 공격에 대한 저항력, 방어 전략의 효과성을 체계적으로 평가하는 데 활용할 수 있다. 특히, 공공 부문, 보안 연구, AI 윤리 연구 등에서 모델의 신뢰성과 안정성을 검증하는 데 유용하게 사용될 수 있다.