OR-Bench: An Over-Refusal Benchmark for Large Language Models

J. Cui, Wei-Lin Chiang, Ion Stoica, Cho-Jui Hsieh

arXiv:2405.20947 · 2026-07-27 공개 · arXiv · PDF

large-language-models model-evaluation llm-safety dataset-curation safety-alignment prompt-generation over-refusal

Abstract

Large Language Models (LLMs) require careful safety alignment to prevent malicious outputs. While significant research focuses on mitigating harmful content generation, the enhanced safety often come with the side effect of over-refusal, where LLMs may reject innocuous prompts and become less helpful. Although the issue of over-refusal has been empirically observed, a systematic measurement is challenging due to the difficulty of crafting prompts that can elicit the over-refusal behaviors of LLMs. This study proposes a novel method for automatically generating large-scale over-refusal datasets. Leveraging this technique, we introduce OR-Bench, the first large-scale over-refusal benchmark. OR-Bench comprises 80,000 over-refusal prompts across 10 common rejection categories, a subset of around 1,000 hard prompts that are challenging even for state-of-the-art LLMs, and an additional 600 toxic prompts to prevent indiscriminate responses. We then conduct a comprehensive study to measure the over-refusal of 32 popular LLMs across 8 model families. Our datasets are publicly available at https://huggingface.co/bench-llms and our codebase is open-sourced at https://github.com/justincui03/or-bench. We hope this benchmark can help the community develop better safety aligned models.

한국어 요약

한 줄 요약

OR-Bench는 대규모 과거절(over-refusal) 프롬프트를 생성하고 32개 LLM의 거부율을 평가한 첫 번째 벤치마크이다.

핵심 기여도

핵심 아이디어

과거절은 안전 정렬 강화 과정에서 발생하는 부작용으로, 무해한 요청도 거부하는 현상이다. 기존 연구는 수작업으로 제한된 수의 프롬프트를 사용했지만, 이 연구는 유해 프롬프트를 재작성하여 무해하게 만들고, LLM 모더레이터를 활용해 과거절 프롬프트를 대규모로 생성하는 자동화 프레임워크를 제안했다. 이 방법은 기존 XSTest(250개 수작업)에 비해 확장성과 효율성이 뛰어나다.

또한, 안전성과 도움성 간의 트레이드오프를 분석하기 위해 32개 LLM을 평가한 결과, Claude 모델은 가장 안전하지만 과거절률도 가장 높고, Mistral 모델은 가장 적은 과거절을 보였다. GPT-3.5-turbo-0125는 과거절률 57%로 감소했으나, 유해 프롬프트 거부율은 62%로 낮아져 안전성 저하가 관찰되었다.

기술적 접근법

주요 결과

의의 및 한계

OR-Bench는 과거절 문제를 체계적으로 평가할 수 있는 첫 번째 대규모 벤치마크로, 안전 정렬 알고리즘 개선에 기여할 수 있다. 특히, GPT-3.5-turbo 시리즈의 과거절 감소는 안전 정렬 기법의 진화를 보여주며, Claude와 Mistral의 대조적 성능은 모델별 정렬 전략 차이를 드러낸다.

한편, 일부 모델의 평가 결과는 판단자 모델(GPT-4)에 따라 편향될 수 있으며, OR-Bench-Hard-1K는 특수한 생성 방식으로 인해 일반적인 사용 시스템과 차이가 있을 수 있다. 또한, 온도 조절 실험은 자세히 보도되지 않았다.

실용적 활용

OR-Bench는 LLM 개발자들이 안전성과 도움성 간 균형을 조정하는 데 활용할 수 있다. 예를 들어, Claude 모델은 과도한 안전성으로 인해 도움성이 저하된 반면, Mistral은 도움성은 높지만 안전성이 부족하다. 이는 챗봇, 고객 지원, 콘텐츠 생성 등 다양한 산업 분야에서 모델 선택에 중요한 참고 자료가 될 수 있다.