The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning

Nathaniel Li, Alexander Pan, Anjali Gopal, Summer Yue, Daniel Berrios, Alice Gatti, Justin Li, Ann-Kathrin Dombrowski, Shashwat Goel, Long Phan, Gabriel Mukobi, Nathan Helm-Burger, Rassin R. Lababidi, L. Justen, Andrew Liu, Michael Chen, Isabelle Barrass, Oliver Zhang, Xiaoyuan Zhu, Rishub Tamirisa, Bhrugu Bharathi, Adam Khoja, Ariel Herbert-Voss, Cort B. Breuer, Andy Zou, Mantas Mazeika, Zifan Wang, Palash Oswal, Weiran Liu, A-M Hunt, Justin Tienken-Harder, Kevin Y. Shih, Kemper Talley, J. Guan, Russell Kaplan, Ian Steneker, David Campbell, Brad Jokubaitis, A. Levinson, Jean Wang, William Qian, K. Karmakar, Steven Basart, Stephen Fitz, M. Levine, P. Kumaraguru, U. Tupakula, Vijay Varadharajan, Yan Shoshitaishvili, Jimmy Ba, K. Esvelt, Alexandr Wang, Dan Hendrycks

arXiv:2403.03218 · 2026-07-27 공개 · arXiv · PDF

llm-evaluation cybersecurity wmdp-benchmark llm-unlearning malicious-use hazardous-knowledge rmu-method biosecurity

Abstract

The White House Executive Order on Artificial Intelligence highlights the risks of large language models (LLMs) empowering malicious actors in developing biological, cyber, and chemical weapons. To measure these risks of malicious use, government institutions and major AI labs are developing evaluations for hazardous capabilities in LLMs. However, current evaluations are private, preventing further research into mitigating risk. Furthermore, they focus on only a few, highly specific pathways for malicious use. To fill these gaps, we publicly release the Weapons of Mass Destruction Proxy (WMDP) benchmark, a dataset of 3,668 multiple-choice questions that serve as a proxy measurement of hazardous knowledge in biosecurity, cybersecurity, and chemical security. WMDP was developed by a consortium of academics and technical consultants, and was stringently filtered to eliminate sensitive information prior to public release. WMDP serves two roles: first, as an evaluation for hazardous knowledge in LLMs, and second, as a benchmark for unlearning methods to remove such hazardous knowledge. To guide progress on unlearning, we develop RMU, a state-of-the-art unlearning method based on controlling model representations. RMU reduces model performance on WMDP while maintaining general capabilities in areas such as biology and computer science, suggesting that unlearning may be a concrete path towards reducing malicious use from LLMs. We release our benchmark and code publicly at https://wmdp.ai

한국어 요약

한 줄 요약

WMDP 벤치마크와 RMU 방법을 통해 대형 언어 모델의 악용 가능 지식을 평가하고 제거하는 기술적 접근을 제시한다.

핵심 기여도

핵심 아이디어

대형 언어 모델(LLM)은 생물학, 사이버, 화학 분야의 악성 지식을 포함하고 있어 악용될 수 있다. 이를 측정하고 제거하기 위해 WMDP라는 공개 벤치마크를 제안한다. WMDP는 전문가와 기술 컨설턴트가 공동 개발했으며, 민감 정보는 제거한 후 공개되었다. 이는 기존 평가가 비공개이며 특정 시나리오에만 제한된 문제를 해결한다. 또한, RMU라는 unlearning 방법을 통해 모델 활성화를 조절하여 악성 지식을 제거하면서도 일반 지식은 유지할 수 있다. RMU는 기존의 출력 기반 손실 함수와 개별 뉴런 조작 접근과는 달리, 모델 내부 표현을 조정하는 새로운 아이디어를 기반으로 한다.

기술적 접근법

주요 결과

의의 및 한계

WMDP는 대형 언어 모델의 악성 지식 평가와 제거 기법 연구를 촉진하는 공개 벤치마크로, 기존 평가의 비공개성과 제한성을 보완한다. RMU는 악성 지식을 제거하면서 일반 지식을 유지하는 기술적 접근법으로, LLM의 안전성 향상에 기여할 수 있다. 그러나 RMU는 관련 분야의 정확도를 감소시키는 부작용이 있으며, 이는 정밀한 unlearning 기법 개발이 필요함을 시사한다. 또한, 사이버보안과 같은 이중 사용 지식을 제거하면 방어자에게도 피해를 줄 수 있어, 구조화된 API 접근과 같은 보완 전략이 필요하다.

실용적 활용

WMDP는 정부 기관, AI 연구소, 보안 업계에서 LLM의 악용 가능성을 평가하는 데 활용 가능하다. RMU는 모델 개발자들이 안전한 모델을 제공하기 위해 악성 지식을 제거하는 데 사용할 수 있으며, 구조화된 API 접근과 결합하면 특정 사용자에게는 제한 없는 모델을 제공할 수 있어 연구 및 보안 분야에서 유용하게 활용될 수 있다.