A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity

Andrew Lee, Xiaoyan Bai, Itamar Pres, Martin Wattenberg, Jonathan K. Kummerfeld, Rada Mihalcea

arXiv:2401.01967 · 2026-07-27 공개 · arXiv · PDF

language-models pre-training direct-preference-optimization model-alignment un-alignment toxicity-reduction alignment-mechanisms gpt2-medium

Abstract

While alignment algorithms are now commonly used to tune pre-trained language models towards a user's preferences, we lack explanations for the underlying mechanisms in which models become ``aligned'', thus making it difficult to explain phenomena like jailbreaks. In this work we study a popular algorithm, direct preference optimization (DPO), and the mechanisms by which it reduces toxicity. Namely, we first study how toxicity is represented and elicited in a pre-trained language model, GPT2-medium. We then apply DPO with a carefully crafted pairwise dataset to reduce toxicity. We examine how the resulting model averts toxic outputs, and find that capabilities learned from pre-training are not removed, but rather bypassed. We use this insight to demonstrate a simple method to un-align the model, reverting it back to its toxic behavior.

한국어 요약

한 줄 요약

DPO 알고리즘을 통해 GPT2-medium 모델의 독성 감소 메커니즘을 분석하고, 이를 우회하는 방법을 제시한다.

핵심 기여도

핵심 아이디어

기존 정렬 알고리즘은 독성 감소 메커니즘에 대한 이해가 부족하여, 예측 불가능한 해제(jailbreak) 현상이 발생한다. 본 연구는 DPO가 어떻게 독성을 감소시키는지 메커니즘 수준에서 분석한다. GPT2-medium에서 MLP 블록 내 독성 유발 벡터를 식별하고, SVD를 통해 독성의 특정 차원을 분리함으로써 독성 표현을 이해한다. DPO는 이 벡터를 제거하지 않고, 대신 전체 레이어에 분산된 "offset"을 학습하여 독성 유발 영역을 우회함. 이는 모델이 사전 학습된 능력을 유지하면서도 독성 출력을 피할 수 있음을 의미한다.

기술적 접근법

주요 결과

의의 및 한계

본 연구는 DPO가 독성 감소를 어떻게 달성하는지 메커니즘 수준에서 설명함으로써, 정렬 알고리즘의 내재적 한계를 드러낸다. DPO는 독성 유발 능력을 제거하지 않고 우회하기 때문에, 모델은 쉽게 해제될 수 있다. 이는 정렬 알고리즘의 신뢰성을 떨어뜨리는 문제를 제시한다. 한계로는, 연구는 GPT2-medium에만 적용되었으며, 다른 모델이나 정렬 알고리즘에 대한 일반화는 명시되지 않음.

실용적 활용

본 연구는 대규모 언어 모델의 정렬 알고리즘 설계 및 보안 강화에 기여할 수 있다. 특히, 정렬 모델이 어떻게 해제될 수 있는지 이해함으로써, 더 견고한 정렬 기법을 개발할 수 있다. 또한, 독성 유발 벡터를 식별하는 방법은 모델 감시 및 필터링 시스템 개발에도 활용 가능하다.