Preference Fine-Tuning of LLMs Should Leverage Suboptimal, On-Policy Data

Fahim Tajwar, Anikait Singh, Archit Sharma, Rafael Rafailov, Jeff G. Schneider, Tengyang Xie, S. Ermon, Chelsea Finn, Aviral Kumar

arXiv:2404.14367 · 2026-07-27 공개 · arXiv · PDF

reinforcement-learning llm-training contrastive-learning supervised-learning preference-fine-tuning on-policy-data negative-gradient mode-seeking-objectives

Abstract

Learning from preference labels plays a crucial role in fine-tuning large language models. There are several distinct approaches for preference fine-tuning, including supervised learning, on-policy reinforcement learning (RL), and contrastive learning. Different methods come with different implementation tradeoffs and performance differences, and existing empirical findings present different conclusions, for instance, some results show that online RL is quite important to attain good fine-tuning results, while others find (offline) contrastive or even purely supervised methods sufficient. This raises a natural question: what kind of approaches are important for fine-tuning with preference data and why? In this paper, we answer this question by performing a rigorous analysis of a number of fine-tuning techniques on didactic and full-scale LLM problems. Our main finding is that, in general, approaches that use on-policy sampling or attempt to push down the likelihood on certain responses (i.e., employ a"negative gradient") outperform offline and maximum likelihood objectives. We conceptualize our insights and unify methods that use on-policy sampling or negative gradient under a notion of mode-seeking objectives for categorical distributions. Mode-seeking objectives are able to alter probability mass on specific bins of a categorical distribution at a fast rate compared to maximum likelihood, allowing them to relocate masses across bins more effectively. Our analysis prescribes actionable insights for preference fine-tuning of LLMs and informs how data should be collected for maximal improvement.

한국어 요약

한 줄 요약

대규모 언어 모델의 선호도 미세조정에서 온정책 샘플링과 음의 경사를 활용하는 접근법이 성능을 향상시킨다.

핵심 기여도

핵심 아이디어

기존의 선호도 기반 미세조정 방법은 감독 학습, 대조 학습, 온정책 강화 학습 등 다양한 접근법이 존재하지만, 이들의 성능 차이는 데이터 커버리지와 목적함수의 성질에 따라 달라진다. 본 연구는 이 문제를 해결하기 위해 **모드 탐색**(mode-seeking) 목적함수라는 개념을 도입한다. 이는 특정 고보상 응답에 대한 확률 질량을 빠르게 재분배할 수 있는 성질을 가진다.

특히, **온정책 샘플링**(on-policy sampling)과 **음의 경사**(negative gradient)를 사용하는 방법은 **모드 커버링**(mode-covering) 목적함수(예: 최대 우도)와 달리, 특정 응답 집합에 확률 질량을 집중시키는 데 효과적이다. 예를 들어, **대조 학습**(contrastive learning)은 음의 경사를 내재적으로 사용하며, **역 KL 발산**(reverse KL divergence)은 모드 탐색적 성질을 가진다. 이는 **Pref-FT** 또는 **Binary Feed-ME**와 같은 순수 감독 학습 방법이 고보상 응답을 효과적으로 학습하지 못하는 이유를 설명한다.

기술적 접근법

주요 결과

의의 및 한계

본 연구는 선호도 기반 미세조정에서 **온정책 샘플링**과 **음의 경사**가 핵심적인 역할을 한다는 이론적 근거를 제시하며, **데이터 커버리지**와 **목적함수 선택**의 중요성을 강조한다. 이는 실제 적용 시 데이터 수집 전략과 학습 알고리즘 선택에 실질적인 가이드라인을 제공한다.

그러나, 본 연구는 **AlpacaFarm**과 **UltraFeedback** 데이터셋에 기반한 실험 결과를 바탕으로 일반화를 추론하고 있으므로, 다른 데이터셋에서 동일한 결과가 나타날지는 추가 연구가 필요하다. 또한, **온정책 샘플링**은 계산 비용이 높고, **정책 갱신** 주기와 **샘플링 빈도** 사이의 트레이드오프를 고려해야 한다는 한계가 있다.

실용적 활용

본 연구의 결과는 **대규모 언어 모델**을 인간 선호도에 맞게 조정하는 데 활용 가능하다. 예를 들어, **AI 챗봇**, **자동 번역**, **컨텐츠 생성** 등에서 **고품질 응답**을 유도하기 위해 **온정책 샘플링**과 **대조 학습**을 결합한 방법을 사용할 수 있다. 또한, **데이터 수집 전략**에서 **고보상 응답이 reference policy의 낮은 확률 영역에 있을 경우** 온정책 샘플링을 우선적으로 고려해야 한다는 점이 실용적 가치를 가진다.