Preference Tuning as Spectral Update Reorganization

Peiyan Zhang, Haibo Jin, Liying Kang, Haohan Wang

arXiv:2607.20438 · 2026-07-25 공개 · arXiv · PDF

lora out-of-distribution rlhf model-alignment parameter-updates spectral-analysis preference-tuning update-reorganization

Abstract

Preference-based post-training is usually understood through endpoint behavior, yet the learned update that produces this behavior remains largely opaque. We study RLHF and related preference optimization through the spectral structure of their induced parameter updates. By decomposing effective LoRA updates and reloading their spectral components as plug-in modules, we turn preference-induced updates into objects that can be isolated, recomposed, and directly intervened on. Across model families, optimization algorithms, and supervision regimes, these updates consistently develop a spectral head--tail organization. A compact head emerges early and carries the dominant endpoint shift, while a heterogeneous residual tail remains. The split is functional rather than merely descriptive. Plug-in intervention shows that the head accounts for the visible behavioral departure from the base model, while the tail is weak in isolation. Cross-run recomposition further shows that mixed adapters follow the source of the head, indicating that the head carries run-level solver bias. This endpoint dominance does not imply learning sufficiency. Head-only learning is non-vacuous but fails to recover the full solution, especially on out-of-distribution behavior. Tail-only learning yields little visible gain, yet the full solution is not recovered without the tail. These findings recast preference post-training as structured update reorganization rather than a monolithic behavioral correction, and suggest that alignment gain and coverage loss are tied to how the learned update itself is organized.

한국어 요약

한 줄 요약

RLHF와 선호도 최적화를 스펙트럼 구조로 분석하여 학습된 업데이트의 재구성 가능성을 밝혀낸 연구.

핵심 기여도

핵심 아이디어

본 연구는 RLHF와 관련된 선호도 기반 후학습을 단순한 행동 변화가 아닌, 학습된 업데이트의 구조적 재조직으로 이해하고자 한다. 기존 연구는 주로 최종 모델의 출력을 기준으로 평가했지만, 이는 업데이트 내부 구조를 명확히 파악하지 못한다. 본 연구는 LoRA 업데이트를 모듈 단위로 SVD를 통해 분해하고, 이를 플러그인 모듈로 재조합함으로써 업데이트의 내부 구조를 분석한다. 이 과정에서 스펙트럼 헤드와 테일이라는 구조적 요소가 나타나며, 이는 엔드포인트 행동과 학습 완전성에 각각 다른 역할을 한다는 점이 밝혀졌다. 특히, 헤드는 주요 행동 변화를 담당하지만, 테일 없이는 완전한 학습이 불가능하다는 점이 강조된다.

기술적 접근법

주요 결과

의의 및 한계

본 연구는 선호도 기반 후학습을 단순한 행동 수정이 아닌, 구조적 업데이트 재조직으로 재해석함으로써, 모델 정렬과 커버리지 손실의 이해를 깊이 있게 확장한다. 특히, 업데이트 내부 구조를 분석함으로써 개입 가능성과 실험 조작성을 높이는 기술적 기반을 제공한다. 그러나 본 연구는 특정 업데이트 구조가 모든 모델에 일반적으로 적용되는지, 또는 특정 조건 하에서만 나타나는지에 대한 일반화 가능성은 명시되지 않았다. 또한, 실제 산업적 적용 시 헤드-테일 분리가 얼마나 효율적으로 이루어지는지도 추가 연구가 필요하다.

실용적 활용

본 연구는 LLM의 후학습 과정에서 업데이트 구조를 분석하고 개입하는 데 활용될 수 있다. 예를 들어, 정렬 개선과 동시에 특정 능력을 유지하거나 복구하기 위해 헤드-테일 구조를 조절할 수 있다. 또한, 모델 개발 과정에서 업데이트의 스펙트럼 구조를 모니터링함으로써 학습 완전성과 커버리지 손실을 예측하고 최적화할 수 있다.