ReFT: Representation Finetuning for Language Models

Zhengxuan Wu, Aryaman Arora, Zheng Wang, Atticus Geiger, Daniel Jurafsky, Christopher D. Manning, Christopher Potts

arXiv:2404.03592 · 2026-07-27 공개 · arXiv · PDF

language-models instruction-tuning parameter-efficient commonsense-reasoning glue arithmetic-reasoning representation-finetuning low-rank-linear

Abstract

Parameter-efficient finetuning (PEFT) methods seek to adapt large neural models via updates to a small number of weights. However, much prior interpretability work has shown that representations encode rich semantic information, suggesting that editing representations might be a more powerful alternative. We pursue this hypothesis by developing a family of Representation Finetuning (ReFT) methods. ReFT methods operate on a frozen base model and learn task-specific interventions on hidden representations. We define a strong instance of the ReFT family, Low-rank Linear Subspace ReFT (LoReFT), and we identify an ablation of this method that trades some performance for increased efficiency. Both are drop-in replacements for existing PEFTs and learn interventions that are 15x--65x more parameter-efficient than LoRA. We showcase LoReFT on eight commonsense reasoning tasks, four arithmetic reasoning tasks, instruction-tuning, and GLUE. In all these evaluations, our ReFTs deliver the best balance of efficiency and performance, and almost always outperform state-of-the-art PEFTs. We release a generic ReFT training library publicly at https://github.com/stanfordnlp/pyreft.

한국어 요약

한 줄 요약

ReFT는 기존 PEFT 대비 15~65배 더 파라미터 효율적인 언어 모델 미세조정 방법으로, LoReFT가 다양한 태스크에서 최고 성능을 보인다.

핵심 기여도

핵심 아이디어

기존 PEFT는 모델의 **가중치(weight)**를 조정하는 반면, ReFT는 **representation**을 조정함으로써 더 효과적인 미세조정을 추구한다. 이는 기존 해석성 연구에서 **representation이 의미 정보를 풍부하게 담고 있음**을 보여주는 결과를 기반으로 한다. ReFT는 기존 모델을 **frozen** 상태로 유지하고, **task-specific intervention**을 학습하여 추론 시 hidden representation을 조정한다. 이는 **LoReFT**에서 구현되며, **low-rank projection matrix**를 사용해 linear subspace 내에서 조정을 수행한다. 이는 기존 DAS 방법을 확장한 형태로, **LoReFT**는 **LoRA 대비 15×–65× 더 적은 파라미터**로 동작한다.

기술적 접근법

주요 결과

의의 및 한계

ReFT는 기존 PEFT 대비 **더 적은 파라미터로 높은 성능**을 보이며, **interpretability**와 **efficiency**를 동시에 달성하는 새로운 방향을 제시한다. 특히, **LoReFT**는 **low-rank subspace**를 활용한 representation 조정을 통해 **가장 효율적인 PEFT 대안**으로 부상하고 있다. 그러나, ReFT는 **representation 조정의 일반화 가능성**이나 **복잡한 다중 태스크 학습**에서의 한계가 명시되지 않았으며, **test set 기반 과적합** 우려도 제기된다.

실용적 활용

ReFT는 **대규모 언어 모델을 저비용으로 미세조정**해야 하는 **교육, 고객 지원, 자동화 시스템** 등에 적용 가능하다. 특히, **LoReFT**는 **GPU 메모리 제한이 있는 환경**에서 유리하며, **instruction-following** 및 **commonsense reasoning**이 필요한 **AI 어시스턴트** 개발에 유용하다.