HuatuoGPT-3: RL-Only Domain Adaptation from Base Models

Junying Chen, Xinyuan Xie, Ziniu Li, Wenyuan Gu, Jianquan Li, Xiang Wan, Guangjun Yu, Ruoyu Sun, Haizhou Li, Benyou Wang

arXiv:2610.05966 · 2026-10-07 공개 · arXiv · PDF

large-language-models domain-adaptation medical-llm healthbench adaptive-objective-evolution gradient-starvation teacher-distribution-anchoring rl-only

Abstract

Domain adaptation aims to turn a general-purpose large language model (LLM) into an expert for a target domain. While the dominant SFT+RL pipeline offers a convenient cold start, it may reduce exploration diversity and introduces additional complexity through multi-stage optimization. These limitations motivate RL-only adaptation. However, pure on-policy RL suffers from a cold-start problem, while mixed-policy RL still falls short: informative tokens in teacher outputs are learned too slowly in early training, and stale teacher outputs can hinder later improvement. We identify these two failure modes as Gradient Starvation and Teacher-Distribution Anchoring. To address them, we propose One-stage Policy Optimization (OnePO), which treats teacher outputs as transient guidance for policy improvement. OnePO combines Adaptive Objective Evolution to strengthen learning on informative low-probability teacher tokens and Teacher Retirement to discard teacher outputs once the current policy can surpass them. On medical adaptation, OnePO achieves 67.2 on HealthBench (Total) with only 20K training samples, outperforming SFT+RL and pure RL by 2.7 and 7.4 points, respectively. We further scale OnePO to produce HuatuoGPT-3, an open-source medical LLM series whose 27B variant reaches 70.1 on HealthBench (Total) and 71.4 on HealthBench Professional, surpassing frontier models such as GPT-6 Astra. Models and code are available at https://github.com/FreedomIntelligence/HuatuoGPT-3.

한국어 요약

한 줄 요약

HuatuoGPT-3는 RL-Only 기반 도메인 적응 프레임워크인 OnePO를 통해 20K 샘플로 67.2의 HealthBench 점수를 달성한 의료 전문 LLM 시리즈다.

핵심 기여도

핵심 아이디어

기존 SFT+RL 파이프라인은 다단계 최적화로 인해 탐색 다양성을 줄이고 복잡도를 증가시킨다. 이에 반해, OnePO는 1단계 RL-Only 접근법을 통해 도메인 적응을 수행한다. 핵심 아이디어는 *teacher outputs*를 일시적인 가이드로 활용하는 데 있다. 초기에는 *Adaptive Objective Evolution*을 통해 낮은 확률의 유용한 토큰에 대한 학습 신호를 강화하고, 이후 *Teacher Retirement*를 통해 더 이상 필요 없는 teacher outputs를 제거함으로써 *Gradient Starvation*와 *Teacher-Distribution Anchoring* 문제를 해결한다. 이는 기존 mixed-policy RL에서 발생하는 학습 지연 및 정체 문제를 극복하는 핵심 통찰이다.

기술적 접근법

주요 결과

의의 및 한계

OnePO는 기존 SFT+RL 파이프라인의 복잡성과 학습 효율 저하 문제를 해결하며, 1단계 RL-Only 접근법의 실용성을 입증한다. 특히, *Adaptive Objective Evolution*과 *Teacher Retirement*는 teacher output의 유용성과 제약을 동적으로 관리함으로써 도메인 적응의 정확도와 효율성을 동시에 향상시킨다. 그러나 100B 이상의 대형 모델에서의 OnePO 검증은 이루어지지 않았으며, 다중 teacher output을 사용하는 경우의 동작 변화도 추가 연구가 필요하다.

실용적 활용

의료 분야에서 20K 샘플로도 높은 성능을 내는 OnePO는 데이터가 제한된 전문 분야 모델 개발에 유용하게 활용될 수 있다. 또한, SFT 단계를 생략함으로써 학습 비용을 절감할 수 있어, 산업 현장에서의 빠른 모델 배포와 유지보수에 적합하다.