reinforcement-learning large-language-models policy-optimization actor-critic language-model-agents hierarchical-rl multi-turn-rl agent-tasks
Abstract
A broad use case of large language models (LLMs) is in goal-directed decision-making tasks (or"agent"tasks), where an LLM needs to not just generate completions for a given prompt, but rather make intelligent decisions over a multi-turn interaction to accomplish a task (e.g., when interacting with the web, using tools, or providing customer support). Reinforcement learning (RL) provides a general paradigm to address such agent tasks, but current RL methods for LLMs largely focus on optimizing single-turn rewards. By construction, most single-turn RL methods cannot endow LLMs with the ability to intelligently seek information over multiple turns, perform credit assignment, or reason about their past actions -- all of which are critical in agent tasks. This raises the question: how can we design effective and efficient multi-turn RL algorithms for LLMs? In this paper, we develop a framework for building multi-turn RL algorithms for fine-tuning LLMs, that preserves the flexibility of existing single-turn RL methods for LLMs (e.g., proximal policy optimization), while accommodating multiple turns, long horizons, and delayed rewards effectively. To do this, our framework adopts a hierarchical RL approach and runs two RL algorithms in parallel: a high-level off-policy value-based RL algorithm to aggregate reward over utterances, and a low-level RL algorithm that utilizes this high-level value function to train a token policy within each utterance or turn. Our hierarchical framework, Actor-Critic Framework with a Hierarchical Structure (ArCHer), can also give rise to other RL methods. Empirically, we find that ArCHer significantly improves efficiency and performance on agent tasks, attaining a sample efficiency of about 100x over existing methods, while also improving with larger model capacity (upto the 7 billion scale that we tested on).
한국어 요약
한 줄 요약
ArCHer는 대규모 언어 모델의 다중 턴 에이전트 작업 학습을 위해 계층적 강화 학습 프레임워크를 제안하며, 기존 방법 대비 100배 높은 샘플 효율성을 달성한다.
핵심 기여도
- ArCHer는 고수준 가치 기반 RL과 저수준 정책 최적화를 결합한 계층적 RL 프레임워크를 제안한다.
- 기존 PPO와 같은 단일 턴 RL 방법 대비 100배 높은 샘플 효율성을 달성한다.
- 7B 파라미터 규모의 모델에서 성능 향상이 관찰된다.
- RoBERTa와 GPT2 모두 고수준 크리틱으로 사용 가능하며 성능 저하 없이 동작한다.
핵심 아이디어
기존 단일 턴 RL 알고리즘은 다중 턴 상호작용에서 정보 수집, 신용 할당, 과거 행동 추론을 효과적으로 학습하지 못한다. 이를 해결하기 위해 ArCHer는 계층적 RL 구조를 도입하여, 고수준에서 가치 기반 RL을 사용해 다중 턴의 보상을 집계하고, 저수준에서는 토큰 단위 정책을 최적화한다. 고수준 크리틱은 RoBERTa나 GPT2와 같은 모델로 구현되며, 저수준 정책은 GPT-2 기반의 액터로 구현된다. 이는 단일 턴 RL의 유연성을 유지하면서도 다중 턴, 장기적 보상, 지연 보상 환경을 효과적으로 처리할 수 있도록 한다.
기술적 접근법
- **ArCHer 프레임워크**: 고수준에서 off-policy TD 학습을 사용한 가치 함수 학습과, 저수준에서 on-policy policy gradient 알고리즘을 사용한 토큰 단위 정책 최적화를 병행한다.
- **모델 구조**: 고수준 크리틱은 RoBERTa-base 또는 GPT2 기반, 저수준 액터는 GPT-2 또는 Mistral 7B 기반.
- **Q-value 과대 추정 방지**: Double Q-learning 기법을 사용해 두 개의 Q-모델과 V-모델을 독립적으로 학습.
- **샘플 재사용**: 고수준 크리틱은 더 긴 시간 간격에서 학습하여 샘플 재사용이 가능하다.
- **하이퍼파라미터**: GPT-2 기반 액터, RoBERTa 기반 크리틱, MLP 헤드를 사용한 임베딩 처리.
주요 결과
- ArCHer는 기존 on-policy 방법(PPO) 대비 약 100배 높은 샘플 효율성을 달성.
- Mistral 7B 모델을 사용할 경우, GPT-2 기반 ArCHer 대비 훨씬 빠르게 학습.
- RoBERTa와 GPT2 기반의 고수준 크리틱 모두 동일한 성능을 보임.
- 48개 샘플만으로는 학습 불안정, 더 큰 버퍼 사용 시 안정성 향상.
의의 및 한계
ArCHer는 다중 턴 에이전트 작업에서 샘플 효율성과 성능을 동시에 향상시키며, 기존 단일 턴 RL 방법을 확장할 수 있는 유연한 프레임워크를 제공한다. 특히, GPT-2와 Mistral 7B 모두에서 성능 향상이 관찰되어, 다양한 모델 아키텍처에 적용 가능하다는 점에서 실용적 가치가 높다. 그러나 실험은 GPT-2와 7B 규모의 Mistral 모델에 제한되었으며, 더 큰 모델에서의 성능 검증이 필요하다. 또한, 인간과의 실시간 상호작용 환경에서는 수천 번의 상호작용이 필요해, 모델 기반 RL 접근법이 유망하다는 점이 언급된다.
실용적 활용
ArCHer는 웹 탐색, 도구 사용, 고객 지원 등과 같은 다중 턴 에이전트 작업에 적용 가능하다. 특히, 샘플 효율성이 높아 대규모 학습 데이터가 부족한 상황에서도 유용하며, 다양한 RL 알고리즘과 모델 아키텍처를 결합할 수 있어 연구 및 산업 현장에서의 활용성이 높다.