ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Yifei Zhou, Andrea Zanette, Jiayi Pan, Sergey Levine, Aviral Kumar

arXiv:2402.19446 · 2026-07-27 공개 · arXiv · PDF

reinforcement-learning large-language-models policy-optimization actor-critic language-model-agents hierarchical-rl multi-turn-rl agent-tasks

Abstract

A broad use case of large language models (LLMs) is in goal-directed decision-making tasks (or"agent"tasks), where an LLM needs to not just generate completions for a given prompt, but rather make intelligent decisions over a multi-turn interaction to accomplish a task (e.g., when interacting with the web, using tools, or providing customer support). Reinforcement learning (RL) provides a general paradigm to address such agent tasks, but current RL methods for LLMs largely focus on optimizing single-turn rewards. By construction, most single-turn RL methods cannot endow LLMs with the ability to intelligently seek information over multiple turns, perform credit assignment, or reason about their past actions -- all of which are critical in agent tasks. This raises the question: how can we design effective and efficient multi-turn RL algorithms for LLMs? In this paper, we develop a framework for building multi-turn RL algorithms for fine-tuning LLMs, that preserves the flexibility of existing single-turn RL methods for LLMs (e.g., proximal policy optimization), while accommodating multiple turns, long horizons, and delayed rewards effectively. To do this, our framework adopts a hierarchical RL approach and runs two RL algorithms in parallel: a high-level off-policy value-based RL algorithm to aggregate reward over utterances, and a low-level RL algorithm that utilizes this high-level value function to train a token policy within each utterance or turn. Our hierarchical framework, Actor-Critic Framework with a Hierarchical Structure (ArCHer), can also give rise to other RL methods. Empirically, we find that ArCHer significantly improves efficiency and performance on agent tasks, attaining a sample efficiency of about 100x over existing methods, while also improving with larger model capacity (upto the 7 billion scale that we tested on).

한국어 요약

한 줄 요약

ArCHer는 대규모 언어 모델의 다중 턴 에이전트 작업 학습을 위해 계층적 강화 학습 프레임워크를 제안하며, 기존 방법 대비 100배 높은 샘플 효율성을 달성한다.

핵심 기여도

핵심 아이디어

기존 단일 턴 RL 알고리즘은 다중 턴 상호작용에서 정보 수집, 신용 할당, 과거 행동 추론을 효과적으로 학습하지 못한다. 이를 해결하기 위해 ArCHer는 계층적 RL 구조를 도입하여, 고수준에서 가치 기반 RL을 사용해 다중 턴의 보상을 집계하고, 저수준에서는 토큰 단위 정책을 최적화한다. 고수준 크리틱은 RoBERTa나 GPT2와 같은 모델로 구현되며, 저수준 정책은 GPT-2 기반의 액터로 구현된다. 이는 단일 턴 RL의 유연성을 유지하면서도 다중 턴, 장기적 보상, 지연 보상 환경을 효과적으로 처리할 수 있도록 한다.

기술적 접근법

주요 결과

의의 및 한계

ArCHer는 다중 턴 에이전트 작업에서 샘플 효율성과 성능을 동시에 향상시키며, 기존 단일 턴 RL 방법을 확장할 수 있는 유연한 프레임워크를 제공한다. 특히, GPT-2와 Mistral 7B 모두에서 성능 향상이 관찰되어, 다양한 모델 아키텍처에 적용 가능하다는 점에서 실용적 가치가 높다. 그러나 실험은 GPT-2와 7B 규모의 Mistral 모델에 제한되었으며, 더 큰 모델에서의 성능 검증이 필요하다. 또한, 인간과의 실시간 상호작용 환경에서는 수천 번의 상호작용이 필요해, 모델 기반 RL 접근법이 유망하다는 점이 언급된다.

실용적 활용

ArCHer는 웹 탐색, 도구 사용, 고객 지원 등과 같은 다중 턴 에이전트 작업에 적용 가능하다. 특히, 샘플 효율성이 높아 대규모 학습 데이터가 부족한 상황에서도 유용하며, 다양한 RL 알고리즘과 모델 아키텍처를 결합할 수 있어 연구 및 산업 현장에서의 활용성이 높다.