QwenGyre: An Elastic Reinforcement Learning Framework for Training xLong-Horizon Agents

Weiqi Wang, Yuxin Zhou, Mouxiang Chen, Siyuan Zhang, Yi Zhang, Yuyan Luo, Zhiyu Yin, Chencan Wu, Jiemin Jiang, Wentao Yao, Chujie Zheng, JianWei Zhang

arXiv:2609.33848 · 2026-09-29 공개 · arXiv · PDF

reinforcement-learning llm-agents large-language-models training-efficiency nl2repobench trajectory-redundancy qwen-gyre xlong-horizon

Abstract

Large language model (LLM) agents increasingly undertake extreme-long (xlong) horizon tasks, where a single execution can span hours, hundreds of model--environment interactions, and nearly 1M tokens per rollout. Applying online reinforcement learning (RL) to such executions poses two fundamental challenges: (1) severe execution variance and prolonged rollout delays cause massive GPU idling; and (2) complex non-linear branching generates massive trajectory redundancy, crippling training efficiency. To address these, we presents QwenGyre, an end-to-end framework for xlong-horizon online RL. QwenGyre elastically reallocates GPUs between rollout and training without interrupting live executions, while its trajectory processor reconstructs branching histories, scores partial progress, and deduplicates redundant paths to bound training costs. Scaled to our flagship model, Qwen~3.8 2.4T, with 700K tokens per rollout, QwenGyre yields a 6.0% absolute gain on NL2RepoBench (52.5% to 58.5%) in 48 steps. Across our evaluations on diverse domains of training datasets, QwenGyre delivers up to 1.85times and 1.78times speedups over Colocate and Async, respectively.

한국어 요약

한 줄 요약

QwenGyre는 xLong-horizon RL에서 GPU 재할당과 트래젝토리 처리를 통해 1.85× 가속을 달성한 엘라스틱 RL 프레임워크이다.

핵심 기여도

핵심 아이디어

기존 온라인 RL은 xLong-horizon 작업에서 GPU 활용률 저하와 트래젝토리 중복 문제를 겪는다. QwenGyre는 이를 해결하기 위해 **Elastic Scheduler**와 **Trajectory Processor**를 결합한 새로운 프레임워크를 제안한다. Elastic Scheduler는 GPU 수요에 따라 롤아웃과 훈련 간 자원을 동적으로 재할당하며, 훈련 중에도 롤아웃이 중단되지 않도록 보장한다. Trajectory Processor는 분기된 트래젝토리를 트리 구조로 재구성하고, 공통된 경로를 제거함으로써 훈련 비용을 제어한다. 이는 특히 1M 토큰 이상의 롤아웃에서 효과적이다.

기술적 접근법

주요 결과

의의 및 한계

QwenGyre는 xLong-horizon 작업에서 GPU 활용률을 극대화하고, 복잡한 트래젝토리 처리를 효율적으로 수행함으로써 온라인 RL의 확장성을 높인다. 특히, 1M 토큰 이상의 롤아웃에서 성능 향상이 두드러지며, 대규모 모델(Qwen 3.8 2.4T)에서도 안정적으로 작동함을 입증했다. 그러나 본 논문은 특정 도메인(NL2RepoBench 등)에 국한된 실험을 수행했으며, 다른 유형의 xLong-horizon 작업에서의 일반화 가능성은 명시되지 않았다.

실용적 활용

QwenGyre는 소프트웨어 개발 자동화, 대규모 코드베이스 리팩토링, 복잡한 시스템 재구현 등 장기 실행이 필요한 AI 에이전트 개발에 적용 가능하다. 특히, GPU 자원을 효율적으로 관리해야 하는 클라우드 기반 RL 훈련 환경에서 유용하다.