Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning

Junyao Yang, Yucheng Shi, Zongxia Li, Zhongzhi Li, Ruhan Wang, Xiangxin Zhou, Kishan Panaganti, Haitao Mi, Leowei Liang

arXiv:2607.18722 · 2026-07-22 공개 · arXiv · PDF

trust-region asynchronous-rl sglang megatron aime24 adaptive-clipping staleness-heterogeneity staleness-adaptive

Abstract

Asynchronous reinforcement learning improves throughput by decoupling rollout generation from optimization, but staleness is an inevitable byproduct compounded by policy lag, engine delays, and mixture-of-experts routing. From a trust-region perspective, this mismatch is critical: training-inference divergence governs approximation error in finite-horizon bounds, whereas PPO clipping only gates sampled outward updates, acting as a sampled surrogate rather than a full-policy constraint. As a result, high-staleness updates remain weakly controlled in the asynchronous regime where stale rollouts matter most. We introduce the Staleness-Adaptive Trust Region (SAT), which uses the detached sampled log-ratio as a practical staleness proxy, identifies high-mismatch tails within each batch via staleness-based kernel scaling, and contracts only the sign-selected endpoint of the nominal PPO interval. This preserves baseline behavior on ordinary tokens while enforcing more conservative updates on newly intercepted outward bands. We prove local interval containment and pointwise pessimism relative to PPO, showing how the adaptive rule reshapes update geometry under heterogeneous staleness. We evaluate SAT in a decoupled asynchronous RL setup built on Qwen3-30B-A3B-Base, using SGLang as the inference engine and Megatron for training. In this setting, SAT-GSPO w/ R3 achieves the best observed AIME24 avg@8, reaching 35.83 at lag 1 and 34.79 at lag 8, while SAT-GSPO reaches 34.17 at lag 1. Adaptive clipping and routing replay act as complementary stabilizers targeting mismatch tails and routing inconsistency, respectively. Overall, aligning clip intervals with staleness heterogeneity effectively stabilizes asynchronous RL.

한국어 요약

한 줄 요약

비동기 강화학습의 안정성을 향상시키기 위해 staleness-적응형 신뢰영역(SAT)을 제안한다.

핵심 기여도

핵심 아이디어

기존 PPO 클리핑은 샘플 기반의 대체적 제약으로, 실제 정책 변화를 제어하지 못한다. 이에 따라, 학습-추론 간의 불일치가 증가할수록 근사 오차가 커지는 문제를 해결하기 위해, SAT는 staleness를 proxy로 활용한 적응형 신뢰영역을 제안한다. SAT는 샘플 로그비를 기반으로 staleness를 추정하고, 커널 스케일링을 통해 높은 mismatch 영역을 식별한다. 이후, PPO 간격의 끝점을 선택적으로 수축하여 보수적 업데이트를 적용한다.

기술적 접근법

주요 결과

의의 및 한계

SAT는 비동기 강화학습에서 staleness에 따른 불안정성을 효과적으로 제어하며, 실제 대규모 모델 학습 환경에서 실현 가능성을 보인다. 특히, adaptive clipping과 routing replay가 서로 다른 불일치 원인을 타겟팅함으로써 보완적 안정화를 제공한다. 그러나, staleness proxy의 정확도나 커널 스케일링의 최적화 조건은 추가 연구가 필요하다.

실용적 활용

대규모 언어 모델의 비동기 학습, 분산 추론-학습 환경, 실시간 강화학습 시스템 등에서 학습 안정성 향상에 활용 가능하다.