ReSPO: Reshaped Sequence Policy Optimization for Gradient Starvation in Off-Policy Learning

Yihang Chen, Yuanhao Ban, Cho-Jui Hsieh

arXiv:2609.35433 · 2026-10-11 공개 · arXiv · PDF

reinforcement-learning qwen3 moe-models off-policy-learning variational-objective policy-drift gradient-starvation alpha-divergence

Abstract

Reinforcement learning from verifiable rewards (RLVR) frequently reuses rollouts across multiple policy updates, increasing the mismatch between the current policy and the data-generating policy. We identify a sign-dependent gradient starvation problem in clipped policy optimization: clipping suppresses under-generated positive responses at the low-importance-weight tail while permitting severely over-generated negative responses to dominate the high-weight tail. To address this, we propose ReSPO (Reshaped Sequence Policy Optimization), which replaces clipping with a smooth, two-branch sequence-level kernel derived from an α-divergence variational objective and an exponential variance-control tilt. The positive branch preserves a nonzero gradient weight for under-generated positive responses, while the negative branch suppresses heavily over-generated negative responses. We demonstrate that ReSPO effectively learns from long positive reasoning trajectories during early training, even when accumulated policy drift relegates them to the low-importance-weight tail. On dense and MoE Qwen3 models, ReSPO accelerates early optimization, improves final training scores, and achieves higher held-out benchmark performance under a rollout reuse, validating our approach on importance-weight tail control in off-policy learning.

한국어 요약

한 줄 요약

ReSPO는 off-policy 학습에서 발생하는 gradient starvation 문제를 해결하기 위해 α-divergence 기반의 sequence-level kernel을 도입한 새로운 정책 최적화 방법이다.

핵심 기여도

핵심 아이디어

기존의 clipped policy optimization(예: GRPO, GSPO)는 importance ratio $ w_t $ 또는 $ s = W^{1/|o|} $를 clipping하여 학습 안정성을 확보하지만, 이는 positive response의 low-importance-weight tail에서 gradient starvation을 유발한다. 반면, negative response의 high-importance-weight tail에서는 over-generated token의 무한 가중치로 학습 신호가 왜곡된다. ReSPO는 이러한 문제를 해결하기 위해 sequence-level importance weight $ W $를 α-divergence 기반의 variational objective로 재구성하고, exponential tilting을 통해 variance를 제어한다. 이로 인해, positive branch는 $ \alpha^{(+)} > 1 $로 $ W \to 0 $에서도 zero가 아닌 gradient를 유지하고, negative branch는 $ \alpha^{(-)} = 1 $로 양쪽 extreme tail을 억제한다. 이는 기존 clipping 기반 방법이 무시하는 low-와 high-tail 신호를 모두 효과적으로 활용할 수 있게 한다.

기술적 접근법

주요 결과

의의 및 한계

ReSPO는 off-policy 학습에서 발생하는 gradient starvation 문제를 체계적으로 해결함으로써, rollout reuse가 증가하는 상황에서도 안정적이고 효과적인 정책 최적화를 가능하게 한다. 특히 긴 긍정적 추론 경로를 early stage에서 학습할 수 있어, LLM의 추론 능력 향상에 기여한다. 또한, 모델 규모가 커질수록 ReSPO의 성능 우위가 두드러지며, 이는 MoE 아키텍처에서도 동일하게 적용 가능하다는 점에서 실용적 가치가 크다. 그러나, ReSPO는 hyperparameter search와 compute budget에 의존적이며, 이에 대한 제한이 명시되어 있다. 또한, Routing Replay와의 조합은 late-stage 성능 향상에 기여하지만, 이는 ReSPO 자체의 핵심 기여도가 아닌 보조적 효과이다.

실용적 활용

ReSPO는 off-policy 학습이 필수적인 LLM fine-tuning, 특히 RL from verifiable rewards(RLVR) 기반의 추론 모델 개선에 적용 가능하다. rollout reuse가 빈번한 대규모 모델 학습 환경에서, ReSPO는 policy drift에 따른 학습 신호 손실을 최소화하고, early-stage에서 긍정적 추론 경로를 효과적으로 학습함으로써 모델 성능을 향상시킬 수 있다. 이는 수학 문제 해결, 논리적 추론, 복잡한 대화 시스템 등에 활용 가능하다.