When Does Muon Help Agentic Reinforcement Learning?

Kai Ruan, Jinghao Lin, Zihe Huang, Ziqi Zhou, Qianshan Wei, Xuan Wang, Hao Sun

arXiv:2607.16169 · 2026-07-20 공개 · arXiv · PDF

reinforcement-learning alfworld muon learning-rate sparse-reward adamw gigpo qwen2-5-0-5b-instruct

Abstract

Muon is competitive with AdamW in large-scale pre-training, but its value for reinforcement-learning (RL) post-training remains unclear. We study vanilla Muon in sparse-reward agentic RL through matched single-seed comparisons with AdamW on ALFWorld using Qwen2.5-0.5B-Instruct. Under Group-in-Group Policy Optimization (GiGPO), applying Muon only to hidden weight matrices raises final-window validation success from 0.290 to 0.546 (+88%); high-rate AdamW controls retain no post-update success. The effect depends on the advantage estimator and learning rate. At 3e-5, Muon improves GRPO from 0.161 to 0.268, whereas GraphGPO's late-window gap narrows near saturation. At 1e-5, GraphGPO Muon reaches 0.901, raises normalized validation AUC from 0.399 to 0.556, and reaches 0.5 and 0.75 success 30 and 60 updates earlier, respectively. These exploratory results show that Muon can benefit agentic RL and motivate studying the policy optimizer, advantage estimator, and learning rate jointly. Multi-seed and cross-task validation remain open.

한국어 요약

한 줄 요약

Muon이 GiGPO 기반의 장기적 희소보상 에이전트 강화학습에서 AdamW 대비 88% 성능 향상을 보인다.

핵심 기여도

핵심 아이디어

Muon은 Adam 계열 최적화기의 요소별 적응 스케일링을 Newton–Schulz (NS) 반복을 통한 모멘텀 행렬의 스펙트럼 정규화로 대체한다. 이는 업데이트 방향의 약한 성분을 강화하는 효과가 있다. 기존 연구에서는 Muon이 단일턴 RLVR에서 실패하는 경우가 많았는데, 이는 샘플링 노이즈가 약한 방향을 지배했기 때문으로 추정된다. 본 연구는 Muon이 장기적 희소보상 에이전트 강화학습에서 성능을 향상시킬 수 있음을 보여주며, 이는 advantage estimator의 구조와 최적화기의 상호작용에 의존한다는 점을 강조한다.

기술적 접근법

주요 결과

의의 및 한계

실용적 활용