TRACE: Rollout-Guided Quantization-Aware Training for FP4 Reinforcement Learning of MoE Language Models

Xin Wang, Hao Yu, Zhengyang Zhuge, Bochao Mao, Zheng Li, Junda Feng, Yuyan Luo, Yi Zhang, Yizhong Cao, Mi Zhang, Dayiheng Liu, Jianwei Zhang

arXiv:2610.07767 · 2026-10-07 공개 · arXiv · PDF

reinforcement-learning large-language-models kv-cache model-compression quantization-aware-training quantization-discrepancy fp4-quantization moe-language-models

Abstract

Reinforcement learning (RL) for post-training large language models (LLMs) incurs substantial computation and memory overhead during rollout generation, which motivates low-precision rollout for efficient RL training. However, existing FP4 RL methods suffer from a key limitation: they primarily optimize quantization accuracy on the training and rollout paths independently rather than directly reducing the discrepancy between the two quantized execution paths. In this work, we propose TRACE (Train-Rollout Quantization Alignment via Compact GuidancE), an FP4 quantization framework for RL training of Mixture-of-Experts (MoE) language models that addresses the limitation of existing FP4 RL methods. TRACE incorporates rollout-guided quantization-aware training that uses rollout-side quantization outcomes to guide training-side FP4 rounding decisions, directly reducing train-rollout discrepancy. Moreover, TRACE adopts an efficient quantization-information caching scheme that selectively retains mantissa and scale information from deeper layers to reduce the storage and communication overhead introduced by rollout guidance. We evaluate TRACE on four large-scale MoE language models across reasoning, coding, and long-horizon RL tasks. Our results demonstrate that TRACE enables joint FP4 weight/activation and FP4 KV-cache rollout with RL performance comparable to BF16 rollout, while achieving up to 5.4xrollout speedup and strong final FP4 performance compared with post-hoc FP4 quantization of BF16-trained policies.

한국어 요약

한 줄 요약

TRACE는 MoE 언어 모델의 FP4 강화학습에서 훈련-롤아웃 불일치를 줄이는 rollout-guided QAT와 효율적인 정보 캐싱을 결합한 양자화 프레임워크이다.

핵심 기여도

핵심 아이디어

기존 FP4 강화학습 방법은 훈련 및 롤아웃 경로에서 독립적으로 양자화 정확도를 최적화하지만, 두 경로 간의 불일치를 직접 줄이지 못한다. TRACE는 **rollout-side 양자화 결과를 훈련-side FP4 반올림 결정에 직접 활용**함으로써, 훈련과 롤아웃 간의 불일치를 줄이는 새로운 접근법을 제안한다. 이는 **FP4 QAT의 한계를 극복**하고, **정책 불일치와 학습 불안정성을 감소**시킨다.

또한, TRACE는 **mantissa와 scale 정보를 선택적으로 보존하는 캐싱 기법**을 도입하여, 롤아웃 가이드가 필요한 정보만 저장함으로써 **저장 및 통신 오버헤드를 최소화**한다. 이는 특히 MoE 모델의 **expert routing과 상호작용하는 수치 차이를 줄이는 데 효과적**이다.

기술적 접근법

주요 결과

의의 및 한계

TRACE는 MoE 언어 모델의 FP4 강화학습에서 **정책 불일치와 학습 불안정성을 효과적으로 완화**하며, **높은 효율성과 성능을 동시에 달성**한다. 특히, **FP4 QAT의 한계를 극복**하고, **정책이 FP4 롤아웃에 점진적으로 적응**하는 현상을 관찰할 수 있다.

하지만, TRACE는 **정책의 오래된 버전으로 인한 활성화 차이를 완전히 제거하지 못**하며, 이는 **정책의 오래됨(staleness)에 따라 성능이 변동될 수 있음**을 의미한다. 또한, **모든 레이어에서 mantissa 정보를 보존하지 않고 선택적으로 보존**하기 때문에, 일부 레이어의 정보 누락이 성능에 영향을 줄 수 있다.

실용적 활용

TRACE는 대규모 MoE 언어 모델의 **효율적인 강화학습 훈련**에 적합하며, 특히 **FP4 양자화를 활용한 저비용, 고속 롤아웃이 필요한 산업 및 연구 환경**에서 활용 가능하다. 예를 들어, **코드 생성, 추론, 장기적 강화학습 작업**에서 TRACE를 적용하면 **성능 저하 없이 훈련 효율성을 크게 향상**시킬 수 있다.