Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization

Weiwen Xu, Jia Liu, Hou Pong Chan, Long Li, Deng Cai, Min Chen, Hao Zhang

arXiv:2607.14614 · 2026-07-20 공개 · arXiv · PDF

reinforcement-learning on-policy-distillation rlvr generalization token-level exploration-exploitation contrastive-policy-optimization correctness-aware

Abstract

Reinforcement learning with verifiable rewards (RLVR) commonly uses entropy for advantage shaping. However, entropy cannot distinguish useful uncertainty from detrimental confusion, limiting its effectiveness as a correctness signal. We propose Contrastive Policy Optimization (CPO), which uses token-level contrastive disagreement between reference-guided and vanilla generation distributions for correctness-aware advantage shaping. Both theoretical and empirical results show that this disagreement reliably indicates token-level correctness. We further show that On-policy Distillation is a special case of CPO, where the posterior distribution is instantiated by an external teacher model. CPO also resolves the zero-advantage problem. Experiments on in-domain and out-of-domain benchmarks demonstrate that CPO substantially outperforms entropy-based RLVR methods while maintaining strong generalization. Further analysis shows that correct and incorrect responses naturally support exploration and exploitation respectively, and balancing both leads to the best performance.

한국어 요약

한 줄 요약

CPO는 참조 기반 생성과 일반 생성 간의 토큰 수준 대조 차이를 활용해 RLVR에서 정확성 인식형 advantage shaping을 구현한다.

핵심 기여도

핵심 아이디어

기존 RLVR에서 사용되는 엔트로피는 유용한 불확실성과 해로운 혼란을 구분하지 못해 정확성 신호로서 한계가 있다. CPO는 참조 기반 생성 분포와 일반 생성 분포 간의 토큰 수준 대조 차이를 활용해 정확성에 민감한 advantage shaping을 구현한다. 이는 학습 도중 각 토큰의 정확성 여부를 명확히 반영할 수 있도록 해준다. 이 접근법은 학습 초기에 발생하는 zero-advantage 문제를 해결하며, 정확한 토큰은 탐색을, 오류 토큰은 활용을 촉진해 전체 성능 향상에 기여한다.

기술적 접근법

주요 결과

의의 및 한계

CPO는 기존 엔트로피 기반 RLVR의 한계를 극복하고, 정확성에 기반한 토큰 수준 학습을 가능하게 하여 수학 추론 및 out-of-domain 일반화를 동시에 향상시킨다. 이론적으로는 OPD를 포함한 다양한 방법을 통합하는 correctness-driven framework를 제시한다. 그러나, CPO는 참조 기반 생성에 의존하므로 참조가 부재하거나 불확실한 상황에서는 효과가 제한될 수 있다. 또한, advantage clipping과 같은 하이퍼파라미터 조정이 성능에 큰 영향을 미치므로, 이를 최적화하는 것이 중요하다.

실용적 활용

CPO는 수학 문제 해결, 프로그래밍, 논리 추론 등 정확성 기반 학습이 필요한 분야에서 활용 가능하다. 특히, 대규모 학습 데이터에서 효과적인 탐색과 활용을 동시에 수행할 수 있어, 교육 AI, 코드 생성, 복잡한 추론 시스템 등에 적용할 수 있다.