엔트로피를 넘어: 대조 정책 최적화를 통한 정확성 인식 어드밴티지 형성
Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization
July 16, 2026
저자: Weiwen Xu, Jia Liu, Hou Pong Chan, Long Li, Deng Cai, Min Chen, Hao Zhang
cs.AI
초록
검증 가능한 보상이 있는 강화 학습(RLVR)은 일반적으로 어드밴티지 셰이핑에 엔트로피를 사용합니다. 그러나 엔트로피는 유용한 불확실성과 해로운 혼란을 구분하지 못하여 정확도 신호로서의 효과가 제한적입니다. 우리는 대조 정책 최적화(CPO)를 제안합니다. 이 방법은 참조 기반 생성 분포와 기본 생성 분포 간의 토큰 수준 대조적 불일치를 활용하여 정확도 인지 어드밴티지 셰이핑을 수행합니다. 이론적 및 경험적 결과 모두 이러한 불일치가 토큰 수준의 정확성을 신뢰성 있게 나타냄을 보여줍니다. 또한 온-정책 증류가 CPO의 특수한 경우임을 보여줍니다. 이 경우 사후 분포는 외부 교사 모델에 의해 구현됩니다. CPO는 또한 제로 어드밴티지 문제를 해결합니다. 인-도메인 및 아웃-오브-도메인 벤치마크 실험에서 CPO가 엔트로피 기반 RLVR 방법보다 훨씬 뛰어난 성능을 보이면서 강력한 일반화를 유지함을 입증합니다. 추가 분석에 따르면 올바른 응답과 잘못된 응답은 각각 자연스럽게 탐험과 활용을 지원하며, 이 둘의 균형을 맞추는 것이 최고의 성능을 가져옵니다.
English
Reinforcement learning with verifiable rewards (RLVR) commonly uses entropy for advantage shaping. However, entropy cannot distinguish useful uncertainty from detrimental confusion, limiting its effectiveness as a correctness signal. We propose Contrastive Policy Optimization (CPO), which uses token-level contrastive disagreement between reference-guided and vanilla generation distributions for correctness-aware advantage shaping. Both theoretical and empirical results show that this disagreement reliably indicates token-level correctness. We further show that On-policy Distillation is a special case of CPO, where the posterior distribution is instantiated by an external teacher model. CPO also resolves the zero-advantage problem. Experiments on in-domain and out-of-domain benchmarks demonstrate that CPO substantially outperforms entropy-based RLVR methods while maintaining strong generalization. Further analysis shows that correct and incorrect responses naturally support exploration and exploitation respectively, and balancing both leads to the best performance.