超越熵:基于对比策略优化的正确性感知优势塑造
Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization
July 16, 2026
作者: Weiwen Xu, Jia Liu, Hou Pong Chan, Long Li, Deng Cai, Min Chen, Hao Zhang
cs.AI
摘要
基于可验证奖励的强化学习(RLVR)通常采用熵进行优势塑形。然而,熵无法区分有益的不确定性与有害的混乱,这限制了其作为正确性信号的有效性。我们提出对比策略优化(CPO),该方法通过参考引导生成分布与原始生成分布之间的词级对比差异,实现正确性感知的优势塑形。理论和实证结果均表明,这种差异能可靠地指示词级正确性。我们进一步证明在线策略蒸馏是CPO的一种特例,其通过外部教师模型实例化后验分布。CPO还解决了零优势问题。在领域内和领域外基准上的实验表明,CPO显著优于基于熵的RLVR方法,同时保持强大的泛化能力。进一步分析显示,正确与错误响应分别自然支持探索与利用,平衡两者可获得最佳性能。
English
Reinforcement learning with verifiable rewards (RLVR) commonly uses entropy for advantage shaping. However, entropy cannot distinguish useful uncertainty from detrimental confusion, limiting its effectiveness as a correctness signal. We propose Contrastive Policy Optimization (CPO), which uses token-level contrastive disagreement between reference-guided and vanilla generation distributions for correctness-aware advantage shaping. Both theoretical and empirical results show that this disagreement reliably indicates token-level correctness. We further show that On-policy Distillation is a special case of CPO, where the posterior distribution is instantiated by an external teacher model. CPO also resolves the zero-advantage problem. Experiments on in-domain and out-of-domain benchmarks demonstrate that CPO substantially outperforms entropy-based RLVR methods while maintaining strong generalization. Further analysis shows that correct and incorrect responses naturally support exploration and exploitation respectively, and balancing both leads to the best performance.