ChatPaper.aiChatPaper

超越熵:基于对比策略优化的正确性感知优势塑造

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization

July 16, 2026
作者: Weiwen Xu, Jia Liu, Hou Pong Chan, Long Li, Deng Cai, Min Chen, Hao Zhang
cs.AI

摘要

基於可驗證獎勵的強化學習(RLVR)通常使用熵來進行優勢塑造。然而,熵無法區分有用的不確定性與有害的混淆,這限制了其作為正確性訊號的有效性。我們提出對比策略優化(CPO),該方法利用參考引導與一般生成分佈之間的詞元級對比分歧,來實現感知正確性的優勢塑造。理論與實驗結果均顯示,此分歧能可靠地指示詞元級的正確性。我們進一步證明,在策略蒸餾是CPO的一個特例,其中後驗分佈由外部教師模型實例化。CPO也解決了零優勢問題。在領域內與領域外基準測試上的實驗表明,CPO在保持強大泛化能力的同時,顯著優於基於熵的RLVR方法。進一步分析顯示,正確與錯誤的回應分別自然支持探索與利用,而平衡兩者能達到最佳性能。
English
Reinforcement learning with verifiable rewards (RLVR) commonly uses entropy for advantage shaping. However, entropy cannot distinguish useful uncertainty from detrimental confusion, limiting its effectiveness as a correctness signal. We propose Contrastive Policy Optimization (CPO), which uses token-level contrastive disagreement between reference-guided and vanilla generation distributions for correctness-aware advantage shaping. Both theoretical and empirical results show that this disagreement reliably indicates token-level correctness. We further show that On-policy Distillation is a special case of CPO, where the posterior distribution is instantiated by an external teacher model. CPO also resolves the zero-advantage problem. Experiments on in-domain and out-of-domain benchmarks demonstrate that CPO substantially outperforms entropy-based RLVR methods while maintaining strong generalization. Further analysis shows that correct and incorrect responses naturally support exploration and exploitation respectively, and balancing both leads to the best performance.