ChatPaper.aiChatPaper

エントロピーを超えて:対比的政策最適化による正しさ認識型アドバンテージ整形

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization

July 16, 2026
著者: Weiwen Xu, Jia Liu, Hou Pong Chan, Long Li, Deng Cai, Min Chen, Hao Zhang
cs.AI

要旨

検証可能な報酬を用いた強化学習(RLVR)では、一般的にアドバンテージの形成にエントロピーが使用される。しかし、エントロピーは有用な不確実性と有害な混乱を区別できないため、正しさのシグナルとしての有効性が制限される。本稿では、参照ガイド付き生成分布とバニラ生成分布の間のトークンレベルの対照的不一致を正しさを考慮したアドバンテージ形成に用いる、対照的方策最適化(CPO)を提案する。理論的および実証的結果の両方から、この不一致がトークンレベルの正しさを確実に示すことが示される。さらに、オン方策蒸留はCPOの特別なケースであり、その場合の事後分布は外部の教師モデルによって具体化されることを示す。CPOはまた、ゼロアドバンテージ問題を解決する。ドメイン内およびドメイン外のベンチマーク実験により、CPOがエントロピーベースのRLVR手法を大幅に上回る性能を示し、かつ強力な汎化性能を維持することが実証された。さらなる分析から、正しい応答と誤った応答はそれぞれ探索と活用を自然に促進し、両者のバランスが最良のパフォーマンスにつながることが明らかになった。
English
Reinforcement learning with verifiable rewards (RLVR) commonly uses entropy for advantage shaping. However, entropy cannot distinguish useful uncertainty from detrimental confusion, limiting its effectiveness as a correctness signal. We propose Contrastive Policy Optimization (CPO), which uses token-level contrastive disagreement between reference-guided and vanilla generation distributions for correctness-aware advantage shaping. Both theoretical and empirical results show that this disagreement reliably indicates token-level correctness. We further show that On-policy Distillation is a special case of CPO, where the posterior distribution is instantiated by an external teacher model. CPO also resolves the zero-advantage problem. Experiments on in-domain and out-of-domain benchmarks demonstrate that CPO substantially outperforms entropy-based RLVR methods while maintaining strong generalization. Further analysis shows that correct and incorrect responses naturally support exploration and exploitation respectively, and balancing both leads to the best performance.